Repository navigation
Disaster Recovery
Document Type: Technical Wiki
Audience: All Employees, Engineering Teams, IT Teams, and Business Stakeholders
Purpose: Provide a clear and practical understanding of Disaster Recovery and how an organization prepares to recover critical systems, applications, and data after a disruption.
Disaster Recovery (DR) is the process an organization uses to prepare for, respond to, and recover from events that disrupt IT systems, applications, infrastructure, or data.
Think of DR like having a spare tire for a car. You hope you never need it, but if something goes wrong, having a prepared replacement helps you get back on the road quickly.
Key Idea: Disaster Recovery is not just about keeping backups. It is about having a planned, documented, and tested way to restore critical business services after a disruption.
Organizations depend on technology for everyday business operations. If a critical system becomes unavailable, the impact can affect employees, customers, revenue, and business operations.
A disaster does not necessarily mean a large natural catastrophe. Any event that significantly disrupts an organization's critical IT services can require disaster recovery.
- Power Outages - Extended loss of electrical power can make systems unavailable.
- Hardware Failures - Servers, storage devices, or network equipment can fail.
- Cyberattacks - Ransomware, malware, or other attacks can make systems or data unavailable.
- Human Error - Accidental deletion or incorrect configuration can cause service or data loss.
- Software Failures - Application failures or faulty deployments can interrupt business services.
- Data Corruption - Data can become damaged or unusable.
- Natural Disasters - Floods, fires, earthquakes, storms, and similar events can affect facilities and infrastructure.
- Third-Party Failures - A critical cloud provider, vendor, or external service may become unavailable.
When creating a Disaster Recovery plan, two questions are especially important:
- How long can we afford to be without the service?
- How much data can we afford to lose?
These questions are answered using RTO and RPO.
RTO = How long can we afford to be down?
The Recovery Time Objective (RTO) is the target amount of time within which a system or service should be restored after a disruption.
Imagine an application has an RTO of 2 hours.
If the application stops working at 10:00 AM, the target is to restore it by approximately 12:00 PM.
10:00 AM 12:00 PM
│ │
▼ ▼
Incident occurs ─────────── 2 hours ────────► Recovery target
Think of RTO like a restaurant saying:
"If our kitchen stops working, we need to be serving customers again within 2 hours."
The 2 hours is the recovery time target.
Important: RTO is a target. Actual recovery time can vary depending on the incident, recovery method, available resources, and other factors.
RPO = How much data can we afford to lose?
The Recovery Point Objective (RPO) defines the maximum acceptable amount of data loss, measured in time.
Imagine a system has an RPO of 30 minutes.
If the system fails at 10:00 AM, the organization should have a recovery point from approximately 9:30 AM or later.
9:30 AM 10:00 AM
│ │
▼ ▼
Last acceptable recovery point ─────────────► System failure
│
│
└── Up to 30 minutes of data
may potentially be lost
Imagine you are writing a document and your computer suddenly stops working.
- If you save your work every 30 minutes, you may lose up to 30 minutes of work.
- If you save every 5 minutes, you may lose much less work.
This saving frequency relates to the idea behind RPO.
Key Idea: A smaller RPO means the organization needs to recover from a more recent copy of the data.
| Metric | Simple Question | Focus |
|---|---|---|
| RTO | How long can we afford to be down? | Downtime |
| RPO | How much data can we afford to lose? | Data loss |
RTO = How quickly do we need to recover?
RPO = How far back can we recover?
Different systems may require different recovery strategies.
The appropriate strategy depends on factors such as:
- Business importance
- RTO
- RPO
- Cost
- Technical requirements
- Recovery complexity
Three common approaches are Backup, Cold Site, and Hot Site.
A backup is a copy of data stored separately from the original data.
Backups are primarily used to protect information from accidental deletion, corruption, system failures, or other forms of data loss.
Original Data
│
▼
Backup
│
▼
Stored Separately
- Protects important data
- Provides a recovery source
- Usually costs less than maintaining a complete duplicate environment
- Can be used to restore lost or corrupted data
A backup alone does not automatically restore the entire application or infrastructure.
The organization still needs a recovery process for restoring systems and services.
Key Idea: A backup protects the data. Disaster Recovery provides the broader process for restoring the business service.
A Cold Site is a recovery location that has the necessary infrastructure or facilities but is not immediately ready to run the organization's complete production environment.
Additional setup and restoration work is required.
Production Failure
│
▼
Activate Cold Site
│
▼
Set Up Systems
│
▼
Restore Data
│
▼
Start Applications
│
▼
Resume Operations
- Lower ongoing cost
- Requires more preparation during an actual disaster
- Usually has a longer recovery time
- Suitable for systems that do not require immediate recovery
A Hot Site is a prepared recovery environment that is maintained in a ready-to-use state.
It is designed to allow critical systems and services to be restored or transferred with minimal delay.
Production Failure
│
▼
Activate Hot Site
│
▼
Restore / Redirect Services
│
▼
Verify Systems
│
▼
Resume Operations
- Faster recovery
- Higher operating cost
- Requires continuous maintenance
- Suitable for systems with strict recovery requirements
| Strategy | Main Purpose | Recovery Speed | Relative Cost |
|---|---|---|---|
| Backup | Protect and restore data | Varies | Lower |
| Cold Site | Provide a recovery location | Slower | Lower / Medium |
| Hot Site | Provide a ready recovery environment | Faster | Higher |
Note: The most appropriate strategy depends on the business requirements of the system. A critical system may require a faster recovery approach than a non-critical system.
A Disaster Recovery plan should be a structured and repeatable process.
The following five steps provide a simple starting point.
First, identify the systems, applications, infrastructure, and data that are important to the business.
Examples include:
- Customer-facing applications
- Databases
- Payment systems
- Internal business applications
- File storage
- Authentication services
- Network infrastructure
For each critical system, identify its important dependencies.
Ask:
"What systems do we absolutely need to keep the business running?"
For each critical system, define its recovery requirements.
At minimum, identify:
- RTO
- RPO
- Business priority
- System dependencies
- Recovery requirements
| System | Priority | RTO | RPO |
|---|---|---|---|
| Customer Application | Critical | 2 hours | 30 minutes |
| Internal Reporting | Medium | 8 hours | 4 hours |
| Archive Storage | Low | 24 hours | 24 hours |
Note: These values are examples only. Actual RTO and RPO values should be determined based on business and technical requirements.
Identify the data that must be protected and establish an appropriate backup strategy.
Define:
- What needs to be backed up
- How frequently backups should run
- How long backups should be retained
- Where backups should be stored
- Who can access the backups
- How backups will be restored
- How backup failures will be detected
A backup is only useful if it can be successfully restored.
Backups should therefore be tested regularly.
The DR plan should clearly explain what people need to do when a major disruption occurs.
The plan should define:
Identify who is responsible for:
- Declaring or escalating the incident
- Coordinating recovery
- Restoring infrastructure
- Restoring applications
- Restoring data
- Communicating with stakeholders
- Confirming service recovery
Document:
- What should happen when an incident occurs
- Who should be contacted
- Which systems should be recovered first
- Where recovery should take place
- How systems and data should be restored
- How recovery should be verified
- How normal operations should be resumed
The instructions should be simple enough to follow during a high-pressure situation.
A DR plan should be tested regularly.
Having a document does not prove that the organization can actually recover its systems.
Testing can reveal:
- Missing backups
- Unusable backups
- Incorrect recovery procedures
- Outdated contact information
- Unexpected recovery times
- Missing system dependencies
- Access or permission problems
- Unclear responsibilities
1. Select a critical system
│
▼
2. Simulate a failure
│
▼
3. Follow the recovery procedure
│
▼
4. Restore the system
│
▼
5. Verify that it works
│
▼
6. Record problems
│
▼
7. Improve the DR plan
Key Idea: DR testing should be an ongoing activity, not a one-time event.
Disaster Recovery is not something an organization creates once and then forgets.
A good DR program follows a continuous improvement cycle.
┌───────────────┐
│ Identify │
└───────┬───────┘
│
▼
┌───────────────┐
│ Plan │
└───────┬───────┘
│
▼
┌───────────────┐
│ Protect │
└───────┬───────┘
│
▼
┌───────────────┐
│ Test │
└───────┬───────┘
│
▼
┌───────────────┐
│ Improve │
└───────┬───────┘
│
└──────────► Repeat
The DR plan should be reviewed whenever there are significant changes to:
- Applications
- Infrastructure
- Business processes
- System dependencies
- Recovery requirements
- Organizational responsibilities
Backup and Disaster Recovery are related, but they are not the same thing.
| Backup | Disaster Recovery |
|---|---|
| Creates copies of data | Provides an overall recovery approach |
| Primarily protects data | Protects business service availability and recoverability |
| Can be part of a DR strategy | Includes backups, people, processes, infrastructure, and procedures |
| Focuses on restoring data | Focuses on restoring business operations |
Having a copy of a database is a backup.
Having:
- The database backup
- Recovery infrastructure
- Documented recovery procedures
- Responsible team members
- Communication procedures
- Tested restoration procedures
is part of Disaster Recovery.
A practical Disaster Recovery program should follow these principles.
Not every system requires the same level of protection.
Focus recovery efforts on the systems that are most important to business operations.
Recovery requirements should be measurable.
Everyone involved should understand:
- How quickly the system needs to be restored
- How much data loss is acceptable
Critical data should have appropriate backup and protection mechanisms.
Recovery procedures should not depend on one person's memory.
Clear documentation makes recovery easier when the situation is stressful.
A recovery plan should be tested to verify that it works as expected.
The DR plan should be updated when systems, applications, infrastructure, or business requirements change.
Problems discovered during DR testing or real incidents should be documented and used to improve the recovery process.
| Term | Meaning |
|---|---|
| DR | Disaster Recovery |
| DR Plan | Documented approach for recovering systems and services |
| Backup | A copy of data used for restoration |
| RTO | Maximum acceptable time to restore a service |
| RPO | Maximum acceptable amount of data loss measured in time |
| Cold Site | Recovery location that requires setup before use |
| Hot Site | Prepared recovery environment designed for rapid recovery |
| Recovery | Process of restoring systems, data, and services |
Disaster Recovery is about being prepared before a disruption happens.
An effective DR capability brings together:
- Critical Asset Identification
- Recovery Requirements
- RTO and RPO
- Data Protection and Backups
- Recovery Infrastructure
- Documented Procedures
- Testing and Validation
- Continuous Improvement
The goal is not simply to recover technology.
The goal of Disaster Recovery is to help the organization restore critical services and continue business operations after a major disruption.
The Disaster Recovery process can be summarized into five simple stages:
| Stage | What It Means |
|---|---|
| 1. Prepare | Identify critical systems, applications, and data that the business depends on. |
| 2. Protect | Create and protect backups and other recovery resources. |
| 3. Recover | Restore critical systems, applications, data, and services after a disruption. |
| 4. Test | Regularly test the recovery plan to make sure it works as expected. |
| 5. Improve | Identify problems found during testing or real incidents and update the DR plan. |
Prepare
↓
Identify Critical Systems and Data
↓
Protect
↓
Backups and Data Protection
↓
Recover
↓
Restore Systems and Services
↓
Test
↓
Verify the Recovery Plan
↓
Improve
↓
Update and Strengthen the DR Plan
↓
Repeat
Remember: The best time to prepare for a disaster is before it happens.