DR / Backup · Cybersecurity & Cloud
Backup and disaster recovery: 3-2-1, ransomware, and the restore you never tested
How to build DR that survives ransomware — immutability, RTO/RPO honesty, and why a snapshot in the same account is not a strategy.
6 min read
Disaster recovery is the ability to restore technology services to a defined state within a defined time. Backups are raw material; tested restoration is the capability the business actually needs.
Ransomware changed the standard. Attackers target backup consoles, identity systems, storage snapshots, and admin credentials because they know recovery pressure creates payment pressure. A backup controlled by the same compromised admin plane may not be a backup when it matters.
The paired checklist helps verify backup and DR hygiene. This guide explains how to judge RTO, RPO, isolation, SaaS recovery, and restore evidence so the checklist does not create confidence that production cannot honor.
What disaster recovery actually is
Disaster recovery is a technology recovery discipline tied to business impact. It identifies critical systems, dependencies, recovery sequence, recovery time objective, recovery point objective, backup architecture, identity dependencies, and validation steps before returning service to users.
Backups support DR but are not the same thing. A nightly database dump, storage snapshot, SaaS recycle bin, immutable object copy, and warm standby all create different recovery options. The right mix depends on business tolerance for downtime and data loss, not on what the storage platform makes easy.
DR planning should include dependency order, not only system priority. A customer portal may depend on identity, DNS, certificates, secrets, databases, message queues, object storage, monitoring, and support tooling. Restoring the application without those dependencies is not recovery, so test plans need to prove the chain can be rebuilt in the right sequence. The same dependency map should identify manual workarounds, customer communication triggers, and finance or support processes that continue while systems are unavailable. Otherwise a technically successful restore can still leave the business unable to operate.
Decisions the checklist will not make for you
The checklist cannot choose honest RTO and RPO values. The business must decide how long each process can be down and how much data can be lost. Technology teams then need to show whether current backup frequency, replication, staffing, and dependencies can meet those numbers.
It also cannot decide how much isolation is enough. Separate credentials, immutable copies, offline storage, separate accounts, and restricted backup administration all reduce ransomware risk, but they add cost and operational complexity. Leadership must choose the resilience target.
The checklist will not decide whether SaaS-native recovery is sufficient. Many SaaS platforms offer retention, recycle bins, or point-in-time features that do not equal full business recovery. Teams must decide when exports or third-party backups are needed.
Where teams actually fail
The most common failure is never restoring. Reports show successful backup jobs, but no one has rebuilt a service, recovered a tenant, restored a database to a clean environment, or validated application integrity. Until restore is tested, the backup is an assumption.
Ransomware exposes identity coupling. Backup administrators use the same Active Directory, SSO, privileged accounts, and management network as production. When the attacker owns identity, they can delete, encrypt, or disable recovery paths before defenders understand the blast radius.
Teams also forget SaaS data and make impossible promises. They claim minute-level RPO while running nightly backups, or they assume a SaaS vendor can export usable data quickly during an outage. DR credibility depends on reconciling promises with actual mechanics.
Restore tests fail when nobody defines acceptance. A database may restore but lose recent transactions, an application may start with broken integrations, or a SaaS export may be readable only by engineers. Business owners should validate whether recovered data is complete enough to operate, not merely whether an IT job completed.
How to use the paired checklist
Start with business services, not backup tools. List critical processes, systems, dependencies, owners, RTO, RPO, recovery order, and manual workarounds. Then use the checklist to validate whether backup design and restore procedure support those targets.
Run restore tests that prove the important paths. Include identity recovery, clean-room restoration, database consistency, application dependencies, SaaS exports, privileged access, and user validation. Record timing and gaps rather than only pass or fail.
Use the checklist after major architecture changes, new SaaS adoption, identity redesign, and ransomware exercises. DR is perishable because dependencies change faster than annual continuity documents. Keep restore evidence with timestamps, owners, defects, and retest results so leadership can see whether resilience is improving or simply being asserted. Include the business validation result, because technical restoration is incomplete until users can perform the critical process under realistic pressure with current data. The best tests also record what manual workaround kept customers served while restoration was underway and who approved the temporary operating mode.
What teams get wrong
- A successful backup job proves we can recover.
- Only a tested restore with validation proves the backup is usable for the service and objective. The test should measure time, data loss, dependency recovery, and business acceptance.
- Snapshots in the production account are enough for ransomware.
- If attackers can reach the same admin plane, snapshots may be deleted, encrypted, or corrupted. Ransomware-ready backup design needs isolation from compromised production identity and management paths.
- SaaS vendors handle all recovery for us.
- SaaS resilience varies; customers may still need exports, configuration backups, retention settings, and recovery procedures. Vendor uptime does not automatically restore customer-deleted, encrypted, or corrupted data.
When the checklist is enough — and when it is not
- RTO or RPO promises do not match backup frequency, architecture, or staffing.
- Backup administration depends on the same identity environment exposed to ransomware.
- Critical SaaS data has no tested export or recovery path.
- A restore test fails, exceeds objectives, or cannot validate data integrity.
Related checklists
Continuity
ISO 22301:2019 Business Continuity Management Checklist
Guide: ISO 22301: BIA, MTPD, and why a DR runbook is not a BCMS
Incident Response
Security Incident Response Readiness Checklist
Guide: Incident response in 2026: severity, legal clocks, and a CSIRT that can page at 2 a.m.
Information Security
ISO 27001:2022 Implementation & Audit Readiness Checklist
Guide: ISO 27001:2022 audit readiness: what Stage 1 and Stage 2 actually test
Related field notes
The checklists and field notes provided on this website are for educational and informational purposes only. They do not constitute legal, financial, or professional advice. Completing a checklist does not guarantee compliance, certification, or immunity from audits. Always consult with a certified auditor or legal counsel for your specific organizational needs. Full disclaimer