Back to Blog
Software

They Had Backups but Couldn't Restore: Rebuilding Recovery Around RTO, RPO, and Immutable Storage

Backup coverage is up, yet backups keep getting encrypted along with production. This guide covers tiering RTO and RPO as concrete numbers, and using immutable storage and recovery drills to verify that a restore actually works.

POLYGLOTSOFT Tech Team2026-09-287 min read3
Backup and RecoveryRTORPOImmutable BackupDisaster Recovery

Backup Coverage Is Rising, but Restores Still Fail

Large data centers with full backup and recovery stacks keep failing to actually restore after an incident. That pattern has shifted the conversation from "are we backing up?" to "can we recover?"

In surveys of Korean ransomware victims, the share of affected companies that held backups rose from 47% in the first half of 2023 to 78.6% in the second half of 2025. Yet among companies that did hold backups, 23.2% reported that the backup data was encrypted along with production. Having a backup and having a backup that survives the incident are two different things.

"Backups are running fine" only reports job status. It says nothing about whether the backup set actually opens, whether the application comes up on that restored data, or how many hours that takes.

Agreeing on RTO and RPO as Numbers

Recovery design starts with two numbers. RTO (Recovery Time Objective) is how many hours a system may stay down. RPO (Recovery Point Objective) is how many hours of data you can afford to lose.

Applying a single pair of numbers across every system fails in one of two directions. Put everything in the top tier and storage and redundancy costs become unsustainable. Loosen everything and the systems that actually stop revenue go unprotected.

A practical approach is a lightweight business impact analysis that sorts systems into three tiers.

  • Tier 1 (revenue or regulatory impact): for example, RTO 4 hours / RPO 15 minutes — ordering, payments, production execution
  • Tier 2 (delay is tolerable): RTO 24 hours / RPO 24 hours — internal groupware, reporting
  • Tier 3 (retention oriented): RTO 72 hours / RPO 1 week — archives, analytical copies
  • The tiers exist to spend money differently, not to justify skipping backups on the lower ones.

    Keeping Backups from Dying with Production

    Ransomware has targeted backup servers as a first objective for years. A backup server that sits on the production network and is managed with domain credentials gets encrypted alongside everything else.

    The backup security guidance published by KISA and similar bodies consistently emphasizes three points.

  • Off-site backups isolated from the service network: keep at least one copy physically or logically disconnected
  • Immutable backups: use storage where data cannot be deleted or altered during the retention window, even by an administrator — object lock on object storage is the common implementation
  • Regular restore testing: verify with restore results, not with backup success logs
  • One more control belongs alongside these: separate the backup admin accounts and console from the production identity system. If a single compromised production admin account can also wipe the backups, immutability settings have already been bypassed at a higher layer.

    The 3-2-1 rule (three copies, two media types, one off-site) still holds in the cloud, but it needs reinterpretation. Three snapshots in the same account and the same region are not three copies — they are closer to one copy that disappears in a single incident. You need copies that cross both the account boundary and the region boundary.

    What a Recovery Drill Exposes

    Most recovery runbooks stop at data restoration. Run an actual drill and the clock burns on everything the runbook never mentioned.

  • Commercial software license keys and who can reissue them
  • Expired SSL certificates and code signing certificates
  • External integration API keys, and source IPs registered on payment or identity provider allowlists
  • DNS change authority and TTLs, plus cron definitions in the batch scheduler
  • This is where "the data is back but the application won't start" comes from. A runbook therefore needs a dependency-ordered recovery sequence: identity and accounts → database → message queue and cache → application → batch jobs, with a note on which stage must complete before the next one means anything.

    The deliverable from a drill should not be pass/fail but a measured RTO. If a system with a 4-hour target actually took 9 hours, the question for the next quarter is where those extra 5 hours went. Only once these measurements accumulate do the target numbers rest on evidence rather than hope.

    The "We're on Cloud and SaaS, So We're Fine" Misconception

    The availability figures a cloud provider publishes describe how often the infrastructure is up, not a promise that your data will be recoverable. Under the shared responsibility model, the provider covers the infrastructure while you cover the data, configuration, and permissions you put on top of it. A dropped table, a bad migration, a compromised admin account — all of those land on your side of the line.

    SaaS is no different. Most SaaS platforms offer a recycle bin and retention policies, but those windows are typically designed in terms of tens of days, which does not cover a quiet deletion or alteration discovered months later. Data subject to statutory retention periods needs a separate backup held outside the SaaS platform.

    Your design should also separate two failure types. A regional outage is solved by a copy in another region. An account-level incident invalidates every copy inside that account. The first calls for geographic separation; the second calls for account and privilege separation plus immutable storage.

    Where to Start Without Blowing the Budget

    Trying to stand up a company-wide disaster recovery program in one pass usually means never starting at all. Scope it small.

  • Pick only three systems tied directly to revenue or regulatory obligations
  • Agree on their RTO and RPO as numbers, and write them down
  • Run one real restore outside business hours
  • Record the elapsed time and every point where you got stuck, exactly as observed
  • A single pass exposes the blanks in the runbook and the gap against your targets. The remaining systems can then be expanded on the basis of that result.

    At POLYGLOTSOFT, runbook authoring, immutable backup configuration, scheduled recovery drills, and measured RTO records are built into our SM (maintenance) and cloud operations work as ongoing operational tasks rather than a one-off consulting deliverable. If your backups are running but the restore has never been verified, start with your top three systems and we will review them with you, then lay out what your current setup needs first.

    Need Technical Consultation?

    Our expert consultants in smart factory, AI, and logistics automation will analyze your requirements.

    Request Free Consultation