Backup & Disaster Recovery

A backup that has never been restored is a hope, not a backup — this page is organised around restore-test recency.
Backups verified · 24h
14 / 14
dumped, stored, and checksum-verified — per manifest
Restore-tested < 30d
12 / 14 2 due
the number an auditor actually asks for
Oldest untested restore
34d
osaka-robotics / prod — test plan drafted below
Measured RTO · fleet median
14m
target ≤ 4h · from real restore tests, not estimates

Fleet backup & restore posture

restore-test age drives the status — not backup age
EnvironmentLast backupCoversPITRLast restore testStatus
meridian-health / prod today 09:00 verified 3 DBs tfstate state secrets inv. 14d 6d ago · RTO 11m passed restore-tested
kestrel-bank / prod today 09:10 verified 3 DBs tfstate state secrets inv. 30d regulated profile 12d ago · RTO 16m passed restore-tested
osaka-robotics / prod today 08:52 verified 3 DBs tfstate state secrets inv. 14d 34d ago · RTO 13m due test overdue
verde-foods / prod today 09:04 verified 3 DBs tfstate state secrets inv. 14d 31d ago · RTO 12m due test overdue
northwind-logistics / prod today 09:02 verified 3 DBs tfstate state secrets inv. 14d 9d ago · RTO 14m passed restore-tested
cobalt-mining / prod on-site 02:00 verified
as of bundle · 3d ago
3 DBs state secrets inv. 7d local media 18d ago · RTO 21m passed
executed on-site · journal via bundle
restore-tested
…8 more environments all backed up < 24h · all restore-tested < 30d
Restore tests run against an isolated scratch namespace — never production. Verification is mechanical: row counts, schema fingerprints, and the product's own smoke hooks against the restored copy.

Recovery objectives — targets vs measured

measured values come from restore tests and DR drills, not estimates
ObjectiveTargetMeasured
RPO — data loss window≤ 15m4m (PITR replay lag, worst case 30d)meets
RTO — single environment≤ 4h11–21m (range across 14 restore tests)meets
RTO — full region failover≤ 8h2h 40m (drill 2026-06-30, simulated)meets
Backup verification lag≤ 24h≤ 4h (checksums + manifest, air-gap via bundle)meets
Objectives are policy per profile — the regulated profile pins PITR at 30d and restore-test cadence at 14d. A missed cadence surfaces here and on the dashboard; it never silently ages out.

Recent restore tests & drills

6d ago
meridian-health/prod · restore test passed · RTO 11m · row counts ✓ · schema fingerprints ✓ · smoke 6/6
9d ago
northwind-logistics/prod · restore test passed · RTO 14m
12d ago
kestrel-bank/prod · passed on 2nd attempt — first restore hit a stale KV secret reference; engine fix shipped in 0.1.90 (KI-2026-013), re-test green
18d ago
cobalt-mining/prod · on-site restore test · RTO 21m · journal arrived in signed bundle, verified here
2026-06-30
Full DR drill · simulated region loss (verde-foods) · failover 2h 40m · 3 runbook gaps found → RB-009 updated, 2 became engine fixes

Restore to production

the real thing — not a test

A production restore is a plan like any other: pick the environment and point-in-time, the engine renders exactly what will be overwritten, and the plan requires a typed confirmation plus an approval token. Backup-before-destructive applies to restores too — the current state is snapshotted before the restore begins.