Tested backups and restore drills, because an untested backup is not a backup
Tested backups and restore drills for small SaaS servers: 3-2-1, database dumps vs snapshots, off-site encrypted storage, alerting and a restore runbook.
By FutureGen Systems
Most small SaaS teams have backups. Far fewer have tested backups, and the difference only shows up on the worst day of the year. A cron job that has been writing empty files for four months looks exactly like a working one until you need it. This post covers a backup strategy that fits a small production setup, a handful of servers and one or two databases, and, more importantly, how to prove it works with scheduled restore drills.
Why tested backups matter more than backups
Backups fail in quiet, boring ways:
- The database password was rotated and the dump script now fails at the first line.
- The disk filled up, so the dump was truncated but the job still exited successfully.
- Backups were written to the same server they were protecting, and the server is what died.
- The encryption key lived only on the server that was lost.
- The dump restores, but only into a database version you no longer run, or it takes nine hours when you assumed one.
None of these show up in a dashboard that says "backup job: OK". They only show up when you try to restore. So the goal is not "we take backups". It is "we restored last month, it took this long, and here is the runbook".
Start with 3-2-1
The 3-2-1 rule is old and still a good baseline:
- 3 copies of your data: production plus two backups.
- 2 different storage types or systems, so one failure mode cannot take out both.
- 1 copy off-site, in a different provider or at least a different region and account.
For a small SaaS, that usually looks like: the live database, a local or same-provider copy for fast restores, and an encrypted copy in S3-compatible object storage at a different provider. Many teams add a further principle: at least one copy that production credentials cannot delete, using object lock or versioning with a separate account. That is what protects you from ransomware or an attacker with your server's keys.
Database dumps vs snapshots
Both are useful. They solve different problems.
| Logical database dumps | Disk or volume snapshots | |
|---|---|---|
| Examples | pg_dump, mysqldump, mongodump |
Provider snapshots, LVM or ZFS snapshots |
| Consistency | Consistent for that database by design | Can be crash-consistent only, unless the database is paused or the tool coordinates with it |
| Portability | Restores to any compatible server or provider | Usually tied to that provider |
| Granularity | Restore one database or even one table | Whole disk |
| Restore speed | Slower for large databases | Fast to roll back a whole machine |
| Point-in-time | No, unless combined with WAL or binlog archiving | No, only at snapshot time |
Our default for small production systems is logical dumps of every database at least nightly, shipped off-site, plus provider snapshots of servers for quick whole-machine recovery. When losing up to a day of data is not acceptable, add continuous archiving, such as PostgreSQL WAL archiving with a tool like pgBackRest or WAL-G, so you can restore to a specific minute.
Do not forget the things that are not in the database: uploaded files, environment and configuration files, TLS and deploy configuration, and anything in object storage that you would miss.
Off-site storage, encryption and retention
- Encrypt before upload. Tools such as restic and Borg encrypt client-side by default. If you write your own scripts, encrypt with a key that is not stored only on the server being backed up.
- Store the keys somewhere else. A password manager the team already uses is fine. Test that someone other than the person who set it up can find them.
- Use separate credentials for the backup bucket, with permission to write but ideally not to delete.
- Set a retention policy you can explain. A common shape is daily backups for 7 to 14 days, weekly for a couple of months, and monthly for a year, adjusted to your contracts and data protection obligations. Retention also means deleting old data you should no longer keep.
A minimal nightly Postgres job might look like this, with restic handling encryption, deduplication and retention:
#!/usr/bin/env bash
set -euo pipefail
pg_dump --format=custom --file=/var/backups/app.dump "$DATABASE_URL"
restic backup /var/backups/app.dump /srv/app/uploads --tag nightly
restic forget --tag nightly --keep-daily 14 --keep-weekly 8 --keep-monthly 12 --prune
curl -fsS "$HEARTBEAT_URL" > /dev/null
The last line matters as much as the rest. It only runs if everything above succeeded.
Alert on failed and missing backups
Alerting on failure is not enough, because the most dangerous backup is the one that silently stopped running. Use a dead man's switch: the job pings a heartbeat URL on success, and the monitoring service alerts you if no ping arrives within the expected window. Healthchecks.io and most uptime monitors support this pattern, and you can self-host the former.
Also alert on:
- Backup size dropping sharply compared with the previous run, which often means a truncated dump.
- The newest object in the off-site bucket being older than expected.
- Disk usage on the backup volume crossing a threshold.
Send alerts to a channel people actually read, and make sure each one is clear about what to do.
Schedule restore drills
This is the part almost everyone skips. Put it in the calendar.
- Monthly, automated: a job restores the latest dump into a throwaway database, runs a few sanity queries (row counts on key tables, the most recent record's timestamp), and reports the result and the restore time.
- Quarterly, by hand: someone restores the full system to a fresh server from off-site backups only, following the runbook, without the person who wrote it helping. Every point where they got stuck becomes a fix to the runbook.
- After any significant change: a database major version upgrade, a new storage provider, or a change in backup tooling.
Record two numbers each time: how much data you would have lost (your recovery point) and how long the restore took (your recovery time). These are the figures you can honestly give a customer who asks.
Write the restore runbook before you need it
During an incident nobody wants to read a wiki page full of background. The runbook should be a short, numbered list that someone stressed at 2 a.m. can follow:
- Where backups live and how to list them.
- Where the encryption keys and storage credentials are.
- Exact commands to restore the database and files to a new server.
- How to point the application at the restored database and verify it.
- Who to tell, and in what order.
- The date of the last successful drill and how long it took.
Keep a copy outside the systems it describes. A runbook stored only on the server that died is not much use.
Make it someone's job
Backups are part of the system administration work we do: Linux servers, hardening, monitoring and backups with rehearsed restores. They are also built into every cloud migration we run, with a tested restore before cutover. If you are not sure when your backups were last restored, get in touch and we can help you find out.