Backups and restore
The backup role (opdns-cp backup run) takes the backups on a
schedule, applies retention and runs restore tests. Its other subcommands
do the same once, from a shell:
opdns-cp backup run schedule backups, retention and restore tests (long-running role)opdns-cp backup now take a backup now (--store postgres|clickhouse|all)opdns-cp backup verify restore a backup into a scratch database and check it (--store, --id latest|random|<id>, --keep)opdns-cp backup restore restore into an empty target (--store postgres --target-database-url URL, or --store clickhouse --target-database NAME) (--id latest|<id>)opdns-cp backup list list complete backups (--store)opdns-cp backup prune apply retention now (--store)What is backed up
Section titled “What is backed up”| Store | How | Where (backup bucket) |
|---|---|---|
| Postgres (every table of the control plane database) | one read-consistent snapshot, COPY of every table in foreign-key order into a gzip tar, encrypted with AES-256-GCM; a plaintext manifest beside it holds the schema version, row counts and a digest per table |
postgres/<id>/ |
| ClickHouse (query logs and aggregates) | the server’s native BACKUP DATABASE ... TO S3(...), with row counts before and after in opdns-manifest.json |
clickhouse/<id>/ |
The backup bucket (--backup-bucket, default opdns-backups) is separate
from the opdns bucket.
Postgres archives are encrypted with OPDNS_BACKUP_KEY (or
OPDNS_BACKUP_KEY_FILE): comma-separated <id>:<base64 of 32 bytes>
entries, of which the first encrypts and all decrypt, so a key can be
rotated by putting the new one first and keeping the old one until its
archives have expired. Generate one with openssl rand -base64 32. It is
required outside env=dev. ClickHouse backups are not encrypted by opdns:
they rely on the bucket’s encryption.
Schedule and retention
Section titled “Schedule and retention”| Postgres | ClickHouse | |
|---|---|---|
| Interval | --postgres-backup-interval, 1 h |
--clickhouse-backup-interval, 24 h |
| Keep every backup for | --postgres-keep-all, 48 h |
--clickhouse-keep-all, 14 days |
| Keep the first of each day for | --postgres-keep-daily, 30 days |
--clickhouse-keep-daily, off |
| Keep the first of each month for | --postgres-keep-monthly, off |
--clickhouse-keep-monthly, 120 days |
| Always keep the newest | --postgres-keep-min, 3 |
--clickhouse-keep-min, 3 |
A failed run is retried after --backup-retry-interval (5 minutes). The
Postgres interval is the recovery point: there is no point-in-time
recovery, a restore goes back to the last archive. WAL archiving is planned.
Restore tests
Section titled “Restore tests”Every --verify-interval (30 days) the role picks a retained backup
(--verify-pick random, or latest), restores it into a scratch
database, compares row counts, content digests and schema version with the
manifest, drops the scratch database and writes the result to
verify/<store>.json. The time the restore took (restore_seconds) is
the recovery-time estimate. opdns-cp backup verify runs one now;
--keep leaves the scratch database for inspection.
Scratch Postgres databases are created on --scratch-database-url
(default the main database server; the user needs CREATEDB). Pointing
it at a separate server in production is planned.
Restore
Section titled “Restore”When: a real loss of the control plane database or the log store, or to
rehearse. Rehearse first when time allows:
opdns-cp backup verify --store postgres --id <id>.
Postgres
Section titled “Postgres”-
Pick the backup:
opdns-cp backup list --store postgres. -
Create an empty database, then restore into it and migrate:
Terminal window opdns-cp backup restore --store postgres --id <id|latest> \--target-database-url postgres://.../opdns_restoredopdns-cp migrate # with OPDNS_DATABASE_URL pointing at the targetThe restore rebuilds the schema from the migrations embedded in the binary up to the backup’s version, so use an
opdns-cpof the same or a newer release. It loads every table in one transaction, checks counts and digests, restores the sequences, and refuses a target that is not empty. -
Point
OPDNS_DATABASE_URLof every role at the restored database and restart them. -
Bring the profile stream and snapshot in line with the restored state:
Terminal window opdns-cp admin profiles reconcile --fix --yes(Profile propagation; drift the command cannot repair, such as
ahead, is listed for you to look at.)
ClickHouse
Section titled “ClickHouse”opdns-cp backup restore --store clickhouse --id <id> --target-database opdns_restoredThe target database must not exist. Then rename it, or point
OPDNS_CLICKHOUSE_URL of cp-ingest and cp-api at it. Records ingested
after the backup are lost unless the logs stream still holds them (it
keeps 24 hours): cp-ingest resumes from its consumer position, so restart
it after the change and watch the aggregates.
Afterwards
Section titled “Afterwards”Record the backup id, the recovery time, the rows restored and what was lost; if user data was affected, it goes into the transparency report.
Alerts
Section titled “Alerts”deploy/dev/prometheus/rules/backup.rules.yml:
| Alert | Severity | Fires when |
|---|---|---|
PostgresBackupStale |
page | no successful Postgres backup for 2 hours |
ClickHouseBackupStale |
warn | no successful ClickHouse backup for 26 hours |
RestoreTestFailing |
warn | a scheduled restore test failed in the last day |
RestoreTestStale |
warn | no successful restore test for 32 days |
Local simulation
Section titled “Local simulation”cp-backup runs in the simulation with Postgres backups every 15 minutes
and a restore test every week. mise run dev:restore-test backs up both
stores now and restores them into scratch databases.