Skip to content

Backups and restore

The backup role (opdns-cp backup run) takes the backups on a schedule, applies retention and runs restore tests. Its other subcommands do the same once, from a shell:

opdns-cp backup run schedule backups, retention and restore tests (long-running role)
opdns-cp backup now take a backup now (--store postgres|clickhouse|all)
opdns-cp backup verify restore a backup into a scratch database and check it (--store, --id latest|random|<id>, --keep)
opdns-cp backup restore restore into an empty target (--store postgres --target-database-url URL,
or --store clickhouse --target-database NAME) (--id latest|<id>)
opdns-cp backup list list complete backups (--store)
opdns-cp backup prune apply retention now (--store)
Store How Where (backup bucket)
Postgres (every table of the control plane database) one read-consistent snapshot, COPY of every table in foreign-key order into a gzip tar, encrypted with AES-256-GCM; a plaintext manifest beside it holds the schema version, row counts and a digest per table postgres/<id>/
ClickHouse (query logs and aggregates) the server’s native BACKUP DATABASE ... TO S3(...), with row counts before and after in opdns-manifest.json clickhouse/<id>/

The backup bucket (--backup-bucket, default opdns-backups) is separate from the opdns bucket.

Postgres archives are encrypted with OPDNS_BACKUP_KEY (or OPDNS_BACKUP_KEY_FILE): comma-separated <id>:<base64 of 32 bytes> entries, of which the first encrypts and all decrypt, so a key can be rotated by putting the new one first and keeping the old one until its archives have expired. Generate one with openssl rand -base64 32. It is required outside env=dev. ClickHouse backups are not encrypted by opdns: they rely on the bucket’s encryption.

Postgres ClickHouse
Interval --postgres-backup-interval, 1 h --clickhouse-backup-interval, 24 h
Keep every backup for --postgres-keep-all, 48 h --clickhouse-keep-all, 14 days
Keep the first of each day for --postgres-keep-daily, 30 days --clickhouse-keep-daily, off
Keep the first of each month for --postgres-keep-monthly, off --clickhouse-keep-monthly, 120 days
Always keep the newest --postgres-keep-min, 3 --clickhouse-keep-min, 3

A failed run is retried after --backup-retry-interval (5 minutes). The Postgres interval is the recovery point: there is no point-in-time recovery, a restore goes back to the last archive. WAL archiving is planned.

Every --verify-interval (30 days) the role picks a retained backup (--verify-pick random, or latest), restores it into a scratch database, compares row counts, content digests and schema version with the manifest, drops the scratch database and writes the result to verify/<store>.json. The time the restore took (restore_seconds) is the recovery-time estimate. opdns-cp backup verify runs one now; --keep leaves the scratch database for inspection.

Scratch Postgres databases are created on --scratch-database-url (default the main database server; the user needs CREATEDB). Pointing it at a separate server in production is planned.

When: a real loss of the control plane database or the log store, or to rehearse. Rehearse first when time allows: opdns-cp backup verify --store postgres --id <id>.

  1. Pick the backup: opdns-cp backup list --store postgres.

  2. Create an empty database, then restore into it and migrate:

    Terminal window
    opdns-cp backup restore --store postgres --id <id|latest> \
    --target-database-url postgres://.../opdns_restored
    opdns-cp migrate # with OPDNS_DATABASE_URL pointing at the target

    The restore rebuilds the schema from the migrations embedded in the binary up to the backup’s version, so use an opdns-cp of the same or a newer release. It loads every table in one transaction, checks counts and digests, restores the sequences, and refuses a target that is not empty.

  3. Point OPDNS_DATABASE_URL of every role at the restored database and restart them.

  4. Bring the profile stream and snapshot in line with the restored state:

    Terminal window
    opdns-cp admin profiles reconcile --fix --yes

    (Profile propagation; drift the command cannot repair, such as ahead, is listed for you to look at.)

Terminal window
opdns-cp backup restore --store clickhouse --id <id> --target-database opdns_restored

The target database must not exist. Then rename it, or point OPDNS_CLICKHOUSE_URL of cp-ingest and cp-api at it. Records ingested after the backup are lost unless the logs stream still holds them (it keeps 24 hours): cp-ingest resumes from its consumer position, so restart it after the change and watch the aggregates.

Record the backup id, the recovery time, the rows restored and what was lost; if user data was affected, it goes into the transparency report.

deploy/dev/prometheus/rules/backup.rules.yml:

Alert Severity Fires when
PostgresBackupStale page no successful Postgres backup for 2 hours
ClickHouseBackupStale warn no successful ClickHouse backup for 26 hours
RestoreTestFailing warn a scheduled restore test failed in the last day
RestoreTestStale warn no successful restore test for 32 days

cp-backup runs in the simulation with Postgres backups every 15 minutes and a restore test every week. mise run dev:restore-test backs up both stores now and restores them into scratch databases.