Profile propagation
The path of a change
Section titled “The path of a change”- The API (or
opdns-cp admin profile set) saves the profile and writes an outbox row in the same Postgres transaction, with the new version. - The publisher that holds the leader lock reads the outbox and publishes
the whole profile document on the JetStream stream
profiles(subjectprofiles.<id>, one message kept per profile, headerOpdns-Version). It marks a whole batch published once NATS has acknowledged it; message-id deduplication makes a repeat after a crash harmless. If the stream was deleted, the publisher recreates it. - Every PoP receives the message through its NATS leaf node and the edge
applies it, ignoring any version not newer than the one it holds. A PoP
that starts reads the newest snapshot (
snapshots/latest.jsonin object storage) and follows the stream from the snapshot’s sequence.
Postgres is the source of truth: every repair below republishes what Postgres holds and never invents a version.
Signed profiles
Section titled “Signed profiles”Since 2026-09-30 every profile the control plane hands out is signed with Ed25519, so a PoP or node that receives a forged or altered document through NATS, object storage or the API refuses it (threat model G-1).
| What | Signed by | Signature |
|---|---|---|
each profiles.<id> message |
the publisher | headers Opdns-Signature (base64) and Opdns-Key-Id over "opdns-profile-sig/1\n" + id + "\n" + version + "\n" + body |
| each snapshot | the publisher | a minisign sidecar, profiles-<seq>.json.sig, written before snapshots/latest.json moves to it |
a node’s profile pull (GET /v1/nodes/self/profile) |
the api | profile_signature and profile_key_id in the response, over the same string with the profile’s compact JSON |
The key id is the first 8 bytes of the public key’s SHA-256, in hex, as for list signing keys.
Keys. The publisher and api roles read OPDNS_PROFILE_SIGNING_KEY
(--profile-signing-key): comma-separated files of hex Ed25519 seeds, one
per line. The first key signs; every key’s public half is announced
to nodes, at enrolment (profile_pubkey in the enrol response) and in
every profile pull, so nodes follow a rotation by themselves. In
development a missing file is generated (with FILE.pub beside it) in the
keys volume. Without a key, outside development, both roles log a warning
and publish unsigned; so do not leave it unset in production.
Verification on the PoPs. Edges verify with OPDNS_PROFILE_PUBLIC_KEYS
(-profile-public-keys): comma-separated hex keys, or files of them, any
of which may verify. With keys set, an unsigned or badly signed message or
snapshot is refused and the edge keeps the last good state it has. Empty
means no check, with a warning at boot outside development. GET /versions on an edge’s admin port lists the key ids it trusts
(profile_key_ids), and replay (below, and fleetctl profiles replay)
copies the signature headers along with each message.
Verification on nodes. Enrolled nodes trust the keys the cloud
announces from their enrolment on (or from their first pull that
announces one, for nodes enrolled earlier), keep them in
profile_keys.json, and may pin them with cloud.profile_public_keys
(Profile signing keys).
| Metric | Counts |
|---|---|
opdns_edge_profile_signature_failures_total{source} |
messages (stream) and snapshots (snapshot) an edge refused |
opdns_node_profile_signature_failures_total |
profile pulls a node refused |
No alert rule watches these counters yet; any increase means a key mismatch after a botched rotation, or forged data in the path.
Rotating the profile key
Section titled “Rotating the profile key”- Generate the new key and add its public half to every edge’s
OPDNS_PROFILE_PUBLIC_KEYS, next to the old one; roll the edges. - Add the new seed to the signing file as the second line and restart the publisher and api: nodes now learn the new key from every pull, while the old key still signs.
- Move the new seed to the first line and restart them: it signs from now on.
opdns-cp admin profiles republish --all --yes, so that the message kept for every profile, and the next snapshot, carry the new signature.- Remove the old key from the signing file and from the edges.
A node keeps trusting a key the cloud stops announcing until its next pull confirms the new set.
Not covered yet
Section titled “Not covered yet”Changes to these are planned:
- The api role holds the signing key, because it signs node responses built from Postgres; a separate node key is the alternative.
- A node trusts the first key it sees unless
cloud.profile_public_keysis pinned, and recovering from a lost cloud key means deleting itsprofile_keys.jsonby hand. - The stream keeps one message per profile, so a refused forgery still
displaces the genuine last message until
republish;reconciledoes not check signatures;snapshots/latest.jsonitself is unsigned; and the edge’s local state file is trusted as it is.
Propagation canary
Section titled “Propagation canary”While leading, the publisher rewrites a reserved system profile, canary,
every --canary-interval (default 1 minute; 0 turns it off) through the
same outbox, so each write crosses the whole path. Every edge that applies
a new canary version live (not while replaying at boot) reports it on the
core NATS subject canary.applied.<pop>. The publisher subtracts its own
write time from the report’s arrival, on one clock, and exports:
| Metric | Is |
|---|---|
opdns_cp_profile_propagation_seconds{pop} |
histogram of commit-to-applied time per PoP |
opdns_cp_profile_canary_lag_seconds{pop,node} |
per PoP server: the lag of the newest canary, rising while the server has not applied it |
opdns_cp_profile_canary_written_timestamp_seconds |
when the canary was last written |
Every profiles message also carries Opdns-Published-At and
Opdns-Committed-At headers, which the edges use for their own
publish-to-apply histogram (opdns_edge_profile_propagation_seconds).
The target is under 2 seconds from commit to every PoP.
Alerts
Section titled “Alerts”deploy/dev/prometheus/rules/profiles.rules.yml (the local simulation’s
Prometheus; production loads the same file):
| Alert | Severity | Fires when |
|---|---|---|
ProfilePublisherNoLeader |
page | no publisher has held the leader lock for 2 minutes: nothing is published |
ProfileOutboxStuck |
page | the oldest unpublished outbox row is over 30 s old for 2 minutes |
ProfileCommitToPublishSlow |
warn | commit-to-publish p99 over 1 s for 10 minutes |
ProfilePropagationCanarySlow |
warn | a PoP’s slowest server over 2 s behind the canary for 5 minutes |
ProfilePropagationCanaryStalled |
page | a PoP has not applied a canary for 5 minutes |
ProfileCanaryNotWritten |
warn | the canary was last written over 5 minutes ago |
ProfilePropagationSlow and ProfileStreamBehind in fleet.rules.yml
watch the same path from the edges’ side.
opdns-cp admin profiles
Section titled “opdns-cp admin profiles”opdns-cp admin profiles statusopdns-cp admin profiles reconcileopdns-cp admin profiles reconcile --fix --yesopdns-cp admin profiles replay --from 1234 --dry-runopdns-cp admin profiles republish --id abc123,def456 --yesopdns-cp admin profiles snapshot| Subcommand | Does |
|---|---|
status |
the stream’s first and last sequence, message and profile counts and last message time; the outbox backlog and its oldest row; the latest snapshot’s sequence, size and age |
reconcile [--json] |
compares Postgres with the stream’s newest message per profile and with the latest snapshot, and lists every drift; exits 3 when there is any |
reconcile --fix [--wait 60s] |
republishes every repairable drift through the outbox, waits for the publisher to drain it, and checks again; exits 3 if drift remains |
replay --from SEQ [--dry-run] |
republishes, unchanged, the newest message of every profile with a message at or after SEQ (stream only, no database) |
republish --id ID[,ID] | --all [--wait 60s] |
re-enqueues the current documents from Postgres (deleted profiles as tombstones) and waits until the publisher drained them |
snapshot |
writes a snapshot now and prints its key, sequence and profile count |
The drift kinds reconcile reports:
| Kind | Means | --fix repairs it |
|---|---|---|
missing |
Postgres has the profile, the stream has no message for it | yes |
behind |
the stream holds an older version, and none is waiting in the outbox | yes |
stale |
deleted in Postgres, the stream still holds a live document | yes |
ahead |
the stream holds a newer version than Postgres | no: investigate |
unknown |
the stream holds a profile Postgres does not know | no: investigate |
snapshot_ahead |
the latest snapshot holds a newer version than Postgres | no: investigate |
A profile whose newer version is still in the outbox is counted as in
flight, not as drift. ahead and snapshot_ahead can follow a
Postgres restore to an older backup; see Backups.
The command takes the database, NATS (--nats-url) and object storage
(--s3-*; without it the snapshot checks are skipped, and snapshot
refuses to run) settings of the other roles. Republished rows go through
the outbox with kind republish, and edges that already hold a version
ignore it. Outside env=dev, republish and reconcile --fix need
--yes; both are logged with the actor cli:<user>@<host>.
fleetctl profiles status shows each PoP’s applied stream sequence (from
Prometheus) and how far it is behind; fleetctl profiles replay is the
same replay as above. When to use which is in the fleet runbook
docs/fleet/runbooks/profile-stream-replay.md.