Skip to content

Profile propagation

  1. The API (or opdns-cp admin profile set) saves the profile and writes an outbox row in the same Postgres transaction, with the new version.
  2. The publisher that holds the leader lock reads the outbox and publishes the whole profile document on the JetStream stream profiles (subject profiles.<id>, one message kept per profile, header Opdns-Version). It marks a whole batch published once NATS has acknowledged it; message-id deduplication makes a repeat after a crash harmless. If the stream was deleted, the publisher recreates it.
  3. Every PoP receives the message through its NATS leaf node and the edge applies it, ignoring any version not newer than the one it holds. A PoP that starts reads the newest snapshot (snapshots/latest.json in object storage) and follows the stream from the snapshot’s sequence.

Postgres is the source of truth: every repair below republishes what Postgres holds and never invents a version.

Since 2026-09-30 every profile the control plane hands out is signed with Ed25519, so a PoP or node that receives a forged or altered document through NATS, object storage or the API refuses it (threat model G-1).

What Signed by Signature
each profiles.<id> message the publisher headers Opdns-Signature (base64) and Opdns-Key-Id over "opdns-profile-sig/1\n" + id + "\n" + version + "\n" + body
each snapshot the publisher a minisign sidecar, profiles-<seq>.json.sig, written before snapshots/latest.json moves to it
a node’s profile pull (GET /v1/nodes/self/profile) the api profile_signature and profile_key_id in the response, over the same string with the profile’s compact JSON

The key id is the first 8 bytes of the public key’s SHA-256, in hex, as for list signing keys.

Keys. The publisher and api roles read OPDNS_PROFILE_SIGNING_KEY (--profile-signing-key): comma-separated files of hex Ed25519 seeds, one per line. The first key signs; every key’s public half is announced to nodes, at enrolment (profile_pubkey in the enrol response) and in every profile pull, so nodes follow a rotation by themselves. In development a missing file is generated (with FILE.pub beside it) in the keys volume. Without a key, outside development, both roles log a warning and publish unsigned; so do not leave it unset in production.

Verification on the PoPs. Edges verify with OPDNS_PROFILE_PUBLIC_KEYS (-profile-public-keys): comma-separated hex keys, or files of them, any of which may verify. With keys set, an unsigned or badly signed message or snapshot is refused and the edge keeps the last good state it has. Empty means no check, with a warning at boot outside development. GET /versions on an edge’s admin port lists the key ids it trusts (profile_key_ids), and replay (below, and fleetctl profiles replay) copies the signature headers along with each message.

Verification on nodes. Enrolled nodes trust the keys the cloud announces from their enrolment on (or from their first pull that announces one, for nodes enrolled earlier), keep them in profile_keys.json, and may pin them with cloud.profile_public_keys (Profile signing keys).

Metric Counts
opdns_edge_profile_signature_failures_total{source} messages (stream) and snapshots (snapshot) an edge refused
opdns_node_profile_signature_failures_total profile pulls a node refused

No alert rule watches these counters yet; any increase means a key mismatch after a botched rotation, or forged data in the path.

  1. Generate the new key and add its public half to every edge’s OPDNS_PROFILE_PUBLIC_KEYS, next to the old one; roll the edges.
  2. Add the new seed to the signing file as the second line and restart the publisher and api: nodes now learn the new key from every pull, while the old key still signs.
  3. Move the new seed to the first line and restart them: it signs from now on.
  4. opdns-cp admin profiles republish --all --yes, so that the message kept for every profile, and the next snapshot, carry the new signature.
  5. Remove the old key from the signing file and from the edges.

A node keeps trusting a key the cloud stops announcing until its next pull confirms the new set.

Changes to these are planned:

  • The api role holds the signing key, because it signs node responses built from Postgres; a separate node key is the alternative.
  • A node trusts the first key it sees unless cloud.profile_public_keys is pinned, and recovering from a lost cloud key means deleting its profile_keys.json by hand.
  • The stream keeps one message per profile, so a refused forgery still displaces the genuine last message until republish; reconcile does not check signatures; snapshots/latest.json itself is unsigned; and the edge’s local state file is trusted as it is.

While leading, the publisher rewrites a reserved system profile, canary, every --canary-interval (default 1 minute; 0 turns it off) through the same outbox, so each write crosses the whole path. Every edge that applies a new canary version live (not while replaying at boot) reports it on the core NATS subject canary.applied.<pop>. The publisher subtracts its own write time from the report’s arrival, on one clock, and exports:

Metric Is
opdns_cp_profile_propagation_seconds{pop} histogram of commit-to-applied time per PoP
opdns_cp_profile_canary_lag_seconds{pop,node} per PoP server: the lag of the newest canary, rising while the server has not applied it
opdns_cp_profile_canary_written_timestamp_seconds when the canary was last written

Every profiles message also carries Opdns-Published-At and Opdns-Committed-At headers, which the edges use for their own publish-to-apply histogram (opdns_edge_profile_propagation_seconds). The target is under 2 seconds from commit to every PoP.

deploy/dev/prometheus/rules/profiles.rules.yml (the local simulation’s Prometheus; production loads the same file):

Alert Severity Fires when
ProfilePublisherNoLeader page no publisher has held the leader lock for 2 minutes: nothing is published
ProfileOutboxStuck page the oldest unpublished outbox row is over 30 s old for 2 minutes
ProfileCommitToPublishSlow warn commit-to-publish p99 over 1 s for 10 minutes
ProfilePropagationCanarySlow warn a PoP’s slowest server over 2 s behind the canary for 5 minutes
ProfilePropagationCanaryStalled page a PoP has not applied a canary for 5 minutes
ProfileCanaryNotWritten warn the canary was last written over 5 minutes ago

ProfilePropagationSlow and ProfileStreamBehind in fleet.rules.yml watch the same path from the edges’ side.

Terminal window
opdns-cp admin profiles status
opdns-cp admin profiles reconcile
opdns-cp admin profiles reconcile --fix --yes
opdns-cp admin profiles replay --from 1234 --dry-run
opdns-cp admin profiles republish --id abc123,def456 --yes
opdns-cp admin profiles snapshot
Subcommand Does
status the stream’s first and last sequence, message and profile counts and last message time; the outbox backlog and its oldest row; the latest snapshot’s sequence, size and age
reconcile [--json] compares Postgres with the stream’s newest message per profile and with the latest snapshot, and lists every drift; exits 3 when there is any
reconcile --fix [--wait 60s] republishes every repairable drift through the outbox, waits for the publisher to drain it, and checks again; exits 3 if drift remains
replay --from SEQ [--dry-run] republishes, unchanged, the newest message of every profile with a message at or after SEQ (stream only, no database)
republish --id ID[,ID] | --all [--wait 60s] re-enqueues the current documents from Postgres (deleted profiles as tombstones) and waits until the publisher drained them
snapshot writes a snapshot now and prints its key, sequence and profile count

The drift kinds reconcile reports:

Kind Means --fix repairs it
missing Postgres has the profile, the stream has no message for it yes
behind the stream holds an older version, and none is waiting in the outbox yes
stale deleted in Postgres, the stream still holds a live document yes
ahead the stream holds a newer version than Postgres no: investigate
unknown the stream holds a profile Postgres does not know no: investigate
snapshot_ahead the latest snapshot holds a newer version than Postgres no: investigate

A profile whose newer version is still in the outbox is counted as in flight, not as drift. ahead and snapshot_ahead can follow a Postgres restore to an older backup; see Backups.

The command takes the database, NATS (--nats-url) and object storage (--s3-*; without it the snapshot checks are skipped, and snapshot refuses to run) settings of the other roles. Republished rows go through the outbox with kind republish, and edges that already hold a version ignore it. Outside env=dev, republish and reconcile --fix need --yes; both are logged with the actor cli:<user>@<host>.

fleetctl profiles status shows each PoP’s applied stream sequence (from Prometheus) and how far it is behind; fleetctl profiles replay is the same replay as above. When to use which is in the fleet runbook docs/fleet/runbooks/profile-stream-replay.md.