Skip to content

Fleet operations

Every procedure here withdraws a server’s anycast route before touching it (a drain, through the server’s anycast-guard) and announces it again only once it is healthy, so clients move to the next PoP instead of failing. Each was rehearsed on 2026-09-30 in the local simulation under load (mise run dev:nomad:rehearse -- <name>); the production-only steps (a real host, the upstream BGP session, a kernel reboot) have not been run yet.

Command Does
fleetctl pop add --pop P [--server S] [--site-env FILE] [--soak 60s] [--dry-run] a new server from inventory entry to announced
fleetctl pop remove POP [--server S] [--grace 30s] [--dry-run] drain, empty and scale away a server or a PoP
fleetctl fleet reboot --all|--wave W|POP... [--grace 30s] [--pause 30s] [--dry-run] reboot servers one at a time through drain
fleetctl fleet update --image TAG --all|--wave W|POP... install a host release and reboot, one server at a time
fleetctl agent status [POP...] the host agents: up since, release, last action
fleetctl pop status, pop drain, pop undrain the guards’ state; withdraw or return a PoP by hand

fleetctl pop add takes a server that is in deploy/fleet/inventory.yaml (validate it first with fleetctl inventory validate) to announced, in 11 steps: inventory, rendered configuration in Nomad variables (BIRD, Unbound and the edge’s per-server settings, plus the site settings from --site-env), the wave’s node pool, the Nomad client joining, the pop-system job placed, the new guard held drained, the wave’s resolver job scaled up by count only (the wave’s other servers keep their allocations), health (20 good probes, both BGP sessions up, a live dead-man lease), an identity check (id.server on the unicast address names the new server), the announcement, and a soak during which it must stay announced.

  • --dry-run prints what each step would write without changing anything.
  • It can be started before the server is up (it waits for the Nomad client) and re-run after any failure: every step checks the current state first. Each waiting step says what it is waiting on.
  • Start it only when every other PoP is announced (fleetctl pop status).
  • In production the server is first installed with deploy/pop/bootstrap.sh (hardening, the Nomad client with mTLS, the host agent), and its upstream BGP session must be configured on the provider’s side; step 8 waits for it.

Rehearsed: a fourth PoP added to the running three-PoP simulation was announced 56 seconds after its container was created (31 s after the command started), and took over its nearest client’s traffic with 0 of 12,058 load-test queries lost.

fleetctl pop remove POP drains the guard(s) server by server and waits --grace for queries in flight, drains the Nomad node with its system jobs and waits until nothing runs there, scales the wave’s resolver job down by count only (checking the wave’s other allocations are untouched) and deletes the server’s three Nomad variables. --server S removes one server of a PoP. It refuses while another PoP is withdrawn (--force overrides that). What is left is printed: remove the entry from the inventory, stop the server’s Nomad client and nomad node purge it.

Rehearsed: the fourth PoP removed in 15 seconds, its client back on the next PoP, 0 of 3,481 queries lost.

fleetctl fleet reboot and fleetctl fleet update --image TAG walk the selected servers in wave order, canary first, one server fleet-wide at a time, never two of a PoP together. For each server:

  1. Guardrails: no resolver rollout running, every other server announced, the server’s host agent answering and idle.
  2. Drain, confirm BIRD withdrew the route, wait --grace.
  3. Act: the host agent reboots, or runs the host update with the tag and reboots only if it succeeded.
  4. Wait for the new boot (and, for an update, the new release).
  5. Health: the Nomad node ready with every allocation running, the guard healthy after 20 good probes (still drained), both BGP sessions up, a live lease.
  6. Undrain, wait for the announcement, pause --pause, next server.

Any failure stops the whole run with a non-zero exit, and the server being worked on stays drained: safe, and nothing further is touched. Undrain it once healthy (fleetctl pop undrain) and re-run for the servers not done yet. Every step is an event on fleet.events (fleetctl events). Start with --wave canary, then --all.

The host agent. pop-agent runs on each server as the root systemd unit opdns-pop-agent, outside Nomad, so a server whose Nomad client or container runtime is broken can still be rebooted away. It answers only status, reboot and update, only for requests naming its own server, on fleet.agent.<pop>, and refuses to reboot or update unless its local guard reports the server drained and withdrawn, so a replayed or mistaken command cannot take an announced server down. It answers first and acts two seconds later. An update runs /usr/local/sbin/opdns-pop-update TAG and records the release in /etc/opdns-release.

Without NATS, drain with the guard’s local socket and reboot over SSH, one server at a time. Without the Nomad servers, do not reboot: see below.

Rehearsed: all three PoPs rebooted in sequence in 102 seconds, 0 of 21,001 load-test queries lost; a canary update in 29 seconds, 0 lost. A first run, on a heavily loaded laptop, lost 589 queries when two other PoPs’ guards withdrew on probe timeouts while the third was rebooting, so no PoP was announced for about 3 seconds; the guardrail cannot prevent a second, independent failure during a maintenance window.

fleet update --image updates the host release, not the edge (which rolls through fleetctl rollout). A separate NATS token for pop-agent in production is planned (the simulation reuses the fleet command token), and so is a review of the guard’s 400 ms health probe with two failures, which withdrew two healthy PoPs in that loaded run.

The PoPs do not need the Nomad servers to serve DNS: allocations keep running, the guards keep checking health and announcing or withdrawing (BIRD is local), and fleetctl pop status, pop drain and pop undrain work over NATS. What stops is everything that schedules or reads Nomad: rollouts, pop add and pop remove, fleet reboot and update (they refuse, unless --no-nomad for a server that is already dark), and any change carried by Nomad variables (certificate or command token rotations, per-server configuration).

During an outage: freeze deploys (no rollouts, PoP adds or removals, certificate or token rotations), and do not reboot a PoP server: after a reboot its Nomad client cannot restore its allocations until it reaches a server, so it stays dark. After the servers return, check every node is ready with the same allocation ids, then run a no-change rollout. The runbook is docs/fleet/runbooks/nomad-outage.md.