Fleet operations
Every procedure here withdraws a server’s anycast route before touching
it (a drain, through the server’s anycast-guard) and announces it
again only once it is healthy, so clients move to the next PoP instead of
failing. Each was rehearsed on 2026-09-30 in the local simulation under
load (mise run dev:nomad:rehearse -- <name>); the production-only steps
(a real host, the upstream BGP session, a kernel reboot) have not been
run yet.
| Command | Does |
|---|---|
fleetctl pop add --pop P [--server S] [--site-env FILE] [--soak 60s] [--dry-run] |
a new server from inventory entry to announced |
fleetctl pop remove POP [--server S] [--grace 30s] [--dry-run] |
drain, empty and scale away a server or a PoP |
fleetctl fleet reboot --all|--wave W|POP... [--grace 30s] [--pause 30s] [--dry-run] |
reboot servers one at a time through drain |
fleetctl fleet update --image TAG --all|--wave W|POP... |
install a host release and reboot, one server at a time |
fleetctl agent status [POP...] |
the host agents: up since, release, last action |
fleetctl pop status, pop drain, pop undrain |
the guards’ state; withdraw or return a PoP by hand |
Add a PoP
Section titled “Add a PoP”fleetctl pop add takes a server that is in deploy/fleet/inventory.yaml
(validate it first with fleetctl inventory validate) to announced, in
11 steps: inventory, rendered configuration in Nomad variables (BIRD,
Unbound and the edge’s per-server settings, plus the site settings from
--site-env), the wave’s node pool, the Nomad client joining, the
pop-system job placed, the new guard held drained, the wave’s
resolver job scaled up by count only (the wave’s other servers keep
their allocations), health (20 good probes, both BGP sessions up, a live
dead-man lease), an identity check (id.server on the unicast address
names the new server), the announcement, and a soak during which it must
stay announced.
--dry-runprints what each step would write without changing anything.- It can be started before the server is up (it waits for the Nomad client) and re-run after any failure: every step checks the current state first. Each waiting step says what it is waiting on.
- Start it only when every other PoP is announced (
fleetctl pop status). - In production the server is first installed with
deploy/pop/bootstrap.sh(hardening, the Nomad client with mTLS, the host agent), and its upstream BGP session must be configured on the provider’s side; step 8 waits for it.
Rehearsed: a fourth PoP added to the running three-PoP simulation was announced 56 seconds after its container was created (31 s after the command started), and took over its nearest client’s traffic with 0 of 12,058 load-test queries lost.
Remove a server or a PoP
Section titled “Remove a server or a PoP”fleetctl pop remove POP drains the guard(s) server by server and waits
--grace for queries in flight, drains the Nomad node with its system
jobs and waits until nothing runs there, scales the wave’s resolver job
down by count only (checking the wave’s other allocations are untouched)
and deletes the server’s three Nomad variables. --server S removes one
server of a PoP. It refuses while another PoP is withdrawn (--force
overrides that). What is left is printed: remove the entry from the
inventory, stop the server’s Nomad client and nomad node purge it.
Rehearsed: the fourth PoP removed in 15 seconds, its client back on the next PoP, 0 of 3,481 queries lost.
Reboot and host updates
Section titled “Reboot and host updates”fleetctl fleet reboot and fleetctl fleet update --image TAG walk the
selected servers in wave order, canary first, one server fleet-wide at
a time, never two of a PoP together. For each server:
- Guardrails: no
resolverrollout running, every other server announced, the server’s host agent answering and idle. - Drain, confirm BIRD withdrew the route, wait
--grace. - Act: the host agent reboots, or runs the host update with the tag and reboots only if it succeeded.
- Wait for the new boot (and, for an update, the new release).
- Health: the Nomad node ready with every allocation running, the guard healthy after 20 good probes (still drained), both BGP sessions up, a live lease.
- Undrain, wait for the announcement, pause
--pause, next server.
Any failure stops the whole run with a non-zero exit, and the server
being worked on stays drained: safe, and nothing further is touched.
Undrain it once healthy (fleetctl pop undrain) and re-run for the
servers not done yet. Every step is an event on fleet.events
(fleetctl events). Start with --wave canary, then --all.
The host agent. pop-agent runs on each server as the root systemd
unit opdns-pop-agent, outside Nomad, so a server whose Nomad client or
container runtime is broken can still be rebooted away. It answers only
status, reboot and update, only for requests naming its own server,
on fleet.agent.<pop>, and refuses to reboot or update unless its local
guard reports the server drained and withdrawn, so a replayed or mistaken
command cannot take an announced server down. It answers first and acts
two seconds later. An update runs /usr/local/sbin/opdns-pop-update TAG
and records the release in /etc/opdns-release.
Without NATS, drain with the guard’s local socket and reboot over SSH, one server at a time. Without the Nomad servers, do not reboot: see below.
Rehearsed: all three PoPs rebooted in sequence in 102 seconds, 0 of 21,001 load-test queries lost; a canary update in 29 seconds, 0 lost. A first run, on a heavily loaded laptop, lost 589 queries when two other PoPs’ guards withdrew on probe timeouts while the third was rebooting, so no PoP was announced for about 3 seconds; the guardrail cannot prevent a second, independent failure during a maintenance window.
fleet update --image updates the host release, not the edge (which
rolls through fleetctl rollout). A separate NATS token for pop-agent
in production is planned (the simulation reuses the fleet command token),
and so is a review of the guard’s 400 ms health probe with two failures,
which withdrew two healthy PoPs in that loaded run.
When the Nomad servers are down
Section titled “When the Nomad servers are down”The PoPs do not need the Nomad servers to serve DNS: allocations keep
running, the guards keep checking health and announcing or withdrawing
(BIRD is local), and fleetctl pop status, pop drain and pop undrain
work over NATS. What stops is everything that schedules or reads Nomad:
rollouts, pop add and pop remove, fleet reboot and update (they
refuse, unless --no-nomad for a server that is already dark), and any
change carried by Nomad variables (certificate or command token
rotations, per-server configuration).
During an outage: freeze deploys (no rollouts, PoP adds or removals,
certificate or token rotations), and do not reboot a PoP server: after
a reboot its Nomad client cannot restore its allocations until it reaches
a server, so it stays dark. After the servers return, check every node is
ready with the same allocation ids, then run a no-change rollout. The
runbook is docs/fleet/runbooks/nomad-outage.md.