Alerting
A small team needs every page to reach a phone, and a separate signal when
alerting itself is broken. Alerts start as Prometheus rules
(deploy/dev/prometheus/rules/*.rules.yml; production loads the same
files), go to Alertmanager, and from there to two receivers. A third path,
the dead-man check, sits outside Alertmanager on purpose.
Prometheus rules --> Alertmanager --> page (webhook to the phone + email) --> warn (email) --> deadman (Watchdog heartbeat, every minute) | dead-man check: no heartbeat for 5 min --> WatchdogMissing, paged on its own pathRoutes
Section titled “Routes”Both Alertmanager configurations, the simulation’s
(deploy/dev/alertmanager/alertmanager.yml) and production’s
(deploy/alertmanager/alertmanager.yml), have the same routing tree and
inhibition rules; only the receivers differ. Alerts are grouped by
alertname and pop.
| Match | Receiver | Wait before the first notification | Between updates | Repeated while firing |
|---|---|---|---|---|
alertname=Watchdog |
deadman |
0 s | 1 min | on every 1-minute tick (repeat interval 50 s) |
severity=page |
page |
10 s | 1 min | every hour |
severity=warn, and anything else |
warn |
30 s | 5 min | every 4 hours |
The default route is warn, so an alert without a severity still arrives
somewhere. Resolved notifications are sent for page and warn.
| Receiver | Simulation | Production |
|---|---|---|
page |
the alert sink’s /webhook/page and an email to [email protected] in Mailpit |
a mobile push webhook (URL in /secrets/page-webhook-url) and an email |
warn |
the alert sink’s /webhook/warn and an email to [email protected] |
an email |
deadman |
fleetctl deadman’s /heartbeat |
an external heartbeat service (URL in /secrets/heartbeat-url) |
The emails use shared templates (deploy/alertmanager/templates): the
subject reads like [WARN] PoPDrained pop=sin (1 firing). Secrets are
files rendered by the deployment, never inline in the configuration.
Inhibition
Section titled “Inhibition”Some alerts mute others while they fire, so one incident pages once:
| While this fires | Muted |
|---|---|
PrefixWithdrawnEverywhere |
every alert with a pop label, and MultiplePoPsWithdrawn |
CertificateExpiresVerySoon |
CertificateExpiresSoon for the same PoP |
HostDiskFull |
HostDiskFilling for the same PoP, server and mount point |
For planned work, silence by label for the window (amtool silence add pop=sin --duration 2h ...), never by disabling rules. A drain already
produces only the PoPDrained warning. Never silence Watchdog: the
dead-man check would page five minutes later.
The dead-man check
Section titled “The dead-man check”Watchdog is a rule that always fires (vector(1)). Alertmanager sends
it to the deadman receiver every minute, and something outside the
alerting path listens: when no heartbeat has arrived for 5 minutes, it
pages WatchdogMissing over its own webhook and email, not through
Alertmanager. That covers every way alerting can die: Prometheus down,
rule evaluation stuck, Alertmanager down, or the path from Alertmanager to
the receivers broken. It pages once, again every hour while the heartbeat
stays missing, and sends a resolved page when it returns.
In the simulation the listener is fleetctl deadman (compose service
deadman; GET /status shows the heartbeat count, the age of the last
one and whether it is paging):
| Flag | Environment | Default |
|---|---|---|
--max-age |
DEADMAN_MAX_AGE |
5m: page when no heartbeat arrived for this long |
--repeat |
DEADMAN_REPEAT |
1h: page again while missing (0: once) |
--check |
DEADMAN_CHECK |
15s: how often it checks |
--alertname |
DEADMAN_ALERTNAME |
Watchdog |
--webhook, --smtp, --mail-from, --mail-to |
DEADMAN_WEBHOOK_URL, DEADMAN_SMTP, DEADMAN_MAIL_FROM, DEADMAN_MAIL_TO |
where it pages (at least one) |
In production the same contract is kept by an external heartbeat service, outside both the control plane’s cloud and the PoPs’ network: the Watchdog pings it, and it notifies the phone when pings stop for 5 minutes.
A WatchdogMissing page means you are operating blind: check that DNS
itself answers, then find the broken hop in order (Prometheus ready, the
alerting-meta rule group healthy, Prometheus sees an active
Alertmanager, Alertmanager ready, the Watchdog present in Alertmanager,
the deadman receiver’s webhook not failing). Afterwards, read the
ALERTS history for the gap: alerts may have fired and resolved unseen.
AlertmanagerNotificationsFailing (warn, 10 minutes) names an integration
(email or webhook) whose notifications fail. Failed notifications are
retried by Alertmanager while the alert keeps firing.
The drill
Section titled “The drill”mise run dev:verify -- --only w # delivery: drain pop-sin, PoPDrained reaches the sink and Mailpit; inhibitionmise run dev:alerting:deadman-drill # stop Alertmanager, WatchdogMissing after 5 min, start it, resolved (about 6 min)mise run test:rules # rule unit tests, and route and inhibition tests against both configurationsRehearsed on 2026-09-30 in the local simulation: a drain reached the
warn webhook in 98.9 s (rule for: 1m, 15 s evaluation, 30 s group
wait) and its email 0.1 s later; a synthetic PrefixWithdrawnEverywhere
reached page in 10.5 s while a per-PoP page sent with it was
suppressed; with Alertmanager stopped, WatchdogMissing paged 262.7 s
after the stop, and the resolved page came 24.3 s after Alertmanager
started again. In production, amtool alert add alertname=DeliveryTest severity=page should reach the phone within the 10 s group wait.
Host alerts
Section titled “Host alerts”Every PoP server runs node_exporter 1.12.1 (pinned by digest), as the
node-exporter task of the pop-system Nomad job, listening on
127.0.0.1:9102 and scraped as job node; in the simulation it is a
service in the PoP image. deploy/dev/prometheus/rules/host.rules.yml
turns it into nine alerts, each with unit tests and a section in the
host alerts runbook;
the Grafana dashboard opdns / hosts shows the same series per PoP and
server.
| Alert | Severity | Fires when |
|---|---|---|
HostExporterDown |
warn | the exporter has not answered scrapes for 5 minutes (no other host alert can fire for that server meanwhile) |
HostDiskFilling |
warn | a file system under 15 % free for 15 minutes |
HostDiskFull |
page | under 5 % free for 5 minutes (mutes the warning) |
HostHighLoad |
warn | 5-minute load over 1.5 per CPU for 15 minutes |
HostNICErrors |
warn | drops and errors over 1 a second on an interface for 10 minutes |
HostUDPReceiveErrors |
warn | the kernel drops over 10 UDP datagrams a second for 10 minutes (DNS queries lost before the edge reads them) |
HostConntrackFull |
warn | the connection-tracking table over 80 % full for 5 minutes (DNS ports should bypass it) |
HostClockUnsynchronised |
warn | the clock unsynchronised, or more than 50 ms off, for 10 minutes |
HostRebootRequired |
warn | a reboot has been pending for a day; reboot through drain with fleetctl fleet reboot (Fleet operations) |
The common escape hatch is the same for all of them: if the server cannot keep serving, drain its PoP and let the anycast cloud absorb the traffic. Unbound’s own statistics are not host metrics: the edge exports them (Resolver operations).
mise run dev:verify -- --only x checks the exporter on every PoP, the
host series, UDP receive errors and the Unbound cache hit ratio during a
dnsperf run, and the recording rules and dashboards.