Skip to content

Alerting

A small team needs every page to reach a phone, and a separate signal when alerting itself is broken. Alerts start as Prometheus rules (deploy/dev/prometheus/rules/*.rules.yml; production loads the same files), go to Alertmanager, and from there to two receivers. A third path, the dead-man check, sits outside Alertmanager on purpose.

Prometheus rules --> Alertmanager --> page (webhook to the phone + email)
--> warn (email)
--> deadman (Watchdog heartbeat, every minute)
|
dead-man check: no heartbeat for 5 min
--> WatchdogMissing, paged on its own path

Both Alertmanager configurations, the simulation’s (deploy/dev/alertmanager/alertmanager.yml) and production’s (deploy/alertmanager/alertmanager.yml), have the same routing tree and inhibition rules; only the receivers differ. Alerts are grouped by alertname and pop.

Match Receiver Wait before the first notification Between updates Repeated while firing
alertname=Watchdog deadman 0 s 1 min on every 1-minute tick (repeat interval 50 s)
severity=page page 10 s 1 min every hour
severity=warn, and anything else warn 30 s 5 min every 4 hours

The default route is warn, so an alert without a severity still arrives somewhere. Resolved notifications are sent for page and warn.

Receiver Simulation Production
page the alert sink’s /webhook/page and an email to [email protected] in Mailpit a mobile push webhook (URL in /secrets/page-webhook-url) and an email
warn the alert sink’s /webhook/warn and an email to [email protected] an email
deadman fleetctl deadman’s /heartbeat an external heartbeat service (URL in /secrets/heartbeat-url)

The emails use shared templates (deploy/alertmanager/templates): the subject reads like [WARN] PoPDrained pop=sin (1 firing). Secrets are files rendered by the deployment, never inline in the configuration.

Some alerts mute others while they fire, so one incident pages once:

While this fires Muted
PrefixWithdrawnEverywhere every alert with a pop label, and MultiplePoPsWithdrawn
CertificateExpiresVerySoon CertificateExpiresSoon for the same PoP
HostDiskFull HostDiskFilling for the same PoP, server and mount point

For planned work, silence by label for the window (amtool silence add pop=sin --duration 2h ...), never by disabling rules. A drain already produces only the PoPDrained warning. Never silence Watchdog: the dead-man check would page five minutes later.

Watchdog is a rule that always fires (vector(1)). Alertmanager sends it to the deadman receiver every minute, and something outside the alerting path listens: when no heartbeat has arrived for 5 minutes, it pages WatchdogMissing over its own webhook and email, not through Alertmanager. That covers every way alerting can die: Prometheus down, rule evaluation stuck, Alertmanager down, or the path from Alertmanager to the receivers broken. It pages once, again every hour while the heartbeat stays missing, and sends a resolved page when it returns.

In the simulation the listener is fleetctl deadman (compose service deadman; GET /status shows the heartbeat count, the age of the last one and whether it is paging):

Flag Environment Default
--max-age DEADMAN_MAX_AGE 5m: page when no heartbeat arrived for this long
--repeat DEADMAN_REPEAT 1h: page again while missing (0: once)
--check DEADMAN_CHECK 15s: how often it checks
--alertname DEADMAN_ALERTNAME Watchdog
--webhook, --smtp, --mail-from, --mail-to DEADMAN_WEBHOOK_URL, DEADMAN_SMTP, DEADMAN_MAIL_FROM, DEADMAN_MAIL_TO where it pages (at least one)

In production the same contract is kept by an external heartbeat service, outside both the control plane’s cloud and the PoPs’ network: the Watchdog pings it, and it notifies the phone when pings stop for 5 minutes.

A WatchdogMissing page means you are operating blind: check that DNS itself answers, then find the broken hop in order (Prometheus ready, the alerting-meta rule group healthy, Prometheus sees an active Alertmanager, Alertmanager ready, the Watchdog present in Alertmanager, the deadman receiver’s webhook not failing). Afterwards, read the ALERTS history for the gap: alerts may have fired and resolved unseen.

AlertmanagerNotificationsFailing (warn, 10 minutes) names an integration (email or webhook) whose notifications fail. Failed notifications are retried by Alertmanager while the alert keeps firing.

Terminal window
mise run dev:verify -- --only w # delivery: drain pop-sin, PoPDrained reaches the sink and Mailpit; inhibition
mise run dev:alerting:deadman-drill # stop Alertmanager, WatchdogMissing after 5 min, start it, resolved (about 6 min)
mise run test:rules # rule unit tests, and route and inhibition tests against both configurations

Rehearsed on 2026-09-30 in the local simulation: a drain reached the warn webhook in 98.9 s (rule for: 1m, 15 s evaluation, 30 s group wait) and its email 0.1 s later; a synthetic PrefixWithdrawnEverywhere reached page in 10.5 s while a per-PoP page sent with it was suppressed; with Alertmanager stopped, WatchdogMissing paged 262.7 s after the stop, and the resolved page came 24.3 s after Alertmanager started again. In production, amtool alert add alertname=DeliveryTest severity=page should reach the phone within the 10 s group wait.

Every PoP server runs node_exporter 1.12.1 (pinned by digest), as the node-exporter task of the pop-system Nomad job, listening on 127.0.0.1:9102 and scraped as job node; in the simulation it is a service in the PoP image. deploy/dev/prometheus/rules/host.rules.yml turns it into nine alerts, each with unit tests and a section in the host alerts runbook; the Grafana dashboard opdns / hosts shows the same series per PoP and server.

Alert Severity Fires when
HostExporterDown warn the exporter has not answered scrapes for 5 minutes (no other host alert can fire for that server meanwhile)
HostDiskFilling warn a file system under 15 % free for 15 minutes
HostDiskFull page under 5 % free for 5 minutes (mutes the warning)
HostHighLoad warn 5-minute load over 1.5 per CPU for 15 minutes
HostNICErrors warn drops and errors over 1 a second on an interface for 10 minutes
HostUDPReceiveErrors warn the kernel drops over 10 UDP datagrams a second for 10 minutes (DNS queries lost before the edge reads them)
HostConntrackFull warn the connection-tracking table over 80 % full for 5 minutes (DNS ports should bypass it)
HostClockUnsynchronised warn the clock unsynchronised, or more than 50 ms off, for 10 minutes
HostRebootRequired warn a reboot has been pending for a day; reboot through drain with fleetctl fleet reboot (Fleet operations)

The common escape hatch is the same for all of them: if the server cannot keep serving, drain its PoP and let the anycast cloud absorb the traffic. Unbound’s own statistics are not host metrics: the edge exports them (Resolver operations).

mise run dev:verify -- --only x checks the exporter on every PoP, the host series, UDP receive errors and the Unbound cache hit ratio during a dnsperf run, and the recording rules and dashboards.