Resolver operations
The pinned Unbound
Section titled “The pinned Unbound”Every image that runs Unbound builds the same release from the NLnet Labs source tarball instead of taking what a distribution packages: 1.26.1 today.
| Image | Used by |
|---|---|
opdns-edge (cmd/opdns-edge/Dockerfile) |
the edge with -unbound-spawn |
opdns-node (cmd/opdns-node/Dockerfile) |
the self-hosted node’s container |
opdns-dev/pop (deploy/dev/pop/Dockerfile) |
the local simulation’s PoPs |
opdns-unbound (deploy/unbound/Dockerfile) |
the fleet’s resolver job with the Docker driver |
The version and the tarball’s SHA-256 live in one file,
internal/edge/unbound/version.env. The build (deploy/unbound/build.sh)
checks the checksum and compiles with the subnet cache (for ECS) and
libevent, without cachedb, HTTP/2 (the edge terminates DoH) or Python. It
installs unbound, unbound-control, unbound-checkconf,
unbound-anchor and unbound-host. unbound-check-version fails the
build, and again each runtime stage, if the binary reports another
version or lacks the subnet cache, validator, iterator or libevent. The
edge exports opdns_edge_unbound_build_info{pinned, running}: the release
the images build and what the running Unbound reports. The static
opdns-node binary uses the host’s own Unbound.
Upgrading
Section titled “Upgrading”- Read the release notes (renamed options, changed defaults, new statistics keys).
- Change
UNBOUND_VERSIONandUNBOUND_SHA256inversion.env, the checksum from the.sha256file next to the tarball, checked once against the PGP-signed.asc. - CI builds the images; a checksum or version mismatch fails them. Check
the new
stats_noresetkeys against the mapping ininternal/edge/unbound/stats.go. - Roll the fleet’s
resolverjob wave by wave (fleetctl rollout).
Renovate watches the Unbound release tags and opens the bump; its pull
request fails until someone updates the checksum by hand, on purpose. A
security advisory is a patch release: follow the same steps the same day.
The full procedure is deploy/unbound/README.md.
Unbound statistics
Section titled “Unbound statistics”The edge runs stats_noreset over Unbound’s remote-control socket every
-unbound-stats-interval (15 s) and exports the totals as about 40
opdns_edge_unbound_* metrics (labelled with the PoP like every edge
series). Per-thread keys are not exported. Among them:
| Metric | Is |
|---|---|
queries_total, cache_hits_total, cache_misses_total, cache_hit_ratio |
load and cache efficiency |
recursion_duration_seconds |
histogram of recursion time for replies that needed it |
queries_by_type_total, answers_by_rcode_total, queries_by_transport_total |
traffic shape |
answers_secure_total, answers_bogus_total, rrsets_bogus_total |
DNSSEC outcomes |
requestlist_current, requestlist_exceeded_total, queries_timed_out_total |
recursion backlog |
memory_bytes{pool}, cache_entries |
memory by pool, cache sizes |
queries_subnet_total, queries_subnet_cache_total |
ECS |
queries_authzone_total |
lookups answered from the local root zone |
uptime_seconds, up, scrape_errors_total |
Unbound’s uptime and the scrape itself |
The socket is the spawned Unbound’s with -unbound-spawn, or
-unbound-control (OPDNS_UNBOUND_CONTROL). Without one, only
build_info is exported. The fleet’s resolver job with the Docker
driver runs the edge and Unbound in separate containers; wiring the
socket through the allocation’s shared directory there is planned
(it works in the local Nomad mode). The self-hosted node does not export
Unbound statistics.
The opdns / edge Grafana dashboard has an Unbound row.
deploy/dev/prometheus/rules/unbound.rules.yml:
| Alert | Severity | Fires when |
|---|---|---|
UnboundCacheHitRatioLow |
ticket | under 60% of Unbound’s queries answered from cache for 15 minutes, at more than 1 query a second |
UnboundRecursionSlow |
ticket | recursion p95 over 500 ms for 15 minutes |
Tuning both thresholds on real fleet traffic is planned.
Local root zone
Section titled “Local root zone”Every PoP’s Unbound keeps a copy of the root zone (RFC 8806) as
auth-zone: "." and answers its own iterator from it, so cold lookups
skip the root servers. Clients cannot query the copy (for-downstream: no). Unbound transfers it from the published root sources (and ICANN’s
HTTPS copy), refreshes it on the root SOA timers, and accepts a copy only
when its ZONEMD digest verifies. The file is /var/lib/unbound/root.zone;
the edge reads its SOA serial and modification time (-root-zone-file)
and exports opdns_edge_root_zone_loaded, _serial and
_updated_timestamp_seconds.
A missing, rejected or expired copy (7 days after the last good refresh)
is not an outage: Unbound queries the root servers as any resolver does
(fallback-enabled: yes), and /healthz and the anycast announcement are
not affected. Do not drain a PoP for it.
| Alert | Severity | Fires when |
|---|---|---|
RootZoneStale |
warn | the copy has not been updated for 2 days |
RootZoneMissing |
warn | no copy loaded for an hour |
unbound-control list_auth_zones shows the serial Unbound holds and
unbound-control auth_zone_transfer . forces a transfer; the fleet
runbook docs/fleet/runbooks/root-zone-stale.md has the full diagnosis.
Production PoPs need outbound TCP and UDP 53 to the root primaries in the
Unbound config and HTTPS to www.internic.net.
A new resolver allocation starts without the file and runs on fallback for the few seconds of the first download (about 2 MB). Keeping Unbound’s state across rollouts (a sticky disk in the Nomad job) is planned.
Self-hosted nodes keep the same copy, on by default since 2026-09-30: the
node renders the same auth-zone "." from the edge’s template, with the
same default sources (the list lives in internal/edge/unbound and is
shared), the zone file in <data_dir>/unbound/root.zone, and the same
opdns_edge_root_zone_* series when the node’s page.metrics is on. The
user can change the sources or turn it off
(Local root zone).
In the local simulation the PoPs transfer from dev-auth, a secondary of
the real root; ROOT_ZONE=off in a PoP’s environment removes the
auth-zone.
Amplification
Section titled “Amplification”The open UDP path (the public resolver path)
is the fleet’s reflection surface: a spoofed question of 40 bytes can ask
for a 4 KiB answer sent to someone else. Since 2026-09-30 the edge bounds
what it sends back to what it receives, instead of relying on the
per-client token bucket alone. All of the mitigations below are on by
default, configured with -rrl-* flags (OPDNS_RRL_*), and apply to UDP
only: TCP, DoT, DoH and DoQ prove the source with their handshake (QUIC
also caps unvalidated sends at 3x by itself).
Target. On UDP, a response is at most 2x the request’s size on the wire for an unidentified client and 4x for an identified one (a query that carries a known profile). A larger answer is truncated (TC=1) so the client retries over TCP.
| Mitigation | Flags and defaults | Does |
|---|---|---|
| Response-size budget | -rrl-max-factor 2, -rrl-max-factor-identified 4, -rrl-minimal on |
caps a UDP answer to the factor times the question. It first strips what a resolver answer does not need (additional records, a positive answer’s authority NS set and their RRSIGs) and truncates only if that is not enough; DNSSEC proofs are kept, so validation survives |
| Response rate limiting (RRL) | -rrl-responses-per-second 50, -rrl-nxdomains-per-second 50, -rrl-errors-per-second 20, -rrl-window 2s, -rrl-slip 2, -rrl-ipv4-prefix 24, -rrl-ipv6-prefix 48, -rrl-identified-scale 4 |
counts UDP responses per client network (a flood aims at a victim network, not one address) and per kind: answers by name and type, NXDOMAIN by parent domain (random subdomains share a bucket), errors by rcode. Over the rate, every second limited response is a bare TC=1 answer, the rest are dropped. Identified responses have their own buckets at 4x the rates, so a flood on the public path cannot throttle a customer |
| ANY | -rrl-minimal-any on |
ANY over UDP gets the RFC 8482 synthesised HINFO; it is forwarded whole only over streams |
| UDP payload cap | fixed, 1232 bytes | advertised on every OPT record and enforced above the budget, even for a cookie-verified source |
| DNS cookies | -rrl-cookies on, -rrl-cookie-secret |
RFC 7873 with RFC 9018 server cookies. A client that echoes a valid server cookie has proven its address: it is exempt from the budget and from RRL (still bounded by the 1232-byte cap and the token bucket). A client cookie alone gets a fresh server cookie; a malformed cookie option is FORMERR. Cookies are answered on every transport and never forwarded to Unbound |
The defaults for RRL are a first pass; tuning them on real PoP traffic is planned. A busy NAT behind one /24 is the case to watch.
The audit
Section titled “The audit”Measured against opdns-edge with the pinned Unbound
(test/conformance, TestAmplificationAudit) with every mitigation on:
UDP bytes on the wire, request and response, and the factor they give.
The last two columns are the same answer whole over TCP, which is what
the reflected volume would be with no UDP cap.
| Query (UDP) | Unidentified req/resp B | Factor | TC | Identified req/resp B | Factor | TC | TCP resp B | Uncapped factor |
|---|---|---|---|---|---|---|---|---|
| ANY (RFC 8482 HINFO) | 37/58 | 1.57x | no | 37/58 | 1.57x | no | 115 | 3.11x |
| TXT ~700 B | 37/37 | 1.00x | yes | 37/37 | 1.00x | yes | 760 | 20.5x |
| TXT ~1,100 B | 39/39 | 1.00x | yes | 39/39 | 1.00x | yes | 1154 | 29.6x |
| TXT ~4 KiB | 37/37 | 1.00x | yes | 37/37 | 1.00x | yes | 4457 | 120x |
| DNSKEY +DO (RRSIG) | 40/40 | 1.00x | yes | 40/40 | 1.00x | yes | 1110 | 27.8x |
| signed CNAME chain +DO | 43/43 | 1.00x | yes | 43/43 | 1.00x | yes | 686 | 16.0x |
| signed NXDOMAIN +DO (NSEC) | 45/45 | 1.00x | yes | 45/45 | 1.00x | yes | 578 | 12.8x |
| A, EDNS buffer 4096 | 35/51 | 1.46x | no | 35/51 | 1.46x | no | 57 | 1.63x |
id.server CH TXT |
38/58 | 1.53x | no | 38/58 | 1.53x | no | 67 | 1.76x |
| REFUSED (unknown class) | 35/35 | 1.00x | no | 35/35 | 1.00x | no | 35 | 1.00x |
| A, small answer | 35/51 | 1.46x | no | 35/51 | 1.46x | no | 57 | 1.63x |
Every response stays under the target. The answers worth reflecting truncate to a TC=1 packet no bigger than the question. Because a plain DNS question is only about 40 bytes, 4x of it (about 160 bytes) still truncates a DNSKEY or signed answer, so identified customers also fetch large answers over TCP: that is the intended outcome, not a regression.
Under load (test/load/reflect, its own compose project, never the shared
stack), a reflection-shaped mix at 20,000 queries a second on a laptop
rig: the edge reflected 13.7x the bytes it received before the
mitigations and 0.5x after (slip and drops send back less than
arrives), while an identified client querying the same edge kept its p99
at 3.3 ms. The rig fails above 2x or above a 20 ms p99.
Metrics
Section titled “Metrics”| Metric | Is |
|---|---|
opdns_edge_amplification_capped_total{identified,action} |
UDP answers minimised or truncated to fit the budget |
opdns_edge_rrl_limited_total{class,action} |
responses RRL held back, by class (answer, nxdomain, error) and action (slipped as TC=1, dropped) |
opdns_edge_cookies_total{state} |
queries with a cookie option: client (a server cookie issued), valid (source verified), malformed (FORMERR) |
opdns_edge_udp_bytes_total{identified,direction} |
UDP bytes in and out, for the reflected ratio actually served |
The full design is docs/resolver/design.md, section 8, addendum
2026-09-30.
Container limits
Section titled “Container limits”At start the edge reads its cgroup v2 limits and sets GOMAXPROCS from
the CPU quota (cpu.max) or CPU set, and the Go soft memory limit
(GOMEMLIMIT) to 90% of memory.max, so it uses exactly its Nomad share
and the garbage collector reacts before the kernel would kill the task.
The remaining 10% covers what the Go limit does not count, such as the
list artifact’s memory-mapped pages. GOMAXPROCS or GOMEMLIMIT in the
environment win.