Skip to content

Offline behaviour and troubleshooting

The node needs opdns only to receive changes. Every connection is outbound from the node.

  • It keeps resolving with the last good profile and list artifact, saved in its data directory (profile.json, lists/). They survive restarts, and so do the list signing keys it learnt (lists_keys.json), so it can still verify its lists after a reboot without the internet.
  • It keeps logging locally.
  • It reconnects its link with backoff (1 second up to 5 minutes, random jitter) and pulls the profile every cloud.profile_poll (15 minutes) as a safety net.
  • The status page and opdns-node status show a warning: “Cloud unreachable since …: resolving with the last profile and lists.”
  • The dashboard shows the node offline once its link is gone. After 10 minutes the account’s owners get an offline alert email, and another when it is back. Cloud logs for it wait in its queue (up to 7 days) and are delivered when it reconnects.

An enrolled node reports its link to opdns as one of these phases, the same everywhere: the local page (Cloud link, with a one-line detail), opdns-node status and /api/status (link.phase, link.summary), /healthz, the metric opdns_node_link_phase{phase} and a log line at start and on every change.

Phase Means
unenrolled no identity: a standalone node, or not enrolled yet
connecting opening the link
connected the link is up (the detail says since when)
backoff the link dropped; the next attempt is at the time shown
offline down for a minute or more (two of the cloud’s health intervals, when the dashboard also shows it offline); since when, and the next attempt
revoked the token was revoked in the dashboard; the node checks again slowly
stopped the node is shutting down

/healthz carries the phase (link.phase, with since, until and offline_since when they apply) but never the error text, so a health check can alert on offline without exposing why. The link being down does not make the node unhealthy: it keeps resolving.

Unbound cannot reach authoritative servers, so new names fail. Names in its cache are served, including stale ones for up to a day (serve-expired). Blocked names are answered locally as usual.

A device without a real-time clock boots with a wrong date. DNSSEC validation then fails, so NTP cannot resolve its server, so the clock is never fixed. The node breaks the loop itself:

  1. While the clock is implausible, it runs Unbound without DNSSEC validation.
  2. It checks the clock against NTP servers given as IP addresses (clock.ntp_servers, default Cloudflare’s anycast NTP), against the signed list artifact’s timestamp, and against its own build time.
  3. Once the clock is within clock.max_skew (1 hour), it turns validation on.

If your network blocks outbound NTP, set clock.ntp_servers to a reachable server by IP.

Symptom Check
bind: address already in use on port 53 another DNS server on the host; see Free port 53
unbound executable "unbound" not found install Unbound or set unbound.binary (binary installs only)
exits with status 2, “this node is not set up yet, so it is not serving DNS” enrol or configure standalone mode
rules apply but no blocklists no trusted list key: with lists.trust_cloud_key: false, lists.public_keys must hold the signing key; on a first start without the internet, the node has not learnt the key yet
/healthz fails Unbound not answering, or no profile or lists loaded yet; see the node’s log
many SERVFAIL right after boot on a Pi the clock gate; wait for the clock to be fixed
a device gets REFUSED with EDE 18 the device’s address is outside the node’s allowed networks (often a global IPv6 address); see Network access
the status page answers 403 or 421 opened from outside your network, or by an unknown host name; see Network access
opdns-node status shows a profile signature error the node refused a profile that was unsigned or signed by a key it does not trust and kept its last good one; see Profile signing keys
the local page’s ipv6 row says IPv4 only the host has no IPv6 route, so Unbound asks authoritative servers over IPv4 only; see IPv6 for upstream queries
opdns-node status shows a lists error about an older version a standalone node found an older latest.json than its artifact and kept its current lists; see List updates and rollbacks

Logs go to standard error, JSON by default (log_format: text for reading by eye, --log-level debug for more).