Skip to content

Architecture and the log model

This page describes what the code does, including the parts that are less flattering. If you find a claim here that the code contradicts, that is a bug in the page or the code: please report it.

How a query and its log record travel through opdnsYour devices send queries by anycast to a PoP. On the PoP, opdns-edge identifies the profile and applies policy; only allowed queries go to Unbound on loopback, which resolves them on the internet. The edge writes one record per query to a log ring; logs off (an hourly count instead of records), client IP off and domain logging off apply there. Records go to the logs stream in the control plane, kept at most 24 hours. Ingest applies destination and retention and keeps hourly counts, then stores records in ClickHouse for the cloud destination, or queues them in nodelogs for your node for up to 7 days. The queue is delivered over the link to your node's SQLite file. The API and dashboard read logs from ClickHouse, or from your node through the link. Devices at home that query your node directly never touch the cloud.PoP (one of many, reached by anycast)opdns control planeYour network (self-hosted, optional)Your devicesphone, laptop, routeropdns-edgeidentify · policy · answer blocksUnboundrecursion · cache · DNSSECThe internetauthoritative serverslog ring+ disk spool, 1 GiB caplogs off: hourly count onlyclient IP off: zeroed heredomains off: blanked herelogs.<pop>stream, 24 h maxingestdestination · retentionClickHousecloud, bothnodelogs.<id>self-hosted, both · 7 dAPIand dashboardlinklink.opdns.netHome devicespointed at the nodeopdns-nodeedge + own UnboundSQLite logsallowed onlylog batches downreads relayed,not storednever touches the cloud
querylog recorddashboard read
  1. Your device sends a query to an opdns address. Anycast routing delivers it to the nearest PoP (point of presence).
  2. On the PoP, opdns-edge first checks the client address against the source block list: a blocked source is refused without any further work. It then identifies the profile and device from the transport (Identification). A source that asks for a profile but does not encode a valid one (a malformed DoT or DoQ server name, a reserved IPv6 address) or names a profile that does not exist is refused, never answered unfiltered. Then it evaluates the profile’s policy (Policy order).
  3. A blocked or rewritten name is answered by the edge itself. So is the reserved check name (check.opdns.io and random names below it), after identification and before policy: its answer describes what the edge saw (PoP, profile, identification, transport, device name), it is never forwarded, logged, counted as profile traffic or charged to rate limits.
  4. An allowed name is forwarded to Unbound, on the same machine over loopback. Unbound does the recursion, caching and DNSSEC validation.
  5. The edge checks the answer (CNAME uncloaking, rebinding protection), writes one log record, and replies. Over UDP it first fits the answer to twice the question’s size (four times for an identified query), dropping records a resolver answer does not need and else truncating it so the client retries over TCP, unless the client proved its address with a DNS cookie; responses are also rate-limited per client network (Amplification).

Only allowed queries, and only from the edge on loopback. It never sees the profile, the device name, or any query the policy blocked, and it holds no per-profile configuration. Its cache is shared by everyone on that PoP.

It does not see the client’s address either, with one exception by design: a profile can opt into EDNS Client Subnet, and the edge then passes a prefix of the client’s address (or, in Full mode, the address) to Unbound to forward to authoritative servers. Forwarding needs the edge’s -unbound-ecs flag, which loads Unbound’s subnet cache; every query of a profile with ECS off then carries a /0 opt-out. Whether a PoP runs with it is set per environment in the fleet inventory (deploy/fleet/inventory.yaml, unbound.ecs, on unless set to false), from which fleetctl renders both the edge’s settings and Unbound’s configuration; the local simulation and production are both on. A PoP deployed with it off sends no client subnet, whatever the profile says.

Every PoP runs the same Unbound release (1.26.1 today), built from the checksum-verified NLnet Labs source rather than taken from a distribution, and the image build fails if the binary reports another version. It keeps a local copy of the root zone (RFC 8806): Unbound transfers the zone from the root servers’ published sources, accepts a copy only when its ZONEMD digest verifies, and answers its own iterator from it, so a cold lookup sends nothing to the root servers. Clients cannot query that copy directly. If the copy is missing or expired, Unbound asks the root servers as any resolver does: slower cold lookups, never an outage. Self-hosted nodes keep the same copy from the same sources, on by default; their owner can point it elsewhere or turn it off (Local root zone).

The edge reads its container’s cgroup v2 limits at start and sets the Go runtime from them: as many scheduler threads as its CPU quota or CPU set allows, and a soft memory limit at 90% of its memory limit, so the garbage collector works harder before the kernel would kill the process. The part left over covers what that limit does not count, such as the list artifact’s memory-mapped pages. GOMAXPROCS or GOMEMLIMIT set in the environment take precedence.

The edge produces one record per identified query (fields in the privacy policy, section 4). The three privacy switches act on the PoP, before the record exists outside the machine. Retention and destination act at ingest, in the control plane, after the record has crossed the internal logs stream.

Switch Applied What leaves the PoP
Logs off PoP no record; an hourly count of the profile’s queries by outcome
Client IP off PoP the record without the client address
Domain logging off PoP (and again at ingest) the record without the name, the answer addresses or rule text; the list name, device, outcome and timing stay
Retention ingest n/a: sets each stored row’s expiry
Destination ingest the record, whatever the destination

Records wait in a per-PoP ring buffer and, if the control plane is unreachable, a disk spool capped at 1 GiB (oldest dropped first). Logging is never on the query’s critical path.

For rows it stores in the cloud database, ingest adds three columns after the privacy switches have been applied, using the deployment’s GeoIP country database, if it has one, and classification data compiled into the control plane (GeoIP and owner classification):

  • the country of each answer address;
  • the company the query reached (Google, Apple, Meta, Amazon, Microsoft, or the CDNs Cloudflare, Akamai, Fastly), from its name or its answer addresses;
  • the client’s country, only when the profile logs client IP addresses. With that switch off, the client address never reaches ingest, and no country is looked up for it.

With domain logging off there are no names or answer addresses, so no answer countries and no company either. The lookups run on the control plane’s own servers; no address is sent to a third party. Records queued for your node get none of these columns.

What the cloud does and keeps for queries your devices send to the cloud addresses:

Destination Cloud resolves the query Record crosses logs Stored in ClickHouse Queued for your node Counters
cloud yes yes yes, until retention no yes
self-hosted yes yes no yes, until your node acknowledges it, 7 days max yes
both yes yes yes yes yes
none yes yes no no yes
logs off yes no no no yes, from the PoP’s hourly counts

Three things follow, and none of them is hidden:

  • Choosing your node as the destination does not mean the cloud never sees your queries. A query sent to a cloud address is resolved by the cloud, and its record passes through the logs stream and a per-profile queue on its way to your node. It is not stored in the cloud database.
  • A query sent to your own node never touches the cloud. Its record is written to the node’s SQLite file and stays there.
  • Dashboard views of node-held logs are relayed through the cloud. When you open Logs or Analytics, the request goes to your node over its link and the result comes back through opdns’s servers, readable in transit (TLS protects each hop). It is not stored. The node has no GeoIP data, so for the destinations view the control plane adds each answer address’s country on the way. While the live tail is open, the cloud asks the node for new rows every second (the stream) or every 2 seconds (the dashboard’s polling fallback).

Counters: ingest keeps hourly per-profile totals by outcome (no names, no addresses) for 400 days, from every record whatever its destination, and from the hourly counts the PoPs send for profiles with logs off and for queries blocked by a CSAM operator entry, which are never logged. Analytics reads its totals and timeline from them whenever nothing but the time range is filtered, whatever the destination, so the overview works with logs off and for the none and self-hosted destinations. Any other filter needs per-query rows and follows the destination (relayed to the node, or empty for none). A query sent straight to your node is not counted by the cloud.

Store What Retention
PoP memory the query while it is answered; rate-limit buckets (per client address, and per client network, name and type for response rate limiting) seconds; buckets until idle
PoP memory a check’s facts (PoP, profile id, identification, transport, ECS, device name) under its random name, readable over HTTPS by whoever has the name 60 seconds
PoP disk spool records while the control plane is unreachable until delivered; 1 GiB cap
logs stream records after PoP-side switches, before ingest-side ones 24 hours max
ClickHouse records for cloud and both your retention (1 to 90 days)
ClickHouse hourly rollups per profile and hour: counts by status, domain, device, answer network, company and answer country, summarised from the stored records expire with the records they summarise
ClickHouse aggregates hourly counts per profile, no names or addresses 400 days
nodelogs queue records for your node until acknowledged; 7 days and 1 GiB fleet-wide max
Your node’s SQLite your node’s own records and the queue’s deliveries the profile’s retention

The control plane runs at OVH; PoPs are rented servers in several countries. Regions and sub-processors are listed in the privacy policy, section 6.

Lists reach the PoPs as one signed, memory-mapped file, the list artifact: every blocked name once, with the set of lists that contain it. Format version 2 gives each name two sets, the lists that match the name itself and the lists that match names below it, so a list can hold suffix rules (example.com: the name and everything below), exact rules (=example.com) and children-only rules (*.example.com). The policy engine probes the query name against the first set and each of its parents against the second. Artifacts of format 1 (suffix rules only) still load. Before a build is published it is checked against a never-block guard list (opdns’s own names, the DNS root and TLD infrastructure and the like) and four gates (how much each list changed, the memory budget, the guard again, and popular sites newly blocked); a build that fails is not published and the PoPs keep the previous one. A build that passes goes to a few canary PoPs before the fleet (List rollouts). Each build also keeps a private record of which source line produced every entry, for answering “why is this blocked?”; it is a separate file, not part of the artifact the PoPs and nodes load.

Profiles reach the PoPs as whole documents on an internal stream, one per profile, newest version wins; a PoP that starts reads the newest snapshot and then follows the stream. Every message and snapshot is signed with Ed25519; a PoP configured with the public keys refuses an unsigned or badly signed one and keeps what it had (Signed profiles). The index that points at the newest snapshot is not signed, and the PoP’s own saved state is trusted as it is. The control plane rewrites a reserved system profile every minute and every PoP reports when it applied it, which measures the time from a change being saved to it being in force on each PoP (the target is under 2 seconds).

An enrolled node keeps one outbound WebSocket to link.opdns.net, authenticated by its node token. Over it:

  • the cloud sends profile changed and lists changed notices (the node then fetches the profile, which is signed, and verifies it against the keys it learnt at enrolment or that you pinned), log batches for the profile, and dashboard queries. Lists changed names the version the fleet serves, never a build still in its canary; after a rollback it names an older version, which the node installs (List updates and rollbacks);
  • the node sends acknowledgements, query results (at most link.max_inflight_queries dashboard queries run at once, 4 by default; one that finds no free slot within a second is answered at once with no rows and marked partial, overload) and periodic health: versions (node, Unbound, profile, lists, log schema), uptime, last sync, SQLite size and oldest record time, pending and dropped record counts, clock offset, whether Unbound is healthy, and the link’s round-trip time;
  • the node rotates its token: it asks for a new one every 90 days (or on opdns-node rotate-token), stores it, confirms, and the old one stops working 10 minutes later.

The node accepts no inbound connections from opdns. It keeps resolving if the link is down. The cloud marks it offline when the link closes, or within about 30 seconds when the link server holding it stops; after 10 minutes offline the account’s owners are emailed, unless they turned the alert off. opdns-node unenrol revokes the node’s token in the cloud before it deletes it locally.

On your network the node answers your local names (home.arpa by default, from its hosts setting and your DHCP leases) and reverse lookups of private addresses itself: they never reach Unbound, the internet or opdns, unless you name your router to forward unknown ones to (Local names). It answers only private, CGNAT, link-local, ULA and loopback sources (plus what you allow), so it is not an open resolver if it is reachable from the internet (Network access). Details: Logs on your node, Offline behaviour.

An operator can suspend a profile (for abuse, for example). The profile stays identified: its addresses, linked IPs and devices still resolve, and its rate limits apply, but on every PoP and enrolled node it compiles to no lists, rules, rewrites, security or parental controls, with logs off. The edge writes no log record and no counter for its queries, as on the public path. Operator blocks still apply. The customer’s changes are refused until it is reinstated; the operator’s reason travels with the profile to the PoPs, which ignore it, and is removed from every copy a customer receives (the profile API, the node’s document, the data export) (Suspended profiles).

A query that identifies no profile (plain IPv4 DNS from an address that is not linked, or DoH to /dns-query) is resolved unfiltered, with no per-query record and counters only, under per-source rate limits. See The public resolver path.

A small operator-level list, for legal orders and abuse, is checked before any profile rule, on the PoPs, for every profile and on the public path. Its blocks are answered NXDOMAIN, whatever the profile’s block mode, with Extended DNS Error 15 (Blocked), never 17, so you can tell them from your own filtering; the extra text is a reason code (legal_order, abuse, csam) and, except for CSAM, a public reference. Some entries apply only in one country. A query blocked by a CSAM entry is never logged, only counted.

A second list blocks client address prefixes that abuse the service: plain DNS and DoT queries from them are refused with EDE 18, and DoH, DoH3 and DoQ connections are closed, before identification and without a log record.

Both files are announced with their version and SHA-256; a PoP refuses a file whose content does not match the announcement, or that is older than it, and keeps its current rules. A file newer than the last announcement, or never announced, is still accepted, so that polling alone keeps working.

Self-hosted nodes receive neither list. How operators use them: Operator and source blocks. The policy and the transparency reports are linked from Policies.