Skip to content

The development environment

The local simulation (deploy/dev/) runs the whole service on one machine with Docker Compose: every control-plane role with hot reload, three PoPs running the real opdns-edge behind BGP anycast, a router, a client on a simulated home network, a self-hosted node enrolled for a test profile, Pebble for certificates, Mailpit for email, and Prometheus, Alertmanager and Grafana. It needs OrbStack on macOS, or Docker Engine 28 or newer with Compose 2.33 or newer on Linux, and mise.

Terminal window
mise run dev # build, start, wait until every service is healthy, then verify
mise run dev:down # stop, keep volumes
mise run dev:reset # stop and delete volumes

mise run dev ends with dev:verify, a smoke and integration test in sections a to x (resolution, BGP failover, the control plane, both log paths, the node, certificates, mail, list rollouts, the local root zone, alert delivery, host metrics and more), with a PASS or FAIL per check. --only e,g runs chosen sections.

Terminal window
mise run dev -- --offline

After one warm online run, the whole topology runs with no internet access, and the mode enforces it: every simulation network is created with IP masquerading off, so no container can reach anything outside the simulation while the published host ports keep working. The mode is written to deploy/dev/.env (not tracked), so later commands keep it; mise run dev without the flag switches back (the networks are recreated, volumes kept).

Normally from the internet Offline
the root zone the last transfer, kept in a volume that survives dev:reset; usable until its signatures expire, about two weeks after the last online run
recursion for internet names fixture zones (example.com, example.org, example.net, nic.fr, ietf.org, denic.de) served by the simulation’s authoritative server; any other name is SERVFAIL
list sources none needed: the list compiler always builds from fixtures
image builds and pulls none: an offline start never builds or pulls, and stops naming any missing image
Go modules a module cache volume, with GOPROXY=off

It needs the network again for the first run on a machine, a Dockerfile change, a new Go dependency, a change to the node (its image is built, not hot-reloaded), or a root zone copy older than about two weeks. When it was introduced, every verify check (234 at the time) and the whole chaos suite passed offline.

Terminal window
mise run dev:solo -- edge|cp|web|node [--no-verify]

Starts only one track’s services, with stubs for the rest, prints the cold start and footprint, then runs that track’s checks. mise run dev brings the full topology back.

Mode What runs Stubbed Cold start (measured)
edge one PoP, the router, the client, the authoritative server no control plane: the edge boots from fixture profiles and the fixture list artifact, query logs go to stdout about 16 s, 5 containers
cp the databases, NATS, object storage, Mailpit, Pebble, Prometheus and every control-plane role but backups no PoPs about 7 s, 14 containers
web what the dashboard needs from the API as cp; run the dashboard on the host with mise run web:dev about 7 s, 11 containers
node the node, its client network, and the control-plane roles it talks to nothing: the node’s cloud is the real subset about 26 s, 16 containers
Terminal window
mise run dev:budget [-- --update] [--no-reset] [--skip-verify]

Measures the loop and the footprint and fails when a metric is over its budget in deploy/dev/budgets.json: reset time, cold start (and its build part), warm restarts of a control-plane role and of an edge, the full verify time, memory after verify, running containers and image sizes. It resets the simulation first unless given --no-reset. Each budget is the recorded baseline plus headroom: 25 % for memory, containers and images, 50 % for wall-clock timings (back-to-back runs on a shared laptop vary by up to 35 %), capped by the design targets (cold start 120 s, warm restart 5 s, 6 GiB). --update records a new baseline, and refuses to when verify failed. Baseline on 2026-09-30: cold start 32.7 s, full verify 212 s, 1.7 GiB resident, 25 containers.

The dashboard has its own budget: pnpm check:bundle in web/ fails when a chunk’s gzipped size exceeds web/bundle-budget.json, with a hard limit of 150 kB for the first load. The first load is at 149.9 kB; raising the limit or moving more code into lazy chunks is planned.

Terminal window
mise run dev:traffic # on: 10 queries a second over every transport
mise run dev:traffic -- on --qps 50 --profile dev001,dev002 --seed 7
mise run dev:traffic -- check # allowed, blocked and rewritten entries per profile within 10 s
mise run dev:traffic -- status # or cpu, run, off

A small generator (deploy/dev/tools/traffic, off by default) sends queries from the simulated home client, so they cross the router and reach the anycast address (or the node) like a real device’s. The mix is seeded, so the same seed and flags give the same sequence: 70 % allowed, 20 % blocked, 10 % rewritten (spread over what each profile has), over Do53, DoT, DoH and DoQ, from three named devices. It costs about 0.7 % of a core and 7 MiB at the default rate. The seed data gives rewrites only to a test profile that logs nothing, so the Logs page shows no rewritten entry yet; extending the seed is an open item.

Containers share the host’s clock, so the chaos suite skews time another way: a development-only go build -overlay replaces the standard library’s time.Now with one that adds OPDNS_DEV_CLOCK_OFFSET to the wall clock (monotonic time untouched). The suite’s clock scenario runs the node at +10 minutes, -10 minutes and +2 days, and one PoP’s edge at plus and minus 5 minutes, and checks what the design says should happen. mise run test:devclock proves release builds carry no such hook. The scenario found that the node’s clock gate has no tolerance below the newest list’s build time, so a node 10 minutes slow turns DNSSEC validation off for 10 minutes after each list publish; adding a tolerance is planned.

Go integration tests (the integration build tag) share one harness, internal/testharness: each test gets fresh state (its own Postgres database, ClickHouse database, NATS and object storage) on the services named by OPDNS_TEST_PG, OPDNS_TEST_CH, OPDNS_TEST_NATS and OPDNS_TEST_S3 (what scripts/ci/services.sh env prints). A test whose service is not set is skipped. With OPDNS_TEST_SERVICES=docker a package starts throw-away containers for the services it needs, on free loopback ports, and removes them afterwards (OPDNS_TEST_KEEP=1 keeps them); they never touch the simulation’s stack. Fakes of an edge and of a self-hosted node are in internal/testharness/fake; cptest, ingesttest and fleettest seed the control plane, ingest and fleet on top of it.

CI has two tiers (2026-10-01), so a push pays only for what it touched:

  • ci-light.yml on every push to main and every pull request, except docs-only changes: Go build, vet, race tests and a short fuzz pass; gofmt, golangci-lint, shellcheck, actionlint and the Nomad job check (plus mise run test:devclock when Go code or deploy/ changes); the OpenAPI lint (mise run lint:spec, see API versioning); the rules tests; the dashboard’s lint, types, unit tests and bundle budget; the status Worker; Terraform with mocked providers; and the gitleaks scan. Each job runs only when its paths changed; ci-result is the one required check.
  • backlog.yml when docs/backlog/ changes.
  • ci-heavy.yml nightly at 03:30 UTC, by hand (gh workflow run ci-heavy.yml -f jobs=simulation), and on a push to main that touches its paths: the simulation, the solo edge, the Docker-backed integration and fleet tests, the images (linux/amd64 only), the dashboard e2e, Lighthouse and visual checks, and the local Talos cluster.

mise run ci:local remains the local gate.

Every night (.github/workflows/ci-heavy.yml, job nightly): the rules tests, the config drift report, the online simulation with verify, the chaos suite, a reset, the offline simulation with verify, and the budgets against the generous deploy/dev/budgets.ci.json (about 90 minutes in all), plus the full control-plane load test in a second job. Failures open nothing: each step’s verdict, times and measurements go to the run summary, and the budget result is kept as an artifact.