The development environment
The local simulation (deploy/dev/) runs the whole service on one
machine with Docker Compose: every control-plane role with hot reload,
three PoPs running the real opdns-edge behind BGP anycast, a router, a
client on a simulated home network, a self-hosted node enrolled for a test
profile, Pebble for certificates, Mailpit for email, and Prometheus,
Alertmanager and Grafana. It needs OrbStack on macOS, or Docker Engine 28
or newer with Compose 2.33 or newer on Linux, and mise.
mise run dev # build, start, wait until every service is healthy, then verifymise run dev:down # stop, keep volumesmise run dev:reset # stop and delete volumesmise run dev ends with dev:verify, a smoke and integration test in
sections a to x (resolution, BGP failover, the control plane, both log
paths, the node, certificates, mail, list rollouts, the local root zone,
alert delivery, host metrics and more), with a PASS or FAIL per check.
--only e,g runs chosen sections.
Offline mode
Section titled “Offline mode”mise run dev -- --offlineAfter one warm online run, the whole topology runs with no internet
access, and the mode enforces it: every simulation network is created
with IP masquerading off, so no container can reach anything outside the
simulation while the published host ports keep working. The mode is
written to deploy/dev/.env (not tracked), so later commands keep it;
mise run dev without the flag switches back (the networks are
recreated, volumes kept).
| Normally from the internet | Offline |
|---|---|
| the root zone | the last transfer, kept in a volume that survives dev:reset; usable until its signatures expire, about two weeks after the last online run |
| recursion for internet names | fixture zones (example.com, example.org, example.net, nic.fr, ietf.org, denic.de) served by the simulation’s authoritative server; any other name is SERVFAIL |
| list sources | none needed: the list compiler always builds from fixtures |
| image builds and pulls | none: an offline start never builds or pulls, and stops naming any missing image |
| Go modules | a module cache volume, with GOPROXY=off |
It needs the network again for the first run on a machine, a Dockerfile change, a new Go dependency, a change to the node (its image is built, not hot-reloaded), or a root zone copy older than about two weeks. When it was introduced, every verify check (234 at the time) and the whole chaos suite passed offline.
Solo modes
Section titled “Solo modes”mise run dev:solo -- edge|cp|web|node [--no-verify]Starts only one track’s services, with stubs for the rest, prints the
cold start and footprint, then runs that track’s checks. mise run dev
brings the full topology back.
| Mode | What runs | Stubbed | Cold start (measured) |
|---|---|---|---|
edge |
one PoP, the router, the client, the authoritative server | no control plane: the edge boots from fixture profiles and the fixture list artifact, query logs go to stdout | about 16 s, 5 containers |
cp |
the databases, NATS, object storage, Mailpit, Pebble, Prometheus and every control-plane role but backups | no PoPs | about 7 s, 14 containers |
web |
what the dashboard needs from the API | as cp; run the dashboard on the host with mise run web:dev |
about 7 s, 11 containers |
node |
the node, its client network, and the control-plane roles it talks to | nothing: the node’s cloud is the real subset | about 26 s, 16 containers |
Loop-speed and footprint budgets
Section titled “Loop-speed and footprint budgets”mise run dev:budget [-- --update] [--no-reset] [--skip-verify]Measures the loop and the footprint and fails when a metric is over its
budget in deploy/dev/budgets.json: reset time, cold start (and its
build part), warm restarts of a control-plane role and of an edge, the
full verify time, memory after verify, running containers and image
sizes. It resets the simulation first unless given --no-reset.
Each budget is the recorded baseline plus headroom: 25 % for memory,
containers and images, 50 % for wall-clock timings (back-to-back runs on a
shared laptop vary by up to 35 %), capped by the design targets (cold start
120 s, warm restart 5 s, 6 GiB). --update records a new baseline, and
refuses to when verify failed. Baseline on 2026-09-30: cold start 32.7 s,
full verify 212 s, 1.7 GiB resident, 25 containers.
The dashboard has its own budget: pnpm check:bundle in web/ fails when
a chunk’s gzipped size exceeds web/bundle-budget.json, with a hard limit
of 150 kB for the first load. The first load is at 149.9 kB; raising the limit or moving more code into
lazy chunks is planned.
Ambient traffic
Section titled “Ambient traffic”mise run dev:traffic # on: 10 queries a second over every transportmise run dev:traffic -- on --qps 50 --profile dev001,dev002 --seed 7mise run dev:traffic -- check # allowed, blocked and rewritten entries per profile within 10 smise run dev:traffic -- status # or cpu, run, offA small generator (deploy/dev/tools/traffic, off by default) sends
queries from the simulated home client, so they cross the router and
reach the anycast address (or the node) like a real device’s. The mix is
seeded, so the same seed and flags give the same sequence: 70 % allowed,
20 % blocked, 10 % rewritten (spread over what each profile has), over
Do53, DoT, DoH and DoQ, from three named devices. It costs about 0.7 % of
a core and 7 MiB at the default rate. The seed data gives rewrites only to
a test profile that logs nothing, so the Logs page shows no rewritten
entry yet; extending the seed is an open item.
Clock skew
Section titled “Clock skew”Containers share the host’s clock, so the chaos suite skews time another
way: a development-only go build -overlay replaces the standard
library’s time.Now with one that adds OPDNS_DEV_CLOCK_OFFSET to the
wall clock (monotonic time untouched). The suite’s clock scenario runs
the node at +10 minutes, -10 minutes and +2 days, and one PoP’s edge at
plus and minus 5 minutes, and checks what the design says should happen.
mise run test:devclock proves release builds carry no such hook. The
scenario found that the node’s clock gate has no tolerance below the
newest list’s build time, so a node 10 minutes slow turns DNSSEC
validation off for 10 minutes after each list publish; adding a tolerance
is planned.
Integration tests
Section titled “Integration tests”Go integration tests (the integration build tag) share one harness,
internal/testharness: each test gets fresh state (its own Postgres
database, ClickHouse database, NATS and object storage) on the services
named by OPDNS_TEST_PG, OPDNS_TEST_CH, OPDNS_TEST_NATS and
OPDNS_TEST_S3 (what scripts/ci/services.sh env prints). A test whose
service is not set is skipped. With OPDNS_TEST_SERVICES=docker a package
starts throw-away containers for the services it needs, on free loopback
ports, and removes them afterwards (OPDNS_TEST_KEEP=1 keeps them); they
never touch the simulation’s stack. Fakes of an edge and of a self-hosted
node are in internal/testharness/fake; cptest, ingesttest and
fleettest seed the control plane, ingest and fleet on top of it.
CI has two tiers (2026-10-01), so a push pays only for what it touched:
ci-light.ymlon every push tomainand every pull request, except docs-only changes: Go build, vet, race tests and a short fuzz pass; gofmt, golangci-lint, shellcheck, actionlint and the Nomad job check (plusmise run test:devclockwhen Go code ordeploy/changes); the OpenAPI lint (mise run lint:spec, see API versioning); the rules tests; the dashboard’s lint, types, unit tests and bundle budget; the status Worker; Terraform with mocked providers; and the gitleaks scan. Each job runs only when its paths changed;ci-resultis the one required check.backlog.ymlwhendocs/backlog/changes.ci-heavy.ymlnightly at 03:30 UTC, by hand (gh workflow run ci-heavy.yml -f jobs=simulation), and on a push tomainthat touches its paths: the simulation, the solo edge, the Docker-backed integration and fleet tests, the images (linux/amd64 only), the dashboard e2e, Lighthouse and visual checks, and the local Talos cluster.
mise run ci:local remains the local gate.
Every night (.github/workflows/ci-heavy.yml, job nightly): the rules
tests, the config drift report, the online simulation with verify, the
chaos suite, a reset, the offline simulation with verify, and the budgets
against the generous deploy/dev/budgets.ci.json (about 90 minutes in
all), plus the full control-plane load test in a second job. Failures open
nothing: each step’s verdict, times and measurements go to the run
summary, and the budget result is kept as an artifact.