Files
echolot/docs/build-status.md
T
mrambossekandClaude Opus 5 0071e00003 app: long runs that watch, and a relay for the port that keeps moving
Long mode starts listeners at t=0 and keeps them running past the
battery: a network-change watcher that finally fills networks[].changes[]
(defined since the schema's first draft, never populated), an RSSI log, a
ping series giving loss and jitter over minutes, and mDNS listening for
the whole window. This is the class of fault a short run cannot see - a
link that drops for four seconds between two probes is reported healthy
by both of them. run.mode records which question was asked, because
silence means different things in the two modes.

The adb relay replaces the retired beacon: AdbRelay watches adbd's own
mDNS with the resolve-once discipline the beacon learned the hard way
(resolving re-arms adbd and pops a notification), a foreground service
keeps it alive with the screen off, and the heartbeat re-posts the cached
endpoint rather than re-resolving. It exists because mDNS does not cross
subnets and the wireless-debug port rotates every few minutes.

Also records why LLDP/CDP cannot follow SSDP into long mode: both are raw
L2 frames, so they need CAP_NET_RAW - root tier, not app, and Shizuku's
shell user does not have it either.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 14:43:33 +02:00

1476 lines
104 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Echolot — build status & next steps
Last updated: 2026-07-29.
## Decided
- Name **Echolot**; domain echo-lot.app; scheme `echolot://`; namespace `app.echo_lot.*`
(hyphen→underscore; appIds/packages can't contain hyphens). Prober appId `app.echo_lot.prober`.
- Stack: native **Kotlin + Jetpack Compose**, no Flutter.
- Tiers: `app` (no root), **`shizuku` in v1** (wireless-ADB pairing), `root` future.
- **License (decided 2026-07-30):** all code **GPL-3.0-or-later**; the specs in `docs/` **CC-BY-4.0**.
Rationale: the moat is the no-root platform research, which is trivially liftable into a
proprietary repackage — copyleft is the only option that prevents that, and the F-Droid /
network-engineer audience reads GPL as a trust signal rather than friction. The specs go
permissive on purpose: a wire protocol only becomes a standard if anyone can implement it.
**Not AGPL for the server** — the ASP loophole it closes is speculative here, while blanket
corporate AGPL bans would hit exactly the enterprise network teams most likely to self-host.
Sole copyright holder, so relicensing the server to AGPL later stays possible.
Open: substitute a real legal copyright holder for "Echolot contributors" in the SPDX headers.
## Specs (in `docs/`, alongside this file)
- `feature-catalog-and-feasibility.md`, `measurement-schema.md`, `probe-protocol.md`. Considered draft-complete and reviewed by the user.
## Capability prober — DELIVERED (as zip, 2026-07-29)
Full Kotlin/Compose project scaffolded: `app.echo_lot.prober`, minSdk 26 / target+compile 35,
AGP 8.7.3, Kotlin 2.0.21, Compose BOM 2024.10, Shizuku api+provider 13.1.5, kotlinx-serialization.
Probes implemented: `link.snapshot`, `icmp.ping4/6` (unprivileged ICMP datagram), `sockopt.matrix`
(TTL/TOS/RECVERR/MTU_DISCOVER), `trace.errqueue_reachable`, `multinetwork.request_and_bind`,
`local.mdns_discover`, `peer.ble_advertise`, `shizuku.command_battery` (ip neigh / ip -6 route /
ip addr / ip monitor / dumpsys network_stack DHCP+IpClient / dumpsys wifi). JSON export via share
intent; results carry verdict + raw evidence.
Could NOT be compiled in the cloud sandbox: dl.google.com (Google Maven) and services.gradle.org
are proxy-blocked, and no device is reachable for on-device runs. Build + iterate locally
(Android Studio / Claude Code). Wrapper is pinned to Gradle 8.14.3.
## Website — `web/` (2026-07-30, LIVE at https://echo-lot.app)
Single-page site for echo-lot.app on Cloudflare Workers: static `public/` from the edge, a small
Worker for the stable short URLs (`/apk`, `/apk.sha256`, `/fdroid`, `/source`) and `/api/latest`.
Version info is resolved from the Gitea releases API at request time (edge-cached 5 min), so
tagging a release is the only publish step — the homepage shows the latest version/date/size via
progressive-enhancement JS and falls back to pre-release copy otherwise. Auto light/dark via
`prefers-color-scheme`: dark = sonar display, light = chart-recorder printout, same amber-accent
token system. Deploy: `npx wrangler deploy` from `web/`, or CI via
`.gitea/workflows/deploy-site.yml` (needs `CLOUDFLARE_API_TOKEN` + `CLOUDFLARE_ACCOUNT_ID`
secrets); full story in `web/README.md`. Branding applied (2026-07-30): real favicon/wordmark
copied into `web/public/assets/` (copies — `assets/branding/` is the source of truth), palette
switched to abyss/teal/amber with teal-as-instrument semantics, sonar-sweep hero replaced by a
Focus-lattice panel (rogue-DHCP finding vignette), tagline "measure, don't guess".
## Branding — `assets/branding/` (2026-07-30)
Identity chosen: the **Focus** mark — a quiet lattice of nodes with exactly one under examination
(amber) inside teal viewfinder brackets. Tone is surgical measurement, not scanning (sonar-sweep
visuals and the "sound out your network" tagline were explicitly rejected); tagline is
**"measure, don't guess"**. Palette: abyss `#071522`, tile `#0E2433`, instrument teal `#35E0C4`,
finding amber `#FFB454`, foam `#E8F4F2` — teal is always the instrument, the single amber point is
the finding. The wordmark is hand-drawn monoline SVG paths (zero font dependency); the second "o"
is the signature teal ring + amber ping. Assets: `icon.svg` (+ `icon.png` 512² — Gitea avatars reject SVG), `icon-adaptive-{foreground,background}.svg`
(108dp adaptive-icon layers, art inside the 66dp safe circle), `wordmark-on-{dark,light}.svg`,
`banner.svg` (1200×300, wired into the repo README), `social-preview.svg`/`.png` (1280×640, for the
Gitea/GitHub social-preview slot). The website's placeholder favicon/wordmark can now be replaced
from these.
## First device report — OnePlus 15, Android 16/SDK 36 (2026-07-30)
Archived at `echolot-prober/reports/CPH2747-android16-sdk36.json`. Collection workflow that works:
deploy over (wireless) adb, but the **user** starts the run and exports the JSON manually — probe
runs reliably kill the adb-over-wifi session (multinetwork/multicast churn the link adb rides on).
OS findings:
- `link.snapshot`, `icmp.ping4` (38 ms), `sockopt.matrix` 4/4, `peer.ble_advertise`: SUPPORTED.
- **`trace.errqueue_reachable` SUPPORTED** — full Os-API errqueue path exists on Android 16;
the C-over-JNI shim (next-steps #2) may be unnecessary on modern devices. Still needs a real
errqueue round-trip test, and older devices may differ.
- `icmp.ping6` EAGAIN — topology, not capability: wifi was v4-only, v6 only on cellular, socket
used the default network. Probe should bind per-network (also covers the ping4 case properly).
- `shizuku.command_battery` SUPPORTED, shell(2000): real `ip -6 route` RA data (lifetimes,
per-table), `ip monitor` NEIGH events, wifi dump with full DhcpResults. Vendor formats look
parseable.
- `local.mdns_discover`: lock acquired, discovery ran, 0 services in the 4 s window — needs a
longer window / known-good reference network before calling it more than formally SUPPORTED.
Prober bugs the run exposed (all fixed same day):
- rotation wiped results + cancelled runs → state moved to a ViewModel;
- `multinetwork.request_and_bind` ERROR → missing `CHANGE_NETWORK_STATE` in the manifest;
- `rtt_ms: "38,1"` → device locale leaked into report numbers, now `Locale.ROOT`;
- 3/7 Shizuku commands `SHIZUKU_BIND_TIMEOUT` → bind/unbind-per-command raced Shizuku; now one
bind per battery (`execBatch`).
Run 2 (same device, fixed APK — `…-run2.json`) confirmed all three fixes:
- `multinetwork.request_and_bind` **SUPPORTED** — wifi + cellular acquired and bound concurrently
(downKbps 39094 / 13231); requesting cellular even brought up a second, dual-stack PDN
(`rmnet_data2`, private v4 /30 + carrier-NAT DNS) next to the IPv6-only default PDN — exactly
the per-network behavior the production app needs.
- Shizuku battery **7/7** in 2.9 s (was 4/7 in 34 s): full ARP/NDP neighbor table, `ip addr`,
IpClient ProvisioningConfiguration incl. SSID/BSSID and APF caps (v6000, 4096 B).
- `rtt_ms: "32.0"` — locale fix holds.
Still open (probe improvements, not blockers): ping6 should bind per-network (still EAGAIN via
v4-only default wifi); mDNS window too short to see real services (0 found on a network that has
them); `ip monitor` 2 s window caught nothing on a quiet net — consider longer/event-triggered
capture; IP_MTU read via getsockoptInt reflection unavailable — reader belongs in the errqueue
shim if that lands.
## Go server skeleton — `server/` (2026-07-30)
Pure-Go stdlib implementation of the spec's core: control plane (enroll with single-use tokens,
profile with SPKI pin, sessions with the §2.4 HKDF key schedule) + UDP data plane (ELT1 header,
HMAC gate, 1024-wide anti-replay, ECHO_RESP with observation block, TIMESYNC, §3.4
anti-amplification). Wire format covered by `go test` (roundtrip, replay/bad-HMAC/unknown-prefix
silent drops); full enroll→profile→session flow smoke-tested live. Two run modes: **container**
(autodetected via /.dockerenv etc. or `--docker`; config from `ECHOLOT_*` env; distroless image,
`network_mode: host` required — Docker NAT would falsify the observed sources the protocol
measures) and **native** (`--install-systemd`/`--uninstall-systemd`, opt-in `--self-update`
against a Gitea releases API). CI: `.gitea/workflows/build-server.yml` — tests on push;
`server-v*` tags push the image to the Gitea registry + attach linux amd64/arm64 binaries that
self-update consumes. Not yet: TCP/TLS echo, STUN, canary DNS, actions, observations API,
real admin UI, self-update signature verification.
Prober build numbering: ascending `versionCode` shown on screen + as `proberBuild` in the JSON
report; current deployed build = **3**. Bump on every deployed change.
## Build-2 reports: OnePlus 15 (A16) + Lenovo TB330FU (A15) — 2026-07-30
Archived as `…-build2.json` in `echolot-prober/reports/`. The findings:
- **`traceroute.udp4` SUPPORTED on BOTH devices** (SDK 35 + 36): 6 real hops via pure-Kotlin
`Os.recvmsg(MSG_ERRQUEUE)` in ~250 ms, identical paths (A1 → Cloudflare). **The C-over-JNI
errqueue shim is dead** for Android 15+; keep it in mind only if pre-15 devices matter.
- Per-network ICMP works as designed: phone ping6 = "cellular only, default has no v6 path"
(SUPPORTED with topology evidence instead of the old bare ERROR); per-network RTTs
(wifi 20 ms vs cellular 201 ms on v4).
- **Network finding (dogfood!): hudeWLAN advertises a v6 default route via RA but hands out no
global v6 address** — both devices have only link-local + a ::/0 route on wifi. That's a
half-broken IPv6 config on the LAN, exactly the class of issue the product should flag.
- mDNS: meta-query returns 0 on BOTH devices (NsdManager limitation, now confirmed twice);
concrete `_http._tcp` finds 4 real services (3× "Magic 2 LAN", EPSON WF-7840) on both.
Production inventory must enumerate concrete types, not trust the meta-query.
- Tablet Shizuku: binder alive, permission granted, but UserService bind timed out (0/7) —
first-spawn dex extraction on slow storage suspected; build 3 raises the bind window to
25 s + one retry. Phone stays 7/7 with rich neighbor/RA/DHCP evidence.
**Build-3 verdict: not timing.** Both 25 s attempts timed out (50 s total) — the UserService
spawn genuinely fails on the Lenovo/A15. Build 4 adds a reflection fallback to the legacy
`Shizuku.newProcess` remote-process API when the bind fails; `exec_path` in the evidence says
which path ran. Whatever the outcome, core-shizuku must not assume UserService works
everywhere. (Phone build 3: still 7/7; `ip_monitor` now catches provoked NEIGH
PROBE→REACHABLE transitions, and a mid-path router dropping one TTL round showed the "*"
hop path works in traceroute.udp4.)
**Build-4 verdict (`…-build4.json`): the fallback WORKS.** Tablet Shizuku now SUPPORTED via
`exec_path: newProcess fallback` — 6/7 commands with real data (full neighbor table, per-netId
v6 route tables, Lenovo's IpClient dump — a distinctly different format from OnePlus's, i.e.
the per-vendor parser sample we wanted). **core-shizuku design consequence: dual-path executor
is mandatory** — UserService where it binds, newProcess where it doesn't; both confirmed live
on real hardware. Known rough edges: `ip monitor` times out under newProcess (subshell trick
doesn't survive that path); the 2×25 s bind wait before falling back is worth shortening once
a device is known-bad. Tablet UserService root cause still open — multi-user (3 users) is the
prime suspect; Shizuku app is current (13.6.0), so the stale-app theory is dead.
- `ip_monitor` returned no events this run even with the provoked gateway ping (gateway was
already REACHABLE, so no NEIGH transition happened). Evidence-dependent, not a bug.
## Pipeline shakeout — 2026-07-30, partially done
Repo pushed to git.rambossek.at/EchoLot/echolot (ssh :2222). Actions runner works:
`server-test` green on push. `server-v0.1.0` release shipped with all assets
(`echolot-server_linux_{amd64,arm64}` + `SHA256SUMS`) after two workflow fixes: mixed
paths+tags triggers never fire (split into test-server.yml / build-server.yml), and a greedy
`sed` grabbed a nested user id instead of the release id → uploads 404ed (first-match grep now;
release creation also falls back to GET-by-tag on re-runs). Binaries verified downloadable →
`--self-update` has its artifact source. **Image job green too** on the dedicated
`compilesau-echolot` runner (label `echolot`): needed a PAT with package read/write as the
`REGISTRY_TOKEN` repo secret — the built-in Actions token is rejected by the container
registry. `git.rambossek.at/echolot/echolot-server:{v0.1.0,latest}` pulls anonymously
(3.9 MB distroless, verified). Full chain proven: tag → tests → binaries → release assets →
registry image → anonymous pull. Note the tag convention: image tags carry the bare version
(`v0.1.0`), release tags the namespaced one (`server-v0.1.0`); selfupdate normalizes.
## Toolchain upgrade + build-5 tablet re-run (2026-07-30)
Moved to the laptop (USB adb — far more stable than wireless, as expected). Upgraded the build
off the JDK-21 ceiling: **AGP 8.7.3 → 9.2.0, Gradle 8.14.3 → 9.6.0, Kotlin 2.0.21 → 2.2.10**,
which lets the toolchain run on JDK 25. Two AGP-9 migration edits were required:
- AGP 9 has **built-in Kotlin support**, so applying `org.jetbrains.kotlin.android` alongside it
fails with "extension with name 'kotlin' already registered" — the alias must be dropped
(`kotlin.compose` / `kotlin.serialization` stay; they are separate compiler plugins).
- The `kotlinOptions { jvmTarget }` block came from that plugin and no longer resolves; AGP
derives jvmTarget from `compileOptions` instead.
`compileSdk`/`targetSdk` 35 → 36, because Android Studio only ships API 36.1 by default and
installing 35 was avoidable churn. **Build-5 re-run on the TB330FU is verdict-identical to
build-4 across all 10 probes** (`…-build5.json`) — the toolchain jump and targetSdk 36 are
behaviour-neutral on this device, including the `newProcess` Shizuku fallback (still 7/7,
`exec_path: newProcess fallback`, UserService still never binds).
Also of note: `traceroute.udp4` is SUPPORTED here (6 hops via errqueue, "no native shim needed"),
which is what retires next-step 2 below — the JNI shim is not required on either known device.
**OnePlus 15 build-5 re-run (`CPH2747-android16-sdk36-build5.json`): all 10 SUPPORTED, identical
to build-3.** `exec_path: UserService` on the phone vs `newProcess fallback` on the tablet, both
7/7 — the dual-path executor requirement is now confirmed on the current toolchain on both
devices. Phone is on Android 16 / SDK 36, i.e. `targetSdk = 36` is also exercised natively there.
Two collection-loop gotchas found while driving the phone over USB:
- **`icmp.ping6`'s verdict is topology-dependent, not build-dependent.** A first run with Shizuku
stopped and only 2 networks up reported UNSUPPORTED; with cellular+wifi present it is the
expected SUPPORTED / "echo reply on cellular only — default network has no v6 path". Do not
read a ping6 verdict without checking `link.snapshot`'s `network_count` in the same report.
- **The Shizuku permission dialog times out after 30 s** (`shizuku.command_battery` logs
`Shizuku permission not granted` at ~30027 ms). Screenshot→read→tap round-trips can exceed
that. Grant the permission *before* starting a run, or accept that the first run after a fresh
install burns the Shizuku probe; the grant is persistent, so the next run is clean.
- The phone auto-rotates; tap coordinates must be recomputed per orientation, and `uiautomator
dump` can report a stale idle state right after a tap — trust the exported JSON's `durationMs`
values as ground truth for whether a run really completed, not the button label.
## Next steps
1. Build locally, run on several physical devices (varied Android versions/vendors), collect the
JSON reports — especially the real per-device Shizuku dump formats.
2. ~~If `trace.errqueue_reachable` = PARTIAL, add a C-over-JNI errqueue shim.~~ **Retired** —
SUPPORTED on both known devices; `traceroute.udp4` reads real hops via `Os.recvmsg` +
`StructMsghdr` reflection, so no `:native` module is needed.
3. ~~Start the Go server skeleton per probe-protocol.md.~~ **Shipped** — live on fmr since
v0.2.0 (2026-07-31); see the server sections below.
4. Fold confirmed capabilities into the production `core-probe` / `core-shizuku` modules.
## Production probe server — LIVE on dedicated VM "fmr" (2026-07-31)
`echolot-server v0.2.0` runs natively (systemd, no docker) on a dedicated VM: 2×IPv4 + 2×IPv6
service addresses (fmr-1/fmr-2.echo-lot.app, dual-stack DNS), a third IPv6 (`::2`) reserved for
SSH only — verified untouched by the daemon (explicit multi-address binds, no wildcard).
Control: fmr-1:8443 (SPKI pin `zRV9qkiLnRexAeh4RrSfJzbPWO+U/2Oj2/NVM/KfXlg=`, verified
externally over v4+v6). UDP data plane on all four service addresses :8442 — the second IP is
the stun-5780 substrate. Daily randomized self-update timer installed (checksum-verified
against SHA256SUMS; signature verification landed 2026-08-02 — see "Release signing" below).
Host config in `/etc/echolot-server.env`. SSH access for sessions: `ssh claude-echolot`.
## Server v0.3.0 — STUN + TCP echo + observations + actions (2026-07-31)
Shipped and deployed to fmr via the server's own `--self-update` (first real exercise:
checksum-verified download v0.2.0→v0.3.0, atomic replace, restart — worked). Added over v0.2.0:
- **STUN** (RFC 5389 + 5780): 4 service addrs × primary/alt-port grid. Externally verified on
v4 AND v6 — binding success with XOR-MAPPED, RESPONSE-ORIGIN, OTHER-ADDRESS present, so the
profile now advertises **`stun-5780`** (the second IP earns its keep).
- **TCP echo** (:8441): JSON greeting with observed src + real Linux TCP_INFO — verified
externally `mss:1440` (v6, 150060), options `[sack,wscale]`, then byte-echo.
- **Observations API** `GET /v1/sessions/{id}/observations` (per-packet UDP view, connect-back
results, TCP records correlated by source IP).
- **Actions** `POST /v1/sessions/{id}/actions`: `delayed_echo` (DELAYED_ECHO at the observed
data-plane source — NAT-lifetime primitive) and `connect_back` (dials the control-plane
source, records connected/refused/timeout+rtt).
- Capabilities computed from what's actually wired: `udp-probe, delayed-echo, connect-back,
tcp-echo, stun-5780`.
Still not implemented: TLS-echo/JA4, HTTP echo, tls-reference, canary DNS (§6.1 reference
records), and the train/big-send/frag/throughput actions. Admin UI still token-mint + health only.
## Canary DNS live — server v0.3.1 on fmr (2026-07-31)
Zone `c.echo-lot.app` delegated (NS → fmr-1/fmr-2) and authoritative on all 4 service IPs
udp+tcp/53. Verified through full public recursion: `ttl-5` A→192.0.2.5 (Cloudflare), `ttl-3600`
AAAA→2001:db8::3600 (Google), `big-txt` TXT returned (TCP fallback, truncated over UDP as
designed). End-to-end session attribution works: a `<nonce>.<session-prefix>.c.echo-lot.app`
query resolved via a public resolver shows up in `GET /v1/sessions/{id}/observations` →
`dns_canary` with the resolver's real egress IP, transport, and EDNS. First real test already
caught a finding: **Google applies 0x20 case randomization** (mixed-case qname), Cloudflare does
not — captured via `case_preserved`. Capabilities now: udp-probe, delayed-echo, connect-back,
tcp-echo, stun-5780, canary-dns. Kept the hand-rolled stdlib DNS (no miekg/dns) — validated
against independent clients. Deployed via `--self-update` (v0.3.0→v0.3.1, checksum-verified).
## Server v0.3.2 + v0.3.3 (2026-07-31)
- **v0.3.2 — control-plane security (live on fmr, externally verified):** `POST /v1/echo`
reflects the received request head+body (b64) and observed TLS (version/cipher/SNI/ALPN) —
captured real SNI `fmr-1.echo-lot.app` and an injected header over public TLS1.3; `GET
/v1/tls-reference` returns the served DER chain + pin (cross-checked against the openssl-derived
pin). Optional cleartext echo listener (default off). Capability `http-echo`.
- **v0.3.3 — MTU probe (data plane):** MTU_PROBE (0x09) → small MTU_ACK (0x0A) carrying the
received datagram size; client DF-probes increasing sizes to find path MTU / black holes. ACK
is tiny → never amplifies. Tested.
- **Note on trains:** upstream trains (TRAIN_DATA 0x03) are already observable — every HMAC-valid
packet is recorded (seq/t_rx/size/type) with no per-packet response, so loss/reordering/inter-
arrival are visible via GET observations. The dedicated data-plane TRAIN_REPORT (0x05) is
deferred: §3.4 anti-amplification means it needs an asymmetric grant + columnar multi-datagram
encoding — a focused batch, not a corner to rush.
Remaining spec: tls-echo (ClientHello+JA4), TRAIN_REPORT, big/frag-send, throughput, downtrain;
real admin UI.
## Server self-test + host tuning — v0.3.4/v0.3.5, fmr proven good (2026-07-31)
The daemon now proves its own host is a clean measurement target:
- **sysctl audit** (`GET /admin/selftest`, startup warnings): on first run it flagged exactly 4
real issues on fmr — accept_ra=1 on a static-v6 host, accept_redirects=1, send_redirects=1,
icmp_ratelimit=1000. Recommended `server/deploy/99-echolot-sysctl.conf` applied (v6 default
route/addrs are proto static with 0 RA-derived routes, so disabling accept_ra is safe —
verified v6 egress intact after). Now sysctl_ok=true, 0 warnings.
- **egress-MTU self-proof**: DF PMTUD via IP_MTU_DISCOVER + getsockopt IP_MTU (v0.3.4 had a bug —
read IP_MTU without connecting → ENOTCONN; v0.3.5 connects first). fmr reports 1500 on both v4
and v6 → mtu_ok=true, so client MTU tests are trustworthy.
- Both signals ride in the profile as `server_selftest{mtu_ok,sysctl_ok}` so a client can skip
MTU testing when the server can't support it honestly.
fmr profile now: `{mtu_ok: true, sysctl_ok: true}`.
## Server v0.3.6 — tls-echo / JA4: spec §4 COMPLETE (2026-07-31)
The elt-echo TLS variant runs on the TCP-echo port (8441), multiplexed by a timed 0x16 peek
(plain echo stays server-speaks-first; a TLS ClientHello routes to the TLS path). It captures
the full ClientHello, returns it raw (b64) + as a JA4 fingerprint (FoxIO), then TLS byte-echoes —
the sec.clienthello_echo evidence. Hand-rolled ClientHello parser (ciphers/exts/ALPN/
supported_versions/sig-algs, GREASE-excluded), unit-tested. **Cross-client verified on fmr**:
openssl → t13d3013eo (30 ciphers), python ssl → t13d1712eo (17) — different stacks, different
fingerprints, correct _a structure both. Capability tls-echo.
Spec §4 (TCP/TLS/HTTP/STUN) is now fully implemented. Server capabilities: udp-probe,
delayed-echo, connect-back, http-echo, tcp-echo, tls-echo, stun-5780, canary-dns.
Remaining spec: §5 heavy actions (downtrain/big_send/frag_send/throughput) + the TRAIN_REPORT
retrieval path — all gated on the anti-amplification grant machinery (§3.4) — and a real admin UI.
## Production app — echolot-app/, core-protocol proven live (2026-07-31)
Started the Android client, bottom-up from the verifiable spine. `echolot-app/` is a multi-module
Gradle build; `core-protocol` is a **pure Kotlin/JVM** module (no Android SDK) implementing the
client half of probe-protocol.md: SPKI-pinned control plane (enroll/profile/session via
HttpsURLConnection — Android-API-1 compatible, hostname verification off since trust is the pin),
HKDF-SHA256 session keys, and the ELT1 UDP data plane (HMAC gate, ECHO+observation, MTU probe) —
byte-compatible with the Go server. Unit tests pass incl. the RFC 5869 HKDF vector (so key
derivation provably matches the server). **Verified END-TO-END against live fmr** via
`scripts/test-fmr.sh` (mint token over SSH → enroll on public control plane → run LiveServerTest):
profile (8 caps), session, ECHO rtt ~11ms with the observation block round-tripping the client's
observed NAT port, MTU probe 1400→1400, observations 298B. Two client bugs found+fixed doing it:
java.net.http did hostname verification (switched to HttpsURLConnection) and ECHO needed ≥72-byte
requests for the full 40-byte observation to survive §3.4 anti-amplification. Next: core-measurement
(schema types), core-probe (port prober probes), core-shizuku (dual-path), Compose UI.
## App: core-measurement + core-engine — full server-facing vertical proven (2026-07-31)
Two more pure-Kotlin/JVM modules, both verifiable without a device:
- **core-measurement**: the measurement-schema.md document model (two-clock, columnar trains,
test-type registry, anonymization types, finding-requires-evidence). The §7.3 deterministic
verdict derivation is implemented + unit-tested; document JSON round-trips.
- **core-engine**: the run engine composing core-protocol probes into core-measurement documents.
Injected clock/UUID source (pure, testable). Runs a server ECHO train → RTT distribution, loss,
and NAT-rebinding detection (from the server's observed source port) as train.udp_updown.
**Verified END-TO-END against fmr**: 20-packet train, 0% loss, RTT 1.7/2.5/6.9ms, single
observed port (no rebinding) → valid MeasurementDocument (2.3kB), overall GREEN.
So the whole server-facing stack — protocol client → engine → schema document → verdict — is now
proven against the live server, no device needed. Next modules (core-probe device-tier,
core-shizuku dual-path, Compose app) are Android + need on-device verification.
## App: installable APK — core-probe + Compose UI (2026-07-31)
The production Android app assembles. Android toolchain in echolot-app mirrors the prober (AGP
9.2 built-in Kotlin — do NOT also apply kotlin.android, it double-registers the `kotlin`
extension; that was the one build gotcha). Modules added:
- **core-probe** (Android lib): Probe→core-measurement Test abstraction; NetworkInventory
(LinkProperties → measurement networks[]), LinkSnapshotProbe (link.snapshot), IcmpProbe
(per-network icmp.ping4/6, ported from the prober's validated per-network logic).
- **app** (Compose): RunViewModel orchestrates probes → assembles a MeasurementDocument with a
§7.3 summary + first-pass findings; Compose UI shows overall/ per-category traffic lights,
networks, tests (status/metrics), findings; JSON export via share intent. Survives rotation
(ViewModel). App-tier only for now; server-facing (core-engine) and Shizuku tier are additive
follow-ups (app degrades gracefully without them, like the prober).
Debug APK: 9.5 MB, `echolot-app/app/build/outputs/apk/debug/app-debug.apk`. Not yet run on device
(needs the user's phone). core-shizuku (dual-path executor) deferred as additive.
## App verified on-device — both phones (2026-07-31)
The production Echolot app runs on real hardware, on BOTH devices, via the beacon-managed adb:
- OnePlus 15 (CPH2747, A16) and Lenovo TB330FU (A15): link.snapshot OK, icmp.ping4 OK
(per-network, RTT ~40ms), icmp.ping6 FAILED → finding "No IPv6 ICMP path on any active network"
→ category ipv6 yellow → **Overall YELLOW**. The verdict is driven by the real broken-LAN IPv6
(RA default route, no global prefix) we first found with the prober — the product now surfaces
it end to end (probe → schema → verdict → traffic-light UI).
Wireless-adb beacon (tools/adb-beacon) made this practical: both devices self-report their
rotating wireless-debug port to fmr:443; a PC connector keeps adb connected. Debugged live
against the restricted LAN (cleartext policy, egress filtering, shared-LAN mDNS crossing, fast
port rotation) — all handled.
## App: core-shizuku (shell tier) built + wired (2026-07-31)
Ported the prober's validated Shizuku executor into the app as a library module:
- AIDL IUserService + UserService (runs `sh -c` as shell/root in the Shizuku-spawned process),
Shizuku provider merged into the app manifest.
- **ShizukuRunner: the build-4 dual-path executor** — bind the UserService (25s + retry) where it
works (OnePlus 7/7), fall back to the legacy `Shizuku.newProcess` reflection API where it never
binds (Lenovo). `exec_path` records which path ran.
- ShizukuProbe: runs the shell battery (ip neigh / ip -6 route / ip addr / ip monitor /
network_stack DHCP / wifi dump) and emits a shizuku-tier `link.ip_monitor` Test with the raw
per-device dumps as evidence + commands_ok/exec_path metrics. Self-degrades to UNSUPPORTED when
Shizuku isn't running.
Wired into RunViewModel (runs after app-tier probes; sets tiers.shizuku). App APK assembles clean.
On-device test deferred: at build time no device was reachable (tablet wifi/beacon dropped on the
churning LAN; phone wireless debugging disabled to stop reconnect notifications). Will verify on a
device later — expecting UserService on the OnePlus, newProcess fallback on the Lenovo, per the
prober.
## App on-device: net.captive_portal verified, Shizuku degrades correctly (2026-07-31)
Installed the app (core-shizuku + net.captive_portal) on the OnePlus 15 and ran it; report archived
at `echolot-app/reports/CPH2747-app-run1.json`. Results:
- **net.captive_portal works** — Android's NetworkMonitor logic reproduced: default + wifi both
returned HTTP **204** on the HTTPS *and* HTTP generate_204 probes → `validated`; **cellular
returned neither (-1/-1) → `no_internet`**. The per-network split immediately surfaces an
asymmetry the OS hides (wifi validated, cellular can't reach the checks at all).
- **Shizuku tier degrades correctly**: `binder_alive:false` → test UNSUPPORTED, `tiers.shizuku:false`
(Shizuku isn't running on the phone). The dual-path *executing* path still needs an on-device
test with Shizuku started (expect UserService on this OnePlus).
- 5 tests now: link.snapshot ok, icmp.ping4 ok, icmp.ping6 failed, net.captive_portal ok,
link.ip_monitor(shizuku) unsupported → overall YELLOW via the IPv6 finding.
Also fixed this session: the beacon app itself caused the "wireless debugging connected"
notification spam (it re-resolved adbd's own mDNS advertisement, making adbd re-arm each time);
now resolves once per service instance and the heartbeat re-POSTs the cached port only.
## App: dns.canary verified against the live server (2026-08-01)
Built the client half of the canary-DNS measurement and verified it on the OnePlus against the
deployed fmr zone (`echolot-app/reports/CPH2747-app-run2-dns.json`):
- All four spec-frozen reference records matched byte-for-byte through the network's own resolver
(ttl-5→192.0.2.5, ttl-60→192.0.2.60, ttl-3600→192.0.2.36, ttl-86400→192.0.2.86) → nothing on
this path rewrites DNS answers (`dns.answer_integrity` in the green case).
- The un-cacheable nonce name `1006ad16.adhoc.c.echo-lot.app` resolved to 192.0.2.21 →
`reached_authoritative: true`, proving the query actually reached the canary server rather than
being answered from a cache or an interceptor.
Findings wired: `dns.answer_rewritten` (high) when a reference mismatches, and
`dns.authoritative_unreachable` (medium) when the nonce isn't answered by the canary server.
This closes the first full client↔server measurement loop: the Kotlin app measures against the Go
server's canary zone on real hardware. 6 tests now run per measurement.
## App: Shizuku shell tier VERIFIED on-device — the app is feature-complete for v1 tiers (2026-08-01)
Ran the app on the OnePlus with Shizuku running (`echolot-app/reports/CPH2747-app-run3-shizuku.json`):
- **`tiers: {app:true, shizuku:true}`**, shizuku test **ok**, **commands_ok 7/7**,
**`exec_path: UserService`** (the OnePlus binds it — matches the prober; the newProcess fallback
stays for the Lenovo).
- Evidence is the real privileged material the production parsers need: live ARP/NDP neighbor
table, per-table IPv6 routes, `ip addr`, an actual `[NEIGH]` netlink event from `ip monitor`,
IpClient DHCP logs incl. APF capabilities, and the wifi state dump — all as shell(2000).
Both privilege tiers now work end to end in the production app on real hardware, alongside the
canary-DNS loop against the live server. 6 tests/run: link.snapshot, icmp.ping4, icmp.ping6,
net.captive_portal, dns.canary, link.ip_monitor(shizuku).
## App: nat.stun_5780 — NAT behavior discovery verified against the live server (2026-08-01)
Client-side RFC 5389/5780 STUN (hand-rolled, stdlib only) exercising the server's `stun-5780`
capability. Three binding requests from ONE socket: primary, the server's OTHER-ADDRESS
(alternate IP), and CHANGE-REQUEST(port). Verified on the OnePlus
(`echolot-app/reports/CPH2747-app-run4-stun.json`):
- local `10.13.102.124` → mapped `178.191.120.247:53259`, `behind_nat: true`
- `other_address 89.185.109.151:3479` — the server's second IP answered, so RFC 5780 works
end to end (client ↔ our own STUN implementation)
- **mapping: endpoint-independent** (same external port toward a different destination → P2P
friendly); **filtering: address/port-dependent** (no reply to CHANGE-REQUEST → unsolicited
inbound is dropped). Classic full-cone-mapping + port-restricted-filtering NAT.
Finding wired: `nat.symmetric` (medium) when mapping is address/port-dependent.
Two bugs caught by running it for real: port preservation was misread as "no NAT" (now compares
ADDRESSES), and an unbound socket reports the wildcard as its local address (now resolved via a
throwaway connected socket). 7 tests/run.
## App: autorun mode + IPv6 severity rework (2026-08-01)
**IPv6 is no longer treated as a defect just for being absent.** The finding now depends on
whether the network actually provisioned IPv6 (a global v6 address or a `::/0` route):
- not provisioned → `ipv6.not_offered`, severity **INFO → green**. Most networks are still
IPv4-only and that is not a fault.
- provisioned but ICMPv6 fails → `ipv6.broken`, severity **MEDIUM → yellow**. Half-configured
IPv6 is worse than none (Happy-Eyeballs stalls). Verified on the OnePlus: our LAN advertises a
v6 default route with no working path, so it correctly reports `ipv6.broken`.
**Autorun mode** — one adb command runs a full measurement unattended and collects the result
without any UI tapping or adb round-trip:
```
adb shell am start -n app.echo_lot.app/.MainActivity --ez autorun true
curl http://<fmr>/reports # list
curl http://<fmr>/report/<name> # fetch
```
The app runs the suite, POSTs the MeasurementDocument to the collection endpoint (receiver.py
gained `POST /report`, `GET /reports`, `GET /report/<name>`), shows the result for 3 s, then
finishes itself — leaving the device as it was found. On upload failure it stays open so the
error is visible. Grant permissions once via `adb shell pm grant app.echo_lot.app
android.permission.ACCESS_FINE_LOCATION` so nothing blocks on a dialog.
## App: router identification, brand icons, DEV build variant (2026-08-01)
**`link.ra_source` — who is advertising IPv6 here, and what box is it?** New app-tier probe
(registry addition). Identification chain, each step recorded as evidence so nothing is guessed:
1. RA source = next-hop of the `::/0` route per network (a `fe80::` link-local).
2. **MAC recovered from the modified-EUI-64 link-local** (strip `ff:fe`, flip the U/L bit) —
e.g. `fe80::7a9a:18ff:fe54:b8f9` → `78:9a:18:54:b8:f9`. RFC 7217/privacy addresses don't encode
a MAC and are reported as such rather than guessed.
3. Vendor via a curated OUI table (`Oui.kt` — SOHO/router vendors; unknown OUIs are printed
verbatim). Locally-administered (randomized) MACs are flagged.
4. **UPnP/SSDP M-SEARCH** → the gateway's `SERVER:` banner + device-description XML gives
manufacturer / model / friendly name. This is what usually names the exact box.
5. Reverse DNS for both gateways.
All SSDP responders are recorded (not just the gateway) so a rogue RA sender that isn't the
gateway can still be matched — and the MAC travels with every identity source, which is the hook
for the future LLDP / mDNS cross-matching.
UI: a "Router / IPv6 advertiser" panel above the network list, leading with the identified
vendor/model.
**Icons + DEV variant.** The branding adaptive icon is now the app icon: `icon-adaptive-*.svg`
converted to Android vector drawables (SVG transform baked in, gradient background, monochrome
layer for themed icons) plus PNG mipmaps for legacy launchers. The **debug build is a separate
app**: `applicationIdSuffix .dev`, label "Echolot DEV", and a DEV-badged icon (layer-list =
production foreground + generated amber DEV ribbon) so it is unmistakable next to a real install
and both can be installed side by side.
NOTE for tooling: the dev package is `app.echo_lot.app.dev`, activity `app.echo_lot.app.MainActivity`.
### link.ra_source verified on-device — it named the actual router (2026-08-01)
Run archived at `echolot-app/reports/CPH2747-app-run5-router-id.json`. On the wifi network the
probe identified the RA sender completely, from an unprivileged app:
- RA source `fe80::7a9a:18ff:fe54:b8f9` → **MAC 78:9A:18:54:B8:F9 recovered via EUI-64**
(matches the Shizuku neighbor table exactly) → vendor **MikroTik** by OUI
- IPv4 gateway `10.13.102.1`, reverse DNS `router.hudelist.local`
- UPnP: `RouterOS/7.23.2 UPnP/1.0 MikroTik` → manufacturer MikroTik, model Router OS,
friendly name "MikroTik Router"
So the box advertising this LAN's broken IPv6 RA is a **MikroTik running RouterOS 7.23.2**, named
by two independent methods (OUI from the address itself + UPnP device description) that corroborate.
On cellular the carrier's RA source is an RFC 7217 privacy address and is correctly reported as
"not EUI-64" rather than guessed.
The SSDP sweep also inventoried the LAN (Synology DS1522+ DSM 7.3, a Sky ES160 gateway) — the
raw material for the planned LLDP/mDNS cross-matching by MAC.
Fixes from this run: added the confirmed MikroTik OUI 78:9A:18 (+ other RouterBOARD ranges), and
an elvis-operator bug that printed "no UPnP response" even when UPnP data was present.
### App UX: progress bar + ETA, cancel, and edge-to-edge insets (2026-08-01)
- **Progress + ETA**: `Probe.estimatedMs` (per-probe, from measured on-device durations — the
timeout-bound probes dominate: icmp.ping6 ~7s on a v4-only net, captive-portal ~9s, SSDP ~7s,
STUN ~6s) drives a determinate bar plus "test N of M · ~Xs left". The Shizuku battery is counted
in the total so the bar covers the whole run.
- **Cancel**: stops an in-flight run and shows what was measured so far, assembled into a normal
document (findings + verdict over the partial set). Deliberately **does not upload** — a partial
run is for the person looking at the screen, not for the record.
- **Insets/cutout**: Android 15 draws edge-to-edge by default, so the title was running under the
status-bar clock and the camera cutout. The root column now uses `safeDrawingPadding()`, which
covers status bar, navigation bar and display cutout.
### Shell-tier readiness shown before a run (2026-08-01)
`ShizukuAvailability` distinguishes four states and the UI only speaks when it is actionable:
- **NOT_INSTALLED → says nothing.** Users who don't use Shizuku are never nagged.
- **INSTALLED_NOT_RUNNING → amber banner** "Shizuku is installed but not running — start it to
include shell-tier tests". This is the case worth reminding about: the user has it, but a
stopped service silently costs them the whole shell tier.
- NEEDS_PERMISSION → "running but not authorised, it will ask on first use".
- READY → green "shell-tier tests will run".
Detection is listener-based (`addBinderReceivedListenerSticky` + binder-dead), because
`pingBinder()` is only truthful once ShizukuProvider has delivered the binder — a one-shot poll at
launch would show a false "not running". Installed-vs-not needs the `<queries>` package-visibility
entry on Android 11+. Verified on-device: with shizuku_server stopped the banner appears correctly.
### Shizuku banner is actionable; progress + cancel verified on-device (2026-08-01)
Tapping the shell-tier banner now does the right thing per state: **installed-but-stopped** →
deep-links into the Shizuku app (a third-party app *cannot* start Shizuku itself; the wireless-
debugging pairing flow is privileged and lives in that app, so taking the user there in one tap is
the best available), **running-but-unauthorised** → fires the Shizuku permission request directly.
The hint line states which action the tap performs.
Verified on-device in one screenshot: progress bar at "test 4 of 8 · icmp.ping6 · ~33s left",
Cancel button beside the disabled Run button, title clear of the status bar/cutout, and the banner
having live-switched from "installed but not running" to "running but not authorised" via the
binder listener when Shizuku was started mid-session.
### Why the Shizuku banner can't start wireless debugging directly (verified, 2026-08-01)
Checked against Shizuku 13.6's own manifest (pulled the APK, `aapt2 dump xmltree`): the
wireless-debugging entry points — `moe.shizuku.manager.adb.AdbPairingTutorialActivity`,
`moe.shizuku.manager.adb.AdbPairingService`, `moe.shizuku.manager.starter.StarterActivity` —
declare **no intent filters**, so they are not exported and a third-party app cannot launch them.
`MainActivity` is the only reachable entry and answers MAIN/LAUNCHER only (no deep link), which is
why the handoff lands on the screen whose primary action is the root start.
Best available behavior, now implemented: the banner still opens Shizuku, but the hint names the
exact steps there ("Pairing", then "Start"), and a second tap target opens **Developer options**
(`Settings.ACTION_APPLICATION_DEVELOPMENT_SETTINGS` — public and exported) since Wireless
debugging must be enabled first for Shizuku's wireless start to work at all.
### Downstream measurements: asymmetric grants, DF-mode big_send (server-v0.4.0 … v0.4.2, 2026-08-01)
The client can measure a round trip and the largest packet it can *send*. It cannot measure the
largest packet it can *receive*, or downstream-only loss — those need the server to push, which is
exactly what §3.4 gates behind an asymmetric grant. Implemented and verified live from the PC:
- **`session.Grant`** — created per action, bound at creation to the session's *observed*
data-plane source (no grant without a verified destination), clamped to server limits, with a
byte budget, an average-rate ceiling and an expiry. Unit-tested for each of those refusals.
- **`downtrain`** — N packets of size S every I µs; the client derives downstream loss,
reordering and inter-arrival spacing.
- **`big_send`** — one datagram per requested size. **DF is on by default**, so the largest size
that arrives *is* the downstream path MTU. Without DF the kernel fragments and the result only
says whether fragments get through — a different fact, and the reason the schema has both
`mtu.pmtud_down` and `mtu.frag_delivery`. Sizes above the server's own egress MTU (from the
startup self-test) are refused up front and reported as `max_df_bytes`, so an absence caused by
our kernel is never read as a limit of the client's path.
Live from the PC against fmr: downstream path MTU **1500** (1472 payload, DF), fragmented delivery
up to **4000**, downstream train **100/100, 0 % loss, 0 reordered**, inter-arrival 3.3 ms for a
3000 µs send interval.
#### Two bugs this shook out, both invisible in a single-homed lab
1. **Granted sends went out from the wrong local address** (fixed in server-v0.4.2). fmr binds two
IPv4 addresses; `connFor` returned whichever socket of the right family came first in the bind
list. A train for a session established on `.150` left from `.151` and every packet was dropped
by the client's NAT, which has no mapping for that pair. tcpdump showed all 50 leaving, the
client saw none — reported as *100 % downstream loss*, a confident measurement of something
that never happened. Sessions now record which of our own bound addresses received their
traffic and granted sends go back through that socket; `connfor_test.go` pins both that and the
family fallback.
2. **A second `ProbeSession` on one server session is silently dead.** Sequence numbers restart at
zero client-side while the server's anti-replay window keeps counting, so every packet is
discarded as a replay — and because the server then never records the new source, the grant
still targets the closed socket. `ServerMeasurement` now uses one ProbeSession for the whole
run; `ProbeSession`'s doc comment states the constraint.
### Run archive, anonymizer and uploads (2026-08-01)
Three pieces, deliberately separate:
- **`core-archive`** — one JSON file per run plus an index entry, in a plain directory the user can
inspect or delete with a file manager. Retention (max runs / max age / max total bytes) is
enforced on every save rather than by a sweeper. `enabled` is a separate flag from the three
ceilings because "no limits" and "keep nothing" are opposite intentions; collapsing them onto
all-zeros is how a user who turns the caps off ends up with an empty history. 13 tests.
- **`core-privacy`** — the schema §8 anonymizer, three levels. `full` (your own server) changes
nothing; `balanced` pseudonymizes SSIDs/hostnames, keeps the OUI half of a MAC and the /16 of a
public IP, keeps RFC1918 verbatim (it describes topology, not a person), and *drops* neighbour
inventories (SSDP/ARP/scan results) rather than mangling them; `strict` keeps only metrics,
statuses and finding codes. Pseudonyms are consistent within a document and — by default — not
across documents, so an upload endpoint cannot link a device's runs; a stable salt is opt-in for
people diffing their own history. Classification is one readable table, not annotations spread
across modules. 14 tests, each pinning a property someone's privacy depends on.
- **Server-side upload policy** — `off | anonymous | account`, plus max size, retention days, max
runs per device, and the *least* anonymization accepted. The profile advertises all of it so the
app presents the choice honestly instead of discovering the rules by being rejected. `account`
refuses today rather than falling back to anonymous: picking the strict setting before OIDC
lands must not silently mean the loose one.
**The archive holds the unredacted document; redaction happens on the way out, per upload.** The
local archive is the user's own data on their own device, and redacting it would destroy exactly
the detail that makes a week-old run worth keeping.
App-side: settings screen (archive limits, privacy level with a plain-language description of what
each keeps, auto-upload off by default, server URL/pin/credential), history screen showing whether
each run left the device, and a **preview of the exact bytes an upload would send** — an anonymizer
the user cannot inspect is only a promise.
Live round trip against fmr: uploaded a run, listed it, fetched it back and asserted the SSID, the
SSDP neighbour name and the free-text note are absent from what the server stores while the
finding code and the metrics survive, then deleted it.
### Still open
- `mtu.pmtud_up` (DF + errqueue), `frag_send`, `throughput`, TRAIN_REPORT retrieval.
- Enrollment UI in the app (server URL/pin/credential are typed in by hand today).
- Accounts/OIDC on the server, which is what `uploads=account` is waiting for.
- Nothing in this entry has been exercised on a phone yet — all of it was verified from the PC
against the live server. On-device verification is the next step.
### SemVer compatibility windows between app and server (server-v0.5.0 … v0.5.2, 2026-08-01)
Both artifacts are SemVer, and each now declares — and enforces — which peer versions it will talk
to. Spec: `docs/probe-protocol.md` §8.
**Two axes, deliberately not conflated.** Release versions are a *proxy* for what actually has to
match, so the real thing is checked first:
- `protocol_version` — **can** these builds talk. Advertised in the profile; a peer in a different
breaking series is refused whatever its release version says. Below 1.0.0 the **minor** is the
breaking axis (SemVer §4).
- release-version window — **may** they, per policy. `[min, max)`, min inclusive, max exclusive,
because the useful bound is always "the version that broke it".
Bounds sit at breaking boundaries, not at releases, so shipping a patch never requires editing a
range. The app requires server `>= 0.4.2` for a stated reason, not caution: earlier multi-homed
servers mis-addressed granted sends and the client measured 100 % downstream loss that never
happened. Operators override the server side with `ECHOLOT_MIN_APP_VERSION` /
`ECHOLOT_MAX_APP_VERSION`; a malformed bound is fatal at startup rather than ignored, so a typo
cannot silently disable a restriction.
Three rules that shaped the implementation:
1. **`GET /v1/profile` is never gated.** It is where a refused client learns which version it needs;
gating it leaves the user with a network error instead of an answer.
2. **An unparseable or absent version is `unknown`, and is allowed.** Dev builds report `dev`, and a
client too old to send the header cannot be identified anyway.
3. **Refusal is 426 with a body naming both versions and the window**, surfaced client-side as a
distinct `VersionRefused` rather than folded into "network error".
The app's `versionCode` is now derived from its SemVer (`major*1e6 + minor*1e4 + patch*10`) instead
of being a second number to remember.
Verified live against fmr (`LiveCompatTest`): profile advertises the window and stays readable for a
refused version; 0.1.0 and 99.0.0 are both refused with actionable messages; 0.2.0 and a missing
header are both served.
One user-visible bug caught in the process: Go's JSON encoder HTML-escapes `<`, `>` and `&` by
default, so the refusal reached the client as `needs \u003e= 0.2.0`. Disabled at the encoder (this
is an API, not a page), and the client now *parses* the error field instead of pattern-matching it,
so it survives whatever a future encoder decides to escape.
### Enrollment: the server mints the bootstrap link (server-v0.5.3 … v0.5.4, 2026-08-01)
Until now a device was configured by hand-typing a control URL, a base64 SPKI pin and a
credential. That is the step that goes wrong, and it goes wrong quietly: a pin off by one
character does not fail loudly, it just never matches, and surfaces days later as an inscrutable
TLS error.
`POST /admin/enroll-tokens` now returns the whole §2.1 bootstrap link alongside the token, because
the server is the only party holding all three parts at once. The app takes it from a paste or an
`echolot://enroll` deep link (so a QR scan configures a server in one action) and writes URL, pin
and credential **together or not at all** — a half-applied server fails later, somewhere else,
with an error pointing at the wrong thing.
The control URL comes from `ECHOLOT_PUBLIC_URL` (set on fmr to `https://fmr-1.echo-lot.app:8443`),
falling back to the first control listen address; a wildcard bind warns rather than emitting a
link to `0.0.0.0`.
**The encoding trap, which is the whole reason this is tested across both languages.** The pin is
base64, so it contains `+`, `/` and `=` — each of which means something else in a query string. An
unencoded `+` decodes to a space, leaving the pin wrong by exactly one character. Base64 has no
spaces, so the parser restores them; that cannot damage a correctly-encoded pin and it rescues
every hand-assembled link. `LiveEnrollmentTest` redeems a link the *server* produced, which is the
only way to catch a disagreement between the Go assembler and the Kotlin parser — a unit test on
either side alone cannot see it. It also asserts the token is refused the second time.
Also fixed a spec divergence found while reading §2.1: the spec names the field
`device_credential`, the first implementation shipped `credential`. The server now sends both and
the client prefers the spec's; the alias goes once nothing reads it.
Two process notes from this round:
- An edit to the admin handler silently failed to apply and the endpoint kept returning just the
token. Caught by deploying and *looking at the response*, not by trusting a green build.
- The live suite is now six tests (`LiveServerTest`, `LiveMeasurement`, `LiveGranted`,
`LiveUpload`, `LiveCompat`, `LiveEnrollment`), all green against fmr from the PC with no device.
### Directional loss: which way is the packet loss? (2026-08-01)
A round trip can only report that *something* was lost somewhere, which is the least useful form
of the answer — "3 % loss" sends an engineer looking in both directions at once. The server
already records every packet it received per sequence number (§6), so the two cases are actually
distinguishable, and `train.udp_updown` now reports them separately:
- sent, never seen by the server → **upstream** loss
- seen by the server, reply never arrived → **downstream** loss
Findings name the direction and say what is *not* implicated, which is half the value:
`connectivity.loss_upstream` ("the return path is not implicated: replies came back for everything
that arrived"), `connectivity.loss_downstream`, `nat.udp_unreachable_upstream`.
Two things the implementation gets deliberately right:
- **Downstream loss is measured against what reached the server**, not against what was sent.
Using "sent" as the denominator counts every upstream loss a second time and overstates the
return path. Pinned by a test with loss in both directions at once.
- **Per-direction jitter without synchronised clocks.** Absolute one-way delay would need clock
sync and we deliberately have none (the two-clock rule). But `server_rx client_tx` carries a
constant unknown offset, and differencing successive samples cancels it — so RFC 3393 one-way
delay variation *is* honestly attributable to a direction even though latency is not. A test
pins that a 10-second clock offset changes nothing.
Correlation is by **wire sequence number**, which is not the loop index: the counter is shared
with every other packet type on the session, so "the nth echo" is not "sequence n". `ProbeSession`
now exposes `lastSeq`, including for a probe that was lost — a lost packet still has a sequence
number, and that number is exactly what tells you which way it was lost.
Live against fmr: 20/20 both ways, and jitter of **0.08 ms upstream vs 0.85 ms downstream** — a
tenfold asymmetry that a round-trip measurement cannot see at all.
10 unit tests on the arithmetic (a wrong denominator here does not crash, it produces a plausible
number pointing at the wrong half of the network) plus the live correlation check.
### frag_send: crafted IP fragments, so *ordering* is testable (server-v0.6.0, 2026-08-01)
`big_send` with `df=false` answers one question — do fragments get through. It cannot answer the
more interesting one, because the kernel always emits fragments in order, first one first.
The classic middlebox fault is exactly about that ordering. Only the **first** fragment carries the
UDP header, and therefore the ports; a stateful firewall or NAT that has not seen it has no flow to
match the rest against, and many simply drop them. That is invisible to every in-order test, and in
the field it looks like "large DNS answers fail on this network" or "the tunnel breaks when the MTU
drops" — it works until the network reorders, then fails intermittently, which is the hardest kind
of fault to chase.
So the server builds the fragments itself (raw socket, `IP_HDRINCL`) and controls their order:
`in_order` (baseline), `reversed` (last fragment first), `first_last` (first fragment held back
250 ms). The datagram is assembled and **signed whole** before being cut up, so what the client
reassembles is indistinguishable from an ordinary packet — otherwise the test would be measuring
our sender rather than the path. New test type `mtu.frag_ordering`; findings
`mtu.fragments_blocked` and `mtu.fragment_reorder_sensitive`.
Two details that would otherwise produce confidently wrong answers:
- **The UDP checksum is computed, not left zero.** Zero is legal in IPv4 and would be less code,
but zero-checksum datagrams are dropped by some middleboxes — and that drop would be recorded as
a fragmentation failure, which is the wrong conclusion entirely.
- **Fragment offsets are in 8-byte units**, so non-final fragments are rounded down to a multiple
of 8. A 100-byte fragment is not an error; it is a datagram no host will ever reassemble.
`frag-send` is advertised only when a raw socket can actually be opened — checked by opening one,
because a permission model has more ways to say no (userns, seccomp, LSM) than a capability bit has
to say yes. fmr runs as root with `cap_net_raw` in its bounding set, so it is available there.
Fragment ordering runs only after `mtu.frag_delivery` shows fragments arrive at all; otherwise the
three orderings would each report "not delivered" and read as three faults instead of one.
The header arithmetic is unit-tested (reassembly coverage with no gaps or double-delivery, MF
flags, shared IP ID, 8-byte offsets, checksum verification over odd and even lengths). Because the
code is `//go:build linux`, the tests are **cross-compiled and run on fmr** — there is no Go
toolchain there, so `go test -c` plus scp is the loop.
Live against fmr: 4 fragments per burst, and all three orderings reassembled — a healthy path, and
the baseline against which a mobile network will be interesting.
### Testing state (2026-08-01)
Six live tests against fmr, all green, no device involved: `LiveServerTest`, `LiveMeasurement`,
`LiveGranted`, `LiveDownstream`, `LiveUpload`, `LiveCompat`, `LiveEnrollment`. Plus 74 client unit
tests and the full Go suite. Everything in the last several entries is verified from the PC; the
app's UI (settings, history, deep-link enrollment) and `mtu.pmtud_up` remain device-only.
### throughput: a rate, plus the qualifier that makes it a measurement (server-v0.6.1 … v0.6.2)
A throughput test reports the *smallest* limit on the path — and the sender's own ceiling is one of
the candidates. If the server is asked for 50 Mbps and 50 Mbps arrives, the network was never the
constraint and "50 Mbps" says nothing about it. So `perf.throughput_udp` always carries
`limited_by` (duration | budget | rate | send_error) and `measures_network`, and a finding is
raised only when the path is actually implicated. The live run against fmr reports 20 Mbit/s with
`measures_network: false`, which is the correct and useful answer.
Loss is computed against the **sender's own count**, fetched from the observations API, not against
the requested rate. A receiver alone cannot tell "the network dropped it" from "the sender never
sent it", and guessing turns a healthy server-side limit into a phantom network fault. The server
keeps one summary per action rather than per-packet records — a ten-second run at 50 Mbps is half a
million packets, and a struct each would turn a measurement into memory exhaustion.
Sending is **paced**, on an absolute schedule. Unpaced would measure the server's NIC and the first
queue it meets, then collapse into loss that reads as a network fault; sleep-per-packet would
accumulate scheduler error and drift the rate down over a ten-second run.
Throughput gets its own grant budget sized from the request, so every *other* action stays bounded
at 8 MiB. When the byte cap binds before the clock does, the **duration is shortened and reported**
rather than the run being truncated: promising thirty seconds and delivering twenty-one is the same
information with a surprise attached, and it keeps "the clock ended the run" as the normal case —
the only case where the rate is a clean property of the path. That behaviour came out of a test
that failed honestly (30 s at 100 Mbps needs 375 MB against a 256 MB cap).
It is **opt-in** in the run config, default off. A 5-second run at 50 Mbps moves ~30 MB; on a
metered mobile connection that is the user's money, and a tool that spends it without being asked
is not one people keep installed.
#### The bug the live test found
The first live run delivered 104 packets and stopped after 50 ms. The grant's rate check exempted
the first 50 ms entirely, meaning to be lenient at startup — the effect was the opposite. A sender
could dump an unbounded burst into that free window, and the instant the check switched on it
compared those bytes against 50 ms worth of allowance and refused everything until real time caught
up. **Every short test passed** (downtrain sends 50 packets, big_send seven); every sustained send
died fifty milliseconds in.
Replaced with a token bucket (`allowance = burst + rate × elapsed`), which is smooth from t=0.
The burst is 100 ms of the allowed rate, floored at one ordinary datagram — deliberately one, since
at 8 kbps a 64 KB floor is sixty-four seconds' worth, exactly the instant dump the ceiling exists to
prevent. The pre-existing rate test caught that when I first tried the generous floor, and it was
right to. Second half of the same bug: callers treated *any* refusal as terminal, so `TryAllow` now
says why — a sender paces through a transient "too fast just now" and still stops dead on a spent
budget or an expired grant. Both halves are pinned by regression tests.
### Findings registry (2026-08-01)
Closes open item 1 of measurement-schema.md §9. A finding code is the stable, machine-readable half
of a result — what a dashboard groups by and what someone greps a year of archived runs for — and
that only holds if a code means exactly one thing forever. Ad-hoc string literals at fifteen call
sites cannot promise that, and by the time the registry was written the failure had already
happened.
**Two emitters had independently produced `connectivity.downstream_loss` and
`connectivity.loss_downstream` for the same claim**, and nothing anywhere objected. Anyone
aggregating either one would have silently seen half their data. Merged into
`connectivity.loss_downstream`, paired with `loss_upstream` so the two directions read as a set.
**Two codes were also renamed out of `nat.*`.** `nat.udp_unreachable` is not about NAT — it means
no replies came back — but the prefix determines the category, and the category determines which
verdict light the finding rolls up into (§7.3). A `nat.*` code landing under *connectivity* is not
a naming quibble; it changes which light turns red. Cheap to fix now, a breaking change later.
Codes are now declared as typed `FindingSpec`s carrying their category and default severity, and
emitters reference the spec instead of retyping the string — so a typo is a compile error and two
call sites cannot disagree about a finding's category.
`docs/findings-registry.md` is the contract, and a test reads it: it fails when the document and
the registry have codes the other lacks, or when a severity differs. Documentation that drifts from
its implementation is worse than none, because it still looks authoritative. The check scopes
itself to table rows, so the prose can keep explaining which codes were retired and why.
Six tests: uniqueness, declared-vs-listed, prefix↔category agreement, naming convention, a
word-order-anagram check (the shape the duplication actually took), and the document agreement.
### A real privacy leak, found by starting on the machine-readable schema (2026-08-01)
The intent was `measurement.schema.json` (§8's promised companion). The first step — checking
whether the anonymizer actually covers the fields the schema declares as sensitive — found that it
did not, so that became the work.
**At the `balanced` level, five identifying values were being uploaded verbatim:**
| value | field | why it matters |
|---|---|---|
| `2001:…::150` | `networks[].link.addresses[].addr` | the device's own global IPv6 address — a strong, geolocatable device identifier |
| `2a02:…::1` | `networks[].link.routes[].gateway` | identifies the ISP allocation |
| `203.0.113.77` | `networks[].link.dns.servers[]` | the configured resolver |
| `nas.example.lan` | `private_dns_hostname` | an internal hostname |
| `example.lan` | `search_domains[]` | the internal domain |
The settings screen describes that level as pseudonymizing addresses. It was not.
**Root cause:** classification keyed on field *names*, and the schema's actual names (`addr`,
`gateway`, `dst`, `servers`, `search_domains`, `private_dns_hostname`) had never been added to the
table. Not a subtle bug — just an unfalsifiable design. The existing tests all passed, because each
one checked a field somebody had remembered to write a case for.
**Two fixes, one of them structural:**
1. The missing names were added.
2. More importantly, a **shape-based backstop**: when a field name is unrecognised, the *value* is
inspected, and anything shaped like an IPv4/IPv6 address or a MAC is treated as one. A name
table can only protect fields someone thought of, which is precisely the wrong property for a
privacy control. Hostnames are deliberately *not* inferred by shape — `train.udp_updown` is
indistinguishable from a domain, and mangling a test type would corrupt the document to protect
nothing.
`LeakTest` is the new guard and is written to fail for fields nobody has considered: it plants
identifying values wherever one can actually occur and asserts none survive, rather than checking
a list of known cases. It also pins that RFC1918 addresses still come through readable, so the
test cannot pass by over-redacting everything.
Route prefixes and the unspecified address needed care in the transform: `0.0.0.0/0` and `::/0`
must stay themselves, or a routing table becomes unreadable for no privacy gain.
**Still outstanding:** `measurement.schema.json` itself. Worth noting what this episode implies for
it — much of a document's payload lives in `evidence`/`metrics`/`params`, which are per-test-type
`JsonObject` by design and therefore *outside* any schema. A schema-driven anonymizer would have
less coverage there than the name-plus-shape one now does, so the schema should be built for
validation and external tooling, not as a replacement for the classifier.
### ULA prefixes are pseudonymized whole (2026-08-01)
Spotted in a real uploaded run from the phone: the server had
`fda1:3fb1:ff92:6696::2662` for a DNS server. The general IPv6 path preserves the leading two
groups (deliberately — for a global address that keeps the ISP allocation, which is the
diagnostically useful part), and for a ULA that passed through **32 of the 40 random bits** of the
global ID.
ULA looks like the v6 equivalent of RFC1918 and the instinct is to treat it the same. That
reasoning does not carry over, and the difference is the whole point: an RFC1918 prefix is shared
by millions of networks and identifies none of them, while a ULA global ID is random and unique to
one network by construction (RFC 4193). The prefix *is* the identifier — it is a network
fingerprint that was surviving redaction.
Now pseudonymized as a unit, so two addresses on the same ULA subnet still land on the same
pseudonymous prefix: "these hosts are on one network" survives, "this is *that* network" does not.
Three tests, one of which uses the exact value observed on the wire.
Worth recording as a reasoning trap: I had originally raised this as "ULA should probably be kept
verbatim, like RFC1918, for consistency". The surface analogy pointed the wrong way, and the
correct answer was the opposite.
### Registry adopted everywhere; v6 findings renamed; Back works (2026-08-01)
The findings registry was only adopted in `core-engine`. The app module still emitted seven codes
as raw strings, so the registry test passed while codes existed outside it — including
`ipv6.broken`, which fired on a real network and was in no registry at all.
All seven now reference registry entries for code, category and severity, so those three cannot
disagree at a call site. A grep for `code = "…"` across the app, engine and probe modules returns
nothing.
**`ipv6.*` → `v6.*`.** The third instance of rule 1: they declared `Category.IPV6` while the prefix
map only knows `v6`, so `TestType.category("ipv6.broken")` fell through to *connectivity* and the
finding rolled up under the wrong verdict light. The test-type registry already used `v6.`.
Two severities reconciled while merging:
- `connectivity.captive_portal` is **medium**, not high. The registry had guessed high; the probe
that emits it had always said medium, and the probe was the considered value — a captive portal
on hotel wifi is what should be there, and logging in clears it. `connectivity.no_internet` is
the high one, because nothing the user does locally fixes that.
- `v6.not_offered` is **info, and the registry says it must stay info**. Most networks still do not
offer IPv6 and that is not a fault; a warning here lights a yellow verdict on a healthy network,
which teaches people to ignore the light.
Also: a `BackHandler` now returns from Settings/History to the run screen. The screen was a plain
state variable with nothing connecting it to the back stack, so the system Back gesture left the
app entirely. Enabled only when there is somewhere to go back to, so Back still exits from the run
screen.
### Upstream throughput (server-v0.6.3, 2026-08-01)
The mirror of the downstream case: the client generates the traffic and the server counts it. No
grant is involved — the client is sending its own packets, so there is nothing to amplify — but it
does need the server's tally, because **only the far end knows how much arrived**. Without that
number a sender measures how fast it can *transmit*, which is usually just the speed of the local
NIC and is a different question from the one being asked.
`TYPE_THROUGHPUT_UP` (0x0F) is counted and deliberately **never answered**: a reply would double
the traffic and drag the return path into a measurement that is specifically about the outbound
one.
The tally is a counter, not a list, and short-circuits **before** the observation log. A
five-second run at 20 Mbps is around ten thousand packets; one struct each would turn a
measurement into an allocation storm on a shared server, and nothing needs the per-packet detail
since the client holds the send-side record. The gap between the two counts is the loss.
`direction=up` on the throughput action sends nothing — it zeroes the counter, so a second run in
one session measures itself rather than inheriting the first one's packets. The live test asserts
`received <= sent`, which is what catches a counter that was never reset.
Live against fmr: **3125 sent, 3125 counted, 0 % loss, 10.0 Mbit/s** at a 10 Mbit/s request, with
`measures_network: false` — correct, since what arrived matched what was offered, so the path was
never the constraint.
### Raw shell dumps leaked the whole LAN (2026-08-01)
Found by running the Shizuku shell tier for the first time. The tier works — `tiers.shizuku: true`,
`exec_path: UserService` (so the UserService binds on the OnePlus, as recorded), `runs_as
shell(2000)`, 7/7 commands — and the run promptly uploaded **every MAC address on the local
network** to fmr at the `balanced` level: router, phones, whatever else was on the wifi. Fourteen
of them.
The probes embed raw command output verbatim (`ip neigh`, `ip route`, `id`), which is genuinely
good evidence and also a complete household device inventory. The anonymizer could not see it:
classification is by field name and by whole-value shape, and `ip_neigh` is one long string that is
itself neither a MAC nor an address. measurement-schema.md §9 item 2 had flagged raw dumps as "hard
to anonymize" and proposed dropping them from exports; nothing enforced either.
**Scrubbing beats dropping.** Identifiers inside any unclassified string are now replaced in place,
using the same pseudonyms as everywhere else — so a MAC that appears both in a parsed field and in
a raw dump still reads as one device. The dump stays readable and auditable: you can still see the
neighbour table's shape, the host count, RFC1918 addresses and vendor prefixes. Dropping the
evidence would have protected the same data while destroying the reason for collecting it.
Two implementation notes worth keeping:
- **One pass, not three.** Sequential passes re-process their own output: once a MAC became
`78:9a:18:xx:yy:zz`, the IPv6 pattern matched it — six hex groups separated by colons *is* an
address — and destroyed the vendor prefix the MAC rule had just preserved. Ordered alternation
resolves each position once, MAC first.
- The patterns are conservative on purpose. A missed address gets caught by another rule or not at
all; an over-eager one mangles timestamps and version strings, corrupting evidence to protect
nothing.
`RealDocumentTest` runs the anonymizer over a captured run when `ECHOLOT_REAL_RUN` points at one,
and fails on any MAC that survives. It self-skips otherwise, so no one's network is committed to the
repo. Against the actual leaked document: **14 MACs in, 0 surviving.**
Also fixed: the Settings *Preview what an upload would send* button did nothing. It read
`UiState.history`, which is empty until the History screen has been opened — the same root cause as
the "0 run(s)" count. It now reads the archive directly, and says so when there is nothing to
preview rather than silently ignoring the tap.
### Security: the admin listener was publicly exposed for ~15 minutes (2026-08-01)
Moving the admin listener to `[::2]:443` for the UI exposed `/admin/enroll-tokens` and
`/admin/selftest` to the internet **with no authentication**. Anyone who could reach
`fmr.echo-lot.app` could mint enrolment tokens.
The listener was designed localhost-only — its own flag help says *"keep localhost"* — and that
assumption travelled with it when the address changed. The compounding error: `checkAdminExposure`,
added the same day, verifies **encryption** and says nothing about **authentication**. It passed,
and a green light on an adjacent property is worse than no check, because it invites you to stop
looking.
Closed by returning to loopback (the TLS and ACME work is retained, just not exposed). All 68 device
enrolments matched the timestamps of test runs, so there is no evidence of abuse — but the window
existed on a freshly published hostname and absence cannot be proven. 39 unused enrolment tokens
were purged, since any could have been minted by someone else and they cost nothing to replace, and
63 test devices removed.
**The admin listener does not become reachable again until it authenticates.** That reorders the UI
work: auth on the listener first, everything else after.
### Open: encrypted uploads, where the operator cannot read the data
Not built. Recorded because the shape is decided by a few early choices, and the current design
happens to leave the door open.
The goal: hand someone an account, let them upload, and be unable to read what they uploaded.
Sketch: a random per-account **master key**, generated on the first device and wrapped under a
key derived from a passphrase (PBKDF2-HMAC-SHA256 — stdlib on both sides). The wrapped key is
stored server-side as an opaque blob, so a new device signs in, fetches it, and unwraps locally;
the server never sees either key. Runs are encrypted client-side with AES-256-GCM, fresh nonce per
run. All of this is stdlib in Go and `javax.crypto` in Kotlin — no dependency either side.
Four consequences that decide whether it is worth it:
1. **What stays readable determines what the UI can do.** The server builds its index by *parsing*
the document — verdict, finding count, started_at. An opaque payload means the client supplies
that metadata or the index disappears, and with it retention-by-verdict and any "runs with
findings" view. The honest version supplies only run id, timestamp and size, and moves the rest
client-side.
2. **Lose the passphrase, lose the data.** That is the feature working, and also the support
burden. It needs a recovery code printed at setup, not a reset flow — there is nothing to reset.
3. **Metadata is not hidden.** The operator still sees which account uploaded, when, how often and
how large. "Cannot see it" is about content, not existence, and saying otherwise would oversell.
4. **It makes `min_anonymization` unenforceable** — a server cannot check a level it cannot read.
That is not a conflict so much as a redundancy: the anonymization floor exists to protect the
user from the operator, and encryption does that better. The two should not both be demanded of
one upload.
What keeps this possible: uploads are already stored byte-for-byte as received, and every index
field is derived in one function (`runs.Put`). The thing to avoid is admin features that *require*
reading content — those would have to be unbuilt later.
### App sign-in, and an undisclosed dependency it surfaced (2026-08-01)
The app can now sign in to the server's identity provider: authorization code with PKCE, a
`Sign in` card in settings, and the `echolot://auth` redirect handled alongside the enrolment one
(told apart by host, since one spends a token and the other completes an authorization).
The detail that decides whether this works on a real phone: **the PKCE verifier is written to
storage before the browser opens**, not held in memory. Handing control to a browser backgrounds
the process and Android may kill it; the callback then arrives at a fresh process. An in-memory
verifier works on a developer's device and fails under memory pressure, which is the worst way for
a sign-in to break.
Nothing from the IdP is retained. The ID token proves who is signing in, once, and the device
credential authenticates everything after — no access tokens stored, no refresh tokens rotated.
**A server remains entirely optional.** All eight probes are device-tier; `serverConfigured` gates
only upload and the account. But answering that question exposed something worth fixing: two probes
hardcode the reference deployment —
```kotlin
DnsCanaryProbe(canaryZone = "c.echo-lot.app", ...) // "Hardcoded to the reference deployment"
StunProbe(serverHost = "fmr-1.echo-lot.app")
```
so a user with no server of their own still sends DNS and STUN traffic to fmr without being told.
For a tool that goes to this much trouble over what leaves the device, an undisclosed dependency on
a third party's infrastructure is the wrong default. It should prefer the configured server, and be
explicit when there is none. **Closed** — both probes take the enrolled server from settings
(canary zone learned from the profile, cleared on re-enroll) and report themselves SKIPPED with
the reason when none is configured.
### v6.broken was a false positive waiting to happen (2026-08-01)
A phone could not open `https://fmr.echo-lot.app` while loading the same server by IP literal
perfectly well. Two things came out of chasing it.
**The admin UI is IPv6-only, by consequence rather than intent.** `fmr.echo-lot.app` has an AAAA
and no A record — verified identical at Cloudflare, Google and Quad9, so DNS itself is healthy.
That follows from reserving all four measurement addresses for testing, which left only `::2` for
management, and `::2` has no IPv4 counterpart. Any client without working IPv6 sees an unreachable
admin interface — a poor property for the interface you reach *from the networks you are debugging*.
**And the app's own `v6.broken` finding was unsound.** It fired on exactly one signal — ICMPv6 echo
getting no reply — with `Confidence.HIGH`. ICMPv6 echo is widely filtered on networks where IPv6
works fine, which is precisely what that phone demonstrated: no ICMPv6 replies, working IPv6 TCP.
The finding asserted a cause it had no evidence for, which is the same class of error as the
multi-homed `100 % downstream loss` earlier: a confident measurement of something that was not
happening.
Now `v6.no_icmp_reply`, severity low, confidence medium, and the text names *both* explanations
instead of choosing one. It is still worth reporting, because filtered ICMPv6 breaks Path MTU
Discovery — large packets vanish rather than being reported as too big — which is a real fault even
when IPv6 works.
The proper fix is corroboration: attempt a real IPv6 connection and only call it broken when that
fails too. That needs a target, which runs into the hardcoded-reference-deployment issue already
open above. **Both closed 2026-08-02** — see "Corroborated IPv6 findings" below.
## Per-network probing is blocked while a VPN is up (2026-08-01)
`Network.bindSocket()` fails with `EPERM` for every underlying network when a VPN holds the
default route — verified on the OnePlus 15 with Netbird active: `Binding socket to network 101
failed: EPERM` for both cellular and wifi. This is Android preventing VPN leaks, not a bug to work
around, and it means the whole per-network measurement approach is unavailable to any user with a
VPN connected. Worth deciding deliberately rather than discovering per report:
- The run currently succeeds and simply measures nothing per network. Honest, but silent — the
document records `attempted: false` and the UI says green.
- A user with a corporate VPN permanently on would get a green run that measured almost nothing.
Options are to detect the VPN and say so plainly ("this network cannot be measured while a VPN is
active"), to measure the tunnel itself as the network under test, or both. **Decided and built
2026-08-02**: say so plainly, everywhere the run is read — see "Constrained runs" below.
Related: `icmp.ping6` now records `attempted` alongside `ok` per network, because collapsing them
made the app report "IPv6 is configured, but ICMPv6 gets no reply" about an interface it had never
succeeded in sending on — a claim about the user's carrier with no evidence behind it.
## Reserved measurement addresses, and the web UI on both families (2026-08-01)
fmr has two IPv4 (.150/.151) and three IPv6 (::150/::151/::2) addresses. `.150`/`::150` now carry
the services; `.151`/`::151` are reserved for measurement, declared in `ECHOLOT_RESERVED_ADDRS`.
Reserved does **not** mean silent. The UDP data plane, the canary DNS and STUN's RFC 5780 alternate
all belong there — reserving an address and then forbidding the measurements that need it would
defeat the purpose. What must never appear is a service, and above all not ports 80 or 443: a
handshake completing on a port known not to be listening is what proves interception, and that
proof survives exactly as long as nothing binds those ports. `config.CheckReserved` enforces it at
startup, refusing wildcard binds outright (every listener defaults to `:port`, so the next one added
will claim reserved addresses without anyone deciding to).
The first version of the guard was too strict and the live config caught it: it would have refused
the existing UDP and DNS binds on `.151`. The rule is about services and web ports, not about
listening at all.
**The adb-beacon receiver was wildcard-bound to `0.0.0.0:443`**, occupying port 443 on every IPv4
address including the reserved one — so the IPv4 interception test had been compromised for as long
as it had been running, silently. It is now `systemctl disable --now echolot-adb-beacon`; restore
with `systemctl enable --now`. Note what this implies: the guard covers this server's own listeners,
and a stray process outside its config can still pollute a reserved address. A startup probe that
*verifies* 80/443 are actually free on the reserved addresses would be a stronger guarantee than
checking our own configuration — built 2026-08-02 (`selftest.ReservedWebPortsFree`, fatal at
startup when anything is listening there).
The admin UI and the ACME responder now take comma-separated addresses like every other listener;
they were single-address, which is why the UI could only ever live on `::2`. It serves on `.150:443` and
`[::150]:443`; sshd on `.150:2322` and `[::150]:2322`.
`::2` is gone entirely — unbound, then removed from `/etc/systemd/network/ext.network`. The
transition kept it bound throughout and dropped it only after the CNAME landed, because removing it
first would have broken both the UI and ACME renewal for the very name the certificate is issued
to. Listeners came off before the address did, in that order, or the services would have failed to
bind on restart.
Verified after a full reboot: `fmr.echo-lot.app` answers 200 over both families, `.151`/`::151` are
closed on 80 and 443, canary DNS is still up on `.151`, and neither `::2` nor the beacon returns.
(`echolot-server` is `After=network-online.target` with `Restart=on-failure`, which is what makes
binding specific addresses safe across a boot — a wildcard bind would not have needed it, and that
is the trade for the reserved addresses being meaningful.)
The point of all this: `fmr.echo-lot.app` gained an A record, so the server stopped being reachable
only over IPv6 — which is what made it unreachable from a phone with no working IPv6, presenting as
"this host does not exist" in two different browsers.
### If the beacon comes back, it belongs in the web UI
Not as a separate listener. The receiver being its own Python service on `0.0.0.0:443` is exactly
what silently compromised the reserved address, and a second process racing for a port is a
recurring problem rather than a one-off: whoever loses the race simply fails to start, and on a
reboot which one that is comes down to unit ordering.
Folding it in costs little and settles several things at once. It would be two routes on the admin
UI (`POST` the observed adb port, `GET /apk` for the staged build), behind the TLS the UI already
terminates and the certificate it already renews, with no extra port and no wildcard. It also gets
authentication for free — the current receiver accepts a port report from anyone who can reach it,
which is tolerable for a dev tool on a trusted network and not something to keep once it lives
beside an admin session.
The one thing that changes on the device side is that the POST becomes HTTPS. That is a real
certificate rather than a self-signed one, so it costs a URL scheme rather than any trust plumbing.
## The control plane shares port 443 (2026-08-01)
`fmr-1.echo-lot.app:443` is the control plane, `fmr.echo-lot.app:443` the admin UI, both on
`.150`/`::150`, one listener, selected by SNI for the certificate and by `Host` for the handler.
The reason is not tidiness, it is reachability. Captive portals, hotel wifi and corporate firewalls
routinely permit only 80 and 443 — which is exactly the population of networks this tool exists to
diagnose. A control plane on 8443 is unreachable precisely when it matters most, and it fails as
"cannot reach server", which tells the user nothing.
They cannot share a certificate, which is why this needs two names. The control plane is trusted by
SPKI pin and so uses a long-lived self-signed certificate; a browser needs one a CA vouches for.
One name on one port is one certificate, so the port can only be shared by splitting the names.
Pinning the Let's Encrypt key instead was considered and rejected: it survives renewal only while
key reuse holds, so a routine key rotation would brick the whole fleet.
Verified per SNI on 443: `fmr.echo-lot.app` serves `issuer=Let's Encrypt`, `fmr-1.echo-lot.app`
serves the self-signed cert whose pin is unchanged (`zRV9…Xlg=`), `/v1/profile` answers 401 on the
control name and 303 to the login page on the UI name.
**8443 stays open.** Devices enrolled before this carry that URL in their settings, and closing it
for the sake of a port number would strand every one of them. It can go once no enrolled device
still points at it — not before.
The rule from the naming change still binds: `fmr` may be a CNAME to exactly one host and never a
multi-address record, because a pinned client that reaches a different key does not fail over.
## Constrained runs: a VPN'd run now says so, everywhere (2026-08-02, app 0.2.1)
The measurement schema gained a top-level `constraints` block (§3) and the app now fills it.
`ConstraintDetector` (core-probe) runs before any probe: one throwaway `Network.bindSocket()` per
non-VPN network, plus a transport check for an active VPN. The result lands in three places, and
all three are deliberate:
- **`run.constraints`** — for machines. A server aggregating thousands of runs can now separate
"measured a healthy network" from "measured almost nothing through a tunnel"; the shapes were
identical before.
- **A `measurement.vpn_constrained` finding** — for the person reading this run, naming the
interfaces that went unmeasured. A constrained run with a quiet findings list still reads as
"nothing wrong here".
- **The §7.3 verdict** — `Verdicts.derive` takes the constraints and returns INCONCLUSIVE
outright for a per-network-blocked run, whatever the category lights say; the run screen shows
an amber "Measured through a VPN" banner above the verdict so INCONCLUSIVE reads as the OS
refusing, not the app failing.
Detection is one bind per network rather than parsing per-test `attempted:false` breadcrumbs, so
it cannot drift when probe evidence formats change.
## Corroborated IPv6 findings: v6.broken is back, with evidence (2026-08-02, app 0.2.1)
The new `V6ConnectProbe` (test type `v6.brokenness`) attempts a real TCP connection over IPv6 to
the configured server's :443, per network that *claims* IPv6 (global address or v6 default
route) — IPv4-only networks are not attempted, since their failure is by design and would
manufacture the exact false positive this exists to kill. The finding derivation is now three-way:
- ICMPv6 silent, TCP works → `v6.no_icmp_reply` at **high** confidence, retitled "ICMPv6 is
filtered here — IPv6 itself works" (still reported: filtered ICMPv6 breaks PMTUD).
- ICMPv6 silent, TCP fails too → **`v6.broken`** (high severity, reinstated in the registry +
findings-registry.md): two independent transports silent on a network advertising IPv6.
- No corroboration (no server configured, or the connect never got as far as sending) → the
two-explanation `v6.no_icmp_reply` at medium confidence, unchanged.
Like STUN and the canary, the probe SKIPs honestly when no server is configured — corroboration
is a benefit of enrollment, not a reason to borrow fmr.
## Server: reserved 80/443 verified against the OS, and signed releases (2026-08-02)
**Reserved-address startup probe.** `serve()` now proves 80/443 are actually free on every
`ECHOLOT_RESERVED_ADDRS` address before starting: a throwaway bind per port
(`selftest.ReservedWebPortsFree`), fatal on EADDRINUSE with the offending address named — the
check `CheckReserved` cannot do, because a stray process outside our config (the adb-beacon
receiver on `0.0.0.0:443` was exactly that) is invisible to configuration checks. Bind errors
that are not "in use" (typo'd address, address not on this host) warn instead of refusing —
they are config problems, not pollution.
**Release signing.** Self-update now trusts a signature, not a host. CI signs `SHA256SUMS` with
an ed25519 key (`relsign` package, `cmd/release-sign`) and the updater refuses any release whose
`SHA256SUMS.sig` is missing or does not verify against the public key baked into the binary
(`selfupdate.DefaultPublicKeyB64`; operators with their own pipeline override via
`ECHOLOT_SELF_UPDATE_PUBKEY`). The private key exists in exactly two places: the Gitea Actions
secret `RELEASE_SIGNING_KEY`, and the offline original on the dev PC at
`~/.echolot/release-signing-key`. It is deliberately NOT on fmr and NOT in the repo — a
compromised release host can withhold updates but no longer inject one. CI hard-fails when the
secret is missing (an unsigned release would strand every verifying server) and cross-checks the
signature against the key in the source it just built.
**ACTION REQUIRED before the next `server-v*` tag:** add the Gitea repo secret
`RELEASE_SIGNING_KEY` (Settings → Actions → Secrets) with the contents of
`~/.echolot/release-signing-key` from the dev PC. Ordering is safe: the currently deployed
v0.3.x updater does not verify, so it will happily install the first signed release; every
release after that is verified. **Done 2026-08-02** — the secret is in place. (The "v0.3.x"
above should read "the currently deployed release": deployments had moved on to v0.9.x by the
time signing landed; the point — the deployed updater predates verification and will accept the
first signed release — is unchanged.)
## Prober fold: traceroute.udp4 and the mDNS inventory go production (2026-08-02, app 0.2.2)
The two highest-value validated capabilities moved from the prober into `core-probe`:
- **`traceroute.udp4`** (`TracerouteProbe`): UDP traceroute reading ICMP time-exceeded off the
socket error queue via `Os.recvmsg(MSG_ERRQUEUE)` through the reflection facade — no root, no
raw socket, no JNI, ~250 ms for six hops. Emits the schema's `TracerouteEvidence` (rtt in ns).
`OsAbi` came with it, including the measured fact that `Os.getsockoptInt` exists on neither
known device, so PMTU must always be read from the errqueue (`ee_info`), never
`getsockopt(IP_MTU)`. The load-bearing line survived the port: EAGAIN out of the reflected
`recvmsg` means "queue empty", not failure.
- **`local.mdns_inventory`** (`MdnsInventoryProbe`): MulticastLock + NSD discovery, the service
inventory that doubles as the VLAN-leakage detector. Both hardware lessons kept: the
`_services._dns-sd._udp.` meta-query returns 0 beside live services on both devices (so the
concrete types are the measurement and the meta-query result is itself evidence), and the
listen window is 10 s because 4 s missed services.
Still to fold, in order: the Shizuku dump *parsers* (the raw `link.ip_monitor` captures already
hold two divergent vendor formats that could feed `link.ra_source` and `sec.arp_watch`);
`multinetwork.request_and_bind` (extend ConstraintDetector to *request* transports rather than
only probing present ones); `peer.ble_advertise` (needs three new permissions and a peer mode to
exist first).
## Server v0.9.2: trains, real TTL/DSCP/ECN, rate limits (2026-08-02)
The spec-vs-implementation gap audit closed its top items; protocol_version 1.0.0 → 1.0.1
(additive — below 1.0.0 the minor is the breaking axis, and nothing here breaks an old client):
- **Upstream trains** (§3.2, types 0x03/0x04/0x05): per-train bounded columnar buffer (8192
rows, head kept on overflow with `Truncated` set — mirrors the schema's `evidence_truncated`
honesty), TRAIN_REPORT split across ≤1200-byte datagrams, grant-free with the §3.4 argument
spelled out (a 17-byte report row answers a ≥36-byte HMAC-valid packet). Unknown train id
gets a zero-row report: "nothing arrived" is an answer. Also surfaced as `udp.trains` in the
observations API.
- **Real TTL/DSCP/ECN observation** (§3.3): the read loop is `ReadMsgUDPAddrPort` with
IP_RECVTTL/IP_RECVTOS/IPV6_RECVHOPLIMIT/IPV6_RECVTCLASS cmsgs on Linux; `0xFF` stays the
"not observed" sentinel elsewhere. This unblocks `sec.dscp_ecn_survival` both directions,
paired with the new `dscp` parameter on `downtrain` (validated 063, refused not clamped,
`dscp_applied` in the response).
- **Rate limiting** (§2.5, was entirely absent): token buckets keyed per credential AND per
source IP; 429 + Retry-After on session/action creation (`/v1/profile` stays ungated), silent
drop on the data plane — charged after the HMAC gate so a spoofed flood cannot drain a
victim's budget, before the replay window so a dropped seq stays usable. UDP ceilings default
above the largest legitimate run (a 200 Mbps throughput test), because a rate limit that
clips a real measurement produces a confidently wrong number.
- **`action_id` in every granted packet** (§5/§9): payload bytes [8:16] across all granted
types, so overlapping actions are attributable. Verified the deployed Kotlin client parses
only ECHO_RESP and MTU_ACK payloads, so the reshuffle strands nobody.
- **Canary log retention**: the stated 24 h privacy default is now enforced
(`ECHOLOT_DNS_LOG_RETENTION_H`), where before the log was time-unbounded.
- **`POST /admin/enroll-tokens`** now answers the spec's JSON shape under content negotiation;
the README's curl works as documented.
- Spec §2.3 registry gained `downtrain` and `tcp-echo`, which the server had been advertising
as strings a conformant client must ignore.
Client-side counterparts still to build: sending 0x03 trains + parsing 0x05 reports
(`train.udp_updown`), and passing `dscp` on downtrain actions.
## ⚠ Version lineage broken: fmr runs v0.11.2, the repo's tags stop at v0.9.x (2026-08-02)
Discovered while preparing to self-update fmr to the freshly released server-v0.9.2:
**fmr runs v0.11.2** (binary installed 2026-08-02 08:51), but this repo's remote has tags only
up to `server-v0.9.1`, master fast-forwarded cleanly from this machine, there is no v0.10/v0.11
release in Gitea, no source checkout or Go toolchain on fmr, and no deploy script in this repo
that stamps versions. Conclusion: v0.11.2 was cross-built from a clone whose commits were never
pushed — presumably another dev machine.
Consequences until resolved:
- **Do NOT run `--self-update` on fmr.** Gitea's `/releases/latest` is the *newest-created*
release, which is now `server-v0.9.2` — semantically older than the deployed binary; the
updater compares strings, not SemVer, and would happily "update" v0.11.2 down to it. No
automatic risk exists (fmr has no update timer installed, only the cert timer), but a manual
run would downgrade.
- The next real release must be tagged **above v0.11.2** (e.g. `server-v0.11.3` or `v0.12.0`)
*after* the missing commits are pushed, so "latest" becomes truly latest again.
- The unpushed v0.10v0.11 work needs to be found and pushed from whichever machine built it,
or the deployed binary's provenance re-established some other way, before the release channel
can be trusted again.
**Resolved same day.** The binary itself settled it: `go version -m` on the deployed executable
shows `vcs.revision=d5b1bab` — a commit on this repo's master — built 2026-08-02 08:36 UTC with a
hand-stamped `-X main.Version=v0.11.2` and a dirty tree (`vcs.modified=true`, the then-uncommitted
schema doc). An earlier session stamped release numbers ahead of the tag line; no code was ever
missing. Current master is tagged and released as **server-v0.11.3** (signed), restoring a
monotonic, tag-backed lineage above the deployed number. The rule going forward: **the version a
binary is stamped with must be a pushed `server-v*` tag** — an ad-hoc stamp above the tag line
poisons `/releases/latest` for the string-comparing updater the moment anyone tags honestly again.
The stale `server-v0.9.2` release (same code lineage, wrong number, created during the confusion)
remains in Gitea but is harmless now that v0.11.3 outranks it as latest.
## LLDP and CDP are root-tier, and that is a hard boundary (2026-08-02)
Asked for alongside SSDP in long mode; they belong to a different tier and no amount of app-side
cleverness moves them. LLDP is an EtherType `0x88CC` frame to `01:80:C2:00:00:0E`; CDP is an
LLC/SNAP frame to `01:00:0C:CC:CC:CC`. Neither is IP, so neither is ever delivered to a socket an
app can open — receiving them needs `AF_PACKET` with `CAP_NET_RAW`, which is root. Shizuku does
not bridge this either: the ADB shell user (uid 2000) has no `CAP_NET_RAW`, and stock devices do
not ship `tcpdump`. Android's unprivileged ICMP sockets are what make `icmp.ping4` work without
root; there is no equivalent back door for raw L2 receive.
Worth building in the root module when it lands, because the payoff is large: LLDP names the
switch, the port and the VLAN a device is attached to, which is the best available answer to
"where in this building am I actually plugged in", and CDP does the same on Cisco gear. Until
then they are recorded as absent capabilities rather than left to look unimplemented.
What IS reachable at app tier, and what long mode now listens for instead: SSDP (passive NOTIFY
plus periodic M-SEARCH), LLMNR, NetBIOS-NS and WS-Discovery — all IP multicast/broadcast, all
sockets an app may open. The security reading matters as much as the inventory: LLMNR and
NetBIOS-NS being live on a segment is a finding in itself, since both are trivially spoofable.
## Design note: what BLE between two devices is actually for (2026-08-02, not built)
Two or more phones running Echolot, talking over Bluetooth LE. The schema already anticipates
this — `Trigger.PEER` and the whole `peer.*` test family (`peer.reachability`, `peer.isolation`,
`peer.multicast`, `peer.lan_train`, `peer.lease_diff`) are in the registry, unused — and the
prober measured `peer.ble_advertise` **SUPPORTED on both known devices**, so the mechanism is
proven; what has been missing is a reason that beats "use the server".
**The reason is that BLE is out-of-band.** Everything else this app does depends on the network
under test being at least partly functional. A second device reachable over a radio that shares
nothing with the wifi turns several measurements from ambiguous into conclusive:
1. **Client isolation becomes measurable at all.** Today, "I sent a packet to the peer and heard
nothing" cannot distinguish AP client isolation from the peer being asleep, gone, or on a
different VLAN — the failure mode is silence, and silence has too many parents. With BLE the
peer confirms out-of-band that it was listening on address X at time T, so silence over IP
becomes *proof* of isolation rather than a guess. This is the single strongest argument for
the feature, and it mirrors the rule this project keeps rediscovering: a measurement that
cannot separate "nothing happened" from "nothing was tried" is not a measurement.
2. **Differential diagnosis: the network or this phone?** Two devices on the same SSID, one
resolving DNS and one not, settles in seconds what a single device cannot settle at all —
and it is the same distinction `system_verdict` exists to draw, only with a second opinion
instead of Android's. Natural finding: *this device fails where a peer on the same link
succeeds* → look at the device (private DNS, ad blocker, per-client router rule, MAC
randomization), not the router.
3. **Two DHCP servers on one L2**, the classic invisible fault: peers compare lease source,
subnet and gateway (`peer.lease_diff`). Disagreement is conclusive and needs no server.
4. **Coverage and roaming**, later: several devices sampling RSSI in different rooms, exchanging
summaries over BLE, gives a picture no single device standing in one place can produce.
**What crosses the link is a summary, never the document.** A measurement document describes
someone's home network in detail; broadcasting it to whoever is nearby would betray the whole
posture of §8. The peer payload should be: a *hashed* network identity (so two devices can agree
they are on the same L2 without either putting the SSID/BSSID on the air in the clear), the §7.3
category verdicts, the finding codes, and an IP endpoint plus a one-shot nonce for the LAN tests.
Findings and verdicts are already the interpretation layer — exactly the right granularity to
share.
**Privacy constraints, which are not optional here.** A BLE advertiser is a tracking beacon: it
must be user-initiated, time-boxed to the run, carry no identifier that is stable across runs
(the resolvable-private-address default plus a per-session ephemeral id), and pair by a code the
two humans can see. "Discoverable by default" would make this app a worse citizen than the
networks it audits.
**Deliberately not doing:** clock synchronisation over BLE. GATT latency is jitter measured in
tens of milliseconds, which is the same order as the one-way delays worth measuring; peers should
sync against the server's `time.server_offset` and use BLE only to correlate run ids. Nor should
BLE become a transport for uploads — it is a *comparison* channel.
Staging when it happens: `peer.isolation` first (highest value, needs only advertise + connect +
a nonce exchange), then `peer.lease_diff` (pure summary comparison, no extra plumbing), then the
rest. Needs `BLUETOOTH_ADVERTISE/CONNECT/SCAN` in the manifest, which the app does not yet
request.
## v0.11.3 live on fmr; trains validated end to end (2026-08-02)
Deployed via `--self-update` (the pre-signing v0.11.2 updater accepted the first signed release,
as planned; every later update verifies). Startup clean on the real host — the reserved-port
check passed against the OS, self-test green, capabilities unchanged plus the new machinery.
**`train.udp_updown` validated against production**: `LiveUpstreamTrainTest` from this PC sent
120 packets; the server's ledger counted 120, both columnar report parts arrived, loss 0.0 %,
`truncated=false`. The 0x03/0x04/0x05 path works, client and server, over the real internet.
Two operational bugs surfaced doing it:
- **CLI-minted tokens are lost while the daemon runs.** `devices.json` is loaded once at startup
and held in memory; `--mint-enroll-token` writes to disk, the running daemon never re-reads,
answers "unknown token", and clobbers the token on its next write. `enroll-link.sh` has only
ever worked by timing luck. Workaround used: mint, `systemctl restart echolot-server`, then
redeem. Real fix belongs server-side (re-read on miss, or route the CLI mint through the
running daemon).
- **`test-fmr.sh` still mints against `127.0.0.1:8444`**, which no longer exists (the admin API
moved to authenticated :443). Needs the same CLI-mint flow enroll-link.sh uses — plus the
restart caveat above until that bug is fixed.