# Echolot — build status & next steps Last updated: 2026-07-29. ## Decided - Name **Echolot**; domain echo-lot.app; scheme `echolot://`; namespace `app.echo_lot.*` (hyphen→underscore; appIds/packages can't contain hyphens). Prober appId `app.echo_lot.prober`. - Stack: native **Kotlin + Jetpack Compose**, no Flutter. - Tiers: `app` (no root), **`shizuku` in v1** (wireless-ADB pairing), `root` future. - **License (decided 2026-07-30):** all code **GPL-3.0-or-later**; the specs in `docs/` **CC-BY-4.0**. Rationale: the moat is the no-root platform research, which is trivially liftable into a proprietary repackage — copyleft is the only option that prevents that, and the F-Droid / network-engineer audience reads GPL as a trust signal rather than friction. The specs go permissive on purpose: a wire protocol only becomes a standard if anyone can implement it. **Not AGPL for the server** — the ASP loophole it closes is speculative here, while blanket corporate AGPL bans would hit exactly the enterprise network teams most likely to self-host. Sole copyright holder, so relicensing the server to AGPL later stays possible. Open: substitute a real legal copyright holder for "Echolot contributors" in the SPDX headers. ## Specs (in `docs/`, alongside this file) - `feature-catalog-and-feasibility.md`, `measurement-schema.md`, `probe-protocol.md`. Considered draft-complete and reviewed by the user. ## Capability prober — DELIVERED (as zip, 2026-07-29) Full Kotlin/Compose project scaffolded: `app.echo_lot.prober`, minSdk 26 / target+compile 35, AGP 8.7.3, Kotlin 2.0.21, Compose BOM 2024.10, Shizuku api+provider 13.1.5, kotlinx-serialization. Probes implemented: `link.snapshot`, `icmp.ping4/6` (unprivileged ICMP datagram), `sockopt.matrix` (TTL/TOS/RECVERR/MTU_DISCOVER), `trace.errqueue_reachable`, `multinetwork.request_and_bind`, `local.mdns_discover`, `peer.ble_advertise`, `shizuku.command_battery` (ip neigh / ip -6 route / ip addr / ip monitor / dumpsys network_stack DHCP+IpClient / dumpsys wifi). JSON export via share intent; results carry verdict + raw evidence. Could NOT be compiled in the cloud sandbox: dl.google.com (Google Maven) and services.gradle.org are proxy-blocked, and no device is reachable for on-device runs. Build + iterate locally (Android Studio / Claude Code). Wrapper is pinned to Gradle 8.14.3. ## Website — `web/` (2026-07-30, LIVE at https://echo-lot.app) Single-page site for echo-lot.app on Cloudflare Workers: static `public/` from the edge, a small Worker for the stable short URLs (`/apk`, `/apk.sha256`, `/fdroid`, `/source`) and `/api/latest`. Version info is resolved from the Gitea releases API at request time (edge-cached 5 min), so tagging a release is the only publish step — the homepage shows the latest version/date/size via progressive-enhancement JS and falls back to pre-release copy otherwise. Auto light/dark via `prefers-color-scheme`: dark = sonar display, light = chart-recorder printout, same amber-accent token system. Deploy: `npx wrangler deploy` from `web/`, or CI via `.gitea/workflows/deploy-site.yml` (needs `CLOUDFLARE_API_TOKEN` + `CLOUDFLARE_ACCOUNT_ID` secrets); full story in `web/README.md`. Branding applied (2026-07-30): real favicon/wordmark copied into `web/public/assets/` (copies — `assets/branding/` is the source of truth), palette switched to abyss/teal/amber with teal-as-instrument semantics, sonar-sweep hero replaced by a Focus-lattice panel (rogue-DHCP finding vignette), tagline "measure, don't guess". ## Branding — `assets/branding/` (2026-07-30) Identity chosen: the **Focus** mark — a quiet lattice of nodes with exactly one under examination (amber) inside teal viewfinder brackets. Tone is surgical measurement, not scanning (sonar-sweep visuals and the "sound out your network" tagline were explicitly rejected); tagline is **"measure, don't guess"**. Palette: abyss `#071522`, tile `#0E2433`, instrument teal `#35E0C4`, finding amber `#FFB454`, foam `#E8F4F2` — teal is always the instrument, the single amber point is the finding. The wordmark is hand-drawn monoline SVG paths (zero font dependency); the second "o" is the signature teal ring + amber ping. Assets: `icon.svg` (+ `icon.png` 512² — Gitea avatars reject SVG), `icon-adaptive-{foreground,background}.svg` (108dp adaptive-icon layers, art inside the 66dp safe circle), `wordmark-on-{dark,light}.svg`, `banner.svg` (1200×300, wired into the repo README), `social-preview.svg`/`.png` (1280×640, for the Gitea/GitHub social-preview slot). The website's placeholder favicon/wordmark can now be replaced from these. ## First device report — OnePlus 15, Android 16/SDK 36 (2026-07-30) Archived at `echolot-prober/reports/CPH2747-android16-sdk36.json`. Collection workflow that works: deploy over (wireless) adb, but the **user** starts the run and exports the JSON manually — probe runs reliably kill the adb-over-wifi session (multinetwork/multicast churn the link adb rides on). OS findings: - `link.snapshot`, `icmp.ping4` (38 ms), `sockopt.matrix` 4/4, `peer.ble_advertise`: SUPPORTED. - **`trace.errqueue_reachable` SUPPORTED** — full Os-API errqueue path exists on Android 16; the C-over-JNI shim (next-steps #2) may be unnecessary on modern devices. Still needs a real errqueue round-trip test, and older devices may differ. - `icmp.ping6` EAGAIN — topology, not capability: wifi was v4-only, v6 only on cellular, socket used the default network. Probe should bind per-network (also covers the ping4 case properly). - `shizuku.command_battery` SUPPORTED, shell(2000): real `ip -6 route` RA data (lifetimes, per-table), `ip monitor` NEIGH events, wifi dump with full DhcpResults. Vendor formats look parseable. - `local.mdns_discover`: lock acquired, discovery ran, 0 services in the 4 s window — needs a longer window / known-good reference network before calling it more than formally SUPPORTED. Prober bugs the run exposed (all fixed same day): - rotation wiped results + cancelled runs → state moved to a ViewModel; - `multinetwork.request_and_bind` ERROR → missing `CHANGE_NETWORK_STATE` in the manifest; - `rtt_ms: "38,1"` → device locale leaked into report numbers, now `Locale.ROOT`; - 3/7 Shizuku commands `SHIZUKU_BIND_TIMEOUT` → bind/unbind-per-command raced Shizuku; now one bind per battery (`execBatch`). Run 2 (same device, fixed APK — `…-run2.json`) confirmed all three fixes: - `multinetwork.request_and_bind` **SUPPORTED** — wifi + cellular acquired and bound concurrently (downKbps 39094 / 13231); requesting cellular even brought up a second, dual-stack PDN (`rmnet_data2`, private v4 /30 + carrier-NAT DNS) next to the IPv6-only default PDN — exactly the per-network behavior the production app needs. - Shizuku battery **7/7** in 2.9 s (was 4/7 in 34 s): full ARP/NDP neighbor table, `ip addr`, IpClient ProvisioningConfiguration incl. SSID/BSSID and APF caps (v6000, 4096 B). - `rtt_ms: "32.0"` — locale fix holds. Still open (probe improvements, not blockers): ping6 should bind per-network (still EAGAIN via v4-only default wifi); mDNS window too short to see real services (0 found on a network that has them); `ip monitor` 2 s window caught nothing on a quiet net — consider longer/event-triggered capture; IP_MTU read via getsockoptInt reflection unavailable — reader belongs in the errqueue shim if that lands. ## Go server skeleton — `server/` (2026-07-30) Pure-Go stdlib implementation of the spec's core: control plane (enroll with single-use tokens, profile with SPKI pin, sessions with the §2.4 HKDF key schedule) + UDP data plane (ELT1 header, HMAC gate, 1024-wide anti-replay, ECHO_RESP with observation block, TIMESYNC, §3.4 anti-amplification). Wire format covered by `go test` (roundtrip, replay/bad-HMAC/unknown-prefix silent drops); full enroll→profile→session flow smoke-tested live. Two run modes: **container** (autodetected via /.dockerenv etc. or `--docker`; config from `ECHOLOT_*` env; distroless image, `network_mode: host` required — Docker NAT would falsify the observed sources the protocol measures) and **native** (`--install-systemd`/`--uninstall-systemd`, opt-in `--self-update` against a Gitea releases API). CI: `.gitea/workflows/build-server.yml` — tests on push; `server-v*` tags push the image to the Gitea registry + attach linux amd64/arm64 binaries that self-update consumes. Not yet: TCP/TLS echo, STUN, canary DNS, actions, observations API, real admin UI, self-update signature verification. Prober build numbering: ascending `versionCode` shown on screen + as `proberBuild` in the JSON report; current deployed build = **3**. Bump on every deployed change. ## Build-2 reports: OnePlus 15 (A16) + Lenovo TB330FU (A15) — 2026-07-30 Archived as `…-build2.json` in `echolot-prober/reports/`. The findings: - **`traceroute.udp4` SUPPORTED on BOTH devices** (SDK 35 + 36): 6 real hops via pure-Kotlin `Os.recvmsg(MSG_ERRQUEUE)` in ~250 ms, identical paths (A1 → Cloudflare). **The C-over-JNI errqueue shim is dead** for Android 15+; keep it in mind only if pre-15 devices matter. - Per-network ICMP works as designed: phone ping6 = "cellular only, default has no v6 path" (SUPPORTED with topology evidence instead of the old bare ERROR); per-network RTTs (wifi 20 ms vs cellular 201 ms on v4). - **Network finding (dogfood!): hudeWLAN advertises a v6 default route via RA but hands out no global v6 address** — both devices have only link-local + a ::/0 route on wifi. That's a half-broken IPv6 config on the LAN, exactly the class of issue the product should flag. - mDNS: meta-query returns 0 on BOTH devices (NsdManager limitation, now confirmed twice); concrete `_http._tcp` finds 4 real services (3× "Magic 2 LAN", EPSON WF-7840) on both. Production inventory must enumerate concrete types, not trust the meta-query. - Tablet Shizuku: binder alive, permission granted, but UserService bind timed out (0/7) — first-spawn dex extraction on slow storage suspected; build 3 raises the bind window to 25 s + one retry. Phone stays 7/7 with rich neighbor/RA/DHCP evidence. **Build-3 verdict: not timing.** Both 25 s attempts timed out (50 s total) — the UserService spawn genuinely fails on the Lenovo/A15. Build 4 adds a reflection fallback to the legacy `Shizuku.newProcess` remote-process API when the bind fails; `exec_path` in the evidence says which path ran. Whatever the outcome, core-shizuku must not assume UserService works everywhere. (Phone build 3: still 7/7; `ip_monitor` now catches provoked NEIGH PROBE→REACHABLE transitions, and a mid-path router dropping one TTL round showed the "*" hop path works in traceroute.udp4.) **Build-4 verdict (`…-build4.json`): the fallback WORKS.** Tablet Shizuku now SUPPORTED via `exec_path: newProcess fallback` — 6/7 commands with real data (full neighbor table, per-netId v6 route tables, Lenovo's IpClient dump — a distinctly different format from OnePlus's, i.e. the per-vendor parser sample we wanted). **core-shizuku design consequence: dual-path executor is mandatory** — UserService where it binds, newProcess where it doesn't; both confirmed live on real hardware. Known rough edges: `ip monitor` times out under newProcess (subshell trick doesn't survive that path); the 2×25 s bind wait before falling back is worth shortening once a device is known-bad. Tablet UserService root cause still open — multi-user (3 users) is the prime suspect; Shizuku app is current (13.6.0), so the stale-app theory is dead. - `ip_monitor` returned no events this run even with the provoked gateway ping (gateway was already REACHABLE, so no NEIGH transition happened). Evidence-dependent, not a bug. ## Pipeline shakeout — 2026-07-30, partially done Repo pushed to git.rambossek.at/EchoLot/echolot (ssh :2222). Actions runner works: `server-test` green on push. `server-v0.1.0` release shipped with all assets (`echolot-server_linux_{amd64,arm64}` + `SHA256SUMS`) after two workflow fixes: mixed paths+tags triggers never fire (split into test-server.yml / build-server.yml), and a greedy `sed` grabbed a nested user id instead of the release id → uploads 404ed (first-match grep now; release creation also falls back to GET-by-tag on re-runs). Binaries verified downloadable → `--self-update` has its artifact source. **Image job green too** on the dedicated `compilesau-echolot` runner (label `echolot`): needed a PAT with package read/write as the `REGISTRY_TOKEN` repo secret — the built-in Actions token is rejected by the container registry. `git.rambossek.at/echolot/echolot-server:{v0.1.0,latest}` pulls anonymously (3.9 MB distroless, verified). Full chain proven: tag → tests → binaries → release assets → registry image → anonymous pull. Note the tag convention: image tags carry the bare version (`v0.1.0`), release tags the namespaced one (`server-v0.1.0`); selfupdate normalizes. ## Toolchain upgrade + build-5 tablet re-run (2026-07-30) Moved to the laptop (USB adb — far more stable than wireless, as expected). Upgraded the build off the JDK-21 ceiling: **AGP 8.7.3 → 9.2.0, Gradle 8.14.3 → 9.6.0, Kotlin 2.0.21 → 2.2.10**, which lets the toolchain run on JDK 25. Two AGP-9 migration edits were required: - AGP 9 has **built-in Kotlin support**, so applying `org.jetbrains.kotlin.android` alongside it fails with "extension with name 'kotlin' already registered" — the alias must be dropped (`kotlin.compose` / `kotlin.serialization` stay; they are separate compiler plugins). - The `kotlinOptions { jvmTarget }` block came from that plugin and no longer resolves; AGP derives jvmTarget from `compileOptions` instead. `compileSdk`/`targetSdk` 35 → 36, because Android Studio only ships API 36.1 by default and installing 35 was avoidable churn. **Build-5 re-run on the TB330FU is verdict-identical to build-4 across all 10 probes** (`…-build5.json`) — the toolchain jump and targetSdk 36 are behaviour-neutral on this device, including the `newProcess` Shizuku fallback (still 7/7, `exec_path: newProcess fallback`, UserService still never binds). Also of note: `traceroute.udp4` is SUPPORTED here (6 hops via errqueue, "no native shim needed"), which is what retires next-step 2 below — the JNI shim is not required on either known device. **OnePlus 15 build-5 re-run (`CPH2747-android16-sdk36-build5.json`): all 10 SUPPORTED, identical to build-3.** `exec_path: UserService` on the phone vs `newProcess fallback` on the tablet, both 7/7 — the dual-path executor requirement is now confirmed on the current toolchain on both devices. Phone is on Android 16 / SDK 36, i.e. `targetSdk = 36` is also exercised natively there. Two collection-loop gotchas found while driving the phone over USB: - **`icmp.ping6`'s verdict is topology-dependent, not build-dependent.** A first run with Shizuku stopped and only 2 networks up reported UNSUPPORTED; with cellular+wifi present it is the expected SUPPORTED / "echo reply on cellular only — default network has no v6 path". Do not read a ping6 verdict without checking `link.snapshot`'s `network_count` in the same report. - **The Shizuku permission dialog times out after 30 s** (`shizuku.command_battery` logs `Shizuku permission not granted` at ~30027 ms). Screenshot→read→tap round-trips can exceed that. Grant the permission *before* starting a run, or accept that the first run after a fresh install burns the Shizuku probe; the grant is persistent, so the next run is clean. - The phone auto-rotates; tap coordinates must be recomputed per orientation, and `uiautomator dump` can report a stale idle state right after a tap — trust the exported JSON's `durationMs` values as ground truth for whether a run really completed, not the button label. ## Next steps 1. Build locally, run on several physical devices (varied Android versions/vendors), collect the JSON reports — especially the real per-device Shizuku dump formats. 2. ~~If `trace.errqueue_reachable` = PARTIAL, add a C-over-JNI errqueue shim.~~ **Retired** — SUPPORTED on both known devices; `traceroute.udp4` reads real hops via `Os.recvmsg` + `StructMsghdr` reflection, so no `:native` module is needed. 3. Start the Go server skeleton (enrollment + profile + sessions + UDP echo with observation blocks + canary-DNS reference records) per probe-protocol.md. 4. Fold confirmed capabilities into the production `core-probe` / `core-shizuku` modules. ## Production probe server — LIVE on dedicated VM "fmr" (2026-07-31) `echolot-server v0.2.0` runs natively (systemd, no docker) on a dedicated VM: 2×IPv4 + 2×IPv6 service addresses (fmr-1/fmr-2.echo-lot.app, dual-stack DNS), a third IPv6 (`::2`) reserved for SSH only — verified untouched by the daemon (explicit multi-address binds, no wildcard). Control: fmr-1:8443 (SPKI pin `zRV9qkiLnRexAeh4RrSfJzbPWO+U/2Oj2/NVM/KfXlg=`, verified externally over v4+v6). UDP data plane on all four service addresses :8442 — the second IP is the stun-5780 substrate. Daily randomized self-update timer installed (checksum-verified against SHA256SUMS; signature verification still TODO before treating the source as untrusted). Host config in `/etc/echolot-server.env`. SSH access for sessions: `ssh claude-echolot`. ## Server v0.3.0 — STUN + TCP echo + observations + actions (2026-07-31) Shipped and deployed to fmr via the server's own `--self-update` (first real exercise: checksum-verified download v0.2.0→v0.3.0, atomic replace, restart — worked). Added over v0.2.0: - **STUN** (RFC 5389 + 5780): 4 service addrs × primary/alt-port grid. Externally verified on v4 AND v6 — binding success with XOR-MAPPED, RESPONSE-ORIGIN, OTHER-ADDRESS present, so the profile now advertises **`stun-5780`** (the second IP earns its keep). - **TCP echo** (:8441): JSON greeting with observed src + real Linux TCP_INFO — verified externally `mss:1440` (v6, 1500−60), options `[sack,wscale]`, then byte-echo. - **Observations API** `GET /v1/sessions/{id}/observations` (per-packet UDP view, connect-back results, TCP records correlated by source IP). - **Actions** `POST /v1/sessions/{id}/actions`: `delayed_echo` (DELAYED_ECHO at the observed data-plane source — NAT-lifetime primitive) and `connect_back` (dials the control-plane source, records connected/refused/timeout+rtt). - Capabilities computed from what's actually wired: `udp-probe, delayed-echo, connect-back, tcp-echo, stun-5780`. Still not implemented: TLS-echo/JA4, HTTP echo, tls-reference, canary DNS (§6.1 reference records), and the train/big-send/frag/throughput actions. Admin UI still token-mint + health only. ## Canary DNS live — server v0.3.1 on fmr (2026-07-31) Zone `c.echo-lot.app` delegated (NS → fmr-1/fmr-2) and authoritative on all 4 service IPs udp+tcp/53. Verified through full public recursion: `ttl-5` A→192.0.2.5 (Cloudflare), `ttl-3600` AAAA→2001:db8::3600 (Google), `big-txt` TXT returned (TCP fallback, truncated over UDP as designed). End-to-end session attribution works: a `..c.echo-lot.app` query resolved via a public resolver shows up in `GET /v1/sessions/{id}/observations` → `dns_canary` with the resolver's real egress IP, transport, and EDNS. First real test already caught a finding: **Google applies 0x20 case randomization** (mixed-case qname), Cloudflare does not — captured via `case_preserved`. Capabilities now: udp-probe, delayed-echo, connect-back, tcp-echo, stun-5780, canary-dns. Kept the hand-rolled stdlib DNS (no miekg/dns) — validated against independent clients. Deployed via `--self-update` (v0.3.0→v0.3.1, checksum-verified). ## Server v0.3.2 + v0.3.3 (2026-07-31) - **v0.3.2 — control-plane security (live on fmr, externally verified):** `POST /v1/echo` reflects the received request head+body (b64) and observed TLS (version/cipher/SNI/ALPN) — captured real SNI `fmr-1.echo-lot.app` and an injected header over public TLS1.3; `GET /v1/tls-reference` returns the served DER chain + pin (cross-checked against the openssl-derived pin). Optional cleartext echo listener (default off). Capability `http-echo`. - **v0.3.3 — MTU probe (data plane):** MTU_PROBE (0x09) → small MTU_ACK (0x0A) carrying the received datagram size; client DF-probes increasing sizes to find path MTU / black holes. ACK is tiny → never amplifies. Tested. - **Note on trains:** upstream trains (TRAIN_DATA 0x03) are already observable — every HMAC-valid packet is recorded (seq/t_rx/size/type) with no per-packet response, so loss/reordering/inter- arrival are visible via GET observations. The dedicated data-plane TRAIN_REPORT (0x05) is deferred: §3.4 anti-amplification means it needs an asymmetric grant + columnar multi-datagram encoding — a focused batch, not a corner to rush. Remaining spec: tls-echo (ClientHello+JA4), TRAIN_REPORT, big/frag-send, throughput, downtrain; real admin UI. ## Server self-test + host tuning — v0.3.4/v0.3.5, fmr proven good (2026-07-31) The daemon now proves its own host is a clean measurement target: - **sysctl audit** (`GET /admin/selftest`, startup warnings): on first run it flagged exactly 4 real issues on fmr — accept_ra=1 on a static-v6 host, accept_redirects=1, send_redirects=1, icmp_ratelimit=1000. Recommended `server/deploy/99-echolot-sysctl.conf` applied (v6 default route/addrs are proto static with 0 RA-derived routes, so disabling accept_ra is safe — verified v6 egress intact after). Now sysctl_ok=true, 0 warnings. - **egress-MTU self-proof**: DF PMTUD via IP_MTU_DISCOVER + getsockopt IP_MTU (v0.3.4 had a bug — read IP_MTU without connecting → ENOTCONN; v0.3.5 connects first). fmr reports 1500 on both v4 and v6 → mtu_ok=true, so client MTU tests are trustworthy. - Both signals ride in the profile as `server_selftest{mtu_ok,sysctl_ok}` so a client can skip MTU testing when the server can't support it honestly. fmr profile now: `{mtu_ok: true, sysctl_ok: true}`. ## Server v0.3.6 — tls-echo / JA4: spec §4 COMPLETE (2026-07-31) The elt-echo TLS variant runs on the TCP-echo port (8441), multiplexed by a timed 0x16 peek (plain echo stays server-speaks-first; a TLS ClientHello routes to the TLS path). It captures the full ClientHello, returns it raw (b64) + as a JA4 fingerprint (FoxIO), then TLS byte-echoes — the sec.clienthello_echo evidence. Hand-rolled ClientHello parser (ciphers/exts/ALPN/ supported_versions/sig-algs, GREASE-excluded), unit-tested. **Cross-client verified on fmr**: openssl → t13d3013eo (30 ciphers), python ssl → t13d1712eo (17) — different stacks, different fingerprints, correct _a structure both. Capability tls-echo. Spec §4 (TCP/TLS/HTTP/STUN) is now fully implemented. Server capabilities: udp-probe, delayed-echo, connect-back, http-echo, tcp-echo, tls-echo, stun-5780, canary-dns. Remaining spec: §5 heavy actions (downtrain/big_send/frag_send/throughput) + the TRAIN_REPORT retrieval path — all gated on the anti-amplification grant machinery (§3.4) — and a real admin UI. ## Production app — echolot-app/, core-protocol proven live (2026-07-31) Started the Android client, bottom-up from the verifiable spine. `echolot-app/` is a multi-module Gradle build; `core-protocol` is a **pure Kotlin/JVM** module (no Android SDK) implementing the client half of probe-protocol.md: SPKI-pinned control plane (enroll/profile/session via HttpsURLConnection — Android-API-1 compatible, hostname verification off since trust is the pin), HKDF-SHA256 session keys, and the ELT1 UDP data plane (HMAC gate, ECHO+observation, MTU probe) — byte-compatible with the Go server. Unit tests pass incl. the RFC 5869 HKDF vector (so key derivation provably matches the server). **Verified END-TO-END against live fmr** via `scripts/test-fmr.sh` (mint token over SSH → enroll on public control plane → run LiveServerTest): profile (8 caps), session, ECHO rtt ~11ms with the observation block round-tripping the client's observed NAT port, MTU probe 1400→1400, observations 298B. Two client bugs found+fixed doing it: java.net.http did hostname verification (switched to HttpsURLConnection) and ECHO needed ≥72-byte requests for the full 40-byte observation to survive §3.4 anti-amplification. Next: core-measurement (schema types), core-probe (port prober probes), core-shizuku (dual-path), Compose UI. ## App: core-measurement + core-engine — full server-facing vertical proven (2026-07-31) Two more pure-Kotlin/JVM modules, both verifiable without a device: - **core-measurement**: the measurement-schema.md document model (two-clock, columnar trains, test-type registry, anonymization types, finding-requires-evidence). The §7.3 deterministic verdict derivation is implemented + unit-tested; document JSON round-trips. - **core-engine**: the run engine composing core-protocol probes into core-measurement documents. Injected clock/UUID source (pure, testable). Runs a server ECHO train → RTT distribution, loss, and NAT-rebinding detection (from the server's observed source port) as train.udp_updown. **Verified END-TO-END against fmr**: 20-packet train, 0% loss, RTT 1.7/2.5/6.9ms, single observed port (no rebinding) → valid MeasurementDocument (2.3kB), overall GREEN. So the whole server-facing stack — protocol client → engine → schema document → verdict — is now proven against the live server, no device needed. Next modules (core-probe device-tier, core-shizuku dual-path, Compose app) are Android + need on-device verification. ## App: installable APK — core-probe + Compose UI (2026-07-31) The production Android app assembles. Android toolchain in echolot-app mirrors the prober (AGP 9.2 built-in Kotlin — do NOT also apply kotlin.android, it double-registers the `kotlin` extension; that was the one build gotcha). Modules added: - **core-probe** (Android lib): Probe→core-measurement Test abstraction; NetworkInventory (LinkProperties → measurement networks[]), LinkSnapshotProbe (link.snapshot), IcmpProbe (per-network icmp.ping4/6, ported from the prober's validated per-network logic). - **app** (Compose): RunViewModel orchestrates probes → assembles a MeasurementDocument with a §7.3 summary + first-pass findings; Compose UI shows overall/ per-category traffic lights, networks, tests (status/metrics), findings; JSON export via share intent. Survives rotation (ViewModel). App-tier only for now; server-facing (core-engine) and Shizuku tier are additive follow-ups (app degrades gracefully without them, like the prober). Debug APK: 9.5 MB, `echolot-app/app/build/outputs/apk/debug/app-debug.apk`. Not yet run on device (needs the user's phone). core-shizuku (dual-path executor) deferred as additive. ## App verified on-device — both phones (2026-07-31) The production Echolot app runs on real hardware, on BOTH devices, via the beacon-managed adb: - OnePlus 15 (CPH2747, A16) and Lenovo TB330FU (A15): link.snapshot OK, icmp.ping4 OK (per-network, RTT ~40ms), icmp.ping6 FAILED → finding "No IPv6 ICMP path on any active network" → category ipv6 yellow → **Overall YELLOW**. The verdict is driven by the real broken-LAN IPv6 (RA default route, no global prefix) we first found with the prober — the product now surfaces it end to end (probe → schema → verdict → traffic-light UI). Wireless-adb beacon (tools/adb-beacon) made this practical: both devices self-report their rotating wireless-debug port to fmr:443; a PC connector keeps adb connected. Debugged live against the restricted LAN (cleartext policy, egress filtering, shared-LAN mDNS crossing, fast port rotation) — all handled. ## App: core-shizuku (shell tier) built + wired (2026-07-31) Ported the prober's validated Shizuku executor into the app as a library module: - AIDL IUserService + UserService (runs `sh -c` as shell/root in the Shizuku-spawned process), Shizuku provider merged into the app manifest. - **ShizukuRunner: the build-4 dual-path executor** — bind the UserService (25s + retry) where it works (OnePlus 7/7), fall back to the legacy `Shizuku.newProcess` reflection API where it never binds (Lenovo). `exec_path` records which path ran. - ShizukuProbe: runs the shell battery (ip neigh / ip -6 route / ip addr / ip monitor / network_stack DHCP / wifi dump) and emits a shizuku-tier `link.ip_monitor` Test with the raw per-device dumps as evidence + commands_ok/exec_path metrics. Self-degrades to UNSUPPORTED when Shizuku isn't running. Wired into RunViewModel (runs after app-tier probes; sets tiers.shizuku). App APK assembles clean. On-device test deferred: at build time no device was reachable (tablet wifi/beacon dropped on the churning LAN; phone wireless debugging disabled to stop reconnect notifications). Will verify on a device later — expecting UserService on the OnePlus, newProcess fallback on the Lenovo, per the prober. ## App on-device: net.captive_portal verified, Shizuku degrades correctly (2026-07-31) Installed the app (core-shizuku + net.captive_portal) on the OnePlus 15 and ran it; report archived at `echolot-app/reports/CPH2747-app-run1.json`. Results: - **net.captive_portal works** — Android's NetworkMonitor logic reproduced: default + wifi both returned HTTP **204** on the HTTPS *and* HTTP generate_204 probes → `validated`; **cellular returned neither (-1/-1) → `no_internet`**. The per-network split immediately surfaces an asymmetry the OS hides (wifi validated, cellular can't reach the checks at all). - **Shizuku tier degrades correctly**: `binder_alive:false` → test UNSUPPORTED, `tiers.shizuku:false` (Shizuku isn't running on the phone). The dual-path *executing* path still needs an on-device test with Shizuku started (expect UserService on this OnePlus). - 5 tests now: link.snapshot ok, icmp.ping4 ok, icmp.ping6 failed, net.captive_portal ok, link.ip_monitor(shizuku) unsupported → overall YELLOW via the IPv6 finding. Also fixed this session: the beacon app itself caused the "wireless debugging connected" notification spam (it re-resolved adbd's own mDNS advertisement, making adbd re-arm each time); now resolves once per service instance and the heartbeat re-POSTs the cached port only. ## App: dns.canary verified against the live server (2026-08-01) Built the client half of the canary-DNS measurement and verified it on the OnePlus against the deployed fmr zone (`echolot-app/reports/CPH2747-app-run2-dns.json`): - All four spec-frozen reference records matched byte-for-byte through the network's own resolver (ttl-5→192.0.2.5, ttl-60→192.0.2.60, ttl-3600→192.0.2.36, ttl-86400→192.0.2.86) → nothing on this path rewrites DNS answers (`dns.answer_integrity` in the green case). - The un-cacheable nonce name `1006ad16.adhoc.c.echo-lot.app` resolved to 192.0.2.21 → `reached_authoritative: true`, proving the query actually reached the canary server rather than being answered from a cache or an interceptor. Findings wired: `dns.answer_rewritten` (high) when a reference mismatches, and `dns.authoritative_unreachable` (medium) when the nonce isn't answered by the canary server. This closes the first full client↔server measurement loop: the Kotlin app measures against the Go server's canary zone on real hardware. 6 tests now run per measurement. ## App: Shizuku shell tier VERIFIED on-device — the app is feature-complete for v1 tiers (2026-08-01) Ran the app on the OnePlus with Shizuku running (`echolot-app/reports/CPH2747-app-run3-shizuku.json`): - **`tiers: {app:true, shizuku:true}`**, shizuku test **ok**, **commands_ok 7/7**, **`exec_path: UserService`** (the OnePlus binds it — matches the prober; the newProcess fallback stays for the Lenovo). - Evidence is the real privileged material the production parsers need: live ARP/NDP neighbor table, per-table IPv6 routes, `ip addr`, an actual `[NEIGH]` netlink event from `ip monitor`, IpClient DHCP logs incl. APF capabilities, and the wifi state dump — all as shell(2000). Both privilege tiers now work end to end in the production app on real hardware, alongside the canary-DNS loop against the live server. 6 tests/run: link.snapshot, icmp.ping4, icmp.ping6, net.captive_portal, dns.canary, link.ip_monitor(shizuku). ## App: nat.stun_5780 — NAT behavior discovery verified against the live server (2026-08-01) Client-side RFC 5389/5780 STUN (hand-rolled, stdlib only) exercising the server's `stun-5780` capability. Three binding requests from ONE socket: primary, the server's OTHER-ADDRESS (alternate IP), and CHANGE-REQUEST(port). Verified on the OnePlus (`echolot-app/reports/CPH2747-app-run4-stun.json`): - local `10.13.102.124` → mapped `178.191.120.247:53259`, `behind_nat: true` - `other_address 89.185.109.151:3479` — the server's second IP answered, so RFC 5780 works end to end (client ↔ our own STUN implementation) - **mapping: endpoint-independent** (same external port toward a different destination → P2P friendly); **filtering: address/port-dependent** (no reply to CHANGE-REQUEST → unsolicited inbound is dropped). Classic full-cone-mapping + port-restricted-filtering NAT. Finding wired: `nat.symmetric` (medium) when mapping is address/port-dependent. Two bugs caught by running it for real: port preservation was misread as "no NAT" (now compares ADDRESSES), and an unbound socket reports the wildcard as its local address (now resolved via a throwaway connected socket). 7 tests/run. ## App: autorun mode + IPv6 severity rework (2026-08-01) **IPv6 is no longer treated as a defect just for being absent.** The finding now depends on whether the network actually provisioned IPv6 (a global v6 address or a `::/0` route): - not provisioned → `ipv6.not_offered`, severity **INFO → green**. Most networks are still IPv4-only and that is not a fault. - provisioned but ICMPv6 fails → `ipv6.broken`, severity **MEDIUM → yellow**. Half-configured IPv6 is worse than none (Happy-Eyeballs stalls). Verified on the OnePlus: our LAN advertises a v6 default route with no working path, so it correctly reports `ipv6.broken`. **Autorun mode** — one adb command runs a full measurement unattended and collects the result without any UI tapping or adb round-trip: ``` adb shell am start -n app.echo_lot.app/.MainActivity --ez autorun true curl http:///reports # list curl http:///report/ # fetch ``` The app runs the suite, POSTs the MeasurementDocument to the collection endpoint (receiver.py gained `POST /report`, `GET /reports`, `GET /report/`), shows the result for 3 s, then finishes itself — leaving the device as it was found. On upload failure it stays open so the error is visible. Grant permissions once via `adb shell pm grant app.echo_lot.app android.permission.ACCESS_FINE_LOCATION` so nothing blocks on a dialog. ## App: router identification, brand icons, DEV build variant (2026-08-01) **`link.ra_source` — who is advertising IPv6 here, and what box is it?** New app-tier probe (registry addition). Identification chain, each step recorded as evidence so nothing is guessed: 1. RA source = next-hop of the `::/0` route per network (a `fe80::` link-local). 2. **MAC recovered from the modified-EUI-64 link-local** (strip `ff:fe`, flip the U/L bit) — e.g. `fe80::7a9a:18ff:fe54:b8f9` → `78:9a:18:54:b8:f9`. RFC 7217/privacy addresses don't encode a MAC and are reported as such rather than guessed. 3. Vendor via a curated OUI table (`Oui.kt` — SOHO/router vendors; unknown OUIs are printed verbatim). Locally-administered (randomized) MACs are flagged. 4. **UPnP/SSDP M-SEARCH** → the gateway's `SERVER:` banner + device-description XML gives manufacturer / model / friendly name. This is what usually names the exact box. 5. Reverse DNS for both gateways. All SSDP responders are recorded (not just the gateway) so a rogue RA sender that isn't the gateway can still be matched — and the MAC travels with every identity source, which is the hook for the future LLDP / mDNS cross-matching. UI: a "Router / IPv6 advertiser" panel above the network list, leading with the identified vendor/model. **Icons + DEV variant.** The branding adaptive icon is now the app icon: `icon-adaptive-*.svg` converted to Android vector drawables (SVG transform baked in, gradient background, monochrome layer for themed icons) plus PNG mipmaps for legacy launchers. The **debug build is a separate app**: `applicationIdSuffix .dev`, label "Echolot DEV", and a DEV-badged icon (layer-list = production foreground + generated amber DEV ribbon) so it is unmistakable next to a real install and both can be installed side by side. NOTE for tooling: the dev package is `app.echo_lot.app.dev`, activity `app.echo_lot.app.MainActivity`. ### link.ra_source verified on-device — it named the actual router (2026-08-01) Run archived at `echolot-app/reports/CPH2747-app-run5-router-id.json`. On the wifi network the probe identified the RA sender completely, from an unprivileged app: - RA source `fe80::7a9a:18ff:fe54:b8f9` → **MAC 78:9A:18:54:B8:F9 recovered via EUI-64** (matches the Shizuku neighbor table exactly) → vendor **MikroTik** by OUI - IPv4 gateway `10.13.102.1`, reverse DNS `router.hudelist.local` - UPnP: `RouterOS/7.23.2 UPnP/1.0 MikroTik` → manufacturer MikroTik, model Router OS, friendly name "MikroTik Router" So the box advertising this LAN's broken IPv6 RA is a **MikroTik running RouterOS 7.23.2**, named by two independent methods (OUI from the address itself + UPnP device description) that corroborate. On cellular the carrier's RA source is an RFC 7217 privacy address and is correctly reported as "not EUI-64" rather than guessed. The SSDP sweep also inventoried the LAN (Synology DS1522+ DSM 7.3, a Sky ES160 gateway) — the raw material for the planned LLDP/mDNS cross-matching by MAC. Fixes from this run: added the confirmed MikroTik OUI 78:9A:18 (+ other RouterBOARD ranges), and an elvis-operator bug that printed "no UPnP response" even when UPnP data was present. ### App UX: progress bar + ETA, cancel, and edge-to-edge insets (2026-08-01) - **Progress + ETA**: `Probe.estimatedMs` (per-probe, from measured on-device durations — the timeout-bound probes dominate: icmp.ping6 ~7s on a v4-only net, captive-portal ~9s, SSDP ~7s, STUN ~6s) drives a determinate bar plus "test N of M · ~Xs left". The Shizuku battery is counted in the total so the bar covers the whole run. - **Cancel**: stops an in-flight run and shows what was measured so far, assembled into a normal document (findings + verdict over the partial set). Deliberately **does not upload** — a partial run is for the person looking at the screen, not for the record. - **Insets/cutout**: Android 15 draws edge-to-edge by default, so the title was running under the status-bar clock and the camera cutout. The root column now uses `safeDrawingPadding()`, which covers status bar, navigation bar and display cutout. ### Shell-tier readiness shown before a run (2026-08-01) `ShizukuAvailability` distinguishes four states and the UI only speaks when it is actionable: - **NOT_INSTALLED → says nothing.** Users who don't use Shizuku are never nagged. - **INSTALLED_NOT_RUNNING → amber banner** "Shizuku is installed but not running — start it to include shell-tier tests". This is the case worth reminding about: the user has it, but a stopped service silently costs them the whole shell tier. - NEEDS_PERMISSION → "running but not authorised, it will ask on first use". - READY → green "shell-tier tests will run". Detection is listener-based (`addBinderReceivedListenerSticky` + binder-dead), because `pingBinder()` is only truthful once ShizukuProvider has delivered the binder — a one-shot poll at launch would show a false "not running". Installed-vs-not needs the `` package-visibility entry on Android 11+. Verified on-device: with shizuku_server stopped the banner appears correctly. ### Shizuku banner is actionable; progress + cancel verified on-device (2026-08-01) Tapping the shell-tier banner now does the right thing per state: **installed-but-stopped** → deep-links into the Shizuku app (a third-party app *cannot* start Shizuku itself; the wireless- debugging pairing flow is privileged and lives in that app, so taking the user there in one tap is the best available), **running-but-unauthorised** → fires the Shizuku permission request directly. The hint line states which action the tap performs. Verified on-device in one screenshot: progress bar at "test 4 of 8 · icmp.ping6 · ~33s left", Cancel button beside the disabled Run button, title clear of the status bar/cutout, and the banner having live-switched from "installed but not running" to "running but not authorised" via the binder listener when Shizuku was started mid-session. ### Why the Shizuku banner can't start wireless debugging directly (verified, 2026-08-01) Checked against Shizuku 13.6's own manifest (pulled the APK, `aapt2 dump xmltree`): the wireless-debugging entry points — `moe.shizuku.manager.adb.AdbPairingTutorialActivity`, `moe.shizuku.manager.adb.AdbPairingService`, `moe.shizuku.manager.starter.StarterActivity` — declare **no intent filters**, so they are not exported and a third-party app cannot launch them. `MainActivity` is the only reachable entry and answers MAIN/LAUNCHER only (no deep link), which is why the handoff lands on the screen whose primary action is the root start. Best available behavior, now implemented: the banner still opens Shizuku, but the hint names the exact steps there ("Pairing", then "Start"), and a second tap target opens **Developer options** (`Settings.ACTION_APPLICATION_DEVELOPMENT_SETTINGS` — public and exported) since Wireless debugging must be enabled first for Shizuku's wireless start to work at all. ### Downstream measurements: asymmetric grants, DF-mode big_send (server-v0.4.0 … v0.4.2, 2026-08-01) The client can measure a round trip and the largest packet it can *send*. It cannot measure the largest packet it can *receive*, or downstream-only loss — those need the server to push, which is exactly what §3.4 gates behind an asymmetric grant. Implemented and verified live from the PC: - **`session.Grant`** — created per action, bound at creation to the session's *observed* data-plane source (no grant without a verified destination), clamped to server limits, with a byte budget, an average-rate ceiling and an expiry. Unit-tested for each of those refusals. - **`downtrain`** — N packets of size S every I µs; the client derives downstream loss, reordering and inter-arrival spacing. - **`big_send`** — one datagram per requested size. **DF is on by default**, so the largest size that arrives *is* the downstream path MTU. Without DF the kernel fragments and the result only says whether fragments get through — a different fact, and the reason the schema has both `mtu.pmtud_down` and `mtu.frag_delivery`. Sizes above the server's own egress MTU (from the startup self-test) are refused up front and reported as `max_df_bytes`, so an absence caused by our kernel is never read as a limit of the client's path. Live from the PC against fmr: downstream path MTU **1500** (1472 payload, DF), fragmented delivery up to **4000**, downstream train **100/100, 0 % loss, 0 reordered**, inter-arrival 3.3 ms for a 3000 µs send interval. #### Two bugs this shook out, both invisible in a single-homed lab 1. **Granted sends went out from the wrong local address** (fixed in server-v0.4.2). fmr binds two IPv4 addresses; `connFor` returned whichever socket of the right family came first in the bind list. A train for a session established on `.150` left from `.151` and every packet was dropped by the client's NAT, which has no mapping for that pair. tcpdump showed all 50 leaving, the client saw none — reported as *100 % downstream loss*, a confident measurement of something that never happened. Sessions now record which of our own bound addresses received their traffic and granted sends go back through that socket; `connfor_test.go` pins both that and the family fallback. 2. **A second `ProbeSession` on one server session is silently dead.** Sequence numbers restart at zero client-side while the server's anti-replay window keeps counting, so every packet is discarded as a replay — and because the server then never records the new source, the grant still targets the closed socket. `ServerMeasurement` now uses one ProbeSession for the whole run; `ProbeSession`'s doc comment states the constraint. ### Run archive, anonymizer and uploads (2026-08-01) Three pieces, deliberately separate: - **`core-archive`** — one JSON file per run plus an index entry, in a plain directory the user can inspect or delete with a file manager. Retention (max runs / max age / max total bytes) is enforced on every save rather than by a sweeper. `enabled` is a separate flag from the three ceilings because "no limits" and "keep nothing" are opposite intentions; collapsing them onto all-zeros is how a user who turns the caps off ends up with an empty history. 13 tests. - **`core-privacy`** — the schema §8 anonymizer, three levels. `full` (your own server) changes nothing; `balanced` pseudonymizes SSIDs/hostnames, keeps the OUI half of a MAC and the /16 of a public IP, keeps RFC1918 verbatim (it describes topology, not a person), and *drops* neighbour inventories (SSDP/ARP/scan results) rather than mangling them; `strict` keeps only metrics, statuses and finding codes. Pseudonyms are consistent within a document and — by default — not across documents, so an upload endpoint cannot link a device's runs; a stable salt is opt-in for people diffing their own history. Classification is one readable table, not annotations spread across modules. 14 tests, each pinning a property someone's privacy depends on. - **Server-side upload policy** — `off | anonymous | account`, plus max size, retention days, max runs per device, and the *least* anonymization accepted. The profile advertises all of it so the app presents the choice honestly instead of discovering the rules by being rejected. `account` refuses today rather than falling back to anonymous: picking the strict setting before OIDC lands must not silently mean the loose one. **The archive holds the unredacted document; redaction happens on the way out, per upload.** The local archive is the user's own data on their own device, and redacting it would destroy exactly the detail that makes a week-old run worth keeping. App-side: settings screen (archive limits, privacy level with a plain-language description of what each keeps, auto-upload off by default, server URL/pin/credential), history screen showing whether each run left the device, and a **preview of the exact bytes an upload would send** — an anonymizer the user cannot inspect is only a promise. Live round trip against fmr: uploaded a run, listed it, fetched it back and asserted the SSID, the SSDP neighbour name and the free-text note are absent from what the server stores while the finding code and the metrics survive, then deleted it. ### Still open - `mtu.pmtud_up` (DF + errqueue), `frag_send`, `throughput`, TRAIN_REPORT retrieval. - Enrollment UI in the app (server URL/pin/credential are typed in by hand today). - Accounts/OIDC on the server, which is what `uploads=account` is waiting for. - Nothing in this entry has been exercised on a phone yet — all of it was verified from the PC against the live server. On-device verification is the next step. ### SemVer compatibility windows between app and server (server-v0.5.0 … v0.5.2, 2026-08-01) Both artifacts are SemVer, and each now declares — and enforces — which peer versions it will talk to. Spec: `docs/probe-protocol.md` §8. **Two axes, deliberately not conflated.** Release versions are a *proxy* for what actually has to match, so the real thing is checked first: - `protocol_version` — **can** these builds talk. Advertised in the profile; a peer in a different breaking series is refused whatever its release version says. Below 1.0.0 the **minor** is the breaking axis (SemVer §4). - release-version window — **may** they, per policy. `[min, max)`, min inclusive, max exclusive, because the useful bound is always "the version that broke it". Bounds sit at breaking boundaries, not at releases, so shipping a patch never requires editing a range. The app requires server `>= 0.4.2` for a stated reason, not caution: earlier multi-homed servers mis-addressed granted sends and the client measured 100 % downstream loss that never happened. Operators override the server side with `ECHOLOT_MIN_APP_VERSION` / `ECHOLOT_MAX_APP_VERSION`; a malformed bound is fatal at startup rather than ignored, so a typo cannot silently disable a restriction. Three rules that shaped the implementation: 1. **`GET /v1/profile` is never gated.** It is where a refused client learns which version it needs; gating it leaves the user with a network error instead of an answer. 2. **An unparseable or absent version is `unknown`, and is allowed.** Dev builds report `dev`, and a client too old to send the header cannot be identified anyway. 3. **Refusal is 426 with a body naming both versions and the window**, surfaced client-side as a distinct `VersionRefused` rather than folded into "network error". The app's `versionCode` is now derived from its SemVer (`major*1e6 + minor*1e4 + patch*10`) instead of being a second number to remember. Verified live against fmr (`LiveCompatTest`): profile advertises the window and stays readable for a refused version; 0.1.0 and 99.0.0 are both refused with actionable messages; 0.2.0 and a missing header are both served. One user-visible bug caught in the process: Go's JSON encoder HTML-escapes `<`, `>` and `&` by default, so the refusal reached the client as `needs \u003e= 0.2.0`. Disabled at the encoder (this is an API, not a page), and the client now *parses* the error field instead of pattern-matching it, so it survives whatever a future encoder decides to escape. ### Enrollment: the server mints the bootstrap link (server-v0.5.3 … v0.5.4, 2026-08-01) Until now a device was configured by hand-typing a control URL, a base64 SPKI pin and a credential. That is the step that goes wrong, and it goes wrong quietly: a pin off by one character does not fail loudly, it just never matches, and surfaces days later as an inscrutable TLS error. `POST /admin/enroll-tokens` now returns the whole §2.1 bootstrap link alongside the token, because the server is the only party holding all three parts at once. The app takes it from a paste or an `echolot://enroll` deep link (so a QR scan configures a server in one action) and writes URL, pin and credential **together or not at all** — a half-applied server fails later, somewhere else, with an error pointing at the wrong thing. The control URL comes from `ECHOLOT_PUBLIC_URL` (set on fmr to `https://fmr-1.echo-lot.app:8443`), falling back to the first control listen address; a wildcard bind warns rather than emitting a link to `0.0.0.0`. **The encoding trap, which is the whole reason this is tested across both languages.** The pin is base64, so it contains `+`, `/` and `=` — each of which means something else in a query string. An unencoded `+` decodes to a space, leaving the pin wrong by exactly one character. Base64 has no spaces, so the parser restores them; that cannot damage a correctly-encoded pin and it rescues every hand-assembled link. `LiveEnrollmentTest` redeems a link the *server* produced, which is the only way to catch a disagreement between the Go assembler and the Kotlin parser — a unit test on either side alone cannot see it. It also asserts the token is refused the second time. Also fixed a spec divergence found while reading §2.1: the spec names the field `device_credential`, the first implementation shipped `credential`. The server now sends both and the client prefers the spec's; the alias goes once nothing reads it. Two process notes from this round: - An edit to the admin handler silently failed to apply and the endpoint kept returning just the token. Caught by deploying and *looking at the response*, not by trusting a green build. - The live suite is now six tests (`LiveServerTest`, `LiveMeasurement`, `LiveGranted`, `LiveUpload`, `LiveCompat`, `LiveEnrollment`), all green against fmr from the PC with no device. ### Directional loss: which way is the packet loss? (2026-08-01) A round trip can only report that *something* was lost somewhere, which is the least useful form of the answer — "3 % loss" sends an engineer looking in both directions at once. The server already records every packet it received per sequence number (§6), so the two cases are actually distinguishable, and `train.udp_updown` now reports them separately: - sent, never seen by the server → **upstream** loss - seen by the server, reply never arrived → **downstream** loss Findings name the direction and say what is *not* implicated, which is half the value: `connectivity.loss_upstream` ("the return path is not implicated: replies came back for everything that arrived"), `connectivity.loss_downstream`, `nat.udp_unreachable_upstream`. Two things the implementation gets deliberately right: - **Downstream loss is measured against what reached the server**, not against what was sent. Using "sent" as the denominator counts every upstream loss a second time and overstates the return path. Pinned by a test with loss in both directions at once. - **Per-direction jitter without synchronised clocks.** Absolute one-way delay would need clock sync and we deliberately have none (the two-clock rule). But `server_rx − client_tx` carries a constant unknown offset, and differencing successive samples cancels it — so RFC 3393 one-way delay variation *is* honestly attributable to a direction even though latency is not. A test pins that a 10-second clock offset changes nothing. Correlation is by **wire sequence number**, which is not the loop index: the counter is shared with every other packet type on the session, so "the nth echo" is not "sequence n". `ProbeSession` now exposes `lastSeq`, including for a probe that was lost — a lost packet still has a sequence number, and that number is exactly what tells you which way it was lost. Live against fmr: 20/20 both ways, and jitter of **0.08 ms upstream vs 0.85 ms downstream** — a tenfold asymmetry that a round-trip measurement cannot see at all. 10 unit tests on the arithmetic (a wrong denominator here does not crash, it produces a plausible number pointing at the wrong half of the network) plus the live correlation check. ### frag_send: crafted IP fragments, so *ordering* is testable (server-v0.6.0, 2026-08-01) `big_send` with `df=false` answers one question — do fragments get through. It cannot answer the more interesting one, because the kernel always emits fragments in order, first one first. The classic middlebox fault is exactly about that ordering. Only the **first** fragment carries the UDP header, and therefore the ports; a stateful firewall or NAT that has not seen it has no flow to match the rest against, and many simply drop them. That is invisible to every in-order test, and in the field it looks like "large DNS answers fail on this network" or "the tunnel breaks when the MTU drops" — it works until the network reorders, then fails intermittently, which is the hardest kind of fault to chase. So the server builds the fragments itself (raw socket, `IP_HDRINCL`) and controls their order: `in_order` (baseline), `reversed` (last fragment first), `first_last` (first fragment held back 250 ms). The datagram is assembled and **signed whole** before being cut up, so what the client reassembles is indistinguishable from an ordinary packet — otherwise the test would be measuring our sender rather than the path. New test type `mtu.frag_ordering`; findings `mtu.fragments_blocked` and `mtu.fragment_reorder_sensitive`. Two details that would otherwise produce confidently wrong answers: - **The UDP checksum is computed, not left zero.** Zero is legal in IPv4 and would be less code, but zero-checksum datagrams are dropped by some middleboxes — and that drop would be recorded as a fragmentation failure, which is the wrong conclusion entirely. - **Fragment offsets are in 8-byte units**, so non-final fragments are rounded down to a multiple of 8. A 100-byte fragment is not an error; it is a datagram no host will ever reassemble. `frag-send` is advertised only when a raw socket can actually be opened — checked by opening one, because a permission model has more ways to say no (userns, seccomp, LSM) than a capability bit has to say yes. fmr runs as root with `cap_net_raw` in its bounding set, so it is available there. Fragment ordering runs only after `mtu.frag_delivery` shows fragments arrive at all; otherwise the three orderings would each report "not delivered" and read as three faults instead of one. The header arithmetic is unit-tested (reassembly coverage with no gaps or double-delivery, MF flags, shared IP ID, 8-byte offsets, checksum verification over odd and even lengths). Because the code is `//go:build linux`, the tests are **cross-compiled and run on fmr** — there is no Go toolchain there, so `go test -c` plus scp is the loop. Live against fmr: 4 fragments per burst, and all three orderings reassembled — a healthy path, and the baseline against which a mobile network will be interesting. ### Testing state (2026-08-01) Six live tests against fmr, all green, no device involved: `LiveServerTest`, `LiveMeasurement`, `LiveGranted`, `LiveDownstream`, `LiveUpload`, `LiveCompat`, `LiveEnrollment`. Plus 74 client unit tests and the full Go suite. Everything in the last several entries is verified from the PC; the app's UI (settings, history, deep-link enrollment) and `mtu.pmtud_up` remain device-only. ### throughput: a rate, plus the qualifier that makes it a measurement (server-v0.6.1 … v0.6.2) A throughput test reports the *smallest* limit on the path — and the sender's own ceiling is one of the candidates. If the server is asked for 50 Mbps and 50 Mbps arrives, the network was never the constraint and "50 Mbps" says nothing about it. So `perf.throughput_udp` always carries `limited_by` (duration | budget | rate | send_error) and `measures_network`, and a finding is raised only when the path is actually implicated. The live run against fmr reports 20 Mbit/s with `measures_network: false`, which is the correct and useful answer. Loss is computed against the **sender's own count**, fetched from the observations API, not against the requested rate. A receiver alone cannot tell "the network dropped it" from "the sender never sent it", and guessing turns a healthy server-side limit into a phantom network fault. The server keeps one summary per action rather than per-packet records — a ten-second run at 50 Mbps is half a million packets, and a struct each would turn a measurement into memory exhaustion. Sending is **paced**, on an absolute schedule. Unpaced would measure the server's NIC and the first queue it meets, then collapse into loss that reads as a network fault; sleep-per-packet would accumulate scheduler error and drift the rate down over a ten-second run. Throughput gets its own grant budget sized from the request, so every *other* action stays bounded at 8 MiB. When the byte cap binds before the clock does, the **duration is shortened and reported** rather than the run being truncated: promising thirty seconds and delivering twenty-one is the same information with a surprise attached, and it keeps "the clock ended the run" as the normal case — the only case where the rate is a clean property of the path. That behaviour came out of a test that failed honestly (30 s at 100 Mbps needs 375 MB against a 256 MB cap). It is **opt-in** in the run config, default off. A 5-second run at 50 Mbps moves ~30 MB; on a metered mobile connection that is the user's money, and a tool that spends it without being asked is not one people keep installed. #### The bug the live test found The first live run delivered 104 packets and stopped after 50 ms. The grant's rate check exempted the first 50 ms entirely, meaning to be lenient at startup — the effect was the opposite. A sender could dump an unbounded burst into that free window, and the instant the check switched on it compared those bytes against 50 ms worth of allowance and refused everything until real time caught up. **Every short test passed** (downtrain sends 50 packets, big_send seven); every sustained send died fifty milliseconds in. Replaced with a token bucket (`allowance = burst + rate × elapsed`), which is smooth from t=0. The burst is 100 ms of the allowed rate, floored at one ordinary datagram — deliberately one, since at 8 kbps a 64 KB floor is sixty-four seconds' worth, exactly the instant dump the ceiling exists to prevent. The pre-existing rate test caught that when I first tried the generous floor, and it was right to. Second half of the same bug: callers treated *any* refusal as terminal, so `TryAllow` now says why — a sender paces through a transient "too fast just now" and still stops dead on a spent budget or an expired grant. Both halves are pinned by regression tests. ### Findings registry (2026-08-01) Closes open item 1 of measurement-schema.md §9. A finding code is the stable, machine-readable half of a result — what a dashboard groups by and what someone greps a year of archived runs for — and that only holds if a code means exactly one thing forever. Ad-hoc string literals at fifteen call sites cannot promise that, and by the time the registry was written the failure had already happened. **Two emitters had independently produced `connectivity.downstream_loss` and `connectivity.loss_downstream` for the same claim**, and nothing anywhere objected. Anyone aggregating either one would have silently seen half their data. Merged into `connectivity.loss_downstream`, paired with `loss_upstream` so the two directions read as a set. **Two codes were also renamed out of `nat.*`.** `nat.udp_unreachable` is not about NAT — it means no replies came back — but the prefix determines the category, and the category determines which verdict light the finding rolls up into (§7.3). A `nat.*` code landing under *connectivity* is not a naming quibble; it changes which light turns red. Cheap to fix now, a breaking change later. Codes are now declared as typed `FindingSpec`s carrying their category and default severity, and emitters reference the spec instead of retyping the string — so a typo is a compile error and two call sites cannot disagree about a finding's category. `docs/findings-registry.md` is the contract, and a test reads it: it fails when the document and the registry have codes the other lacks, or when a severity differs. Documentation that drifts from its implementation is worse than none, because it still looks authoritative. The check scopes itself to table rows, so the prose can keep explaining which codes were retired and why. Six tests: uniqueness, declared-vs-listed, prefix↔category agreement, naming convention, a word-order-anagram check (the shape the duplication actually took), and the document agreement. ### A real privacy leak, found by starting on the machine-readable schema (2026-08-01) The intent was `measurement.schema.json` (§8's promised companion). The first step — checking whether the anonymizer actually covers the fields the schema declares as sensitive — found that it did not, so that became the work. **At the `balanced` level, five identifying values were being uploaded verbatim:** | value | field | why it matters | |---|---|---| | `2001:…::150` | `networks[].link.addresses[].addr` | the device's own global IPv6 address — a strong, geolocatable device identifier | | `2a02:…::1` | `networks[].link.routes[].gateway` | identifies the ISP allocation | | `203.0.113.77` | `networks[].link.dns.servers[]` | the configured resolver | | `nas.example.lan` | `private_dns_hostname` | an internal hostname | | `example.lan` | `search_domains[]` | the internal domain | The settings screen describes that level as pseudonymizing addresses. It was not. **Root cause:** classification keyed on field *names*, and the schema's actual names (`addr`, `gateway`, `dst`, `servers`, `search_domains`, `private_dns_hostname`) had never been added to the table. Not a subtle bug — just an unfalsifiable design. The existing tests all passed, because each one checked a field somebody had remembered to write a case for. **Two fixes, one of them structural:** 1. The missing names were added. 2. More importantly, a **shape-based backstop**: when a field name is unrecognised, the *value* is inspected, and anything shaped like an IPv4/IPv6 address or a MAC is treated as one. A name table can only protect fields someone thought of, which is precisely the wrong property for a privacy control. Hostnames are deliberately *not* inferred by shape — `train.udp_updown` is indistinguishable from a domain, and mangling a test type would corrupt the document to protect nothing. `LeakTest` is the new guard and is written to fail for fields nobody has considered: it plants identifying values wherever one can actually occur and asserts none survive, rather than checking a list of known cases. It also pins that RFC1918 addresses still come through readable, so the test cannot pass by over-redacting everything. Route prefixes and the unspecified address needed care in the transform: `0.0.0.0/0` and `::/0` must stay themselves, or a routing table becomes unreadable for no privacy gain. **Still outstanding:** `measurement.schema.json` itself. Worth noting what this episode implies for it — much of a document's payload lives in `evidence`/`metrics`/`params`, which are per-test-type `JsonObject` by design and therefore *outside* any schema. A schema-driven anonymizer would have less coverage there than the name-plus-shape one now does, so the schema should be built for validation and external tooling, not as a replacement for the classifier. ### ULA prefixes are pseudonymized whole (2026-08-01) Spotted in a real uploaded run from the phone: the server had `fda1:3fb1:ff92:6696::2662` for a DNS server. The general IPv6 path preserves the leading two groups (deliberately — for a global address that keeps the ISP allocation, which is the diagnostically useful part), and for a ULA that passed through **32 of the 40 random bits** of the global ID. ULA looks like the v6 equivalent of RFC1918 and the instinct is to treat it the same. That reasoning does not carry over, and the difference is the whole point: an RFC1918 prefix is shared by millions of networks and identifies none of them, while a ULA global ID is random and unique to one network by construction (RFC 4193). The prefix *is* the identifier — it is a network fingerprint that was surviving redaction. Now pseudonymized as a unit, so two addresses on the same ULA subnet still land on the same pseudonymous prefix: "these hosts are on one network" survives, "this is *that* network" does not. Three tests, one of which uses the exact value observed on the wire. Worth recording as a reasoning trap: I had originally raised this as "ULA should probably be kept verbatim, like RFC1918, for consistency". The surface analogy pointed the wrong way, and the correct answer was the opposite. ### Registry adopted everywhere; v6 findings renamed; Back works (2026-08-01) The findings registry was only adopted in `core-engine`. The app module still emitted seven codes as raw strings, so the registry test passed while codes existed outside it — including `ipv6.broken`, which fired on a real network and was in no registry at all. All seven now reference registry entries for code, category and severity, so those three cannot disagree at a call site. A grep for `code = "…"` across the app, engine and probe modules returns nothing. **`ipv6.*` → `v6.*`.** The third instance of rule 1: they declared `Category.IPV6` while the prefix map only knows `v6`, so `TestType.category("ipv6.broken")` fell through to *connectivity* and the finding rolled up under the wrong verdict light. The test-type registry already used `v6.`. Two severities reconciled while merging: - `connectivity.captive_portal` is **medium**, not high. The registry had guessed high; the probe that emits it had always said medium, and the probe was the considered value — a captive portal on hotel wifi is what should be there, and logging in clears it. `connectivity.no_internet` is the high one, because nothing the user does locally fixes that. - `v6.not_offered` is **info, and the registry says it must stay info**. Most networks still do not offer IPv6 and that is not a fault; a warning here lights a yellow verdict on a healthy network, which teaches people to ignore the light. Also: a `BackHandler` now returns from Settings/History to the run screen. The screen was a plain state variable with nothing connecting it to the back stack, so the system Back gesture left the app entirely. Enabled only when there is somewhere to go back to, so Back still exits from the run screen. ### Upstream throughput (server-v0.6.3, 2026-08-01) The mirror of the downstream case: the client generates the traffic and the server counts it. No grant is involved — the client is sending its own packets, so there is nothing to amplify — but it does need the server's tally, because **only the far end knows how much arrived**. Without that number a sender measures how fast it can *transmit*, which is usually just the speed of the local NIC and is a different question from the one being asked. `TYPE_THROUGHPUT_UP` (0x0F) is counted and deliberately **never answered**: a reply would double the traffic and drag the return path into a measurement that is specifically about the outbound one. The tally is a counter, not a list, and short-circuits **before** the observation log. A five-second run at 20 Mbps is around ten thousand packets; one struct each would turn a measurement into an allocation storm on a shared server, and nothing needs the per-packet detail since the client holds the send-side record. The gap between the two counts is the loss. `direction=up` on the throughput action sends nothing — it zeroes the counter, so a second run in one session measures itself rather than inheriting the first one's packets. The live test asserts `received <= sent`, which is what catches a counter that was never reset. Live against fmr: **3125 sent, 3125 counted, 0 % loss, 10.0 Mbit/s** at a 10 Mbit/s request, with `measures_network: false` — correct, since what arrived matched what was offered, so the path was never the constraint. ### Raw shell dumps leaked the whole LAN (2026-08-01) Found by running the Shizuku shell tier for the first time. The tier works — `tiers.shizuku: true`, `exec_path: UserService` (so the UserService binds on the OnePlus, as recorded), `runs_as shell(2000)`, 7/7 commands — and the run promptly uploaded **every MAC address on the local network** to fmr at the `balanced` level: router, phones, whatever else was on the wifi. Fourteen of them. The probes embed raw command output verbatim (`ip neigh`, `ip route`, `id`), which is genuinely good evidence and also a complete household device inventory. The anonymizer could not see it: classification is by field name and by whole-value shape, and `ip_neigh` is one long string that is itself neither a MAC nor an address. measurement-schema.md §9 item 2 had flagged raw dumps as "hard to anonymize" and proposed dropping them from exports; nothing enforced either. **Scrubbing beats dropping.** Identifiers inside any unclassified string are now replaced in place, using the same pseudonyms as everywhere else — so a MAC that appears both in a parsed field and in a raw dump still reads as one device. The dump stays readable and auditable: you can still see the neighbour table's shape, the host count, RFC1918 addresses and vendor prefixes. Dropping the evidence would have protected the same data while destroying the reason for collecting it. Two implementation notes worth keeping: - **One pass, not three.** Sequential passes re-process their own output: once a MAC became `78:9a:18:xx:yy:zz`, the IPv6 pattern matched it — six hex groups separated by colons *is* an address — and destroyed the vendor prefix the MAC rule had just preserved. Ordered alternation resolves each position once, MAC first. - The patterns are conservative on purpose. A missed address gets caught by another rule or not at all; an over-eager one mangles timestamps and version strings, corrupting evidence to protect nothing. `RealDocumentTest` runs the anonymizer over a captured run when `ECHOLOT_REAL_RUN` points at one, and fails on any MAC that survives. It self-skips otherwise, so no one's network is committed to the repo. Against the actual leaked document: **14 MACs in, 0 surviving.** Also fixed: the Settings *Preview what an upload would send* button did nothing. It read `UiState.history`, which is empty until the History screen has been opened — the same root cause as the "0 run(s)" count. It now reads the archive directly, and says so when there is nothing to preview rather than silently ignoring the tap. ### Security: the admin listener was publicly exposed for ~15 minutes (2026-08-01) Moving the admin listener to `[::2]:443` for the UI exposed `/admin/enroll-tokens` and `/admin/selftest` to the internet **with no authentication**. Anyone who could reach `fmr.echo-lot.app` could mint enrolment tokens. The listener was designed localhost-only — its own flag help says *"keep localhost"* — and that assumption travelled with it when the address changed. The compounding error: `checkAdminExposure`, added the same day, verifies **encryption** and says nothing about **authentication**. It passed, and a green light on an adjacent property is worse than no check, because it invites you to stop looking. Closed by returning to loopback (the TLS and ACME work is retained, just not exposed). All 68 device enrolments matched the timestamps of test runs, so there is no evidence of abuse — but the window existed on a freshly published hostname and absence cannot be proven. 39 unused enrolment tokens were purged, since any could have been minted by someone else and they cost nothing to replace, and 63 test devices removed. **The admin listener does not become reachable again until it authenticates.** That reorders the UI work: auth on the listener first, everything else after. ### Open: encrypted uploads, where the operator cannot read the data Not built. Recorded because the shape is decided by a few early choices, and the current design happens to leave the door open. The goal: hand someone an account, let them upload, and be unable to read what they uploaded. Sketch: a random per-account **master key**, generated on the first device and wrapped under a key derived from a passphrase (PBKDF2-HMAC-SHA256 — stdlib on both sides). The wrapped key is stored server-side as an opaque blob, so a new device signs in, fetches it, and unwraps locally; the server never sees either key. Runs are encrypted client-side with AES-256-GCM, fresh nonce per run. All of this is stdlib in Go and `javax.crypto` in Kotlin — no dependency either side. Four consequences that decide whether it is worth it: 1. **What stays readable determines what the UI can do.** The server builds its index by *parsing* the document — verdict, finding count, started_at. An opaque payload means the client supplies that metadata or the index disappears, and with it retention-by-verdict and any "runs with findings" view. The honest version supplies only run id, timestamp and size, and moves the rest client-side. 2. **Lose the passphrase, lose the data.** That is the feature working, and also the support burden. It needs a recovery code printed at setup, not a reset flow — there is nothing to reset. 3. **Metadata is not hidden.** The operator still sees which account uploaded, when, how often and how large. "Cannot see it" is about content, not existence, and saying otherwise would oversell. 4. **It makes `min_anonymization` unenforceable** — a server cannot check a level it cannot read. That is not a conflict so much as a redundancy: the anonymization floor exists to protect the user from the operator, and encryption does that better. The two should not both be demanded of one upload. What keeps this possible: uploads are already stored byte-for-byte as received, and every index field is derived in one function (`runs.Put`). The thing to avoid is admin features that *require* reading content — those would have to be unbuilt later.