Compare commits

...
Author SHA1 Message Date
mrambossekandClaude Opus 5 9ee7d554a6 app: listen to what the segment says unprompted
server-release / image (push) Successful in 14s
server-release / release (push) Successful in 31s
Four passive collectors join long mode: SSDP (passive NOTIFY plus paced
M-SEARCH from the capture socket, so unicast replies land in the same
capture), LLMNR, NetBIOS-NS and WS-Discovery. All are periodic and sparse
by nature, which is exactly why they belong to a window rather than a
probe - a thirty-second run mostly hears silence and would report an
empty network as confidently as a quiet one.

NetBIOS reports unsupported on the app tier and says why: UDP 137 is
privileged. The decoder and evidence shape are tested and waiting for the
Shizuku tier; a recorded reason beats a missing test.

Silence without a multicast lock or a group join is PARTIAL, never OK -
that case is a fact about this app, not about the network. Hostnames,
banners and device UUIDs are classified in core-privacy so the anonymizer
treats them like every other identifier rather than letting a neighbour's
device model ride out in an upload.

No active WSD Probe and no NBSTAT sweep: the app listens to what a
network broadcasts, it does not announce itself to strangers or
interrogate its neighbours.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 14:58:41 +02:00
mrambossekandClaude Opus 5 93125a4b1d app: 0.3.0 — run modes and the dev relay
server-test / test (push) Successful in 37s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 14:46:25 +02:00
mrambossekandClaude Opus 5 0071e00003 app: long runs that watch, and a relay for the port that keeps moving
Long mode starts listeners at t=0 and keeps them running past the
battery: a network-change watcher that finally fills networks[].changes[]
(defined since the schema's first draft, never populated), an RSSI log, a
ping series giving loss and jitter over minutes, and mDNS listening for
the whole window. This is the class of fault a short run cannot see - a
link that drops for four seconds between two probes is reported healthy
by both of them. run.mode records which question was asked, because
silence means different things in the two modes.

The adb relay replaces the retired beacon: AdbRelay watches adbd's own
mDNS with the resolve-once discipline the beacon learned the hard way
(resolving re-arms adbd and pops a notification), a foreground service
keeps it alive with the screen off, and the heartbeat re-posts the cached
endpoint rather than re-resolving. It exists because mDNS does not cross
subnets and the wireless-debug port rotates every few minutes.

Also records why LLDP/CDP cannot follow SSDP into long mode: both are raw
L2 frames, so they need CAP_NET_RAW - root tier, not app, and Shizuku's
shell user does not have it either.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 14:43:33 +02:00
mrambossekandClaude Opus 5 ae63bd7c7f server: relay a test device's adb endpoint, because mDNS does not cross subnets
The beacon this replaces was a separate service wildcard-bound to
0.0.0.0:443 - it silently occupied port 443 on the reserved measurement
addresses, voiding the IPv4 interception proof for as long as it ran, and
it accepted a port report from anyone who could reach it. So this lives
where the repo's own post-mortem said it belongs: POST on the control
plane authenticated by the device credential, GET on the admin UI behind
the existing apiAdmin helper. No new listener, no new port, no wildcard.

Entries expire after 24h (ECHOLOT_ADB_ENDPOINT_RETENTION_H) on both write
and read - a LAN address is a breadcrumb for driving a test device, not
measurement data worth keeping.

Also records the BLE peer-comparison design: the case for it is that BLE
is out-of-band, which is what makes client isolation measurable at all -
silence over IP cannot distinguish an isolating AP from an absent peer,
and a peer confirming out-of-band that it was listening turns that
silence into proof.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 14:31:01 +02:00
mrambossekandClaude Opus 5 ab6e278272 app: an IMS network is not a blocked measurement
rmnet_data1 on the OnePlus is the carrier's IMS/VoLTE network: it carries
IMS&MMTEL but neither INTERNET nor NOT_RESTRICTED, so binding needs
CONNECTIVITY_USE_RESTRICTED_NETWORKS - signature-level, unobtainable.
EPERM there is permanent and says nothing about any network's health, but
the constraint detector counted it, so a phone with VoLTE reported
'measurement blocked' on every run and every verdict came out
INCONCLUSIVE. Capabilities now decide: networks an app may never bind are
recorded as app_usable:false in networks[] and skipped, rather than
reported as something the run failed to do.

Also retires the 'while it tears down' wording, which was a guess that
turned out to be wrong - the refusal outlasts the VPN indefinitely, so
the text now names the VPN as the usual cause without asserting a
timeline it cannot know.

App version 0.2.4.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 14:11:08 +02:00
mrambossekandClaude Opus 5 d420a71a01 docs: v0.11.3 live on fmr, trains validated 120/120; two mint bugs logged
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 14:04:46 +02:00
mrambossekandClaude Opus 5 d621611bdd docs: the v0.11.2 mystery was an unstamped-tag problem, not lost code
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 13:54:42 +02:00
mrambossekandClaude Opus 5 c5f3c2e7af app: learn probe durations per device; stop claiming VPN after teardown
server-release / image (push) Successful in 15s
server-release / release (push) Successful in 31s
The ETA table was fixed at compile time, but real durations are a property
of this phone and the network it stands in - ICMPv6 answers in
milliseconds where IPv6 works and waits out its timeout where it does
not. Estimates now prefer an EMA (70/30) of what this device actually
measured, seeded by the old constants on first run.

VPN detection had two lies in it, both found on hardware: a disconnected
tunnel lingers in allNetworks while tearing down, so 'VPN active' is now
judged from the ACTIVE network only; and any bind failure counted as
'per-network blocked', so a network dying mid-run flipped a healthy run
to INCONCLUSIVE - now only a genuine EPERM refusal counts. Banner and
finding split three ways (VPN + blocked / blocked only / VPN only) so the
app never names a VPN the user just turned off.

App version 0.2.3.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 13:41:10 +02:00
mrambossekandClaude Opus 5 3214cc877a app: parse the Shizuku dumps into link.ra_source and sec.arp_watch
Pure-Kotlin parsers over the captures the battery already makes, tested
against the real archived dumps from both devices - including the
Lenovo's 1000000015 route tables (table ids stay strings), the OnePlus's
stray uid=2000 prefix and mid-line hoplimit, and an ip_monitor that only
ever said EXEC_TIMEOUT. ShizukuProbe now returns three tests; a failed
capture is SKIPPED with its reason so no test type silently vanishes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 13:34:52 +02:00
mrambossekandClaude Opus 5 515a6aef04 app: speak train (0x03-0x05) and mark downtrains with DSCP
ProbeSession sends paced upstream trains and fetches the server's columnar
received view; UpstreamTrainMeasurement lines both ledgers up per sequence
number into TrainEvidence - the directional loss attribution a round trip
cannot make. Loss is computed against the server's total count, not its
row list, so a buffer-capped report can never invent loss on big trains.
A missing report stays ambiguous by name (report lost, or server predates
trains) instead of being blamed on the train.

downtrain actions can now request a DSCP marking, recording dscp_applied
so a survival comparison never blames the path for a marking the sender
skipped.

Also logged: fmr runs v0.11.2 built from commits this repo's remote never
saw, so --self-update on fmr is OFF-LIMITS until that lineage is repaired
- the fresh server-v0.9.2 release is newest-created but semantically
older, and the updater compares strings, not SemVer.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 13:34:52 +02:00
mrambossekandClaude Opus 5 d574c76630 docs: log the prober fold and server v0.9.2 feature work
server-release / image (push) Successful in 16s
server-test / test (push) Successful in 48s
server-release / release (push) Successful in 1m2s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 13:06:09 +02:00
mrambossekandClaude Opus 5 8118e213ae server: upstream trains, observed TTL/DSCP/ECN, rate limits, action ids
Types 0x03/0x04/0x05 land with a bounded columnar train buffer (head kept,
truncation declared) and grant-free multi-part reports - a report row is
smaller than the packet it answers, so $3.4 holds without a grant. The
read loop now collects TTL/TOS cmsgs on Linux, replacing the 0xFF stubs in
the observation block with what the kernel saw; downtrain gained a dscp
parameter, so DSCP survival is measurable in both directions.

Rate limiting ($2.5) exists now: per-credential AND per-source buckets,
429 on the control plane, silent drop on the data plane after the HMAC
gate and before the replay window. UDP ceilings default above the largest
legitimate run - a limit that clips a real measurement produces a
confidently wrong number.

Every granted packet carries its action_id at payload[8:16]; overlapping
actions were unattributable before. Canary DNS logs now honor the stated
24h privacy default. /admin/enroll-tokens answers the spec's JSON shape.
protocol_version 1.0.1 (additive).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 13:04:54 +02:00
mrambossekandClaude Opus 5 f6849f8e6a app: fold the prober's validated capabilities into core-probe
traceroute.udp4 lands as TracerouteProbe: ICMP time-exceeded read off the
socket error queue via Os.recvmsg(MSG_ERRQUEUE) through the reflection
facade the prober validated on both devices - no root, no raw socket, and
the C-over-JNI shim stays dead. Emits the schema's TracerouteEvidence with
rtt_ns. OsAbi carries the hardcoded sockopt ABI numbers across, including
the measured fact that Os.getsockoptInt exists on neither device, so PMTU
must come from the errqueue, never getsockopt(IP_MTU).

local.mdns_inventory lands as MdnsInventoryProbe with both hardware-bought
lessons intact: the meta-query lies (0 results beside live services on both
devices), and 4s of listening misses what 10s catches.

App version 0.2.2.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 12:39:52 +02:00
mrambossekandClaude Opus 5 7d98a43866 docs: close the VPN, v6, reserved-probe and signing items in the build log
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 12:32:02 +02:00
mrambossekandClaude Opus 5 a49bef5821 server: refuse unsigned releases and polluted reserved addresses
Self-update now verifies SHA256SUMS.sig (ed25519, relsign package) against
a public key baked into the binary; the private key exists only in the CI
secret store, so a compromised release host can withhold updates but not
inject one. CI signs on every server-v* tag and hard-fails without the
secret. Operators with their own pipeline override the key via
ECHOLOT_SELF_UPDATE_PUBKEY (mint a pair with release-sign -gen).

Startup also now proves 80/443 are actually free on the reserved
measurement addresses by asking the OS (throwaway bind), not the config -
CheckReserved could never see a stray process, and the adb-beacon receiver
on 0.0.0.0:443 was exactly that.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 12:32:02 +02:00
mrambossekandClaude Opus 5 20cfecf566 app: name what a VPN blocked, and prove v6 broken before saying so
Constraints are detected up front (one throwaway bind per network) and land
in run.constraints, a measurement.vpn_constrained finding, the $7.3 verdict
(INCONCLUSIVE outright) and a banner on the run screen - a VPN'd run looked
exactly like a clean run of a healthy network before this.

v6.broken returns to the registry now that it can be earned: V6ConnectProbe
(v6.brokenness) makes a real TCP connection over IPv6 to the enrolled
server, and only both transports failing on a network that advertises IPv6
justifies the claim. TCP succeeding turns the finding into 'ICMPv6 is
filtered, IPv6 works' at high confidence instead.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 12:31:44 +02:00
mrambossekandClaude Opus 5 987b2ceb47 server: read the verdict from where the schema puts it
Every uploaded run showed "not recorded" in the web UI because the meta
extractor read summary.verdict. Schema §7.3 calls that field
summary.overall; "verdict" is the per-category field one level down. So
the verdict was never stored, and the UI faithfully reported a gap that
was this parser's doing rather than the document's.

The test encoded the same mistake — its fixture posted
summary.verdict:"warn" — so it passed throughout against a parser that
read a field nothing writes. Corrected to summary.overall, and to a
verdict that exists: §7.3 defines green|yellow|red|inconclusive, and
"warn" was never one of them.

The eleven runs already stored had their meta backfilled from the
documents, which are kept byte-for-byte and still carry the real value.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 10:52:25 +02:00
mrambossekandClaude Opus 5 d5b1bab577 app: stop the autorun upload pointing at a listener that is gone
"upload failed: HTTP 400 client sent an HTTP request to an HTTPS server"
is a TLS listener rejecting cleartext, and the cleartext was ours:
REPORT_UPLOAD_URL was http://89.185.109.150:443/report, which used to be
the adb-beacon receiver holding 0.0.0.0:443 in plaintext. Disabling that
receiver and giving 443 to echolot-server left this posting plain HTTP at
a TLS port.

Blanked rather than repointed. The endpoint existed so an unattended run
could be collected without adb, and autorun reports are now read straight
off the device with `run-as cat` — so it was buying nothing and emitting
an alarming error for a debugging convenience. Deliberately not aimed at
/v1/runs either: that is the consent-gated upload, and a debugging
shortcut must not be able to satisfy it by accident.

The message says what happened instead of implying something broke.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 10:36:49 +02:00
mrambossekandClaude Opus 5 a17c3fd9e6 app: catch a search domain that swallows DNS queries
A tablet on a healthy network could not resolve anything. The DNS server
answered the bare name correctly — NOERROR, two records, A and AAAA, with
and without EDNS0 — so the earlier finding blamed the device's resolver.
It was wrong. The network advertised hudelist.local as a search domain and
the server silently dropped every query under it: not NXDOMAIN, nothing at
all. Resolvers append search domains, so they waited for a reply that was
never coming.

Silence is the part that makes this vicious. A negative answer moves a
resolver on; no answer looks like packet loss, so it retries, and some
give up on the lookup entirely. It also explains how two devices on one
network can disagree about whether DNS works — the phone tried the plain
name first and never noticed.

The probe now asks about a nonce name under each advertised search domain,
where the wanted answer is NXDOMAIN and only silence is a fault. The
finding is ordered ahead of dns.system_resolver_broken so the two cannot
both fire: without that, this exact network gets told its device is
broken.

Severity follows the harm rather than the shape. HIGH when resolution is
actually failing, MEDIUM when the domain is a black hole but this resolver
happens to try the plain name first — calling that HIGH would be crying
wolf on a network that works. The message names the fix and notes that
.local is reserved for mDNS by RFC 6762 and widely dropped by design,
while home.arpa (RFC 8375) is the name reserved for this.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 10:29:04 +02:00
mrambossekandClaude Opus 5 cfa58e8d60 app: a DNS reply is not a DNS answer
The probe counted any well-formed packet from the server as "the server
answers", checking only the transaction id and a minimum length. REFUSED
and SERVFAIL are well-formed packets. So a server actively refusing this
client would have been reported as healthy, and the finding — whose whole
output is "the network is fine, your device is not" — would have pointed
confidently at the wrong component.

It now requires rcode 0 and at least one record, and reports a refusal as
what it is: a working server saying no, which points back at the network.
The rcode is named rather than numbered, because "REFUSED" is a fact an
operator can act on and "rcode 5" is a lookup.

Caught by decoding what fmr's router actually replied — ab cd 81 80 00 01
00 02, NOERROR with two answers — after realising the earlier check only
counted bytes. The reply was genuinely good, so the finding on the tablet
stands; the check was wrong regardless.

The advice is broader too. That tablet's fault survived a reboot, which
makes "toggle wifi and it clears" wrong as a flat claim: it now says what
to look at when a restart does not fix it — something on the device
filtering DNS, or a per-client rule on the router.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 10:10:36 +02:00
mrambossekandClaude Opus 5 2ed4d1f478 Reach the server by address when its name will not resolve
A measurement tool that cannot report from a broken network is useless
exactly when it matters, and a wedged resolver is one of the faults this
app is built to find — it should not also be the thing that stops the
finding being delivered. The profile already carries the server's
addresses; they are now kept and used when the name fails.

Safe because the pin is the trust and the name is not part of it: the
server presents the same certificate however it was reached, and a wrong
address fails the pin like anything else. Only the primaries are cached —
the alternate pair exists for NAT behaviour discovery and does not carry
the control plane, so falling back to one would fail for a second,
unrelated reason.

Substituted only when the name genuinely does not resolve, and only after
checking the candidate answers on the port: on a v4-only network a v6
address would otherwise be chosen and fail slowly, which is the wrong
answer delivered late.

The server had to meet it halfway. Sharing 443 by SNI meant a client
arriving by IP sent no server name and got the Let's Encrypt certificate,
failing the pin. A numeric host — or no SNI at all — now selects the
pinned certificate and routes to the control plane. That is sound because
the admin UI is only ever reached by name: browsers always send SNI, and
nobody bookmarks an IP for a site with a CA-issued certificate.

Verified against fmr: by IP on both families the served pin is the
control one and /v1/profile answers 401, while fmr.echo-lot.app still
serves the Let's Encrypt certificate and the admin UI.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 10:03:14 +02:00
mrambossekandClaude Opus 5 d65dbbc75a app: tell a broken device resolver apart from a broken network
Diagnosing a tablet that claimed "no internet" took twenty adb commands
to establish something the app should have said in one run: ping to
1.1.1.1 worked, the configured DNS server answered a raw UDP query in
65 bytes, and Android still could not resolve a hostname. The network was
fine; netd had wedged.

Those two failures look identical to a user and want opposite responses —
"look at your router" against "toggle your wifi" — so dns.resolver asks
the network's own servers directly and compares the answer against what
the platform returns for the same name. The query is hand-rolled over a
plain DatagramSocket on purpose: anything routed through a resolver API
would inherit the very fault being looked for.

dns.system_resolver_broken fires only on the pairing that is otherwise
unattributable: server answered, platform did not. Per network, because a
phone can have wedged wifi and working cellular at once.

Also records Android's own verdict per network — validated, captive
portal, partial connectivity — which the app reproduced with its own HTTP
probes but never stored. It is free, it is what the user sees in the
status bar, and its disagreement with our measurements is exactly what
identified the tablet. NET_CAPABILITY_PARTIAL_CONNECTIVITY is @SystemApi
so the constant is inlined with its rationale, in the manner of OsAbi.kt,
and read defensively enough to report unknown rather than false.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 09:44:49 +02:00
mrambossekandClaude Opus 5 40e76c52ca adminui: show the enrolment link as a QR code
Enrolling a device that cannot reach the admin UI meant transcribing a
200-character link with a base64 pin in it — the step the link format
exists to avoid, and the one where a pin wrong by one character fails
later as an inscrutable TLS error.

Rendered as inline SVG rather than a PNG data: URI, because the page's
CSP is default-src 'none' and means it: a data: image would need img-src
opened, markup needs nothing. One path rather than a rect per module,
since a link this long encodes to about 60x60 and two thousand elements
is a lot of DOM for a picture of a square. It is generated from the same
validated value as the href, so a rejected link produces neither.

This relaxes the stdlib-only rule, deliberately and recorded in CLAUDE.md.
The rule bought one self-contained binary with no supply chain to audit,
which one small pure-Go package barely dents; F-Droid never applied to
the server, only the app ships there. A correct QR encoder is ~500 lines
of Reed-Solomon that nobody should be hand-writing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 08:48:02 +02:00
mrambossekandClaude Opus 5 5b02d40802 app: notice when Shizuku is authorised
The banner watched only for the binder arriving or dying, and granting
permission does neither. So after the user tapped the banner, answered
the dialog and came back, it still read "running but not authorised" —
wrong at precisely the moment they were looking for confirmation that it
had worked.

Two listeners were missing, because there are two ways this changes.
Answering our own request now fires OnRequestPermissionResultListener.
That is not enough on its own: permission can equally be granted inside
Shizuku's own app, and Shizuku started or stopped there, none of which
calls back into this process — so the state is re-read whenever the
screen comes forward, which is the only thing that covers every route.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 08:37:45 +02:00
mrambossekandClaude Opus 5 416b783647 app: lay the server facts out in real columns
The block was rendered as one padded string, which only lines up in a
monospaced font — and the monospace never took, so the values sat at
ragged offsets and the point of the list was lost. Padding text to fake a
table makes the layout depend on a typeface decision made elsewhere.

It is a label column of fixed width and a value column that takes the
rest, so the addresses align whatever the font does and a long value wraps
inside its own column instead of under the labels. The capability list
goes back to plain comma-separated text: hand-wrapping it at three per
line was working around the same missing alignment.

Placement too. The facts now sit under the "Check server" button that
fetches them, rather than among the input fields, and the sentence
explaining the endpoint sits directly under the URL field it describes —
it had ended up orphaned between the two, reading as a comment on nothing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 08:24:00 +02:00
mrambossekandClaude Opus 5 c4f2a10790 app: show what the server reports as facts, not as inputs
The settings card offered three editable boxes and said nothing about the
server itself — which addresses a test will actually use, on which ports,
what it can measure. That is the part a person checks before trusting a
result, and "which address did this come from" is precisely the question
a report leaves open.

The server now publishes it. The profile's targets carried one IPv4 and a
TODO; it reports both families and both alternates, derived from the UDP
listen spec rather than configured separately, so the list cannot drift
from what is actually bound. No reservation means no alternate is
claimed: announcing a second address as the RFC 5780 alternate when none
was set aside would promise a redirect the server will not send.

The app renders them read-only, in a panel visibly distinct from the
fields above. An editable box that changes nothing is worse than no box,
and these are facts to read rather than settings to apply.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 08:14:42 +02:00
mrambossekandClaude Opus 5 fe4ec23ba1 app: show the server by the name its operator handed out
The Server URL field showed the endpoint the app dials, which after
discovery is not the name anyone was given — enrolling against
fmr.echo-lot.app left the settings reading fmr-1.echo-lot.app, with no
explanation of where the -1 came from.

It shows the public name now. The endpoint is not hidden, just demoted to
a line beneath that says where the connection actually goes and why the
two differ: a network engineer debugging a failed connection wants that,
and burying it would trade one confusion for another.

Typing a URL by hand sets both, since there is no discovery to consult in
that case — setting only the public one would leave the app still dialling
the previous server, which is the kind of half-applied change that fails
much later and somewhere else.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 08:07:55 +02:00
mrambossekandClaude Opus 5 33799b8135 adminui: the enrolment link was rendered as a dead anchor
html/template rewrites an href whose scheme it does not recognise to
"#ZgotmplZ". echolot:// is not on its list, so "Open in the Echolot app"
was not a link at all — tapping it did nothing, and nothing showed it:
the markup reads correctly, the app resolves the scheme, and only the
sanitised attribute in the served HTML gives it away.

Marking the value template.URL opts out of that sanitising, which is only
safe because the shape is now checked first. The link arrives in a query
parameter, so without the check a crafted /devices?link=javascript:… would
put a script URL into the page for an admin to click.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 08:04:56 +02:00
mrambossekandClaude Opus 5 f6e093944c build: document the poisoned build cache; compare like with like
An entire module was missing from the APK. The app died with
ClassNotFoundException for app.echo_lot.protocol.EnrollmentLink while the
build was green, the module's jar was correct, and :app:dependencies
listed it on debugRuntimeClasspath — its code simply never reached AGP's
intermediates.

The cause was a poisoned Gradle build cache entry, which is why nothing
obvious fixed it: clean, rm -rf */build and --rerun-tasks all leave the
build cache alone. Only --no-build-cache did. Every app build made in this
session shipped without core-protocol, so enrolment, sign-in and upload
would all have crashed identically; several hours of "the tap does
nothing" were this, misread as a UI problem.

CLAUDE.md now carries the symptom, the fix, and the verification —
grepping the dex for a string literal only that module defines, because
grepping for a class *name* proves nothing: callers carry the name as a
reference whether or not the class is packaged. That false check is what
let me believe an earlier rebuild had fixed it.

Also: the enrolment dialog compared the stored endpoint against the link's
public URL, so re-enrolling with the same server announced itself as a
move to a different one. Those are deliberately different strings now that
discovery exists; the comparison uses the public name on both sides.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 08:01:58 +02:00
mrambossekandClaude Opus 5 bc384531e6 app: enrollment feedback beside its own button
The enrollment result was written to the same status line as everything
else, which renders at the far end of the server card below three text
fields — and on a fresh install renders nowhere at all, because that line
only appears once a run exists. So enrolling looked identical whether it
worked or not.

It has its own line now, directly under the Enroll button that caused it,
and its own state rather than sharing one with "Check server": two
actions, two results.

The server fields were also re-read the instant the button was pressed,
before the enrollment coroutine had done anything, so they showed the
previous server's values. They now refresh when the result lands, which
is the point at which there is something new to show.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 00:48:36 +02:00
mrambossekandClaude Opus 5 082a2314ef protocol: enrollment links carry the public name, not the endpoint
Sharing port 443 between the admin UI and the control plane forces two
hostnames — one port and one name is one certificate, and the two need
different ones. That difference had been leaking into every enrollment
link, so an operator handed out fmr-1.echo-lot.app when the thing they
and their users know is fmr.echo-lot.app.

The link now carries the public name and the app asks GET /v1/discover
where to actually connect. The endpoint is plumbing: it exists to select
a certificate, and nobody needs to see it.

Discovery hands out an address and never a pin. The pin stays in the
link. Fetching it over an ordinary TLS connection would make pinning
worth exactly what the certificate authorities are worth, and pinning is
there to survive one the operator does not control — a root injected by
corporate device management, say, which is unremarkable on the networks
this tool gets pointed at. With the pin pre-shared, an intercepted
discovery can only send a device somewhere the pin will not match: an
outage, not a compromise.

Optional on both sides. A server that does not answer, or a link that
already names the control endpoint, works unchanged — enrollment must not
start failing because a lookup did.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 00:13:27 +02:00
mrambossekandClaude Opus 5 a720e84411 app: one activity for deep links, and say what re-enrolling actually does
MainActivity had no launchMode, so every echolot:// link stacked a fresh
activity with its own ViewModel. The enrolment then ran in a throwaway
copy and pressing back returned to the original screen showing none of
it — silent, and indistinguishable from the link not working at all.
singleTask plus onNewIntent means the link reaches the screen already in
front of the user.

The confirmation dialog also read as nonsense when re-enrolling with the
server already configured: "already enrolled with X ... enrolling with X
replaces that". Naming one URL twice looks like a bug and buries the
consequence that does apply — the credential is replaced, the old one
stops working at once, and the device appears on the server as a second
entry next to the first, which is worth revoking afterwards.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 23:54:08 +02:00
mrambossekandClaude Opus 5 30d9501507 app: confirm before an enrollment link replaces an existing one
The scheme was already registered and the deep link already worked — it
enrolled on arrival, with no confirmation. Now that the web UI offers the
link as something to follow, that is one tap between a working enrollment
and a replaced one, from a page that might be showing a link minted for a
different device entirely.

Enrolling is not additive: the new credential replaces the old, and on
the previous server this device simply stops reporting. So the link is
held and the user is asked, with both server URLs named — the question is
"which server", and it cannot be answered without seeing both.

The dialog says what survives, because that is the part someone hesitates
over: uploads already on the old server stay there, runs stored on the
phone are untouched, and the device reappears on the new server as a new
device rather than carrying its history across.

Enrolling also clears the cached canary zone. It describes the old
server's deployment, and querying it against the new one would measure
somebody else's zone and file the answer under this network.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 23:42:37 +02:00
mrambossekandClaude Opus 5 c639d58043 adminui: make the enrolment link followable on the phone being enrolled
The link was printed as text next to an adb command, so enrolling a phone
meant getting a 200-character string from one device onto another —
which is the step that goes wrong, and it does not have to happen at all.
The app registers the echolot:// scheme, so on the phone being enrolled
following the link is the entire procedure.

The raw string and the adb form stay for every other case: reading it on
a laptop, or enrolling a device that is not the one holding the browser.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 23:39:55 +02:00
mrambossekandClaude Opus 5 6d042a9d89 server: control plane shares port 443 with the admin UI
Captive portals, hotel wifi and corporate firewalls routinely permit only
80 and 443 — exactly the networks this tool exists to diagnose. A control
plane on 8443 is unreachable precisely when it matters most, and it fails
as "cannot reach server", which tells the user nothing about why.

The two cannot share a certificate, so sharing the port needs two names.
The control plane is trusted by SPKI pin and uses a long-lived
self-signed certificate; a browser needs one a CA vouches for. One name
on one port is one certificate. Pinning the Let's Encrypt key instead was
considered and rejected: it survives renewal only while key reuse holds,
so a routine key rotation would brick the fleet.

One listener now picks the certificate by SNI and the handler by Host.
Both have to agree, or a client gets the pinned certificate with the
admin UI behind it.

8443 stays open. Devices enrolled before this carry that URL in their
settings, and closing it for the sake of a port number would strand every
one of them; it can go once nothing points at it.

Verified per SNI on 443: fmr.echo-lot.app serves the Let's Encrypt cert,
fmr-1.echo-lot.app serves the self-signed one whose pin is unchanged, and
/v1/profile answers 401 on the control name against 303 to the login page
on the UI name.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 23:25:23 +02:00
mrambossekandClaude Opus 5 4aaaa5f5d4 app: probe the server this device is enrolled with, not ours
The canary-DNS zone and the STUN host were compiled in as c.echo-lot.app
and fmr-1.echo-lot.app, so every copy of the app measured against this
particular deployment whatever server its owner had enrolled with. On
someone else's install those two tests describe our infrastructure and
report the result as a fact about their network.

The zone comes from the server's own profile, which has advertised
canary_zone all along — the app simply never read it. It is cached in
settings because the canary probe runs at device tier, before anything
has contacted the control plane, and a probe that had to make a call
first would fail on exactly the networks worth measuring. The STUN host
is derived from the configured server URL rather than stored, since a
second copy of the server's name goes stale the moment someone
re-enrolls elsewhere.

With no server configured both now report SKIPPED. StunProbe previously
would have reported FAILED on a blank host, which reads as a finding
about the network when the truth is that no packet was ever sent — the
same conflation between "measured nothing" and "measured a fault" that
the ICMPv6 finding had.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 23:20:52 +02:00
mrambossekandClaude Opus 5 e0428b4c84 server: enroll against fmr.echo-lot.app; mint links from the CLI
The control plane now advertises the same name the web UI answers on.
That is safe because the client authenticates by SPKI pin and explicitly
does not verify the hostname — "pin is the trust, not the name" — so no
certificate covers or needs to cover either name.

The per-host name still means something, though, and the rule it encodes
has to survive: pinning binds a client to one server's key, so fmr may be
a CNAME to exactly one host and never a multi-address service record. A
second server gets enrolled as fmr-2 explicitly, because a client that
reaches a different key does not fail over, it fails.

Minting a link was broken and had been since the authenticated admin UI
replaced the old admin API: enroll-link.sh still posted to
127.0.0.1:8444/admin/enroll-tokens, an endpoint that no longer exists on
a listener that no longer binds loopback. Rather than add a second
unauthenticated door — which is how the old one ended up briefly reachable
from the network — the binary mints its own link. Whoever can run it
against the state directory already holds every privilege the server has,
so authenticating them to themselves would be theatre.

EnrollmentURI is shared with the running server's EnrollmentLink rather
than reimplemented. Two copies of that encoding would eventually disagree,
and the failure mode is a pin that looks right and surfaces as an
inscrutable TLS error rather than as a bad pin.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 23:16:23 +02:00
mrambossekandClaude Opus 5 8001234e8b adminui: render the self-test instead of dumping it
The box under "Self-test" was `%+v` of a Go struct on one line, read
through a horizontal scrollbar — on a phone you could see about six words
of it, from the middle. The design pass had polished the frame around it
and left the contents a debug dump.

The report was structured the whole time: each sysctl check carries the
name, what was found, what was wanted, a severity, and a sentence
explaining why the setting matters to measurement. All of that was being
flattened into one string. It now renders as records like everything
else, with the explanation set as prose across the full row, because it
is a sentence and not a fourth column.

Two faults the render caught: the desktop row grid applied to every
readout, so the standalone summary panel was chopped into four narrow
columns and "full 1500" broke into "ful/l/150/0"; and a fixed first
column wrapped `net.ipv6.conf.all.accept_ra` mid-word. The grid is now
scoped to readouts inside a row, and the label column may grow to 18rem
before it wraps.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 23:04:55 +02:00
mrambossekandClaude Opus 5 332b6f3589 adminui: give the admin UI the instrument's own visual language
Echolot is the German word for an echo sounder — an instrument that emits
a ping and reads what comes back — and the UI now looks like one instead
of like a dashboard. The three big-number stat cards went first: that
layout is the stock answer for any admin page, and it told a network
engineer nothing they could act on.

Palette is a water column rather than a neutral near-black, with a
desaturated sea-green return for the accent. Verdict colours come from
the domain, so they carry meaning rather than decorate. No web fonts —
the CSP forbids loading anything and shipping font files would trade
what makes this a single pleasant binary for a typeface — so the
character comes from treatment: machine-set headings, tracked small
caps, hairlines.

The one ornament is a trace of returns across time on the runs page, one
bar per run coloured by verdict, oldest to newest. It is real data, pure
CSS, and it is what an echo sounder actually draws.

The mobile fix is the same idea rather than a fallback. Tables become
label-and-value records with dotted leaders, which is how a sounding log
prints and is easier to read on a phone than any table that scrolls
sideways. Above 46rem every row shares one grid so the columns agree by
construction; the first attempt used table-cell and each row wrapped
independently, which produces a table that does not align — a list
paying for borders.

Found by rendering it rather than reading the CSS: the trace stranded
itself against the right edge when runs were few, "Open run" broke across
two lines, equal columns wrapped device names while a one-digit count
kept a quarter of the row, and the sign-in page carried no wordmark at
all, so you arrived somewhere that never said what it was.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 22:45:33 +02:00
mrambossekandClaude Opus 5 621ee99b77 adminui: make the web UI usable on a phone
The header was a rigid flex row, so on a narrow screen the account name
and the sign-out button were pushed off the side of the viewport where
they could not be reached at all — not merely ugly, unusable. It wraps
now, and below 40rem the account block takes its own full-width row so a
long display name cannot crowd out the navigation.

Wide content scrolls inside its own box rather than dragging the page
sideways with it. Tables sit in an overflow-x container and <pre> is
capped at the viewport width; without that, one long self-test line or
one device table makes every other column of text unreadable, and on a
phone it is not obvious that the page has moved at all. Long opaque
strings — device ids, enrolment links — wrap anywhere rather than
insisting on a width nothing has.

Also: box-sizing on everything, stat cards that share a row instead of
each claiming the full width, and larger touch targets on buttons, where
.4rem is comfortable with a mouse and fiddly with a thumb.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 22:33:41 +02:00
mrambossekandClaude Opus 5 89b084338f docs: fmr is on .150/::150 only; ::2 removed and verified across a reboot
Listeners came off ::2 before the address did — the other order fails to
bind on restart — and the address stayed up until the CNAME to fmr-1 had
landed, since dropping it earlier would have broken ACME renewal for the
name the certificate is issued to.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 22:27:57 +02:00
mrambossekandClaude Opus 5 23ed835e00 docs: record that a revived beacon belongs in the web UI, not its own listener
Its Python service wildcard-bound 0.0.0.0:443, which is what silently
compromised the reserved measurement address. Two routes on the admin UI
would inherit the TLS and certificate already in place, need no extra
port, and pick up authentication the standalone receiver never had.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 22:17:55 +02:00
mrambossekandClaude Opus 5 5b5a54db7d server: reserve measurement addresses; serve the UI on both families
fmr keeps .151/::151 for measurement. Their diagnostic value is entirely
in their listening state being known: a TLS handshake that completes on a
port nothing listens on proves interception, with no competing
explanation. One stray bind turns that proof into a shrug, and nothing
about the failure is visible — the run still says the network is clean.

CheckReserved refuses to start when a listener would take one. Wildcards
are refused outright, because that is how this actually happens: every
listener defaults to ":port" and the next one added gets copied from an
existing default, claiming every address without anyone deciding to.

Reserved is not silent, though. The first version of the guard would have
refused the live config's UDP and canary-DNS binds on .151, which are
deliberate — as is STUN's RFC 5780 alternate. Reserving an address and
then forbidding the measurements that need it defeats the purpose. The
rule is narrower: no services, and never ports 80 or 443.

The adb-beacon receiver was wildcard-bound to 0.0.0.0:443, holding port
443 on every IPv4 address including the reserved one, so the IPv4
interception test had been compromised for as long as it had run. It is
disabled; restore with systemctl enable --now echolot-adb-beacon. This
also marks the guard's limit: it governs this server's listeners, and a
process outside its config can still pollute a reserved address.

The admin UI and ACME responder were single-address, which is why the UI
could only live on ::2 and why the server was reachable over IPv6 alone —
the thing that made it look nonexistent from a phone without working
IPv6. Both now take address lists like every other listener.

Verified from outside: .150/::150/::2 answer on 443 with a valid cert for
fmr.echo-lot.app, .151/::151 are closed on 80 and 443, and canary DNS is
still up on .151.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 22:17:29 +02:00
mrambossekandClaude Opus 5 d9ee8bc2ae app: don't report ICMPv6 silence that was never measured
The per-network attribution fix worked — the finding named rmnet_data1
instead of the IPv4-only wifi — and immediately exposed a worse problem
underneath. Cellular's result was `ok: false` because binding a socket to
it failed with EPERM, so no echo request was ever sent; the finding then
reported "IPv6 is configured, but ICMPv6 gets no reply" about a network
the app had never pinged. That is an assertion about the user's carrier
with nothing behind it.

`attempted` now travels beside `ok`, set only once sendto has returned,
and the finding requires both. Failing to bind is a fact about this app's
permissions on this device; it says nothing about the network, and the
two must not share a boolean.

Verified on hardware with a VPN active: every network fails to bind with
EPERM, nothing is sent, and no ICMPv6 finding is emitted — where the
previous build would have blamed the carrier. Recorded in build-status:
Android blocks per-network binding entirely while a VPN holds the default
route, so per-network measurement is unavailable to anyone with one
connected. That needs a deliberate answer rather than a silently green run.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 21:32:34 +02:00
mrambossekandClaude Opus 5 7a5004f293 server: a non-admin account can manage its own uploads
Signing in and being allowed to administer the server were the same
question: the OIDC callback refused a session outright to anyone outside
the admin group. A legitimate user could authenticate, be told what they
could not do, and be left with no way to see or delete the data their own
devices had uploaded.

They are separate questions now. Everyone who authenticates gets a
session; the admin flag rides inside the MAC'd payload, so promoting
yourself means forging a signature rather than editing a cookie, and a
role that does not parse fails closed to "user".

Pages scope themselves through visibleDevices/mayTouchRun rather than
filtering individually — per-page scoping is what the next page added
will be missing, and that failure is silent, since a listing that leaks
other people's uploads looks exactly like one that does not. Someone
else's run answers 404, not 403: a distinguishable refusal would confirm
the run exists. Revoking devices and minting enrolment tokens affect the
whole server and stay behind adminOnly at the route table, where someone
looking for who-may-do-what will actually find it.

Ownership is re-read per request instead of captured at sign-in, so
unlinking an account takes effect immediately rather than at session
expiry. Tests cover that, plus the degenerate case of an empty subject,
which must own nothing rather than everything with an empty account id.

Also: attribute the ICMPv6 finding per network. It compared "is IPv6
configured anywhere on this device" against "did any network answer",
which on a phone reports IPv6-is-broken about a network where IPv6 was
never configured. network_ref is null on every test, so the probe now
records per-network outcomes structurally rather than as prose a finding
would have to parse.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 21:10:38 +02:00
mrambossekandClaude Opus 5 7eaf0c4190 app: name the two half-configured IPv6 shapes, per network
v6.no_icmp_reply infers trouble from silence, which is ambiguous by
construction: a firewall dropping echo requests looks the same as a
network that cannot carry IPv6 at all. Two much stronger signals were
already sitting unread in the link snapshot, and a test device on a
Netbird tunnel surfaced both at once.

v6.route_without_address — a ::/0 route with no global address. The
router advertises itself as an IPv6 gateway while SLAAC produces nothing
usable. Hosts believe IPv6 is available and pay a connection timeout on
every dual-stack destination before falling back, which is felt as
general slowness with no packet loss to explain it.

v6.no_default_route — the mirror: a global address with nothing to route
it. A VPN installing host routes to specific destinations produces this
deliberately and it works, so a VPN transport reports it as INFO rather
than as a fault; without one it means the network handed out an address
it does not carry traffic for.

Both are read from the routing table, so neither is inferred from
silence, and both are reported per interface — "IPv6 is broken" is
useless advice when wifi is the broken one and cellular is fine.

Classification lives in core-measurement rather than the ViewModel so it
can be tested without a device; the fixtures are a real dumpsys table
(wifi advertising a route it cannot source from, working cellular, a
tunnel with two host routes) because the risk here is not bad boolean
logic but imagining shapes real networks do not produce.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 20:42:03 +02:00
mrambossekandClaude Fable 5 fec374abf5 findings: v6.broken claimed a cause it had no evidence for
A phone reported "IPv6 is configured but not working" while loading an IPv6-only
site over TCP perfectly well. The finding fired on one signal - ICMPv6 echo
getting no reply - at HIGH confidence. ICMPv6 echo is widely filtered on
networks where IPv6 works, so the two cases are indistinguishable from where the
app stands, and it was picking one.

Same class of error as the multi-homed downstream-loss bug: a confident
measurement of something that was not happening. Now v6.no_icmp_reply, low
severity, medium confidence, naming both explanations. Still reported, because
filtered ICMPv6 breaks Path MTU Discovery - large packets vanish instead of
being reported as too big - which is a fault in its own right.

Corroborating with a real IPv6 connection would separate the two properly, but
needs a target, which runs into the hardcoded-deployment issue already open.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 20:16:37 +02:00
mrambossekandClaude Fable 5 b5c8dda9a2 app: sign in to the server's identity provider
Authorization code with PKCE, a Sign in card in settings, and the echolot://auth
redirect handled next to the enrolment one - told apart by host, because one
spends a token and the other completes an authorization, and confusing them
would fail obscurely.

The detail that decides whether this survives a real phone: the PKCE verifier is
written to storage before the browser opens rather than held in memory. Handing
control to a browser backgrounds the process and Android may kill it, so the
callback arrives at a fresh one. An in-memory verifier works on a developer's
device and fails under memory pressure.

Pending state is cleared before the exchange is attempted, whatever the outcome:
it is single-use, and leaving it behind would let a later callback complete a
flow nobody started.

Also records an open issue the question about server requirements surfaced: two
probes hardcode the reference deployment, so a user with no server still sends
DNS and STUN traffic to fmr without being told. For a tool this careful about
what leaves the device, that is the wrong default.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 19:53:38 +02:00
mrambossekandClaude Fable 5 4e6f2da3fb runs: scope by account; app: the PKCE half of signing in
server-release / image (push) Successful in 15s
server-test / test (push) Successful in 36s
server-release / release (push) Successful in 38s
Three phones on one account now produce one history, which is the main reason to
have accounts beyond upload permission. GET /v1/runs returns the account's runs
and says how many devices contributed; fetching and deleting resolve a run id
against the caller's own devices, so an id from another account is not found
rather than fetched from wherever it happens to live.

The rule that needed stating: the empty account is never a group. Devices nobody
has signed in on are unrelated devices that share the absence of an owner, and
matching on "" would let any anonymous device read every other one's runs.
Tested, along with sibling-device access working and cross-account access not.

App side: authorization code with PKCE. The app is a public client - anything
compiled into an APK can be read out with unzip and strings - and the redirect
returns through a custom URI scheme that any app on the device may register, so
an intercepted code is a real risk. PKCE makes a stolen code worthless: it can
only be exchanged by presenting a verifier that never left the process.

A callback whose state does not match is refused before the code is spent and
before any network call, since that is exactly how someone gets a victim to
complete the attacker's sign-in.

Nothing from the IdP is retained. The ID token is used once to prove who is
signing in and then discarded; the device credential authenticates everything
afterwards. No access tokens to store, no refresh tokens to rotate.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 19:46:32 +02:00
mrambossekandClaude Fable 5 0eaba6150b adminui: an admin interface, behind authentication without exception
server-release / image (push) Successful in 15s
server-test / test (push) Successful in 36s
server-release / release (push) Successful in 38s
Replaces the unauthenticated admin mux. Everything but /healthz requires a
session, and that is the point: the previous arrangement relied on binding to
loopback, which worked exactly until the address changed and then failed
silently and publicly. A binding address is a deployment detail, not an access
control, and this package does not treat it as one.

Two ways in. OIDC through the confidential client, with state and PKCE - PKCE
even here, because it costs one hash and closes code interception independently
of the secret. And the break-glass password, throttled, for when the IdP is the
thing that is broken. Signing in without the admin group is refused with the
group named, because "you are not an admin" is a different problem from "your
password is wrong" and the remedy is elsewhere.

Sessions are MAC-checked cookies: HttpOnly, SameSite=Lax, Secure when TLS is on.
CSRF tokens are derived from the session rather than stored, so there is no
server-side table to keep in sync, and they are required on every state-changing
POST - SameSite already blocks cross-site posts in current browsers, but this is
the control that does not depend on the browser being current.

Server-rendered with html/template and no JavaScript: the pages are lists and
forms, and a framework would add a build step, a dependency tree and an update
treadmill to a program that has none of those. The CSP is default-src 'none'
accordingly.

Pages: overview, devices (with revocation and enrolment-link minting), uploaded
runs and a run viewer. Revocations and deletions are logged with who did them.
Runs are shown exactly as uploaded, at the privacy level their uploader chose -
nothing in the UI can un-redact one.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 19:31:06 +02:00
mrambossekandClaude Fable 5 7bb54e1ec8 docs: record the admin-listener exposure, and the encrypted-upload design
The incident is written down with its cause rather than just its fix: the admin
listener was built localhost-only, and that assumption travelled with it when I
changed the address. The compounding error is the one worth remembering -
checkAdminExposure verifies encryption and says nothing about authentication, so
it passed and gave false confidence. A green light on an adjacent property is
worse than no check.

Also records the encrypted-upload idea while the reasoning is fresh, including
the four consequences that decide whether it is worth building: what metadata
must stay readable (and what the UI loses if it does not), that losing the
passphrase loses the data by design, that metadata is not hidden regardless, and
that it makes a server-side anonymization floor unenforceable - which is fine,
since encryption serves the same purpose better.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 19:04:09 +02:00
mrambossekandClaude Fable 5 c7750fbf0b oidc: one verifier per issuer, because IdPs mint one per application
server-release / image (push) Successful in 15s
server-test / test (push) Successful in 34s
server-release / release (push) Successful in 35s
Authentik derives the issuer from the application slug, so two applications mean
two issuers - and a token's `iss` must match whoever signed it. A single pinned
issuer could therefore only ever serve one of the two clients.

So there is a verifier per issuer, and each accepts only the client belonging to
it. That is tighter than the previous arrangement as well as more general: a
token minted for the phone cannot be replayed at the admin login, and vice
versa, because they arrive at different verifiers with different audiences.

ECHOLOT_OIDC_APP_ISSUER is optional - empty means both clients share
ECHOLOT_OIDC_ISSUER, which is what IdPs with one global issuer do.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 18:47:34 +02:00
mrambossekandClaude Fable 5 5d7f59a66a acme: answer HTTP-01 from the server itself, on port 80
server-release / image (push) Successful in 15s
server-test / test (push) Successful in 34s
server-release / release (push) Successful in 35s
HTTP-01 always arrives on port 80 - the CA chooses the port, not the operator -
so it never collides with an admin UI on 443. The conflict only exists for
TLS-ALPN-01, which is the challenge type that does use 443.

Given that, the server keeps a permanent listener on 80 that answers challenges
from a webroot and redirects everything else to the admin UI. Same arrangement
as the webroot plugins for Apache and nginx, and better than letting the ACME
client bind 80 per renewal: nothing binds and unbinds, so a renewal cannot fail
because the port was briefly busy, and the client needs only write access to a
directory instead of the privilege to bind a low port. Port 80 also gets a use
it would want anyway.

The ACME client stays an external program. lego is also a Go library, but
importing it would put a large dependency tree into a server that deliberately
has none, and the CLI does the same job from a timer.

Tokens are validated by *shape* before any filesystem call, so traversal never
reaches the disk - a stronger guarantee than sanitising a path and trusting the
sanitiser.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 18:16:09 +02:00
mrambossekandClaude Fable 5 6afcb131ef admin: terminate TLS in the binary, with a certificate that reloads itself
server-test / test (push) Successful in 33s
Direct rather than behind Caddy or nginx. This binary already serves TLS for the
control plane, so it is reuse rather than new machinery; one process with one
config file is most of what makes this thing pleasant to run; and a proxy on the
box would invite someone to eventually front the control plane too, which would
break SPKI pinning because clients pin that certificate's key.

The hard part of TLS is not termination, it is renewal - so the certificate is
re-read when the files change. No reload hook to write, and none to quietly stop
working months later and be noticed only after the certificate has expired. A
torn write (renewal tools write cert and key separately) keeps the previous
certificate rather than taking the listener down.

Not applied to the control plane, on purpose: clients pin that key, so replacing
it should cost an operator a moment's thought and a restart, not happen because
a file changed. Two listeners, two different right answers.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 17:51:58 +02:00
mrambossekandClaude Fable 5 cd187f9ef5 config: the OIDC client secret, admin TLS, and a stop on plaintext admin
server-test / test (push) Successful in 34s
Two gaps found while answering where configuration lives.

The confidential admin client needs a secret and there was nowhere to put one -
I had added the issuer and both client ids but not the secret the admin login
actually needs. It now reads from ECHOLOT_OIDC_CLIENT_SECRET, and preferably
from ECHOLOT_OIDC_CLIENT_SECRET_FILE: a secret in the environment is readable by
anything that can see /proc/<pid>/environ and lands in every dump of the unit's
config, whereas a path is one file whose permissions an operator can reason
about. (/etc/echolot-server.env was also 0644; now 0600 on fmr.)

And the server now refuses to serve the admin UI in plaintext on a non-loopback
address. The session cookie is a bearer credential for everything the server can
do, and the OIDC authorization code arrives in a URL; in the clear, both belong
to anyone on the path - and on a globally routable address that is the internet.
A hard stop rather than a warning, because a warning in a log is not read by the
person who most needs it, and because the safe answers are cheap: bind to
loopback and tunnel, or supply a certificate. ECHOLOT_ADMIN_INSECURE=1 overrides
it, so the decision is made rather than stumbled into.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 17:49:26 +02:00
mrambossekandClaude Fable 5 3a4cb1c327 cli: survive the --serve transition when nobody is watching
server-release / image (push) Successful in 16s
server-test / test (push) Successful in 33s
server-release / release (push) Successful in 34s
Deploying v0.8.0 broke fmr, and the reason is a flaw I should have seen:
self-update is executed by the OLD binary, so the unit repair I put in the new
binary's updater cannot fix the very update that installs it. The unit kept its
argument-less ExecStart, the new binary answered that with usage and exit 2, and
the service went into a restart loop.

Fixed on fmr by hand, but that is not a fix for anyone else - and the whole
premise of an unattended self-update is that nobody is watching when it happens.

So: when started with no verb *and* systemd started us, the server repairs the
unit and serves anyway, loudly. systemd sets INVOCATION_ID for every service
invocation and nothing else does, so a person at a terminal still gets usage and
a non-zero exit. Marked as a one-release shim to remove once no deployment
predates --serve.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 17:44:33 +02:00
mrambossekandClaude Fable 5 3cdbccee18 cli: serving is an explicit verb; no arguments prints usage
server-release / image (push) Successful in 15s
server-test / test (push) Successful in 33s
server-release / release (push) Successful in 34s
Running an unfamiliar binary by name should tell you what it does, not bind a
dozen ports and start answering the internet. --serve (or --daemon) now does
that, and a bare invocation prints usage and exits 2 - non-zero on purpose, so a
service manager sees a failure rather than concluding the server ran and
finished cleanly.

The hazard this creates is worth spelling out, because it bites once and
silently: three places started the binary with no arguments - the systemd unit,
the unit template, and the Dockerfile - and --self-update replaces the binary
but never the unit. A routine update would therefore leave a service that cannot
start, discovered whenever the host next rebooted.

So the updater repairs it: after replacing the binary it appends --serve to an
ExecStart that has no flags, but only in a unit this program wrote (identified
by its description). Editing an operator's hand-written unit would be overreach;
leaving ours broken would be negligence.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 17:42:20 +02:00
mrambossekandClaude Fable 5 80d2092f1b oidc: accept both the app's public client and the server's confidential one
server-test / test (push) Successful in 37s
Explaining public vs confidential clients surfaced a gap in my own design: I had
assumed a single client id, but there are two clients here with genuinely
different properties.

  the Android app     public + PKCE, because an APK cannot keep a secret
  the admin UI        confidential, because the server can keep one in
                      /etc/echolot-server.env and weakening it to public buys
                      nothing

So the audience check now accepts either registered client id - and only those
two. "Any client of this issuer" would let every other application registered
with the same IdP authenticate here, which is the entire reason the check
exists. Either id alone is enough to enable sign-in, since an operator may
register only the app or only the admin UI.

The profile advertises the *app's* client id, since that is what a phone should
authorize as.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 17:31:33 +02:00
mrambossekandClaude Fable 5 89a5ff9139 adminauth: a break-glass local admin alongside OIDC
server-test / test (push) Successful in 44s
If the IdP is misconfigured, unreachable, or the admin group is a typo, the
operator is locked out of their own server with no way back short of editing
JSON on disk. A fallback that only matters when everything else is broken is
exactly the thing you cannot add later - by then you cannot get in to add it.

Stored as PBKDF2-HMAC-SHA256 from the standard library (Go 1.24+ has it, so no
dependency), 600k iterations, per-credential salt. A password rather than a
bearer token on purpose: a break-glass credential is the one most likely to end
up in a backup or a config-management repo, and a hash survives that where a
token does not. There is no email reset flow and should not be -
--set-admin-password on the host is the reset, and whoever can run it already
has the machine.

The password is read from stdin, never a flag, so it stays out of shell history
and the process list; piping still works for automation.

Details the tests pin, each for a reason:
  - the username is compared in constant time too, or a fast rejection is a
    timing oracle for which usernames exist;
  - the *stored* iteration count is used, so raising the constant later does not
    lock out existing passwords;
  - the throttle grows with consecutive failures but stays bounded and forgives
    after a quiet minute - a break-glass credential an attacker can lock out is
    a denial of service against the one person who needs it;
  - sessions are MAC-checked before anything in them is read, and rotating the
    secret invalidates every one at once, which is how they are revoked.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 17:12:54 +02:00
mrambossekandClaude Fable 5 ce6d0c2f64 oidc: the server becomes a relying party, and devices can carry an account
server-release / image (push) Successful in 15s
server-test / test (push) Successful in 35s
server-release / release (push) Successful in 34s
Echolot delegates identity to whatever IdP the operator already runs and stores
no passwords - no hashing, no reset flow, no lockout policy, and no credential
database to lose. For a tool people self-host next to other services, that is
the difference between one more service and one more thing that can leak
someone's password.

Verification is stdlib-only, matching the server's no-dependency rule. Longer
than jwt.Parse, and auditable in one sitting. The part that matters is the
algorithm allow-list: taking `alg` from the token is the classic forgery, so it
is fixed in code. Tests cover the real attacks against a genuine signer - a
self-contained IdP with real keys, because a mock that returns success proves
nothing about a verifier:

  alg=none, HS256/RS256 confusion, a payload swapped under a valid signature,
  a token addressed to another client, a token from another issuer, expired
  and future-dated tokens, and discovery that renames the issuer (which would
  otherwise have us fetch a stranger's keys believing they were the provider's).

With no admin group configured nobody is an admin. An operator who has not said
who may administer the server has not thereby said "anyone who can log in".

Device and account stay separate concepts: enrollment admits a device (operator's
token), signing in attributes it to a person (POST /v1/account/link, device
credential plus ID token - both required, neither substitutes). uploads=account
now means what it says instead of refusing everyone, and signing in does not
override uploads=off.

The profile advertises the sign-in configuration so the app can offer the button
only when there is something behind it, and drive PKCE without anyone typing an
issuer URL. A discovery failure is reported rather than hidden, so "configured
but the provider is not answering" is distinguishable from "not configured".

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 16:52:56 +02:00
mrambossekandClaude Fable 5 57a5ef8796 app: the stable-pseudonym switch was live at a level that pseudonymizes nothing
At `full` the anonymizer returns the document unchanged, so the salt has nothing
to act on - but the switch was enabled and looked like it did something. A
control that silently does nothing is the same class of fault as the preview
button and the archived-level label: the screen implying more than is true.

Shown disabled with the reason rather than hidden. The setting is still stored
and applies the moment the level changes, so making it vanish would hide state
that is still there; and a settings screen whose controls appear and disappear
as you touch other controls is harder to trust, not easier. The label dims with
the switch so "not active right now" reads at a glance.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 16:16:36 +02:00
mrambossekandClaude Fable 5 c19f382640 privacy: scrub identifiers inside raw shell output
Running the Shizuku tier for the first time uploaded every MAC address on the
local network to the server at the balanced level - fourteen of them, router and
all. The probes embed raw command output verbatim (ip neigh, ip route), which is
good evidence and also a complete household device inventory, and the anonymizer
could not see it: classification is by field name and whole-value shape, and
ip_neigh is one long string that is itself neither a MAC nor an address.

measurement-schema.md flagged raw dumps as hard to anonymize and proposed
dropping them from exports. Scrubbing is better: identifiers inside unclassified
strings are replaced in place with the same pseudonyms used elsewhere, so a MAC
appearing in both a parsed field and a raw dump still reads as one device, and
the dump stays readable - neighbour-table shape, host count, RFC1918 addresses
and vendor prefixes all survive. Dropping it would have protected the same data
by destroying the reason for collecting it.

One pass, not three: sequential passes re-process their own output. Once a MAC
became 78:9a:18:xx:yy:zz the IPv6 pattern matched it - six hex groups separated
by colons is an address - and destroyed the vendor prefix the MAC rule had just
preserved. Ordered alternation resolves each position once, MAC first.

RealDocumentTest runs the anonymizer over a captured run when ECHOLOT_REAL_RUN
points at one and fails on any surviving MAC; it self-skips otherwise so no
one's network lands in the repo. Against the document that leaked: 14 in, 0 out.

Also: the Settings preview button did nothing, reading UiState.history which is
empty until the History screen has been opened - same root cause as the "0
run(s)" count. It reads the archive now, and says when there is nothing to show.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 16:12:13 +02:00
mrambossekandClaude Fable 5 d04babff51 engine: live test for upstream throughput
3125 sent, 3125 counted by the server, 0% loss. The assertion that earns its
keep is received <= sent: that is what catches a counter that was never reset
between runs, which would otherwise look like a suspiciously good result.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 15:56:51 +02:00
mrambossekandClaude Fable 5 892e952a8e throughput: the upstream direction, counted by the only party that can
server-release / image (push) Successful in 15s
server-test / test (push) Successful in 33s
server-release / release (push) Successful in 33s
The client generates the traffic and the server counts it. No grant is involved
- the client is sending its own packets, so there is nothing to amplify - but it
does need the server's tally, because only the far end knows how much arrived.
Without that number a sender measures how fast it can transmit, which is usually
just the speed of the local NIC and is not the question being asked.

A new wire type the server counts and deliberately never answers: a reply would
double the traffic and drag the return path into a measurement that is
specifically about the outbound one.

The tally is a counter, not a list, and short-circuits before the observation
log. A five-second run at 20 Mbps is around ten thousand packets; one struct
each would turn a measurement into an allocation storm on a shared server, and
nothing needs the per-packet detail since the client holds the send-side record.
The gap between the two counts is the loss.

direction=up on the throughput action sends nothing - it zeroes the counter, so
a second run in one session measures itself instead of inheriting the first.

Same honesty rule as downstream: measures_network is false when what arrived
matches what was offered, because then the path was never the constraint.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 15:54:09 +02:00
mrambossekandClaude Fable 5 8646bab52d findings: adopt the registry in the app module; rename ipv6.* to v6.*
The registry was only used in core-engine. The app still emitted seven codes as
raw strings, so the registry test passed while codes lived outside it - among
them ipv6.broken, which fired on a real network and was in no registry at all.

All seven now take their code, category and severity from a registry entry, so
those three cannot disagree at a call site. Grepping for code = "..." across the
app, engine and probe modules now returns nothing.

ipv6.* -> v6.* is the third instance of the same rule being broken: they
declared Category.IPV6 while the prefix map only knows "v6", so
TestType.category("ipv6.broken") fell through to connectivity and the finding
rolled up under the wrong verdict light. The test-type registry already used v6.

Two severities reconciled rather than assumed:

  connectivity.captive_portal is medium, not high. The registry had guessed
  high; the probe emitting it had always said medium, and the probe was the
  considered value - a captive portal on hotel wifi is what should be there.
  no_internet keeps high, since nothing local fixes that.

  v6.not_offered stays info, and the registry now says why it must. Most
  networks still do not offer IPv6; a warning there lights a yellow verdict on a
  healthy network and teaches people to ignore the light.

Plus a BackHandler: the screen was a plain state variable with nothing tying it
to the back stack, so Back left the app from Settings/History instead of
returning to the run screen.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 15:47:52 +02:00
mrambossekandClaude Fable 5 6bba420845 app: the history row was naming the wrong document's privacy level
A row read "22 kB · full" directly beneath "uploaded to fmr", while the status
line above said the upload went as BALANCED. Both were true and they described
different documents: the row showed ArchivedRun.anonymization, which describes
the *archived* copy - deliberately unredacted, so always "full" - and the status
line described the *uploaded* copy.

Read together, that says the complete data was uploaded when a redacted copy was
sent. A privacy display that overstates what left the device is worse than none,
and telling the user what left the device is the one thing this screen is for.

The level a run was uploaded at is now recorded separately (uploaded_as) and the
row says "kept complete on this device" / "uploaded to fmr as balanced" - each
label naming the copy it belongs to.

Two more from the same screenshot:

  - Every row showed no verdict. The archive read summary.verdict; the schema
    calls it summary.overall. Silently null on every run, so the list's most
    prominent element was blank while everything else looked fine. The test
    fixture had the same wrong field name, which is why it passed.
  - The status line rendered the server's raw JSON index entry into the UI.

Verified on device: a fresh run archives with verdict "yellow" and
uploaded_as "balanced" beside anonymization "full".

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 15:37:05 +02:00
mrambossekandClaude Fable 5 ac6c653115 privacy: pseudonymize the whole ULA prefix, not just its tail
Found in a real uploaded run from the phone: the server held
fda1:3fb1:ff92:6696::2662 for a DNS server. The general IPv6 path keeps the
leading two groups on purpose - for a global address that preserves the ISP
allocation, which is the useful part - but for a ULA that passes through 32 of
the 40 random bits of the global ID.

A ULA looks like the v6 RFC1918 and the instinct is to treat it the same. It is
not analogous, and the difference is the point: an RFC1918 prefix is shared by
millions of networks and identifies none of them, while a ULA global ID is
random and unique to one network by construction (RFC 4193). The prefix IS the
identifier, so it was a network fingerprint surviving redaction.

Pseudonymized as a unit now, so two addresses on one ULA subnet still share a
pseudonymous prefix - "these hosts are on one network" survives, "this is that
network" does not. RFC1918 stays readable, and the contrast is what justifies
it; a test pins both halves.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 15:26:09 +02:00
mrambossekandClaude Fable 5 305d21f8a7 app: insets on the two newer screens, and one source for the run count
On-device verification found both.

safeDrawingPadding() was on the run screen but not on Settings or History -
they were added later and never got it - so "< Back  Settings" sat under the
status-bar clock. The same fault the run screen had already fixed, reintroduced
by new code that did not know about it.

Settings also read "0 run(s), 23 kB stored": the count came from
UiState.history, which stays empty until the History screen has been opened,
while the size read the archive directly. Two sources for one fact; the count
now reads the archive too.

Verified on a OnePlus 15 (A16): header clears the status bar, count reads
"1 run(s), 23 kB stored".

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 15:20:57 +02:00
mrambossekandClaude Fable 5 172afb421d privacy: fix a real leak - global IPv6 addresses were uploaded verbatim
Setting out to build the machine-readable schema, the first step was checking
whether the anonymizer covers the fields the schema declares sensitive. It did
not, and five identifying values were going out at the `balanced` level:

  networks[].link.addresses[].addr   the device's own global IPv6 address
  networks[].link.routes[].gateway   the ISP allocation
  networks[].link.dns.servers[]      the configured resolver
  private_dns_hostname               an internal hostname
  search_domains[]                   the internal domain

The settings screen describes that level as pseudonymizing addresses.

Root cause: classification keyed on field names, and the schema's actual names
were never added to the table. Every existing test passed, because each checked
a field somebody had remembered to write a case for - an unfalsifiable design
for a privacy control.

So beyond adding the names, classification now falls back to the *value* when
the name is unknown: anything shaped like an IPv4/IPv6 address or a MAC is
treated as one. Hostnames deliberately are not inferred by shape, since
train.udp_updown is indistinguishable from a domain and mangling a test type
would corrupt the document to protect nothing.

LeakTest is the guard, and is written to fail for fields nobody thought of: it
plants identifying values wherever one can occur and asserts none survive. It
also pins that RFC1918 addresses stay readable, so it cannot pass by
over-redacting. Route prefixes and :: needed care - 0.0.0.0/0 must stay itself
or a routing table becomes unreadable for no privacy gain.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 14:45:52 +02:00
mrambossekandClaude Fable 5 e7afc2210f findings: a registry, because the codes had already drifted
A finding code is the stable half of a result - what a dashboard groups by and
what someone greps a year of archived runs for. That only holds if a code means
exactly one thing forever, which fifteen ad-hoc string literals cannot promise.

By the time this was written the failure had happened twice:

  - Two emitters independently produced connectivity.downstream_loss and
    connectivity.loss_downstream for the same claim. Nothing objected. Anyone
    aggregating either would have silently seen half their data.
  - Two codes sat under nat.* while being declared Category.CONNECTIVITY.
    nat.udp_unreachable is not about NAT, and the prefix decides the category,
    which decides which verdict light the finding rolls up into. Renamed while
    that is still cheap.

Codes are now typed FindingSpecs carrying category and default severity;
emitters reference the spec rather than retyping the string, so a typo is a
compile error and two call sites cannot disagree about a finding's category.

docs/findings-registry.md is the contract and a test reads it, failing when the
document and the code disagree on which codes exist or how severe they are.
Documentation that drifts from its implementation is worse than none, because it
still looks authoritative. The check reads table rows only, so the prose can go
on explaining which codes were retired and why.

Closes open item 1 of measurement-schema.md section 9.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 14:25:54 +02:00
mrambossekandClaude Fable 5 f7701c2d2f engine: throughput is opt-in in the run config; document the work
A 5-second run at 50 Mbps moves ~30 MB. On a metered connection that is the
user's money, and a measurement tool that spends it unasked is not one people
keep installed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 14:14:20 +02:00
mrambossekandClaude Fable 5 3c9af04e6f grant: replace the rate check with a token bucket
server-release / image (push) Successful in 15s
server-test / test (push) Successful in 32s
server-release / release (push) Successful in 33s
The live throughput test found it: a 3-second run delivered 104 packets and
stopped after 50 milliseconds.

The rate check exempted the first 50 ms entirely, meaning to be lenient at
startup. The effect was the opposite. A sender could dump an unbounded burst
into that free window, and the instant the check switched on it compared those
bytes against 50 ms worth of allowance and refused everything until real time
caught up. Every short test passed — downtrain sends 50 packets, big_send seven
— and every sustained send died about fifty milliseconds in.

A token bucket (allowance = burst + rate x elapsed) has no such cliff; it is
smooth from t=0. The burst is 100 ms of the allowed rate, floored at one
ordinary datagram so a single packet is never refused outright. The floor is
deliberately one datagram: at 8 kbps a 64 KB floor would be sixty-four seconds'
worth, which is precisely the instant dump the ceiling exists to prevent. The
existing rate test caught that when I first tried it, and it was right.

Second half of the same bug: callers treated any refusal as terminal. TryAllow
now says why, so a sender can pace through a transient "too fast just now" and
still stop dead on a spent budget or an expired grant.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 14:10:30 +02:00
mrambossekandClaude Fable 5 3333788d9e throughput: paced downstream rate, with the qualifier that makes it honest
server-release / image (push) Successful in 15s
server-test / test (push) Successful in 32s
server-release / release (push) Successful in 32s
A throughput number reports the smallest limit on the path, and the sender's own
ceiling is one of the candidates. If the server was asked for 50 Mbps and 50
Mbps arrived, the network was never the constraint and "50 Mbps" says nothing
about it. So the result always carries limited_by and measures_network, and a
finding is raised only when the path is actually implicated.

Loss is computed against the *sender's* count, not the requested rate: the
server reports what it put on the wire, and the gap is the loss. A receiver
alone cannot tell "the network dropped it" from "the sender never sent it", and
guessing turns a healthy server-side limit into a phantom network fault. The
count is stored per action, not per packet — half a million packets of structs
would turn a measurement into memory exhaustion.

Sending is paced rather than flat out. An unpaced burst measures the server's
NIC and the first queue it meets, then collapses into loss that reads as a
network fault. The schedule is absolute rather than sleep-per-packet, which
would accumulate scheduler error and drift the rate down over a ten-second run.

Throughput gets its own grant budget sized from the request, so every other
action stays bounded at 8 MiB. When the byte cap binds before the clock does,
the *duration* is shortened and reported, rather than the run being truncated
halfway: promising thirty seconds and delivering twenty-one is the same
information with a surprise attached, and it keeps "the clock ended the run" as
the normal case — the only case where the rate is a clean property of the path.

That last behaviour came out of a test that failed honestly: 30 s at 100 Mbps
needs 375 MB against a 256 MB cap.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 14:05:23 +02:00
mrambossekandClaude Fable 5 35744c609e docs: record frag_send and the current testing state
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 13:49:43 +02:00
mrambossekandClaude Fable 5 a7dccf7da2 frag_send: crafted IP fragments, so ordering can be tested and not just delivery
server-release / image (push) Successful in 15s
server-test / test (push) Successful in 32s
server-release / release (push) Successful in 32s
Letting the kernel fragment an oversized datagram answers one question — do
fragments get through. It cannot answer the more interesting one, because the
kernel always emits them in order, first one first.

The classic middlebox fault is exactly about that ordering. Only the first
fragment carries the UDP header, and therefore the ports; a stateful firewall
or NAT that has not seen it has no flow to match the rest against, and many
drop them. That is invisible to any in-order test and shows up in the field as
"large DNS answers fail on this network" or "the tunnel breaks when the MTU
drops" — it works until the network reorders, then fails intermittently, which
is the hardest kind of fault to chase.

So the server now builds the fragments itself (raw socket, IP_HDRINCL) and
controls their order: in_order as a baseline, reversed, and first-fragment-last.
The datagram is assembled and signed whole before being cut up, so what the
client reassembles is indistinguishable from an ordinary packet — otherwise it
would be measuring our sender rather than the path.

Two details that would silently produce wrong answers:
  - The UDP checksum is computed rather than left zero. A zero-checksum datagram
    is dropped by some middleboxes, and that drop would be recorded as a
    fragmentation failure, which is the wrong conclusion entirely.
  - Fragment offsets are in 8-byte units, so non-final fragments are rounded to
    a multiple of 8. A 100-byte fragment is not an error, it is a datagram no
    host will ever reassemble.

frag-send is advertised only when a raw socket can actually be opened — checked
by opening one, since a permission model has more ways to say no than a
capability bit has to say yes.

Fragment header arithmetic is unit-tested (reassembly coverage, MF flags, shared
IP ID, 8-byte offsets, checksum verification), cross-compiled and run on Linux
since the code is build-tagged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 13:45:09 +02:00
mrambossekandClaude Fable 5 4ffa6e4ae2 engine: split packet loss by direction using the server's observations
"3 % loss" sends an engineer looking in both directions at once. The server
records every packet it received per sequence number, so the two cases are
distinguishable: sent-but-never-seen is upstream loss, seen-but-no-reply is
downstream. The findings say which, and say what is not implicated.

Downstream loss is measured against what reached the server, not against what
was sent — the other denominator counts every upstream loss twice and
overstates the return path.

Per-direction jitter comes out of the same records without needing synchronised
clocks: (server_rx - client_tx) carries a constant unknown offset, and
differencing successive samples cancels it, so RFC 3393 variation is honestly
attributable to a direction even though absolute latency is not.

Correlation is by wire sequence number, not loop index — the counter is shared
with every packet type on the session. ProbeSession exposes it even for a lost
probe, since that is precisely the packet whose direction is in question.

Live against fmr: 0.08 ms upstream jitter vs 0.85 ms downstream, an asymmetry a
round-trip test cannot see.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 13:09:59 +02:00
mrambossekandClaude Fable 5 3e7e3b8d33 scripts: one command to mint an enrollment link, QR included
Scanning beats pasting a 200-character string onto a phone, and with a device
attached the deep link can be delivered by adb with no typing at all.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 12:12:28 +02:00
mrambossekandClaude Fable 5 199807a8c9 docs: enrollment link encoding rules in the spec, session log in build-status
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 12:11:37 +02:00
mrambossekandClaude Fable 5 5291bdd045 enrollment: actually emit enroll_uri from the admin endpoint
server-release / image (push) Successful in 15s
server-test / test (push) Successful in 30s
server-release / release (push) Successful in 31s
The previous commit's edit to the admin handler silently did not apply, so the
endpoint still returned just the token. Caught by deploying and looking at the
response rather than by trusting the build to have picked it up.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 12:08:11 +02:00
mrambossekandClaude Fable 5 fe3658e009 chore: ignore the Kotlin compiler's .kotlin scratch directory
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 12:06:47 +02:00
mrambossekandClaude Fable 5 ad85f3bfcd enrollment: the server mints the §2.1 bootstrap link, the app consumes it
server-release / image (push) Successful in 15s
server-test / test (push) Successful in 29s
server-release / release (push) Successful in 31s
POST /admin/enroll-tokens now returns the whole link, not just the token:

  echolot://enroll?v=1&u=<control URL>&p=pin-sha256:<b64>&t=<token>

The server is the only party that knows all three parts at once, and the part
an operator gets wrong by hand is the base64 pin — which does not fail loudly,
it just never matches, surfacing days later as an inscrutable TLS error. The
app takes the link from a paste or from an echolot:// deep link (QR scan), and
writes URL, pin and credential together or not at all.

One trap the tests pin: an unencoded "+" in a query string decodes to a space,
so a hand-assembled link arrives with a pin wrong by one character. Base64 has
no spaces, so they are restored — unambiguous, and it cannot damage a correctly
encoded pin.

Also fixes a spec divergence: §2.1 names the field device_credential and the
first implementation shipped "credential". Both are sent now and the client
prefers the spec's; the alias goes once nothing reads it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 12:06:22 +02:00
mrambossekandClaude Fable 5 8166611af1 docs: record the compatibility-window work in build-status
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 11:43:02 +02:00
mrambossekandClaude Fable 5 33a6acb0bf compat: stop mangling refusal messages with HTML escapes
server-release / image (push) Successful in 14s
server-test / test (push) Successful in 29s
server-release / release (push) Successful in 31s
The server's 426 body reached the user as "needs \u003e= 0.2.0, \u003c 1.0.0":
Go escapes <, > and & by default for JSON destined for a page, which this is
not. Disabled at the encoder. The client now parses the error field rather than
pattern-matching it, so it survives whatever a future encoder decides to escape.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 11:41:00 +02:00
mrambossekandClaude Fable 5 9d6572bc33 compat: fix the too-new message's grammar, add a live gate test
server-release / image (push) Successful in 14s
server-test / test (push) Successful in 29s
server-release / release (push) Successful in 30s
The generated refusal read "point at a app within range". Also adds
LiveCompatTest, which checks the half a unit test cannot reach: that two
independently-built artifacts agree on the window, that the profile stays
readable for a version the server refuses, and that both bounds are enforced.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 11:38:45 +02:00
mrambossekandClaude Fable 5 0c5b021b63 compat: SemVer version windows between app and server
server-release / image (push) Successful in 14s
server-test / test (push) Successful in 30s
server-release / release (push) Successful in 30s
Both sides now declare what they will talk to, and enforce it. Two axes kept
deliberately separate, because conflating them is the trap:

  protocol_version  — CAN these builds talk. The correctness axis. Below 1.0.0
                      the minor is the breaking axis, per SemVer §4.
  release window    — MAY they, per policy. [min, max), advertised in the
                      profile, overridable by the operator.

The server refuses out-of-window apps with 426 and a body naming both versions
and the accepted range; the app checks the profile in both directions before a
run rather than discovering mid-measurement that it will be refused.

Three rules that shape the rest:

  - GET /v1/profile is never gated. It is where a refused client learns which
    version it needs; gating it leaves the user with a network error instead of
    an answer, which is precisely the confusion this exists to remove.
  - An unparseable or absent version is "unknown", and is allowed. Development
    builds report "dev", and a client too old to send the header cannot be
    identified anyway.
  - Bounds sit at breaking boundaries, not at releases, so shipping a patch
    never requires editing a range. The app's server minimum is 0.4.2 for a
    stated reason: earlier multi-homed servers mis-addressed granted sends and
    the client measured 100% downstream loss that never happened.

The app's versionCode is now derived from its SemVer instead of being a second
number someone has to remember to bump.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 11:36:34 +02:00
mrambossekandClaude Fable 5 277e33da75 chore: drop core-measurement/bin from the index too
deploy-site / deploy (push) Failing after 1m31s
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 10:54:48 +02:00
mrambossekandClaude Fable 5 14e5fad1b2 engine: downstream MTU and downstream train in the measurement document
Three facts the client cannot produce alone, kept deliberately separate:
mtu.pmtud_down (largest datagram that arrives unfragmented — meaningful only
because the server sets DF), mtu.frag_delivery (whether larger ones arrive once
fragmentation is allowed), and train.udp_downstream (loss, reordering and
arrival spacing in the download direction, which a round trip cannot separate
from upstream loss).

ServerMeasurement now runs them on the same ProbeSession as the echo train. It
had to: a fresh session restarts client-side sequence numbers and the server's
anti-replay window discards the lot, so the re-primed source is never recorded
and every granted send goes to a socket that has already closed. That produced
four confidently-wrong FAILED tests and a RED verdict on a healthy network.

Live against fmr: path MTU 1500, fragments to 4000, 100/100 downstream, GREEN.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 10:54:38 +02:00
mrambossekandClaude Fable 5 ce1aaa332a server: send granted traffic from the address the session actually used
server-release / image (push) Successful in 15s
server-test / test (push) Successful in 30s
server-release / release (push) Successful in 30s
fmr binds two IPv4 addresses. connFor picked whichever socket of the right
family came first in the bind list, so a downtrain for a session established on
.150 went out from .151 — and every packet was dropped by the client's NAT,
which has no mapping for that pair. tcpdump on the server showed all 50 leaving;
the client saw none. Read as "100% downstream loss", which is the worst kind of
wrong: a confident measurement of something that never happened.

Sessions now record which of our own bound addresses received their traffic, and
granted sends (and delayed echo) go back out through that socket. The fallback
to a family match is kept for the case where nothing has been received yet, and
the test pins both paths — a single-homed lab can never reproduce this.

Also: the client-side halves of the same work — anonymizer (core-privacy), local
run archive with retention (core-archive), upload client, and the app's settings
and history screens.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 10:45:43 +02:00
mrambossekandClaude Fable 5 7a94c9a3d7 chore: ignore the VSCodium Java extension's bin/ output
server-release / image (push) Successful in 15s
server-test / test (push) Successful in 28s
server-release / release (push) Successful in 29s
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 10:26:30 +02:00
mrambossekandClaude Fable 5 2521d39989 server: DF-mode big_send + uploaded-run storage with an operator policy
big_send now forces the Don't-Fragment bit for the whole burst by default, so
the largest size that arrives IS the downstream path MTU rather than "fragments
got through" — two different measurements the schema already separates. Sizes
above our own egress MTU (from the startup self-test) are refused up front and
reported as max_df_bytes, because absence caused by our kernel must not be read
as a limit of the client's path.

Uploads: one JSON file per run under the state dir, with the policy the operator
actually cares about — who may upload (off / anonymous / account), how large,
how long to keep, and the least anonymization accepted. The profile advertises
all of it so the app can present the switch honestly instead of discovering the
rules by failing. `account` refuses today rather than falling back to anonymous:
picking the strict setting before OIDC lands must not silently mean the loose one.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 10:26:19 +02:00
204 changed files with 24613 additions and 2163 deletions
+24 -6
View File
@@ -10,12 +10,18 @@
# releases don't trigger each other's pipelines. # releases don't trigger each other's pipelines.
# #
# Required secrets: # Required secrets:
# REGISTRY_TOKEN personal access token with read+write package scope — # REGISTRY_TOKEN personal access token with read+write package scope —
# the built-in Actions token is NOT accepted by the # the built-in Actions token is NOT accepted by the
# container registry (docker login → unauthorized). # container registry (docker login → unauthorized).
# Create: user Settings → Applications → Generate token. # Create: user Settings → Applications → Generate token.
# REGISTRY_USER optional; defaults to the pushing actor's username. # REGISTRY_USER optional; defaults to the pushing actor's username.
# The release job needs only the built-in GITHUB_TOKEN. # RELEASE_SIGNING_KEY base64 ed25519 seed that signs SHA256SUMS. Self-updating
# servers verify the signature against the public key baked
# into the binary (selfupdate.DefaultPublicKeyB64) and REFUSE
# unsigned releases, so this job hard-fails without it —
# a release nobody can install is better failed loudly here.
# Mint a pair with: go run ./cmd/release-sign -gen
# The release job otherwise needs only the built-in GITHUB_TOKEN.
name: server-release name: server-release
on: on:
@@ -44,6 +50,18 @@ jobs:
done done
(cd ../dist && sha256sum * > SHA256SUMS) (cd ../dist && sha256sum * > SHA256SUMS)
- name: Sign SHA256SUMS
working-directory: server
env:
RELEASE_SIGNING_KEY: ${{ secrets.RELEASE_SIGNING_KEY }}
run: |
[ -n "$RELEASE_SIGNING_KEY" ] || { echo "::error::secret RELEASE_SIGNING_KEY is missing — self-updating servers refuse unsigned releases, so publishing one would strand the fleet. Add it under Settings → Actions → Secrets."; exit 1; }
go run ./cmd/release-sign ../dist/SHA256SUMS
# Verify with the key baked into the binary we just built — catches a
# secret that does not match DefaultPublicKeyB64 before it ships.
PUB=$(grep -o 'DefaultPublicKeyB64 = "[^"]*"' internal/selfupdate/selfupdate.go | cut -d'"' -f2)
go run ./cmd/release-sign -verify -pub "$PUB" ../dist/SHA256SUMS
- name: Create release + attach binaries - name: Create release + attach binaries
env: env:
TOKEN: ${{ secrets.GITHUB_TOKEN }} TOKEN: ${{ secrets.GITHUB_TOKEN }}
+6
View File
@@ -40,3 +40,9 @@ keystore.properties
# wrangler build/dev artifacts # wrangler build/dev artifacts
web/.wrangler/ web/.wrangler/
# Eclipse/JDT output from the VSCodium Java extension — not a build artifact we own
echolot-app/*/bin/
# Kotlin compiler scratch/error logs
echolot-app/.kotlin/
+67 -4
View File
@@ -32,10 +32,25 @@ Keep prober result IDs aligned with the measurement-schema test-type registry.
## Versioning ## Versioning
Prefer **patch** bumps (`server-v0.3.1`) for additive/incremental work; reserve **minor** bumps Both artifacts are **SemVer**. Prefer **patch** bumps (`server-v0.3.1`) for additive/incremental
for real milestones. Don't burn through minor versions. Tags are namespaced: `server-v*` for the work; reserve **minor** bumps for real milestones. Don't burn through minor versions. Tags are
Go server, `v*` for the app. Pushing a `server-v*` tag runs CI → binaries + Gitea release + namespaced: `server-v*` for the Go server, `v*` for the app. Pushing a `server-v*` tag runs CI →
registry image; the server on fmr can `--self-update` from those releases. binaries + Gitea release + registry image; the server on fmr can `--self-update` from those
releases.
The app's version lives once, as `appVersionName` in `app/build.gradle.kts`; **`versionCode` is
derived from it** (`major*1e6 + minor*1e4 + patch*10`). Never set it by hand — a second number a
human has to remember to bump eventually disagrees with the first.
**Versions are load-bearing** (probe-protocol.md §8): the server refuses apps outside its window
with `426`, and the app refuses servers outside its own. Two axes, kept separate:
- `protocol_version`*can* they talk. The correctness axis; below 1.0.0 the **minor** is the
breaking axis.
- release-version window — *may* they, per policy. `[min, max)`, bounds at breaking boundaries so
a patch never strands a fleet. Client bounds: `Compat.kt`. Server: `ECHOLOT_MIN/MAX_APP_VERSION`.
Raise a minimum only when older peers are actively harmful, and say why in the constant's comment.
`GET /v1/profile` must stay ungated — it is how a refused client learns what it needs.
## Layout ## Layout
@@ -52,6 +67,29 @@ echolot-prober/ the capability prober (self-contained Gradle buil
ui/ProberScreen.kt result cards colored by verdict ui/ProberScreen.kt result cards colored by verdict
``` ```
## Client modules (echolot-app/)
Pure Kotlin/JVM where possible, so the interesting logic is unit-testable without a device and can
be exercised against the live server from the PC:
- `core-protocol` — control plane (pinned TLS) + ELT1 UDP data plane. **One `ProbeSession` per
server session, for its whole lifetime**: a second one restarts sequence numbers, the server's
anti-replay window discards every packet, and granted sends then target the closed socket.
- `core-measurement` — the schema types. `core-engine` — composes probes into documents.
- `core-privacy` — the §8 anonymizer (`full` / `balanced` / `strict`). Field classification lives
in one table (`Classification.kt`); keep it there rather than annotating models.
- `core-archive` — on-device run storage + retention. `enabled` is separate from the three
ceilings: all-zeros means "no limits", not "keep nothing".
**The local archive keeps the unredacted document; anonymization happens per upload, on the way
out.** Never redact what is stored locally.
### Live testing without a device
`echolot-app/scripts/test-fmr.sh [gradle-task] [test-filter]` mints an enrollment token over SSH,
enrolls, computes the SPKI pin from the served cert and runs a `Live*Test` against fmr. This covers
the whole server-facing vertical (granted sends, downstream MTU, uploads) with no phone involved —
use it before asking the user to test on hardware.
## Conventions ## Conventions
- Every probe returns a `ProbeResult` — never throws to the caller (MainActivity wraps anyway). - Every probe returns a `ProbeResult` — never throws to the caller (MainActivity wraps anyway).
@@ -107,6 +145,23 @@ First build downloads AGP/Compose/Shizuku from Google Maven + Maven Central.
Shizuku, **toggle Wireless debugging off/on** — Shizuku keeps running (separate process), a fresh Shizuku, **toggle Wireless debugging off/on** — Shizuku keeps running (separate process), a fresh
port + mDNS record appear, and the beacon/connector recover. Plan the Shizuku-tier dev loop port + mDNS record appear, and the beacon/connector recover. Plan the Shizuku-tier dev loop
around this (or USB, if ever available). around this (or USB, if ever available).
- **A poisoned Gradle *build cache* entry can silently drop a whole module from the APK.**
Symptom: the app dies with `ClassNotFoundException` for a class that plainly exists, while the
build is green and `./gradlew :app:dependencies` lists the module on `debugRuntimeClasspath`.
The module's own jar is correct; its code simply never reaches AGP's intermediates. `clean`,
`rm -rf */build` and `--rerun-tasks` all fail to fix it, because **none of them touch the build
cache** — look for `compileKotlin FROM-CACHE` in the log. Fix: rebuild with `--no-build-cache`.
Verify by grepping the APK's dex for a string literal that only that module defines; grepping for
a *class name* proves nothing, because callers carry the name as a reference whether or not the
class is packaged:
`unzip -o -q app-debug.apk "classes*.dex" && grep -a "pin-sha256:" *.dex`
Suspect this whenever a runtime failure contradicts a successful build.
- **Empty-jar race with the IDE.** VSCodium's Java/Kotlin extension runs its own Gradle daemon on
the same project; when it overlaps a CLI build, a module's `build/libs/*.jar` can end up
containing only a manifest, and Gradle then considers `jar` up-to-date. Dependent modules fail
with "Unresolved reference" on symbols that plainly exist. Fix: `rm -f <module>/build/libs/*.jar`
and re-run the `jar` task. Suspect this whenever a reference resolves in one module but not in
its consumer.
- As of the AGP 9.2.0 / Gradle 9.6.0 / Kotlin 2.2.10 bump, JDK 17+ (including 25) works — - As of the AGP 9.2.0 / Gradle 9.6.0 / Kotlin 2.2.10 bump, JDK 17+ (including 25) works —
AGP 9 requires Gradle 9.1.0+ and Kotlin 2.2.10+ as its minimum KGP version. On machines with AGP 9 requires Gradle 9.1.0+ and Kotlin 2.2.10+ as its minimum KGP version. On machines with
Android Studio, its `jbr` directory works as `JAVA_HOME`. Android Studio, its `jbr` directory works as `JAVA_HOME`.
@@ -118,3 +173,11 @@ First build downloads AGP/Compose/Shizuku from Google Maven + Maven Central.
are SUPPORTED on both known devices via `Os.recvmsg` + `StructMsghdr` reflection. are SUPPORTED on both known devices via `Os.recvmsg` + `StructMsghdr` reflection.
3. Fold the confirmed capabilities + Shizuku dump-format samples back into the production 3. Fold the confirmed capabilities + Shizuku dump-format samples back into the production
`core-probe` / `core-shizuku` modules. `core-probe` / `core-shizuku` modules.
## Enrolling a device with a server
`echolot-app/scripts/enroll-link.sh [note]` mints a §2.1 bootstrap link on fmr over SSH and prints
it (plus a QR if `qrencode` is installed, plus the `adb shell am start -a …VIEW -d '<uri>'` command
when a device is attached). The link carries a single-use token — treat it as a secret until spent.
Never hand-assemble one: the base64 pin needs percent-encoding, and a pin wrong by one character
fails as an inscrutable TLS error rather than as a bad pin.
+947 -3
View File
@@ -211,8 +211,8 @@ Two collection-loop gotchas found while driving the phone over USB:
2. ~~If `trace.errqueue_reachable` = PARTIAL, add a C-over-JNI errqueue shim.~~ **Retired** — 2. ~~If `trace.errqueue_reachable` = PARTIAL, add a C-over-JNI errqueue shim.~~ **Retired** —
SUPPORTED on both known devices; `traceroute.udp4` reads real hops via `Os.recvmsg` + SUPPORTED on both known devices; `traceroute.udp4` reads real hops via `Os.recvmsg` +
`StructMsghdr` reflection, so no `:native` module is needed. `StructMsghdr` reflection, so no `:native` module is needed.
3. Start the Go server skeleton (enrollment + profile + sessions + UDP echo with observation 3. ~~Start the Go server skeleton per probe-protocol.md.~~ **Shipped** — live on fmr since
blocks + canary-DNS reference records) per probe-protocol.md. v0.2.0 (2026-07-31); see the server sections below.
4. Fold confirmed capabilities into the production `core-probe` / `core-shizuku` modules. 4. Fold confirmed capabilities into the production `core-probe` / `core-shizuku` modules.
## Production probe server — LIVE on dedicated VM "fmr" (2026-07-31) ## Production probe server — LIVE on dedicated VM "fmr" (2026-07-31)
@@ -222,7 +222,7 @@ SSH only — verified untouched by the daemon (explicit multi-address binds, no
Control: fmr-1:8443 (SPKI pin `zRV9qkiLnRexAeh4RrSfJzbPWO+U/2Oj2/NVM/KfXlg=`, verified Control: fmr-1:8443 (SPKI pin `zRV9qkiLnRexAeh4RrSfJzbPWO+U/2Oj2/NVM/KfXlg=`, verified
externally over v4+v6). UDP data plane on all four service addresses :8442 — the second IP is externally over v4+v6). UDP data plane on all four service addresses :8442 — the second IP is
the stun-5780 substrate. Daily randomized self-update timer installed (checksum-verified the stun-5780 substrate. Daily randomized self-update timer installed (checksum-verified
against SHA256SUMS; signature verification still TODO before treating the source as untrusted). against SHA256SUMS; signature verification landed 2026-08-02 — see "Release signing" below).
Host config in `/etc/echolot-server.env`. SSH access for sessions: `ssh claude-echolot`. Host config in `/etc/echolot-server.env`. SSH access for sessions: `ssh claude-echolot`.
## Server v0.3.0 — STUN + TCP echo + observations + actions (2026-07-31) ## Server v0.3.0 — STUN + TCP echo + observations + actions (2026-07-31)
@@ -543,3 +543,947 @@ Best available behavior, now implemented: the banner still opens Shizuku, but th
exact steps there ("Pairing", then "Start"), and a second tap target opens **Developer options** exact steps there ("Pairing", then "Start"), and a second tap target opens **Developer options**
(`Settings.ACTION_APPLICATION_DEVELOPMENT_SETTINGS` — public and exported) since Wireless (`Settings.ACTION_APPLICATION_DEVELOPMENT_SETTINGS` — public and exported) since Wireless
debugging must be enabled first for Shizuku's wireless start to work at all. debugging must be enabled first for Shizuku's wireless start to work at all.
### Downstream measurements: asymmetric grants, DF-mode big_send (server-v0.4.0 … v0.4.2, 2026-08-01)
The client can measure a round trip and the largest packet it can *send*. It cannot measure the
largest packet it can *receive*, or downstream-only loss — those need the server to push, which is
exactly what §3.4 gates behind an asymmetric grant. Implemented and verified live from the PC:
- **`session.Grant`** — created per action, bound at creation to the session's *observed*
data-plane source (no grant without a verified destination), clamped to server limits, with a
byte budget, an average-rate ceiling and an expiry. Unit-tested for each of those refusals.
- **`downtrain`** — N packets of size S every I µs; the client derives downstream loss,
reordering and inter-arrival spacing.
- **`big_send`** — one datagram per requested size. **DF is on by default**, so the largest size
that arrives *is* the downstream path MTU. Without DF the kernel fragments and the result only
says whether fragments get through — a different fact, and the reason the schema has both
`mtu.pmtud_down` and `mtu.frag_delivery`. Sizes above the server's own egress MTU (from the
startup self-test) are refused up front and reported as `max_df_bytes`, so an absence caused by
our kernel is never read as a limit of the client's path.
Live from the PC against fmr: downstream path MTU **1500** (1472 payload, DF), fragmented delivery
up to **4000**, downstream train **100/100, 0 % loss, 0 reordered**, inter-arrival 3.3 ms for a
3000 µs send interval.
#### Two bugs this shook out, both invisible in a single-homed lab
1. **Granted sends went out from the wrong local address** (fixed in server-v0.4.2). fmr binds two
IPv4 addresses; `connFor` returned whichever socket of the right family came first in the bind
list. A train for a session established on `.150` left from `.151` and every packet was dropped
by the client's NAT, which has no mapping for that pair. tcpdump showed all 50 leaving, the
client saw none — reported as *100 % downstream loss*, a confident measurement of something
that never happened. Sessions now record which of our own bound addresses received their
traffic and granted sends go back through that socket; `connfor_test.go` pins both that and the
family fallback.
2. **A second `ProbeSession` on one server session is silently dead.** Sequence numbers restart at
zero client-side while the server's anti-replay window keeps counting, so every packet is
discarded as a replay — and because the server then never records the new source, the grant
still targets the closed socket. `ServerMeasurement` now uses one ProbeSession for the whole
run; `ProbeSession`'s doc comment states the constraint.
### Run archive, anonymizer and uploads (2026-08-01)
Three pieces, deliberately separate:
- **`core-archive`** — one JSON file per run plus an index entry, in a plain directory the user can
inspect or delete with a file manager. Retention (max runs / max age / max total bytes) is
enforced on every save rather than by a sweeper. `enabled` is a separate flag from the three
ceilings because "no limits" and "keep nothing" are opposite intentions; collapsing them onto
all-zeros is how a user who turns the caps off ends up with an empty history. 13 tests.
- **`core-privacy`** — the schema §8 anonymizer, three levels. `full` (your own server) changes
nothing; `balanced` pseudonymizes SSIDs/hostnames, keeps the OUI half of a MAC and the /16 of a
public IP, keeps RFC1918 verbatim (it describes topology, not a person), and *drops* neighbour
inventories (SSDP/ARP/scan results) rather than mangling them; `strict` keeps only metrics,
statuses and finding codes. Pseudonyms are consistent within a document and — by default — not
across documents, so an upload endpoint cannot link a device's runs; a stable salt is opt-in for
people diffing their own history. Classification is one readable table, not annotations spread
across modules. 14 tests, each pinning a property someone's privacy depends on.
- **Server-side upload policy** — `off | anonymous | account`, plus max size, retention days, max
runs per device, and the *least* anonymization accepted. The profile advertises all of it so the
app presents the choice honestly instead of discovering the rules by being rejected. `account`
refuses today rather than falling back to anonymous: picking the strict setting before OIDC
lands must not silently mean the loose one.
**The archive holds the unredacted document; redaction happens on the way out, per upload.** The
local archive is the user's own data on their own device, and redacting it would destroy exactly
the detail that makes a week-old run worth keeping.
App-side: settings screen (archive limits, privacy level with a plain-language description of what
each keeps, auto-upload off by default, server URL/pin/credential), history screen showing whether
each run left the device, and a **preview of the exact bytes an upload would send** — an anonymizer
the user cannot inspect is only a promise.
Live round trip against fmr: uploaded a run, listed it, fetched it back and asserted the SSID, the
SSDP neighbour name and the free-text note are absent from what the server stores while the
finding code and the metrics survive, then deleted it.
### Still open
- `mtu.pmtud_up` (DF + errqueue), `frag_send`, `throughput`, TRAIN_REPORT retrieval.
- Enrollment UI in the app (server URL/pin/credential are typed in by hand today).
- Accounts/OIDC on the server, which is what `uploads=account` is waiting for.
- Nothing in this entry has been exercised on a phone yet — all of it was verified from the PC
against the live server. On-device verification is the next step.
### SemVer compatibility windows between app and server (server-v0.5.0 … v0.5.2, 2026-08-01)
Both artifacts are SemVer, and each now declares — and enforces — which peer versions it will talk
to. Spec: `docs/probe-protocol.md` §8.
**Two axes, deliberately not conflated.** Release versions are a *proxy* for what actually has to
match, so the real thing is checked first:
- `protocol_version` — **can** these builds talk. Advertised in the profile; a peer in a different
breaking series is refused whatever its release version says. Below 1.0.0 the **minor** is the
breaking axis (SemVer §4).
- release-version window — **may** they, per policy. `[min, max)`, min inclusive, max exclusive,
because the useful bound is always "the version that broke it".
Bounds sit at breaking boundaries, not at releases, so shipping a patch never requires editing a
range. The app requires server `>= 0.4.2` for a stated reason, not caution: earlier multi-homed
servers mis-addressed granted sends and the client measured 100 % downstream loss that never
happened. Operators override the server side with `ECHOLOT_MIN_APP_VERSION` /
`ECHOLOT_MAX_APP_VERSION`; a malformed bound is fatal at startup rather than ignored, so a typo
cannot silently disable a restriction.
Three rules that shaped the implementation:
1. **`GET /v1/profile` is never gated.** It is where a refused client learns which version it needs;
gating it leaves the user with a network error instead of an answer.
2. **An unparseable or absent version is `unknown`, and is allowed.** Dev builds report `dev`, and a
client too old to send the header cannot be identified anyway.
3. **Refusal is 426 with a body naming both versions and the window**, surfaced client-side as a
distinct `VersionRefused` rather than folded into "network error".
The app's `versionCode` is now derived from its SemVer (`major*1e6 + minor*1e4 + patch*10`) instead
of being a second number to remember.
Verified live against fmr (`LiveCompatTest`): profile advertises the window and stays readable for a
refused version; 0.1.0 and 99.0.0 are both refused with actionable messages; 0.2.0 and a missing
header are both served.
One user-visible bug caught in the process: Go's JSON encoder HTML-escapes `<`, `>` and `&` by
default, so the refusal reached the client as `needs \u003e= 0.2.0`. Disabled at the encoder (this
is an API, not a page), and the client now *parses* the error field instead of pattern-matching it,
so it survives whatever a future encoder decides to escape.
### Enrollment: the server mints the bootstrap link (server-v0.5.3 … v0.5.4, 2026-08-01)
Until now a device was configured by hand-typing a control URL, a base64 SPKI pin and a
credential. That is the step that goes wrong, and it goes wrong quietly: a pin off by one
character does not fail loudly, it just never matches, and surfaces days later as an inscrutable
TLS error.
`POST /admin/enroll-tokens` now returns the whole §2.1 bootstrap link alongside the token, because
the server is the only party holding all three parts at once. The app takes it from a paste or an
`echolot://enroll` deep link (so a QR scan configures a server in one action) and writes URL, pin
and credential **together or not at all** — a half-applied server fails later, somewhere else,
with an error pointing at the wrong thing.
The control URL comes from `ECHOLOT_PUBLIC_URL` (set on fmr to `https://fmr-1.echo-lot.app:8443`),
falling back to the first control listen address; a wildcard bind warns rather than emitting a
link to `0.0.0.0`.
**The encoding trap, which is the whole reason this is tested across both languages.** The pin is
base64, so it contains `+`, `/` and `=` — each of which means something else in a query string. An
unencoded `+` decodes to a space, leaving the pin wrong by exactly one character. Base64 has no
spaces, so the parser restores them; that cannot damage a correctly-encoded pin and it rescues
every hand-assembled link. `LiveEnrollmentTest` redeems a link the *server* produced, which is the
only way to catch a disagreement between the Go assembler and the Kotlin parser — a unit test on
either side alone cannot see it. It also asserts the token is refused the second time.
Also fixed a spec divergence found while reading §2.1: the spec names the field
`device_credential`, the first implementation shipped `credential`. The server now sends both and
the client prefers the spec's; the alias goes once nothing reads it.
Two process notes from this round:
- An edit to the admin handler silently failed to apply and the endpoint kept returning just the
token. Caught by deploying and *looking at the response*, not by trusting a green build.
- The live suite is now six tests (`LiveServerTest`, `LiveMeasurement`, `LiveGranted`,
`LiveUpload`, `LiveCompat`, `LiveEnrollment`), all green against fmr from the PC with no device.
### Directional loss: which way is the packet loss? (2026-08-01)
A round trip can only report that *something* was lost somewhere, which is the least useful form
of the answer — "3 % loss" sends an engineer looking in both directions at once. The server
already records every packet it received per sequence number (§6), so the two cases are actually
distinguishable, and `train.udp_updown` now reports them separately:
- sent, never seen by the server → **upstream** loss
- seen by the server, reply never arrived → **downstream** loss
Findings name the direction and say what is *not* implicated, which is half the value:
`connectivity.loss_upstream` ("the return path is not implicated: replies came back for everything
that arrived"), `connectivity.loss_downstream`, `nat.udp_unreachable_upstream`.
Two things the implementation gets deliberately right:
- **Downstream loss is measured against what reached the server**, not against what was sent.
Using "sent" as the denominator counts every upstream loss a second time and overstates the
return path. Pinned by a test with loss in both directions at once.
- **Per-direction jitter without synchronised clocks.** Absolute one-way delay would need clock
sync and we deliberately have none (the two-clock rule). But `server_rx client_tx` carries a
constant unknown offset, and differencing successive samples cancels it — so RFC 3393 one-way
delay variation *is* honestly attributable to a direction even though latency is not. A test
pins that a 10-second clock offset changes nothing.
Correlation is by **wire sequence number**, which is not the loop index: the counter is shared
with every other packet type on the session, so "the nth echo" is not "sequence n". `ProbeSession`
now exposes `lastSeq`, including for a probe that was lost — a lost packet still has a sequence
number, and that number is exactly what tells you which way it was lost.
Live against fmr: 20/20 both ways, and jitter of **0.08 ms upstream vs 0.85 ms downstream** — a
tenfold asymmetry that a round-trip measurement cannot see at all.
10 unit tests on the arithmetic (a wrong denominator here does not crash, it produces a plausible
number pointing at the wrong half of the network) plus the live correlation check.
### frag_send: crafted IP fragments, so *ordering* is testable (server-v0.6.0, 2026-08-01)
`big_send` with `df=false` answers one question — do fragments get through. It cannot answer the
more interesting one, because the kernel always emits fragments in order, first one first.
The classic middlebox fault is exactly about that ordering. Only the **first** fragment carries the
UDP header, and therefore the ports; a stateful firewall or NAT that has not seen it has no flow to
match the rest against, and many simply drop them. That is invisible to every in-order test, and in
the field it looks like "large DNS answers fail on this network" or "the tunnel breaks when the MTU
drops" — it works until the network reorders, then fails intermittently, which is the hardest kind
of fault to chase.
So the server builds the fragments itself (raw socket, `IP_HDRINCL`) and controls their order:
`in_order` (baseline), `reversed` (last fragment first), `first_last` (first fragment held back
250 ms). The datagram is assembled and **signed whole** before being cut up, so what the client
reassembles is indistinguishable from an ordinary packet — otherwise the test would be measuring
our sender rather than the path. New test type `mtu.frag_ordering`; findings
`mtu.fragments_blocked` and `mtu.fragment_reorder_sensitive`.
Two details that would otherwise produce confidently wrong answers:
- **The UDP checksum is computed, not left zero.** Zero is legal in IPv4 and would be less code,
but zero-checksum datagrams are dropped by some middleboxes — and that drop would be recorded as
a fragmentation failure, which is the wrong conclusion entirely.
- **Fragment offsets are in 8-byte units**, so non-final fragments are rounded down to a multiple
of 8. A 100-byte fragment is not an error; it is a datagram no host will ever reassemble.
`frag-send` is advertised only when a raw socket can actually be opened — checked by opening one,
because a permission model has more ways to say no (userns, seccomp, LSM) than a capability bit has
to say yes. fmr runs as root with `cap_net_raw` in its bounding set, so it is available there.
Fragment ordering runs only after `mtu.frag_delivery` shows fragments arrive at all; otherwise the
three orderings would each report "not delivered" and read as three faults instead of one.
The header arithmetic is unit-tested (reassembly coverage with no gaps or double-delivery, MF
flags, shared IP ID, 8-byte offsets, checksum verification over odd and even lengths). Because the
code is `//go:build linux`, the tests are **cross-compiled and run on fmr** — there is no Go
toolchain there, so `go test -c` plus scp is the loop.
Live against fmr: 4 fragments per burst, and all three orderings reassembled — a healthy path, and
the baseline against which a mobile network will be interesting.
### Testing state (2026-08-01)
Six live tests against fmr, all green, no device involved: `LiveServerTest`, `LiveMeasurement`,
`LiveGranted`, `LiveDownstream`, `LiveUpload`, `LiveCompat`, `LiveEnrollment`. Plus 74 client unit
tests and the full Go suite. Everything in the last several entries is verified from the PC; the
app's UI (settings, history, deep-link enrollment) and `mtu.pmtud_up` remain device-only.
### throughput: a rate, plus the qualifier that makes it a measurement (server-v0.6.1 … v0.6.2)
A throughput test reports the *smallest* limit on the path — and the sender's own ceiling is one of
the candidates. If the server is asked for 50 Mbps and 50 Mbps arrives, the network was never the
constraint and "50 Mbps" says nothing about it. So `perf.throughput_udp` always carries
`limited_by` (duration | budget | rate | send_error) and `measures_network`, and a finding is
raised only when the path is actually implicated. The live run against fmr reports 20 Mbit/s with
`measures_network: false`, which is the correct and useful answer.
Loss is computed against the **sender's own count**, fetched from the observations API, not against
the requested rate. A receiver alone cannot tell "the network dropped it" from "the sender never
sent it", and guessing turns a healthy server-side limit into a phantom network fault. The server
keeps one summary per action rather than per-packet records — a ten-second run at 50 Mbps is half a
million packets, and a struct each would turn a measurement into memory exhaustion.
Sending is **paced**, on an absolute schedule. Unpaced would measure the server's NIC and the first
queue it meets, then collapse into loss that reads as a network fault; sleep-per-packet would
accumulate scheduler error and drift the rate down over a ten-second run.
Throughput gets its own grant budget sized from the request, so every *other* action stays bounded
at 8 MiB. When the byte cap binds before the clock does, the **duration is shortened and reported**
rather than the run being truncated: promising thirty seconds and delivering twenty-one is the same
information with a surprise attached, and it keeps "the clock ended the run" as the normal case —
the only case where the rate is a clean property of the path. That behaviour came out of a test
that failed honestly (30 s at 100 Mbps needs 375 MB against a 256 MB cap).
It is **opt-in** in the run config, default off. A 5-second run at 50 Mbps moves ~30 MB; on a
metered mobile connection that is the user's money, and a tool that spends it without being asked
is not one people keep installed.
#### The bug the live test found
The first live run delivered 104 packets and stopped after 50 ms. The grant's rate check exempted
the first 50 ms entirely, meaning to be lenient at startup — the effect was the opposite. A sender
could dump an unbounded burst into that free window, and the instant the check switched on it
compared those bytes against 50 ms worth of allowance and refused everything until real time caught
up. **Every short test passed** (downtrain sends 50 packets, big_send seven); every sustained send
died fifty milliseconds in.
Replaced with a token bucket (`allowance = burst + rate × elapsed`), which is smooth from t=0.
The burst is 100 ms of the allowed rate, floored at one ordinary datagram — deliberately one, since
at 8 kbps a 64 KB floor is sixty-four seconds' worth, exactly the instant dump the ceiling exists to
prevent. The pre-existing rate test caught that when I first tried the generous floor, and it was
right to. Second half of the same bug: callers treated *any* refusal as terminal, so `TryAllow` now
says why — a sender paces through a transient "too fast just now" and still stops dead on a spent
budget or an expired grant. Both halves are pinned by regression tests.
### Findings registry (2026-08-01)
Closes open item 1 of measurement-schema.md §9. A finding code is the stable, machine-readable half
of a result — what a dashboard groups by and what someone greps a year of archived runs for — and
that only holds if a code means exactly one thing forever. Ad-hoc string literals at fifteen call
sites cannot promise that, and by the time the registry was written the failure had already
happened.
**Two emitters had independently produced `connectivity.downstream_loss` and
`connectivity.loss_downstream` for the same claim**, and nothing anywhere objected. Anyone
aggregating either one would have silently seen half their data. Merged into
`connectivity.loss_downstream`, paired with `loss_upstream` so the two directions read as a set.
**Two codes were also renamed out of `nat.*`.** `nat.udp_unreachable` is not about NAT — it means
no replies came back — but the prefix determines the category, and the category determines which
verdict light the finding rolls up into (§7.3). A `nat.*` code landing under *connectivity* is not
a naming quibble; it changes which light turns red. Cheap to fix now, a breaking change later.
Codes are now declared as typed `FindingSpec`s carrying their category and default severity, and
emitters reference the spec instead of retyping the string — so a typo is a compile error and two
call sites cannot disagree about a finding's category.
`docs/findings-registry.md` is the contract, and a test reads it: it fails when the document and
the registry have codes the other lacks, or when a severity differs. Documentation that drifts from
its implementation is worse than none, because it still looks authoritative. The check scopes
itself to table rows, so the prose can keep explaining which codes were retired and why.
Six tests: uniqueness, declared-vs-listed, prefix↔category agreement, naming convention, a
word-order-anagram check (the shape the duplication actually took), and the document agreement.
### A real privacy leak, found by starting on the machine-readable schema (2026-08-01)
The intent was `measurement.schema.json` (§8's promised companion). The first step — checking
whether the anonymizer actually covers the fields the schema declares as sensitive — found that it
did not, so that became the work.
**At the `balanced` level, five identifying values were being uploaded verbatim:**
| value | field | why it matters |
|---|---|---|
| `2001:…::150` | `networks[].link.addresses[].addr` | the device's own global IPv6 address — a strong, geolocatable device identifier |
| `2a02:…::1` | `networks[].link.routes[].gateway` | identifies the ISP allocation |
| `203.0.113.77` | `networks[].link.dns.servers[]` | the configured resolver |
| `nas.example.lan` | `private_dns_hostname` | an internal hostname |
| `example.lan` | `search_domains[]` | the internal domain |
The settings screen describes that level as pseudonymizing addresses. It was not.
**Root cause:** classification keyed on field *names*, and the schema's actual names (`addr`,
`gateway`, `dst`, `servers`, `search_domains`, `private_dns_hostname`) had never been added to the
table. Not a subtle bug — just an unfalsifiable design. The existing tests all passed, because each
one checked a field somebody had remembered to write a case for.
**Two fixes, one of them structural:**
1. The missing names were added.
2. More importantly, a **shape-based backstop**: when a field name is unrecognised, the *value* is
inspected, and anything shaped like an IPv4/IPv6 address or a MAC is treated as one. A name
table can only protect fields someone thought of, which is precisely the wrong property for a
privacy control. Hostnames are deliberately *not* inferred by shape — `train.udp_updown` is
indistinguishable from a domain, and mangling a test type would corrupt the document to protect
nothing.
`LeakTest` is the new guard and is written to fail for fields nobody has considered: it plants
identifying values wherever one can actually occur and asserts none survive, rather than checking
a list of known cases. It also pins that RFC1918 addresses still come through readable, so the
test cannot pass by over-redacting everything.
Route prefixes and the unspecified address needed care in the transform: `0.0.0.0/0` and `::/0`
must stay themselves, or a routing table becomes unreadable for no privacy gain.
**Still outstanding:** `measurement.schema.json` itself. Worth noting what this episode implies for
it — much of a document's payload lives in `evidence`/`metrics`/`params`, which are per-test-type
`JsonObject` by design and therefore *outside* any schema. A schema-driven anonymizer would have
less coverage there than the name-plus-shape one now does, so the schema should be built for
validation and external tooling, not as a replacement for the classifier.
### ULA prefixes are pseudonymized whole (2026-08-01)
Spotted in a real uploaded run from the phone: the server had
`fda1:3fb1:ff92:6696::2662` for a DNS server. The general IPv6 path preserves the leading two
groups (deliberately — for a global address that keeps the ISP allocation, which is the
diagnostically useful part), and for a ULA that passed through **32 of the 40 random bits** of the
global ID.
ULA looks like the v6 equivalent of RFC1918 and the instinct is to treat it the same. That
reasoning does not carry over, and the difference is the whole point: an RFC1918 prefix is shared
by millions of networks and identifies none of them, while a ULA global ID is random and unique to
one network by construction (RFC 4193). The prefix *is* the identifier — it is a network
fingerprint that was surviving redaction.
Now pseudonymized as a unit, so two addresses on the same ULA subnet still land on the same
pseudonymous prefix: "these hosts are on one network" survives, "this is *that* network" does not.
Three tests, one of which uses the exact value observed on the wire.
Worth recording as a reasoning trap: I had originally raised this as "ULA should probably be kept
verbatim, like RFC1918, for consistency". The surface analogy pointed the wrong way, and the
correct answer was the opposite.
### Registry adopted everywhere; v6 findings renamed; Back works (2026-08-01)
The findings registry was only adopted in `core-engine`. The app module still emitted seven codes
as raw strings, so the registry test passed while codes existed outside it — including
`ipv6.broken`, which fired on a real network and was in no registry at all.
All seven now reference registry entries for code, category and severity, so those three cannot
disagree at a call site. A grep for `code = "…"` across the app, engine and probe modules returns
nothing.
**`ipv6.*` → `v6.*`.** The third instance of rule 1: they declared `Category.IPV6` while the prefix
map only knows `v6`, so `TestType.category("ipv6.broken")` fell through to *connectivity* and the
finding rolled up under the wrong verdict light. The test-type registry already used `v6.`.
Two severities reconciled while merging:
- `connectivity.captive_portal` is **medium**, not high. The registry had guessed high; the probe
that emits it had always said medium, and the probe was the considered value — a captive portal
on hotel wifi is what should be there, and logging in clears it. `connectivity.no_internet` is
the high one, because nothing the user does locally fixes that.
- `v6.not_offered` is **info, and the registry says it must stay info**. Most networks still do not
offer IPv6 and that is not a fault; a warning here lights a yellow verdict on a healthy network,
which teaches people to ignore the light.
Also: a `BackHandler` now returns from Settings/History to the run screen. The screen was a plain
state variable with nothing connecting it to the back stack, so the system Back gesture left the
app entirely. Enabled only when there is somewhere to go back to, so Back still exits from the run
screen.
### Upstream throughput (server-v0.6.3, 2026-08-01)
The mirror of the downstream case: the client generates the traffic and the server counts it. No
grant is involved — the client is sending its own packets, so there is nothing to amplify — but it
does need the server's tally, because **only the far end knows how much arrived**. Without that
number a sender measures how fast it can *transmit*, which is usually just the speed of the local
NIC and is a different question from the one being asked.
`TYPE_THROUGHPUT_UP` (0x0F) is counted and deliberately **never answered**: a reply would double
the traffic and drag the return path into a measurement that is specifically about the outbound
one.
The tally is a counter, not a list, and short-circuits **before** the observation log. A
five-second run at 20 Mbps is around ten thousand packets; one struct each would turn a
measurement into an allocation storm on a shared server, and nothing needs the per-packet detail
since the client holds the send-side record. The gap between the two counts is the loss.
`direction=up` on the throughput action sends nothing — it zeroes the counter, so a second run in
one session measures itself rather than inheriting the first one's packets. The live test asserts
`received <= sent`, which is what catches a counter that was never reset.
Live against fmr: **3125 sent, 3125 counted, 0 % loss, 10.0 Mbit/s** at a 10 Mbit/s request, with
`measures_network: false` — correct, since what arrived matched what was offered, so the path was
never the constraint.
### Raw shell dumps leaked the whole LAN (2026-08-01)
Found by running the Shizuku shell tier for the first time. The tier works — `tiers.shizuku: true`,
`exec_path: UserService` (so the UserService binds on the OnePlus, as recorded), `runs_as
shell(2000)`, 7/7 commands — and the run promptly uploaded **every MAC address on the local
network** to fmr at the `balanced` level: router, phones, whatever else was on the wifi. Fourteen
of them.
The probes embed raw command output verbatim (`ip neigh`, `ip route`, `id`), which is genuinely
good evidence and also a complete household device inventory. The anonymizer could not see it:
classification is by field name and by whole-value shape, and `ip_neigh` is one long string that is
itself neither a MAC nor an address. measurement-schema.md §9 item 2 had flagged raw dumps as "hard
to anonymize" and proposed dropping them from exports; nothing enforced either.
**Scrubbing beats dropping.** Identifiers inside any unclassified string are now replaced in place,
using the same pseudonyms as everywhere else — so a MAC that appears both in a parsed field and in
a raw dump still reads as one device. The dump stays readable and auditable: you can still see the
neighbour table's shape, the host count, RFC1918 addresses and vendor prefixes. Dropping the
evidence would have protected the same data while destroying the reason for collecting it.
Two implementation notes worth keeping:
- **One pass, not three.** Sequential passes re-process their own output: once a MAC became
`78:9a:18:xx:yy:zz`, the IPv6 pattern matched it — six hex groups separated by colons *is* an
address — and destroyed the vendor prefix the MAC rule had just preserved. Ordered alternation
resolves each position once, MAC first.
- The patterns are conservative on purpose. A missed address gets caught by another rule or not at
all; an over-eager one mangles timestamps and version strings, corrupting evidence to protect
nothing.
`RealDocumentTest` runs the anonymizer over a captured run when `ECHOLOT_REAL_RUN` points at one,
and fails on any MAC that survives. It self-skips otherwise, so no one's network is committed to the
repo. Against the actual leaked document: **14 MACs in, 0 surviving.**
Also fixed: the Settings *Preview what an upload would send* button did nothing. It read
`UiState.history`, which is empty until the History screen has been opened — the same root cause as
the "0 run(s)" count. It now reads the archive directly, and says so when there is nothing to
preview rather than silently ignoring the tap.
### Security: the admin listener was publicly exposed for ~15 minutes (2026-08-01)
Moving the admin listener to `[::2]:443` for the UI exposed `/admin/enroll-tokens` and
`/admin/selftest` to the internet **with no authentication**. Anyone who could reach
`fmr.echo-lot.app` could mint enrolment tokens.
The listener was designed localhost-only — its own flag help says *"keep localhost"* — and that
assumption travelled with it when the address changed. The compounding error: `checkAdminExposure`,
added the same day, verifies **encryption** and says nothing about **authentication**. It passed,
and a green light on an adjacent property is worse than no check, because it invites you to stop
looking.
Closed by returning to loopback (the TLS and ACME work is retained, just not exposed). All 68 device
enrolments matched the timestamps of test runs, so there is no evidence of abuse — but the window
existed on a freshly published hostname and absence cannot be proven. 39 unused enrolment tokens
were purged, since any could have been minted by someone else and they cost nothing to replace, and
63 test devices removed.
**The admin listener does not become reachable again until it authenticates.** That reorders the UI
work: auth on the listener first, everything else after.
### Open: encrypted uploads, where the operator cannot read the data
Not built. Recorded because the shape is decided by a few early choices, and the current design
happens to leave the door open.
The goal: hand someone an account, let them upload, and be unable to read what they uploaded.
Sketch: a random per-account **master key**, generated on the first device and wrapped under a
key derived from a passphrase (PBKDF2-HMAC-SHA256 — stdlib on both sides). The wrapped key is
stored server-side as an opaque blob, so a new device signs in, fetches it, and unwraps locally;
the server never sees either key. Runs are encrypted client-side with AES-256-GCM, fresh nonce per
run. All of this is stdlib in Go and `javax.crypto` in Kotlin — no dependency either side.
Four consequences that decide whether it is worth it:
1. **What stays readable determines what the UI can do.** The server builds its index by *parsing*
the document — verdict, finding count, started_at. An opaque payload means the client supplies
that metadata or the index disappears, and with it retention-by-verdict and any "runs with
findings" view. The honest version supplies only run id, timestamp and size, and moves the rest
client-side.
2. **Lose the passphrase, lose the data.** That is the feature working, and also the support
burden. It needs a recovery code printed at setup, not a reset flow — there is nothing to reset.
3. **Metadata is not hidden.** The operator still sees which account uploaded, when, how often and
how large. "Cannot see it" is about content, not existence, and saying otherwise would oversell.
4. **It makes `min_anonymization` unenforceable** — a server cannot check a level it cannot read.
That is not a conflict so much as a redundancy: the anonymization floor exists to protect the
user from the operator, and encryption does that better. The two should not both be demanded of
one upload.
What keeps this possible: uploads are already stored byte-for-byte as received, and every index
field is derived in one function (`runs.Put`). The thing to avoid is admin features that *require*
reading content — those would have to be unbuilt later.
### App sign-in, and an undisclosed dependency it surfaced (2026-08-01)
The app can now sign in to the server's identity provider: authorization code with PKCE, a
`Sign in` card in settings, and the `echolot://auth` redirect handled alongside the enrolment one
(told apart by host, since one spends a token and the other completes an authorization).
The detail that decides whether this works on a real phone: **the PKCE verifier is written to
storage before the browser opens**, not held in memory. Handing control to a browser backgrounds
the process and Android may kill it; the callback then arrives at a fresh process. An in-memory
verifier works on a developer's device and fails under memory pressure, which is the worst way for
a sign-in to break.
Nothing from the IdP is retained. The ID token proves who is signing in, once, and the device
credential authenticates everything after — no access tokens stored, no refresh tokens rotated.
**A server remains entirely optional.** All eight probes are device-tier; `serverConfigured` gates
only upload and the account. But answering that question exposed something worth fixing: two probes
hardcode the reference deployment —
```kotlin
DnsCanaryProbe(canaryZone = "c.echo-lot.app", ...) // "Hardcoded to the reference deployment"
StunProbe(serverHost = "fmr-1.echo-lot.app")
```
so a user with no server of their own still sends DNS and STUN traffic to fmr without being told.
For a tool that goes to this much trouble over what leaves the device, an undisclosed dependency on
a third party's infrastructure is the wrong default. It should prefer the configured server, and be
explicit when there is none. **Closed** — both probes take the enrolled server from settings
(canary zone learned from the profile, cleared on re-enroll) and report themselves SKIPPED with
the reason when none is configured.
### v6.broken was a false positive waiting to happen (2026-08-01)
A phone could not open `https://fmr.echo-lot.app` while loading the same server by IP literal
perfectly well. Two things came out of chasing it.
**The admin UI is IPv6-only, by consequence rather than intent.** `fmr.echo-lot.app` has an AAAA
and no A record — verified identical at Cloudflare, Google and Quad9, so DNS itself is healthy.
That follows from reserving all four measurement addresses for testing, which left only `::2` for
management, and `::2` has no IPv4 counterpart. Any client without working IPv6 sees an unreachable
admin interface — a poor property for the interface you reach *from the networks you are debugging*.
**And the app's own `v6.broken` finding was unsound.** It fired on exactly one signal — ICMPv6 echo
getting no reply — with `Confidence.HIGH`. ICMPv6 echo is widely filtered on networks where IPv6
works fine, which is precisely what that phone demonstrated: no ICMPv6 replies, working IPv6 TCP.
The finding asserted a cause it had no evidence for, which is the same class of error as the
multi-homed `100 % downstream loss` earlier: a confident measurement of something that was not
happening.
Now `v6.no_icmp_reply`, severity low, confidence medium, and the text names *both* explanations
instead of choosing one. It is still worth reporting, because filtered ICMPv6 breaks Path MTU
Discovery — large packets vanish rather than being reported as too big — which is a real fault even
when IPv6 works.
The proper fix is corroboration: attempt a real IPv6 connection and only call it broken when that
fails too. That needs a target, which runs into the hardcoded-reference-deployment issue already
open above. **Both closed 2026-08-02** — see "Corroborated IPv6 findings" below.
## Per-network probing is blocked while a VPN is up (2026-08-01)
`Network.bindSocket()` fails with `EPERM` for every underlying network when a VPN holds the
default route — verified on the OnePlus 15 with Netbird active: `Binding socket to network 101
failed: EPERM` for both cellular and wifi. This is Android preventing VPN leaks, not a bug to work
around, and it means the whole per-network measurement approach is unavailable to any user with a
VPN connected. Worth deciding deliberately rather than discovering per report:
- The run currently succeeds and simply measures nothing per network. Honest, but silent — the
document records `attempted: false` and the UI says green.
- A user with a corporate VPN permanently on would get a green run that measured almost nothing.
Options are to detect the VPN and say so plainly ("this network cannot be measured while a VPN is
active"), to measure the tunnel itself as the network under test, or both. **Decided and built
2026-08-02**: say so plainly, everywhere the run is read — see "Constrained runs" below.
Related: `icmp.ping6` now records `attempted` alongside `ok` per network, because collapsing them
made the app report "IPv6 is configured, but ICMPv6 gets no reply" about an interface it had never
succeeded in sending on — a claim about the user's carrier with no evidence behind it.
## Reserved measurement addresses, and the web UI on both families (2026-08-01)
fmr has two IPv4 (.150/.151) and three IPv6 (::150/::151/::2) addresses. `.150`/`::150` now carry
the services; `.151`/`::151` are reserved for measurement, declared in `ECHOLOT_RESERVED_ADDRS`.
Reserved does **not** mean silent. The UDP data plane, the canary DNS and STUN's RFC 5780 alternate
all belong there — reserving an address and then forbidding the measurements that need it would
defeat the purpose. What must never appear is a service, and above all not ports 80 or 443: a
handshake completing on a port known not to be listening is what proves interception, and that
proof survives exactly as long as nothing binds those ports. `config.CheckReserved` enforces it at
startup, refusing wildcard binds outright (every listener defaults to `:port`, so the next one added
will claim reserved addresses without anyone deciding to).
The first version of the guard was too strict and the live config caught it: it would have refused
the existing UDP and DNS binds on `.151`. The rule is about services and web ports, not about
listening at all.
**The adb-beacon receiver was wildcard-bound to `0.0.0.0:443`**, occupying port 443 on every IPv4
address including the reserved one — so the IPv4 interception test had been compromised for as long
as it had been running, silently. It is now `systemctl disable --now echolot-adb-beacon`; restore
with `systemctl enable --now`. Note what this implies: the guard covers this server's own listeners,
and a stray process outside its config can still pollute a reserved address. A startup probe that
*verifies* 80/443 are actually free on the reserved addresses would be a stronger guarantee than
checking our own configuration — built 2026-08-02 (`selftest.ReservedWebPortsFree`, fatal at
startup when anything is listening there).
The admin UI and the ACME responder now take comma-separated addresses like every other listener;
they were single-address, which is why the UI could only ever live on `::2`. It serves on `.150:443` and
`[::150]:443`; sshd on `.150:2322` and `[::150]:2322`.
`::2` is gone entirely — unbound, then removed from `/etc/systemd/network/ext.network`. The
transition kept it bound throughout and dropped it only after the CNAME landed, because removing it
first would have broken both the UI and ACME renewal for the very name the certificate is issued
to. Listeners came off before the address did, in that order, or the services would have failed to
bind on restart.
Verified after a full reboot: `fmr.echo-lot.app` answers 200 over both families, `.151`/`::151` are
closed on 80 and 443, canary DNS is still up on `.151`, and neither `::2` nor the beacon returns.
(`echolot-server` is `After=network-online.target` with `Restart=on-failure`, which is what makes
binding specific addresses safe across a boot — a wildcard bind would not have needed it, and that
is the trade for the reserved addresses being meaningful.)
The point of all this: `fmr.echo-lot.app` gained an A record, so the server stopped being reachable
only over IPv6 — which is what made it unreachable from a phone with no working IPv6, presenting as
"this host does not exist" in two different browsers.
### If the beacon comes back, it belongs in the web UI
Not as a separate listener. The receiver being its own Python service on `0.0.0.0:443` is exactly
what silently compromised the reserved address, and a second process racing for a port is a
recurring problem rather than a one-off: whoever loses the race simply fails to start, and on a
reboot which one that is comes down to unit ordering.
Folding it in costs little and settles several things at once. It would be two routes on the admin
UI (`POST` the observed adb port, `GET /apk` for the staged build), behind the TLS the UI already
terminates and the certificate it already renews, with no extra port and no wildcard. It also gets
authentication for free — the current receiver accepts a port report from anyone who can reach it,
which is tolerable for a dev tool on a trusted network and not something to keep once it lives
beside an admin session.
The one thing that changes on the device side is that the POST becomes HTTPS. That is a real
certificate rather than a self-signed one, so it costs a URL scheme rather than any trust plumbing.
## The control plane shares port 443 (2026-08-01)
`fmr-1.echo-lot.app:443` is the control plane, `fmr.echo-lot.app:443` the admin UI, both on
`.150`/`::150`, one listener, selected by SNI for the certificate and by `Host` for the handler.
The reason is not tidiness, it is reachability. Captive portals, hotel wifi and corporate firewalls
routinely permit only 80 and 443 — which is exactly the population of networks this tool exists to
diagnose. A control plane on 8443 is unreachable precisely when it matters most, and it fails as
"cannot reach server", which tells the user nothing.
They cannot share a certificate, which is why this needs two names. The control plane is trusted by
SPKI pin and so uses a long-lived self-signed certificate; a browser needs one a CA vouches for.
One name on one port is one certificate, so the port can only be shared by splitting the names.
Pinning the Let's Encrypt key instead was considered and rejected: it survives renewal only while
key reuse holds, so a routine key rotation would brick the whole fleet.
Verified per SNI on 443: `fmr.echo-lot.app` serves `issuer=Let's Encrypt`, `fmr-1.echo-lot.app`
serves the self-signed cert whose pin is unchanged (`zRV9…Xlg=`), `/v1/profile` answers 401 on the
control name and 303 to the login page on the UI name.
**8443 stays open.** Devices enrolled before this carry that URL in their settings, and closing it
for the sake of a port number would strand every one of them. It can go once no enrolled device
still points at it — not before.
The rule from the naming change still binds: `fmr` may be a CNAME to exactly one host and never a
multi-address record, because a pinned client that reaches a different key does not fail over.
## Constrained runs: a VPN'd run now says so, everywhere (2026-08-02, app 0.2.1)
The measurement schema gained a top-level `constraints` block (§3) and the app now fills it.
`ConstraintDetector` (core-probe) runs before any probe: one throwaway `Network.bindSocket()` per
non-VPN network, plus a transport check for an active VPN. The result lands in three places, and
all three are deliberate:
- **`run.constraints`** — for machines. A server aggregating thousands of runs can now separate
"measured a healthy network" from "measured almost nothing through a tunnel"; the shapes were
identical before.
- **A `measurement.vpn_constrained` finding** — for the person reading this run, naming the
interfaces that went unmeasured. A constrained run with a quiet findings list still reads as
"nothing wrong here".
- **The §7.3 verdict** — `Verdicts.derive` takes the constraints and returns INCONCLUSIVE
outright for a per-network-blocked run, whatever the category lights say; the run screen shows
an amber "Measured through a VPN" banner above the verdict so INCONCLUSIVE reads as the OS
refusing, not the app failing.
Detection is one bind per network rather than parsing per-test `attempted:false` breadcrumbs, so
it cannot drift when probe evidence formats change.
## Corroborated IPv6 findings: v6.broken is back, with evidence (2026-08-02, app 0.2.1)
The new `V6ConnectProbe` (test type `v6.brokenness`) attempts a real TCP connection over IPv6 to
the configured server's :443, per network that *claims* IPv6 (global address or v6 default
route) — IPv4-only networks are not attempted, since their failure is by design and would
manufacture the exact false positive this exists to kill. The finding derivation is now three-way:
- ICMPv6 silent, TCP works → `v6.no_icmp_reply` at **high** confidence, retitled "ICMPv6 is
filtered here — IPv6 itself works" (still reported: filtered ICMPv6 breaks PMTUD).
- ICMPv6 silent, TCP fails too → **`v6.broken`** (high severity, reinstated in the registry +
findings-registry.md): two independent transports silent on a network advertising IPv6.
- No corroboration (no server configured, or the connect never got as far as sending) → the
two-explanation `v6.no_icmp_reply` at medium confidence, unchanged.
Like STUN and the canary, the probe SKIPs honestly when no server is configured — corroboration
is a benefit of enrollment, not a reason to borrow fmr.
## Server: reserved 80/443 verified against the OS, and signed releases (2026-08-02)
**Reserved-address startup probe.** `serve()` now proves 80/443 are actually free on every
`ECHOLOT_RESERVED_ADDRS` address before starting: a throwaway bind per port
(`selftest.ReservedWebPortsFree`), fatal on EADDRINUSE with the offending address named — the
check `CheckReserved` cannot do, because a stray process outside our config (the adb-beacon
receiver on `0.0.0.0:443` was exactly that) is invisible to configuration checks. Bind errors
that are not "in use" (typo'd address, address not on this host) warn instead of refusing —
they are config problems, not pollution.
**Release signing.** Self-update now trusts a signature, not a host. CI signs `SHA256SUMS` with
an ed25519 key (`relsign` package, `cmd/release-sign`) and the updater refuses any release whose
`SHA256SUMS.sig` is missing or does not verify against the public key baked into the binary
(`selfupdate.DefaultPublicKeyB64`; operators with their own pipeline override via
`ECHOLOT_SELF_UPDATE_PUBKEY`). The private key exists in exactly two places: the Gitea Actions
secret `RELEASE_SIGNING_KEY`, and the offline original on the dev PC at
`~/.echolot/release-signing-key`. It is deliberately NOT on fmr and NOT in the repo — a
compromised release host can withhold updates but no longer inject one. CI hard-fails when the
secret is missing (an unsigned release would strand every verifying server) and cross-checks the
signature against the key in the source it just built.
**ACTION REQUIRED before the next `server-v*` tag:** add the Gitea repo secret
`RELEASE_SIGNING_KEY` (Settings → Actions → Secrets) with the contents of
`~/.echolot/release-signing-key` from the dev PC. Ordering is safe: the currently deployed
v0.3.x updater does not verify, so it will happily install the first signed release; every
release after that is verified. **Done 2026-08-02** — the secret is in place. (The "v0.3.x"
above should read "the currently deployed release": deployments had moved on to v0.9.x by the
time signing landed; the point — the deployed updater predates verification and will accept the
first signed release — is unchanged.)
## Prober fold: traceroute.udp4 and the mDNS inventory go production (2026-08-02, app 0.2.2)
The two highest-value validated capabilities moved from the prober into `core-probe`:
- **`traceroute.udp4`** (`TracerouteProbe`): UDP traceroute reading ICMP time-exceeded off the
socket error queue via `Os.recvmsg(MSG_ERRQUEUE)` through the reflection facade — no root, no
raw socket, no JNI, ~250 ms for six hops. Emits the schema's `TracerouteEvidence` (rtt in ns).
`OsAbi` came with it, including the measured fact that `Os.getsockoptInt` exists on neither
known device, so PMTU must always be read from the errqueue (`ee_info`), never
`getsockopt(IP_MTU)`. The load-bearing line survived the port: EAGAIN out of the reflected
`recvmsg` means "queue empty", not failure.
- **`local.mdns_inventory`** (`MdnsInventoryProbe`): MulticastLock + NSD discovery, the service
inventory that doubles as the VLAN-leakage detector. Both hardware lessons kept: the
`_services._dns-sd._udp.` meta-query returns 0 beside live services on both devices (so the
concrete types are the measurement and the meta-query result is itself evidence), and the
listen window is 10 s because 4 s missed services.
Still to fold, in order: the Shizuku dump *parsers* (the raw `link.ip_monitor` captures already
hold two divergent vendor formats that could feed `link.ra_source` and `sec.arp_watch`);
`multinetwork.request_and_bind` (extend ConstraintDetector to *request* transports rather than
only probing present ones); `peer.ble_advertise` (needs three new permissions and a peer mode to
exist first).
## Server v0.9.2: trains, real TTL/DSCP/ECN, rate limits (2026-08-02)
The spec-vs-implementation gap audit closed its top items; protocol_version 1.0.0 → 1.0.1
(additive — below 1.0.0 the minor is the breaking axis, and nothing here breaks an old client):
- **Upstream trains** (§3.2, types 0x03/0x04/0x05): per-train bounded columnar buffer (8192
rows, head kept on overflow with `Truncated` set — mirrors the schema's `evidence_truncated`
honesty), TRAIN_REPORT split across ≤1200-byte datagrams, grant-free with the §3.4 argument
spelled out (a 17-byte report row answers a ≥36-byte HMAC-valid packet). Unknown train id
gets a zero-row report: "nothing arrived" is an answer. Also surfaced as `udp.trains` in the
observations API.
- **Real TTL/DSCP/ECN observation** (§3.3): the read loop is `ReadMsgUDPAddrPort` with
IP_RECVTTL/IP_RECVTOS/IPV6_RECVHOPLIMIT/IPV6_RECVTCLASS cmsgs on Linux; `0xFF` stays the
"not observed" sentinel elsewhere. This unblocks `sec.dscp_ecn_survival` both directions,
paired with the new `dscp` parameter on `downtrain` (validated 063, refused not clamped,
`dscp_applied` in the response).
- **Rate limiting** (§2.5, was entirely absent): token buckets keyed per credential AND per
source IP; 429 + Retry-After on session/action creation (`/v1/profile` stays ungated), silent
drop on the data plane — charged after the HMAC gate so a spoofed flood cannot drain a
victim's budget, before the replay window so a dropped seq stays usable. UDP ceilings default
above the largest legitimate run (a 200 Mbps throughput test), because a rate limit that
clips a real measurement produces a confidently wrong number.
- **`action_id` in every granted packet** (§5/§9): payload bytes [8:16] across all granted
types, so overlapping actions are attributable. Verified the deployed Kotlin client parses
only ECHO_RESP and MTU_ACK payloads, so the reshuffle strands nobody.
- **Canary log retention**: the stated 24 h privacy default is now enforced
(`ECHOLOT_DNS_LOG_RETENTION_H`), where before the log was time-unbounded.
- **`POST /admin/enroll-tokens`** now answers the spec's JSON shape under content negotiation;
the README's curl works as documented.
- Spec §2.3 registry gained `downtrain` and `tcp-echo`, which the server had been advertising
as strings a conformant client must ignore.
Client-side counterparts still to build: sending 0x03 trains + parsing 0x05 reports
(`train.udp_updown`), and passing `dscp` on downtrain actions.
## ⚠ Version lineage broken: fmr runs v0.11.2, the repo's tags stop at v0.9.x (2026-08-02)
Discovered while preparing to self-update fmr to the freshly released server-v0.9.2:
**fmr runs v0.11.2** (binary installed 2026-08-02 08:51), but this repo's remote has tags only
up to `server-v0.9.1`, master fast-forwarded cleanly from this machine, there is no v0.10/v0.11
release in Gitea, no source checkout or Go toolchain on fmr, and no deploy script in this repo
that stamps versions. Conclusion: v0.11.2 was cross-built from a clone whose commits were never
pushed — presumably another dev machine.
Consequences until resolved:
- **Do NOT run `--self-update` on fmr.** Gitea's `/releases/latest` is the *newest-created*
release, which is now `server-v0.9.2` — semantically older than the deployed binary; the
updater compares strings, not SemVer, and would happily "update" v0.11.2 down to it. No
automatic risk exists (fmr has no update timer installed, only the cert timer), but a manual
run would downgrade.
- The next real release must be tagged **above v0.11.2** (e.g. `server-v0.11.3` or `v0.12.0`)
*after* the missing commits are pushed, so "latest" becomes truly latest again.
- The unpushed v0.10v0.11 work needs to be found and pushed from whichever machine built it,
or the deployed binary's provenance re-established some other way, before the release channel
can be trusted again.
**Resolved same day.** The binary itself settled it: `go version -m` on the deployed executable
shows `vcs.revision=d5b1bab` — a commit on this repo's master — built 2026-08-02 08:36 UTC with a
hand-stamped `-X main.Version=v0.11.2` and a dirty tree (`vcs.modified=true`, the then-uncommitted
schema doc). An earlier session stamped release numbers ahead of the tag line; no code was ever
missing. Current master is tagged and released as **server-v0.11.3** (signed), restoring a
monotonic, tag-backed lineage above the deployed number. The rule going forward: **the version a
binary is stamped with must be a pushed `server-v*` tag** — an ad-hoc stamp above the tag line
poisons `/releases/latest` for the string-comparing updater the moment anyone tags honestly again.
The stale `server-v0.9.2` release (same code lineage, wrong number, created during the confusion)
remains in Gitea but is harmless now that v0.11.3 outranks it as latest.
## LLDP and CDP are root-tier, and that is a hard boundary (2026-08-02)
Asked for alongside SSDP in long mode; they belong to a different tier and no amount of app-side
cleverness moves them. LLDP is an EtherType `0x88CC` frame to `01:80:C2:00:00:0E`; CDP is an
LLC/SNAP frame to `01:00:0C:CC:CC:CC`. Neither is IP, so neither is ever delivered to a socket an
app can open — receiving them needs `AF_PACKET` with `CAP_NET_RAW`, which is root. Shizuku does
not bridge this either: the ADB shell user (uid 2000) has no `CAP_NET_RAW`, and stock devices do
not ship `tcpdump`. Android's unprivileged ICMP sockets are what make `icmp.ping4` work without
root; there is no equivalent back door for raw L2 receive.
Worth building in the root module when it lands, because the payoff is large: LLDP names the
switch, the port and the VLAN a device is attached to, which is the best available answer to
"where in this building am I actually plugged in", and CDP does the same on Cisco gear. Until
then they are recorded as absent capabilities rather than left to look unimplemented.
What IS reachable at app tier, and what long mode now listens for instead: SSDP (passive NOTIFY
plus periodic M-SEARCH), LLMNR, NetBIOS-NS and WS-Discovery — all IP multicast/broadcast, all
sockets an app may open. The security reading matters as much as the inventory: LLMNR and
NetBIOS-NS being live on a segment is a finding in itself, since both are trivially spoofable.
**One of those four cannot run at app tier either, and says so.** `local.netbios_inventory`
reports `unsupported` with the bind error attached: UDP 137 is below 1024, and Android reserves
privileged ports exactly like any other Linux. The decoder, the evidence shape and the registry
id are built and tested, waiting for the Shizuku tier to supply a socket. Recorded as a result
rather than dropped, so nobody later reads its absence as an oversight.
Two deliberate restraints in that work, both worth keeping: no active WS-Discovery Probe (an
M-SEARCH is traffic every SSDP device expects constantly, whereas a WSD Probe from an unknown
host announces *this* device to the segment), and no NBSTAT sweep (that is host scanning, not
measurement — the app listens to what a network broadcasts, it does not interrogate its
neighbours). The parsers are also hardened against the input they will actually meet: a DNS
compression pointer is refused rather than followed (the classic parser hang), and the WSD
extractor is string-based on purpose, tested against an entity bomb and 20 000-deep nesting.
## Design note: what BLE between two devices is actually for (2026-08-02, not built)
Two or more phones running Echolot, talking over Bluetooth LE. The schema already anticipates
this — `Trigger.PEER` and the whole `peer.*` test family (`peer.reachability`, `peer.isolation`,
`peer.multicast`, `peer.lan_train`, `peer.lease_diff`) are in the registry, unused — and the
prober measured `peer.ble_advertise` **SUPPORTED on both known devices**, so the mechanism is
proven; what has been missing is a reason that beats "use the server".
**The reason is that BLE is out-of-band.** Everything else this app does depends on the network
under test being at least partly functional. A second device reachable over a radio that shares
nothing with the wifi turns several measurements from ambiguous into conclusive:
1. **Client isolation becomes measurable at all.** Today, "I sent a packet to the peer and heard
nothing" cannot distinguish AP client isolation from the peer being asleep, gone, or on a
different VLAN — the failure mode is silence, and silence has too many parents. With BLE the
peer confirms out-of-band that it was listening on address X at time T, so silence over IP
becomes *proof* of isolation rather than a guess. This is the single strongest argument for
the feature, and it mirrors the rule this project keeps rediscovering: a measurement that
cannot separate "nothing happened" from "nothing was tried" is not a measurement.
2. **Differential diagnosis: the network or this phone?** Two devices on the same SSID, one
resolving DNS and one not, settles in seconds what a single device cannot settle at all —
and it is the same distinction `system_verdict` exists to draw, only with a second opinion
instead of Android's. Natural finding: *this device fails where a peer on the same link
succeeds* → look at the device (private DNS, ad blocker, per-client router rule, MAC
randomization), not the router.
3. **Two DHCP servers on one L2**, the classic invisible fault: peers compare lease source,
subnet and gateway (`peer.lease_diff`). Disagreement is conclusive and needs no server.
4. **Coverage and roaming**, later: several devices sampling RSSI in different rooms, exchanging
summaries over BLE, gives a picture no single device standing in one place can produce.
**What crosses the link is a summary, never the document.** A measurement document describes
someone's home network in detail; broadcasting it to whoever is nearby would betray the whole
posture of §8. The peer payload should be: a *hashed* network identity (so two devices can agree
they are on the same L2 without either putting the SSID/BSSID on the air in the clear), the §7.3
category verdicts, the finding codes, and an IP endpoint plus a one-shot nonce for the LAN tests.
Findings and verdicts are already the interpretation layer — exactly the right granularity to
share.
**Privacy constraints, which are not optional here.** A BLE advertiser is a tracking beacon: it
must be user-initiated, time-boxed to the run, carry no identifier that is stable across runs
(the resolvable-private-address default plus a per-session ephemeral id), and pair by a code the
two humans can see. "Discoverable by default" would make this app a worse citizen than the
networks it audits.
**Deliberately not doing:** clock synchronisation over BLE. GATT latency is jitter measured in
tens of milliseconds, which is the same order as the one-way delays worth measuring; peers should
sync against the server's `time.server_offset` and use BLE only to correlate run ids. Nor should
BLE become a transport for uploads — it is a *comparison* channel.
Staging when it happens: `peer.isolation` first (highest value, needs only advertise + connect +
a nonce exchange), then `peer.lease_diff` (pure summary comparison, no extra plumbing), then the
rest. Needs `BLUETOOTH_ADVERTISE/CONNECT/SCAN` in the manifest, which the app does not yet
request.
## v0.11.3 live on fmr; trains validated end to end (2026-08-02)
Deployed via `--self-update` (the pre-signing v0.11.2 updater accepted the first signed release,
as planned; every later update verifies). Startup clean on the real host — the reserved-port
check passed against the OS, self-test green, capabilities unchanged plus the new machinery.
**`train.udp_updown` validated against production**: `LiveUpstreamTrainTest` from this PC sent
120 packets; the server's ledger counted 120, both columnar report parts arrived, loss 0.0 %,
`truncated=false`. The 0x03/0x04/0x05 path works, client and server, over the real internet.
Two operational bugs surfaced doing it:
- **CLI-minted tokens are lost while the daemon runs.** `devices.json` is loaded once at startup
and held in memory; `--mint-enroll-token` writes to disk, the running daemon never re-reads,
answers "unknown token", and clobbers the token on its next write. `enroll-link.sh` has only
ever worked by timing luck. Workaround used: mint, `systemctl restart echolot-server`, then
redeem. Real fix belongs server-side (re-read on miss, or route the CLI mint through the
running daemon).
- **`test-fmr.sh` still mints against `127.0.0.1:8444`**, which no longer exists (the admin API
moved to authenticated :443). Needs the same CLI-mint flow enroll-link.sh uses — plus the
restart caveat above until that bug is fixed.
+133
View File
@@ -0,0 +1,133 @@
<!--
SPDX-FileCopyrightText: 2026 Echolot contributors
SPDX-License-Identifier: CC-BY-4.0
-->
# Echolot findings registry
Closes open item 1 of `measurement-schema.md` §9.
A **finding code** is the stable, machine-readable half of a result. The prose around it changes
freely; the code is what a dashboard groups by, what a diff between two runs keys on, and what
someone greps a year of archived runs for. That only works if a code means exactly one thing,
forever.
This document is the contract. It is kept in step with
`echolot-app/core-measurement/.../FindingRegistry.kt` by a test that fails when either side has a
code the other does not — a registry that drifts from its documentation is worse than none,
because it looks authoritative.
## Rules
1. **The prefix determines the category**, and the category determines which verdict light the
finding rolls up into (§7.3). A `nat.*` code appearing under *connectivity* is not a naming
quibble; it changes which light turns red. Two codes were renamed from `nat.*` to
`connectivity.*` for exactly this reason.
2. **One code per concept.** Two emitters independently produced `connectivity.downstream_loss`
and `connectivity.loss_downstream` for the same claim before this registry existed. Anyone
aggregating either would have silently seen half their data.
3. **Codes are declared, not typed.** Emitters reference a `FindingSpec`, so a typo is a compile
error and no two call sites can disagree about a finding's category or default severity.
4. **Severity in the registry is the default.** An emitter may escalate for a specific run; it may
not quietly reclassify the finding in general.
5. **Say what is ruled out**, where that is the useful half. "Loss upstream" is worth far more
when it also states that the return path is clean, because that halves where to look next.
6. **Renaming a code is a breaking change** once runs are archived at scale. Before 1.0 it is
cheap; after, it needs an alias and a deprecation window.
## Registry
### connectivity
| code | severity | means | rules out |
|---|---|---|---|
| `connectivity.udp_unreachable` | high | No UDP echo replies came back from the server at all. | — |
| `connectivity.udp_unreachable_upstream` | high | The server received none of the probes, so traffic is dropped on the way out. | The return path: nothing arrived to be replied to. |
| `connectivity.udp_loss` | medium | A large fraction of round-trip probes were lost, direction unknown. | — |
| `connectivity.loss_upstream` | medium | Probes were lost on the way to the server. | The return path: replies came back for everything that arrived. |
| `connectivity.loss_downstream` | medium | Packets were lost on the way back from the server. | The outbound path: the server received what it was answering. |
| `connectivity.downstream_blocked` | high | Server-initiated packets never arrive, although round trips work. | Basic reachability: the path forwards replies, just not unsolicited traffic. |
| `connectivity.downstream_reorder` | low | Downstream packets arrive in a different order than they were sent. | — |
| `connectivity.captive_portal` | medium | A captive portal is intercepting connectivity checks. | — |
| `connectivity.no_internet` | high | Android's own connectivity checks fail on this network. | — |
| `connectivity.link_flapping` | medium | A network dropped and came back one or more times during the run. | A momentary probe failure: the drop was watched happening, not inferred from silence. |
`connectivity.link_flapping` is only reachable from a **long run** (`run.mode: "long"`,
measurement-schema §3). It is derived from `networks[].changes[]` rather than from any test's
evidence, because no one-shot probe can produce it: the probes before and after a four-second drop
both succeed. The emitter escalates to *high* from three completed drop-and-return cycles, and
requires the cycle to complete — a network switched off partway through a run is not flapping.
### mtu
| code | severity | means | rules out |
|---|---|---|---|
| `mtu.reduced_downstream` | low | The downstream path MTU is below the usual 1500 bytes. | — |
| `mtu.downstream_blackhole` | medium | Datagrams above the path MTU are dropped downstream, fragmented or not. | — |
| `mtu.fragments_blocked` | medium | IP fragments do not reach this device even when sent in order. | — |
| `mtu.fragment_reorder_sensitive` | low | Fragments are delivered in order but dropped when reordered or delayed. | Fragmentation itself: in-order fragments arrive fine. |
### nat
| code | severity | means | rules out |
|---|---|---|---|
| `nat.udp_rebinding` | medium | A NAT remapped the UDP source port mid-flow. | — |
| `nat.symmetric` | medium | The NAT assigns a different external port per destination. | — |
### perf
| code | severity | means | rules out |
|---|---|---|---|
| `perf.throughput_no_delivery` | high | No throughput traffic arrived, although the server sent it. | — |
| `perf.throughput_below_offered` | low | Less throughput arrived than the server sent for the whole run. | — |
### dns
| code | severity | means | rules out |
|---|---|---|---|
| `dns.answer_rewritten` | high | A resolver returned an answer that differs from the authoritative record. | — |
| `dns.authoritative_unreachable` | medium | The canary zone's authoritative server could not be reached. | — |
### v6
The prefix is `v6.`, matching the test-type registry (`v6.brokenness`, `v6.happy_eyeballs`, …).
These were `ipv6.*` while declaring `Category.IPV6`; since the prefix map only knows `v6`, they
rolled up under *connectivity* instead — the third occurrence of rule 1 being broken.
| code | severity | means | rules out |
|---|---|---|---|
| `dns.search_domain_unanswered` | high | The network advertises a DNS search domain that its own server does not answer for. | A fault on this device: the same server answers ordinary names normally. |
| `dns.system_resolver_broken` | high | The network's DNS server answers, but this device cannot resolve names through it. | A network fault: the server replied to a query sent from this device. |
| `measurement.vpn_constrained` | info | A VPN was active, so the networks underneath it could not be measured. | Nothing — this run says little about the underlying network either way. |
| `v6.no_default_route` | medium | The device has a global IPv6 address but no IPv6 default route. | Guesswork: this is read from the routing table, not inferred from silence. |
| `v6.route_without_address` | medium | The network advertises an IPv6 default route but the device has no global IPv6 address. | A working IPv6 setup: SLAAC did not produce a usable address on this link. |
| `v6.no_icmp_reply` | low | IPv6 is configured but ICMPv6 echo gets no reply. | Nothing on its own: IPv6 may work fine with ICMP filtered. |
| `v6.broken` | high | IPv6 is advertised on this network but carries no traffic. | ICMP filtering as the benign explanation: a TCP connection over IPv6 failed too. |
| `v6.not_offered` | info | This network does not offer IPv6. | — |
`v6.no_icmp_reply` was `v6.broken` until a phone reported it while loading an IPv6-only site over
TCP perfectly well. The only evidence behind it is ICMPv6 echo, which is widely filtered on
networks where IPv6 works — so the finding now states what was observed and names both
explanations instead of choosing one. It is still worth reporting: filtered ICMPv6 breaks Path MTU
Discovery.
`v6.broken` returned once that corroboration existed: the `v6.brokenness` test attempts a real TCP
connection over IPv6 to the configured server, and only when *both* transports fail on a network
that advertises IPv6 is the brokenness claim made — at high severity, because every dual-stack
destination pays a timeout before falling back to IPv4. When the TCP connect *succeeds*,
`v6.no_icmp_reply` is emitted at high confidence instead, now able to say plainly that ICMPv6 is
filtered while IPv6 works. With no server configured there is no corroboration target and the
two-explanation `v6.no_icmp_reply` stands unchanged.
`v6.not_offered` is **info and must stay info**. Most networks still do not offer IPv6 and that is
not a fault; reporting it as a warning lights a yellow verdict on a healthy network, which teaches
people to ignore the light — the one thing a diagnostic must never do.
## Adding a finding
1. Add a `FindingSpec` to `FindingRegistry`, and to its `all` list.
2. Add the row here, under the section its prefix names.
3. Emit it with `finding(FindingRegistry.YOUR_CODE, …)`.
The registry test checks 1 and 2 agree, that every prefix maps to the category it claims, and that
no two entries share a code.
+52 -3
View File
@@ -38,6 +38,7 @@ Export encoding: UTF-8 JSON, gzip for files (`.echolot.json.gz`), share intent u
{ {
"id": "0198c5f2-...-uuidv7", "id": "0198c5f2-...-uuidv7",
"trigger": "manual | scheduled | monitor | peer", "trigger": "manual | scheduled | monitor | peer",
"mode": "short | long",
"started_at": "2026-07-29T14:03:21.114Z", "started_at": "2026-07-29T14:03:21.114Z",
"ended_at": "2026-07-29T14:07:44.902Z", "ended_at": "2026-07-29T14:07:44.902Z",
"clock": { "clock": {
@@ -52,12 +53,47 @@ Export encoding: UTF-8 JSON, gzip for files (`.echolot.json.gz`), share intent u
}, },
"tiers": { "app": true, "shizuku": true, "root": false }, "tiers": { "app": true, "shizuku": true, "root": false },
"profiles_used": ["profile-uuid", ...], "profiles_used": ["profile-uuid", ...],
"constraints": {
"vpn_active": true,
"per_network_blocked": true,
"unmeasured_networks": ["net-0", "net-1"]
},
"notes": "free-text user annotation" "notes": "free-text user annotation"
} }
``` ```
`tiers` records what was *available*; each test records what it *used*. `tiers` records what was *available*; each test records what it *used*.
`mode` records how long the run watched, and it exists because **it changes what a reader may
conclude from absence**. A `short` run is a sequence of one-shot probes — each looks at the network
for a second or two and moves on — which characterises the network's *configuration* well and is
structurally blind to anything intermittent. A `long` run starts continuous listeners at t=0, runs
the same battery beside them, and keeps sampling until its window closes; the window's length is
recorded in the `params` of the tests the listeners produce, not here.
The consequence is asymmetric and matters more than the field looks. A finding is worth the same in
either mode: a drop that was observed, was observed. Silence is not. "No link changes were seen" is
evidence of a stable link after five minutes of watching and is evidence of nothing at all after a
thirty-second run, in which a link could drop and return between two consecutive probes without
leaving a mark anywhere in the document. Consumers — a diff between two runs, a dashboard counting
how often a fault occurs, a person reading one report — must therefore not treat the absence of a
time-dependent finding in a `short` run as its refutation, and must not compare the two modes as if
they had asked the same question. `connectivity.link_flapping` is the first finding that only a
`long` run can reach; `networks[].changes[]` (§4) is likewise populated only by a long run's
listener, and an empty `changes[]` in a short run means "not watched", never "nothing happened".
Absent `mode` means `short`: it was added after the first documents were written, and every one of
them was a battery of one-shot probes.
`constraints` records what was *prevented*. A constrained run is neither a failed run nor a normal
one, and the distinction has to survive into the data: a run taken through a VPN has the same shape
and the same green verdict as a clean run of a healthy network, so without this a reader — or a
server aggregating thousands of them — cannot tell that almost nothing was measured. The known case
is `per_network_blocked`: Android refuses `Network.bindSocket()` on the underlying networks while a
VPN holds the default route, so every per-network test measures the tunnel or nothing at all, and
any conclusion about the link underneath is unfounded. Consumers should treat findings from a
constrained run as scoped to what was actually reachable, and `unmeasured_networks` names the rest.
## 4. `networks[]` — one entry per Android `Network` in play ## 4. `networks[]` — one entry per Android `Network` in play
A run may exercise several networks simultaneously (Wi-Fi + cellular + USB ethernet). Everything is a snapshot at run start; a `changes[]` list captures mid-run deltas. A run may exercise several networks simultaneously (Wi-Fi + cellular + USB ethernet). Everything is a snapshot at run start; a `changes[]` list captures mid-run deltas.
@@ -97,12 +133,24 @@ A run may exercise several networks simultaneously (Wi-Fi + cellular + USB ether
"changes": [ "changes": [
{ "at_mono_ns": 91000000000, "kind": "lost | gained | link_changed", { "at_mono_ns": 91000000000, "kind": "lost | gained | link_changed",
"detail": { /* new link snapshot or diff */ } } "detail": { /* new link snapshot or diff */ } }
] ],
"app_usable": true
} }
``` ```
`routes[].proto` and lifetime fields are Shizuku-tier data (`ip route`/`ip addr`); app-tier snapshots leave them absent — absence means "not observed", never "not present". `routes[].proto` and lifetime fields are Shizuku-tier data (`ip route`/`ip addr`); app-tier snapshots leave them absent — absence means "not observed", never "not present".
`app_usable` records whether an ordinary app may send on this network at all. Android lists the
carrier's special-purpose networks — IMS/VoLTE, MMS, XCAP — alongside the real ones, and they
carry neither `INTERNET` nor `NOT_RESTRICTED`; binding to one needs
`CONNECTIVITY_USE_RESTRICTED_NETWORKS`, which is signature-level and unobtainable for a normal
app. Those networks are therefore permanently unmeasurable, and that is a property of Android's
permission model rather than of the link. They stay in `networks[]` because they are genuinely
present — an interface silently missing from the inventory is its own kind of lie — but a
consumer must not read the absence of tests against them as a fault, and they are **not**
`constraints.unmeasured_networks` (§3): nothing was prevented, the run was never entitled to
measure them.
## 5. `server_sessions[]` ## 5. `server_sessions[]`
```json ```json
@@ -269,7 +317,7 @@ The JSON Schema (machine-readable companion, `measurement.schema.json`, generate
| type | example fields | v2 anonymizer transform | | type | example fields | v2 anonymizer transform |
|---|---|---| |---|---|---|
| `ip4`, `ip6` | addresses, routes, hops, DNS answers | prefix-preserving pseudonymization, consistent per document; well-known/reserved ranges kept verbatim | | `ip4`, `ip6` | addresses, routes, hops, DNS answers | prefix-preserving pseudonymization, consistent per document; well-known/reserved ranges kept verbatim. **Exception: ULA (`fc00::/7`) has its whole prefix pseudonymized as a unit.** It resembles RFC1918 but is not analogous: a ULA global ID is 40 random bits, unique to one network by construction (RFC 4193), so the prefix *is* the identifier, whereas `192.168.0.0/16` is shared by millions of networks and identifies none. Pseudonymizing it as a unit keeps "these hosts are on one subnet" while dropping "this is that subnet". |
| `mac`, `bssid` | wifi, arp_watch | OUI kept, NIC part pseudonymized | | `mac`, `bssid` | wifi, arp_watch | OUI kept, NIC part pseudonymized |
| `fqdn` | DNS names, reverse lookups | per-label pseudonyms, public-suffix kept | | `fqdn` | DNS names, reverse lookups | per-label pseudonyms, public-suffix kept |
| `ssid` | wifi | pseudonym | | `ssid` | wifi | pseudonym |
@@ -279,7 +327,8 @@ Free-text fields (`notes`, `error.detail`, dump excerpts from Shizuku parsers) c
## 9. Open items ## 9. Open items
1. Findings registry document — start alongside the first implemented tests. 1. ~~Findings registry document~~ — done: `findings-registry.md`, kept in step with
`FindingRegistry.kt` by a test that fails when the two disagree.
2. Whether Shizuku raw-dump excerpts (dumpsys/ip output) are embedded in `evidence` verbatim (auditable, but large and hard to anonymize) or parsed-only with an optional "attach raw dumps" toggle. Proposal: toggle, default on for local archive, default off for export. 2. Whether Shizuku raw-dump excerpts (dumpsys/ip output) are embedded in `evidence` verbatim (auditable, but large and hard to anonymize) or parsed-only with an optional "attach raw dumps" toggle. Proposal: toggle, default on for local archive, default off for export.
3. Peer-mode documents: each device produces its own run; the coordinator embeds the peer's findings summary and cross-references by `run.id`. Full merge format deferred. 3. Peer-mode documents: each device produces its own run; the coordinator embeds the peer's findings summary and cross-references by `run.id`. Full merge format deferred.
4. Size guardrails: soft cap 20 MB uncompressed per run; trains beyond that downsample evidence (keep aggregates + first/last N + all anomalies) and record `"evidence_truncated": true`. 4. Size guardrails: soft cap 20 MB uncompressed per run; trains beyond that downsample evidence (keep aggregates + first/last N + all anomalies) and record `"evidence_truncated": true`.
+122 -6
View File
@@ -24,11 +24,34 @@ echolot://enroll?v=1&u=<control-URL, urlencoded>&p=pin-sha256:<b64 SPKI hash>&t=
``` ```
POST /v1/enroll Authorization: Bearer <enrollment-token> POST /v1/enroll Authorization: Bearer <enrollment-token>
→ 200 { "device_credential": "<random 256-bit, b64url>", → 201 { "device_credential": "<random 256-bit, b64url>",
"device_id": "uuid", "device_id": "uuid" }
"profile": { ... §2.2 ... } }
``` ```
The **server assembles the bootstrap link**, because it is the only party holding all three parts
at once, and the part an operator gets wrong by hand is the base64 pin — which does not fail
loudly, it just never matches, and surfaces later as an inscrutable TLS error:
```
POST /admin/enroll-tokens
→ { "token": "…", "expires_in_s": 86400,
"enroll_uri": "echolot://enroll?v=1&u=…&p=…&t=…" }
```
The control URL in the link comes from `ECHOLOT_PUBLIC_URL`, falling back to the first control
listen address. A wildcard bind has no single right answer, so it warns rather than guessing.
Encoding notes that matter in practice:
- `u`, `p` and `t` are **percent-encoded**. The pin is base64, so it contains `+`, `/` and `=`,
every one of which means something else in a query string.
- A `+` that was *not* encoded decodes to a space. Base64 contains no spaces, so a parser SHOULD
restore them — the alternative is a pin wrong by one character and a failure that points nowhere
near the cause.
- The control URL MUST be `https://`. The pin only protects a TLS connection; a cleartext URL
would hand the token to anyone on the path.
- **The link is a secret** while it is live: it carries a bearer token, so anyone who sees it
before the device does can enroll instead.
Enrollment tokens are single-use with expiry, created in the admin UI, scoped `enroll`. The device credential is a long-lived bearer secret, scoped `run-tests`; it is also the HKDF input for session keys. Revocation = deleting the device in the admin UI. Enrollment tokens are single-use with expiry, created in the admin UI, scoped `enroll`. The device credential is a long-lived bearer secret, scoped `run-tests`; it is also the HKDF input for session keys. Revocation = deleting the device in the admin UI.
### 2.2 Profile ### 2.2 Profile
@@ -61,7 +84,12 @@ The app re-fetches the profile at the start of every run (falling back to the ca
### 2.3 Capabilities (v1 registry) ### 2.3 Capabilities (v1 registry)
`udp-probe`, `stun-basic`, `stun-5780`, `canary-dns`, `recursive-dns`, `connect-back`, `delayed-echo`, `big-send`, `frag-send`, `tls-echo`, `http-echo`, `throughput`, `ntp`. A server omits what it can't offer (e.g. `stun-5780` without a second IP degrades to `stun-basic`). Clients must skip, and record as `unsupported`, any test whose capability is absent. Unknown capability strings are ignored. `udp-probe`, `stun-basic`, `stun-5780`, `canary-dns`, `recursive-dns`, `connect-back`, `delayed-echo`, `big-send`, `frag-send`, `tls-echo`, `http-echo`, `throughput`, `ntp`, plus:
- `downtrain` — server-sent downstream trains via the §5 `downtrain` action. Upstream trains need no capability of their own: they are plain client-sent data-plane packets and ride `udp-probe`.
- `tcp-echo` — the plain-TCP echo endpoint (§4); `tls-echo` is its ALPN variant on the same port.
A server omits what it can't offer (e.g. `stun-5780` without a second IP degrades to `stun-basic`). Clients must skip, and record as `unsupported`, any test whose capability is absent. Unknown capability strings are ignored.
### 2.4 Sessions ### 2.4 Sessions
@@ -88,6 +116,31 @@ DELETE /v1/sessions/{id}
Per-credential and per-source-IP token buckets on: session creation, actions, UDP packets, bytes. `429` on control plane; silent drop on data plane (probes must tolerate loss anyway). All reflected/generated traffic goes **only** to the session's observed source address (or, for connect-back, the source address of the session-creating request). Data-plane responses to unauthenticated packets are never larger than the request (§3.4). Per-credential and per-source-IP token buckets on: session creation, actions, UDP packets, bytes. `429` on control plane; silent drop on data plane (probes must tolerate loss anyway). All reflected/generated traffic goes **only** to the session's observed source address (or, for connect-back, the source address of the session-creating request). Data-plane responses to unauthenticated packets are never larger than the request (§3.4).
### 2.2 `GET /v1/discover` — where the control plane lives
Unauthenticated, and says almost nothing: the control-plane URL and the server's display name.
```json
{ "control_url": "https://probe.example.net", "name": "example" }
```
It exists so an enrollment link can carry the name a person recognises while the app still connects
to the name that selects the pinned certificate. When a server shares port 443 between its admin UI
and its control plane, those must be different hostnames — one port and one name is one certificate,
and the two need different ones (a browser-trusted certificate, and a long-lived self-signed one the
client pins). Without discovery, the difference leaks into every enrollment link an operator hands
out.
**It hands out an address, never a pin.** The pin travels in the link itself. Serving it here would
reduce pinning to whatever the certificate authorities are worth, and pinning exists precisely to
survive one the operator does not control — a root injected by corporate device management, for
instance, which is unremarkable on the networks this tool is pointed at. Because the pin is
pre-shared, an intercepted discovery response can only send a device to the wrong host, where the
pin will not match: an outage, not a compromise.
Clients treat it as optional. A server that does not answer, or a link that already names the
control endpoint, works unchanged — enrollment must not begin failing because a lookup did.
## 3. UDP probe protocol ## 3. UDP probe protocol
### 3.1 Packet header (fixed 32 bytes, network byte order) ### 3.1 Packet header (fixed 32 bytes, network byte order)
@@ -192,14 +245,77 @@ Note: exact RDATA constants to be frozen in the implementation's `dns_reference.
Same daemon, separate listener (default localhost-only): health + self-test (are both IPs live, is the canary zone delegated correctly, is UDP reachable from outside — tested via a public echolot "mirror" if configured), enrollment token management (create/expire/scope), device list + revocation, retention settings, QR rendering (client-side JS). Out of scope for this spec beyond the endpoints above. Same daemon, separate listener (default localhost-only): health + self-test (are both IPs live, is the canary zone delegated correctly, is UDP reachable from outside — tested via a public echolot "mirror" if configured), enrollment token management (create/expire/scope), device list + revocation, retention settings, QR rendering (client-side JS). Out of scope for this spec beyond the endpoints above.
## 8. Cross-references to the measurement schema ## 8. Version compatibility
Both artifacts are versioned with **SemVer**: the Go server (`server-vX.Y.Z` tags) and the Android
app (`versionName`; `versionCode` is derived from it, never maintained separately). Two independent
things are checked, and conflating them is the mistake this section exists to prevent.
### 8.1 Protocol version — *can* these builds talk?
`protocol_version` is the version of **this document**. It is advertised in the profile
(`compat.protocol_version`) and is the correctness axis: a peer in a different breaking series
cannot be talked to, whatever its release version says. Below `1.0.0` the **minor** is the breaking
axis (SemVer §4); at and above it, the major is. A patch bump of the protocol never splits a fleet.
### 8.2 Release-version window — *should* they, per policy?
Each side declares the range of peer release versions it will work with, as `[min, max)`
**minimum inclusive, maximum exclusive**, because the useful bound is always "the version that
broke it" and writing that literally is unambiguous. An empty maximum means unbounded.
The server advertises its window and enforces it:
```jsonc
"compat": {
"protocol_version": "1.0.0",
"schema_version": "1.0.0",
"app_min": "0.2.0",
"app_max": "1.0.0" // exclusive; "" = no upper bound
}
```
Operators override it with `ECHOLOT_MIN_APP_VERSION` / `ECHOLOT_MAX_APP_VERSION` (or
`--min-app-version` / `--max-app-version`). A malformed bound is **fatal at startup**, not ignored:
a typo must not silently disable a restriction the operator meant to set.
The app sends its version on every control-plane request:
```
X-Echolot-App-Version: 0.2.0
```
and carries its own bounds for the server (`MIN_SERVER` / `MAX_SERVER` in `Compat.kt`). It checks
the profile in **both** directions — is the server in our range, and are we in the server's — so a
mismatch is reported before a run starts rather than discovered halfway through one.
### 8.3 Rules
1. **`GET /v1/profile` is never gated.** It is where a refused client learns which version it needs.
Gating it leaves the user with a network error instead of an answer, defeating the check.
2. **Refusal is `426 Upgrade Required`**, with a body naming both versions and the accepted window:
```json
{ "error": "app 0.1.0 is older than this build supports (needs >= 0.2.0, < 1.0.0). Update the app.",
"app_version": "0.1.0", "accepts_app": ">= 0.2.0, < 1.0.0",
"server_version": "0.5.0", "protocol_version": "1.0.0" }
```
3. **An unparseable or absent version is `unknown`, and is allowed.** Development builds report
`dev`, and a client too old to send the header cannot be identified anyway. The check exists to
turn confusing failures into clear ones; refusing what it cannot identify does the opposite.
4. **Bounds move at breaking boundaries, not at releases.** Shipping a patch must never require
editing a range. A minimum is raised only when older peers are actually harmful — e.g. the app
requires server `>= 0.4.2` because earlier multi-homed servers sent granted traffic from an
address the session never used, which the client measured as 100 % downstream loss. A
confidently wrong measurement is worse than a refused one.
## 9. Cross-references to the measurement schema
- Observation-block fields (§3.3) appear as `*_seen_by_server` columns in `train` evidence (schema §6.2). - Observation-block fields (§3.3) appear as `*_seen_by_server` columns in `train` evidence (schema §6.2).
- `TIMESYNC` (§3.2) produces the `time.server_offset` test; without it, cross-clock fields must not be compared (schema §6.2 note). - `TIMESYNC` (§3.2) produces the `time.server_offset` test; without it, cross-clock fields must not be compared (schema §6.2 note).
- Capability strings (§2.3) are copied verbatim into `server_sessions[].capabilities` (schema §5); tests skipped for missing capability get `status: "unsupported"`. - Capability strings (§2.3) are copied verbatim into `server_sessions[].capabilities` (schema §5); tests skipped for missing capability get `status: "unsupported"`.
- `session_id` maps to `server_sessions[].session_id`; `action_id`s appear in test `params`. - `session_id` maps to `server_sessions[].session_id`; `action_id`s appear in test `params`.
## 9. Open items ## 10. Open items
1. Whether TRAIN_REPORT should also stream during long trains (partial reports every N packets) for live UI feedback — leaning yes, same type with a `flags` bit. 1. Whether TRAIN_REPORT should also stream during long trains (partial reports every N packets) for live UI feedback — leaning yes, same type with a `flags` bit.
2. Throughput methodology (fixed streams vs BBR-style ramp) — decide with `perf.*` test design. 2. Throughput methodology (fixed streams vs BBR-style ramp) — decide with `perf.*` test design.
+31 -4
View File
@@ -8,6 +8,20 @@ plugins {
alias(libs.plugins.kotlin.serialization) alias(libs.plugins.kotlin.serialization)
} }
// The app's version is SemVer and lives here, once. versionCode is derived from it rather than
// maintained alongside: Play/F-Droid need a monotonically increasing integer, but a second number
// that a human has to remember to bump is a number that eventually disagrees with the first — and
// the version is now load-bearing, since the server decides whether to serve us by it.
//
// major*1_000_000 + minor*10_000 + patch*10 leaves room for 9 patch-level rebuilds (the trailing
// digit) without disturbing the mapping, and stays inside the 2_100_000_000 ceiling until major 2100.
val appVersionName = "0.3.0"
fun versionCodeOf(semver: String): Int {
val (major, minor, patch) = semver.substringBefore('-').split(".").map(String::toInt)
return major * 1_000_000 + minor * 10_000 + patch * 10
}
android { android {
namespace = "app.echo_lot.app" namespace = "app.echo_lot.app"
compileSdk = 36 compileSdk = 36
@@ -16,12 +30,23 @@ android {
applicationId = "app.echo_lot.app" applicationId = "app.echo_lot.app"
minSdk = 26 minSdk = 26
targetSdk = 36 targetSdk = 36
versionCode = 1 versionCode = versionCodeOf(appVersionName)
versionName = "0.1.0" versionName = appVersionName
// Automation: `adb shell am start -n app.echo_lot.app/.MainActivity --ez autorun true` // Automation: `adb shell am start -n app.echo_lot.app/.MainActivity --ez autorun true`
// runs a measurement immediately and POSTs the report here (dev collection endpoint). // runs a measurement immediately and POSTs the report here (dev collection endpoint).
buildConfigField("String", "REPORT_UPLOAD_URL", "\"http://89.185.109.150:443/report\"") // Empty: the collection endpoint this pointed at was the adb-beacon receiver, which held
buildConfigField("String", "REPORT_UPLOAD_SECRET", "\"D4OmG5gGJsElqVVbtYIZbR\"") // 0.0.0.0:443 in cleartext. That service is gone and echolot-server owns 443 with TLS, so
// posting plaintext there now fails as "client sent an HTTP request to an HTTPS server" —
// an alarming error for a debugging convenience that is no longer needed, since autorun
// reports are read straight off the device with `run-as cat`.
//
// Deliberately not repointed at /v1/runs. That is the consent-gated upload, and a
// debugging shortcut must not be able to satisfy it by accident.
buildConfigField("String", "REPORT_UPLOAD_URL", "\"\"")
buildConfigField("String", "REPORT_UPLOAD_SECRET", "\"\"")
// The bare SemVer, without the debug build's "-dev" suffix stripped away by the server's
// parser anyway — sent to servers so they can apply their compatibility window.
buildConfigField("String", "APP_SEMVER", "\"$appVersionName\"")
} }
buildTypes { buildTypes {
release { isMinifyEnabled = false } release { isMinifyEnabled = false }
@@ -48,6 +73,8 @@ dependencies {
implementation(project(":core-engine")) implementation(project(":core-engine"))
implementation(project(":core-probe")) implementation(project(":core-probe"))
implementation(project(":core-shizuku")) implementation(project(":core-shizuku"))
implementation(project(":core-privacy"))
implementation(project(":core-archive"))
implementation(libs.kotlinx.serialization.json) implementation(libs.kotlinx.serialization.json)
implementation(libs.kotlinx.coroutines.android) implementation(libs.kotlinx.coroutines.android)
+45 -1
View File
@@ -10,6 +10,17 @@
<uses-permission android:name="android.permission.CHANGE_NETWORK_STATE" /> <uses-permission android:name="android.permission.CHANGE_NETWORK_STATE" />
<uses-permission android:name="android.permission.CHANGE_WIFI_MULTICAST_STATE" /> <uses-permission android:name="android.permission.CHANGE_WIFI_MULTICAST_STATE" />
<uses-permission android:name="android.permission.ACCESS_FINE_LOCATION" /> <uses-permission android:name="android.permission.ACCESS_FINE_LOCATION" />
<!--
The adb relay (AdbRelayService) only. It runs in the foreground because it must keep
watching while the tablet sits unattended with its screen off, and dataSync is the type
that describes it: it carries an observation off the LAN, nothing more. Android 15 caps
dataSync at a few hours a day, which is acceptable for a tool that is switched on for a
debugging session rather than left running forever.
-->
<uses-permission android:name="android.permission.FOREGROUND_SERVICE" />
<uses-permission android:name="android.permission.FOREGROUND_SERVICE_DATA_SYNC" />
<!-- Only so the relay's ongoing status is visible; the service runs either way. -->
<uses-permission android:name="android.permission.POST_NOTIFICATIONS" />
<application <application
android:allowBackup="false" android:allowBackup="false"
@@ -22,13 +33,46 @@
<activity <activity
android:name=".MainActivity" android:name=".MainActivity"
android:exported="true"> android:exported="true"
android:launchMode="singleTask">
<intent-filter> <intent-filter>
<action android:name="android.intent.action.MAIN" /> <action android:name="android.intent.action.MAIN" />
<category android:name="android.intent.category.LAUNCHER" /> <category android:name="android.intent.category.LAUNCHER" />
</intent-filter> </intent-filter>
<!--
Enrollment bootstrap (probe-protocol.md §2.1): echolot://enroll?v=1&u=…&p=…&t=…
Scanning a QR or tapping a link the operator sent configures the server in one
action, instead of transcribing a URL, a base64 pin and a token by hand — the pin
in particular fails silently when it is wrong by one character.
-->
<intent-filter android:autoVerify="false">
<action android:name="android.intent.action.VIEW" />
<category android:name="android.intent.category.DEFAULT" />
<category android:name="android.intent.category.BROWSABLE" />
<data android:scheme="echolot" android:host="enroll" />
</intent-filter>
<!--
Sign-in redirect. The browser hands the authorization code back through this, which
is exactly why the flow uses PKCE: any app may register this scheme, so the code
alone must not be enough to complete a sign-in.
-->
<intent-filter android:autoVerify="false">
<action android:name="android.intent.action.VIEW" />
<category android:name="android.intent.category.DEFAULT" />
<category android:name="android.intent.category.BROWSABLE" />
<data android:scheme="echolot" android:host="auth" />
</intent-filter>
</activity> </activity>
<!--
Not exported: nothing outside this app has any business starting a relay that reports
where this device can be reached.
-->
<service
android:name=".AdbRelayService"
android:exported="false"
android:foregroundServiceType="dataSync" />
<provider <provider
android:name="androidx.core.content.FileProvider" android:name="androidx.core.content.FileProvider"
android:authorities="${applicationId}.fileprovider" android:authorities="${applicationId}.fileprovider"
@@ -0,0 +1,125 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.app
import app.echo_lot.protocol.AuthInfo
import app.echo_lot.protocol.ControlClient
import app.echo_lot.protocol.OidcLogin
import kotlinx.serialization.json.Json
import kotlinx.serialization.json.jsonObject
import kotlinx.serialization.json.jsonPrimitive
/**
* Signing in to the configured server's identity provider.
*
* The awkward part of a browser-based sign-in on Android is that the app is not running while it
* happens. Handing control to a browser puts this process in the background, where it may be
* killed at any moment; the callback then arrives at a fresh process with none of the state the
* exchange needs. So the PKCE verifier and state are written to storage before the browser opens,
* not held in memory — an in-memory value works on a developer's device and fails on a phone under
* memory pressure, which is the worst way for this to break.
*
* Nothing from the identity provider is kept afterwards. The ID token proves who is signing in,
* once; the device credential authenticates everything from then on.
*/
class Account(private val settings: Settings) {
private val json = Json { ignoreUnknownKeys = true }
sealed interface SignInStart {
/** Open this in a browser. */
data class Browser(val url: String) : SignInStart
data class Unavailable(val reason: String) : SignInStart
}
/** Fetches the server's auth configuration and builds the authorization URL. */
fun begin(): SignInStart {
if (!settings.serverConfigured) {
return SignInStart.Unavailable(
"Enrol with a server first — sign-in belongs to the server's identity provider."
)
}
val auth = runCatching { client().profile(settings.serverCredential).auth }.getOrNull()
?: return SignInStart.Unavailable("Could not reach the server to ask how to sign in.")
auth.discoveryError?.let {
// The distinction matters: "the operator configured an IdP that is not answering" is
// their problem to fix, and is not the same as "this server has no accounts".
return SignInStart.Unavailable("The server's identity provider is not responding: $it")
}
if (!auth.enabled) {
return SignInStart.Unavailable("This server does not offer accounts.")
}
return try {
val pending = OidcLogin.begin(auth)
// Written before the browser opens, because after that this process may not survive.
settings.pendingVerifier = pending.verifier
settings.pendingState = pending.state
SignInStart.Browser(pending.authorizationUrl)
} catch (t: Throwable) {
SignInStart.Unavailable(t.message ?: "Could not start sign-in.")
}
}
/**
* Completes sign-in from the `echolot://auth` redirect.
*
* Blocking; callers run it off the main thread.
*/
fun complete(callbackUri: String): String {
val verifier = settings.pendingVerifier
val state = settings.pendingState
// Cleared first, whatever happens next: these are single-use, and leaving them behind
// would let a later callback be completed against a flow nobody started.
settings.clearPendingAuth()
if (verifier.isBlank() || state.isBlank()) {
return "That sign-in did not start on this device."
}
return try {
val auth = client().profile(settings.serverCredential).auth
val idToken = OidcLogin.complete(
auth, OidcLogin.Pending("", verifier, state), callbackUri,
)
val reply = client().linkAccount(settings.serverCredential, idToken)
val o = json.parseToJsonElement(reply).jsonObject
val name = o["display_name"]?.jsonPrimitive?.content ?: "signed in"
settings.accountName = name
settings.accountId = o["account_id"]?.jsonPrimitive?.content ?: ""
val admin = o["admin"]?.jsonPrimitive?.content == "true"
"Signed in as $name" + if (admin) " (administrator)" else ""
} catch (e: OidcLogin.LoginFailed) {
e.message ?: "Sign-in failed."
} catch (t: Throwable) {
"Sign-in failed: ${t.message ?: t.javaClass.simpleName}"
}
}
/** Signs out. The device stays enrolled — signing out should not cost an enrolment. */
fun signOut(): String = try {
client().unlinkAccount(settings.serverCredential)
settings.accountName = ""
settings.accountId = ""
"Signed out. This device is still enrolled."
} catch (t: Throwable) {
"Could not sign out: ${t.message ?: t.javaClass.simpleName}"
}
/** Asks the server who it thinks is signed in, so the UI is not trusting stale local state. */
fun refresh(): String? = runCatching {
val o = json.parseToJsonElement(client().accountStatus(settings.serverCredential)).jsonObject
val signedIn = o["signed_in"]?.jsonPrimitive?.content == "true"
settings.accountName = if (signedIn) {
o["display_name"]?.jsonPrimitive?.content ?: ""
} else {
""
}
settings.accountName.takeIf { it.isNotBlank() }
}.getOrNull()
private fun client() = ControlClient(
settings.serverUrl, setOf(settings.serverPin), BuildConfig.APP_SEMVER,
fallbackAddrs = settings.serverAddrList(),
)
}
@@ -0,0 +1,136 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.app
import android.content.Context
import android.net.nsd.NsdManager
import android.net.nsd.NsdServiceInfo
import android.net.wifi.WifiManager
import java.net.Inet4Address
/**
* Watches adbd's own mDNS advertisement on this device and reports the endpoint to the server.
*
* This replaces the retired Python adb-beacon, and exists for one reason: **mDNS does not cross
* subnets**. A developer on another network cannot see `_adb-tls-connect._tcp` at all, while the
* wireless-debug port rotates every few minutes — so the port has to be carried out of the LAN by
* something sitting inside it. That is this. A tablet parked on the test network relays; the
* developer reads the endpoint back from the server.
*
* Two hard-won rules from the beacon, both load-bearing:
*
* - **Resolve each service instance exactly ONCE.** Resolving adbd's advertisement makes adbd
* re-arm its connection and post a "wireless debugging connected" notification; re-resolving on
* every heartbeat turns that into a stream of them. The guard is cleared only when the service
* is *lost*, which is also what catches rotation: the new advertisement is a new instance, gets
* resolved once, and is reported within seconds.
* - **Do not run this on the OnePlus.** On network churn that device drops and re-publishes its
* advertisement repeatedly, so lost/found cycles keep clearing the guard and each resolve
* re-arms adbd. Guarding reduces but cannot eliminate the noise; the Lenovo tablet is the
* intended host, which is also why relaying is a mode rather than something always on.
*/
class AdbRelay(
private val ctx: Context,
private val onEvent: (String) -> Unit,
) {
private val nsd = ctx.getSystemService(NsdManager::class.java)
/** Instances already resolved, by service name — the re-arm guard described above. */
private val resolved = HashSet<String>()
/** Last endpoint reported, so the heartbeat re-posts from cache instead of re-resolving. */
@Volatile var lastEndpoint: Endpoint? = null
private set
data class Endpoint(val host: String, val port: Int, val serviceName: String)
private var listener: NsdManager.DiscoveryListener? = null
fun start() {
if (nsd == null) {
onEvent("mDNS unavailable on this device")
return
}
if (listener != null) return
val l = object : NsdManager.DiscoveryListener {
override fun onStartDiscoveryFailed(type: String?, code: Int) {
onEvent("discovery failed to start (code $code)")
}
override fun onStopDiscoveryFailed(type: String?, code: Int) {}
override fun onDiscoveryStarted(type: String?) {
onEvent("watching for adbd on this network")
}
override fun onDiscoveryStopped(type: String?) {}
override fun onServiceFound(info: NsdServiceInfo?) {
val name = info?.serviceName ?: return
// The guard: one resolve per instance, ever. adbd re-arms on every resolve.
if (!resolved.add(name)) return
resolve(info)
}
override fun onServiceLost(info: NsdServiceInfo?) {
// Rotation: the old instance is gone, so allow the replacement to be resolved.
info?.serviceName?.let { resolved.remove(it) }
}
}
listener = l
runCatching { nsd.discoverServices(ADB_SERVICE, NsdManager.PROTOCOL_DNS_SD, l) }
.onFailure { onEvent("could not start discovery: ${it.message}") }
}
fun stop() {
listener?.let { l -> runCatching { nsd?.stopServiceDiscovery(l) } }
listener = null
resolved.clear()
}
@Suppress("DEPRECATION") // the callback-based resolve is the one that exists across our minSdk
private fun resolve(info: NsdServiceInfo) {
val cb = object : NsdManager.ResolveListener {
override fun onResolveFailed(i: NsdServiceInfo?, code: Int) {
// Let a failed instance be retried: the guard exists to stop *successful*
// re-resolution, not to give up on a transient failure.
i?.serviceName?.let { resolved.remove(it) }
onEvent("resolve failed (code $code)")
}
override fun onServiceResolved(i: NsdServiceInfo?) {
val host = i?.host?.hostAddress ?: return
// adbd advertises on every address it listens on, including link-local v6. The
// reachable one from a developer's subnet is the routable v4 address, and it is
// also the only one worth relaying — a link-local address means nothing off-link.
if (i.host !is Inet4Address) return
// Only this device's own advertisement: on a shared network several phones may
// have wireless debugging on, and relaying a neighbour's port would send a
// developer to the wrong device.
if (host != localIp()) return
lastEndpoint = Endpoint(host, i.port, i.serviceName ?: "adb")
onEvent("found adbd at $host:${i.port}")
}
}
runCatching { nsd?.resolveService(info, cb) }
.onFailure {
resolved.remove(info.serviceName)
onEvent("resolve threw: ${it.message}")
}
}
/** This device's own IPv4 address on the wifi it is relaying from. */
private fun localIp(): String? = runCatching {
val wifi = ctx.getSystemService(WifiManager::class.java) ?: return null
@Suppress("DEPRECATION")
val ip = wifi.connectionInfo.ipAddress
if (ip == 0) return null
@Suppress("DEPRECATION")
String.format(
"%d.%d.%d.%d",
ip and 0xff, ip shr 8 and 0xff, ip shr 16 and 0xff, ip shr 24 and 0xff,
)
}.getOrNull()
private companion object {
const val ADB_SERVICE = "_adb-tls-connect._tcp"
}
}
@@ -0,0 +1,156 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.app
import android.app.Notification
import android.app.NotificationChannel
import android.app.NotificationManager
import android.app.Service
import android.content.Context
import android.content.Intent
import android.os.Build
import android.os.IBinder
import app.echo_lot.protocol.ControlClient
import kotlinx.coroutines.CoroutineScope
import kotlinx.coroutines.Dispatchers
import kotlinx.coroutines.SupervisorJob
import kotlinx.coroutines.cancel
import kotlinx.coroutines.delay
import kotlinx.coroutines.launch
import kotlinx.coroutines.withContext
/**
* Keeps [AdbRelay] running and posts what it finds to the enrolled server.
*
* A foreground service because the whole point is to be useful while nobody is looking at the
* tablet: a background process is frozen within minutes of the screen going off, and a relay that
* stops relaying the moment it is left alone would be worse than none — it would be trusted right
* up until the moment it went quiet.
*
* The heartbeat re-posts the CACHED endpoint and never re-resolves. Resolving adbd's advertisement
* makes adbd re-arm its connection and raise a "wireless debugging connected" notification, so a
* heartbeat that re-resolved would turn a background convenience into a stream of notifications on
* a device sitting on a shelf. Rotation is still caught, because losing the old advertisement
* clears the resolve guard in [AdbRelay] and the replacement is resolved once, within seconds.
*/
class AdbRelayService : Service() {
private val scope = CoroutineScope(SupervisorJob() + Dispatchers.IO)
private var relay: AdbRelay? = null
@Volatile private var status: String = "starting"
@Volatile private var lastPosted: String? = null
override fun onBind(intent: Intent?): IBinder? = null
override fun onCreate() {
super.onCreate()
startForeground(NOTIFICATION_ID, notification("starting"))
val settings = Settings(this)
val r = AdbRelay(this) { msg ->
status = msg
notify(msg)
}
relay = r
r.start()
scope.launch {
while (true) {
val ep = r.lastEndpoint
if (ep != null) {
val wire = "${ep.host}:${ep.port}"
// Re-post on a heartbeat even when unchanged: the server stamps a received-at
// time, and a developer needs to tell "this endpoint is current" from "this
// endpoint is what the tablet saw before it went out of range".
val result = post(settings, ep)
status = if (result == null) {
lastPosted = wire
"reported $wire"
} else {
"found $wire, but reporting failed: $result"
}
notify(status)
}
delay(HEARTBEAT_MS)
}
}
}
/** Posts one endpoint; returns null on success or a short reason on failure. */
private suspend fun post(settings: Settings, ep: AdbRelay.Endpoint): String? =
withContext(Dispatchers.IO) {
if (!settings.serverConfigured) return@withContext "no server enrolled"
runCatching {
ControlClient(settings.serverUrl, setOf(settings.serverPin), BuildConfig.APP_SEMVER)
.reportAdbEndpoint(
credential = settings.serverCredential,
host = ep.host,
port = ep.port,
deviceName = Build.MODEL,
note = "echolot relay",
)
null
}.getOrElse { it.message?.take(120) ?: it.javaClass.simpleName }
}
override fun onStartCommand(intent: Intent?, flags: Int, startId: Int): Int {
// Restarted by the system if it is killed: a relay that quietly does not come back after
// a low-memory kill is the failure mode this exists to avoid.
return START_STICKY
}
override fun onDestroy() {
relay?.stop()
scope.cancel()
super.onDestroy()
}
private fun notify(text: String) {
val nm = getSystemService(NotificationManager::class.java)
nm?.notify(NOTIFICATION_ID, notification(text))
}
private fun notification(text: String): Notification {
val nm = getSystemService(NotificationManager::class.java)
if (Build.VERSION.SDK_INT >= Build.VERSION_CODES.O) {
// LOW: this is a status line for a tool the user deliberately started, not news.
val ch = NotificationChannel(CHANNEL, "adb relay", NotificationManager.IMPORTANCE_LOW)
ch.description = "Reports this device's wireless-debug endpoint to the Echolot server"
nm?.createNotificationChannel(ch)
}
val open = android.app.PendingIntent.getActivity(
this, 0, Intent(this, MainActivity::class.java),
android.app.PendingIntent.FLAG_IMMUTABLE,
)
return Notification.Builder(this, CHANNEL)
.setContentTitle("Echolot adb relay")
.setContentText(text)
.setSmallIcon(android.R.drawable.stat_sys_download_done)
.setOngoing(true)
.setContentIntent(open)
.build()
}
companion object {
private const val CHANNEL = "adb-relay"
private const val NOTIFICATION_ID = 4711
/**
* Two minutes. The port rotates on roughly that cadence, and the freshness of the answer
* is the whole product — but this only re-posts a cached value, so it costs one small
* HTTPS request and never touches mDNS.
*/
private const val HEARTBEAT_MS = 120_000L
fun start(ctx: Context) {
val i = Intent(ctx, AdbRelayService::class.java)
if (Build.VERSION.SDK_INT >= Build.VERSION_CODES.O) ctx.startForegroundService(i)
else ctx.startService(i)
}
fun stop(ctx: Context) {
ctx.stopService(Intent(ctx, AdbRelayService::class.java))
}
}
}
@@ -0,0 +1,116 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.app
import androidx.compose.foundation.layout.Arrangement
import androidx.compose.foundation.layout.safeDrawingPadding
import androidx.compose.foundation.layout.Column
import androidx.compose.foundation.layout.Row
import androidx.compose.foundation.layout.fillMaxWidth
import androidx.compose.foundation.layout.padding
import androidx.compose.foundation.lazy.LazyColumn
import androidx.compose.foundation.lazy.items
import androidx.compose.material3.Card
import androidx.compose.material3.MaterialTheme
import androidx.compose.material3.Text
import androidx.compose.material3.TextButton
import androidx.compose.runtime.Composable
import androidx.compose.ui.Alignment
import androidx.compose.ui.Modifier
import androidx.compose.ui.graphics.Color
import androidx.compose.ui.unit.dp
import app.echo_lot.archive.ArchivedRun
import java.time.Instant
import java.time.ZoneId
import java.time.format.DateTimeFormatter
/**
* Archived runs, newest first.
*
* Each row states plainly whether the run left the device, because "is this backed up / did I
* share this?" is the question a history list actually gets asked.
*/
@Composable
fun HistoryScreen(
runs: List<ArchivedRun>,
status: String?,
onOpen: (String) -> Unit,
onUpload: (String) -> Unit,
onDelete: (String) -> Unit,
onBack: () -> Unit,
) {
Column(Modifier.fillMaxWidth().safeDrawingPadding().padding(16.dp), verticalArrangement = Arrangement.spacedBy(10.dp)) {
Row(verticalAlignment = Alignment.CenterVertically) {
TextButton(onClick = onBack) { Text(" Back") }
Text("History", style = MaterialTheme.typography.titleLarge)
}
status?.let { Text(it, style = MaterialTheme.typography.bodySmall) }
if (runs.isEmpty()) {
Text(
"No archived runs yet. Finished runs are kept here automatically unless you turn " +
"archiving off in settings.",
style = MaterialTheme.typography.bodyMedium,
)
return@Column
}
LazyColumn(verticalArrangement = Arrangement.spacedBy(8.dp)) {
items(runs, key = { it.id }) { r ->
Card(Modifier.fillMaxWidth()) {
Column(Modifier.padding(12.dp), verticalArrangement = Arrangement.spacedBy(4.dp)) {
Row(verticalAlignment = Alignment.CenterVertically) {
Text(
r.verdict?.uppercase() ?: "",
color = verdictTint(r.verdict),
style = MaterialTheme.typography.titleMedium,
)
Text(
" " + humanTime(r.savedAtEpochMs),
style = MaterialTheme.typography.bodyMedium,
)
}
Text(
"${r.findingCount} finding(s) · ${r.sizeBytes / 1024} kB · " +
"kept complete on this device",
style = MaterialTheme.typography.bodySmall,
)
// The upload line names the level the upload was made at, not the
// archive's. They describe different documents, and showing the archive's
// level here claimed more had left the device than actually did.
Text(
if (r.uploaded) {
"uploaded to ${r.uploadedTo ?: "a server"}" +
(r.uploadedAs?.let { " as $it" } ?: "")
} else {
"on this device only"
},
style = MaterialTheme.typography.bodySmall,
color = if (r.uploaded) Color(0xFF7FD17F) else Color(0xFFBBBBBB),
)
Row(horizontalArrangement = Arrangement.spacedBy(4.dp)) {
TextButton(onClick = { onOpen(r.id) }) { Text("Export") }
TextButton(onClick = { onUpload(r.id) }) {
Text(if (r.uploaded) "Upload again" else "Upload")
}
TextButton(onClick = { onDelete(r.id) }) { Text("Delete") }
}
}
}
}
}
}
}
private fun verdictTint(v: String?): Color = when (v?.lowercase()) {
"green", "ok", "pass" -> Color(0xFF7FD17F)
"yellow", "warn" -> Color(0xFFE0C060)
"red", "fail" -> Color(0xFFE07070)
else -> Color(0xFFBBBBBB)
}
private val stamp: DateTimeFormatter =
DateTimeFormatter.ofPattern("yyyy-MM-dd HH:mm").withZone(ZoneId.systemDefault())
private fun humanTime(epochMs: Long): String = stamp.format(Instant.ofEpochMilli(epochMs))
@@ -18,6 +18,10 @@ import androidx.compose.foundation.shape.RoundedCornerShape
import androidx.compose.foundation.verticalScroll import androidx.compose.foundation.verticalScroll
import androidx.compose.material3.* import androidx.compose.material3.*
import androidx.compose.runtime.Composable import androidx.compose.runtime.Composable
import androidx.compose.runtime.getValue
import androidx.compose.runtime.mutableStateOf
import androidx.compose.runtime.remember
import androidx.compose.runtime.setValue
import androidx.compose.ui.Alignment import androidx.compose.ui.Alignment
import androidx.compose.ui.Modifier import androidx.compose.ui.Modifier
import androidx.compose.ui.graphics.Color import androidx.compose.ui.graphics.Color
@@ -26,28 +30,157 @@ import androidx.compose.ui.text.font.FontWeight
import androidx.compose.ui.unit.dp import androidx.compose.ui.unit.dp
import androidx.compose.ui.unit.sp import androidx.compose.ui.unit.sp
import androidx.core.content.ContextCompat import androidx.core.content.ContextCompat
import androidx.lifecycle.lifecycleScope
import kotlinx.coroutines.launch
import androidx.lifecycle.viewmodel.compose.viewModel import androidx.lifecycle.viewmodel.compose.viewModel
import app.echo_lot.measurement.* import app.echo_lot.measurement.*
/** The app's three top-level screens. */
private enum class Screen { RUN, HISTORY, SETTINGS }
class MainActivity : ComponentActivity() { class MainActivity : ComponentActivity() {
private val permissionLauncher = private val permissionLauncher =
registerForActivityResult(ActivityResultContracts.RequestMultiplePermissions()) { /* proceed regardless */ } registerForActivityResult(ActivityResultContracts.RequestMultiplePermissions()) { /* proceed regardless */ }
/**
* The intent currently being acted on, so a deep link that arrives while the app is running
* is seen by the screen the user is already looking at.
*
* The activity is singleTask for the same reason. As a standard activity it stacked a second
* instance per link, each with its own ViewModel: the enrolment then happened in a throwaway
* copy, and pressing back returned to the original screen showing none of it. Silent, and
* indistinguishable from the link simply not working.
*/
private val liveIntent = mutableStateOf<android.content.Intent?>(null)
override fun onNewIntent(intent: android.content.Intent) {
super.onNewIntent(intent)
setIntent(intent)
liveIntent.value = intent
}
override fun onCreate(savedInstanceState: Bundle?) { override fun onCreate(savedInstanceState: Bundle?) {
super.onCreate(savedInstanceState) super.onCreate(savedInstanceState)
requestRuntimePermissions() requestRuntimePermissions()
resumeRelayIfEnabled()
liveIntent.value = intent
setContent { setContent {
MaterialTheme(colorScheme = darkColorScheme()) { MaterialTheme(colorScheme = darkColorScheme()) {
Surface(color = MaterialTheme.colorScheme.background) { Surface(color = MaterialTheme.colorScheme.background) {
val vm: RunViewModel = viewModel() val vm: RunViewModel = viewModel()
// Three flat screens, so a plain state variable beats a navigation library:
// there is no back stack to model beyond "return to the run screen".
var screen by remember { mutableStateOf(Screen.RUN) }
var preview by remember { mutableStateOf<String?>(null) }
// Automation entry point: // Automation entry point:
// adb shell am start -n app.echo_lot.app/.MainActivity --ez autorun true // adb shell am start -n app.echo_lot.app/.MainActivity --ez autorun true
// starts a run immediately and uploads the report, so an unattended // starts a run immediately and uploads the report, so an unattended
// measurement needs no UI tapping and no adb round-trip to collect. // measurement needs no UI tapping and no adb round-trip to collect.
val autorun = intent?.getBooleanExtra("autorun", false) == true val autorun = intent?.getBooleanExtra("autorun", false) == true
// An echolot://enroll link (QR scan, or a link the operator sent) opens the
// app straight into settings with the enrollment already done, so the user
// sees the result rather than a form they still have to fill in.
// Both deep links land here. They are told apart by host, so a sign-in
// redirect is never mistaken for an enrolment link — one spends a token, the
// other completes an authorization, and confusing them would fail obscurely.
val incoming = liveIntent.value?.takeIf { it.action == Intent.ACTION_VIEW }?.dataString
val authUri = incoming?.takeIf { it.startsWith("echolot://auth") }
val enrollUri = incoming?.takeIf { it.startsWith("echolot://enroll") }
androidx.compose.runtime.LaunchedEffect(enrollUri) {
if (enrollUri != null) {
vm.enroll(enrollUri)
screen = Screen.SETTINGS
}
}
// Replacing an existing enrollment is asked about, never assumed. Following a
// link from a web page is one tap, and the old credential does not survive it.
vm.state.pendingEnroll?.let { pending ->
androidx.compose.material3.AlertDialog(
onDismissRequest = { vm.cancelEnroll() },
title = {
androidx.compose.material3.Text(
if (pending.sameServer) "Enroll again with this server?"
else "Replace this device's server?"
)
},
text = {
androidx.compose.material3.Text(
// Naming the same URL twice reads as a mistake and buries the
// one consequence that actually applies: the device is issued a
// fresh credential and shows up as a second entry.
if (pending.sameServer) {
"This device is already enrolled with " +
"${pending.currentServer}.\n\n" +
"Enrolling again replaces its credential. The old one " +
"stops working immediately, and the device appears on " +
"the server as a new entry alongside the current one — " +
"which you may want to revoke afterwards.\n\n" +
"Runs already uploaded, and runs stored on this phone, " +
"are not affected."
} else {
"This device is already enrolled with " +
"${pending.currentServer}.\n\n" +
"Enrolling with ${pending.newServer} replaces that. Runs " +
"already uploaded stay where they are, but this device " +
"stops reporting to the old server and appears on the new " +
"one as a new device.\n\n" +
"Runs stored on this phone are not affected."
}
)
},
confirmButton = {
androidx.compose.material3.TextButton(onClick = { vm.confirmEnroll() }) {
androidx.compose.material3.Text(
if (pending.sameServer) "Enroll again" else "Enroll here"
)
}
},
dismissButton = {
androidx.compose.material3.TextButton(onClick = { vm.cancelEnroll() }) {
androidx.compose.material3.Text("Keep current server")
}
},
)
}
androidx.compose.runtime.LaunchedEffect(authUri) {
if (authUri != null) {
vm.completeSignIn(authUri)
screen = Screen.SETTINGS
}
}
// A long run samples for minutes, and Android starts throttling timers and
// network access within moments of the screen going off — so a run left to
// itself would measure the device's power management rather than the network,
// and would do it silently. Held only while a run is in flight, and released
// on the way out.
val view = androidx.compose.ui.platform.LocalView.current
androidx.compose.runtime.DisposableEffect(vm.state.running) {
view.keepScreenOn = vm.state.running
onDispose { view.keepScreenOn = false }
}
// Shizuku can be started, stopped or authorised in its own app, where nothing
// calls back into this process. Asking again each time this screen comes
// forward is what makes the banner right after the user has been away to fix
// it — which is exactly the moment they look at it.
val lifecycleOwner = androidx.compose.ui.platform.LocalLifecycleOwner.current
androidx.compose.runtime.DisposableEffect(lifecycleOwner) {
val obs = androidx.lifecycle.LifecycleEventObserver { _, event ->
if (event == androidx.lifecycle.Lifecycle.Event.ON_RESUME) {
vm.refreshShizuku()
}
}
lifecycleOwner.lifecycle.addObserver(obs)
onDispose { lifecycleOwner.lifecycle.removeObserver(obs) }
}
// Autorun stays a quick run: it is an unattended batch job driven over adb, and
// an automation that silently held the device for five minutes would be a
// surprise. `--es mode long` asks for the other one explicitly.
val autorunMode =
if (intent?.getStringExtra("mode") == "long") RunMode.LONG else RunMode.SHORT
androidx.compose.runtime.LaunchedEffect(autorun) { androidx.compose.runtime.LaunchedEffect(autorun) {
if (autorun) vm.run(upload = true) if (autorun) vm.run(autorunMode, devUpload = true)
} }
// In autorun the app is a batch job: once the run is done AND the upload // In autorun the app is a batch job: once the run is done AND the upload
// succeeded, show the result briefly, then close so the device is left as it // succeeded, show the result briefly, then close so the device is left as it
@@ -61,9 +194,72 @@ class MainActivity : ComponentActivity() {
finish() finish()
} }
} }
EcholotScreen( // Without this, the system Back gesture leaves the activity from Settings or
// History instead of returning to the run screen — the screen is a plain state
// variable, so nothing connects it to the back stack. Registered only when
// there is somewhere to go back to, so Back still exits from the run screen.
androidx.activity.compose.BackHandler(enabled = screen != Screen.RUN) {
screen = Screen.RUN
}
when (screen) {
Screen.SETTINGS -> SettingsScreen(
settings = vm.settings,
archivedRuns = vm.archivedRunCount(),
archivedBytes = vm.archivedBytes(),
onApplyRetention = vm::applyRetention,
onDeleteAll = vm::deleteAllRuns,
onPreviewUpload = {
// Straight from the archive: the newest run is the one the user just
// made and the one they are deciding about. Always shows something,
// even when there is nothing to preview yet.
lifecycleScope.launch { preview = vm.previewNewestRun() }
},
onCheckServer = vm::checkServer,
accountName = vm.accountName,
onSignIn = {
vm.beginSignIn { url ->
// A plain VIEW intent rather than a Custom Tab: the browser is
// where the user's existing IdP session already lives, and
// androidx.browser would be a dependency for a rounded corner.
runCatching {
startActivity(Intent(Intent.ACTION_VIEW, android.net.Uri.parse(url)))
}
}
},
onSignOut = vm::signOut,
onEnroll = vm::enroll,
serverStatus = vm.state.archiveStatus,
enrollStatus = vm.state.enrollStatus,
onRelayChange = { on ->
if (on) AdbRelayService.start(this@MainActivity)
else AdbRelayService.stop(this@MainActivity)
},
onBack = { screen = Screen.RUN },
)
Screen.HISTORY -> HistoryScreen(
runs = vm.state.history,
status = vm.state.archiveStatus,
onOpen = { id ->
lifecycleScope.launch {
vm.readRun(id)?.let { text ->
startActivity(
Intent.createChooser(
Report.shareJson(this@MainActivity, id, text),
"Export Echolot run",
)
)
}
}
},
onUpload = vm::uploadRun,
onDelete = vm::deleteRun,
onBack = { screen = Screen.RUN },
)
Screen.RUN -> EcholotScreen(
state = vm.state, state = vm.state,
onRun = { vm.run() }, longMinutes = vm.settings.longRunMinutes,
onRun = { mode -> vm.run(mode) },
onCancel = vm::cancel, onCancel = vm::cancel,
onDeveloperOptions = { onDeveloperOptions = {
runCatching { runCatching {
@@ -86,7 +282,13 @@ class MainActivity : ComponentActivity() {
} }
}, },
onExport = { doc -> startActivity(Intent.createChooser(Report.share(this, doc), "Export Echolot run")) }, onExport = { doc -> startActivity(Intent.createChooser(Report.share(this, doc), "Export Echolot run")) },
) onOpenSettings = { screen = Screen.SETTINGS },
onOpenHistory = { vm.refreshHistory(); screen = Screen.HISTORY },
)
}
preview?.let { text ->
UploadPreviewDialog(text) { preview = null }
}
} }
} }
} }
@@ -94,11 +296,29 @@ class MainActivity : ComponentActivity() {
private fun requestRuntimePermissions() { private fun requestRuntimePermissions() {
val perms = mutableListOf(Manifest.permission.ACCESS_FINE_LOCATION) val perms = mutableListOf(Manifest.permission.ACCESS_FINE_LOCATION)
// Only so the relay's ongoing notification is visible. The service runs either way, but a
// foreground service the user cannot see is worse than one they can dismiss knowingly.
if (Build.VERSION.SDK_INT >= Build.VERSION_CODES.TIRAMISU) {
perms.add(Manifest.permission.POST_NOTIFICATIONS)
}
val missing = perms.filter { val missing = perms.filter {
ContextCompat.checkSelfPermission(this, it) != PackageManager.PERMISSION_GRANTED ContextCompat.checkSelfPermission(this, it) != PackageManager.PERMISSION_GRANTED
} }
if (missing.isNotEmpty()) permissionLauncher.launch(missing.toTypedArray()) if (missing.isNotEmpty()) permissionLauncher.launch(missing.toTypedArray())
} }
/**
* Restarts the relay if it was left on.
*
* A relay that silently fails to come back after a reboot or a process kill is worse than one
* that was never enabled: it is trusted right up to the moment it goes quiet, and the symptom
* is a stale endpoint that sends a developer to a port nothing is listening on.
*/
private fun resumeRelayIfEnabled() {
if (Settings(this).adbRelayEnabled) {
runCatching { AdbRelayService.start(this) }
}
}
} }
private fun verdictColor(v: Verdict): Color = when (v) { private fun verdictColor(v: Verdict): Color = when (v) {
@@ -118,11 +338,14 @@ private fun statusColor(s: TestStatus): Color = when (s) {
@Composable @Composable
private fun EcholotScreen( private fun EcholotScreen(
state: UiState, state: UiState,
onRun: () -> Unit, longMinutes: Int,
onRun: (RunMode) -> Unit,
onCancel: () -> Unit, onCancel: () -> Unit,
onShizukuAction: () -> Unit, onShizukuAction: () -> Unit,
onDeveloperOptions: () -> Unit, onDeveloperOptions: () -> Unit,
onExport: (MeasurementDocument) -> Unit, onExport: (MeasurementDocument) -> Unit,
onOpenSettings: () -> Unit,
onOpenHistory: () -> Unit,
) { ) {
Column( Column(
Modifier Modifier
@@ -135,8 +358,15 @@ private fun EcholotScreen(
.verticalScroll(rememberScrollState()), .verticalScroll(rememberScrollState()),
verticalArrangement = Arrangement.spacedBy(12.dp), verticalArrangement = Arrangement.spacedBy(12.dp),
) { ) {
Text("Echolot", fontSize = 26.sp, fontWeight = FontWeight.SemiBold) Row(Modifier.fillMaxWidth(), verticalAlignment = Alignment.CenterVertically) {
Text("measure, don't guess", color = MaterialTheme.colorScheme.onSurfaceVariant, fontSize = 13.sp) Column(Modifier.weight(1f)) {
Text("Echolot", fontSize = 26.sp, fontWeight = FontWeight.SemiBold)
Text("measure, don't guess",
color = MaterialTheme.colorScheme.onSurfaceVariant, fontSize = 13.sp)
}
TextButton(onClick = onOpenHistory) { Text("History") }
TextButton(onClick = onOpenSettings) { Text("Settings") }
}
// Shell-tier readiness, before the run. Nothing is shown when Shizuku isn't installed — // Shell-tier readiness, before the run. Nothing is shown when Shizuku isn't installed —
// only users who actually use it get reminded that it must be running. // only users who actually use it get reminded that it must be running.
@@ -173,8 +403,36 @@ private fun EcholotScreen(
} }
} }
// The choice is made before the run, not after, because it is a choice about how long the
// user is willing to stand still — and because the two modes answer different questions.
var mode by remember { mutableStateOf(RunMode.SHORT) }
Row(horizontalArrangement = Arrangement.spacedBy(8.dp), verticalAlignment = Alignment.CenterVertically) {
FilterChip(
selected = mode == RunMode.SHORT,
onClick = { mode = RunMode.SHORT },
enabled = !state.running,
label = { Text("Quick") },
)
FilterChip(
selected = mode == RunMode.LONG,
onClick = { mode = RunMode.LONG },
enabled = !state.running,
label = { Text("Long ($longMinutes min)") },
)
}
Text(
if (mode == RunMode.SHORT) {
"About 30 seconds. Describes how the network is configured right now."
} else {
"Listens for $longMinutes minutes while it measures. Finds what a quick run " +
"structurally cannot: links that drop and come back, signal that decays, " +
"loss that arrives in bursts."
},
fontSize = 12.sp, color = MaterialTheme.colorScheme.onSurfaceVariant,
)
Row(horizontalArrangement = Arrangement.spacedBy(12.dp), verticalAlignment = Alignment.CenterVertically) { Row(horizontalArrangement = Arrangement.spacedBy(12.dp), verticalAlignment = Alignment.CenterVertically) {
Button(onClick = onRun, enabled = !state.running) { Button(onClick = { onRun(mode) }, enabled = !state.running) {
Text(if (state.running) "Running…" else "Run measurement") Text(if (state.running) "Running…" else "Run measurement")
} }
if (state.running) { if (state.running) {
@@ -185,31 +443,65 @@ private fun EcholotScreen(
} }
} }
state.archiveStatus?.let {
Text(it, fontSize = 12.sp, color = MaterialTheme.colorScheme.onSurfaceVariant)
}
state.uploadStatus?.let { state.uploadStatus?.let {
Text(it, fontSize = 12.sp, color = MaterialTheme.colorScheme.onSurfaceVariant) Text(it, fontSize = 12.sp, color = MaterialTheme.colorScheme.onSurfaceVariant)
} }
if (state.running) { if (state.running) {
Column(verticalArrangement = Arrangement.spacedBy(6.dp)) { Column(verticalArrangement = Arrangement.spacedBy(6.dp)) {
val frac = if (state.stepsTotal > 0) // In a long run the window is the run: once the battery is done, four minutes of
state.stepsDone.toFloat() / state.stepsTotal else 0f // listening remain, and a bar driven by the step count would sit at 100 % through
// all of it — which reads as an app that has hung, not one that is working.
val listening = state.windowTotalS > 0
val frac = when {
listening -> state.windowElapsedS.toFloat() / state.windowTotalS
state.stepsTotal > 0 -> state.stepsDone.toFloat() / state.stepsTotal
else -> 0f
}
LinearProgressIndicator( LinearProgressIndicator(
progress = { frac }, progress = { frac.coerceIn(0f, 1f) },
modifier = Modifier.fillMaxWidth(), modifier = Modifier.fillMaxWidth(),
) )
if (listening) {
Row(horizontalArrangement = Arrangement.spacedBy(8.dp)) {
Text(
"listening ${clock(state.windowElapsedS)} of ${clock(state.windowTotalS)}",
fontSize = 12.sp, fontWeight = FontWeight.Medium,
)
Text(
"${clock(state.windowTotalS - state.windowElapsedS)} left",
fontSize = 12.sp, modifier = Modifier.weight(1f),
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
}
}
Row(horizontalArrangement = Arrangement.spacedBy(8.dp)) { Row(horizontalArrangement = Arrangement.spacedBy(8.dp)) {
// Still shown during a long run's listening phase, where it stops at the last
// test: the battery's progress is real information, it is simply not the whole
// run any more.
Text( Text(
if (state.stepsTotal > 0) if (state.stepsTotal > 0)
"test ${state.stepsDone + 1} of ${state.stepsTotal}" else "starting", "test ${(state.stepsDone + 1).coerceAtMost(state.stepsTotal)} of ${state.stepsTotal}"
else "starting",
fontSize = 12.sp, fontWeight = FontWeight.Medium, fontSize = 12.sp, fontWeight = FontWeight.Medium,
) )
Text(state.currentStep ?: "", fontSize = 12.sp, Text(state.currentStep ?: "", fontSize = 12.sp,
fontFamily = FontFamily.Monospace, modifier = Modifier.weight(1f)) fontFamily = FontFamily.Monospace, modifier = Modifier.weight(1f))
if (state.etaSeconds > 0) { if (!listening && state.etaSeconds > 0) {
Text("~${state.etaSeconds}s left", fontSize = 12.sp, Text("~${state.etaSeconds}s left", fontSize = 12.sp,
color = MaterialTheme.colorScheme.onSurfaceVariant) color = MaterialTheme.colorScheme.onSurfaceVariant)
} }
} }
if (listening) {
Text(
"Cancelling keeps what has been collected so far.",
fontSize = 11.sp, color = MaterialTheme.colorScheme.onSurfaceVariant,
)
}
} }
} }
@@ -219,11 +511,57 @@ private fun EcholotScreen(
@Composable @Composable
private fun Results(doc: MeasurementDocument) { private fun Results(doc: MeasurementDocument) {
// A constrained run is answered before the lights are: the verdict below is INCONCLUSIVE by
// §7.3, and without this banner "inconclusive" reads as the app failing rather than the OS
// (correctly) refusing to let anything past the VPN be measured.
val constraints = doc.run.constraints
if (constraints.constrained) {
val blocked = constraints.unmeasuredNetworks
.mapNotNull { id -> doc.networks.firstOrNull { it.id == id } }
.joinToString(", ") { it.iface?.takeIf { s -> s.isNotBlank() } ?: it.transport.name.lowercase() }
.ifBlank { "the networks beneath it" }
// Same three-way split as the finding: saying "VPN" when the user just disconnected
// theirs (the wall lingers during teardown) reads as the app being wrong, not the OS.
val (headline, body) = when {
constraints.vpnActive && constraints.perNetworkBlocked ->
"Measured through a VPN" to
("Android does not let apps send on the networks beneath an active VPN, so " +
"$blocked could not be measured — these results describe the tunnel. " +
"Disconnect the VPN and run again to measure the networks themselves.")
constraints.perNetworkBlocked ->
"Some networks could not be measured" to
("Android refused this app permission to send on $blocked, so they went " +
"unmeasured. A connected VPN is the usual cause; the refusal can also " +
"outlast one. Everything else in this run is unaffected.")
else ->
"A VPN holds the default route" to
("Default-route results describe the tunnel; per-network measurements " +
"reached the underlying networks.")
}
Card(colors = CardDefaults.cardColors(containerColor = Color(0xFF3A2E12))) {
Column(Modifier.fillMaxWidth().padding(12.dp)) {
Text(headline, color = Color(0xFFFFD08A), fontWeight = FontWeight.SemiBold)
Text(body, fontSize = 12.sp, color = Color(0xFFFFD08A))
}
}
}
val summary = doc.summary val summary = doc.summary
if (summary != null) { if (summary != null) {
Card(colors = CardDefaults.cardColors(containerColor = verdictColor(summary.overall))) { Card(colors = CardDefaults.cardColors(containerColor = verdictColor(summary.overall))) {
Column(Modifier.fillMaxWidth().padding(16.dp)) { Column(Modifier.fillMaxWidth().padding(16.dp)) {
Text("Overall: ${summary.overall}", color = Color.White, fontWeight = FontWeight.Bold, fontSize = 18.sp) Text("Overall: ${summary.overall}", color = Color.White, fontWeight = FontWeight.Bold, fontSize = 18.sp)
// Which question this document answers. A green light from a quick run does not
// mean the same thing as a green light from a long one, and the report should not
// let the two look identical.
Text(
if (doc.run.mode == RunMode.LONG) {
"long run — the network was watched continuously as well as probed"
} else {
"quick run — a snapshot; nothing here rules out an intermittent fault"
},
color = Color.White, fontSize = 12.sp,
)
} }
} }
FlowCategories(summary.categories) FlowCategories(summary.categories)
@@ -329,6 +667,12 @@ private fun FlowCategories(categories: Map<String, CategorySummary>) {
} }
} }
/** m:ss — minutes are how a five-minute wait is read; "247s left" is a number to convert. */
private fun clock(seconds: Int): String {
val s = seconds.coerceAtLeast(0)
return "${s / 60}:${"%02d".format(s % 60)}"
}
@Composable @Composable
private fun Dot(color: Color) { private fun Dot(color: Color) {
Surface(color = color, shape = RoundedCornerShape(50), modifier = Modifier.size(12.dp)) {} Surface(color = color, shape = RoundedCornerShape(50), modifier = Modifier.size(12.dp)) {}
@@ -338,3 +682,24 @@ private fun Dot(color: Color) {
private fun SectionTitle(text: String) { private fun SectionTitle(text: String) {
Text(text, fontWeight = FontWeight.SemiBold, fontSize = 15.sp, modifier = Modifier.padding(top = 8.dp)) Text(text, fontWeight = FontWeight.SemiBold, fontSize = 15.sp, modifier = Modifier.padding(top = 8.dp))
} }
/**
* Shows the exact JSON an upload would send.
*
* This exists because an anonymizer the user cannot inspect is just a promise. Being able to
* read the outgoing document — and find their own SSID absent from it — is what makes the
* privacy setting checkable rather than merely stated.
*/
@Composable
private fun UploadPreviewDialog(text: String, onDismiss: () -> Unit) {
AlertDialog(
onDismissRequest = onDismiss,
confirmButton = { TextButton(onClick = onDismiss) { Text("Close") } },
title = { Text("This is what would be uploaded") },
text = {
Column(Modifier.heightIn(max = 420.dp).verticalScroll(rememberScrollState())) {
Text(text, fontSize = 10.sp, fontFamily = FontFamily.Monospace)
}
},
)
}
@@ -17,10 +17,18 @@ object Report {
fun toJson(doc: MeasurementDocument): String = fun toJson(doc: MeasurementDocument): String =
json.encodeToString(MeasurementDocument.serializer(), doc) json.encodeToString(MeasurementDocument.serializer(), doc)
fun share(ctx: Context, doc: MeasurementDocument): Intent { fun share(ctx: Context, doc: MeasurementDocument): Intent =
shareJson(ctx, doc.run.id, toJson(doc))
/**
* Shares an already-serialized run — an archived one, whose bytes must go out exactly as
* stored rather than being re-serialized through the model (which would silently drop
* anything a newer schema version added).
*/
fun shareJson(ctx: Context, runId: String, json: String): Intent {
val dir = File(ctx.cacheDir, "reports").apply { mkdirs() } val dir = File(ctx.cacheDir, "reports").apply { mkdirs() }
val file = File(dir, "echolot-run-${doc.run.id}.json") val file = File(dir, "echolot-run-$runId.json")
file.writeText(toJson(doc)) file.writeText(json)
val uri = FileProvider.getUriForFile(ctx, "${ctx.packageName}.fileprovider", file) val uri = FileProvider.getUriForFile(ctx, "${ctx.packageName}.fileprovider", file)
return Intent(Intent.ACTION_SEND).apply { return Intent(Intent.ACTION_SEND).apply {
type = "application/json" type = "application/json"
@@ -0,0 +1,228 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.app
import android.content.Context
import app.echo_lot.archive.ArchivedRun
import app.echo_lot.archive.RunArchive
import app.echo_lot.measurement.MeasurementDocument
import app.echo_lot.privacy.Anonymizer
import app.echo_lot.privacy.PrivacyLevel
import app.echo_lot.privacy.Salt
import app.echo_lot.protocol.Compat
import app.echo_lot.protocol.ControlClient
import app.echo_lot.protocol.UploadRefused
import app.echo_lot.protocol.VersionRefused
import kotlinx.serialization.json.Json
import kotlinx.serialization.json.JsonObject
import kotlinx.serialization.json.jsonObject
import java.io.File
import java.security.SecureRandom
/**
* Ties together the three things that happen to a finished run: it gets archived, it may get
* anonymized, and it may get uploaded — in that order, and with the archive always holding the
* *unredacted* document.
*
* That ordering is the important decision. The local archive is the user's own data on their own
* device, and redacting it would destroy exactly the detail that makes a week-old run worth
* keeping; the anonymizer exists for the moment data crosses to someone else's machine. So
* redaction happens on the way out, per upload, and the archive is never the lossy copy.
*/
class RunStore(context: Context, private val settings: Settings) {
private val archive = RunArchive(File(context.filesDir, "runs"))
private val json = Json { encodeDefaults = true; explicitNulls = true }
fun list(): List<ArchivedRun> = archive.list()
fun read(id: String): String? = archive.read(id)
fun delete(id: String) = archive.delete(id)
fun deleteAll(): Int = archive.deleteAll()
fun totalBytes(): Long = archive.totalBytes()
/** Archives a finished run under the user's retention policy. Null when archiving is off. */
fun archive(doc: MeasurementDocument): ArchivedRun? =
archive.save(Report.toJson(doc), settings.retention())
/** Applies retention now — e.g. after the user tightens the limits in settings. */
fun purgeNow() = archive.purge(settings.retention())
/**
* Produces exactly the bytes an upload would send, so the UI can show the user their own
* document as the server will see it *before* it goes. "Preview what you're about to share"
* is the only honest way to present an anonymizer: its correctness is not something a user
* should have to take on faith.
*/
fun redactedForUpload(docJson: String, level: PrivacyLevel = settings.privacyLevel): String {
val parsed = runCatching { json.parseToJsonElement(docJson).jsonObject }.getOrNull()
?: return docJson
return json.encodeToString(JsonObject.serializer(), Anonymizer(level, salt()).anonymize(parsed))
}
private fun salt(): Salt =
if (settings.stableSalt) Salt.stable(settings.saltSecret())
else Salt.perRun(ByteArray(32).also { SecureRandom().nextBytes(it) })
sealed interface UploadOutcome {
data class Sent(val serverName: String, val detail: String) : UploadOutcome
/** The operator's policy says no. Not retryable, and not the user's fault. */
data class Refused(val reason: String) : UploadOutcome
/**
* The two builds do not go together. Kept apart from [Refused] and [Failed] because the
* remedy is different and specific — install a particular version — and a message that
* says so is worth more than one that says "upload failed".
*/
data class Incompatible(val reason: String) : UploadOutcome
data class Failed(val detail: String) : UploadOutcome
data object NotConfigured : UploadOutcome
}
private fun client() = ControlClient(
settings.serverUrl, setOf(settings.serverPin), BuildConfig.APP_SEMVER,
fallbackAddrs = settings.serverAddrList(),
)
/** Remembers where the server lives, so a later run can reach it without DNS. */
private fun rememberAddrs(p: app.echo_lot.protocol.Profile) {
val addrs = p.targets.flatMap { listOfNotNull(it.ip4, it.ip6) }
.filter { it.isNotBlank() }
if (addrs.isNotEmpty()) settings.serverAddrs = addrs.joinToString(",")
}
/**
* Checks the configured server without uploading anything: reachable, pinned, compatible, and
* willing to accept runs. Lets the user find out in settings rather than from a failed run.
*/
fun checkServer(): String {
if (!settings.serverConfigured) return "Fill in the server URL, pin and credential first."
return try {
val profile = client().profile(settings.serverCredential)
// Learned here so the next run's canary probe knows what to ask for.
profile.canaryZone.takeIf { it.isNotBlank() }?.let { settings.canaryZone = it }
settings.serverFacts = describeFacts(profile)
rememberAddrs(profile)
val compat = Compat.check(profile, BuildConfig.APP_SEMVER)
val head = "${profile.name} · server ${profile.serverVersion} · " +
"protocol ${profile.compat.protocolVersion.ifBlank { "unstated" }}"
when {
compat.message != null -> head + "\n" + compat.message
else -> {
val uploads = profile.uploads.refusalReason()
?: "uploads accepted (min anonymization: ${profile.uploads.minAnonymization})"
head + "\nCompatible. " + uploads
}
}
} catch (e: VersionRefused) {
"This server will not serve this app: ${e.message}"
} catch (t: Throwable) {
"Could not reach the server: ${t.message ?: t.javaClass.simpleName}"
}
}
/**
* Redeems an enrollment link and stores the resulting server configuration (§2.1).
*
* Everything is written at once or not at all: a half-applied server — say a URL and pin with
* no credential — fails later, somewhere else, with an error that points at the wrong thing.
* Blocking; callers run it off the main thread.
*/
/**
* Renders what the server says about itself, for display.
*
* Only what a person measuring against it would want to check: which addresses the tests will
* actually use, on which ports, and what the server admits it can do. Addresses first, because
* "which address did this result come from" is the question a report leaves open.
*/
private fun describeFacts(p: app.echo_lot.protocol.Profile): String {
val lines = ArrayList<String>()
// "label|value" per line, laid out as real columns by the UI rather than padded with
// spaces here. Space padding only lines up in a monospaced font, which makes the layout
// depend on a typeface choice made somewhere else entirely.
fun row(label: String, value: String) = lines.add("$label|$value")
row("server", "${p.name} · ${p.serverVersion}")
for (t in p.targets) {
t.ip4?.let { row("IPv4", it) }
t.ip6?.let { row("IPv6", it) }
// Marked rather than listed apart: it is the same server, and what matters is being
// able to tell which address a NAT-behaviour result came from.
t.ip4Alt?.let { row("IPv4 alt", it) }
t.ip6Alt?.let { row("IPv6 alt", it) }
row("ports", "udp ${t.udpPort} · tcp ${t.tcpPort} · stun ${t.stunPort}")
}
if (p.canaryZone.isNotBlank()) row("dns zone", p.canaryZone)
if (p.capabilities.isNotEmpty()) row("measures", p.capabilities.joinToString(", "))
return lines.joinToString(System.lineSeparator())
}
fun enroll(link: String, deviceName: String?): String {
val parsed = app.echo_lot.protocol.EnrollmentLink.parse(link)
?: return "That does not look like an Echolot enrollment link. It should start with " +
"echolot://enroll and carry a URL, a pin and a token."
return try {
val enrolled = parsed.redeem(deviceName, BuildConfig.APP_SEMVER)
val compat = Compat.check(enrolled.profile, BuildConfig.APP_SEMVER)
settings.serverFacts = describeFacts(enrolled.profile)
rememberAddrs(enrolled.profile)
enrolled.profile.canaryZone.takeIf { it.isNotBlank() }?.let { settings.canaryZone = it }
settings.serverUrl = enrolled.controlUrl
settings.serverPublicUrl = enrolled.publicUrl
settings.serverPin = enrolled.pin
settings.serverCredential = enrolled.credential
val head = "Enrolled with ${enrolled.profile.name} " +
"(server ${enrolled.profile.serverVersion})."
if (compat.message != null) head + " " + compat.message else head
} catch (e: VersionRefused) {
"That server will not serve this app: ${e.message}"
} catch (t: Throwable) {
// The commonest causes are a spent token and a wrong pin, and they look nothing alike
// in the message — so pass it through rather than flattening it to "enrollment failed".
"Enrollment failed: ${t.message ?: t.javaClass.simpleName}"
}
}
/**
* Uploads one archived run to the configured server, redacting first.
*
* The server's advertised minimum wins over the user's preference when it is stricter — a
* server may demand more anonymization than the user chose, never less. Blocking; callers
* run it off the main thread.
*/
fun upload(runId: String): UploadOutcome {
if (!settings.serverConfigured) return UploadOutcome.NotConfigured
val docJson = read(runId) ?: return UploadOutcome.Failed("run $runId is not in the archive")
return try {
val client = client()
val profile = client.profile(settings.serverCredential)
profile.canaryZone.takeIf { it.isNotBlank() }?.let { settings.canaryZone = it }
settings.serverFacts = describeFacts(profile)
rememberAddrs(profile)
// Compatibility before policy: an incompatible server may well advertise an upload
// policy it would never actually apply to us.
val compat = Compat.check(profile, BuildConfig.APP_SEMVER)
if (!compat.usable) return UploadOutcome.Incompatible(compat.message ?: "incompatible versions")
profile.uploads.refusalReason()?.let { return UploadOutcome.Refused(it) }
val level = PrivacyLevel.max(
settings.privacyLevel,
PrivacyLevel.fromWire(profile.uploads.minAnonymization),
)
val body = redactedForUpload(docJson, level)
val reply = client.uploadRun(settings.serverCredential, body)
archive.markUploaded(runId, profile.name, level.wire)
// Deliberately not echoing `reply`: it is the server's index entry as raw JSON, and
// it ended up rendered verbatim in the UI. Size and level are what a person wants.
UploadOutcome.Sent(profile.name, "as $level, ${body.toByteArray().size} bytes")
} catch (e: VersionRefused) {
UploadOutcome.Incompatible(e.message ?: "the server refused this app's version")
} catch (e: UploadRefused) {
UploadOutcome.Refused(e.message ?: "refused by the server")
} catch (t: Throwable) {
UploadOutcome.Failed(t.message ?: t.javaClass.simpleName)
}
}
}
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,281 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.app
import android.content.Context
import android.content.SharedPreferences
import app.echo_lot.archive.RetentionPolicy
import app.echo_lot.privacy.PrivacyLevel
import java.security.SecureRandom
/**
* User settings for archiving, uploading and anonymization.
*
* Defaults are the conservative reading of "an engineer's tool that still respects the person
* holding it": keep history (that is the point of the archive), never upload without being asked,
* and when uploading, strip identifiers unless the user says this is their own server.
*
* SharedPreferences rather than DataStore because these are a dozen scalars read synchronously at
* the start of a run; a coroutine-flow store would add a dependency and a lifecycle for nothing.
*/
class Settings(context: Context) {
private val prefs: SharedPreferences =
context.getSharedPreferences("echolot-settings", Context.MODE_PRIVATE)
// ---- archive ---------------------------------------------------------------------
var archiveEnabled: Boolean
get() = prefs.getBoolean(ARCHIVE_ENABLED, true)
set(v) = prefs.edit().putBoolean(ARCHIVE_ENABLED, v).apply()
/** 0 = no ceiling. */
var maxRuns: Int
get() = prefs.getInt(MAX_RUNS, 100)
set(v) = prefs.edit().putInt(MAX_RUNS, v.coerceAtLeast(0)).apply()
var maxAgeDays: Int
get() = prefs.getInt(MAX_AGE_DAYS, 90)
set(v) = prefs.edit().putInt(MAX_AGE_DAYS, v.coerceAtLeast(0)).apply()
var maxTotalMb: Int
get() = prefs.getInt(MAX_TOTAL_MB, 64)
set(v) = prefs.edit().putInt(MAX_TOTAL_MB, v.coerceAtLeast(0)).apply()
fun retention(): RetentionPolicy = RetentionPolicy(
enabled = archiveEnabled,
maxRuns = maxRuns,
maxAgeDays = maxAgeDays,
maxTotalBytes = maxTotalMb.toLong() * 1024 * 1024,
)
// ---- measurement -------------------------------------------------------------------
/**
* How long a long run listens, in minutes. Offered as 1 / 5 / 15.
*
* 5 is the default because it is the shortest window in which the things long mode exists to
* catch — a link that flaps, a signal that decays as someone walks around, loss that comes in
* bursts — have a fair chance of happening at least twice. One minute is for checking that the
* mode works at all; fifteen is for chasing something already suspected.
*
* Clamped rather than trusted: a zero-minute long run would produce a document claiming a
* window it never watched, which is the one thing `run.mode` exists to prevent.
*/
var longRunMinutes: Int
get() = prefs.getInt(LONG_RUN_MINUTES, 5).coerceIn(1, 60)
set(v) = prefs.edit().putInt(LONG_RUN_MINUTES, v.coerceIn(1, 60)).apply()
// ---- dev relay -----------------------------------------------------------------------
/**
* Whether this device relays adbd's wireless-debug endpoint to the enrolled server.
*
* Off by default and never implied by anything else: it publishes where this device can be
* reached for debugging, which is a decision rather than a side effect. Intended for a spare
* device parked on a test network — see AdbRelay for why the OnePlus is a poor host for it.
*/
var adbRelayEnabled: Boolean
get() = prefs.getBoolean(ADB_RELAY, false)
set(v) = prefs.edit().putBoolean(ADB_RELAY, v).apply()
// ---- run-duration learning ---------------------------------------------------------
/**
* Learned duration of one test type on THIS device, or null before the first run.
*
* The static Probe.estimatedMs values are only cold-start seeds: real durations depend on
* the phone and the network it stands in (ICMPv6 answers in milliseconds where IPv6 works
* and waits out full timeouts where it does not), so a fixed table is wrong for almost
* everyone almost always. What was measured last time is the only estimate that tracks
* reality.
*/
fun learnedDurationMs(type: String): Long? =
prefs.getLong("$DURATION_PREFIX$type", -1L).takeIf { it > 0 }
/**
* Feeds one measured duration into the estimate — EMA, 70 % old / 30 % new. Heavy enough
* on history that a single odd run (a captive portal stalling DNS) does not whipsaw the
* bar, light enough that a real change (enrolling with a server un-skips three probes)
* converges within a few runs. Recorded whatever the test's status: a probe that skips in
* 2 ms will keep skipping in 2 ms until circumstances change, and then the EMA follows.
*/
fun recordDurationMs(type: String, ms: Long) {
if (ms < 0) return
val key = "$DURATION_PREFIX$type"
val old = prefs.getLong(key, -1L)
val next = if (old <= 0) ms else (old * 7 + ms * 3) / 10
prefs.edit().putLong(key, next).apply()
}
// ---- upload ----------------------------------------------------------------------
/**
* Off by default. Measurement data describes the network the user is standing in; sending it
* anywhere is a decision they make, not one they discover after the fact.
*/
var autoUpload: Boolean
get() = prefs.getBoolean(AUTO_UPLOAD, false)
set(v) = prefs.edit().putBoolean(AUTO_UPLOAD, v).apply()
/** Anonymization applied before a run leaves the device. Never applied to the local archive. */
var privacyLevel: PrivacyLevel
get() = PrivacyLevel.fromWire(prefs.getString(PRIVACY_LEVEL, PrivacyLevel.BALANCED.wire))
set(v) = prefs.edit().putString(PRIVACY_LEVEL, v.wire).apply()
/**
* Whether pseudonyms stay stable across runs. That makes history diffable ("same SSID as
* last week") and is what someone wants on their own server — but it also produces an
* identifier that links a device's uploads, so it is off unless chosen.
*/
var stableSalt: Boolean
get() = prefs.getBoolean(STABLE_SALT, false)
set(v) = prefs.edit().putBoolean(STABLE_SALT, v).apply()
/**
* The device-local secret behind stable pseudonyms. Generated once, never leaves the device,
* and clearing it (via [resetSalt]) breaks the link to everything uploaded before.
*/
fun saltSecret(): ByteArray {
prefs.getString(SALT_SECRET, null)?.let { return hex(it) }
val fresh = ByteArray(32).also { SecureRandom().nextBytes(it) }
prefs.edit().putString(SALT_SECRET, fresh.joinToString("") { "%02x".format(it) }).apply()
return fresh
}
fun resetSalt() = prefs.edit().remove(SALT_SECRET).apply()
// ---- server ----------------------------------------------------------------------
var serverUrl: String
get() = prefs.getString(SERVER_URL, "") ?: ""
set(v) = prefs.edit().putString(SERVER_URL, v.trim()).apply()
var serverPin: String
get() = prefs.getString(SERVER_PIN, "") ?: ""
set(v) = prefs.edit().putString(SERVER_PIN, v.trim()).apply()
/**
* The address the operator handed out, for showing to a person.
*
* Separate from [serverUrl], which is the endpoint actually dialled. They differ when the
* server publishes one public name and points devices at another to select its pinned
* certificate — a detail worth keeping out of the user's face but not out of the settings.
*/
var serverPublicUrl: String
get() = (prefs.getString(SERVER_PUBLIC_URL, "") ?: "").ifBlank { serverUrl }
set(v) = prefs.edit().putString(SERVER_PUBLIC_URL, v.trim()).apply()
/**
* What the server said about itself, last time it was asked: addresses, ports, capabilities.
*
* Cached as a rendered block rather than as fields, because it is shown and never acted on —
* these are facts to read, not settings to apply, and storing them as settings would invite
* exactly the confusion of an editable box that changes nothing.
*/
var serverFacts: String
get() = prefs.getString(SERVER_FACTS, "") ?: ""
set(v) = prefs.edit().putString(SERVER_FACTS, v).apply()
/**
* The server's own addresses, learned from its profile, for reaching it when DNS will not.
*
* Only the primaries: the alternate pair exists for NAT behaviour discovery and does not carry
* the control plane, so falling back to one would fail for a second, unrelated reason.
*/
var serverAddrs: String
get() = prefs.getString(SERVER_ADDRS, "") ?: ""
set(v) = prefs.edit().putString(SERVER_ADDRS, v).apply()
fun serverAddrList(): List<String> =
serverAddrs.split(',').map { it.trim() }.filter { it.isNotEmpty() }
var serverCredential: String
get() = prefs.getString(SERVER_CRED, "") ?: ""
set(v) = prefs.edit().putString(SERVER_CRED, v.trim()).apply()
val serverConfigured: Boolean
get() = serverUrl.isNotBlank() && serverPin.isNotBlank() && serverCredential.isNotBlank()
/**
* The DNS zone this server is authoritative for, learned from its profile.
*
* Cached because the canary probe runs at device tier, before anything has talked to the
* server, and a probe that had to make a control-plane call first would fail on exactly the
* networks worth measuring. Empty means "not known yet", and the probe reports itself as
* skipped rather than inventing a zone.
*/
var canaryZone: String
get() = prefs.getString(CANARY_ZONE, "") ?: ""
set(v) = prefs.edit().putString(CANARY_ZONE, v.trim()).apply()
/**
* Host part of the configured server URL, for probes that address it directly (STUN).
*
* Derived rather than stored: a second copy of the server's name is a second thing to keep in
* step, and it would go stale the moment someone re-enrolled against a different server.
*/
fun serverHost(): String = runCatching {
java.net.URI(serverUrl).host?.takeIf { it.isNotBlank() }
}.getOrNull() ?: ""
// ---- account ---------------------------------------------------------------------
/**
* The PKCE verifier and state for a sign-in that is out at the browser.
*
* Persisted rather than held in memory because handing control to a browser backgrounds this
* process, and Android may kill it before the callback returns. An in-memory value works on a
* developer's device and fails on a phone under memory pressure.
*/
var pendingVerifier: String
get() = prefs.getString(PENDING_VERIFIER, "") ?: ""
set(v) = prefs.edit().putString(PENDING_VERIFIER, v).apply()
var pendingState: String
get() = prefs.getString(PENDING_STATE, "") ?: ""
set(v) = prefs.edit().putString(PENDING_STATE, v).apply()
fun clearPendingAuth() = prefs.edit().remove(PENDING_VERIFIER).remove(PENDING_STATE).apply()
/** Display name of whoever is signed in on this device; empty when nobody is. */
var accountName: String
get() = prefs.getString(ACCOUNT_NAME, "") ?: ""
set(v) = prefs.edit().putString(ACCOUNT_NAME, v).apply()
var accountId: String
get() = prefs.getString(ACCOUNT_ID, "") ?: ""
set(v) = prefs.edit().putString(ACCOUNT_ID, v).apply()
val signedIn: Boolean get() = accountName.isNotBlank()
private fun hex(s: String) = ByteArray(s.length / 2) {
((Character.digit(s[it * 2], 16) shl 4) or Character.digit(s[it * 2 + 1], 16)).toByte()
}
private companion object {
const val ARCHIVE_ENABLED = "archive_enabled"
const val MAX_RUNS = "archive_max_runs"
const val MAX_AGE_DAYS = "archive_max_age_days"
const val MAX_TOTAL_MB = "archive_max_total_mb"
const val AUTO_UPLOAD = "auto_upload"
const val PRIVACY_LEVEL = "privacy_level"
const val STABLE_SALT = "stable_salt"
const val SALT_SECRET = "salt_secret"
const val SERVER_URL = "server_url"
const val SERVER_PIN = "server_pin"
const val SERVER_CRED = "server_credential"
const val SERVER_PUBLIC_URL = "server_public_url"
const val SERVER_FACTS = "server_facts"
const val SERVER_ADDRS = "server_addrs"
const val CANARY_ZONE = "server_canary_zone"
const val PENDING_VERIFIER = "pending_auth_verifier"
const val PENDING_STATE = "pending_auth_state"
const val ACCOUNT_NAME = "account_name"
const val ACCOUNT_ID = "account_id"
const val DURATION_PREFIX = "duration_ms."
const val ADB_RELAY = "adb_relay_enabled"
const val LONG_RUN_MINUTES = "long_run_minutes"
}
}
@@ -0,0 +1,468 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.app
import androidx.compose.foundation.layout.Arrangement
import androidx.compose.foundation.layout.safeDrawingPadding
import androidx.compose.foundation.layout.Column
import androidx.compose.foundation.layout.Row
import androidx.compose.foundation.layout.Spacer
import androidx.compose.foundation.layout.fillMaxWidth
import androidx.compose.foundation.layout.height
import androidx.compose.foundation.layout.width
import androidx.compose.foundation.layout.padding
import androidx.compose.foundation.shape.RoundedCornerShape
import androidx.compose.foundation.rememberScrollState
import androidx.compose.foundation.verticalScroll
import androidx.compose.material3.Button
import androidx.compose.material3.Card
import androidx.compose.material3.FilterChip
import androidx.compose.material3.LocalContentColor
import androidx.compose.material3.MaterialTheme
import androidx.compose.material3.OutlinedTextField
import androidx.compose.material3.Surface
import androidx.compose.material3.Switch
import androidx.compose.material3.Text
import androidx.compose.material3.TextButton
import androidx.compose.runtime.Composable
import androidx.compose.runtime.getValue
import androidx.compose.runtime.mutableStateOf
import androidx.compose.runtime.remember
import androidx.compose.runtime.setValue
import androidx.compose.ui.Alignment
import androidx.compose.ui.Modifier
import androidx.compose.ui.text.font.FontFamily
import androidx.compose.ui.unit.dp
import androidx.compose.ui.unit.sp
import app.echo_lot.privacy.PrivacyLevel
/**
* Archiving, upload and anonymization settings.
*
* The screen is written to make the consequences legible rather than to look tidy: every toggle
* says what it means for the user's data in a sentence, and the privacy levels are described by
* what survives them, because "balanced" on its own tells nobody anything.
*/
@Composable
fun SettingsScreen(
settings: Settings,
archivedRuns: Int,
archivedBytes: Long,
onApplyRetention: () -> Unit,
onDeleteAll: () -> Unit,
onPreviewUpload: () -> Unit,
onCheckServer: () -> Unit,
accountName: String,
onSignIn: () -> Unit,
onSignOut: () -> Unit,
onEnroll: (String) -> Unit,
serverStatus: String?,
enrollStatus: String?,
/** Starts or stops the adb relay service; the toggle only records the preference. */
onRelayChange: (Boolean) -> Unit,
onBack: () -> Unit,
) {
// SharedPreferences is not observable, so mirror each value into Compose state and write
// through on change. A dozen scalars; a store with flows would be ceremony for nothing.
var archiveEnabled by remember { mutableStateOf(settings.archiveEnabled) }
var maxRuns by remember { mutableStateOf(settings.maxRuns.toString()) }
var maxAgeDays by remember { mutableStateOf(settings.maxAgeDays.toString()) }
var maxTotalMb by remember { mutableStateOf(settings.maxTotalMb.toString()) }
var longMinutes by remember { mutableStateOf(settings.longRunMinutes) }
var autoUpload by remember { mutableStateOf(settings.autoUpload) }
var privacy by remember { mutableStateOf(settings.privacyLevel) }
var stableSalt by remember { mutableStateOf(settings.stableSalt) }
var relayOn by remember { mutableStateOf(settings.adbRelayEnabled) }
var enrollLink by remember { mutableStateOf("") }
// The public name, which is what the operator handed out and what a person recognises. The
// endpoint actually dialled is shown beneath it when the two differ, rather than hidden — a
// network engineer debugging a connection wants to see where it really goes.
var serverUrl by remember { mutableStateOf(settings.serverPublicUrl) }
var serverPin by remember { mutableStateOf(settings.serverPin) }
var serverCred by remember { mutableStateOf(settings.serverCredential) }
// Enrolling is asynchronous, so these are re-read when its result lands rather than when the
// button is pressed — reading them immediately showed the previous server's values and looked
// exactly like an enrollment that had silently done nothing.
var serverFacts by remember { mutableStateOf(settings.serverFacts) }
androidx.compose.runtime.LaunchedEffect(enrollStatus, serverStatus) {
serverFacts = settings.serverFacts
serverUrl = settings.serverPublicUrl
serverPin = settings.serverPin
serverCred = settings.serverCredential
}
Column(
Modifier.fillMaxWidth().safeDrawingPadding().verticalScroll(rememberScrollState()).padding(16.dp),
verticalArrangement = Arrangement.spacedBy(12.dp),
) {
Row(verticalAlignment = Alignment.CenterVertically) {
TextButton(onClick = onBack) { Text(" Back") }
Text("Settings", style = MaterialTheme.typography.titleLarge)
}
// ---- measurement ----
Card(Modifier.fillMaxWidth()) {
Column(Modifier.padding(14.dp), verticalArrangement = Arrangement.spacedBy(8.dp)) {
Text("Long runs", style = MaterialTheme.typography.titleMedium)
Text(
"How long a long run keeps listening. The measurements themselves take about " +
"30 seconds either way; the rest of the window is spent watching for " +
"things that only happen sometimes.",
style = MaterialTheme.typography.bodySmall,
)
Row(horizontalArrangement = Arrangement.spacedBy(8.dp)) {
for (minutes in listOf(1, 5, 15)) {
FilterChip(
selected = longMinutes == minutes,
onClick = { longMinutes = minutes; settings.longRunMinutes = minutes },
label = { Text("$minutes min") },
)
}
}
Text(
when (longMinutes) {
1 -> "Barely longer than a quick run — enough to confirm the listeners " +
"work, rarely enough to catch anything intermittent."
15 -> "For a fault you already suspect and have to prove. Keep the screen " +
"on and the device where the problem happens."
else -> "Long enough for a link that drops every couple of minutes to do " +
"it at least once, short enough to wait out."
},
style = MaterialTheme.typography.bodySmall,
)
}
}
// ---- archive ----
Card(Modifier.fillMaxWidth()) {
Column(Modifier.padding(14.dp), verticalArrangement = Arrangement.spacedBy(8.dp)) {
Text("Archive", style = MaterialTheme.typography.titleMedium)
Toggle(
label = "Keep finished runs on this device",
detail = "History is what makes a run comparable later. Archived runs are " +
"stored complete and unredacted — anonymization only applies to uploads.",
checked = archiveEnabled,
) { archiveEnabled = it; settings.archiveEnabled = it }
Text(
"Purge automatically when a run exceeds any of these. 0 turns that limit off.",
style = MaterialTheme.typography.bodySmall,
)
NumberField("Keep at most (runs)", maxRuns) {
maxRuns = it; settings.maxRuns = it.toIntOrNull() ?: 0
}
NumberField("Delete older than (days)", maxAgeDays) {
maxAgeDays = it; settings.maxAgeDays = it.toIntOrNull() ?: 0
}
NumberField("Keep at most (MB)", maxTotalMb) {
maxTotalMb = it; settings.maxTotalMb = it.toIntOrNull() ?: 0
}
Text(
"$archivedRuns run(s), ${archivedBytes / 1024} kB stored",
style = MaterialTheme.typography.bodySmall,
)
Row(horizontalArrangement = Arrangement.spacedBy(8.dp)) {
Button(onClick = onApplyRetention) { Text("Apply now") }
TextButton(onClick = onDeleteAll) { Text("Delete all runs") }
}
}
}
// ---- privacy ----
Card(Modifier.fillMaxWidth()) {
Column(Modifier.padding(14.dp), verticalArrangement = Arrangement.spacedBy(8.dp)) {
Text("What leaves the device", style = MaterialTheme.typography.titleMedium)
Text(
"Applied to uploads only. Measurements, verdicts and finding codes survive " +
"every level — only the parts that identify you or your network change.",
style = MaterialTheme.typography.bodySmall,
)
Row(horizontalArrangement = Arrangement.spacedBy(8.dp)) {
for (level in PrivacyLevel.entries) {
FilterChip(
selected = privacy == level,
onClick = { privacy = level; settings.privacyLevel = level },
label = { Text(level.wire) },
)
}
}
Text(privacyExplanation(privacy), style = MaterialTheme.typography.bodySmall)
// At FULL nothing is pseudonymized, so a salt has nothing to act on. Shown
// disabled rather than hidden: the setting is still stored and still applies the
// moment the level changes, and a control that vanishes hides that fact.
Toggle(
label = "Stable pseudonyms across runs",
detail = if (privacy == PrivacyLevel.FULL) {
"Not used at this level — nothing is pseudonymized, so there is nothing " +
"to keep stable. Choose balanced or strict to use this."
} else {
"Lets you compare uploaded runs over time (same SSID reads the same " +
"each time). It also links your uploads together, so leave it off on " +
"a server you don't run yourself."
},
checked = stableSalt && privacy != PrivacyLevel.FULL,
enabled = privacy != PrivacyLevel.FULL,
) { stableSalt = it; settings.stableSalt = it }
TextButton(onClick = onPreviewUpload) { Text("Preview what an upload would send") }
}
}
// ---- account ----
Card(Modifier.fillMaxWidth()) {
Column(Modifier.padding(14.dp), verticalArrangement = Arrangement.spacedBy(8.dp)) {
Text("Account", style = MaterialTheme.typography.titleMedium)
if (accountName.isNotBlank()) {
Text("Signed in as $accountName", style = MaterialTheme.typography.bodyMedium)
Text(
"Runs from every device signed in to this account share one history.",
style = MaterialTheme.typography.bodySmall,
)
TextButton(onClick = onSignOut) { Text("Sign out") }
} else {
Text(
"Signing in is optional. It links this device to an account on your " +
"server, so several devices share one history — and some servers only " +
"accept uploads from a signed-in device.",
style = MaterialTheme.typography.bodySmall,
)
Button(onClick = onSignIn, enabled = settings.serverConfigured) {
Text("Sign in")
}
if (!settings.serverConfigured) {
Text(
"Enrol with a server first — the account belongs to the server, not " +
"to the app.",
style = MaterialTheme.typography.bodySmall,
)
}
}
}
}
// ---- upload ----
Card(Modifier.fillMaxWidth()) {
Column(Modifier.padding(14.dp), verticalArrangement = Arrangement.spacedBy(8.dp)) {
Text("Upload", style = MaterialTheme.typography.titleMedium)
Toggle(
label = "Upload finished runs automatically",
detail = "Sends each completed run to the server below, anonymized to the " +
"level above. The server may require more anonymization than you chose; " +
"it can never require less.",
checked = autoUpload,
) { autoUpload = it; settings.autoUpload = it }
// Enrollment first, because it is the path that works: one link carries the
// URL, the pin and a single-use token. The three fields below exist for when
// someone has to reconstruct a configuration by hand, not as the normal route.
Text(
"Paste an enrollment link from your server operator, or scan its QR code. " +
"It fills in all three fields below. The link contains a one-time token — " +
"treat it like a password until it is used.",
style = MaterialTheme.typography.bodySmall,
)
OutlinedTextField(
value = enrollLink, onValueChange = { enrollLink = it },
label = { Text("echolot://enroll?…") }, singleLine = true,
textStyle = MaterialTheme.typography.bodySmall.copy(fontFamily = FontFamily.Monospace),
modifier = Modifier.fillMaxWidth(),
)
Button(
onClick = {
onEnroll(enrollLink)
enrollLink = "" // spent either way; leaving it around invites a retry
},
enabled = enrollLink.isNotBlank(),
) { Text("Enroll") }
// Beside the button that caused it. Enrolling is asynchronous, so without this the
// only sign of success is three fields quietly changing further down the card.
enrollStatus?.let {
Text(it, style = MaterialTheme.typography.bodySmall)
}
OutlinedTextField(
value = serverUrl,
onValueChange = {
serverUrl = it
// Typed by hand there is no discovery to consult, so what was entered is
// both the public name and the endpoint. Setting only one of them would
// leave the app dialling the previous server.
settings.serverUrl = it
settings.serverPublicUrl = it
},
label = { Text("Server URL") }, singleLine = true, modifier = Modifier.fillMaxWidth(),
)
// Directly under the field it explains. Anywhere else it reads as a stray sentence
// about some other part of the screen.
if (settings.serverUrl.isNotBlank() && settings.serverUrl != settings.serverPublicUrl) {
Text(
"Connects to ${settings.serverUrl} — this server publishes one name and " +
"points devices at another, so its pinned certificate can share a port " +
"with its web interface.",
style = MaterialTheme.typography.bodySmall,
color = LocalContentColor.current.copy(alpha = 0.7f),
)
}
OutlinedTextField(
value = serverPin, onValueChange = { serverPin = it; settings.serverPin = it },
label = { Text("Certificate pin (SPKI, base64)") }, singleLine = true,
textStyle = MaterialTheme.typography.bodySmall.copy(fontFamily = FontFamily.Monospace),
modifier = Modifier.fillMaxWidth(),
)
OutlinedTextField(
value = serverCred,
onValueChange = { serverCred = it; settings.serverCredential = it },
label = { Text("Device credential") }, singleLine = true,
textStyle = MaterialTheme.typography.bodySmall.copy(fontFamily = FontFamily.Monospace),
modifier = Modifier.fillMaxWidth(),
)
Row(verticalAlignment = Alignment.CenterVertically) {
Button(onClick = onCheckServer) { Text("Check server") }
Text(
" " + if (settings.serverConfigured) "Configured."
else "Uploads stay off until all three fields are set.",
style = MaterialTheme.typography.bodySmall,
)
}
// Version compatibility is checked here rather than discovered mid-run: a server
// that will refuse this build should say so before a measurement is wasted.
serverStatus?.let {
Text(it, style = MaterialTheme.typography.bodySmall)
}
// What the server reported, placed under the button that asks it rather than among
// the fields above: these are facts to read, not settings to apply, and an
// editable-looking box that changes nothing is worse than no box at all.
//
// Monospaced so the addresses line up under each other — column alignment is most
// of what makes a list of IPs quicker to read than prose.
if (serverFacts.isNotBlank()) {
Surface(
color = MaterialTheme.colorScheme.surfaceVariant,
shape = RoundedCornerShape(8.dp),
modifier = Modifier.fillMaxWidth(),
) {
Column(
Modifier.padding(horizontal = 12.dp, vertical = 10.dp),
verticalArrangement = Arrangement.spacedBy(2.dp),
) {
Text(
"WHAT THIS SERVER REPORTS",
style = MaterialTheme.typography.labelSmall,
color = LocalContentColor.current.copy(alpha = 0.7f),
)
// Real columns rather than padded text: the label column has a fixed
// width, so values line up whatever the font does, and a long value
// wraps inside its own column instead of under the labels.
for (line in serverFacts.lines()) {
val label = line.substringBefore('|')
val value = line.substringAfter('|', "")
Row(Modifier.fillMaxWidth()) {
Text(
label,
style = MaterialTheme.typography.bodySmall,
color = LocalContentColor.current.copy(alpha = 0.7f),
modifier = Modifier.width(72.dp),
)
Text(
value,
style = MaterialTheme.typography.bodySmall.copy(
fontFamily = FontFamily.Monospace,
),
modifier = Modifier.weight(1f),
)
}
}
}
}
}
Text(
"This app is ${BuildConfig.APP_SEMVER} and speaks probe protocol " +
"${app.echo_lot.protocol.Compat.PROTOCOL_VERSION}. It works with servers " +
"${app.echo_lot.protocol.Compat.serverRange}.",
style = MaterialTheme.typography.bodySmall,
)
}
}
// ---- dev relay ------------------------------------------------------------------
//
// Last, and deliberately plain: this is scaffolding for driving a test device, not a
// measurement. It publishes where this device can be reached over adb, which is why it
// is off until someone decides otherwise.
Card(Modifier.fillMaxWidth()) {
Column(Modifier.padding(16.dp), verticalArrangement = Arrangement.spacedBy(8.dp)) {
Text("Developer relay", style = MaterialTheme.typography.titleMedium)
Toggle(
label = "Relay this device's adb endpoint",
detail = "Watches adbd's own mDNS announcement and reports host:port to the " +
"enrolled server, so a developer on another network can reach this " +
"device — mDNS does not cross subnets, and the port rotates every few " +
"minutes. Runs in the foreground with a notification while on.",
checked = relayOn,
enabled = settings.serverConfigured,
onChange = { on ->
relayOn = on
settings.adbRelayEnabled = on
onRelayChange(on)
},
)
if (!settings.serverConfigured) {
Text(
"Needs an enrolled server: the report is authenticated with this " +
"device's credential.",
style = MaterialTheme.typography.bodySmall,
color = LocalContentColor.current.copy(alpha = 0.7f),
)
}
}
}
Spacer(Modifier.height(24.dp))
}
}
private fun privacyExplanation(level: PrivacyLevel): String = when (level) {
PrivacyLevel.FULL ->
"Nothing is removed: SSIDs, MAC addresses, hostnames and discovered neighbours are sent " +
"as measured. Appropriate for a server you run yourself."
PrivacyLevel.BALANCED ->
"Network names and hostnames become pseudonyms, MAC addresses keep only their vendor " +
"prefix, public IP addresses keep only their /16, and discovered neighbours (SSDP, " +
"ARP, nearby networks) are dropped entirely. Private addresses stay readable, since " +
"192.168.1.1 describes the topology and not the person."
PrivacyLevel.STRICT ->
"Only numbers: test results, metrics and finding codes. No network description, no raw " +
"evidence, no finding text. Nothing left can identify a network."
}
@Composable
private fun Toggle(
label: String,
detail: String,
checked: Boolean,
enabled: Boolean = true,
onChange: (Boolean) -> Unit,
) {
Row(Modifier.fillMaxWidth(), verticalAlignment = Alignment.Top) {
Column(Modifier.weight(1f)) {
// Dimmed together with the switch, so "this does nothing right now" reads at a glance
// instead of only on close inspection.
val alpha = if (enabled) 1f else 0.5f
Text(label, style = MaterialTheme.typography.bodyMedium, color = LocalContentColor.current.copy(alpha = alpha))
Text(detail, style = MaterialTheme.typography.bodySmall, color = LocalContentColor.current.copy(alpha = alpha))
}
Switch(checked = checked, onCheckedChange = onChange, enabled = enabled)
}
}
@Composable
private fun NumberField(label: String, value: String, onChange: (String) -> Unit) {
OutlinedTextField(
value = value,
onValueChange = { s -> onChange(s.filter { it.isDigit() }.take(7)) },
label = { Text(label) },
singleLine = true,
modifier = Modifier.fillMaxWidth(),
)
}
+23
View File
@@ -0,0 +1,23 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
plugins {
alias(libs.plugins.kotlin.jvm)
alias(libs.plugins.kotlin.serialization)
}
// The on-device run archive: measurement documents on disk, with retention.
// Pure Kotlin/JVM (it takes a directory, not a Context) so the retention rules
// — the part with edge cases — are unit-testable without a device.
dependencies {
implementation(libs.kotlinx.serialization.json)
testImplementation(kotlin("test"))
}
kotlin {
jvmToolchain(21)
compilerOptions { jvmTarget.set(org.jetbrains.kotlin.gradle.dsl.JvmTarget.JVM_17) }
}
java { sourceCompatibility = JavaVersion.VERSION_17; targetCompatibility = JavaVersion.VERSION_17 }
tasks.test { useJUnitPlatform() }
@@ -0,0 +1,225 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
// Package archive keeps completed measurement runs on the device.
//
// The point of an archive is the second run: "this network was fine on Tuesday" is only
// answerable if Tuesday was kept. But an app that silently accumulates network dumps forever is
// its own privacy problem, so retention is a first-class part of the type rather than a cleanup
// job somebody remembers to write — every save enforces it.
//
// Storage is one JSON file per run plus a small index entry, in a plain directory. Nothing here
// needs a database, and a plain directory is something a user can inspect, copy off, or delete
// with a file manager. Files are written to a temp name and renamed, so a run interrupted mid-
// write never leaves a half-document that reads as real.
package app.echo_lot.archive
import kotlinx.serialization.SerialName
import kotlinx.serialization.Serializable
import kotlinx.serialization.json.Json
import kotlinx.serialization.json.JsonObject
import kotlinx.serialization.json.jsonObject
import kotlinx.serialization.json.jsonPrimitive
import java.io.File
/** Index entry for one archived run — enough for a history list without opening the documents. */
@Serializable
data class ArchivedRun(
val id: String,
@SerialName("saved_at_epoch_ms") val savedAtEpochMs: Long,
@SerialName("started_at") val startedAt: String? = null,
val verdict: String? = null,
@SerialName("finding_count") val findingCount: Int = 0,
@SerialName("size_bytes") val sizeBytes: Long = 0,
/**
* How the *archived* document is redacted. Always "full" in practice, because the archive
* deliberately keeps the unredacted run - see the package doc. This is not what was uploaded.
*/
val anonymization: String = "full",
/** Whether this run has been accepted by a server, so history can show what is backed up. */
val uploaded: Boolean = false,
@SerialName("uploaded_to") val uploadedTo: String? = null,
/**
* The level the run was *uploaded* at, which is a different document from the archived one.
*
* Kept separately because conflating the two is actively misleading: the history row showed
* the archive's own level ("full") directly beneath "uploaded to fmr", which reads as "the
* complete data was uploaded" when a redacted copy had been sent. A privacy display that
* overstates what left the device is worse than none.
*/
@SerialName("uploaded_as") val uploadedAs: String? = null,
)
/**
* Retention limits. All three are independent ceilings; a run is dropped when it violates any of
* them. Zero disables that limit.
*
* The default keeps a hundred runs or three months, whichever comes first. That is enough to see
* a pattern ("it degrades every evening") without turning the phone into an archive of every
* network its owner ever walked past.
*/
@Serializable
data class RetentionPolicy(
/**
* Whether to archive at all. Separate from the limits because "no limits" (every limit zero)
* and "keep nothing" are opposite intentions, and collapsing them onto the same value is how
* a user who turns all the caps off ends up with an empty history.
*/
val enabled: Boolean = true,
@SerialName("max_runs") val maxRuns: Int = 100,
@SerialName("max_age_days") val maxAgeDays: Int = 90,
@SerialName("max_total_bytes") val maxTotalBytes: Long = 64L * 1024 * 1024,
) {
companion object {
/** Archiving off: runs are shown once and never written. */
val KeepNothing = RetentionPolicy(enabled = false)
/** Archiving on with no ceilings. Every run is kept until the user deletes it. */
val Unlimited = RetentionPolicy(maxRuns = 0, maxAgeDays = 0, maxTotalBytes = 0)
val Default = RetentionPolicy()
}
val keepsAnything: Boolean get() = enabled
}
/** What a purge removed, so the UI can say "dropped 3 old runs" instead of silently deleting. */
data class PurgeResult(val removed: List<String>, val freedBytes: Long) {
val isEmpty: Boolean get() = removed.isEmpty()
}
class RunArchive(private val dir: File, private val now: () -> Long = System::currentTimeMillis) {
private val json = Json { ignoreUnknownKeys = true; encodeDefaults = true }
init {
dir.mkdirs()
}
/**
* Writes one run and applies retention. Returns the index entry, or null when the policy
* keeps nothing at all — in which case nothing is written, rather than written and instantly
* deleted (the difference matters on flash storage and to anyone watching the filesystem).
*/
fun save(runJson: String, policy: RetentionPolicy = RetentionPolicy.Default): ArchivedRun? {
if (!policy.keepsAnything) return null
val doc = runCatching { json.parseToJsonElement(runJson).jsonObject }.getOrNull() ?: return null
val meta = indexOf(doc, runJson.toByteArray().size.toLong()) ?: return null
writeAtomically(File(dir, meta.id + EXT), runJson)
writeAtomically(File(dir, meta.id + META_EXT), json.encodeToString(ArchivedRun.serializer(), meta))
purge(policy)
return meta
}
/** History, newest first. */
fun list(): List<ArchivedRun> =
(dir.listFiles { f -> f.name.endsWith(META_EXT) } ?: emptyArray())
.mapNotNull { f ->
runCatching { json.decodeFromString(ArchivedRun.serializer(), f.readText()) }.getOrNull()
}
.sortedByDescending { it.savedAtEpochMs }
fun read(id: String): String? = File(dir, safe(id) + EXT).takeIf { it.isFile }?.readText()
fun delete(id: String): Boolean {
val s = safe(id)
val doc = File(dir, s + EXT).delete()
File(dir, s + META_EXT).delete()
return doc
}
fun deleteAll(): Int = list().count { delete(it.id) }
/** Records that a server accepted this run, so history can distinguish backed-up from local. */
fun markUploaded(id: String, serverName: String, uploadedAs: String? = null) {
val f = File(dir, safe(id) + META_EXT)
val meta = runCatching { json.decodeFromString(ArchivedRun.serializer(), f.readText()) }.getOrNull()
?: return
writeAtomically(
f,
json.encodeToString(
ArchivedRun.serializer(),
meta.copy(uploaded = true, uploadedTo = serverName, uploadedAs = uploadedAs),
),
)
}
fun totalBytes(): Long = list().sumOf { it.sizeBytes }
/**
* Enforces the policy. Age first, then total size, then count: dropping stale runs may already
* satisfy the other two, and it is the limit a user reasons about ("keep three months"), so it
* should not be pre-empted by a size sweep deleting last week instead.
*/
fun purge(policy: RetentionPolicy): PurgeResult {
val removed = ArrayList<String>()
var freed = 0L
fun drop(r: ArchivedRun) {
if (delete(r.id)) {
removed.add(r.id)
freed += r.sizeBytes
}
}
var kept = list()
if (policy.maxAgeDays > 0) {
val cutoff = now() - policy.maxAgeDays * 24L * 60 * 60 * 1000
val (fresh, stale) = kept.partition { it.savedAtEpochMs >= cutoff }
stale.forEach(::drop)
kept = fresh
}
if (policy.maxTotalBytes > 0) {
var total = kept.sumOf { it.sizeBytes }
// Oldest first until we are under the ceiling.
for (r in kept.reversed()) {
if (total <= policy.maxTotalBytes) break
drop(r)
total -= r.sizeBytes
}
kept = kept.filter { it.id !in removed }
}
if (policy.maxRuns > 0 && kept.size > policy.maxRuns) {
kept.drop(policy.maxRuns).forEach(::drop) // list() is newest-first
}
return PurgeResult(removed, freed)
}
// ---- internals ---------------------------------------------------------------------
private fun indexOf(doc: JsonObject, size: Long): ArchivedRun? {
val run = doc["run"]?.jsonObject ?: return null
val id = run["id"]?.jsonPrimitive?.content?.let(::safe)?.takeIf { it.isNotEmpty() } ?: return null
return ArchivedRun(
id = id,
savedAtEpochMs = now(),
startedAt = run["started_at"]?.jsonPrimitive?.content,
// The schema calls it `overall` (Summary.overall); reading `verdict` here silently
// yielded null for every run, so the history list's most prominent element - the
// coloured verdict - was blank on every row.
verdict = doc["summary"]?.jsonObject?.get("overall")?.jsonPrimitive?.content,
findingCount = (doc["findings"] as? kotlinx.serialization.json.JsonArray)?.size ?: 0,
sizeBytes = size,
anonymization = run["privacy"]?.jsonObject?.get("anonymization")?.jsonPrimitive?.content ?: "full",
)
}
private fun writeAtomically(target: File, content: String) {
val tmp = File(target.parentFile, target.name + ".tmp")
tmp.writeText(content)
if (!tmp.renameTo(target)) {
target.delete()
tmp.renameTo(target)
}
}
/** Run ids reach the filesystem; keep them to characters that cannot climb out of [dir]. */
private fun safe(id: String): String = buildString {
for (c in id) if (c.isLetterOrDigit() || c == '-' || c == '_') append(c)
}.take(64)
private companion object {
const val EXT = ".json"
const val META_EXT = ".meta.json"
}
}
@@ -0,0 +1,197 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.archive
import java.io.File
import java.nio.file.Files
import kotlin.test.AfterTest
import kotlin.test.Test
import kotlin.test.assertEquals
import kotlin.test.assertFalse
import kotlin.test.assertNotNull
import kotlin.test.assertNull
import kotlin.test.assertTrue
class RunArchiveTest {
private val dir: File = Files.createTempDirectory("echolot-archive").toFile()
private var clock = 1_000_000_000_000L // fixed: retention is time arithmetic, not wall time
private fun archive() = RunArchive(dir) { clock }
@AfterTest fun cleanup() { dir.deleteRecursively() }
private fun doc(id: String, findings: Int = 1, pad: Int = 0): String {
val f = (1..findings).joinToString(",") { """{"id":"f$it"}""" }
return """{"run":{"id":"$id","started_at":"2026-08-01T10:00:00Z","privacy":{"anonymization":"balanced"}},""" +
""""findings":[$f],"summary":{"overall":"warn"},"pad":"${"x".repeat(pad)}"}"""
}
@Test
fun savedRunsComeBackNewestFirst() {
val a = archive()
for (i in 1..3) {
a.save(doc("run-$i"))
clock += 60_000
}
assertEquals(listOf("run-3", "run-2", "run-1"), a.list().map { it.id })
}
@Test
fun theIndexSummarisesTheDocument() {
val meta = assertNotNull(archive().save(doc("run-1", findings = 4)))
assertEquals("warn", meta.verdict)
assertEquals(4, meta.findingCount)
assertEquals("balanced", meta.anonymization)
assertEquals("2026-08-01T10:00:00Z", meta.startedAt)
assertFalse(meta.uploaded)
}
@Test
fun theDocumentComesBackByteForByte() {
val a = archive()
val original = doc("run-1")
a.save(original)
assertEquals(original, a.read("run-1"))
}
@Test
fun countLimitKeepsTheNewest() {
val a = archive()
val policy = RetentionPolicy(maxRuns = 3, maxAgeDays = 0, maxTotalBytes = 0)
for (i in 1..7) {
a.save(doc("run-$i"), policy)
clock += 60_000
}
assertEquals(listOf("run-7", "run-6", "run-5"), a.list().map { it.id })
assertNull(a.read("run-1"), "purged run's document should be gone, not just its index entry")
}
@Test
fun ageLimitDropsRunsPastTheWindow() {
val a = archive()
val policy = RetentionPolicy(maxRuns = 0, maxAgeDays = 7, maxTotalBytes = 0)
a.save(doc("old"), policy)
clock += 30L * 24 * 60 * 60 * 1000 // a month later
a.save(doc("new"), policy)
assertEquals(listOf("new"), a.list().map { it.id })
}
@Test
fun sizeLimitDropsOldestUntilUnderTheCeiling() {
val a = archive()
val one = doc("x", pad = 900).toByteArray().size.toLong()
val policy = RetentionPolicy(maxRuns = 0, maxAgeDays = 0, maxTotalBytes = one * 2 + 10)
for (i in 1..5) {
a.save(doc("run-$i", pad = 900), policy)
clock += 60_000
}
val kept = a.list()
assertTrue(kept.size <= 2, "size ceiling not enforced: kept ${kept.size}")
assertEquals("run-5", kept.first().id, "the newest run must always survive")
assertTrue(a.totalBytes() <= policy.maxTotalBytes)
}
// A policy that keeps nothing must not write-then-delete: the run should never touch storage.
@Test
fun keepNothingWritesNothing() {
val a = archive()
assertNull(a.save(doc("run-1"), RetentionPolicy.KeepNothing))
assertTrue(a.list().isEmpty())
assertEquals(0, dir.listFiles()?.size ?: 0, "files were written for a keep-nothing policy")
}
// "no ceilings" and "keep nothing" must not be the same policy, however a user arrives at
// one: turning every limit off should keep everything, not wipe the history.
@Test
fun unlimitedKeepsEverythingWhileKeepNothingKeepsNone() {
val a = archive()
for (i in 1..5) {
a.save(doc("run-$i"), RetentionPolicy.Unlimited)
clock += 60_000
}
assertEquals(5, a.list().size)
assertNull(a.save(doc("run-6"), RetentionPolicy.KeepNothing))
assertEquals(5, a.list().size, "keep-nothing must not touch what is already archived")
}
@Test
fun uploadStateIsRecorded() {
val a = archive()
a.save(doc("run-1"))
a.markUploaded("run-1", "fmr", "balanced")
val meta = a.list().single()
assertTrue(meta.uploaded)
assertEquals("fmr", meta.uploadedTo)
assertEquals("run-1", meta.id, "marking upload must not disturb the rest of the entry")
}
// The archive's own level and the level a run was uploaded at describe *different documents*.
// Showing the archive's ("full", because the archive is deliberately unredacted) next to
// "uploaded to fmr" reads as "the complete data was uploaded" when a redacted copy was sent —
// a privacy display that overstates what left the device is worse than none.
@Test
fun theUploadedLevelIsRecordedSeparatelyFromTheArchivedOne() {
val a = archive()
// A real archived document carries no privacy stamp: the anonymizer never runs on the
// archive. The shared doc() fixture has one, which is exactly the unrealism that let this
// confusion through in the first place.
a.save("""{"run":{"id":"run-1"},"findings":[],"summary":{"overall":"green"}}""")
a.markUploaded("run-1", "fmr", "balanced")
val meta = a.list().single()
assertEquals("full", meta.anonymization, "the archived copy is unredacted, by design")
assertEquals("balanced", meta.uploadedAs, "the uploaded copy was redacted, and must say so")
}
// The verdict is read from `summary.overall` — the schema's actual field name. Reading
// `summary.verdict` silently yielded null for every run, so the history list's most prominent
// element was blank on every row while everything else looked fine.
@Test
fun theVerdictComesFromTheSchemasOverallField() {
val a = archive()
a.save("""{"run":{"id":"r1"},"findings":[],"summary":{"overall":"yellow"}}""")
assertEquals("yellow", a.list().single().verdict)
}
@Test
fun deleteRemovesBothFiles() {
val a = archive()
a.save(doc("run-1"))
assertTrue(a.delete("run-1"))
assertTrue(a.list().isEmpty())
assertNull(a.read("run-1"))
assertEquals(0, dir.listFiles()?.size ?: 0)
}
@Test
fun malformedInputIsRejectedRatherThanArchived() {
val a = archive()
assertNull(a.save("not json"))
assertNull(a.save("""{"summary":{"verdict":"ok"}}"""), "a document with no run id has no identity")
assertTrue(a.list().isEmpty())
}
// Run ids come from a document that may have been produced elsewhere; they must not be able
// to write outside the archive directory.
@Test
fun runIdsCannotEscapeTheArchiveDirectory() {
val a = archive()
a.save(doc("../../evil"))
val strays = dir.parentFile.listFiles { f -> f.name.contains("evil") } ?: emptyArray()
assertTrue(strays.isEmpty(), "wrote outside the archive: ${strays.toList()}")
}
@Test
fun purgeReportsWhatItRemoved() {
val a = archive()
for (i in 1..5) {
a.save(doc("run-$i"), RetentionPolicy.Unlimited)
clock += 60_000
}
val result = a.purge(RetentionPolicy(maxRuns = 2, maxAgeDays = 0, maxTotalBytes = 0))
assertEquals(3, result.removed.size)
assertTrue(result.freedBytes > 0)
assertEquals(2, a.list().size)
}
}
@@ -1,166 +0,0 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
// Package engine composes core-protocol probes into core-measurement documents — the run engine
// the app drives. This file covers the server-facing vertical (control plane + UDP data plane);
// device-tier probes (link snapshot, Shizuku, local discovery) plug in from the Android modules.
package app.echo_lot.engine
import app.echo_lot.measurement.*
import app.echo_lot.protocol.ControlClient
import app.echo_lot.protocol.ProbeSession
import kotlinx.serialization.json.Json
import kotlinx.serialization.json.JsonObject
import kotlinx.serialization.json.encodeToJsonElement
/**
* Runs the server-facing measurements against one target and assembles a [MeasurementDocument]:
* an ECHO train (RTT distribution, loss, and NAT-rebinding detection from the server's observed
* source port) as `train.udp_updown`. Everything is real evidence with recomputable metrics, and
* findings are derived deterministically. IDs/timestamps are injected so the engine stays pure
* (no clocks/UUIDs of its own) and unit-testable.
*/
class ServerMeasurement(
private val ids: IdSource,
private val app: AppInfo,
private val device: DeviceInfo,
) {
private val json = Json { encodeDefaults = true; explicitNulls = true }
data class Config(
val controlUrl: String,
val pins: Set<String>,
val credential: String,
val target: String,
val udpHost: String,
val udpPort: Int,
val echoCount: Int = 20,
val echoPaddingBytes: Int = 64,
)
fun run(cfg: Config): MeasurementDocument {
val runId = ids.uuid()
val startWall = ids.nowWall()
val startMono = ids.monoNs()
val control = ControlClient(cfg.controlUrl, cfg.pins)
val profile = control.profile(cfg.credential)
val session = control.createSession(cfg.credential, cfg.target)
val serverSession = ServerSession(
id = "sess-1",
profileName = profile.name,
controlUrl = cfg.controlUrl,
serverVersion = profile.serverVersion,
capabilities = profile.capabilities,
sessionId = session.sessionId,
target = SessionTarget(ip4 = cfg.udpHost, udpPort = cfg.udpPort),
)
val (test, findings) = echoTrain(cfg, control, session, startMono)
control.deleteSession(cfg.credential, session.sessionId)
val summary = Verdicts.derive(listOf(test), findings)
return MeasurementDocument(
run = Run(
id = runId, trigger = Trigger.MANUAL, startedAt = startWall, endedAt = ids.nowWall(),
clock = Clock(monoOriginWall = startWall),
app = app, device = device,
tiers = Tiers(app = true),
),
serverSessions = listOf(serverSession),
tests = listOf(test),
findings = findings,
summary = summary,
)
}
private fun echoTrain(
cfg: Config, control: ControlClient, session: app.echo_lot.protocol.SessionResponse, startMono: Long,
): Pair<Test, List<Finding>> {
val testId = ids.uuid()
val seqs = ArrayList<Int>()
val tTx = ArrayList<Long?>()
val tRx = ArrayList<Long?>()
val sizes = ArrayList<Int>()
val rtts = ArrayList<Double>()
val observedPorts = LinkedHashSet<Int>()
ProbeSession(cfg.credential, session, cfg.udpHost, cfg.udpPort).use { ps ->
for (i in 0 until cfg.echoCount) {
val txMono = ids.monoNs() - startMono
val r = ps.echo(cfg.echoPaddingBytes)
seqs.add(i)
tTx.add(txMono)
sizes.add(Wire_HEADER + cfg.echoPaddingBytes)
if (r != null) {
tRx.add(ids.monoNs() - startMono)
rtts.add(r.rttMs)
r.observation?.observedPort?.let { observedPorts.add(it) }
} else {
tRx.add(null)
}
}
}
val sent = cfg.echoCount
val received = rtts.size
val lossPct = if (sent == 0) 0.0 else (sent - received) * 100.0 / sent
val natRebinding = observedPorts.size > 1
val evidence: JsonObject = TrainEvidence(
epochMonoNs = startMono, seq = seqs, tTxNs = tTx, tRxNs = tRx, sizeBytes = sizes,
).toEvidence()
val metrics: JsonObject = json.encodeToJsonElement(
EchoMetrics(
sent = sent, received = received, lossPct = round1(lossPct),
rttMsMin = rtts.minOrNull()?.let(::round1),
rttMsAvg = rtts.average().takeIf { received > 0 }?.let(::round1),
rttMsMax = rtts.maxOrNull()?.let(::round1),
observedPorts = observedPorts.toList(),
natRebindingDetected = natRebinding,
)
) as JsonObject
val status = when {
received == 0 -> TestStatus.FAILED
received < sent -> TestStatus.PARTIAL
else -> TestStatus.OK
}
val test = Test(
id = testId, type = TestType.TRAIN_UDP_UPDOWN, sessionRef = "sess-1", tier = Tier.APP,
startedMonoNs = startMono, endedMonoNs = ids.monoNs(), status = status,
evidence = evidence, metrics = metrics,
)
val findings = ArrayList<Finding>()
if (received == 0) {
findings.add(finding("nat.udp_unreachable", Category.CONNECTIVITY, Severity.HIGH, testId,
"No UDP echo replies from the server",
"Every ECHO probe to the server's UDP data plane was lost — the path blocks or drops the session's UDP traffic."))
} else if (lossPct >= 20.0) {
findings.add(finding("connectivity.udp_loss", Category.CONNECTIVITY, Severity.MEDIUM, testId,
"High UDP loss to the server (${round1(lossPct)}%)",
"A large fraction of ECHO probes were lost, indicating an unreliable UDP path."))
}
if (natRebinding) {
findings.add(finding("nat.udp_rebinding", Category.NAT, Severity.MEDIUM, testId,
"NAT remapped the UDP source port mid-flow",
"The server observed more than one source port for this session (${observedPorts.joinToString()}), i.e. a NAT with a short UDP mapping or per-packet remapping."))
}
return test to findings
}
private fun finding(code: String, cat: Category, sev: Severity, testId: String, title: String, desc: String) =
Finding(
id = ids.uuid(), code = code, category = cat, severity = sev, confidence = Confidence.HIGH,
title = title, description = desc, evidenceRefs = listOf(EvidenceRef(testId)),
)
private companion object {
const val Wire_HEADER = 32
fun round1(v: Double) = Math.round(v * 10.0) / 10.0
}
}
@@ -1,38 +0,0 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.engine
import kotlinx.serialization.SerialName
import kotlinx.serialization.Serializable
import java.time.Instant
import java.util.UUID
/**
* Clock/ID source, injected so the engine has no hidden nondeterminism and stays unit-testable.
* The default uses wall + monotonic clocks and random UUIDs; tests supply deterministic ones.
*/
interface IdSource {
fun uuid(): String
fun monoNs(): Long
fun nowWall(): String
}
class SystemIdSource : IdSource {
override fun uuid(): String = UUID.randomUUID().toString()
override fun monoNs(): Long = System.nanoTime()
override fun nowWall(): String = Instant.now().toString()
}
/** Metrics for train.udp_updown; recomputable from the columnar evidence. */
@Serializable
data class EchoMetrics(
val sent: Int,
val received: Int,
@SerialName("loss_pct") val lossPct: Double,
@SerialName("rtt_ms_min") val rttMsMin: Double? = null,
@SerialName("rtt_ms_avg") val rttMsAvg: Double? = null,
@SerialName("rtt_ms_max") val rttMsMax: Double? = null,
@SerialName("observed_ports") val observedPorts: List<Int> = emptyList(),
@SerialName("nat_rebinding_detected") val natRebindingDetected: Boolean = false,
)
@@ -1,63 +0,0 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.engine
import app.echo_lot.measurement.*
import kotlinx.serialization.json.Json
import kotlin.test.Test
import kotlin.test.assertEquals
import kotlin.test.assertTrue
/**
* Runs the full server-facing engine against a live server and validates the produced
* MeasurementDocument. Self-skips without ECHOLOT_LIVE_* (same contract as core-protocol's live
* test). This is the whole vertical: protocol client → engine → schema document → verdict.
*/
class LiveMeasurementTest {
private val url = System.getenv("ECHOLOT_LIVE_URL")
private val pin = System.getenv("ECHOLOT_LIVE_PIN")
private val cred = System.getenv("ECHOLOT_LIVE_CRED")
private val udp = System.getenv("ECHOLOT_LIVE_UDP")
private val target = System.getenv("ECHOLOT_LIVE_TARGET") ?: "fmr"
@Test
fun producesValidDocumentFromLiveServer() {
if (url == null || pin == null || cred == null || udp == null) {
println("LiveMeasurementTest skipped (no ECHOLOT_LIVE_* env)")
return
}
val (host, port) = udp.split(":").let { it[0] to it[1].toInt() }
val engine = ServerMeasurement(
ids = SystemIdSource(),
app = AppInfo(version = "0.1.0", build = 1, flavor = "test"),
device = DeviceInfo("test", "jvm", 0, "n/a"),
)
val doc = engine.run(
ServerMeasurement.Config(
controlUrl = url, pins = setOf(pin), credential = cred,
target = target, udpHost = host, udpPort = port, echoCount = 20,
)
)
// The document must round-trip and carry the expected structure.
val encoded = Json { encodeDefaults = true }.encodeToString(MeasurementDocument.serializer(), doc)
println("document (${encoded.length} bytes): overall=${doc.summary?.overall}")
assertEquals(1, doc.serverSessions.size)
assertTrue(doc.serverSessions[0].capabilities.contains("udp-probe"))
val test = doc.tests.single()
assertEquals(TestType.TRAIN_UDP_UPDOWN, test.type)
assertTrue(test.status == TestStatus.OK || test.status == TestStatus.PARTIAL,
"expected replies from live server, got ${test.status}")
val metrics = Json.parseToJsonElement(test.metrics.toString())
println("metrics: $metrics")
assertTrue(metrics.toString().contains("rtt_ms_avg"))
assertTrue(doc.summary != null)
// A healthy local->fmr path should be green (no loss, no rebinding) or yellow.
println("summary: ${doc.summary}")
}
}
+2 -1
View File
@@ -12,6 +12,7 @@ plugins {
dependencies { dependencies {
implementation(project(":core-protocol")) implementation(project(":core-protocol"))
implementation(project(":core-measurement")) implementation(project(":core-measurement"))
implementation(project(":core-privacy"))
implementation(libs.kotlinx.serialization.json) implementation(libs.kotlinx.serialization.json)
testImplementation(kotlin("test")) testImplementation(kotlin("test"))
} }
@@ -24,6 +25,6 @@ java { sourceCompatibility = JavaVersion.VERSION_17; targetCompatibility = JavaV
tasks.test { tasks.test {
useJUnitPlatform() useJUnitPlatform()
listOf("ECHOLOT_LIVE_URL","ECHOLOT_LIVE_PIN","ECHOLOT_LIVE_CRED","ECHOLOT_LIVE_UDP","ECHOLOT_LIVE_TARGET") listOf("ECHOLOT_LIVE_URL","ECHOLOT_LIVE_PIN","ECHOLOT_LIVE_CRED","ECHOLOT_LIVE_UDP","ECHOLOT_LIVE_TARGET","ECHOLOT_ENROLL_URI")
.forEach { k -> System.getenv(k)?.let { environment(k, it) } } .forEach { k -> System.getenv(k)?.let { environment(k, it) } }
} }
@@ -0,0 +1,107 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.engine
import kotlinx.serialization.SerialName
import kotlinx.serialization.Serializable
/**
* Splits a round-trip train into its two directions using what the server witnessed.
*
* A round trip can only report that *something* was lost somewhere. That is the least useful form
* of the answer: "3 % loss" sends an engineer looking in both directions at once. The server
* records every packet it received, per sequence number (probe-protocol.md §6), so the two cases
* are actually distinguishable:
*
* - sent, never seen by the server → **upstream** loss
* - seen by the server, reply never arrived → **downstream** loss
*
* The same records give one-way delay *variation* per direction. Absolute one-way delay would
* need synchronised clocks and we deliberately have none (measurement-schema.md's two-clock rule),
* but the variation does not: (server_rx client_tx) contains an unknown constant clock offset,
* and differencing successive samples cancels it. So jitter is honestly attributable to a
* direction even though latency is not.
*/
object Directional {
/** One probe as the client saw it. [tRxNs] null means no reply came back. */
data class Sample(val seq: Int, val tTxNs: Long, val tRxNs: Long?)
/** One probe as the server saw it: its own receive and transmit stamps, on its own clock. */
data class ServerSighting(val seq: Int, val tRxNs: Long, val tTxNs: Long)
fun analyse(sent: List<Sample>, seen: List<ServerSighting>): DirectionalMetrics {
val byServerSeq = seen.associateBy { it.seq }
// Only sequences we actually sent count. A server record for a sequence we have no note
// of is not evidence about this train — it is a bug or a stray, and silently folding it
// in would produce loss percentages above 100 or below zero.
val relevant = sent.filter { byServerSeq.containsKey(it.seq) }
val nSent = sent.size
val nSeen = relevant.size
val nReplied = sent.count { it.tRxNs != null }
// A reply can only exist if the request arrived, so downstream loss is measured against
// what the server saw, not against what we sent — otherwise upstream loss is counted twice.
val lostUp = nSent - nSeen
val lostDown = (nSeen - nReplied).coerceAtLeast(0)
val upDeltas = relevant.sortedBy { it.seq }
.map { byServerSeq.getValue(it.seq).tRxNs - it.tTxNs }
val downDeltas = sent.filter { it.tRxNs != null && byServerSeq.containsKey(it.seq) }
.sortedBy { it.seq }
.map { it.tRxNs!! - byServerSeq.getValue(it.seq).tTxNs }
return DirectionalMetrics(
sent = nSent,
seenByServer = nSeen,
repliesReceived = nReplied,
lostUpstream = lostUp,
lostDownstream = lostDown,
lossUpstreamPct = pct(lostUp, nSent),
// Denominator is what reached the server: of the packets that got there, how many
// replies came back.
lossDownstreamPct = pct(lostDown, nSeen),
jitterUpstreamMs = jitterMs(upDeltas),
jitterDownstreamMs = jitterMs(downDeltas),
/** True when the server saw nothing at all, which is a different fault from loss. */
noneReachedServer = nSent > 0 && nSeen == 0,
)
}
/**
* Mean absolute difference between consecutive one-way samples (RFC 3393 IPDV, averaged).
*
* Differencing is what makes this legitimate without synchronised clocks: each sample carries
* the same unknown offset between the two clocks, and the difference cancels it. Fewer than
* two samples yields null rather than zero — "no jitter" and "not enough data to say" are
* different claims and only one of them is true here.
*/
private fun jitterMs(oneWayNs: List<Long>): Double? {
if (oneWayNs.size < 2) return null
val deltas = oneWayNs.zipWithNext { a, b -> kotlin.math.abs(b - a) }
return round2(deltas.average() / 1_000_000.0)
}
private fun pct(part: Int, whole: Int): Double =
if (whole <= 0) 0.0 else round2(part * 100.0 / whole)
private fun round2(v: Double) = Math.round(v * 100.0) / 100.0
}
/** Directional metrics for train.udp_updown; recomputable from the columnar evidence. */
@Serializable
data class DirectionalMetrics(
val sent: Int,
@SerialName("seen_by_server") val seenByServer: Int,
@SerialName("replies_received") val repliesReceived: Int,
@SerialName("lost_upstream") val lostUpstream: Int,
@SerialName("lost_downstream") val lostDownstream: Int,
@SerialName("loss_upstream_pct") val lossUpstreamPct: Double,
@SerialName("loss_downstream_pct") val lossDownstreamPct: Double,
/** One-way delay variation (RFC 3393), per direction. Null when there were too few samples. */
@SerialName("jitter_upstream_ms") val jitterUpstreamMs: Double? = null,
@SerialName("jitter_downstream_ms") val jitterDownstreamMs: Double? = null,
@SerialName("none_reached_server") val noneReachedServer: Boolean = false,
)
@@ -0,0 +1,510 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.engine
import app.echo_lot.measurement.*
import app.echo_lot.protocol.ControlClient
import app.echo_lot.protocol.ProbeSession
import app.echo_lot.protocol.Wire
import kotlinx.serialization.SerialName
import kotlinx.serialization.Serializable
import kotlinx.serialization.json.Json
import kotlinx.serialization.json.JsonObject
import kotlinx.serialization.json.encodeToJsonElement
/**
* The measurements only the far end can make: what the *downstream* path does to traffic the
* client never asked for packet-by-packet.
*
* A client alone can measure a round trip, and it can find the largest packet it can *send*. It
* cannot find the largest packet it can *receive*, or whether the network drops downstream
* packets independently of upstream ones — those need a server willing to push, which is why the
* protocol gates them behind an asymmetric grant (probe-protocol.md §3.4).
*
* Three separate facts come out, and keeping them separate is the point:
* - `mtu.pmtud_down` — the largest datagram that arrives *unfragmented*. This is the number
* that matters for anything setting DF, and it is only meaningful because the server sets DF.
* - `mtu.frag_delivery` — whether larger datagrams arrive once the network is allowed to
* fragment them. A path can be fine for one and broken for the other; conflating them is how
* you get "MTU is 4000" on a link that drops every DF packet over 1400.
* - `train.udp_downstream` — loss, reordering and arrival spacing in the download direction.
*/
class DownstreamMeasurement(private val ids: IdSource) {
private val json = Json { encodeDefaults = true; explicitNulls = true }
/** How long to wait for a granted burst after the server accepts the action. */
private val collectWindowMs = 4_000L
/**
* Shorter, but long enough to cover the first_last mode's deliberate 250 ms hold plus a
* reassembly. A fragment burst is one datagram: it is here quickly or not at all.
*/
private val fragWindowMs = 1_500L
/**
* Asks the server to send one deliberately-fragmented datagram per ordering, and reports
* which orderings survive the path.
*
* Kernel fragmentation always emits fragments in order, first one first, so an oversized
* datagram can only answer "do fragments get through at all". The interesting fault is about
* ordering: only the *first* fragment carries the UDP ports, so a stateful firewall that has
* not seen it has nothing to match the rest against, and many drop them. That failure is
* invisible to every in-order test and shows up in the field as "large DNS answers fail here"
* or "the tunnel breaks when the MTU drops".
*/
fun fragmentOrdering(
credential: String,
sessionId: String,
control: ControlClient,
probe: ProbeSession,
sessionRef: String,
sizeBytes: Int = 2000,
fragBytes: Int = 576,
): Pair<Test, List<Finding>> {
val testId = ids.uuid()
val started = ids.monoNs()
val delivered = LinkedHashMap<String, Boolean>()
val fragmentCounts = LinkedHashMap<String, Int>()
var unsupported = false
for (mode in FRAG_MODES) {
val reply = runCatching {
control.action(
credential, sessionId,
"""{"action":"frag_send","size_bytes":$sizeBytes,"mode":"$mode","frag_bytes":$fragBytes}""",
)
}
if (reply.isFailure) {
// A server without a raw socket says so; that is a missing capability, not a
// property of the network, and must not be recorded as a failed delivery.
unsupported = true
break
}
parseInt(reply.getOrNull(), "fragments")?.let { fragmentCounts[mode] = it }
// The burst is already on the wire when the action returns (it is sent
// synchronously), so anything that survived is either here or lost.
val got = probe.collectGranted(fragWindowMs).any { it.type == Wire.TYPE_FRAG_DATA }
delivered[mode] = got
}
if (unsupported) {
return Test(
id = testId, type = TestType.MTU_FRAG_ORDERING, sessionRef = sessionRef, tier = Tier.APP,
startedMonoNs = started, endedMonoNs = ids.monoNs(),
status = TestStatus.UNSUPPORTED,
error = TestError("no_raw_socket", "this server cannot craft fragments"),
) to emptyList()
}
val metrics = json.encodeToJsonElement(
FragOrderingMetrics(
sizeBytes = sizeBytes,
fragBytes = fragBytes,
fragmentsPerBurst = fragmentCounts,
deliveredByMode = delivered,
inOrderDelivered = delivered[FRAG_IN_ORDER] == true,
reorderedDelivered = delivered[FRAG_REVERSED] == true,
delayedFirstDelivered = delivered[FRAG_FIRST_LAST] == true,
),
) as JsonObject
val findings = ArrayList<Finding>()
val inOrder = delivered[FRAG_IN_ORDER] == true
val reversed = delivered[FRAG_REVERSED] == true
val firstLast = delivered[FRAG_FIRST_LAST] == true
if (!inOrder) {
findings.add(
finding(
FindingRegistry.FRAGMENTS_BLOCKED, testId,
"IP fragments do not reach this device",
"A fragmented datagram sent in the normal order never arrived. Anything that " +
"relies on fragmentation — large DNS answers over UDP, some VPN traffic — " +
"will fail here rather than slow down.",
),
)
} else if (!reversed || !firstLast) {
// The precise and useful finding: fragments work, but only if they arrive tidily.
val which = buildList {
if (!reversed) add("out of order")
if (!firstLast) add("with the first fragment delayed")
}.joinToString(" or ")
findings.add(
finding(
FindingRegistry.FRAGMENT_REORDER_SENSITIVE, testId,
"Fragments are dropped when they arrive $which",
"In-order fragments are delivered, but the same datagram sent $which is not. " +
"Something on the path only reassembles when the first fragment (the one " +
"carrying the UDP ports) arrives first — typical of a stateful firewall " +
"or NAT. It works until the network reorders, then fails intermittently, " +
"which is the hardest kind of fault to chase.",
),
)
}
return Test(
id = testId, type = TestType.MTU_FRAG_ORDERING, sessionRef = sessionRef, tier = Tier.APP,
startedMonoNs = started, endedMonoNs = ids.monoNs(),
status = if (inOrder) TestStatus.OK else TestStatus.PARTIAL,
metrics = metrics,
) to findings
}
/**
* Runs all three against an already-primed session.
*
* [session] must already have sent at least one ECHO: the grant is bound to the source the
* server has actually observed, so an unprimed session gets a 409 rather than a grant. That
* is the anti-amplification rule doing its job, not an error to work around.
*/
fun run(
credential: String,
sessionId: String,
control: ControlClient,
probe: ProbeSession,
sessionRef: String,
sizes: List<Int> = DEFAULT_SIZES,
trainCount: Int = 100,
trainSizeBytes: Int = 300,
trainIntervalUs: Int = 3_000,
// DSCP to mark the train with (0-63), or -1 to leave packets unmarked. Pairing a marked
// downtrain with the server-observed DSCP of an upstream train is the two-direction
// sec.dscp_ecn_survival measurement.
trainDscp: Int = -1,
): Pair<List<Test>, List<Finding>> {
val tests = ArrayList<Test>()
val findings = ArrayList<Finding>()
val df = bigSend(credential, sessionId, control, probe, sessionRef, sizes, df = true)
val frag = bigSend(credential, sessionId, control, probe, sessionRef, sizes, df = false)
val train = downTrain(
credential, sessionId, control, probe, sessionRef, trainCount, trainSizeBytes,
trainIntervalUs, trainDscp,
)
tests.add(df.test); tests.add(frag.test); tests.add(train.test)
// Fragment ordering only makes sense once we know fragments arrive at all; when they do
// not, the ordering variants would all report "not delivered" and read as three faults
// instead of one.
if (frag.largestDelivered != null) {
val (fragTest, fragFindings) =
fragmentOrdering(credential, sessionId, control, probe, sessionRef)
tests.add(fragTest)
findings.addAll(fragFindings)
}
// A downstream MTU below the classic 1500-byte Ethernet payload is worth saying out loud:
// it is the usual cause of "small requests work, large responses hang".
val pathMtu = df.largestDelivered
if (pathMtu != null && pathMtu > 0) {
val ipMtu = pathMtu + IP_UDP_OVERHEAD4
if (ipMtu < 1500) {
findings.add(
finding(
FindingRegistry.MTU_REDUCED_DOWNSTREAM, df.test.id,
"Downstream path MTU is $ipMtu bytes, below 1500",
"The largest datagram that reached this device without fragmenting was " +
"$pathMtu bytes of payload ($ipMtu on the wire). Tunnels (PPPoE, VPN, " +
"IPv6-in-IPv4) commonly do this; it is only a fault when something " +
"on the path also blocks the ICMP messages that let senders discover it.",
),
)
}
// The dangerous combination: unfragmented large packets vanish AND fragments do too,
// so a sender that never gets told will retransmit into a black hole.
val fragLargest = frag.largestDelivered ?: 0
if (fragLargest <= pathMtu && sizes.any { it > pathMtu }) {
findings.add(
finding(
FindingRegistry.MTU_DOWNSTREAM_BLACKHOLE, frag.test.id,
"Datagrams above $pathMtu bytes are dropped downstream, fragmented or not",
"Nothing larger than $pathMtu bytes arrived, even when the network was " +
"free to fragment it. Traffic that relies on large responses will " +
"stall rather than fail cleanly.",
),
)
}
}
if (train.received == 0) {
findings.add(
finding(
FindingRegistry.DOWNSTREAM_BLOCKED, train.test.id,
"No server-initiated packets arrived",
"The server sent ${train.sent} packets toward this device and none arrived, " +
"while the round-trip echo worked. Something on the path forwards replies " +
"but drops traffic the device did not individually solicit.",
),
)
} else if (train.lossPct >= 5.0) {
findings.add(
finding(
FindingRegistry.LOSS_DOWNSTREAM, train.test.id,
"Downstream loss of ${round1(train.lossPct)}%",
"${train.sent - train.received} of ${train.sent} packets sent toward this " +
"device were lost. Downstream loss is invisible to a round-trip test, " +
"which reports only that *something* was lost somewhere.",
),
)
}
if (train.reordered > 0) {
findings.add(
finding(
FindingRegistry.DOWNSTREAM_REORDER, train.test.id,
"${train.reordered} downstream packet(s) arrived out of order",
"Packets arrived in a different order than they were sent. Usually per-packet " +
"load balancing across links; harmless for most traffic, not for all of it.",
),
)
}
return tests to findings
}
// ---- big_send ---------------------------------------------------------------------
private class SizeResult(val test: Test, val largestDelivered: Int?)
private fun bigSend(
credential: String, sessionId: String, control: ControlClient, probe: ProbeSession,
sessionRef: String, sizes: List<Int>, df: Boolean,
): SizeResult {
val testId = ids.uuid()
val started = ids.monoNs()
val requested = sizes.joinToString(",")
val reply = runCatching {
control.action(
credential, sessionId,
"""{"action":"big_send","df":$df,"sizes_bytes":[$requested]}""",
)
}
if (reply.isFailure) {
return SizeResult(
Test(
id = testId, type = if (df) TestType.MTU_PMTUD_DOWN else TestType.MTU_FRAG_DELIVERY,
sessionRef = sessionRef, tier = Tier.APP,
startedMonoNs = started, endedMonoNs = ids.monoNs(),
status = TestStatus.UNSUPPORTED,
error = TestError("action_refused", reply.exceptionOrNull()?.message ?: "big_send refused"),
),
null,
)
}
// The server tells us which sizes it actually put on the wire. With DF it refuses
// anything above its own egress MTU, and treating those as "lost downstream" would
// blame the client's network for our own limit.
val accepted = parseIntArray(reply.getOrNull(), "sizes_bytes").ifEmpty { sizes }
val serverMaxDf = parseInt(reply.getOrNull(), "max_df_bytes")
val arrived = probe.collectGranted(collectWindowMs)
.filter { it.type == Wire.TYPE_BIG_SEND }
.map { it.sizeBytes }
.distinct()
.sorted()
val largest = arrived.maxOrNull()
val metrics = json.encodeToJsonElement(
BigSendMetrics(
requestedBytes = sizes,
sentBytes = accepted,
deliveredBytes = arrived,
largestDeliveredBytes = largest,
dontFragment = df,
serverMaxDfBytes = serverMaxDf,
// Only meaningful for the DF run; the IP-level MTU is the payload plus headers.
pathMtuBytes = if (df && largest != null) largest + IP_UDP_OVERHEAD4 else null,
),
) as JsonObject
val status = when {
arrived.isEmpty() -> TestStatus.FAILED
arrived.size < accepted.size -> TestStatus.PARTIAL
else -> TestStatus.OK
}
return SizeResult(
Test(
id = testId, type = if (df) TestType.MTU_PMTUD_DOWN else TestType.MTU_FRAG_DELIVERY,
sessionRef = sessionRef, tier = Tier.APP,
startedMonoNs = started, endedMonoNs = ids.monoNs(),
status = status, metrics = metrics,
),
largest,
)
}
// ---- downtrain --------------------------------------------------------------------
private class TrainResult(
val test: Test, val sent: Int, val received: Int, val lossPct: Double, val reordered: Int,
)
private fun downTrain(
credential: String, sessionId: String, control: ControlClient, probe: ProbeSession,
sessionRef: String, count: Int, sizeBytes: Int, intervalUs: Int, dscp: Int = -1,
): TrainResult {
val testId = ids.uuid()
val started = ids.monoNs()
// dscp is only sent when requested: an older server rejects unknown-value problems
// louder than absent keys, and unmarked is the correct default for a plain loss train.
val dscpField = if (dscp in 0..63) ""","dscp":$dscp""" else ""
val reply = runCatching {
control.action(
credential, sessionId,
"""{"action":"downtrain","count":$count,"size_bytes":$sizeBytes,""" +
""""interval_us":$intervalUs$dscpField}""",
)
}
if (reply.isFailure) {
return TrainResult(
Test(
id = testId, type = TestType.TRAIN_UDP_DOWNSTREAM, sessionRef = sessionRef,
tier = Tier.APP, startedMonoNs = started, endedMonoNs = ids.monoNs(),
status = TestStatus.UNSUPPORTED,
error = TestError("action_refused", reply.exceptionOrNull()?.message ?: "downtrain refused"),
),
0, 0, 0.0, 0,
)
}
val sent = parseInt(reply.getOrNull(), "count") ?: count
val got = probe.collectGranted(collectWindowMs).filter { it.type == Wire.TYPE_DOWNTRAIN_DATA }
val received = got.size
val lossPct = if (sent == 0) 0.0 else (sent - received) * 100.0 / sent
// Reordering: a packet whose sequence is below the highest already seen. Counting
// inversions rather than "not sorted" keeps one late packet from being reported as
// dozens of reorder events.
var highest = -1
var reordered = 0
for (p in got) {
if (p.seq < highest) reordered++ else highest = p.seq
}
// Columnar evidence per the schema: what arrived, when, and how big — so every metric
// above is recomputable by a reader who does not trust our arithmetic.
val evidence = TrainEvidence(
epochMonoNs = started,
seq = got.map { it.seq },
tTxNs = got.map { null },
tRxNs = got.map { it.tRxNs },
sizeBytes = got.map { it.sizeBytes },
).toEvidence()
val interArrival = got.zipWithNext { a, b -> (b.tRxNs - a.tRxNs) / 1_000_000.0 }
val metrics = json.encodeToJsonElement(
DownTrainMetrics(
sent = sent, received = received, lossPct = round1(lossPct),
reorderedPackets = reordered,
sizeBytes = sizeBytes,
interArrivalMsAvg = interArrival.average().takeIf { interArrival.isNotEmpty() }?.let(::round1),
interArrivalMsMax = interArrival.maxOrNull()?.let(::round1),
sendIntervalUs = intervalUs,
dscpRequested = dscp.takeIf { it in 0..63 },
// The server says whether it could actually mark (dscp_applied); recorded so a
// survival comparison never blames the path for a marking the sender skipped.
dscpApplied = reply.getOrNull()?.let {
Regex("\"dscp_applied\"\\s*:\\s*(true|false)").find(it)?.groupValues?.get(1)?.toBoolean()
},
),
) as JsonObject
val status = when {
received == 0 -> TestStatus.FAILED
received < sent -> TestStatus.PARTIAL
else -> TestStatus.OK
}
return TrainResult(
Test(
id = testId, type = TestType.TRAIN_UDP_DOWNSTREAM, sessionRef = sessionRef, tier = Tier.APP,
startedMonoNs = started, endedMonoNs = ids.monoNs(),
status = status, evidence = evidence, metrics = metrics,
),
sent, received, lossPct, reordered,
)
}
// ---- helpers ----------------------------------------------------------------------
/**
* Builds a finding from a registry entry, which supplies the code, category and severity.
*
* Taking a [FindingSpec] rather than three loose values is the point: a typo becomes a
* compile error, and two call sites cannot disagree about which category a finding belongs
* to - a disagreement that would split one fault across two verdict lights.
*/
private fun finding(spec: FindingSpec, testId: String, title: String, desc: String) =
Finding(
id = ids.uuid(), code = spec.code, category = spec.category, severity = spec.severity,
confidence = Confidence.HIGH,
title = title, description = desc, evidenceRefs = listOf(EvidenceRef(testId)),
)
/** Minimal scalar extraction from the action reply; the shape is small and server-owned. */
private fun parseInt(body: String?, key: String): Int? =
body?.let { Regex("\"$key\"\\s*:\\s*(-?\\d+)").find(it)?.groupValues?.get(1)?.toIntOrNull() }
private fun parseIntArray(body: String?, key: String): List<Int> =
body?.let { b ->
Regex("\"$key\"\\s*:\\s*\\[([^\\]]*)\\]").find(b)?.groupValues?.get(1)
?.split(",")?.mapNotNull { it.trim().toIntOrNull() }
} ?: emptyList()
private companion object {
/** IPv4 (20) + UDP (8). The v6 case is 48; reported per-family once v6 sessions land. */
const val IP_UDP_OVERHEAD4 = 28
const val FRAG_IN_ORDER = "in_order"
const val FRAG_REVERSED = "reversed"
const val FRAG_FIRST_LAST = "first_last"
val FRAG_MODES = listOf(FRAG_IN_ORDER, FRAG_REVERSED, FRAG_FIRST_LAST)
/** Straddles the usual suspects: 1500 Ethernet, 1492 PPPoE, 1400-ish tunnels. */
val DEFAULT_SIZES = listOf(600, 1200, 1372, 1400, 1450, 1472, 1500, 2000, 4000)
fun round1(v: Double) = Math.round(v * 10.0) / 10.0
}
}
/** Metrics for mtu.pmtud_down / mtu.frag_delivery. */
@Serializable
data class BigSendMetrics(
@SerialName("requested_bytes") val requestedBytes: List<Int>,
@SerialName("sent_bytes") val sentBytes: List<Int>,
@SerialName("delivered_bytes") val deliveredBytes: List<Int>,
@SerialName("largest_delivered_bytes") val largestDeliveredBytes: Int? = null,
@SerialName("dont_fragment") val dontFragment: Boolean,
/** The server's own DF ceiling; sizes above it were never sent and are not path evidence. */
@SerialName("server_max_df_bytes") val serverMaxDfBytes: Int? = null,
@SerialName("path_mtu_bytes") val pathMtuBytes: Int? = null,
)
/** Metrics for mtu.frag_ordering. */
@Serializable
data class FragOrderingMetrics(
@SerialName("size_bytes") val sizeBytes: Int,
@SerialName("frag_bytes") val fragBytes: Int,
@SerialName("fragments_per_burst") val fragmentsPerBurst: Map<String, Int>,
@SerialName("delivered_by_mode") val deliveredByMode: Map<String, Boolean>,
@SerialName("in_order_delivered") val inOrderDelivered: Boolean,
@SerialName("reordered_delivered") val reorderedDelivered: Boolean,
@SerialName("delayed_first_delivered") val delayedFirstDelivered: Boolean,
)
/** Metrics for train.udp_downstream. */
@Serializable
data class DownTrainMetrics(
val sent: Int,
val received: Int,
@SerialName("loss_pct") val lossPct: Double,
@SerialName("reordered_packets") val reorderedPackets: Int,
@SerialName("size_bytes") val sizeBytes: Int,
@SerialName("inter_arrival_ms_avg") val interArrivalMsAvg: Double? = null,
@SerialName("inter_arrival_ms_max") val interArrivalMsMax: Double? = null,
@SerialName("send_interval_us") val sendIntervalUs: Int,
@SerialName("dscp_requested") val dscpRequested: Int? = null,
@SerialName("dscp_applied") val dscpApplied: Boolean? = null,
)
@@ -10,8 +10,13 @@ import app.echo_lot.measurement.*
import app.echo_lot.protocol.ControlClient import app.echo_lot.protocol.ControlClient
import app.echo_lot.protocol.ProbeSession import app.echo_lot.protocol.ProbeSession
import kotlinx.serialization.json.Json import kotlinx.serialization.json.Json
import kotlinx.serialization.json.JsonArray
import kotlinx.serialization.json.JsonObject import kotlinx.serialization.json.JsonObject
import kotlinx.serialization.json.encodeToJsonElement import kotlinx.serialization.json.encodeToJsonElement
import kotlinx.serialization.json.intOrNull
import kotlinx.serialization.json.jsonObject
import kotlinx.serialization.json.jsonPrimitive
import kotlinx.serialization.json.longOrNull
/** /**
* Runs the server-facing measurements against one target and assembles a [MeasurementDocument]: * Runs the server-facing measurements against one target and assembles a [MeasurementDocument]:
@@ -36,6 +41,20 @@ class ServerMeasurement(
val udpPort: Int, val udpPort: Int,
val echoCount: Int = 20, val echoCount: Int = 20,
val echoPaddingBytes: Int = 64, val echoPaddingBytes: Int = 64,
/**
* Whether to ask the server to push traffic back (downstream MTU and downstream train).
* Costs a few hundred kB of download and needs a server that advertises the grants, so
* it is a flag rather than an assumption.
*/
val downstream: Boolean = true,
/**
* Throughput moves real data — a 5-second run at 50 Mbps is about 30 MB — so it is off
* unless asked for. On a metered mobile connection that is the user's money, and a
* measurement tool that spends it without being told to is not one people keep installed.
*/
val throughput: Boolean = false,
@Suppress("unused") val throughputSeconds: Int = 5,
@Suppress("unused") val throughputKbps: Int = 50_000,
) )
fun run(cfg: Config): MeasurementDocument { fun run(cfg: Config): MeasurementDocument {
@@ -57,11 +76,49 @@ class ServerMeasurement(
target = SessionTarget(ip4 = cfg.udpHost, udpPort = cfg.udpPort), target = SessionTarget(ip4 = cfg.udpHost, udpPort = cfg.udpPort),
) )
val (test, findings) = echoTrain(cfg, control, session, startMono) val tests = ArrayList<Test>()
val allFindings = ArrayList<Finding>()
// One ProbeSession for the whole run. A second one would open a new socket and restart
// the sequence counter, which the server's anti-replay window correctly rejects — so the
// re-primed source is never recorded and every granted send goes to the old, closed port.
// Session identity lives on the server; the socket must live as long as it does.
ProbeSession(cfg.credential, session, cfg.udpHost, cfg.udpPort).use { ps ->
val (test, findings) = echoTrain(cfg, ps, startMono, control, session.sessionId)
tests.add(test)
allFindings.addAll(findings)
// The upstream train needs no grant and no capability beyond udp-probe itself; a
// server that predates trains simply never answers the report request, which the
// measurement reports as exactly that ambiguity rather than as network loss.
val (utTest, utFindings) = UpstreamTrainMeasurement(ids)
.run(ps, sessionRef = "sess-1")
tests.add(utTest)
allFindings.addAll(utFindings)
// Downstream needs a session the server has already seen traffic from — the echo
// train just provided that — and a server that advertises the grants. Skipped
// quietly against an older server rather than reported as a failure of the network.
if (cfg.downstream && profile.supports("downtrain") && profile.supports("big-send")) {
val (dsTests, dsFindings) = DownstreamMeasurement(ids)
.run(cfg.credential, session.sessionId, control, ps, sessionRef = "sess-1")
tests.addAll(dsTests)
allFindings.addAll(dsFindings)
}
if (cfg.throughput && profile.supports("throughput")) {
val (tpTest, tpFindings) = ThroughputMeasurement(ids).run(
cfg.credential, session.sessionId, control, ps, sessionRef = "sess-1",
durationS = cfg.throughputSeconds, kbps = cfg.throughputKbps,
)
tests.add(tpTest)
allFindings.addAll(tpFindings)
}
}
control.deleteSession(cfg.credential, session.sessionId) control.deleteSession(cfg.credential, session.sessionId)
val summary = Verdicts.derive(listOf(test), findings) val summary = Verdicts.derive(tests, allFindings)
return MeasurementDocument( return MeasurementDocument(
run = Run( run = Run(
id = runId, trigger = Trigger.MANUAL, startedAt = startWall, endedAt = ids.nowWall(), id = runId, trigger = Trigger.MANUAL, startedAt = startWall, endedAt = ids.nowWall(),
@@ -70,14 +127,15 @@ class ServerMeasurement(
tiers = Tiers(app = true), tiers = Tiers(app = true),
), ),
serverSessions = listOf(serverSession), serverSessions = listOf(serverSession),
tests = listOf(test), tests = tests,
findings = findings, findings = allFindings,
summary = summary, summary = summary,
) )
} }
private fun echoTrain( private fun echoTrain(
cfg: Config, control: ControlClient, session: app.echo_lot.protocol.SessionResponse, startMono: Long, cfg: Config, ps: ProbeSession, startMono: Long,
control: ControlClient? = null, sessionId: String? = null,
): Pair<Test, List<Finding>> { ): Pair<Test, List<Finding>> {
val testId = ids.uuid() val testId = ids.uuid()
val seqs = ArrayList<Int>() val seqs = ArrayList<Int>()
@@ -87,23 +145,41 @@ class ServerMeasurement(
val rtts = ArrayList<Double>() val rtts = ArrayList<Double>()
val observedPorts = LinkedHashSet<Int>() val observedPorts = LinkedHashSet<Int>()
ProbeSession(cfg.credential, session, cfg.udpHost, cfg.udpPort).use { ps -> // Wire sequence numbers, kept so the server's observations can be correlated packet by
for (i in 0 until cfg.echoCount) { // packet. They are not 0..n-1: the counter is shared with every other packet type on the
val txMono = ids.monoNs() - startMono // session, so "the nth echo" is not "sequence n".
val r = ps.echo(cfg.echoPaddingBytes) val wireSeqs = ArrayList<Int>()
seqs.add(i)
tTx.add(txMono) for (i in 0 until cfg.echoCount) {
sizes.add(Wire_HEADER + cfg.echoPaddingBytes) val txMono = ids.monoNs() - startMono
if (r != null) { val r = ps.echo(cfg.echoPaddingBytes)
tRx.add(ids.monoNs() - startMono) wireSeqs.add(ps.lastSeq)
rtts.add(r.rttMs) seqs.add(i)
r.observation?.observedPort?.let { observedPorts.add(it) } tTx.add(txMono)
} else { sizes.add(Wire_HEADER + cfg.echoPaddingBytes)
tRx.add(null) if (r != null) {
} tRx.add(ids.monoNs() - startMono)
rtts.add(r.rttMs)
r.observation?.observedPort?.let { observedPorts.add(it) }
} else {
tRx.add(null)
} }
} }
// Ask the server what it actually received. This is what turns "3 % loss somewhere" into
// "3 % loss upstream" - the least useful form of the answer into a usable one.
val directional: DirectionalMetrics? =
if (control != null && sessionId != null) {
runCatching {
val samples = wireSeqs.indices.map {
Directional.Sample(wireSeqs[it], tTx[it] ?: 0L, tRx[it])
}
Directional.analyse(samples, serverSightings(control, cfg, sessionId))
}.getOrNull() // an older server without the endpoint simply yields no split
} else {
null
}
val sent = cfg.echoCount val sent = cfg.echoCount
val received = rtts.size val received = rtts.size
val lossPct = if (sent == 0) 0.0 else (sent - received) * 100.0 / sent val lossPct = if (sent == 0) 0.0 else (sent - received) * 100.0 / sent
@@ -113,6 +189,9 @@ class ServerMeasurement(
epochMonoNs = startMono, seq = seqs, tTxNs = tTx, tRxNs = tRx, sizeBytes = sizes, epochMonoNs = startMono, seq = seqs, tTxNs = tTx, tRxNs = tRx, sizeBytes = sizes,
).toEvidence() ).toEvidence()
val directionalJson = directional?.let {
json.encodeToJsonElement(DirectionalMetrics.serializer(), it) as JsonObject
}
val metrics: JsonObject = json.encodeToJsonElement( val metrics: JsonObject = json.encodeToJsonElement(
EchoMetrics( EchoMetrics(
sent = sent, received = received, lossPct = round1(lossPct), sent = sent, received = received, lossPct = round1(lossPct),
@@ -122,7 +201,7 @@ class ServerMeasurement(
observedPorts = observedPorts.toList(), observedPorts = observedPorts.toList(),
natRebindingDetected = natRebinding, natRebindingDetected = natRebinding,
) )
) as JsonObject ).let { base -> JsonObject((base as JsonObject) + (directionalJson ?: JsonObject(emptyMap()))) }
val status = when { val status = when {
received == 0 -> TestStatus.FAILED received == 0 -> TestStatus.FAILED
@@ -137,30 +216,91 @@ class ServerMeasurement(
val findings = ArrayList<Finding>() val findings = ArrayList<Finding>()
if (received == 0) { if (received == 0) {
findings.add(finding("nat.udp_unreachable", Category.CONNECTIVITY, Severity.HIGH, testId, findings.add(finding(FindingRegistry.UDP_UNREACHABLE, testId,
"No UDP echo replies from the server", "No UDP echo replies from the server",
"Every ECHO probe to the server's UDP data plane was lost — the path blocks or drops the session's UDP traffic.")) "Every ECHO probe to the server's UDP data plane was lost — the path blocks or drops the session's UDP traffic."))
} else if (lossPct >= 20.0) { } else if (lossPct >= 20.0) {
findings.add(finding("connectivity.udp_loss", Category.CONNECTIVITY, Severity.MEDIUM, testId, findings.add(finding(FindingRegistry.UDP_LOSS, testId,
"High UDP loss to the server (${round1(lossPct)}%)", "High UDP loss to the server (${round1(lossPct)}%)",
"A large fraction of ECHO probes were lost, indicating an unreliable UDP path.")) "A large fraction of ECHO probes were lost, indicating an unreliable UDP path."))
} }
// Naming the direction is the entire value of the split, so the findings do.
directional?.let { d ->
when {
d.noneReachedServer && received == 0 -> findings.add(
finding(FindingRegistry.UDP_UNREACHABLE_UPSTREAM, testId,
"Nothing reached the server",
"The server received none of the ${d.sent} probes, so the traffic is being " +
"dropped on the way out, not on the way back. A firewall or NAT on " +
"this side of the path is the place to look."),
)
d.lossUpstreamPct >= 2.0 -> findings.add(
finding(FindingRegistry.LOSS_UPSTREAM, testId,
"${d.lossUpstreamPct} % of probes were lost on the way to the server",
"${d.lostUpstream} of ${d.sent} probes never reached the server. The " +
"return path is not implicated: replies came back for everything that " +
"arrived."),
)
}
if (d.lossDownstreamPct >= 2.0) {
findings.add(
finding(FindingRegistry.LOSS_DOWNSTREAM, testId,
"${d.lossDownstreamPct} % of replies were lost on the way back",
"The server received ${d.seenByServer} probes and answered them, but " +
"${d.lostDownstream} of those replies never arrived. The outbound path " +
"is fine; the fault is on the return leg."),
)
}
}
if (natRebinding) { if (natRebinding) {
findings.add(finding("nat.udp_rebinding", Category.NAT, Severity.MEDIUM, testId, findings.add(finding(FindingRegistry.NAT_UDP_REBINDING, testId,
"NAT remapped the UDP source port mid-flow", "NAT remapped the UDP source port mid-flow",
"The server observed more than one source port for this session (${observedPorts.joinToString()}), i.e. a NAT with a short UDP mapping or per-packet remapping.")) "The server observed more than one source port for this session (${observedPorts.joinToString()}), i.e. a NAT with a short UDP mapping or per-packet remapping."))
} }
return test to findings return test to findings
} }
private fun finding(code: String, cat: Category, sev: Severity, testId: String, title: String, desc: String) = /**
* The server's per-packet record of this session's echoes (spec section 6). Filtered to
* ECHO_REQ, because the observation list also holds MTU probes and anything else we sent -
* counting those as train packets would invent loss that is not there.
*/
private fun serverSightings(
control: ControlClient, cfg: Config, sessionId: String,
): List<Directional.ServerSighting> {
val body = control.observations(cfg.credential, sessionId)
val packets = Json.parseToJsonElement(body).jsonObject["udp"]
?.jsonObject?.get("packets") as? JsonArray ?: return emptyList()
return packets.mapNotNull { el ->
val o = el as? JsonObject ?: return@mapNotNull null
val type = o["type"]?.jsonPrimitive?.intOrNull ?: return@mapNotNull null
if (type != ECHO_REQ_TYPE) return@mapNotNull null
Directional.ServerSighting(
seq = o["seq"]?.jsonPrimitive?.intOrNull ?: return@mapNotNull null,
tRxNs = o["t_rx_ns"]?.jsonPrimitive?.longOrNull ?: return@mapNotNull null,
tTxNs = o["t_tx_ns"]?.jsonPrimitive?.longOrNull ?: return@mapNotNull null,
)
}
}
/**
* Builds a finding from a registry entry, which supplies the code, category and severity.
*
* Taking a [FindingSpec] rather than three loose values is the point: a typo becomes a
* compile error, and two call sites cannot disagree about which category a finding belongs
* to - a disagreement that would split one fault across two verdict lights.
*/
private fun finding(spec: FindingSpec, testId: String, title: String, desc: String) =
Finding( Finding(
id = ids.uuid(), code = code, category = cat, severity = sev, confidence = Confidence.HIGH, id = ids.uuid(), code = spec.code, category = spec.category, severity = spec.severity,
confidence = Confidence.HIGH,
title = title, description = desc, evidenceRefs = listOf(EvidenceRef(testId)), title = title, description = desc, evidenceRefs = listOf(EvidenceRef(testId)),
) )
private companion object { private companion object {
const val Wire_HEADER = 32 const val Wire_HEADER = 32
const val ECHO_REQ_TYPE = 0x01
fun round1(v: Double) = Math.round(v * 10.0) / 10.0 fun round1(v: Double) = Math.round(v * 10.0) / 10.0
} }
} }
@@ -0,0 +1,338 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.engine
import app.echo_lot.measurement.*
import app.echo_lot.protocol.ControlClient
import app.echo_lot.protocol.ProbeSession
import app.echo_lot.protocol.Wire
import kotlinx.serialization.SerialName
import kotlinx.serialization.Serializable
import kotlinx.serialization.json.Json
import kotlinx.serialization.json.JsonObject
import kotlinx.serialization.json.encodeToJsonElement
import kotlinx.serialization.json.jsonArray
import kotlinx.serialization.json.jsonObject
import kotlinx.serialization.json.jsonPrimitive
/**
* Downstream throughput: the server sends at a paced rate for a bounded time and the client
* measures what arrives (`perf.throughput_udp`).
*
* The number this produces is only meaningful with a qualifier attached, and getting that
* qualifier right is most of the work here. A throughput test reports the *smallest* limit on the
* path, and the sender's own ceiling is one of the candidates: if the server was asked for 50 Mbps
* and 50 Mbps arrived, the network was never the constraint and "50 Mbps" says nothing about it.
* Reporting that as a capacity measurement would be a confident lie, so the result always carries
* [ThroughputMetrics.limitedBy] and a finding is only raised when the network is actually
* implicated.
*
* Comparing against the *sender's* count rather than the requested rate is the other half: the
* server reports how much it actually put on the wire, and the gap between that and what arrived
* is the loss. A receiver alone cannot tell "the network dropped it" from "the sender never sent
* it", and guessing turns a healthy server-side limit into a phantom network fault.
*/
class ThroughputMeasurement(private val ids: IdSource) {
private val json = Json { encodeDefaults = true; explicitNulls = true }
fun run(
credential: String,
sessionId: String,
control: ControlClient,
probe: ProbeSession,
sessionRef: String,
durationS: Int = 5,
kbps: Int = 50_000,
sizeBytes: Int = 1200,
): Pair<Test, List<Finding>> {
val testId = ids.uuid()
val started = ids.monoNs()
val reply = runCatching {
control.action(
credential, sessionId,
"""{"action":"throughput","direction":"down","duration_s":$durationS,""" +
""""kbps":$kbps,"size_bytes":$sizeBytes}""",
)
}
if (reply.isFailure) {
return Test(
id = testId, type = TestType.PERF_THROUGHPUT_UDP, sessionRef = sessionRef, tier = Tier.APP,
startedMonoNs = started, endedMonoNs = ids.monoNs(),
status = TestStatus.UNSUPPORTED,
error = TestError("action_refused", reply.exceptionOrNull()?.message ?: "throughput refused"),
) to emptyList()
}
// The server may have shortened the run to fit its own byte budget; listen for what it
// actually promised, not for what we asked.
val plannedMs = parseInt(reply.getOrNull(), "duration_ms") ?: (durationS * 1000)
// A margin past the planned end so the tail of the run is not counted as loss: packets
// still in flight when we stop listening were not dropped, they were merely late.
val received = probe.collectGranted(plannedMs + 1_500L)
.filter { it.type == Wire.TYPE_THROUGHPUT_DATA }
val bytes = received.sumOf { it.sizeBytes.toLong() }
val spanNs = if (received.size >= 2) {
received.maxOf { it.tRxNs } - received.minOf { it.tRxNs }
} else {
0L
}
// Measured over the arrival span rather than our listening window, which includes the
// request round trip and the trailing margin and would understate the rate.
val receivedKbps = if (spanNs > 0) (bytes * 8 * 1_000_000 / spanNs).toInt() else 0
val sender = senderReport(control, credential, sessionId)
val sentPackets = sender?.packets ?: 0
val lossPct = if (sentPackets > 0) {
round2((sentPackets - received.size).coerceAtLeast(0) * 100.0 / sentPackets)
} else {
null
}
// Only a run the *clock* ended measured the network. One stopped by our own byte budget
// or rate ceiling measured this server.
val limitedBy = sender?.limitedBy ?: "unknown"
val networkLimited = limitedBy == "duration" &&
sender != null && receivedKbps > 0 && receivedKbps < sender.kbps * 9 / 10
val metrics = json.encodeToJsonElement(
ThroughputMetrics(
requestedKbps = kbps,
plannedDurationMs = plannedMs,
packetsReceived = received.size,
bytesReceived = bytes,
receivedKbps = receivedKbps,
senderPackets = sender?.packets,
senderBytes = sender?.bytes,
senderKbps = sender?.kbps,
lossPct = lossPct,
limitedBy = limitedBy,
measuresNetwork = networkLimited,
),
) as JsonObject
val findings = ArrayList<Finding>()
when {
sender == null -> Unit // no sender report: nothing can be concluded, so nothing is
received.isEmpty() -> findings.add(
finding(
FindingRegistry.THROUGHPUT_NO_DELIVERY, testId,
"No throughput traffic arrived",
"The server sent ${sender.packets} packets and none arrived. This is a " +
"connectivity fault rather than a slow link.",
),
)
networkLimited -> findings.add(
finding(
FindingRegistry.THROUGHPUT_BELOW_OFFERED, testId,
"Downstream throughput ${receivedKbps / 1000} Mbit/s, below the " +
"${sender.kbps / 1000} Mbit/s offered",
"The server sent at ${sender.kbps / 1000} Mbit/s for the full run and " +
"${receivedKbps / 1000} Mbit/s arrived" +
(lossPct?.let { ", losing $it % of packets" } ?: "") +
". The path could not carry what was offered.",
),
)
}
return Test(
id = testId, type = TestType.PERF_THROUGHPUT_UDP, sessionRef = sessionRef, tier = Tier.APP,
startedMonoNs = started, endedMonoNs = ids.monoNs(),
status = if (received.isEmpty()) TestStatus.FAILED else TestStatus.OK,
metrics = metrics,
) to findings
}
/**
* Upstream throughput: the client sends, the server counts.
*
* The mirror image of the downstream case, and it needs no grant — the client is generating
* its own traffic, so there is no amplification to gate. What it does need is the server's
* count: only the far end knows how much arrived, and without that number a sender can
* measure how fast it can *transmit*, which is not the same question and is usually just the
* speed of the local NIC.
*/
fun runUpstream(
credential: String,
sessionId: String,
control: ControlClient,
probe: ProbeSession,
sessionRef: String,
durationS: Int = 5,
kbps: Int = 20_000,
sizeBytes: Int = 1200,
): Pair<Test, List<Finding>> {
val testId = ids.uuid()
val started = ids.monoNs()
// Zeroes the server's counter so this run measures itself rather than inheriting the
// packets of an earlier one on the same session.
val reply = runCatching {
control.action(credential, sessionId, """{"action":"throughput","direction":"up"}""")
}
if (reply.isFailure) {
return Test(
id = testId, type = TestType.PERF_THROUGHPUT_UDP, sessionRef = sessionRef, tier = Tier.APP,
startedMonoNs = started, endedMonoNs = ids.monoNs(),
status = TestStatus.UNSUPPORTED,
error = TestError("action_refused", reply.exceptionOrNull()?.message ?: "refused"),
) to emptyList()
}
val sent = probe.sendThroughput(durationS * 1000L, kbps, sizeBytes)
// A moment for the tail of the run to arrive; counting still-in-flight packets as lost
// would inflate the loss figure by whatever the path's delay happens to be.
Thread.sleep(500)
val seen = upstreamCount(control, credential, sessionId)
val lossPct = if (sent.packets > 0 && seen != null) {
round2((sent.packets - seen.packets).coerceAtLeast(0) * 100.0 / sent.packets)
} else {
null
}
// The receiver's rate is the measurement. The sender's is what we managed to emit, which
// is a property of this phone and its radio, not of the network.
val achievedKbps = seen?.kbps ?: 0
val metrics = json.encodeToJsonElement(
UpstreamThroughputMetrics(
requestedKbps = kbps,
sentPackets = sent.packets,
sentBytes = sent.bytes,
sentKbps = sent.kbps,
receivedPackets = seen?.packets,
receivedBytes = seen?.bytes,
receivedKbps = achievedKbps,
lossPct = lossPct,
// Same honesty rule as downstream: if what arrived matches what we offered, the
// path was never the constraint and this number says nothing about it.
measuresNetwork = seen != null && achievedKbps > 0 && achievedKbps < sent.kbps * 9 / 10,
),
) as JsonObject
val findings = ArrayList<Finding>()
if (seen != null && seen.packets == 0 && sent.packets > 0) {
findings.add(
finding(
FindingRegistry.THROUGHPUT_NO_DELIVERY, testId,
"No upstream traffic reached the server",
"This device sent ${sent.packets} packets and the server received none. " +
"That is a connectivity fault on the outbound path rather than a slow link.",
),
)
} else if (lossPct != null && lossPct >= 2.0) {
findings.add(
finding(
FindingRegistry.THROUGHPUT_BELOW_OFFERED, testId,
"Upstream loss of $lossPct % at ${sent.kbps / 1000} Mbit/s",
"The server received ${seen?.packets} of the ${sent.packets} packets this " +
"device sent. The outbound path could not carry what was offered.",
),
)
}
return Test(
id = testId, type = TestType.PERF_THROUGHPUT_UDP, sessionRef = sessionRef, tier = Tier.APP,
startedMonoNs = started, endedMonoNs = ids.monoNs(),
status = if (seen == null || seen.packets == 0) TestStatus.FAILED else TestStatus.OK,
metrics = metrics,
) to findings
}
private data class UpstreamCount(val packets: Int, val bytes: Long, val kbps: Int)
/** The server's tally for this session's upstream run. */
private fun upstreamCount(
control: ControlClient, credential: String, sessionId: String,
): UpstreamCount? = runCatching {
val o = Json.parseToJsonElement(control.observations(credential, sessionId))
.jsonObject["throughput_up"]?.jsonObject ?: return null
UpstreamCount(
packets = o["packets"]?.jsonPrimitive?.content?.toIntOrNull() ?: 0,
bytes = o["bytes"]?.jsonPrimitive?.content?.toLongOrNull() ?: 0,
kbps = o["kbps"]?.jsonPrimitive?.content?.toIntOrNull() ?: 0,
)
}.getOrNull()
private data class SenderReport(
val packets: Int, val bytes: Long, val kbps: Int, val limitedBy: String,
)
/** The server's own account of the run, from the observations API. */
private fun senderReport(
control: ControlClient, credential: String, sessionId: String,
): SenderReport? = runCatching {
val arr = Json.parseToJsonElement(control.observations(credential, sessionId))
.jsonObject["throughput"]?.jsonArray ?: return null
val last = arr.lastOrNull()?.jsonObject ?: return null
SenderReport(
packets = last["packets"]?.jsonPrimitive?.content?.toIntOrNull() ?: 0,
bytes = last["bytes"]?.jsonPrimitive?.content?.toLongOrNull() ?: 0,
kbps = last["kbps"]?.jsonPrimitive?.content?.toIntOrNull() ?: 0,
limitedBy = last["limited_by"]?.jsonPrimitive?.content ?: "unknown",
)
}.getOrNull()
private fun parseInt(body: String?, key: String): Int? =
body?.let { Regex("\"$key\"\\s*:\\s*(-?\\d+)").find(it)?.groupValues?.get(1)?.toIntOrNull() }
/**
* Builds a finding from a registry entry, which supplies the code, category and severity.
*
* Taking a [FindingSpec] rather than three loose values is the point: a typo becomes a
* compile error, and two call sites cannot disagree about which category a finding belongs
* to - a disagreement that would split one fault across two verdict lights.
*/
private fun finding(spec: FindingSpec, testId: String, title: String, desc: String) =
Finding(
id = ids.uuid(), code = spec.code, category = spec.category, severity = spec.severity,
confidence = Confidence.HIGH,
title = title, description = desc, evidenceRefs = listOf(EvidenceRef(testId)),
)
private fun round2(v: Double) = Math.round(v * 100.0) / 100.0
}
/** Metrics for perf.throughput_udp in the upstream direction. */
@Serializable
data class UpstreamThroughputMetrics(
val direction: String = "up",
@SerialName("requested_kbps") val requestedKbps: Int,
@SerialName("sent_packets") val sentPackets: Int,
@SerialName("sent_bytes") val sentBytes: Long,
/** What this device managed to emit — a property of the phone and its radio, not the path. */
@SerialName("sent_kbps") val sentKbps: Int,
@SerialName("received_packets") val receivedPackets: Int? = null,
@SerialName("received_bytes") val receivedBytes: Long? = null,
/** What arrived, measured by the only party that can measure it. This is the result. */
@SerialName("received_kbps") val receivedKbps: Int,
@SerialName("loss_pct") val lossPct: Double? = null,
@SerialName("measures_network") val measuresNetwork: Boolean,
)
/** Metrics for perf.throughput_udp. */
@Serializable
data class ThroughputMetrics(
val direction: String = "down",
@SerialName("requested_kbps") val requestedKbps: Int,
@SerialName("planned_duration_ms") val plannedDurationMs: Int,
@SerialName("packets_received") val packetsReceived: Int,
@SerialName("bytes_received") val bytesReceived: Long,
@SerialName("received_kbps") val receivedKbps: Int,
@SerialName("sender_packets") val senderPackets: Int? = null,
@SerialName("sender_bytes") val senderBytes: Long? = null,
@SerialName("sender_kbps") val senderKbps: Int? = null,
/** Against the sender's count, so a server-side limit is never counted as network loss. */
@SerialName("loss_pct") val lossPct: Double? = null,
/** What ended the run: duration | budget | rate | send_error | unknown. */
@SerialName("limited_by") val limitedBy: String,
/**
* Whether this number says anything about the network. False when the sender's own ceiling
* was the binding constraint — in which case the rate is a property of the test, not the path.
*/
@SerialName("measures_network") val measuresNetwork: Boolean,
)
@@ -0,0 +1,159 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.engine
import app.echo_lot.measurement.*
import app.echo_lot.protocol.ProbeSession
import kotlinx.serialization.SerialName
import kotlinx.serialization.Serializable
import kotlinx.serialization.json.Json
import kotlinx.serialization.json.JsonObject
import kotlinx.serialization.json.encodeToJsonElement
/**
* train.udp_updown the client sends a paced train (types 0x03), then asks the server what
* arrived (0x04 0x05) and lines both views up per sequence number.
*
* This is the measurement a round trip cannot make: an echo run only says "lost somewhere", the
* train's two ledgers say lost on the way OUT, specifically, because the server's report names
* exactly which sequence numbers reached it. The downstream direction has its own test
* (train.udp_downstream) under a grant; this one needs none, since the client generates all the
* traffic itself.
*
* The evidence is the schema's columnar TrainEvidence: one index per sent packet, with the
* server-side columns null where a packet never arrived. Server timestamps are on the server's
* own clock only differences within that clock mean anything unless time.server_offset maps
* them (two-clock rule).
*/
class UpstreamTrainMeasurement(private val ids: IdSource) {
private val json = Json { encodeDefaults = true; explicitNulls = true }
fun run(
probe: ProbeSession,
sessionRef: String,
count: Int = 200,
sizeBytes: Int = 200,
interPacketMs: Long = 5,
): Pair<Test, List<Finding>> {
val testId = ids.uuid()
val started = ids.monoNs()
// The id only needs to be unique within this session; a clash across sessions is
// meaningless because trains are buffered per session on the server.
val trainId = (System.nanoTime() and 0x7FFFFFFF).toInt()
val sent = probe.sendTrain(trainId, count, sizeBytes, interPacketMs)
// Let the tail arrive before asking for the ledger; packets still in flight when the
// report is cut would read as upstream loss.
Thread.sleep(300)
val report = probe.trainReport(trainId)
if (report == null) {
return Test(
id = testId, type = TestType.TRAIN_UDP_UPDOWN, sessionRef = sessionRef,
tier = Tier.APP, startedMonoNs = started, endedMonoNs = ids.monoNs(),
status = TestStatus.FAILED,
// Honest ambiguity: an old server drops 0x04 silently, and a lost report looks
// identical from here. Neither says anything about the train itself.
error = TestError(
"no_report",
"no train report arrived — the report was lost, or the server predates trains",
),
) to emptyList()
}
val bySeq = report.rows.associateBy { it.seq }
fun col255(v: Int): Int? = v.takeIf { it != 255 } // 255 = "not observed" on the wire
val evidence = TrainEvidence(
epochMonoNs = started,
seq = sent.map { it.seq },
tTxNs = sent.map { it.tTxNs },
tSrvRxNs = sent.map { bySeq[it.seq]?.tRxNs },
tRxNs = sent.map { null }, // upstream only: nothing comes back per packet
sizeBytes = sent.map { it.sizeBytes },
ttlSeenByServer = sent.map { bySeq[it.seq]?.let { r -> col255(r.ttl) } },
dscpSeenByServer = sent.map { bySeq[it.seq]?.let { r -> col255(r.dscp) } },
ecnSeenByServer = sent.map { bySeq[it.seq]?.let { r -> col255(r.ecn) } },
evidenceTruncated = report.truncated,
).toEvidence()
// Loss against the server's total count, not its row list: rows past the server's buffer
// cap are counted but not kept, and treating them as lost would invent loss exactly on
// the biggest trains.
val lossPct = if (sent.isEmpty()) 0.0 else {
(sent.size - report.received).coerceAtLeast(0) * 100.0 / sent.size
}
val metrics = json.encodeToJsonElement(
UpstreamTrainMetrics(
sent = sent.size,
receivedByServer = report.received,
lossPct = round1(lossPct),
reportPartsExpected = report.partsExpected,
reportPartsReceived = report.partsReceived,
truncated = report.truncated,
),
) as JsonObject
val findings = ArrayList<Finding>()
if (sent.isNotEmpty() && report.received == 0) {
findings.add(
Finding(
id = ids.uuid(),
code = FindingRegistry.UDP_UNREACHABLE_UPSTREAM.code,
category = FindingRegistry.UDP_UNREACHABLE_UPSTREAM.category,
severity = FindingRegistry.UDP_UNREACHABLE_UPSTREAM.severity,
confidence = Confidence.HIGH,
title = "The server received none of ${sent.size} upstream packets",
description = "Every train packet vanished on the way out, while the " +
"report request's reply made it back — the outbound path drops this " +
"traffic, the return path works.",
evidenceRefs = listOf(EvidenceRef(testId)),
),
)
} else if (lossPct >= 2.0) {
findings.add(
Finding(
id = ids.uuid(),
code = FindingRegistry.LOSS_UPSTREAM.code,
category = FindingRegistry.LOSS_UPSTREAM.category,
severity = FindingRegistry.LOSS_UPSTREAM.severity,
confidence = Confidence.HIGH,
title = "Upstream loss of ${round1(lossPct)} %",
description = "The server received ${report.received} of the ${sent.size} " +
"packets this device sent, and its per-sequence ledger names the " +
"missing ones. This is outbound loss specifically; the return path " +
"delivered the report.",
evidenceRefs = listOf(EvidenceRef(testId)),
),
)
}
val status = when {
report.received == 0 && sent.isNotEmpty() -> TestStatus.FAILED
report.partsReceived < report.partsExpected -> TestStatus.PARTIAL
else -> TestStatus.OK
}
return Test(
id = testId, type = TestType.TRAIN_UDP_UPDOWN, sessionRef = sessionRef, tier = Tier.APP,
startedMonoNs = started, endedMonoNs = ids.monoNs(),
status = status, evidence = evidence, metrics = metrics,
) to findings
}
private fun round1(v: Double) = Math.round(v * 10.0) / 10.0
}
/** Metrics for train.udp_updown. */
@Serializable
data class UpstreamTrainMetrics(
val sent: Int,
/** The server's total count — includes packets past its row buffer (counted, not listed). */
@SerialName("received_by_server") val receivedByServer: Int,
@SerialName("loss_pct") val lossPct: Double,
@SerialName("report_parts_expected") val reportPartsExpected: Int,
@SerialName("report_parts_received") val reportPartsReceived: Int,
/** The server's row buffer overflowed: rows are a sample, the count is still complete. */
val truncated: Boolean,
)
@@ -0,0 +1,155 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.engine
import app.echo_lot.engine.Directional.Sample
import app.echo_lot.engine.Directional.ServerSighting
import kotlin.test.Test
import kotlin.test.assertEquals
import kotlin.test.assertFalse
import kotlin.test.assertNotNull
import kotlin.test.assertNull
import kotlin.test.assertTrue
/**
* The arithmetic that turns "3 % loss somewhere" into "3 % loss upstream". Getting a denominator
* wrong here does not crash anything it produces a plausible number pointing at the wrong half
* of the network, which is worse than no number at all. Hence a test per claim.
*/
class DirectionalTest {
/** A clean train: every packet sent, seen and answered. Server clock offset by a constant. */
private fun clean(n: Int, offsetNs: Long = 5_000_000_000L): Pair<List<Sample>, List<ServerSighting>> {
val sent = (1..n).map { Sample(it, tTxNs = it * 10_000_000L, tRxNs = it * 10_000_000L + 4_000_000L) }
val seen = (1..n).map {
ServerSighting(it, tRxNs = offsetNs + it * 10_000_000L + 2_000_000L,
tTxNs = offsetNs + it * 10_000_000L + 2_100_000L)
}
return sent to seen
}
@Test
fun aCleanTrainReportsNoLossInEitherDirection() {
val (sent, seen) = clean(10)
val m = Directional.analyse(sent, seen)
assertEquals(10, m.sent)
assertEquals(10, m.seenByServer)
assertEquals(10, m.repliesReceived)
assertEquals(0.0, m.lossUpstreamPct)
assertEquals(0.0, m.lossDownstreamPct)
assertFalse(m.noneReachedServer)
}
// The whole point: a packet the server never saw was lost on the way there.
@Test
fun packetsTheServerNeverSawAreUpstreamLoss() {
val (sent, seen) = clean(10)
val m = Directional.analyse(sent, seen.filter { it.seq !in setOf(3, 7) })
assertEquals(2, m.lostUpstream)
assertEquals(0, m.lostDownstream)
assertEquals(20.0, m.lossUpstreamPct)
assertEquals(0.0, m.lossDownstreamPct, "a packet that never arrived cannot be lost coming back")
}
@Test
fun repliesThatNeverArrivedAreDownstreamLoss() {
val (sent, seen) = clean(10)
val withHoles = sent.map { if (it.seq in setOf(2, 5)) it.copy(tRxNs = null) else it }
val m = Directional.analyse(withHoles, seen)
assertEquals(0, m.lostUpstream)
assertEquals(2, m.lostDownstream)
assertEquals(20.0, m.lossDownstreamPct)
}
// Downstream loss is measured against what actually reached the server. Using "sent" as the
// denominator would count every upstream loss a second time and overstate the return path.
@Test
fun downstreamLossIsRelativeToWhatReachedTheServer() {
val (sent, seen) = clean(10)
// 5 lost on the way there; of the 5 that arrived, 1 reply is lost coming back.
val seenPartial = seen.filter { it.seq > 5 }
val withHole = sent.map {
when {
it.seq <= 5 -> it.copy(tRxNs = null) // never got there, so never came back
it.seq == 6 -> it.copy(tRxNs = null) // arrived, reply lost
else -> it
}
}
val m = Directional.analyse(withHole, seenPartial)
assertEquals(5, m.lostUpstream)
assertEquals(50.0, m.lossUpstreamPct)
assertEquals(1, m.lostDownstream)
assertEquals(20.0, m.lossDownstreamPct, "1 of the 5 that arrived, not 1 of 10")
}
@Test
fun aServerThatSawNothingIsCalledOutSeparately() {
val (sent, _) = clean(6)
val m = Directional.analyse(sent.map { it.copy(tRxNs = null) }, emptyList())
assertTrue(m.noneReachedServer)
assertEquals(100.0, m.lossUpstreamPct)
assertEquals(0.0, m.lossDownstreamPct, "with nothing arriving there is no return path to blame")
}
// Jitter is legitimate without synchronised clocks because the offset cancels when successive
// one-way samples are differenced. This pins that: a huge constant offset must not show up.
@Test
fun jitterIsUnaffectedByTheClockOffsetBetweenTheTwoMachines() {
val (sent, near) = clean(10, offsetNs = 0)
val (_, far) = clean(10, offsetNs = 9_999_999_999L)
val a = Directional.analyse(sent, near)
val b = Directional.analyse(sent, far)
assertEquals(a.jitterUpstreamMs, b.jitterUpstreamMs,
"a constant clock offset must cancel when consecutive samples are differenced")
assertEquals(0.0, assertNotNull(a.jitterUpstreamMs), "an evenly spaced train has no jitter")
}
@Test
fun jitterReflectsUnevenArrival() {
val sent = listOf(
Sample(1, 0, 10_000_000),
Sample(2, 10_000_000, 20_000_000),
Sample(3, 20_000_000, 30_000_000),
)
// Server receive times drift: +2ms, +7ms, +3ms relative to send.
val seen = listOf(
ServerSighting(1, 2_000_000, 2_100_000),
ServerSighting(2, 17_000_000, 17_100_000),
ServerSighting(3, 23_000_000, 23_100_000),
)
val m = Directional.analyse(sent, seen)
// one-way samples: 2ms, 7ms, 3ms → |7-2| and |3-7| → mean 4.5ms
assertEquals(4.5, assertNotNull(m.jitterUpstreamMs))
}
// "No jitter" and "not enough data to say" are different claims, and only one is true here.
@Test
fun tooFewSamplesReportsNoJitterRatherThanZero() {
val m = Directional.analyse(
listOf(Sample(1, 0, 10_000_000)),
listOf(ServerSighting(1, 2_000_000, 2_100_000)),
)
assertNull(m.jitterUpstreamMs)
assertNull(m.jitterDownstreamMs)
}
// A server record for a sequence we never sent is not evidence about this train; folding it
// in would yield loss percentages outside 0100.
@Test
fun strayServerRecordsAreIgnored() {
val (sent, seen) = clean(5)
val m = Directional.analyse(sent, seen + ServerSighting(99, 1, 2) + ServerSighting(100, 3, 4))
assertEquals(5, m.seenByServer)
assertEquals(0.0, m.lossUpstreamPct)
assertTrue(m.lossDownstreamPct in 0.0..100.0)
}
@Test
fun anEmptyTrainDoesNotDivideByZero() {
val m = Directional.analyse(emptyList(), emptyList())
assertEquals(0.0, m.lossUpstreamPct)
assertEquals(0.0, m.lossDownstreamPct)
assertFalse(m.noneReachedServer, "nothing sent is not the same as nothing arriving")
}
}
@@ -0,0 +1,71 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.engine
import app.echo_lot.protocol.Compat
import app.echo_lot.protocol.ControlClient
import app.echo_lot.protocol.VersionRefused
import kotlin.test.Test
import kotlin.test.assertEquals
import kotlin.test.assertNotNull
import kotlin.test.assertTrue
import kotlin.test.fail
/**
* Checks the version gate against a LIVE server the half that unit tests cannot reach, because
* the whole point is that two independently-built artifacts agree. Self-skips without
* ECHOLOT_LIVE_*.
*/
class LiveCompatTest {
private val url = System.getenv("ECHOLOT_LIVE_URL")
private val pin = System.getenv("ECHOLOT_LIVE_PIN")
private val cred = System.getenv("ECHOLOT_LIVE_CRED")
private fun clientAs(version: String) = ControlClient(url!!, setOf(pin!!), version)
@Test
fun theServerAdvertisesAndEnforcesItsWindow() {
if (url == null || pin == null || cred == null) {
println("LiveCompatTest skipped (no ECHOLOT_LIVE_* env)"); return
}
// The profile must state the window — without it the app cannot pre-empt a refusal.
val profile = clientAs("0.2.0").profile(cred)
println("server ${profile.serverVersion} protocol=${profile.compat.protocolVersion} " +
"accepts app [${profile.compat.appMin}, ${profile.compat.appMax})")
assertTrue(profile.compat.protocolVersion.isNotBlank(), "profile omits protocol_version")
assertTrue(profile.compat.appMin.isNotBlank(), "profile omits app_min")
// This build must be inside it, or every other live test here is meaningless.
val verdict = Compat.check(profile, "0.2.0")
assertEquals(Compat.Verdict.OK, verdict.verdict, verdict.message ?: "")
// The profile stays reachable for a version the server would otherwise refuse: that is
// how a refused client discovers what it needs.
val ancient = clientAs("0.1.0")
val stillReadable = ancient.profile(cred)
assertEquals(profile.serverVersion, stillReadable.serverVersion,
"the profile endpoint must never be gated on app version")
// And a gated endpoint refuses it, with a message naming the window.
try {
ancient.createSession(cred, System.getenv("ECHOLOT_LIVE_TARGET") ?: "fmr")
fail("server accepted a session from an out-of-window app")
} catch (e: VersionRefused) {
val msg = assertNotNull(e.message)
println("refused as expected: $msg")
assertTrue(msg.contains("0.1.0"), "refusal should name the offending version: $msg")
assertTrue(msg.contains(profile.compat.appMin), "refusal should name the window: $msg")
}
// Too new is refused the same way — the window is a range, not a floor.
try {
clientAs("99.0.0").createSession(cred, System.getenv("ECHOLOT_LIVE_TARGET") ?: "fmr")
fail("server accepted a session from an app above its window")
} catch (e: VersionRefused) {
println("too-new refused as expected: ${e.message}")
}
}
}
@@ -0,0 +1,81 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.engine
import app.echo_lot.measurement.TestStatus
import app.echo_lot.measurement.TestType
import app.echo_lot.protocol.ControlClient
import app.echo_lot.protocol.ProbeSession
import kotlin.test.Test as JTest
import kotlin.test.assertEquals
import kotlin.test.assertNotNull
import kotlin.test.assertTrue
/**
* Runs DownstreamMeasurement against a LIVE server and checks the *documents* it produces, not
* just that packets moved: the tests must carry recomputable metrics and land on the right test
* types, because that is what an archived run is read back as. Self-skips without ECHOLOT_LIVE_*.
*/
class LiveDownstreamTest {
private val url = System.getenv("ECHOLOT_LIVE_URL")
private val pin = System.getenv("ECHOLOT_LIVE_PIN")
private val cred = System.getenv("ECHOLOT_LIVE_CRED")
private val udp = System.getenv("ECHOLOT_LIVE_UDP")
private val target = System.getenv("ECHOLOT_LIVE_TARGET") ?: "fmr"
@JTest
fun producesDownstreamTestsAndFindings() {
if (url == null || pin == null || cred == null || udp == null) {
println("LiveDownstreamTest skipped (no ECHOLOT_LIVE_* env)"); return
}
val control = ControlClient(url, setOf(pin))
val session = control.createSession(cred, target)
val (host, port) = udp.split(":").let { it[0] to it[1].toInt() }
val (tests, findings) = ProbeSession(cred, session, host, port).use { ps ->
ps.echo() // prime: the grant binds to the source the server has actually observed
DownstreamMeasurement(SystemIdSource())
.run(cred, session.sessionId, control, ps, sessionRef = "sess-1")
}
control.deleteSession(cred, session.sessionId)
for (t in tests) println("${t.type} status=${t.status} metrics=${t.metrics}")
for (f in findings) println("finding ${f.code} [${f.severity}] ${f.title}")
// Assert on what is present, not on how many: adding a measurement should not be a
// test edit. (It was, once — hence the note.)
assertTrue(tests.size >= 3, "expected at least the three downstream tests, got ${tests.size}")
val byType = tests.associateBy { it.type }
val pmtud = assertNotNull(byType[TestType.MTU_PMTUD_DOWN], "no mtu.pmtud_down test")
assertTrue(pmtud.status == TestStatus.OK || pmtud.status == TestStatus.PARTIAL,
"DF probe did not deliver anything: ${pmtud.status}")
val pathMtu = pmtud.metrics?.get("path_mtu_bytes")?.toString()?.toIntOrNull()
assertNotNull(pathMtu, "pmtud_down must report a path MTU")
assertTrue(pathMtu in 576..9000, "implausible downstream path MTU: $pathMtu")
println("downstream path MTU = $pathMtu bytes")
val frag = assertNotNull(byType[TestType.MTU_FRAG_DELIVERY], "no mtu.frag_delivery test")
assertNotNull(frag.metrics?.get("largest_delivered_bytes"))
// Fragment ordering runs only when fragments arrive at all, and only against a server
// that can craft them — so it is checked when present rather than required.
byType[TestType.MTU_FRAG_ORDERING]?.let { fo ->
val m = fo.metrics?.toString() ?: ""
println("fragment ordering: ${fo.status} $m")
if (fo.status != TestStatus.UNSUPPORTED) {
assertTrue(m.contains("in_order"), "no per-ordering result: $m")
assertTrue(m.contains("reversed"), "reversed ordering was never attempted: $m")
}
}
val train = assertNotNull(byType[TestType.TRAIN_UDP_DOWNSTREAM], "no downstream train")
assertNotNull(train.evidence, "a train without columnar evidence is not recomputable")
val received = train.metrics?.get("received")?.toString()?.toIntOrNull() ?: 0
assertTrue(received > 0, "no downstream train packets arrived")
println("downstream train: $received received, loss=${train.metrics?.get("loss_pct")}, " +
"reordered=${train.metrics?.get("reordered_packets")}")
}
}
@@ -0,0 +1,63 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.engine
import app.echo_lot.protocol.EnrollmentLink
import kotlin.test.Test
import kotlin.test.assertEquals
import kotlin.test.assertNotNull
import kotlin.test.assertTrue
import kotlin.test.fail
/**
* Enrolls against a LIVE server using the link the server itself minted (probe-protocol.md §2.1).
*
* This is the test that matters for enrollment, because the failure mode it guards against is a
* *disagreement* between two programs: the Go side assembles the link, the Kotlin side takes it
* apart, and if they differ by one percent-encoding the pin is wrong by one character which
* does not fail loudly, it fails as an inscrutable TLS error days later. A unit test on either
* side alone cannot see that.
*
* Needs ECHOLOT_ENROLL_URI (minted over SSH by scripts/test-fmr.sh); self-skips without it.
*/
class LiveEnrollmentTest {
private val enrollUri = System.getenv("ECHOLOT_ENROLL_URI")
@Test
fun enrollsFromTheServersOwnLink() {
if (enrollUri.isNullOrBlank()) {
println("LiveEnrollmentTest skipped (no ECHOLOT_ENROLL_URI)"); return
}
println("link: ${enrollUri.take(60)}")
val link = assertNotNull(
EnrollmentLink.parse(enrollUri),
"the client could not parse a link the server produced — the two sides disagree",
)
println("parsed: url=${link.controlUrl} pin=${link.pin.take(12)}… token=${link.token.take(8)}")
// Redeeming applies the pin to the very request that spends the token, so a wrong pin
// fails here at the handshake rather than after the token is gone.
val enrolled = link.redeem(deviceName = "live-test", appVersion = "0.2.0")
assertTrue(enrolled.credential.isNotBlank(), "no credential came back")
assertTrue(enrolled.deviceId.isNotBlank(), "no device id came back")
println("enrolled: device=${enrolled.deviceId} server=${enrolled.profile.name} " +
"${enrolled.profile.serverVersion}")
// The credential must actually work, and the pin from the link must be the one that
// verifies the server — that is the whole claim the link is making.
assertEquals(link.controlUrl, enrolled.controlUrl)
assertTrue(enrolled.profile.capabilities.contains("udp-probe"),
"profile fetched with the new credential looks wrong: ${enrolled.profile.capabilities}")
// Single-use: a token that still works after redemption is a token an attacker can reuse.
try {
link.redeem(deviceName = "should-not-happen", appVersion = "0.2.0")
fail("the enrollment token was accepted twice — it must be single-use")
} catch (t: Throwable) {
println("second redemption correctly refused: ${t.message?.take(120)}")
}
}
}
@@ -0,0 +1,89 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.engine
import app.echo_lot.protocol.ControlClient
import app.echo_lot.protocol.ProbeSession
import app.echo_lot.protocol.Wire
import kotlin.test.Test
import kotlin.test.assertTrue
/**
* Exercises the server's §5 granted sends against a LIVE server: downtrain (downstream loss /
* ordering) and big_send (downstream MTU). Self-skips without ECHOLOT_LIVE_*.
*
* This is the direction a client cannot measure alone only the far end can push large or
* numerous packets toward it so it is also the direction that needs the anti-amplification
* grant, and this test is the proof that the grant path works end to end.
*/
class LiveGrantedTest {
private val url = System.getenv("ECHOLOT_LIVE_URL")
private val pin = System.getenv("ECHOLOT_LIVE_PIN")
private val cred = System.getenv("ECHOLOT_LIVE_CRED")
private val udp = System.getenv("ECHOLOT_LIVE_UDP")
private val target = System.getenv("ECHOLOT_LIVE_TARGET") ?: "fmr"
@Test
fun downstreamTrainAndBigSend() {
if (url == null || pin == null || cred == null || udp == null) {
println("LiveGrantedTest skipped (no ECHOLOT_LIVE_* env)"); return
}
val control = ControlClient(url, setOf(pin))
val profile = control.profile(cred)
println("capabilities: ${profile.capabilities}")
val session = control.createSession(cred, target)
val (host, port) = udp.split(":").let { it[0] to it[1].toInt() }
ProbeSession(cred, session, host, port).use { ps ->
// The grant is bound to the OBSERVED data-plane source, so we must be seen first.
val echo = ps.echo()
println("primed with echo rtt=${echo?.rttMs}")
// --- downtrain: 50 packets of 300 bytes, 5ms apart ---
val dtResp = control.action(
cred, session.sessionId,
"""{"action":"downtrain","count":50,"size_bytes":300,"interval_us":5000}""",
)
println("downtrain accepted: ${dtResp.take(160)}")
val down = ps.collectGranted(windowMs = 4000)
.filter { it.type == Wire.TYPE_DOWNTRAIN_DATA }
val seqs = down.map { it.seq }.toSet()
println("downtrain received ${down.size}/50 packets, distinct seqs=${seqs.size}, " +
"sizes=${down.map { it.sizeBytes }.distinct()}")
assertTrue(down.isNotEmpty(), "no DOWNTRAIN_DATA arrived — granted send path is broken")
// --- big_send with DF: the largest size that arrives is the downstream path MTU ---
val sizes = listOf(600, 1200, 1400, 1472, 1500, 2000, 4000)
val dfResp = control.action(
cred, session.sessionId,
"""{"action":"big_send","df":true,"sizes_bytes":${sizes}}""",
)
println("big_send(df) accepted: ${dfResp.take(200)}")
val dfArrived = ps.collectGranted(windowMs = 4000)
.filter { it.type == Wire.TYPE_BIG_SEND }.map { it.sizeBytes }.sorted()
println("big_send(df) arrived: $dfArrived")
assertTrue(dfArrived.isNotEmpty(), "no unfragmented BIG_SEND packets arrived")
val pathMtu = dfArrived.max()
// --- and without DF, to see whether fragments get through above that ---
val fragResp = control.action(
cred, session.sessionId,
"""{"action":"big_send","df":false,"sizes_bytes":${sizes}}""",
)
println("big_send(frag) accepted: ${fragResp.take(200)}")
val fragArrived = ps.collectGranted(windowMs = 4000)
.filter { it.type == Wire.TYPE_BIG_SEND }.map { it.sizeBytes }.sorted()
println("big_send(frag) arrived: $fragArrived")
// The distinction the DF flag exists for: fragmented delivery may exceed the
// unfragmented path MTU, and reporting the former as the latter would be a lie.
println("downstream path MTU (payload bytes) = $pathMtu; " +
"largest fragmented delivery = ${fragArrived.maxOrNull()}")
assertTrue((fragArrived.maxOrNull() ?: 0) >= pathMtu,
"fragmented delivery should reach at least as far as unfragmented")
}
control.deleteSession(cred, session.sessionId)
}
}
@@ -7,6 +7,7 @@ import app.echo_lot.measurement.*
import kotlinx.serialization.json.Json import kotlinx.serialization.json.Json
import kotlin.test.Test import kotlin.test.Test
import kotlin.test.assertEquals import kotlin.test.assertEquals
import kotlin.test.assertNotNull
import kotlin.test.assertTrue import kotlin.test.assertTrue
/** /**
@@ -47,8 +48,11 @@ class LiveMeasurementTest {
assertEquals(1, doc.serverSessions.size) assertEquals(1, doc.serverSessions.size)
assertTrue(doc.serverSessions[0].capabilities.contains("udp-probe")) assertTrue(doc.serverSessions[0].capabilities.contains("udp-probe"))
val test = doc.tests.single() // A full run is the echo train plus the three downstream tests; assert on the one this
assertEquals(TestType.TRAIN_UDP_UPDOWN, test.type) // test is about rather than on the count, so adding a measurement is not a test edit.
for (t in doc.tests) println(" ${t.type}${t.status}")
for (f in doc.findings) println(" finding ${f.code} [${f.severity}] ${f.title}")
val test = doc.tests.first { it.type == TestType.TRAIN_UDP_UPDOWN }
assertTrue(test.status == TestStatus.OK || test.status == TestStatus.PARTIAL, assertTrue(test.status == TestStatus.OK || test.status == TestStatus.PARTIAL,
"expected replies from live server, got ${test.status}") "expected replies from live server, got ${test.status}")
@@ -56,6 +60,18 @@ class LiveMeasurementTest {
println("metrics: $metrics") println("metrics: $metrics")
assertTrue(metrics.toString().contains("rtt_ms_avg")) assertTrue(metrics.toString().contains("rtt_ms_avg"))
// The directional split is the point of asking the server what it saw: without it a
// lossy path is reported as "loss" with no direction, which sends an engineer looking
// in both at once. Correlation is by wire sequence number, so a mismatch here means the
// two sides disagree about which packet is which.
val m = metrics.toString()
assertTrue(m.contains("seen_by_server"), "no directional split in the metrics: $m")
val seen = Regex(""""seen_by_server":(\d+)""").find(m)?.groupValues?.get(1)?.toInt()
assertNotNull(seen, "seen_by_server missing")
assertEquals(20, seen, "the server should have seen every probe on a healthy path")
assertTrue(m.contains("jitter_upstream_ms"), "no per-direction jitter: $m")
println("directional: $m")
assertTrue(doc.summary != null) assertTrue(doc.summary != null)
// A healthy local->fmr path should be green (no loss, no rebinding) or yellow. // A healthy local->fmr path should be green (no loss, no rebinding) or yellow.
println("summary: ${doc.summary}") println("summary: ${doc.summary}")
@@ -0,0 +1,105 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.engine
import app.echo_lot.measurement.TestStatus
import app.echo_lot.protocol.ControlClient
import app.echo_lot.protocol.ProbeSession
import kotlin.test.Test
import kotlin.test.assertEquals
import kotlin.test.assertNotNull
import kotlin.test.assertTrue
/**
* Downstream throughput against a LIVE server. Self-skips without ECHOLOT_LIVE_*.
*
* The assertions are about *honesty* rather than speed: a rate is only a measurement if the run
* was ended by the clock and the sender's own count backs it up. A test that just asserted "some
* Mbps arrived" would pass equally well against a broken implementation.
*/
class LiveThroughputTest {
private val url = System.getenv("ECHOLOT_LIVE_URL")
private val pin = System.getenv("ECHOLOT_LIVE_PIN")
private val cred = System.getenv("ECHOLOT_LIVE_CRED")
private val udp = System.getenv("ECHOLOT_LIVE_UDP")
private val target = System.getenv("ECHOLOT_LIVE_TARGET") ?: "fmr"
@Test
fun measuresDownstreamRateAndSaysWhatLimitedIt() {
if (url == null || pin == null || cred == null || udp == null) {
println("LiveThroughputTest skipped (no ECHOLOT_LIVE_* env)"); return
}
val control = ControlClient(url, setOf(pin), "0.2.0")
val session = control.createSession(cred, target)
val (host, port) = udp.split(":").let { it[0] to it[1].toInt() }
val (test, findings) = ProbeSession(cred, session, host, port).use { ps ->
ps.echo() // prime: the grant binds to the observed source
ThroughputMeasurement(SystemIdSource()).run(
cred, session.sessionId, control, ps, sessionRef = "sess-1",
durationS = 3, kbps = 20_000,
)
}
control.deleteSession(cred, session.sessionId)
val m = assertNotNull(test.metrics).toString()
println("throughput: ${test.status} $m")
for (f in findings) println("finding ${f.code} [${f.severity}] ${f.title}")
assertEquals(TestStatus.OK, test.status, "no throughput traffic arrived: $m")
// The sender's own count must be present — without it, loss cannot be attributed and the
// number is not a measurement.
assertTrue(m.contains("sender_packets"), "no sender report to compare against: $m")
assertTrue(m.contains("limited_by"), "the result must say what ended the run: $m")
val received = Regex(""""received_kbps":(\d+)""").find(m)?.groupValues?.get(1)?.toInt()
assertNotNull(received)
assertTrue(received > 0, "measured 0 kbps: $m")
println("received ${received / 1000} Mbit/s")
// A run this short and this far below the ceiling should end on the clock. Anything else
// means the grant was the constraint, and then the rate says nothing about the path.
assertTrue(m.contains(""""limited_by":"duration""""),
"the run did not end on the clock, so the rate measures the server, not the path: $m")
}
// Upstream is the direction only the far end can measure. The assertion that matters is that
// the server's count is present and plausible against what we sent — a test that only checked
// "we transmitted some Mbps" would pass against a server that counted nothing at all.
@Test
fun measuresUpstreamAgainstTheServersCount() {
if (url == null || pin == null || cred == null || udp == null) {
println("LiveThroughputTest(up) skipped"); return
}
val control = ControlClient(url, setOf(pin), "0.2.0")
val session = control.createSession(cred, target)
val (host, port) = udp.split(":").let { it[0] to it[1].toInt() }
val (test, findings) = ProbeSession(cred, session, host, port).use { ps ->
ps.echo()
ThroughputMeasurement(SystemIdSource()).runUpstream(
cred, session.sessionId, control, ps, sessionRef = "sess-1",
durationS = 3, kbps = 10_000,
)
}
control.deleteSession(cred, session.sessionId)
val m = assertNotNull(test.metrics).toString()
println("upstream: ${test.status} $m")
for (f in findings) println("finding ${f.code} [${f.severity}] ${f.title}")
assertEquals(TestStatus.OK, test.status, "the server counted nothing: $m")
val recv = Regex(""""received_packets":(\d+)""").find(m)?.groupValues?.get(1)?.toInt()
val sent = Regex(""""sent_packets":(\d+)""").find(m)?.groupValues?.get(1)?.toInt()
assertNotNull(recv); assertNotNull(sent)
assertTrue(sent > 100, "barely anything was sent, so the rate means nothing: $m")
assertTrue(recv > 0, "the server received none of $sent packets: $m")
// The counts should be close on a healthy path; wildly different means the two sides are
// counting different things rather than the network losing packets.
assertTrue(recv <= sent, "the server counted MORE than we sent — the counter is not being reset")
println("sent $sent, server saw $recv")
}
}
@@ -0,0 +1,102 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.engine
import app.echo_lot.privacy.Anonymizer
import app.echo_lot.privacy.PrivacyLevel
import app.echo_lot.privacy.Salt
import app.echo_lot.protocol.ControlClient
import app.echo_lot.protocol.UploadRefused
import kotlinx.serialization.json.Json
import kotlinx.serialization.json.JsonObject
import kotlinx.serialization.json.jsonObject
import kotlin.test.Test
import kotlin.test.assertFalse
import kotlin.test.assertTrue
/**
* Drives the upload path against a LIVE server: anonymize, upload, list, fetch back, delete.
*
* The point is not that the HTTP works it is that what comes *back off the server* has been
* stripped. Uploading and then re-reading the stored document is the only check that proves the
* anonymizer ran on the bytes that actually left, rather than on a copy. Self-skips without
* ECHOLOT_LIVE_*.
*/
class LiveUploadTest {
private val url = System.getenv("ECHOLOT_LIVE_URL")
private val pin = System.getenv("ECHOLOT_LIVE_PIN")
private val cred = System.getenv("ECHOLOT_LIVE_CRED")
private val json = Json { prettyPrint = false }
private fun sampleRun(id: String) = """
{
"schema": "echolot/measurement",
"run": {
"id": "$id", "trigger": "manual", "started_at": "2026-08-01T10:00:00Z",
"notes": "kitchen table",
"device": {"manufacturer": "OnePlus", "model": "CPH2747"}
},
"networks": [{
"id": "net-1", "ssid": "Rambossek WLAN", "bssid": "78:9a:18:aa:bb:cc",
"gateway_ip4": "192.168.1.1", "public_ip4": "203.0.113.77",
"ssdp_responders": [{"friendly_name": "Living Room TV"}]
}],
"tests": [{"id": "t1", "type": "train.udp_updown", "status": "ok",
"metrics": {"rtt_ms_avg": 12.4, "loss_pct": 0.0}}],
"findings": [{"id": "f1", "code": "nat.udp_rebinding", "severity": "medium"}],
"summary": {"verdict": "warn"}
}
""".trimIndent()
@Test
fun uploadRoundTrip() {
if (url == null || pin == null || cred == null) {
println("LiveUploadTest skipped (no ECHOLOT_LIVE_* env)"); return
}
val control = ControlClient(url, setOf(pin))
val profile = control.profile(cred)
val policy = profile.uploads
println("upload policy: mode=${policy.mode} min_anon=${policy.minAnonymization} " +
"max_bytes=${policy.maxBytes} retention_days=${policy.retentionDays}")
val runId = "livetest-" + System.nanoTime().toString().takeLast(10)
val level = PrivacyLevel.max(PrivacyLevel.BALANCED, PrivacyLevel.fromWire(policy.minAnonymization))
val redacted = json.encodeToString(
JsonObject.serializer(),
Anonymizer(level, Salt.perRun(ByteArray(32) { 9 }))
.anonymize(json.parseToJsonElement(sampleRun(runId)).jsonObject),
)
assertFalse(redacted.contains("Rambossek"), "the anonymizer did not strip the SSID before upload")
if (!policy.accepted) {
// A server configured to refuse must refuse — that is the behaviour worth asserting.
try {
control.uploadRun(cred, redacted)
throw AssertionError("server advertises mode=${policy.mode} but accepted an upload")
} catch (e: UploadRefused) {
println("upload correctly refused: ${e.message?.take(140)}")
return
}
}
val created = control.uploadRun(cred, redacted)
println("stored: ${created.take(200)}")
val listed = control.listRuns(cred)
assertTrue(listed.contains(runId), "uploaded run is missing from the server's list")
val fetched = control.getRun(cred, runId)
assertFalse(fetched.contains("Rambossek"), "the SSID is sitting on the server")
assertFalse(fetched.contains("Living Room TV"), "an SSDP neighbour name is sitting on the server")
assertFalse(fetched.contains("kitchen table"), "a free-text note is sitting on the server")
assertTrue(fetched.contains("nat.udp_rebinding"), "the finding code should survive — it is the point")
assertTrue(fetched.contains("12.4"), "metrics should survive anonymization")
println("round trip verified: identifiers stripped, measurements intact")
control.deleteRun(cred, runId)
assertFalse(control.listRuns(cred).contains(runId), "delete did not remove the run")
println("deleted")
}
}
@@ -0,0 +1,65 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.engine
import app.echo_lot.measurement.TestStatus
import app.echo_lot.protocol.ControlClient
import app.echo_lot.protocol.ProbeSession
import kotlin.test.Test
import kotlin.test.assertEquals
import kotlin.test.assertNotNull
import kotlin.test.assertTrue
/**
* Upstream train (types 0x03-0x05) against a LIVE server. Self-skips without ECHOLOT_LIVE_*.
*
* What is asserted is the ledger property: the server's report must account for what was sent,
* per sequence number, because directional loss attribution is the entire reason trains exist
* a test that only checked "a report came back" would pass against a server that counts nothing.
*/
class LiveUpstreamTrainTest {
private val url = System.getenv("ECHOLOT_LIVE_URL")
private val pin = System.getenv("ECHOLOT_LIVE_PIN")
private val cred = System.getenv("ECHOLOT_LIVE_CRED")
private val udp = System.getenv("ECHOLOT_LIVE_UDP")
private val target = System.getenv("ECHOLOT_LIVE_TARGET") ?: "fmr"
@Test
fun serverLedgerAccountsForTheTrain() {
if (url == null || pin == null || cred == null || udp == null) {
println("LiveUpstreamTrainTest skipped (no ECHOLOT_LIVE_* env)"); return
}
val control = ControlClient(url, setOf(pin), "0.2.0")
val session = control.createSession(cred, target)
val (host, port) = udp.split(":").let { it[0] to it[1].toInt() }
val (test, findings) = ProbeSession(cred, session, host, port).use { ps ->
ps.echo() // prime the session so its source is known
UpstreamTrainMeasurement(SystemIdSource()).run(
ps, sessionRef = "sess-1", count = 120, sizeBytes = 200, interPacketMs = 3,
)
}
control.deleteSession(cred, session.sessionId)
val m = assertNotNull(test.metrics).toString()
println("updown: ${test.status} $m")
for (f in findings) println("finding ${f.code} [${f.severity}] ${f.title}")
assertEquals(TestStatus.OK, test.status, "train report incomplete or absent: $m")
val sent = Regex(""""sent":(\d+)""").find(m)?.groupValues?.get(1)?.toInt()
val received = Regex(""""received_by_server":(\d+)""").find(m)?.groupValues?.get(1)?.toInt()
assertNotNull(sent); assertNotNull(received)
assertTrue(sent > 0, "nothing was sent: $m")
// Over a working path the ledger must be near-complete; a lossy wifi may drop a few, but
// a server that fails to count would show up as massive phantom loss here.
assertTrue(received >= sent * 9 / 10, "server counted $received of $sent: $m")
// The columnar evidence must carry a server timestamp for arrived packets — that column
// is what one-way delay math consumes after timesync.
val ev = assertNotNull(test.evidence).toString()
assertTrue(ev.contains("t_srv_rx_ns"), "no server rx column in evidence")
}
}
@@ -1,100 +0,0 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
// Package measurement models one measurement run (measurement-schema.md) — the archived,
// diffable, exportable unit. Design rules honored in the types: observation/interpretation
// separated (tests[] vs findings[]), two clocks (wall RFC3339 for humans, *_mono_ns for math),
// units in field names, columnar trains. params/evidence/metrics are per-test-type, so they are
// carried as JsonObject (the probe engine fills them; consumers ignore unknown fields).
package app.echo_lot.measurement
import kotlinx.serialization.SerialName
import kotlinx.serialization.Serializable
import kotlinx.serialization.json.JsonObject
@Serializable
data class MeasurementDocument(
val schema: String = "echolot/measurement",
@SerialName("schema_version") val schemaVersion: String = "1.0.0",
val run: Run,
val networks: List<Network> = emptyList(),
@SerialName("server_sessions") val serverSessions: List<ServerSession> = emptyList(),
val tests: List<Test> = emptyList(),
val findings: List<Finding> = emptyList(),
val summary: Summary? = null,
)
@Serializable
data class Run(
val id: String, // UUIDv7
val trigger: Trigger,
@SerialName("started_at") val startedAt: String, // RFC3339 UTC, human correlation only
@SerialName("ended_at") val endedAt: String? = null,
val clock: Clock,
val app: AppInfo,
val device: DeviceInfo,
val tiers: Tiers,
@SerialName("profiles_used") val profilesUsed: List<String> = emptyList(),
val notes: String? = null,
)
@Serializable
enum class Trigger {
@SerialName("manual") MANUAL,
@SerialName("scheduled") SCHEDULED,
@SerialName("monitor") MONITOR,
@SerialName("peer") PEER,
}
/** The two-clock anchor: mono_origin_wall maps the monotonic epoch to a wall time for humans;
* all math uses *_mono_ns relative to that monotonic origin. */
@Serializable
data class Clock(
@SerialName("mono_origin_wall") val monoOriginWall: String,
@SerialName("ntp_offset_ms") val ntpOffsetMs: Double? = null,
@SerialName("ntp_offset_source") val ntpOffsetSource: String? = null,
)
@Serializable
data class AppInfo(
val version: String,
val build: Int,
val git: String? = null,
val flavor: String? = null,
)
@Serializable
data class DeviceInfo(
val manufacturer: String,
val model: String,
@SerialName("android_sdk") val androidSdk: Int,
@SerialName("android_release") val androidRelease: String,
@SerialName("security_patch") val securityPatch: String? = null,
)
/** What each tier was *available*; each test records what it *used*. */
@Serializable
data class Tiers(
val app: Boolean = true,
val shizuku: Boolean = false,
val root: Boolean = false,
)
@Serializable
data class ServerSession(
val id: String,
@SerialName("profile_id") val profileId: String? = null,
@SerialName("profile_name") val profileName: String? = null,
@SerialName("control_url") val controlUrl: String,
@SerialName("server_version") val serverVersion: String? = null,
val capabilities: List<String> = emptyList(),
@SerialName("session_id") val sessionId: String,
val target: SessionTarget,
)
@Serializable
data class SessionTarget(
val ip4: String? = null,
val ip6: String? = null,
@SerialName("udp_port") val udpPort: Int = 0,
)
@@ -1,85 +0,0 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.measurement
import kotlinx.serialization.SerialName
import kotlinx.serialization.Serializable
import kotlinx.serialization.json.Json
import kotlinx.serialization.json.JsonObject
import kotlinx.serialization.json.encodeToJsonElement
/**
* Typed builders for the per-test-type evidence shapes the schema fixes (§6.2/§6.3/§6.4). Probe
* code fills these and folds them into [Test.evidence] via [toEvidence]; keeping them typed here
* means the columnar/traceroute/resolver contracts live in one place.
*/
@PublishedApi
internal val evidenceJson = Json { encodeDefaults = true; explicitNulls = true }
/** Serialize any typed evidence object into the JsonObject the Test envelope carries. */
inline fun <reified T> T.toEvidence(): JsonObject =
evidenceJson.encodeToJsonElement(this) as JsonObject
/**
* Packet-train evidence (§6.2): columnar parallel arrays, one index per probe packet. Missing
* observations are null at that index a 10k-packet train stays in the hundreds of kB. Server
* columns use the server session epoch; only differences within one clock are meaningful unless a
* time.server_offset test maps them.
*/
@Serializable
data class TrainEvidence(
@SerialName("epoch_mono_ns") val epochMonoNs: Long,
val seq: List<Int>,
@SerialName("t_tx_ns") val tTxNs: List<Long?>,
@SerialName("t_srv_rx_ns") val tSrvRxNs: List<Long?> = emptyList(),
@SerialName("t_srv_tx_ns") val tSrvTxNs: List<Long?> = emptyList(),
@SerialName("t_rx_ns") val tRxNs: List<Long?>,
@SerialName("size_bytes") val sizeBytes: List<Int>,
@SerialName("dscp_sent") val dscpSent: Int? = null,
@SerialName("dscp_seen_by_server") val dscpSeenByServer: List<Int?> = emptyList(),
@SerialName("ecn_sent") val ecnSent: Int? = null,
@SerialName("ecn_seen_by_server") val ecnSeenByServer: List<Int?> = emptyList(),
@SerialName("ttl_seen_by_server") val ttlSeenByServer: List<Int?> = emptyList(),
@SerialName("evidence_truncated") val evidenceTruncated: Boolean = false,
)
/** Traceroute evidence (§6.3): fixed-tuple flow + per-TTL probe replies. */
@Serializable
data class TracerouteEvidence(val flow: Flow, val hops: List<Hop>)
@Serializable
data class Flow(
@SerialName("src_port") val srcPort: Int,
@SerialName("dst_port") val dstPort: Int,
@SerialName("fixed_tuple") val fixedTuple: Boolean = true,
)
@Serializable
data class Hop(val ttl: Int, val probes: List<HopProbe>)
@Serializable
data class HopProbe(
@SerialName("reply_from") val replyFrom: String? = null,
@SerialName("rtt_ns") val rttNs: Long? = null,
val icmp: String? = null,
@SerialName("reply_ttl") val replyTtl: Int? = null,
)
/** Resolver under test (§6.4); every dns.* test carries this in params. */
@Serializable
data class ResolverSpec(
val source: ResolverSource,
val address: String? = null,
val port: Int = 53,
val transport: String, // do53-udp | do53-tcp | dot | doh
@SerialName("doh_url") val dohUrl: String? = null,
)
@Serializable
enum class ResolverSource {
@SerialName("system") SYSTEM,
@SerialName("manual") MANUAL,
@SerialName("server-recursive") SERVER_RECURSIVE,
}
@@ -1,68 +0,0 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.measurement
import kotlinx.serialization.SerialName
import kotlinx.serialization.Serializable
/** Interpretation with references back to evidence (measurement-schema.md §7.1). A finding with
* no evidence_refs is invalid every finding must be re-derivable from the evidence alone. */
@Serializable
data class Finding(
val id: String, // UUIDv7
val code: String, // stable registry (findings-registry.md), lint-rule style
val category: Category,
val severity: Severity,
val confidence: Confidence,
@SerialName("network_ref") val networkRef: String? = null,
val title: String,
val description: String,
@SerialName("evidence_refs") val evidenceRefs: List<EvidenceRef>,
val recommendation: String? = null,
) {
init {
require(evidenceRefs.isNotEmpty()) { "a finding must reference at least one piece of evidence" }
}
}
@Serializable
data class EvidenceRef(val test: String, val pointer: String? = null)
/** Fixed §7.2 categories; each maps to one traffic light. */
@Serializable
enum class Category {
@SerialName("connectivity") CONNECTIVITY,
@SerialName("dns") DNS,
@SerialName("nat") NAT,
@SerialName("mtu") MTU,
@SerialName("ipv6") IPV6,
@SerialName("security") SECURITY,
@SerialName("performance") PERFORMANCE,
@SerialName("local") LOCAL,
@SerialName("wifi") WIFI,
}
/** Ordered worst→best via [rank]; drives the §7.3 light mapping. */
@Serializable
enum class Severity(val rank: Int) {
@SerialName("critical") CRITICAL(4),
@SerialName("high") HIGH(3),
@SerialName("medium") MEDIUM(2),
@SerialName("low") LOW(1),
@SerialName("info") INFO(0);
/** §7.3: critical|high → red, medium|low → yellow, info → green. */
fun toLight(): Verdict = when (this) {
CRITICAL, HIGH -> Verdict.RED
MEDIUM, LOW -> Verdict.YELLOW
INFO -> Verdict.GREEN
}
}
@Serializable
enum class Confidence {
@SerialName("high") HIGH,
@SerialName("medium") MEDIUM,
@SerialName("low") LOW,
}
@@ -1,113 +0,0 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.measurement
import kotlinx.serialization.SerialName
import kotlinx.serialization.Serializable
import kotlinx.serialization.json.JsonObject
/** One Android Network in play (measurement-schema.md §4). Shizuku-tier fields (route proto,
* lifetimes) are absent at app tier absence means "not observed", never "not present". */
@Serializable
data class Network(
val id: String,
val transport: Transport,
@SerialName("interface") val iface: String? = null,
val link: Link,
val wifi: Wifi? = null,
val cellular: Cellular? = null,
val changes: List<NetworkChange> = emptyList(),
)
@Serializable
enum class Transport {
@SerialName("wifi") WIFI,
@SerialName("cellular") CELLULAR,
@SerialName("ethernet") ETHERNET,
@SerialName("vpn") VPN,
@SerialName("other") OTHER,
}
@Serializable
data class Link(
val mtu: Int? = null,
val addresses: List<Address> = emptyList(),
val routes: List<Route> = emptyList(),
val dns: DnsConfig? = null,
val dhcp: Dhcp? = null,
@SerialName("captive_portal") val captivePortal: CaptivePortal? = null,
)
@Serializable
data class Address(
val addr: String, // ip4 | ip6 (logical type, §8)
@SerialName("prefix_len") val prefixLen: Int,
val scope: String? = null,
val flags: List<String> = emptyList(),
@SerialName("valid_lft_s") val validLftS: Long? = null,
@SerialName("pref_lft_s") val prefLftS: Long? = null,
)
@Serializable
data class Route(
val dst: String,
val gateway: String? = null,
val iface: String? = null,
val proto: RouteProto? = null, // shizuku tier; null = not observed
@SerialName("expires_s") val expiresS: Long? = null,
)
@Serializable
enum class RouteProto {
@SerialName("dhcp") DHCP,
@SerialName("ra") RA,
@SerialName("static") STATIC,
@SerialName("unknown") UNKNOWN,
}
@Serializable
data class DnsConfig(
val servers: List<String> = emptyList(),
@SerialName("private_dns_mode") val privateDnsMode: String? = null,
@SerialName("private_dns_hostname") val privateDnsHostname: String? = null,
@SerialName("search_domains") val searchDomains: List<String> = emptyList(),
@SerialName("nat64_prefix") val nat64Prefix: String? = null,
)
@Serializable
data class Dhcp(val server: String? = null, @SerialName("lease_s") val leaseS: Long? = null)
@Serializable
data class CaptivePortal(
val detected: Boolean = false,
@SerialName("api_url") val apiUrl: String? = null,
@SerialName("venue_url") val venueUrl: String? = null,
)
@Serializable
data class Wifi(
val ssid: String? = null, // ssid (logical type)
val bssid: String? = null, // bssid (logical type)
@SerialName("rssi_dbm") val rssiDbm: Int? = null,
@SerialName("link_speed_mbps") val linkSpeedMbps: Int? = null,
@SerialName("frequency_mhz") val frequencyMhz: Int? = null,
@SerialName("channel_width_mhz") val channelWidthMhz: Int? = null,
val standard: String? = null,
val security: String? = null,
@SerialName("mac_randomization") val macRandomization: Boolean? = null,
)
@Serializable
data class Cellular(
val rat: String? = null,
val operator: String? = null,
val band: String? = null,
)
@Serializable
data class NetworkChange(
@SerialName("at_mono_ns") val atMonoNs: Long,
val kind: String, // lost | gained | link_changed
val detail: JsonObject? = null,
)
@@ -1,90 +0,0 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.measurement
import kotlinx.serialization.SerialName
import kotlinx.serialization.Serializable
@Serializable
enum class Verdict {
@SerialName("green") GREEN,
@SerialName("yellow") YELLOW,
@SerialName("red") RED,
@SerialName("inconclusive") INCONCLUSIVE,
}
@Serializable
data class Summary(
val overall: Verdict,
val categories: Map<String, CategorySummary>,
)
@Serializable
data class CategorySummary(
val verdict: Verdict,
@SerialName("worst_finding") val worstFinding: String? = null,
@SerialName("tests_run") val testsRun: Int,
@SerialName("tests_failed") val testsFailed: Int,
)
/**
* Deterministic verdict derivation, fixed by measurement-schema.md §7.3:
*
* - A category's verdict = the light of its worst-severity finding
* (critical|high red, medium|low yellow, info/none green).
* - A category is `inconclusive` when > 50% of its tests are failed/unsupported.
* - Overall = the worst category light; `inconclusive` only when ALL categories are.
*
* The mapping test-type category comes from [TestType.category]. Only categories that have
* findings or tests appear in the summary.
*/
object Verdicts {
private fun isInconclusiveTest(s: TestStatus) =
s == TestStatus.FAILED || s == TestStatus.UNSUPPORTED
fun derive(tests: List<Test>, findings: List<Finding>): Summary {
val testsByCat = tests.groupBy { TestType.category(it.type) }
val findingsByCat = findings.groupBy { it.category }
val categories = (testsByCat.keys + findingsByCat.keys)
val perCat = LinkedHashMap<String, CategorySummary>()
for (cat in Category.entries) {
if (cat !in categories) continue
val catTests = testsByCat[cat].orEmpty()
val catFindings = findingsByCat[cat].orEmpty()
val failed = catTests.count { isInconclusiveTest(it.status) }
val inconclusive = catTests.isNotEmpty() && failed * 2 > catTests.size
val worst = catFindings.maxByOrNull { it.severity.rank }
val verdict = when {
inconclusive -> Verdict.INCONCLUSIVE
worst == null -> Verdict.GREEN
else -> worst.severity.toLight()
}
perCat[serialName(cat)] = CategorySummary(
verdict = verdict,
worstFinding = worst?.id,
testsRun = catTests.size,
testsFailed = failed,
)
}
val overall = deriveOverall(perCat.values)
return Summary(overall = overall, categories = perCat)
}
/** Overall = worst light; inconclusive only if every category is inconclusive. */
private fun deriveOverall(cats: Collection<CategorySummary>): Verdict {
if (cats.isEmpty()) return Verdict.INCONCLUSIVE
if (cats.all { it.verdict == Verdict.INCONCLUSIVE }) return Verdict.INCONCLUSIVE
val rank = mapOf(Verdict.GREEN to 0, Verdict.YELLOW to 1, Verdict.RED to 2)
// Non-inconclusive categories decide the overall light.
return cats.filter { it.verdict != Verdict.INCONCLUSIVE }
.maxByOrNull { rank.getValue(it.verdict) }!!.verdict
}
private fun serialName(cat: Category): String = cat.name.lowercase()
}
@@ -1,149 +0,0 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.measurement
import kotlinx.serialization.SerialName
import kotlinx.serialization.Serializable
import kotlinx.serialization.json.JsonObject
/** The generic test envelope (measurement-schema.md §6). params must fully reproduce the test;
* evidence is append-only raw truth; metrics must be recomputable from evidence. All three are
* per-test-type JSON, so they are carried as JsonObject. */
@Serializable
data class Test(
val id: String, // UUIDv7
val type: String, // TestType registry (§6.1)
@SerialName("network_ref") val networkRef: String? = null,
@SerialName("session_ref") val sessionRef: String? = null, // null for local-only tests
val tier: Tier,
@SerialName("started_mono_ns") val startedMonoNs: Long,
@SerialName("ended_mono_ns") val endedMonoNs: Long,
val status: TestStatus,
val error: TestError? = null,
val params: JsonObject? = null,
val evidence: JsonObject? = null,
val metrics: JsonObject? = null,
)
@Serializable
enum class Tier {
@SerialName("app") APP,
@SerialName("shizuku") SHIZUKU,
@SerialName("root") ROOT,
}
@Serializable
enum class TestStatus {
@SerialName("ok") OK,
@SerialName("failed") FAILED,
@SerialName("unsupported") UNSUPPORTED,
@SerialName("skipped") SKIPPED,
@SerialName("partial") PARTIAL,
}
@Serializable
data class TestError(val code: String, val detail: String? = null)
/**
* The v1 test-type registry (§6.1). String constants (dotted, family-first) so probe code and the
* server's measurement-schema test-type registry stay aligned. [category] maps a type to one of
* the fixed §7.2 categories for verdict rollup.
*/
object TestType {
// link
const val LINK_SNAPSHOT = "link.snapshot"
const val LINK_DHCP_RENEWAL_WATCH = "link.dhcp_renewal_watch"
const val LINK_IP_MONITOR = "link.ip_monitor"
/** Who advertises IPv6 on this link (+ gateway identity). Registry addition, v1.1. */
const val LINK_RA_SOURCE = "link.ra_source"
// net — connectivity validation (reproduces Android's NetworkMonitor generate_204 checks)
const val NET_CAPTIVE_PORTAL = "net.captive_portal"
// icmp
const val ICMP_PING4 = "icmp.ping4"
const val ICMP_PING6 = "icmp.ping6"
// trace
const val TRACEROUTE_UDP4 = "traceroute.udp4"
const val TRACEROUTE_UDP6 = "traceroute.udp6"
const val TRACEROUTE_ICMP4 = "traceroute.icmp4"
const val TRACEROUTE_ICMP6 = "traceroute.icmp6"
// train
const val TRAIN_UDP_UPDOWN = "train.udp_updown"
// mtu
const val MTU_PMTUD_UP = "mtu.pmtud_up"
const val MTU_PMTUD_DOWN = "mtu.pmtud_down"
const val MTU_BLACKHOLE = "mtu.blackhole"
const val MTU_MSS_OBSERVED = "mtu.mss_observed"
const val MTU_FRAG_DELIVERY = "mtu.frag_delivery"
// nat
const val NAT_STUN_5780 = "nat.stun_5780"
const val NAT_MAPPING_LIFETIME_UDP = "nat.mapping_lifetime_udp"
const val NAT_MAPPING_LIFETIME_TCP = "nat.mapping_lifetime_tcp"
const val NAT_HAIRPIN = "nat.hairpin"
const val NAT_CONNECT_BACK = "nat.connect_back"
const val NAT_CGNAT_DETECT = "nat.cgnat_detect"
// dns
const val DNS_RESOLVER_INVENTORY = "dns.resolver_inventory"
const val DNS_CANARY = "dns.canary"
const val DNS_INTERCEPTION = "dns.interception"
const val DNS_TTL_INTEGRITY = "dns.ttl_integrity"
const val DNS_ANSWER_INTEGRITY = "dns.answer_integrity"
const val DNS_DNSSEC = "dns.dnssec"
const val DNS_NXDOMAIN_WILDCARD = "dns.nxdomain_wildcard"
const val DNS_REBIND_FILTER = "dns.rebind_filter"
const val DNS_AAAA_FILTER = "dns.aaaa_filter"
const val DNS_DNS64 = "dns.dns64"
const val DNS_COMPARE = "dns.compare"
// sec
const val SEC_TLS_REFERENCE = "sec.tls_reference"
const val SEC_CLIENTHELLO_ECHO = "sec.clienthello_echo"
const val SEC_HTTP_ECHO = "sec.http_echo"
const val SEC_SNI_FILTER = "sec.sni_filter"
const val SEC_DSCP_ECN_SURVIVAL = "sec.dscp_ecn_survival"
const val SEC_ARP_WATCH = "sec.arp_watch"
// port
const val PORT_REACH_SWEEP = "port.reach_sweep"
const val PORT_UDP_USABILITY = "port.udp_usability"
// perf
const val PERF_THROUGHPUT_TCP = "perf.throughput_tcp"
const val PERF_THROUGHPUT_UDP = "perf.throughput_udp"
const val PERF_BUFFERBLOAT = "perf.bufferbloat"
const val PERF_RRC_LATENCY = "perf.rrc_latency"
// v6
const val V6_DUALSTACK_COMPARE = "v6.dualstack_compare"
const val V6_HAPPY_EYEBALLS = "v6.happy_eyeballs"
const val V6_BROKENNESS = "v6.brokenness"
const val V6_NAT64_CLAT = "v6.nat64_clat"
// wifi
const val WIFI_ENVIRONMENT_SCAN = "wifi.environment_scan"
const val WIFI_ROAM_LOG = "wifi.roam_log"
const val WIFI_SIGNAL_LOG = "wifi.signal_log"
// local
const val LOCAL_MDNS_INVENTORY = "local.mdns_inventory"
const val LOCAL_SSDP_INVENTORY = "local.ssdp_inventory"
const val LOCAL_LLMNR_INVENTORY = "local.llmnr_inventory"
const val LOCAL_GATEWAY_SERVICES = "local.gateway_services"
const val LOCAL_NTP = "local.ntp"
// peer
const val PEER_REACHABILITY = "peer.reachability"
const val PEER_ISOLATION = "peer.isolation"
const val PEER_MULTICAST = "peer.multicast"
const val PEER_LAN_TRAIN = "peer.lan_train"
const val PEER_LEASE_DIFF = "peer.lease_diff"
// time
const val TIME_SERVER_OFFSET = "time.server_offset"
/** Maps a dotted test type to its §7.2 category for verdict rollup. */
fun category(type: String): Category = when (type.substringBefore('.')) {
"link", "icmp", "trace", "traceroute", "train", "port", "time", "net" -> Category.CONNECTIVITY
"dns" -> Category.DNS
"nat" -> Category.NAT
"mtu" -> Category.MTU
"v6" -> Category.IPV6
"sec" -> Category.SECURITY
"perf" -> Category.PERFORMANCE
"local", "peer" -> Category.LOCAL
"wifi" -> Category.WIFI
else -> Category.CONNECTIVITY
}
}
@@ -1,71 +0,0 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.measurement
import kotlinx.serialization.json.Json
import kotlin.test.Test as JTest
import kotlin.test.assertEquals
import kotlin.test.assertTrue
class SerializationTest {
private val json = Json { ignoreUnknownKeys = true; encodeDefaults = true }
@JTest
fun documentRoundTrips() {
val doc = MeasurementDocument(
run = Run(
id = "0198c5f2-0000-7000-8000-000000000000",
trigger = Trigger.MANUAL,
startedAt = "2026-07-31T14:03:21.114Z",
clock = Clock(monoOriginWall = "2026-07-31T14:03:21.114Z"),
app = AppInfo(version = "0.1.0", build = 1),
device = DeviceInfo("OnePlus", "CPH2747", 36, "16"),
tiers = Tiers(app = true, shizuku = true),
),
networks = listOf(
Network(
id = "net-1", transport = Transport.WIFI, iface = "wlan0",
link = Link(mtu = 1500, addresses = listOf(Address("192.0.2.23", 24, "global"))),
wifi = Wifi(ssid = "example", rssiDbm = -54),
),
),
tests = listOf(
Test(
id = "t-1", type = TestType.ICMP_PING4, networkRef = "net-1", tier = Tier.APP,
startedMonoNs = 0, endedMonoNs = 38_000_000, status = TestStatus.OK,
evidence = TrainEvidence(
epochMonoNs = 0, seq = listOf(0, 1), tTxNs = listOf(0L, 20_000_000L),
tRxNs = listOf(16_500_000L, null), sizeBytes = listOf(64, 64),
).toEvidence(),
),
),
)
val encoded = json.encodeToString(MeasurementDocument.serializer(), doc)
val decoded = json.decodeFromString(MeasurementDocument.serializer(), encoded)
assertEquals(doc.run.id, decoded.run.id)
assertEquals(Transport.WIFI, decoded.networks[0].transport)
assertEquals(TestType.ICMP_PING4, decoded.tests[0].type)
// snake_case field names on the wire
assertTrue(encoded.contains("\"schema_version\""))
assertTrue(encoded.contains("\"mono_origin_wall\""))
assertTrue(encoded.contains("\"t_tx_ns\""))
// null preserved at train index 1
assertTrue(encoded.contains("[16500000,null]"))
}
@JTest
fun findingRequiresEvidence() {
try {
Finding(
id = "f-1", code = "x", category = Category.DNS, severity = Severity.INFO,
confidence = Confidence.LOW, title = "t", description = "d", evidenceRefs = emptyList(),
)
throw AssertionError("expected IllegalArgumentException for empty evidence_refs")
} catch (e: IllegalArgumentException) {
// expected — a finding with no evidence is invalid (§7.1)
}
}
}
@@ -1,109 +0,0 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.measurement
import kotlin.test.Test as JTest
import kotlin.test.assertEquals
class VerdictsTest {
private fun test(type: String, status: TestStatus, id: String = type): Test =
Test(id = id, type = type, tier = Tier.APP, startedMonoNs = 0, endedMonoNs = 1, status = status)
private fun finding(cat: Category, sev: Severity, id: String = "f-$cat-$sev"): Finding =
Finding(
id = id, code = "x.$cat", category = cat, severity = sev, confidence = Confidence.HIGH,
title = "t", description = "d", evidenceRefs = listOf(EvidenceRef("some-test")),
)
@JTest
fun categoryLightFromWorstSeverity() {
val tests = listOf(test(TestType.DNS_CANARY, TestStatus.OK))
val findings = listOf(
finding(Category.DNS, Severity.LOW),
finding(Category.DNS, Severity.HIGH), // worst → red
finding(Category.DNS, Severity.INFO),
)
val s = Verdicts.derive(tests, findings)
assertEquals(Verdict.RED, s.categories["dns"]!!.verdict)
assertEquals("f-DNS-HIGH", s.categories["dns"]!!.worstFinding)
assertEquals(Verdict.RED, s.overall)
}
@JTest
fun noFindingsIsGreen() {
val s = Verdicts.derive(listOf(test(TestType.MTU_BLACKHOLE, TestStatus.OK)), emptyList())
assertEquals(Verdict.GREEN, s.categories["mtu"]!!.verdict)
assertEquals(Verdict.GREEN, s.overall)
}
@JTest
fun mediumAndLowAreYellow() {
val s = Verdicts.derive(
listOf(test(TestType.SEC_HTTP_ECHO, TestStatus.OK)),
listOf(finding(Category.SECURITY, Severity.MEDIUM)),
)
assertEquals(Verdict.YELLOW, s.categories["security"]!!.verdict)
}
@JTest
fun majorityFailedIsInconclusive() {
// 2 of 3 dns tests failed → > 50% → inconclusive, even with a finding present.
val tests = listOf(
test(TestType.DNS_CANARY, TestStatus.FAILED, "a"),
test(TestType.DNS_TTL_INTEGRITY, TestStatus.UNSUPPORTED, "b"),
test(TestType.DNS_COMPARE, TestStatus.OK, "c"),
)
val s = Verdicts.derive(tests, listOf(finding(Category.DNS, Severity.HIGH)))
assertEquals(Verdict.INCONCLUSIVE, s.categories["dns"]!!.verdict)
assertEquals(2, s.categories["dns"]!!.testsFailed)
assertEquals(3, s.categories["dns"]!!.testsRun)
}
@JTest
fun exactlyHalfFailedIsNotInconclusive() {
// 1 of 2 failed → not > 50% → the finding decides.
val tests = listOf(
test(TestType.NAT_HAIRPIN, TestStatus.FAILED, "a"),
test(TestType.NAT_CONNECT_BACK, TestStatus.OK, "b"),
)
val s = Verdicts.derive(tests, listOf(finding(Category.NAT, Severity.CRITICAL)))
assertEquals(Verdict.RED, s.categories["nat"]!!.verdict)
}
@JTest
fun overallIsWorstCategory() {
val tests = listOf(
test(TestType.DNS_CANARY, TestStatus.OK, "d"),
test(TestType.MTU_BLACKHOLE, TestStatus.OK, "m"),
)
val findings = listOf(
finding(Category.DNS, Severity.MEDIUM), // yellow
finding(Category.MTU, Severity.CRITICAL), // red
)
val s = Verdicts.derive(tests, findings)
assertEquals(Verdict.RED, s.overall)
}
@JTest
fun overallInconclusiveOnlyWhenAllAre() {
val tests = listOf(
test(TestType.DNS_CANARY, TestStatus.FAILED, "d"), // dns inconclusive
test(TestType.MTU_BLACKHOLE, TestStatus.OK, "m"), // mtu green
)
val s = Verdicts.derive(tests, emptyList())
assertEquals(Verdict.INCONCLUSIVE, s.categories["dns"]!!.verdict)
assertEquals(Verdict.GREEN, s.categories["mtu"]!!.verdict)
assertEquals(Verdict.GREEN, s.overall) // not all inconclusive → mtu decides
}
@JTest
fun categoryMappingCoversFamilies() {
assertEquals(Category.CONNECTIVITY, TestType.category(TestType.TRACEROUTE_UDP4))
assertEquals(Category.IPV6, TestType.category(TestType.V6_BROKENNESS))
assertEquals(Category.LOCAL, TestType.category(TestType.PEER_MULTICAST))
assertEquals(Category.PERFORMANCE, TestType.category(TestType.PERF_BUFFERBLOAT))
assertEquals(Category.CONNECTIVITY, TestType.category(TestType.TIME_SERVER_OFFSET))
}
}
@@ -28,6 +28,7 @@ data class MeasurementDocument(
data class Run( data class Run(
val id: String, // UUIDv7 val id: String, // UUIDv7
val trigger: Trigger, val trigger: Trigger,
val mode: RunMode = RunMode.SHORT,
@SerialName("started_at") val startedAt: String, // RFC3339 UTC, human correlation only @SerialName("started_at") val startedAt: String, // RFC3339 UTC, human correlation only
@SerialName("ended_at") val endedAt: String? = null, @SerialName("ended_at") val endedAt: String? = null,
val clock: Clock, val clock: Clock,
@@ -35,9 +36,58 @@ data class Run(
val device: DeviceInfo, val device: DeviceInfo,
val tiers: Tiers, val tiers: Tiers,
@SerialName("profiles_used") val profilesUsed: List<String> = emptyList(), @SerialName("profiles_used") val profilesUsed: List<String> = emptyList(),
val constraints: Constraints = Constraints(),
val notes: String? = null, val notes: String? = null,
) )
/**
* What limited this run the counterpart to [Tiers], which records what was available.
*
* A constrained run is not a failed run, and it is not a normal one either. Without this, a run
* taken through a VPN looks exactly like a clean run of a healthy network: the same shape, the
* same green verdict, and no way for a reader or a server aggregating thousands of these to
* know that almost nothing was actually measured.
*/
@Serializable
data class Constraints(
/** A VPN held the default route while this ran. */
@SerialName("vpn_active") val vpnActive: Boolean = false,
/**
* Per-network probing was refused by the OS.
*
* Android blocks `Network.bindSocket()` on the underlying networks whenever a VPN is up, to
* stop apps leaking around the tunnel. Every per-network test then measures nothing, so any
* conclusion drawn about the wifi or cellular link underneath is unfounded.
*/
@SerialName("per_network_blocked") val perNetworkBlocked: Boolean = false,
/** Networks that could not be measured, by id. */
@SerialName("unmeasured_networks") val unmeasuredNetworks: List<String> = emptyList(),
) {
/** True when this run's results mean something different from an unconstrained one. */
val constrained: Boolean get() = vpnActive || perNetworkBlocked
}
/**
* How long the run watched the network and therefore what its silence is worth.
*
* A [SHORT] run is a sequence of one-shot probes: each looks at the network for a second or two and
* moves on. That is enough to characterise a network's *configuration*, and it is structurally
* incapable of seeing anything intermittent. A wifi link that drops for four seconds every two
* minutes, a resolver that stalls under load, an AP that roams none of these leave a trace in
* thirty seconds of probing unless the run happened to coincide with one.
*
* A [LONG] run starts continuous listeners at t=0, runs the same battery beside them, and keeps
* sampling until the window closes. It answers a different question, so a reader must not treat the
* two alike: **the mode is what licenses an argument from absence**. "No drops were observed" means
* something after five minutes of watching and nothing at all after a thirty-second run, and
* without this field the two documents are indistinguishable.
*/
@Serializable
enum class RunMode {
@SerialName("short") SHORT,
@SerialName("long") LONG,
}
@Serializable @Serializable
enum class Trigger { enum class Trigger {
@SerialName("manual") MANUAL, @SerialName("manual") MANUAL,
@@ -0,0 +1,341 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.measurement
/**
* The registry of finding codes (measurement-schema.md §9, open item 1).
*
* A finding code is the stable, machine-readable half of a result: the prose changes, the code is
* what a dashboard groups by and what someone greps a year of archived runs for. That only holds
* if a code means exactly one thing forever which is not something ad-hoc string literals at
* fifteen call sites can promise.
*
* The failure this exists to prevent had already happened by the time it was written. Two
* independently-added emitters produced `connectivity.downstream_loss` and
* `connectivity.loss_downstream` for the same concept, and nothing anywhere objected. Anyone
* aggregating either one would have silently seen half their data.
*
* So codes are declared here as typed specs, each carrying its category and default severity, and
* emitters reference the spec rather than retyping the string. That makes a typo a compile error,
* and makes it impossible for two call sites to disagree about which category a finding belongs
* to a disagreement that would otherwise split one fault across two verdict lights.
*/
data class FindingSpec(
val code: String,
val category: Category,
/** Severity when nothing about the specific run argues otherwise; emitters may escalate. */
val severity: Severity,
/** One line: what this finding asserts. Present tense, no hedging. */
val meaning: String,
/**
* What the finding rules *out*, where that is the useful half. "Loss upstream" is worth much
* more when it also says the return path is fine, because that halves where to look next.
*/
val rulesOut: String? = null,
)
object FindingRegistry {
// ---- connectivity ----------------------------------------------------------------
// Renamed from nat.* before anything shipped: neither of these is about NAT, and the
// prefix is what decides which category - and therefore which verdict light - a finding
// rolls up into. A nat.* code landing under connectivity would be a permanent puzzle.
val UDP_UNREACHABLE = FindingSpec(
"connectivity.udp_unreachable", Category.CONNECTIVITY, Severity.HIGH,
"No UDP echo replies came back from the server at all.",
)
val UDP_UNREACHABLE_UPSTREAM = FindingSpec(
"connectivity.udp_unreachable_upstream", Category.CONNECTIVITY, Severity.HIGH,
"The server received none of the probes, so traffic is dropped on the way out.",
rulesOut = "The return path: nothing arrived to be replied to.",
)
val UDP_LOSS = FindingSpec(
"connectivity.udp_loss", Category.CONNECTIVITY, Severity.MEDIUM,
"A large fraction of round-trip probes were lost, direction unknown.",
)
val LOSS_UPSTREAM = FindingSpec(
"connectivity.loss_upstream", Category.CONNECTIVITY, Severity.MEDIUM,
"Probes were lost on the way to the server.",
rulesOut = "The return path: replies came back for everything that arrived.",
)
/**
* The single code for "lost on the return path", whichever measurement found it.
*
* Two emitters had independently invented `connectivity.downstream_loss` and
* `connectivity.loss_downstream` for this, and nothing objected. Anyone aggregating either
* one would have silently seen half their data. Paired with [LOSS_UPSTREAM] so the two
* directions read as a set.
*/
val LOSS_DOWNSTREAM = FindingSpec(
"connectivity.loss_downstream", Category.CONNECTIVITY, Severity.MEDIUM,
"Packets were lost on the way back from the server.",
rulesOut = "The outbound path: the server received what it was answering.",
)
val DOWNSTREAM_BLOCKED = FindingSpec(
"connectivity.downstream_blocked", Category.CONNECTIVITY, Severity.HIGH,
"Server-initiated packets never arrive, although round trips work.",
rulesOut = "Basic reachability: the path forwards replies, just not unsolicited traffic.",
)
val DOWNSTREAM_REORDER = FindingSpec(
"connectivity.downstream_reorder", Category.CONNECTIVITY, Severity.LOW,
"Downstream packets arrive in a different order than they were sent.",
)
// MEDIUM, not HIGH: a captive portal is a condition to report, not necessarily a fault - on
// hotel or cafe wifi it is exactly what should be there, and logging in clears it. NO_INTERNET
// is the HIGH one, because nothing the user does locally fixes that. The registry first said
// HIGH; the probe emitting it had always said MEDIUM, and the probe was the considered value.
val CAPTIVE_PORTAL = FindingSpec(
"connectivity.captive_portal", Category.CONNECTIVITY, Severity.MEDIUM,
"A captive portal is intercepting connectivity checks.",
)
val NO_INTERNET = FindingSpec(
"connectivity.no_internet", Category.CONNECTIVITY, Severity.HIGH,
"Android's own connectivity checks fail on this network.",
)
/**
* The finding a short run cannot make.
*
* Every one-shot probe describes the network during its own two seconds. A link that drops and
* returns between two of them leaves no trace anywhere in the document the probes before and
* after both succeed, and the run reports a healthy network. Only a listener that watches the
* whole window sees the gap, which is why this is emitted from `networks[].changes[]` (§4)
* rather than from any test's evidence.
*
* MEDIUM by default and escalated by the emitter on repeat: one drop in five minutes is worth
* knowing about, three is the difference between "the wifi hiccuped" and "this link is why
* calls keep dropping". Deliberately claims a *completed* cycle — lost and then regained — so
* it never fires for a network that was simply turned off partway through the run.
*/
val LINK_FLAPPING = FindingSpec(
"connectivity.link_flapping", Category.CONNECTIVITY, Severity.MEDIUM,
"A network dropped and came back one or more times during the run.",
rulesOut = "A momentary probe failure: the drop was watched happening, not inferred from silence.",
)
// ---- mtu -------------------------------------------------------------------------
val MTU_REDUCED_DOWNSTREAM = FindingSpec(
"mtu.reduced_downstream", Category.MTU, Severity.LOW,
"The downstream path MTU is below the usual 1500 bytes.",
)
val MTU_DOWNSTREAM_BLACKHOLE = FindingSpec(
"mtu.downstream_blackhole", Category.MTU, Severity.MEDIUM,
"Datagrams above the path MTU are dropped downstream, fragmented or not.",
)
val FRAGMENTS_BLOCKED = FindingSpec(
"mtu.fragments_blocked", Category.MTU, Severity.MEDIUM,
"IP fragments do not reach this device even when sent in order.",
)
val FRAGMENT_REORDER_SENSITIVE = FindingSpec(
"mtu.fragment_reorder_sensitive", Category.MTU, Severity.LOW,
"Fragments are delivered in order but dropped when reordered or delayed.",
rulesOut = "Fragmentation itself: in-order fragments arrive fine.",
)
// ---- nat -------------------------------------------------------------------------
val NAT_UDP_REBINDING = FindingSpec(
"nat.udp_rebinding", Category.NAT, Severity.MEDIUM,
"A NAT remapped the UDP source port mid-flow.",
)
val NAT_SYMMETRIC = FindingSpec(
"nat.symmetric", Category.NAT, Severity.MEDIUM,
"The NAT assigns a different external port per destination.",
)
// ---- perf ------------------------------------------------------------------------
val THROUGHPUT_NO_DELIVERY = FindingSpec(
"perf.throughput_no_delivery", Category.PERFORMANCE, Severity.HIGH,
"No throughput traffic arrived, although the server sent it.",
)
val THROUGHPUT_BELOW_OFFERED = FindingSpec(
"perf.throughput_below_offered", Category.PERFORMANCE, Severity.LOW,
"Less throughput arrived than the server sent for the whole run.",
)
// ---- dns -------------------------------------------------------------------------
val DNS_ANSWER_REWRITTEN = FindingSpec(
"dns.answer_rewritten", Category.DNS, Severity.HIGH,
"A resolver returned an answer that differs from the authoritative record.",
)
val DNS_AUTHORITATIVE_UNREACHABLE = FindingSpec(
"dns.authoritative_unreachable", Category.DNS, Severity.MEDIUM,
"The canary zone's authoritative server could not be reached.",
)
// ---- v6 ----------------------------------------------------------------------------
//
// Prefix is `v6.`, matching the test-type registry (v6.brokenness, v6.happy_eyeballs, ...).
// These were `ipv6.*` while declaring Category.IPV6, but the prefix map only knows "v6", so
// they silently rolled up under connectivity: the third instance of a prefix disagreeing with
// its category and quietly moving a fault to a different verdict light.
/**
* Renamed from `v6.broken`, which claimed more than the evidence supports.
*
* The only signal behind it is ICMPv6 echo getting no reply and ICMPv6 echo is widely
* filtered on networks where IPv6 otherwise works perfectly. A phone that reported this while
* happily loading an IPv6-only site over TCP is what caught it. From here the two cases look
* identical, so the finding now says what was observed and names both explanations rather than
* picking one.
*
* It is worth reporting either way: filtered ICMPv6 breaks Path MTU Discovery, which is its
* own fault even when IPv6 works.
*/
/**
* A global IPv6 address with no default route.
*
* This is the structural version of the same complaint, and it is worth far more than the
* ICMP one because it admits no other explanation: the device has an address it cannot route
* with. Nothing is filtered, nothing is inferred the routing table says so directly, and it
* is already in the link snapshot.
*
* Not always a fault. A VPN that installs host routes to specific destinations produces
* exactly this shape on purpose, and it works. What makes it worth reporting either way is
* that applications cannot tell: having a global address, they will try IPv6 first and stall
* for every destination the routes do not cover.
*/
/**
* An IPv6 default route with no global address to use it from the mirror of
* [V6_NO_DEFAULT_ROUTE], and the more common misconfiguration of the two.
*
* The router is sending RAs that name it as a default gateway, but SLAAC produced no address:
* no prefix information option, or a prefix without the autonomous flag, or DHCPv6-only
* addressing the device did not complete. The network is announcing IPv6 service it does not
* actually deliver.
*
* This is worth flagging above the ICMP signal because it is both certain and consequential.
* Hosts see router advertisements, believe IPv6 is available, and pay a connection-attempt
* timeout on every dual-stack destination before falling back to IPv4 the classic "the
* internet feels slow" complaint with no packet loss anywhere to explain it.
*/
val V6_ROUTE_WITHOUT_ADDRESS = FindingSpec(
"v6.route_without_address", Category.IPV6, Severity.MEDIUM,
"The network advertises an IPv6 default route but the device has no global IPv6 address.",
rulesOut = "A working IPv6 setup: SLAAC did not produce a usable address on this link.",
)
/**
* A VPN prevented the underlying networks from being measured.
*
* Reported rather than worked around: Android refuses `Network.bindSocket()` on the networks
* beneath a VPN precisely so apps cannot leak around the tunnel, and that is correct
* behaviour. What is not acceptable is a run that quietly measures nothing and calls the
* result healthy, so this says plainly which networks went unmeasured and why.
*/
/**
* The network's DNS server answers, but this device cannot resolve through it.
*
* Worth separating from every other DNS failure because the remedy is somewhere else entirely.
* A name that will not resolve looks identical to a user whatever the cause, and the two causes
* pull in opposite directions: a server that does not answer means the network is broken and
* the router is the thing to examine, while a server that answers a direct query on a device
* that still cannot resolve means the platform resolver has wedged fixed by toggling wifi,
* and nothing to do with the network at all.
*
* Proven rather than inferred: the probe sends its own UDP query, bypassing the component under
* suspicion, and compares that against what the platform returns for the same name.
*/
/**
* The network hands out a search domain its DNS server will not answer for.
*
* A resolver appends search domains to lookups, so every name a client asks about can stall on
* a domain the server ignores. The failure mode is silence rather than a negative answer, and
* silence is indistinguishable from packet loss: clients retry instead of moving on, and some
* give up on the lookup entirely. That makes it look like the device is broken when the
* network is.
*
* Whether it bites depends on the resolver some try the bare name first and never notice
* which is why two devices on the same network can disagree about whether DNS works.
*/
val DNS_SEARCH_DOMAIN_UNANSWERED = FindingSpec(
"dns.search_domain_unanswered", Category.DNS, Severity.HIGH,
"The network advertises a DNS search domain that its own server does not answer for.",
rulesOut = "A fault on this device: the same server answers ordinary names normally.",
)
val DNS_SYSTEM_RESOLVER_BROKEN = FindingSpec(
"dns.system_resolver_broken", Category.DNS, Severity.HIGH,
"The network's DNS server answers, but this device cannot resolve names through it.",
rulesOut = "A network fault: the server replied to a query sent from this device.",
)
val MEASUREMENT_VPN_CONSTRAINED = FindingSpec(
"measurement.vpn_constrained", Category.CONNECTIVITY, Severity.INFO,
"A VPN was active, so the networks underneath it could not be measured.",
rulesOut = "Nothing — this run says little about the underlying network either way.",
)
val V6_NO_DEFAULT_ROUTE = FindingSpec(
"v6.no_default_route", Category.IPV6, Severity.MEDIUM,
"The device has a global IPv6 address but no IPv6 default route.",
rulesOut = "Guesswork: this is read from the routing table, not inferred from silence.",
)
val V6_NO_ICMP_REPLY = FindingSpec(
"v6.no_icmp_reply", Category.IPV6, Severity.LOW,
"IPv6 is configured but ICMPv6 echo gets no reply.",
rulesOut = "Nothing on its own: IPv6 may work fine with ICMP filtered.",
)
/**
* IPv6 is advertised and does not work the claim `v6.broken` originally made on ICMP
* silence alone, now reinstated because it can finally be backed: it is only emitted when a
* real IPv6 TCP connection (v6.brokenness) failed on the same network whose ICMPv6 went
* unanswered. Two independent transports failing on a network that advertises IPv6 is what
* "broken" actually means; either signal alone still gets [V6_NO_ICMP_REPLY].
*/
val V6_BROKEN = FindingSpec(
"v6.broken", Category.IPV6, Severity.HIGH,
"IPv6 is advertised on this network but carries no traffic.",
rulesOut = "ICMP filtering as the benign explanation: a TCP connection over IPv6 failed too.",
)
/**
* INFO deliberately, and it needs to stay that way.
*
* Most networks still do not offer IPv6, and that is not a fault. Reporting it as a warning
* lights a yellow verdict on a perfectly healthy network, which teaches people to ignore the
* light the one thing a diagnostic must never do.
*/
val V6_NOT_OFFERED = FindingSpec(
"v6.not_offered", Category.IPV6, Severity.INFO,
"This network does not offer IPv6.",
)
/** Every registered finding, in declaration order. */
val all: List<FindingSpec> = listOf(
UDP_UNREACHABLE, UDP_UNREACHABLE_UPSTREAM, UDP_LOSS, LOSS_UPSTREAM, LOSS_DOWNSTREAM,
DOWNSTREAM_BLOCKED, DOWNSTREAM_REORDER, CAPTIVE_PORTAL, NO_INTERNET, LINK_FLAPPING,
MTU_REDUCED_DOWNSTREAM, MTU_DOWNSTREAM_BLACKHOLE, FRAGMENTS_BLOCKED,
FRAGMENT_REORDER_SENSITIVE,
NAT_UDP_REBINDING, NAT_SYMMETRIC,
THROUGHPUT_NO_DELIVERY, THROUGHPUT_BELOW_OFFERED,
DNS_ANSWER_REWRITTEN, DNS_AUTHORITATIVE_UNREACHABLE,
DNS_SEARCH_DOMAIN_UNANSWERED, DNS_SYSTEM_RESOLVER_BROKEN, MEASUREMENT_VPN_CONSTRAINED,
V6_NO_DEFAULT_ROUTE, V6_ROUTE_WITHOUT_ADDRESS, V6_NO_ICMP_REPLY, V6_BROKEN, V6_NOT_OFFERED,
)
private val byCode: Map<String, FindingSpec> = all.associateBy { it.code }
fun byCode(code: String): FindingSpec? = byCode[code]
}
@@ -18,6 +18,39 @@ data class Network(
val wifi: Wifi? = null, val wifi: Wifi? = null,
val cellular: Cellular? = null, val cellular: Cellular? = null,
val changes: List<NetworkChange> = emptyList(), val changes: List<NetworkChange> = emptyList(),
@SerialName("system_verdict") val systemVerdict: SystemVerdict? = null,
/**
* Whether an ordinary app may send on this network at all.
*
* False for the carrier's special-purpose networks IMS/VoLTE, MMS, XCAP which appear
* beside the real ones in Android's list and carry neither `INTERNET` nor `NOT_RESTRICTED`.
* Binding to those needs `CONNECTIVITY_USE_RESTRICTED_NETWORKS`, a privileged permission no
* normal app can hold, so the refusal is permanent and says nothing about the network's
* health. Recorded rather than hidden: a reader seeing an interface with no measurements
* against it deserves to know the OS forbade them, instead of concluding the link is dead.
*/
@SerialName("app_usable") val appUsable: Boolean? = null,
)
/**
* What Android itself concluded about a network, as opposed to what we measured.
*
* Recorded because it is the verdict the user can see the "no internet" warning in the status
* bar and because it is free: the platform has already done the work by the time a run starts.
*
* Its real value is disagreement. When Android says a network is unusable and our own probes reach
* the internet regardless, the fault is in the device rather than the network, and that distinction
* is the difference between "fix your router" and "toggle your wifi". Neither number alone can say
* that; only the two together.
*/
@Serializable
data class SystemVerdict(
/** Android's own connectivity check passed. Null when the platform did not say. */
val validated: Boolean? = null,
/** Android believes a captive portal is intercepting this network. */
@SerialName("captive_portal") val captivePortal: Boolean? = null,
/** Some traffic works and some does not — Android's own hedge. */
@SerialName("partial_connectivity") val partialConnectivity: Boolean? = null,
) )
@Serializable @Serializable
@@ -111,3 +144,38 @@ data class NetworkChange(
val kind: String, // lost | gained | link_changed val kind: String, // lost | gained | link_changed
val detail: JsonObject? = null, val detail: JsonObject? = null,
) )
/**
* What a network's `changes[]` add up to.
*
* Lives beside the type rather than in the collector that produces it because two independent
* consumers ask the same question the watcher, computing its metrics, and the run engine,
* deciding whether to emit `connectivity.link_flapping` and a document whose metric and finding
* disagreed about how many times the link dropped would be worse than one reporting neither.
*/
object NetworkChanges {
const val LOST = "lost"
const val GAINED = "gained"
const val LINK_CHANGED = "link_changed"
/**
* Completed drop-and-return cycles: a `lost` with a later `gained` on the same network.
*
* A cycle has to *complete*. A link that goes away at minute four and is still gone when the
* window closes was not flapping it was switched off, or the device was carried out of
* range, and calling that the same fault would put a phone in a lift beside a failing access
* point.
*/
fun flapCycles(kinds: List<String>): Int {
var cycles = 0
var down = false
for (k in kinds) {
if (k == LOST) down = true
else if (k == GAINED && down) { cycles++; down = false }
}
return cycles
}
fun flapCyclesOf(changes: List<NetworkChange>): Int = flapCycles(changes.map { it.kind })
}
@@ -35,6 +35,9 @@ data class CategorySummary(
* (critical|high red, medium|low yellow, info/none green). * (critical|high red, medium|low yellow, info/none green).
* - A category is `inconclusive` when > 50% of its tests are failed/unsupported. * - A category is `inconclusive` when > 50% of its tests are failed/unsupported.
* - Overall = the worst category light; `inconclusive` only when ALL categories are. * - Overall = the worst category light; `inconclusive` only when ALL categories are.
* - A run whose per-network probing was blocked is `inconclusive` outright, whatever the
* categories say. The lights describe what the tests found; when the OS refused to let the
* tests run, a green light would describe nothing at all.
* *
* The mapping test-type category comes from [TestType.category]. Only categories that have * The mapping test-type category comes from [TestType.category]. Only categories that have
* findings or tests appear in the summary. * findings or tests appear in the summary.
@@ -44,7 +47,10 @@ object Verdicts {
private fun isInconclusiveTest(s: TestStatus) = private fun isInconclusiveTest(s: TestStatus) =
s == TestStatus.FAILED || s == TestStatus.UNSUPPORTED s == TestStatus.FAILED || s == TestStatus.UNSUPPORTED
fun derive(tests: List<Test>, findings: List<Finding>): Summary { fun derive(tests: List<Test>, findings: List<Finding>): Summary =
derive(tests, findings, Constraints())
fun derive(tests: List<Test>, findings: List<Finding>, constraints: Constraints): Summary {
val testsByCat = tests.groupBy { TestType.category(it.type) } val testsByCat = tests.groupBy { TestType.category(it.type) }
val findingsByCat = findings.groupBy { it.category } val findingsByCat = findings.groupBy { it.category }
val categories = (testsByCat.keys + findingsByCat.keys) val categories = (testsByCat.keys + findingsByCat.keys)
@@ -72,7 +78,14 @@ object Verdicts {
) )
} }
val overall = deriveOverall(perCat.values) // A run that could not measure the networks it was asked about has not found them
// healthy; it has found out nothing. Reporting that as green is the single most
// misleading thing this function could do, so the constraint outranks the lights.
val overall = if (constraints.perNetworkBlocked) {
Verdict.INCONCLUSIVE
} else {
deriveOverall(perCat.values)
}
return Summary(overall = overall, categories = perCat) return Summary(overall = overall, categories = perCat)
} }
@@ -69,12 +69,16 @@ object TestType {
const val TRACEROUTE_ICMP6 = "traceroute.icmp6" const val TRACEROUTE_ICMP6 = "traceroute.icmp6"
// train // train
const val TRAIN_UDP_UPDOWN = "train.udp_updown" const val TRAIN_UDP_UPDOWN = "train.udp_updown"
/** Server-to-client train under a §3.4 grant: the direction a round trip cannot separate. */
const val TRAIN_UDP_DOWNSTREAM = "train.udp_downstream"
// mtu // mtu
const val MTU_PMTUD_UP = "mtu.pmtud_up" const val MTU_PMTUD_UP = "mtu.pmtud_up"
const val MTU_PMTUD_DOWN = "mtu.pmtud_down" const val MTU_PMTUD_DOWN = "mtu.pmtud_down"
const val MTU_BLACKHOLE = "mtu.blackhole" const val MTU_BLACKHOLE = "mtu.blackhole"
const val MTU_MSS_OBSERVED = "mtu.mss_observed" const val MTU_MSS_OBSERVED = "mtu.mss_observed"
const val MTU_FRAG_DELIVERY = "mtu.frag_delivery" const val MTU_FRAG_DELIVERY = "mtu.frag_delivery"
/** Whether fragments survive arriving out of order, not merely whether they survive. */
const val MTU_FRAG_ORDERING = "mtu.frag_ordering"
// nat // nat
const val NAT_STUN_5780 = "nat.stun_5780" const val NAT_STUN_5780 = "nat.stun_5780"
const val NAT_MAPPING_LIFETIME_UDP = "nat.mapping_lifetime_udp" const val NAT_MAPPING_LIFETIME_UDP = "nat.mapping_lifetime_udp"
@@ -85,6 +89,13 @@ object TestType {
// dns // dns
const val DNS_RESOLVER_INVENTORY = "dns.resolver_inventory" const val DNS_RESOLVER_INVENTORY = "dns.resolver_inventory"
const val DNS_CANARY = "dns.canary" const val DNS_CANARY = "dns.canary"
/**
* Does this device's own resolver work, as distinct from the network's DNS.
*
* Registry addition, v1.1. Kept apart from [DNS_CANARY], which asks whether answers are being
* tampered with; this asks whether answers arrive at all, and where the failure sits.
*/
const val DNS_RESOLVER = "dns.resolver"
const val DNS_INTERCEPTION = "dns.interception" const val DNS_INTERCEPTION = "dns.interception"
const val DNS_TTL_INTEGRITY = "dns.ttl_integrity" const val DNS_TTL_INTEGRITY = "dns.ttl_integrity"
const val DNS_ANSWER_INTEGRITY = "dns.answer_integrity" const val DNS_ANSWER_INTEGRITY = "dns.answer_integrity"
@@ -122,6 +133,16 @@ object TestType {
const val LOCAL_MDNS_INVENTORY = "local.mdns_inventory" const val LOCAL_MDNS_INVENTORY = "local.mdns_inventory"
const val LOCAL_SSDP_INVENTORY = "local.ssdp_inventory" const val LOCAL_SSDP_INVENTORY = "local.ssdp_inventory"
const val LOCAL_LLMNR_INVENTORY = "local.llmnr_inventory" const val LOCAL_LLMNR_INVENTORY = "local.llmnr_inventory"
/**
* WS-Discovery (UDP 3702) and NetBIOS name service (UDP 137). Registry additions, v1.2.
*
* Both are passive: the traffic is broadcast to the segment whether or not anyone asks, so
* listening is the whole measurement. They earn their own ids rather than folding into
* [LOCAL_SSDP_INVENTORY] because what they imply differs WS-Discovery inventories printers
* and cameras, while NetBIOS/LLMNR chatter is a security finding in its own right.
*/
const val LOCAL_WSD_INVENTORY = "local.wsd_inventory"
const val LOCAL_NETBIOS_INVENTORY = "local.netbios_inventory"
const val LOCAL_GATEWAY_SERVICES = "local.gateway_services" const val LOCAL_GATEWAY_SERVICES = "local.gateway_services"
const val LOCAL_NTP = "local.ntp" const val LOCAL_NTP = "local.ntp"
// peer // peer
@@ -0,0 +1,73 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.measurement
/**
* The two ways a network can be half-configured for IPv6, read from the link snapshot.
*
* Pure model logic rather than something a ViewModel does, because "is this network's IPv6
* broken, and in which direction" is exactly the kind of judgement that should be checkable
* against a captured routing table without a phone in the loop.
*/
object V6Analysis {
/** Linux tunnel interfaces: WireGuard/Netbird (tun*, wg*), plus the usual VPN names. */
private val TUNNEL_IFACE = Regex("""^(tun|tap|wg|ppp|ipsec|utun)\d*$""")
/** What one network's IPv6 configuration looks like. */
data class Shape(
val iface: String,
/** A global address with no ::/0 route: an address the device cannot route with. */
val addressWithoutRoute: Boolean,
/** A ::/0 route with no global address: a route the device cannot source from. */
val routeWithoutAddress: Boolean,
/** The routes belong to a tunnel, so a partial view of IPv6 is likely deliberate. */
val tunnel: Boolean,
)
/**
* Classifies each network's IPv6 configuration.
*
* Both shapes are read straight from the link snapshot rather than inferred from silence, so
* unlike an ICMP signal there is no competing explanation for what was observed and both
* matter for the same reason: an application cannot tell in advance, so it tries IPv6 first
* and waits.
*
* They differ in what they mean. An address with no route is what a VPN installing host routes
* to specific destinations produces on purpose, and it works; calling that a fault would be the
* "lack of IPv6 is a yellow condition" mistake in a new costume, so a tunnel downgrades it to
* information. A route with no address is the opposite: the router advertised itself as a
* default gateway but SLAAC produced nothing usable, so the network is announcing IPv6 service
* it does not deliver. That one is a real misconfiguration however it arises.
*/
fun classify(networks: List<Network>): List<Shape> = networks.map { n ->
val globalV6 = n.link.addresses.any { isGlobalV6(it.addr) }
val v6Routes = n.link.routes.filter { it.dst.contains(':') }
val hasDefault = v6Routes.any { it.dst == "::/0" }
Shape(
iface = n.iface ?: v6Routes.firstOrNull()?.iface.orEmpty(),
addressWithoutRoute = globalV6 && !hasDefault,
routeWithoutAddress = hasDefault && !globalV6,
// Android labels the transport itself, which beats guessing from a name; the regex
// stays as a backstop for tunnels Android does not own (a userspace WireGuard, say,
// or anything seen through the shell tier).
tunnel = n.transport == Transport.VPN ||
v6Routes.any { TUNNEL_IFACE.containsMatchIn(it.iface.orEmpty()) },
)
}
/**
* Whether an address is IPv6 and usable as a source for off-link traffic.
*
* ULAs count. A ULA is not globally routable, but it is a global-*scope* address the stack
* will happily select as a source, which is the property that matters here an overlay
* network handing out fc00::/7 addresses is providing working IPv6 to the destinations it
* carries, and treating that as "no address" would misreport every VPN as broken.
*/
private fun isGlobalV6(addr: String): Boolean {
if (!addr.contains(':')) return false
val a = addr.substringBefore('%').lowercase() // strip any zone index
return !a.startsWith("fe80") && a != "::1" && a != "::"
}
}
@@ -0,0 +1,133 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.measurement
import java.io.File
import kotlin.test.Test
import kotlin.test.assertEquals
import kotlin.test.assertTrue
import kotlin.test.fail
/**
* Keeps the finding registry honest.
*
* The interesting test is the last one: it reads `docs/findings-registry.md` and fails when the
* document and the code disagree. Documentation that drifts from its implementation is worse than
* none, because it still looks authoritative and a finding registry is precisely the artifact
* other people build tooling against.
*/
class FindingRegistryTest {
@Test
fun codesAreUnique() {
val dupes = FindingRegistry.all.groupBy { it.code }.filterValues { it.size > 1 }.keys
assertTrue(dupes.isEmpty(), "duplicate finding codes: $dupes")
}
@Test
fun everyDeclaredSpecIsInTheAllList() {
// Reflection over the object's properties: a spec that is declared but left out of `all`
// is invisible to the doc check and to any consumer enumerating the registry.
val declared = FindingRegistry::class.java.declaredMethods
.filter { it.parameterCount == 0 && it.returnType == FindingSpec::class.java }
.mapNotNull { runCatching { it.invoke(FindingRegistry) as FindingSpec }.getOrNull() }
.map { it.code }
.toSet()
val listed = FindingRegistry.all.map { it.code }.toSet()
assertEquals(declared, listed, "declared specs and the `all` list disagree")
}
// The prefix decides the category, and the category decides which verdict light the finding
// rolls up into. A code whose prefix disagrees with its category silently moves a fault to a
// different light — the exact bug that got two codes renamed out of nat.*.
@Test
fun everyPrefixMatchesItsCategory() {
for (spec in FindingRegistry.all) {
val fromPrefix = TestType.category(spec.code)
assertEquals(
fromPrefix, spec.category,
"${spec.code} is declared as ${spec.category} but its prefix maps to $fromPrefix",
)
}
}
@Test
fun codesFollowTheNamingConvention() {
val shape = Regex("^[a-z0-9]+\\.[a-z0-9_]+$")
for (spec in FindingRegistry.all) {
assertTrue(shape.matches(spec.code), "malformed code: ${spec.code}")
assertTrue(spec.meaning.isNotBlank(), "${spec.code} has no meaning")
assertTrue(
spec.meaning.trimEnd().endsWith("."),
"${spec.code}'s meaning should be a sentence: '${spec.meaning}'",
)
}
}
// Two near-identical codes are how one fault ends up split across two dashboards. This is a
// blunt check — it will not catch every synonym — but it catches the shape that already
// happened: the same words in a different order.
@Test
fun noTwoCodesAreAnagramsOfEachOther() {
val normalised = FindingRegistry.all.associate { spec ->
spec.code to spec.code.substringAfter('.').split('_').sorted().joinToString("_")
}
val clashes = normalised.entries.groupBy { it.value }.filterValues { it.size > 1 }
if (clashes.isNotEmpty()) {
fail("codes differing only in word order: ${clashes.values.map { g -> g.map { it.key } }}")
}
}
@Test
fun theDocumentAndTheRegistryAgree() {
val doc = findDoc() ?: run {
println("findings-registry.md not found from ${File(".").absolutePath} — skipping")
return
}
val text = doc.readText()
// Only table rows count as "documented". Prose may legitimately mention a code that no
// longer exists — the rules section explains why two were merged — and treating that as
// a registry entry would force the document to forget its own history.
val documented = text.lines()
.filter { it.trimStart().startsWith("|") }
.flatMap { row -> Regex("`([a-z0-9]+\\.[a-z0-9_]+)`").findAll(row).map { it.groupValues[1] } }
.toSet()
val registered = FindingRegistry.all.map { it.code }.toSet()
val missingFromDoc = registered - documented
val missingFromCode = documented - registered
assertTrue(
missingFromDoc.isEmpty(),
"these codes exist in FindingRegistry but not in docs/findings-registry.md: $missingFromDoc",
)
assertTrue(
missingFromCode.isEmpty(),
"docs/findings-registry.md documents codes that no longer exist: $missingFromCode",
)
// And the severities must match, or the document is describing a different system.
for (spec in FindingRegistry.all) {
val row = text.lines().firstOrNull {
it.trimStart().startsWith("|") && it.contains("`${spec.code}`")
} ?: continue
val severity = spec.severity.name.lowercase()
assertTrue(
row.contains("| $severity |"),
"${spec.code} is ${severity} in code but the doc row says otherwise: $row",
)
}
}
/** Walks up from the test's working directory to find the repo's docs/ folder. */
private fun findDoc(): File? {
var dir: File? = File(".").absoluteFile
repeat(6) {
val candidate = File(dir, "docs/findings-registry.md")
if (candidate.isFile) return candidate
dir = dir?.parentFile
}
return null
}
}
@@ -0,0 +1,65 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.measurement
import kotlin.test.Test
import kotlin.test.assertEquals
/**
* Pins the counting behind `connectivity.link_flapping`.
*
* The finding claims a link went away and came back, and its severity escalates on repetition, so
* this is arithmetic a person reading a report will act on. The cases that matter are the ones
* where the naive count is wrong: a link still down when the window closed, and a run that started
* while the link was already gone.
*/
class NetworkChangesTest {
@Test
fun aQuietWindowHasNoCycles() {
assertEquals(0, NetworkChanges.flapCycles(emptyList()))
assertEquals(0, NetworkChanges.flapCycles(listOf("link_changed", "link_changed")))
}
@Test
fun oneDropAndReturnIsOneCycle() {
assertEquals(1, NetworkChanges.flapCycles(listOf("lost", "gained")))
assertEquals(
1,
NetworkChanges.flapCycles(listOf("link_changed", "lost", "link_changed", "gained")),
)
}
@Test
fun repeatedDropsCountSeparately() {
assertEquals(3, NetworkChanges.flapCycles(listOf("lost", "gained", "lost", "gained", "lost", "gained")))
}
// A link that is still down when the run ends was not flapping — it was switched off, or the
// device left its range. Counting that as a cycle would put a phone in a lift beside a failing
// access point.
@Test
fun aDropThatNeverReturnsIsNotACycle() {
assertEquals(0, NetworkChanges.flapCycles(listOf("lost")))
assertEquals(1, NetworkChanges.flapCycles(listOf("lost", "gained", "lost")))
}
// The mirror case: the window opened while the network was already gone, so its return is the
// first thing seen. Nothing was watched dropping, so nothing is claimed.
@Test
fun aReturnWithNoObservedDropIsNotACycle() {
assertEquals(0, NetworkChanges.flapCycles(listOf("gained")))
assertEquals(0, NetworkChanges.flapCycles(listOf("gained", "link_changed")))
}
@Test
fun theChangeOverloadAgreesWithTheKindsOverload() {
val changes = listOf(
NetworkChange(atMonoNs = 1, kind = NetworkChanges.LOST),
NetworkChange(atMonoNs = 2, kind = NetworkChanges.GAINED),
NetworkChange(atMonoNs = 3, kind = NetworkChanges.LINK_CHANGED),
)
assertEquals(1, NetworkChanges.flapCyclesOf(changes))
}
}
@@ -0,0 +1,125 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.measurement
import kotlin.test.Test
import kotlin.test.assertEquals
import kotlin.test.assertFalse
import kotlin.test.assertTrue
/**
* The fixtures here are a real device's routing table, transcribed from `dumpsys connectivity`
* on a OnePlus 15 with a Netbird tunnel up: wifi advertising a default route it cannot source
* from, cellular working properly, and a VPN carrying host routes to two destinations.
*
* Using a captured table rather than invented ones matters, because the bug this guards against
* is not "the boolean logic is wrong" it is "the shapes I imagined are not the shapes real
* networks produce".
*/
class V6AnalysisTest {
private fun net(
id: String,
transport: Transport,
iface: String,
addrs: List<String>,
routes: List<Pair<String, String>>,
) = Network(
id = id,
transport = transport,
iface = iface,
link = Link(
addresses = addrs.map { Address(addr = it.substringBefore('/'), prefixLen = 64) },
routes = routes.map { (dst, dev) -> Route(dst = dst, iface = dev) },
),
)
/** wlan0: an IPv6 default route via a link-local gateway, but SLAAC produced no address. */
private val wifi = net(
"w", Transport.WIFI, "wlan0",
addrs = listOf("fe80::bcf6:edff:fe67:b139", "10.13.102.122"),
routes = listOf(
"fe80::/64" to "wlan0",
"::/0" to "wlan0",
"0.0.0.0/0" to "wlan0",
),
)
/** rmnet_data1: a properly configured cellular link — global address and a default route. */
private val cellular = net(
"c", Transport.CELLULAR, "rmnet_data1",
addrs = listOf("2001:4bb8:46a:e724:289d:87ff:feb6:ebd3"),
routes = listOf("::/0" to "rmnet_data1", "2001:4bb8:46a:e724::/64" to "rmnet_data1"),
)
/** tun1: Netbird, with a ULA and host routes to exactly two destinations. */
private val vpn = net(
"v", Transport.VPN, "tun1",
addrs = listOf("100.64.158.131", "fdfd:c4fe:c4fe:c4fe:1f3c:98a0:dd66:ac7"),
routes = listOf(
"2001:1ad0:c4fe:6767::2/128" to "tun1",
"2001:1ad0:c4fe:a::136/128" to "tun1",
"fdfd:c4fe:c4fe:c4fe::/64" to "tun1",
),
)
@Test
fun `wifi advertising a route it cannot source from is reported`() {
val s = V6Analysis.classify(listOf(wifi)).single()
assertTrue(s.routeWithoutAddress, "::/0 with only a link-local address is the RA-without-SLAAC case")
assertFalse(s.addressWithoutRoute)
assertFalse(s.tunnel, "wifi is not a tunnel")
assertEquals("wlan0", s.iface)
}
@Test
fun `a properly configured link produces no finding`() {
val s = V6Analysis.classify(listOf(cellular)).single()
assertFalse(s.routeWithoutAddress)
assertFalse(s.addressWithoutRoute)
}
@Test
fun `a tunnel with host routes is deliberate, not broken`() {
val s = V6Analysis.classify(listOf(vpn)).single()
assertTrue(s.addressWithoutRoute, "a ULA and no ::/0 is an address with nothing to route it")
assertTrue(s.tunnel, "so it must be reported as information, not as a fault")
assertFalse(s.routeWithoutAddress)
}
@Test
fun `each network is judged on its own`() {
// The whole point of per-network classification: "IPv6 is broken" is useless advice when
// wifi is the broken one and cellular is fine.
val shapes = V6Analysis.classify(listOf(wifi, cellular, vpn)).associateBy { it.iface }
assertTrue(shapes.getValue("wlan0").routeWithoutAddress)
assertFalse(shapes.getValue("rmnet_data1").routeWithoutAddress)
assertFalse(shapes.getValue("rmnet_data1").addressWithoutRoute)
assertTrue(shapes.getValue("tun1").addressWithoutRoute)
}
@Test
fun `a link-local-only network with no v6 route says nothing either way`() {
// Plain IPv4-only wifi: no IPv6 offered at all. That is v6.not_offered's business, and
// reporting it here as well would double up on a network that is merely legacy, not broken.
val v4only = net(
"4", Transport.WIFI, "wlan0",
addrs = listOf("fe80::1", "192.168.1.5"),
routes = listOf("0.0.0.0/0" to "wlan0"),
)
val s = V6Analysis.classify(listOf(v4only)).single()
assertFalse(s.routeWithoutAddress)
assertFalse(s.addressWithoutRoute)
}
@Test
fun `a zone index does not hide a link-local address`() {
val zoned = net(
"z", Transport.WIFI, "wlan0",
addrs = listOf("fe80::1%wlan0"),
routes = listOf("::/0" to "wlan0"),
)
assertTrue(V6Analysis.classify(listOf(zoned)).single().routeWithoutAddress)
}
}
+27
View File
@@ -0,0 +1,27 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
plugins {
alias(libs.plugins.kotlin.jvm)
alias(libs.plugins.kotlin.serialization)
}
// The anonymizer (measurement-schema.md §8). Pure Kotlin/JVM and deliberately
// dependency-free beyond JSON: it must be trivially auditable, because a bug
// here leaks a user's network onto someone else's server.
dependencies {
implementation(libs.kotlinx.serialization.json)
testImplementation(kotlin("test"))
}
kotlin {
jvmToolchain(21)
compilerOptions { jvmTarget.set(org.jetbrains.kotlin.gradle.dsl.JvmTarget.JVM_17) }
}
java { sourceCompatibility = JavaVersion.VERSION_17; targetCompatibility = JavaVersion.VERSION_17 }
tasks.test {
useJUnitPlatform()
// Opt-in: point this at a captured run to check the anonymizer against real data.
System.getenv("ECHOLOT_REAL_RUN")?.let { environment("ECHOLOT_REAL_RUN", it) }
}
@@ -0,0 +1,327 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
// Package privacy implements the anonymization contract of measurement-schema.md §8.
//
// The threat model is specific. An engineer running their own server wants the full document —
// SSIDs and MACs are what make a run useful a week later. Someone measuring against a stranger's
// server wants the numbers to survive and the identifiers not to. So this is a *transform*, not a
// filter: the output is still a valid measurement document with the same tests, metrics and
// findings; only the identifying scalars change, and consistently, so "same SSID as last run" is
// still answerable from pseudonyms alone.
//
// Two properties are load-bearing and are what the tests pin:
// - Consistency within a document: one input value always maps to one pseudonym, so
// correlations inside a run survive.
// - No consistency *across* documents unless the user asks for it: the salt is per-run by
// default, so pseudonyms cannot be used to track a device between uploads. A stable salt is
// opt-in (`Salt.stable`) for people diffing their own history on their own server.
package app.echo_lot.privacy
import kotlinx.serialization.json.*
import java.security.MessageDigest
import java.util.Locale
/** How much to strip. Ordered: FULL < BALANCED < STRICT. Wire values match the server's. */
enum class PrivacyLevel(val wire: String) {
/** Nothing removed. The right choice for your own server. */
FULL("full"),
/**
* Identifiers pseudonymized, neighbour inventory dropped. Topology and timing survive:
* you can still see that the gateway is a MikroTik at a /24 boundary with 3 % loss, but not
* which MikroTik, on which SSID, next to whose Chromecast.
*/
BALANCED("balanced"),
/**
* Numbers only: tests keep their metrics and status, evidence is dropped, findings keep their
* codes and severities but lose descriptions (which quote real names). What is left cannot
* identify a network, and is still enough for aggregate "how common is this fault" work.
*/
STRICT("strict");
companion object {
fun fromWire(s: String?): PrivacyLevel =
entries.firstOrNull { it.wire == s?.lowercase(Locale.ROOT) } ?: FULL
/** The stricter of two levels — used to honour a server's minimum. */
fun max(a: PrivacyLevel, b: PrivacyLevel): PrivacyLevel = if (a.ordinal >= b.ordinal) a else b
}
}
/**
* The pseudonymization salt. Per-run by default: a fresh random salt means the same SSID uploaded
* twice yields two different pseudonyms, so an upload endpoint cannot link runs to a device.
* A stable salt trades that away for cross-run diffing and is only appropriate on a server you
* own the app makes that an explicit choice, not a default.
*/
class Salt private constructor(internal val bytes: ByteArray, val stable: Boolean) {
companion object {
fun perRun(random: ByteArray): Salt = Salt(random.copyOf(), stable = false)
fun stable(secret: ByteArray): Salt = Salt(secret.copyOf(), stable = true)
}
}
/**
* Transforms a measurement document to [level].
*
* Field classification is by JSON key name, because the schema names things consistently
* (`ssid`, `bssid`, `mac`, `ip4`, `ip6`, `fqdn`, ) and a name-driven pass is auditable by
* reading one table. Anything unrecognized is treated as identifying when it is a string inside
* a known-sensitive container, and left alone otherwise see [Classification].
*/
class Anonymizer(private val level: PrivacyLevel, private val salt: Salt) {
private val cache = HashMap<String, String>()
fun anonymize(doc: JsonObject): JsonObject {
if (level == PrivacyLevel.FULL) return stamp(doc)
val walked = walkObject(doc, path = emptyList())
val out = if (level == PrivacyLevel.STRICT) strip(walked) else walked
return stamp(out)
}
/** Records what was done, so a reader of the archived/uploaded document is never guessing. */
private fun stamp(doc: JsonObject): JsonObject {
val run = doc["run"]?.jsonObject ?: return doc
val privacy = buildJsonObject {
put("anonymization", level.wire)
put("salt", if (salt.stable) "stable" else "per_run")
}
return JsonObject(doc + ("run" to JsonObject(run + ("privacy" to privacy))))
}
// ---- the tree walk -------------------------------------------------------------------
private fun walkObject(obj: JsonObject, path: List<String>): JsonObject = buildJsonObject {
for ((k, v) in obj) {
val childPath = path + k
when {
Classification.dropAtBalanced(childPath) -> Unit // omit entirely
else -> put(k, walk(k, v, childPath))
}
}
}
private fun walk(key: String, v: JsonElement, path: List<String>): JsonElement = when (v) {
is JsonObject -> walkObject(v, path)
is JsonArray -> JsonArray(v.map { walk(key, it, path) })
is JsonPrimitive ->
if (v.isString) {
// Name first (it is precise), then shape (it is exhaustive). A field nobody
// classified must not be a field that leaks.
val type = Classification.typeOf(key, path) ?: Classification.inferFromValue(v.content)
JsonPrimitive(transform(type, v.content))
} else {
v
}
}
private fun transform(type: LogicalType?, value: String): String = when (type) {
// Unclassified strings still get their *embedded* identifiers scrubbed. A whole-value
// check cannot see them: raw shell output is one long string that is neither a MAC nor an
// address, so it sailed through both the name table and the shape check carrying every
// MAC on the user's LAN.
null -> scrubEmbedded(value)
LogicalType.SSID -> pseudo("ssid", value) { "net-" + it.take(6) }
LogicalType.MAC, LogicalType.BSSID -> macPreservingOui(value)
LogicalType.IP4 -> ip4(value)
LogicalType.IP6 -> ip6(value)
LogicalType.FQDN -> fqdn(value)
LogicalType.OPAQUE_ID -> "redacted"
LogicalType.FREETEXT -> "[removed: may contain identifying text]"
}
/**
* Replaces addresses and MACs found *inside* a longer string.
*
* Shizuku probes embed raw command output verbatim `ip neigh`, `ip route`, `dumpsys` which
* is genuinely valuable evidence and also a complete inventory of every device on the user's
* network, with hardware addresses. measurement-schema.md §9 flagged these as "hard to
* anonymize" and proposed dropping them from exports.
*
* Scrubbing beats dropping: the output stays readable and auditable you can still see the
* shape of the neighbour table and how many hosts there were while the identifiers become
* the same pseudonyms used everywhere else in the document. So a MAC appearing both in a
* parsed field and in a raw dump still reads as one device.
*
* Only addresses and MACs are touched, for the same reason as [Classification.inferFromValue]:
* they are the patterns that cannot be mistaken for something else in free text.
*/
private fun scrubEmbedded(value: String): String {
// Cheap bail-out: the overwhelming majority of strings are short and contain neither.
if (value.length < 7 || (!value.contains(':') && !value.contains('.'))) return value
// One pass, not three. Sequential passes re-process their own output: after a MAC became
// 78:9a:18:xx:yy:zz the IPv6 pattern matched it — six hex groups separated by colons is
// exactly an address — and mangled the vendor prefix that the MAC rule had just taken
// care to preserve. Ordered alternation resolves each position once, MAC first.
return EMBEDDED.replace(value) { m ->
when {
m.groups[1] != null -> macPreservingOui(m.value)
m.groups[2] != null -> ip6(m.value)
else -> ip4(m.value)
}
}
}
// ---- per-type transforms -------------------------------------------------------------
/**
* Keeps the OUI (which vendor) and pseudonymizes the NIC part (which unit). Vendor is the
* diagnostically valuable half "the RA comes from a MikroTik" survives, "…from THAT
* MikroTik" does not.
*/
private fun macPreservingOui(value: String): String {
val sep = if (value.contains('-')) '-' else ':'
val parts = value.split(sep)
if (parts.size != 6 || parts.any { it.length != 2 }) return pseudo("mac", value) { "mac-" + it.take(8) }
val nic = pseudo("mac", value) { it }
return (parts.take(3) + listOf(nic.substring(0, 2), nic.substring(2, 4), nic.substring(4, 6)))
.joinToString(sep.toString())
.lowercase(Locale.ROOT)
}
/**
* Prefix-preserving within the same class, with reserved ranges kept verbatim: RFC1918 and
* CGNAT addresses say something about the topology and nothing about the person, and a run
* where 192.168.1.1 became a random public address would be actively misleading to read.
* Public addresses keep only their /16 so the network is still locatable at ISP granularity.
*/
private fun ip4(value: String): String {
// A route destination carries a prefix length; pseudonymize the address and put it back,
// or "0.0.0.0/0" turns into nonsense and the routing table becomes unreadable.
value.substringAfter('/', "").takeIf { it.isNotEmpty() && value.contains('/') }?.let { len ->
return ip4(value.substringBefore('/')) + "/" + len
}
val o = value.split(".")
if (o.size != 4 || o.any { it.toIntOrNull() == null }) return value
val n = o.map { it.toInt() }
val reserved = n[0] == 10 ||
(n[0] == 172 && n[1] in 16..31) ||
(n[0] == 192 && n[1] == 168) ||
(n[0] == 169 && n[1] == 254) ||
(n[0] == 100 && n[1] in 64..127) ||
n[0] == 127 || n[0] == 0 || n[0] >= 224
if (reserved) return value
val h = pseudo("ip4", value) { it }
return "${n[0]}.${n[1]}.${h.substring(0, 2).toInt(16)}.${h.substring(2, 4).toInt(16)}"
}
/**
* IPv6 keeps the scope and the first 32 bits (so 2001:db8: still reads as global unicast in
* the same allocation) and pseudonymizes the rest the interface identifier is the part that
* is a device fingerprint, especially with EUI-64.
*/
private fun ip6(value: String): String {
// Dotted quads reach here through the family-agnostic field names (addr, gateway, dst);
// hand them to the IPv4 path rather than mangling them as if they were v6.
if (value.count { it == ':' } < 2) return ip4(value)
if (value.contains('/')) {
return ip6(value.substringBefore('/')) + "/" + value.substringAfter('/')
}
val v = value.lowercase(Locale.ROOT)
// The unspecified address and the default route are not identities; mangling them would
// make a routing table unreadable for no privacy gain.
if (v == "::1" || v == "::" || v.startsWith("fe80:") || v.startsWith("ff")) return v
// Unique local addresses (fc00::/7) need the *whole* prefix replaced, not the tail.
//
// They look like the v6 equivalent of RFC1918, and the first instinct is to keep them for
// the same reason: private, topological, says nothing about anyone. That reasoning does
// not carry over. An RFC1918 prefix is shared by millions of networks and identifies
// none of them; a ULA global ID is 40 *random* bits, unique to one network by
// construction (RFC 4193). It is a network fingerprint. Passing the leading groups
// through - which is what the general path does - leaked 32 of those 40 bits.
//
// The prefix is pseudonymized as a unit, so two addresses on the same ULA subnet still
// land on the same pseudonymous prefix. "These hosts are on one network" survives;
// "this is *that* network" does not.
if (v.startsWith("fc") || v.startsWith("fd")) {
val groups = v.substringBefore('%').split(":")
val prefix = pseudo("ula-prefix", groups.take(3).joinToString(":")) { it }
val host = pseudo("ula-host", v) { it }
return "fd${prefix.substring(0, 2)}:${prefix.substring(2, 6)}:${prefix.substring(6, 10)}" +
"::${host.substring(0, 4)}"
}
val groups = v.substringBefore('%').split(":")
if (groups.size < 3) return v
val h = pseudo("ip6", value) { it }
return "${groups[0]}:${groups[1]}:${h.substring(0, 4)}:${h.substring(4, 8)}::${h.substring(8, 12)}"
}
/**
* Per-label pseudonyms with the public suffix kept, so "it resolved somewhere under .local"
* or "…under example.com" survives without naming the host. The suffix list is deliberately
* short: guessing wrong keeps *more* pseudonymized, never less.
*/
private fun fqdn(value: String): String {
if (value.isEmpty()) return value
val trailing = value.endsWith(".")
val labels = value.trimEnd('.').split(".")
if (labels.size == 1) return pseudo("fqdn", value) { "host-" + it.take(6) }
val keep = if (labels.last() in publicSuffixes) 1 else 0
val head = labels.dropLast(keep).map { l -> pseudo("label", l) { "l-" + it.take(6) } }
return (head + labels.takeLast(keep)).joinToString(".") + if (trailing) "." else ""
}
// ---- STRICT ---------------------------------------------------------------------------
/**
* STRICT keeps the shape of the document and the numbers, and nothing that quotes the
* network back. Evidence goes (trains carry addresses and hostnames), finding prose goes
* (it interpolates real names), networks go entirely.
*/
private fun strip(doc: JsonObject): JsonObject = buildJsonObject {
for ((k, v) in doc) {
when (k) {
"networks", "server_sessions" -> Unit
"tests" -> put(k, JsonArray((v as? JsonArray ?: JsonArray(emptyList())).map { t ->
val o = t.jsonObject
JsonObject(o.filterKeys { it != "evidence" && it != "params" })
}))
"findings" -> put(k, JsonArray((v as? JsonArray ?: JsonArray(emptyList())).map { f ->
val o = f.jsonObject
JsonObject(o.filterKeys { it != "description" && it != "title" && it != "evidence_refs" })
}))
"run" -> put(k, JsonObject(v.jsonObject.filterKeys { it != "notes" }))
else -> put(k, v)
}
}
}
// ---- pseudonym machinery ---------------------------------------------------------------
/** Deterministic per (domain, value, salt); memoized so one value maps to one pseudonym. */
private fun pseudo(domain: String, value: String, shape: (String) -> String): String =
cache.getOrPut("$domain\u0000$value") {
val md = MessageDigest.getInstance("SHA-256")
md.update(salt.bytes)
md.update(domain.toByteArray())
md.update(0)
md.update(value.lowercase(Locale.ROOT).toByteArray())
shape(md.digest().joinToString("") { "%02x".format(it) })
}
private companion object {
/**
* MAC | IPv6 | IPv4, in that order alternation is ordered, so a MAC-shaped token is
* claimed by the MAC rule before the IPv6 rule can see it.
*
* The patterns are deliberately conservative. A missed address is scrubbed by another
* rule or not at all; an over-eager one mangles timestamps, version strings and log
* prefixes, corrupting evidence to protect nothing.
*/
val EMBEDDED = Regex(
// Raw strings: a regex written with escaped escapes is a regex nobody can check.
"""(\b[0-9a-fA-F]{2}(?:[:-][0-9a-fA-F]{2}){5}\b)""" +
"""|(\b(?:[0-9a-fA-F]{1,4}:){2,7}(?::|[0-9a-fA-F]{1,4})(?:[0-9a-fA-F:]*))""" +
"""|(\b(?:\d{1,3}\.){3}\d{1,3}\b)"""
)
val publicSuffixes = setOf(
"local", "lan", "home", "internal", "arpa",
"com", "net", "org", "io", "app", "dev", "at", "de", "eu", "uk",
)
}
}
@@ -0,0 +1,166 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.privacy
/** The logical types of measurement-schema.md §8. */
enum class LogicalType { IP4, IP6, MAC, BSSID, SSID, FQDN, OPAQUE_ID, FREETEXT }
/**
* Which fields hold which logical type, and which whole subtrees are dropped below FULL.
*
* This is a table on purpose. The alternative annotating the Kotlin models and reflecting over
* them spreads the answer across every module and makes "what exactly gets uploaded?" a
* question you answer by reading the whole app. Here it is one file a reviewer can check against
* the spec in a sitting, and a new field that nobody classified stays visible in the output
* rather than being silently mangled.
*
* The bias is toward over-classifying: a metric wrongly pseudonymized is a bug someone reports;
* an SSID wrongly kept is a leak nobody notices.
*/
object Classification {
private val byKey: Map<String, LogicalType> = buildMap {
listOf(
"ip4", "ipv4", "gateway_ip4", "dns_ip4", "src_ip4", "dst_ip4", "public_ip4",
"observed_ip4", "hop_ip4", "answer_ip4", "address_ip4", "server_ip4",
).forEach { put(it, LogicalType.IP4) }
listOf(
"ip6", "ipv6", "gateway_ip6", "dns_ip6", "src_ip6", "dst_ip6", "public_ip6",
"observed_ip6", "hop_ip6", "answer_ip6", "address_ip6", "server_ip6",
"link_local", "ra_source", "prefix",
).forEach { put(it, LogicalType.IP6) }
// Family-agnostic address fields — the names the models actually use (Address.addr,
// Route.gateway, Route.dst, DnsConfig.servers). Their absence here was a real leak: the
// device's own global IPv6 address went out verbatim at the level whose description
// promises addresses are pseudonymized. Typed IP6 because the transform detects the
// family from the value, falling through to the IPv4 path for a dotted quad.
listOf(
"addr", "address", "gateway", "dst", "src", "servers", "server", "resolver",
"next_hop", "via", "public_ip", "observed_ip",
// Who sent each passive-discovery announcement (SSDP / LLMNR / NetBIOS / WS-Discovery).
// Family-agnostic on purpose: the same field carries a dotted quad from a v4 group and
// a link-local from ff02::c, and the v6 transform hands dotted quads to the v4 path.
"source_ip",
).forEach { put(it, LogicalType.IP6) }
listOf("mac", "hw_addr", "gateway_mac", "router_mac", "sender_mac", "peer_mac")
.forEach { put(it, LogicalType.MAC) }
listOf("bssid", "ap_mac").forEach { put(it, LogicalType.BSSID) }
listOf("ssid", "network_name", "wifi_ssid").forEach { put(it, LogicalType.SSID) }
listOf(
"fqdn", "hostname", "host", "name", "reverse_dns", "ptr", "domain", "query_name",
"friendly_name", "server_name", "sni", "cname", "search_domain", "device_name",
// Plural and prefixed variants the models actually use.
"search_domains", "private_dns_hostname", "domains", "hostnames",
// Names a device shouts at the whole segment. An LLMNR question and a NetBIOS
// registration are hostnames in every sense that matters here — they name a machine on
// somebody's home network — so they get the same per-label treatment as any other.
"netbios_name",
).forEach { put(it, LogicalType.FQDN) }
listOf(
"session_id", "credential", "token", "device_id", "android_id", "serial", "imsi", "iccid",
// Discovery identities. A UPnP USN and a WS-Discovery endpoint UUID are stable,
// globally unique per device and frequently derived from a serial number — exactly the
// thing that lets two uploads be recognised as the same household.
"usn", "device_uuid",
).forEach { put(it, LogicalType.OPAQUE_ID) }
listOf(
"notes", "detail", "raw", "excerpt", "location", "model_description",
// The make/model a device volunteers, and the URLs it points at. `server_banner` is
// deliberately not named `server`, which the family-agnostic address block already
// claims — a SERVER header run through the address transform would be mangled into
// nonsense while protecting nothing.
"server_banner", "product_hint", "wsd_types", "wsd_xaddrs",
).forEach { put(it, LogicalType.FREETEXT) }
}
/**
* Whole subtrees that BALANCED removes rather than pseudonymizes.
*
* Neighbour inventories (SSDP/UPnP responders, ARP tables, discovered peers) are the clearest
* case: they describe other people's devices, they are a household fingerprint even with the
* names hashed, and no metric depends on them. Dropping beats mangling.
*/
private val droppedPaths: List<List<String>> = listOf(
listOf("networks", "neighbors"),
listOf("networks", "arp"),
listOf("networks", "wifi", "scan_results"),
listOf("run", "device", "security_patch"),
)
/** Key suffixes whose whole value is a neighbour inventory wherever they appear. */
private val droppedKeys = setOf(
"ssdp_responders", "upnp", "neighbors", "arp_table", "scan_results",
"nearby_networks", "peers", "raw_dump", "dumpsys",
// The long run's passive-discovery inventories. Same argument as `ssdp_responders`, only
// more so: these are minutes of everything the segment said about itself — device models,
// hostnames, printers, who is looking for whom. Even fully pseudonymized the *shape* of a
// household is a fingerprint, no metric depends on the list (the counts live in `metrics`,
// which survives), and the collectors' own status fields stay behind to say the capture
// worked. Dropping beats mangling.
"ssdp_devices", "llmnr_queries", "netbios_names", "wsd_devices",
)
fun typeOf(key: String, path: List<String>): LogicalType? {
byKey[key]?.let { return it }
// Inside a discovery/neighbour container every string is someone's device name until
// proven otherwise, so classify unknown strings there as free text rather than passing
// them through.
if (path.any { it in droppedKeys }) return LogicalType.FREETEXT
return null
}
/**
* Last-resort classification from the *value*, when the field name is unrecognised.
*
* A name table can only protect fields somebody remembered to add, which is the wrong
* property for a privacy control: the dangerous field is the one nobody thought of. This
* exists because that failed once already `addresses[].addr` holds the device's own global
* IPv6 address, the table had never heard of the name, and it went out verbatim.
*
* Only addresses and MACs are inferred, because only those have shapes that cannot be
* mistaken for something else. Hostnames deliberately are not: `train.udp_updown` is
* indistinguishable from a domain by shape, and mangling a test type would corrupt the
* document to protect nothing.
*/
fun inferFromValue(value: String): LogicalType? {
val v = value.trim()
if (v.isEmpty() || v.length > 64) return null
if (looksLikeMac(v)) return LogicalType.MAC
if (looksLikeIp6(v)) return LogicalType.IP6
if (looksLikeIp4(v)) return LogicalType.IP4
return null
}
private fun isHex(c: Char) = c in '0'..'9' || c in 'a'..'f' || c in 'A'..'F'
private fun looksLikeMac(v: String): Boolean {
val parts = v.split(':', '-')
return parts.size == 6 && parts.all { p -> p.length == 2 && p.all(::isHex) }
}
private fun looksLikeIp4(v: String): Boolean {
val parts = v.substringBefore('/').split('.')
return parts.size == 4 && parts.all { p ->
p.isNotEmpty() && p.length <= 3 && p.all(Char::isDigit) && p.toInt() <= 255
}
}
private fun looksLikeIp6(v: String): Boolean {
val core = v.substringBefore('/').substringBefore('%')
// Two colons minimum, so a time or a MAC fragment does not qualify, and nothing but the
// characters an address may contain.
return core.count { it == ':' } >= 2 && core.all { it == ':' || isHex(it) }
}
fun dropAtBalanced(path: List<String>): Boolean {
if (path.isNotEmpty() && path.last() in droppedKeys) return true
return droppedPaths.any { dropped -> dropped.all { path.contains(it) } }
}
}
@@ -0,0 +1,228 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.privacy
import kotlinx.serialization.json.*
import kotlin.test.Test
import kotlin.test.assertEquals
import kotlin.test.assertFalse
import kotlin.test.assertNotEquals
import kotlin.test.assertNull
import kotlin.test.assertTrue
/**
* These tests are the audit of the anonymizer: each one states a property someone's privacy
* depends on, so a regression here fails loudly rather than quietly leaking.
*/
class AnonymizerTest {
private val json = Json { prettyPrint = false }
private val salt = Salt.perRun(ByteArray(32) { it.toByte() })
private fun sample(): JsonObject = json.parseToJsonElement(
"""
{
"schema": "echolot/measurement",
"run": {
"id": "0190-run", "trigger": "manual", "notes": "at Anna's flat",
"device": {"manufacturer": "OnePlus", "model": "CPH2747", "security_patch": "2026-06-05"}
},
"networks": [{
"id": "net-1", "ssid": "Rambossek WLAN", "bssid": "78:9a:18:aa:bb:cc",
"gateway_ip4": "192.168.1.1", "public_ip4": "89.185.109.150",
"gateway_ip6": "2001:1ad0:c4fe:6767::1", "link_local": "fe80::7a9a:18ff:feaa:bbcc",
"neighbors": [{"name": "Anna's Chromecast", "mac": "aa:bb:cc:dd:ee:ff"}],
"ssdp_responders": [{"friendly_name": "Living Room TV", "location": "http://192.168.1.44:8060/"}]
}],
"tests": [{
"id": "t1", "type": "train.udp_updown", "status": "ok",
"metrics": {"rtt_ms_avg": 12.4, "loss_pct": 0.0},
"evidence": {"seq": [0,1,2], "t_rx_ns": [1,2,3]}
}],
"findings": [{
"id": "f1", "code": "nat.udp_rebinding", "severity": "medium",
"title": "NAT remapped the port", "description": "server saw 89.185.109.150:41000"
}],
"summary": {"verdict": "warn"}
}
""".trimIndent(),
).jsonObject
private fun anon(level: PrivacyLevel, doc: JsonObject = sample()) = Anonymizer(level, salt).anonymize(doc)
private fun flat(e: JsonElement): String = e.toString()
@Test
fun fullLeavesTheDocumentAloneButRecordsThat() {
val out = anon(PrivacyLevel.FULL)
assertEquals("Rambossek WLAN", out["networks"]!!.jsonArray[0].jsonObject["ssid"]!!.jsonPrimitive.content)
assertEquals("full", out["run"]!!.jsonObject["privacy"]!!.jsonObject["anonymization"]!!.jsonPrimitive.content)
}
@Test
fun balancedRemovesTheSsidAndTheNotes() {
val text = flat(anon(PrivacyLevel.BALANCED))
assertFalse(text.contains("Rambossek"), "SSID survived: $text")
assertFalse(text.contains("Anna"), "free-text note or neighbour name survived: $text")
}
@Test
fun balancedDropsNeighbourInventoriesEntirely() {
val net = anon(PrivacyLevel.BALANCED)["networks"]!!.jsonArray[0].jsonObject
assertNull(net["neighbors"], "neighbour list should be dropped, not pseudonymized")
assertNull(net["ssdp_responders"], "SSDP responders should be dropped, not pseudonymized")
}
@Test
fun balancedKeepsTheVendorHalfOfAMac() {
val bssid = anon(PrivacyLevel.BALANCED)["networks"]!!.jsonArray[0].jsonObject["bssid"]!!.jsonPrimitive.content
assertTrue(bssid.startsWith("78:9a:18"), "OUI should survive so the vendor is still known: $bssid")
assertFalse(bssid.endsWith("aa:bb:cc"), "NIC part should be pseudonymized: $bssid")
}
@Test
fun privateAddressesAreKeptVerbatimAndPublicOnesAreNot() {
val net = anon(PrivacyLevel.BALANCED)["networks"]!!.jsonArray[0].jsonObject
assertEquals("192.168.1.1", net["gateway_ip4"]!!.jsonPrimitive.content,
"RFC1918 says nothing about the user and everything about the topology")
assertNotEquals("89.185.109.150", net["public_ip4"]!!.jsonPrimitive.content)
assertTrue(net["public_ip4"]!!.jsonPrimitive.content.startsWith("89.185."),
"the /16 should survive for ISP-level context")
}
@Test
fun linkLocalIsKeptButGlobalV6IsNot() {
val net = anon(PrivacyLevel.BALANCED)["networks"]!!.jsonArray[0].jsonObject
assertEquals("fe80::7a9a:18ff:feaa:bbcc", net["link_local"]!!.jsonPrimitive.content)
assertNotEquals("2001:1ad0:c4fe:6767::1", net["gateway_ip6"]!!.jsonPrimitive.content)
}
@Test
fun metricsAndVerdictsAreNeverTouched() {
for (level in PrivacyLevel.entries) {
val out = anon(level)
val t = out["tests"]!!.jsonArray[0].jsonObject
assertEquals(12.4, t["metrics"]!!.jsonObject["rtt_ms_avg"]!!.jsonPrimitive.double, 1e-9,
"$level changed a metric")
assertEquals("ok", t["status"]!!.jsonPrimitive.content)
assertEquals("warn", out["summary"]!!.jsonObject["verdict"]!!.jsonPrimitive.content)
}
}
@Test
fun findingCodesSurviveEveryLevelSoAggregationStillWorks() {
for (level in PrivacyLevel.entries) {
val f = anon(level)["findings"]!!.jsonArray[0].jsonObject
assertEquals("nat.udp_rebinding", f["code"]!!.jsonPrimitive.content, "$level lost the finding code")
assertEquals("medium", f["severity"]!!.jsonPrimitive.content)
}
}
@Test
fun strictDropsEvidenceAndProse() {
val out = anon(PrivacyLevel.STRICT)
assertNull(out["networks"], "STRICT should not describe the network at all")
assertNull(out["tests"]!!.jsonArray[0].jsonObject["evidence"])
assertNull(out["findings"]!!.jsonArray[0].jsonObject["description"])
assertFalse(flat(out).contains("89.185.109.150"), "an address leaked through finding prose")
}
@Test
fun pseudonymsAreConsistentWithinADocument() {
val doc = json.parseToJsonElement(
"""{"run":{"id":"r"},"networks":[{"ssid":"Home"},{"ssid":"Home"},{"ssid":"Other"}]}"""
).jsonObject
val nets = anon(PrivacyLevel.BALANCED, doc)["networks"]!!.jsonArray
val a = nets[0].jsonObject["ssid"]!!.jsonPrimitive.content
val b = nets[1].jsonObject["ssid"]!!.jsonPrimitive.content
val c = nets[2].jsonObject["ssid"]!!.jsonPrimitive.content
assertEquals(a, b, "the same SSID must map to the same pseudonym inside one run")
assertNotEquals(a, c, "different SSIDs must not collide")
}
@Test
fun perRunSaltsDoNotLinkTwoUploadsOfTheSameNetwork() {
val doc = json.parseToJsonElement("""{"run":{"id":"r"},"networks":[{"ssid":"Home"}]}""").jsonObject
val one = Anonymizer(PrivacyLevel.BALANCED, Salt.perRun(ByteArray(32) { 1 })).anonymize(doc)
val two = Anonymizer(PrivacyLevel.BALANCED, Salt.perRun(ByteArray(32) { 2 })).anonymize(doc)
assertNotEquals(
one["networks"]!!.jsonArray[0].jsonObject["ssid"],
two["networks"]!!.jsonArray[0].jsonObject["ssid"],
"a per-run salt must not produce a cross-run tracking identifier",
)
}
@Test
fun aStableSaltDoesLinkThemBecauseThatIsWhatItIsFor() {
val doc = json.parseToJsonElement("""{"run":{"id":"r"},"networks":[{"ssid":"Home"}]}""").jsonObject
val secret = ByteArray(32) { 7 }
val one = Anonymizer(PrivacyLevel.BALANCED, Salt.stable(secret)).anonymize(doc)
val two = Anonymizer(PrivacyLevel.BALANCED, Salt.stable(secret)).anonymize(doc)
assertEquals(
one["networks"]!!.jsonArray[0].jsonObject["ssid"],
two["networks"]!!.jsonArray[0].jsonObject["ssid"],
)
assertEquals("stable", one["run"]!!.jsonObject["privacy"]!!.jsonObject["salt"]!!.jsonPrimitive.content)
}
@Test
fun theDeclaredLevelMatchesWhatWasApplied() {
for (level in PrivacyLevel.entries) {
assertEquals(
level.wire,
anon(level)["run"]!!.jsonObject["privacy"]!!.jsonObject["anonymization"]!!.jsonPrimitive.content,
)
}
}
@Test
fun serverMinimumWins() {
assertEquals(PrivacyLevel.STRICT, PrivacyLevel.max(PrivacyLevel.FULL, PrivacyLevel.STRICT))
assertEquals(PrivacyLevel.BALANCED, PrivacyLevel.max(PrivacyLevel.BALANCED, PrivacyLevel.FULL))
assertEquals(PrivacyLevel.FULL, PrivacyLevel.fromWire("nonsense"))
}
// A ULA looks like the v6 RFC1918 and is not. Its global ID is 40 random bits, unique to one
// network by construction (RFC 4193), so the prefix IS the identifier - unlike 192.168.x,
// which millions of networks share. Passing the leading groups through leaked most of it.
@Test
fun ulaPrefixesArePseudonymizedWhole() {
val doc = json.parseToJsonElement(
"""{"run":{"id":"r"},"networks":[{"link":{"dns":{"servers":["fda1:3fb1:ff92:6696::2662"]}}}]}"""
).jsonObject
val out = flat(anon(PrivacyLevel.BALANCED, doc))
assertFalse(out.contains("fda1"), "the ULA global ID survived: $out")
assertFalse(out.contains("3fb1"), "part of the ULA global ID survived: $out")
assertTrue(out.contains("fd"), "the result should still read as a ULA: $out")
}
// Pseudonymizing the prefix as a unit keeps the one fact that is diagnostically useful:
// whether two addresses sit on the same network.
@Test
fun addressesOnOneUlaSubnetStayRelated() {
val doc = json.parseToJsonElement(
"""{"run":{"id":"r"},"networks":[{"link":{"dns":{"servers":[
"fda1:3fb1:ff92:6696::1","fda1:3fb1:ff92:6696::2","fdff:9999:8888:7777::1"]}}}]}"""
).jsonObject
val servers = anon(PrivacyLevel.BALANCED, doc)["networks"]!!.jsonArray[0].jsonObject["link"]!!
.jsonObject["dns"]!!.jsonObject["servers"]!!.jsonArray.map { it.jsonPrimitive.content }
val prefixOf = { s: String -> s.substringBeforeLast("::") }
assertEquals(prefixOf(servers[0]), prefixOf(servers[1]),
"two addresses on one ULA subnet should share a pseudonymous prefix")
assertNotEquals(prefixOf(servers[0]), prefixOf(servers[2]),
"a different ULA network must not collide with the first")
}
// RFC1918 stays readable, and this is the contrast that justifies it: a shared, meaningless
// prefix is topology; a unique random one is identity.
@Test
fun rfc1918StaysReadableUnlikeUla() {
val doc = json.parseToJsonElement(
"""{"run":{"id":"r"},"networks":[{"link":{"dns":{"servers":["192.168.1.1","10.13.102.1"]}}}]}"""
).jsonObject
val out = flat(anon(PrivacyLevel.BALANCED, doc))
assertTrue(out.contains("192.168.1.1"), "RFC1918 should survive: $out")
assertTrue(out.contains("10.13.102.1"), "RFC1918 should survive: $out")
}
}
@@ -0,0 +1,177 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.privacy
import kotlinx.serialization.json.Json
import kotlinx.serialization.json.JsonObject
import kotlinx.serialization.json.jsonArray
import kotlinx.serialization.json.jsonObject
import kotlinx.serialization.json.jsonPrimitive
import kotlin.test.Test
import kotlin.test.assertFalse
import kotlin.test.assertTrue
/**
* The blunt instrument: build a document with identifying values in every place one can actually
* occur, anonymize it, and assert none of them survive.
*
* [AnonymizerTest] checks that the fields the classification table knows about are handled
* correctly. This checks the other half the fields it does *not* know about. A per-field test
* can only fail for a field someone remembered to write a case for, which is exactly the wrong
* property for a privacy check: the dangerous field is the one nobody thought of.
*
* Concretely, this is written the way it is because the schema's own field names disagree with
* the classifier's. `Address.addr` carries an IP and is documented as such in
* measurement-schema.md §8, but the classifier keys on names like `ip4` and `gateway_ip4` and had
* never heard of `addr`.
*/
class LeakTest {
private val json = Json { prettyPrint = false }
private val salt = Salt.perRun(ByteArray(32) { 3 })
/**
* Every string here is something that identifies a person, a household or a device, placed
* where the real models actually put it (`core-measurement`'s Network/Link/Address/DnsConfig).
*/
private val secrets = listOf(
"Rambossek WLAN", // ssid
"78:9a:18:aa:bb:cc", // bssid
"aa:bb:cc:dd:ee:11", // gateway mac
"2001:1ad0:c4fe:6767::150", // global v6 address on the interface
"2a02:1748:dead:beef::1", // v6 default gateway
"203.0.113.77", // public v4
"nas.rambossek.lan", // private-dns hostname
"rambossek.lan", // search domain
"Anna's Chromecast", // neighbour name
"kitchen table", // free-text note
)
private fun document(): String = """
{
"schema": "echolot/measurement",
"run": {
"id": "run-1", "trigger": "manual", "notes": "${secrets[9]}",
"device": {"manufacturer": "OnePlus", "model": "CPH2747"}
},
"networks": [{
"id": "net-1", "transport": "wifi",
"link": {
"mtu": 1500,
"addresses": [
{"addr": "${secrets[3]}", "prefix_len": 64, "scope": "global"},
{"addr": "192.168.1.44", "prefix_len": 24, "scope": "global"}
],
"routes": [
{"dst": "::/0", "gateway": "${secrets[4]}", "iface": "wlan0"},
{"dst": "0.0.0.0/0", "gateway": "192.168.1.1", "iface": "wlan0"}
],
"dns": {
"servers": ["${secrets[5]}", "192.168.1.1"],
"private_dns_hostname": "${secrets[6]}",
"search_domains": ["${secrets[7]}"]
}
},
"wifi": {"ssid": "${secrets[0]}", "bssid": "${secrets[1]}"},
"neighbors": [{"name": "${secrets[8]}", "mac": "${secrets[2]}"}]
}],
"tests": [{"id": "t1", "type": "train.udp_updown", "status": "ok",
"metrics": {"rtt_ms_avg": 12.4}}],
"findings": [],
"summary": {"verdict": "ok"}
}
""".trimIndent()
private fun anonymized(level: PrivacyLevel): String =
json.encodeToString(
kotlinx.serialization.json.JsonObject.serializer(),
Anonymizer(level, salt).anonymize(json.parseToJsonElement(document()).jsonObject),
)
@Test
fun nothingIdentifyingSurvivesBalanced() {
val out = anonymized(PrivacyLevel.BALANCED)
val leaked = secrets.filter { out.contains(it) }
assertTrue(
leaked.isEmpty(),
"these identifying values were uploaded verbatim at BALANCED: $leaked\n\n$out",
)
}
@Test
fun nothingIdentifyingSurvivesStrict() {
val out = anonymized(PrivacyLevel.STRICT)
val leaked = secrets.filter { out.contains(it) }
assertTrue(leaked.isEmpty(), "leaked at STRICT: $leaked\n\n$out")
}
// Private addresses are kept on purpose — they describe the topology and not the person — so
// this pins that the leak test above is not passing by accident of over-redaction.
@Test
fun privateAddressesAreStillReadable() {
val out = anonymized(PrivacyLevel.BALANCED)
assertTrue(out.contains("192.168.1.1"), "RFC1918 gateway should survive: $out")
assertTrue(out.contains("192.168.1.44"), "RFC1918 interface address should survive: $out")
}
/**
* Raw shell output embeds a complete inventory of the local network, and neither the field-name
* table nor the whole-value shape check can see it: `ip_neigh` is one long string that is
* itself neither a MAC nor an address.
*
* This is not hypothetical. The blob below is (abridged) real output that reached the server
* at the `balanced` level from a test device, carrying the hardware address of every host on
* the network. measurement-schema.md §9 had flagged raw dumps as "hard to anonymize"; nothing
* enforced it.
*/
@Test
fun identifiersInsideRawShellOutputAreScrubbed() {
// Joined rather than written with escapes, so the fixture stays readable and there is no
// chance of an escape being mangled on its way into the JSON below.
val dump = listOf(
"uid=2000",
"10.13.102.5 dev wlan0 lladdr 90:09:d0:1a:83:e4 REACHABLE",
"10.13.102.1 dev wlan0 lladdr 78:9a:18:54:b8:f9 REACHABLE",
"10.13.102.111 dev wlan0 lladdr dc:a2:66:08:69:95 STALE",
"2001:4bb8:46a:e724:289d:87ff:feb6:ebd3 dev wlan0 lladdr b8:be:f4:bc:ca:cf STALE",
).joinToString(" | ")
val doc = json.parseToJsonElement(
"""{"run":{"id":"r"},"tests":[{"id":"t","type":"link.ip_monitor",
"evidence":{"ip_neigh":"$dump"}}]}"""
).jsonObject
val out = json.encodeToString(
kotlinx.serialization.json.JsonObject.serializer(),
Anonymizer(PrivacyLevel.BALANCED, salt).anonymize(doc),
)
for (mac in listOf("90:09:d0:1a:83:e4", "78:9a:18:54:b8:f9", "dc:a2:66:08:69:95", "b8:be:f4:bc:ca:cf")) {
assertFalse(out.contains(mac), "a neighbour's MAC survived inside the raw dump: $mac")
}
assertFalse(out.contains("2001:4bb8:46a:e724:289d:87ff:feb6:ebd3"),
"a global IPv6 survived inside the raw dump")
// Scrubbed, not dropped: the evidence must still be readable, or the raw dump stops being
// evidence at all. Structure, hostnames of the fields, and RFC1918 addresses stay.
assertTrue(out.contains("REACHABLE") && out.contains("STALE"), "the dump lost its structure")
assertTrue(out.contains("10.13.102.1"), "RFC1918 addresses should stay readable: $out")
assertTrue(out.contains("78:9a:18"), "the vendor prefix should survive for identification")
}
// A MAC in a raw dump and the same MAC in a parsed field must land on the same pseudonym, or
// the document stops being internally consistent and one device reads as two.
@Test
fun theSameIdentifierMatchesAcrossParsedAndRawFields() {
val doc = json.parseToJsonElement(
"""{"run":{"id":"r"},
"networks":[{"wifi":{"bssid":"78:9a:18:54:b8:f9"}}],
"tests":[{"id":"t","evidence":{"ip_neigh":"gw dev wlan0 lladdr 78:9a:18:54:b8:f9 REACHABLE"}}]}"""
).jsonObject
val out = Anonymizer(PrivacyLevel.BALANCED, salt).anonymize(doc)
val parsed = out["networks"]!!.jsonArray[0].jsonObject["wifi"]!!.jsonObject["bssid"]!!
.jsonPrimitive.content
val raw = json.encodeToString(kotlinx.serialization.json.JsonObject.serializer(), out)
assertTrue(raw.contains(parsed),
"the parsed BSSID pseudonym ($parsed) does not appear in the scrubbed dump")
}
}
@@ -0,0 +1,48 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.privacy
import kotlinx.serialization.json.Json
import kotlinx.serialization.json.JsonObject
import kotlinx.serialization.json.jsonObject
import java.io.File
import kotlin.test.Test
import kotlin.test.assertTrue
/**
* Runs the anonymizer over a real captured document when one is supplied via ECHOLOT_REAL_RUN,
* and reports every MAC and public address that survives.
*
* Fixtures only contain the identifiers somebody thought to put in them. A real run off a real
* phone contains whatever the probes actually produce which is how the raw-shell-output leak was
* found in the first place. Self-skips when no document is supplied, so nobody's network ends up
* committed to the repository.
*/
class RealDocumentTest {
@Test
fun noIdentifiersSurviveInARealDocument() {
val path = System.getenv("ECHOLOT_REAL_RUN")
if (path.isNullOrBlank() || !File(path).isFile) {
println("RealDocumentTest skipped (set ECHOLOT_REAL_RUN to a captured run)"); return
}
val json = Json { prettyPrint = false }
val doc = json.parseToJsonElement(File(path).readText()).jsonObject
val out = json.encodeToString(
JsonObject.serializer(),
Anonymizer(PrivacyLevel.BALANCED, Salt.perRun(ByteArray(32) { 5 })).anonymize(doc),
)
val macs = Regex("""\b[0-9a-fA-F]{2}(?::[0-9a-fA-F]{2}){5}\b""").findAll(out)
.map { it.value.lowercase() }
.filter { it != "00:00:00:00:00:00" }
.toSet()
val original = Regex("""\b[0-9a-fA-F]{2}(?::[0-9a-fA-F]{2}){5}\b""")
.findAll(File(path).readText()).map { it.value.lowercase() }.toSet()
val survived = macs intersect original
println("MACs in the original: ${original.size}; unchanged after anonymizing: ${survived.size}")
assertTrue(survived.isEmpty(), "these real MAC addresses survived anonymization: $survived")
}
}
+5
View File
@@ -28,4 +28,9 @@ dependencies {
implementation(project(":core-measurement")) implementation(project(":core-measurement"))
implementation(libs.kotlinx.serialization.json) implementation(libs.kotlinx.serialization.json)
implementation(libs.kotlinx.coroutines.android) implementation(libs.kotlinx.coroutines.android)
// JVM unit tests for the pure wire-format decoders (DiscoveryParsers.kt) against realistic and
// deliberately malformed payloads — no device, no Android runtime. Same setup as core-shizuku's
// dump-parser tests: JUnit 4, because that is what AGP's unit-test source set runs.
testImplementation(libs.kotlin.test.junit)
testImplementation(libs.junit4)
} }
@@ -0,0 +1,78 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.probe
import android.content.Context
import app.echo_lot.measurement.Test
import app.echo_lot.measurement.TestError
import app.echo_lot.measurement.TestStatus
import app.echo_lot.measurement.Tier
/**
* A measurement that watches, rather than one that asks the long-run counterpart to [Probe].
*
* The difference is not duration but what is observable at all. A [Probe] describes the network
* during its own two seconds, so anything intermittent is invisible to the whole battery unless it
* happens to coincide with a probe: a wifi link that drops for four seconds every two minutes is
* reported as perfectly healthy by every one-shot test, both before and after the gap. A collector
* starts at t=0, keeps sampling while the battery runs beside it and after it finishes, and yields
* one [Test] when the window closes.
*
* Contract, and all of it is load-bearing for a run that can be cancelled at any second:
* - [start] must return promptly, having launched whatever it needs on its own scope. The battery
* runs concurrently and must not wait for a listener.
* - [stop] must be callable after a failed [start], must never throw, and must return whatever was
* gathered so far. A cancelled long run still owes the user the two minutes it did watch.
* - Neither may throw to the caller; a collector that could not register its listener reports that
* as an `unsupported` Test, which is a result rather than an absence.
*/
interface Collector {
/** A TestType registry id — collectors do not get their own namespace. */
val type: String
val tier: Tier get() = Tier.APP
suspend fun start(ctx: Context, ids: ProbeIds)
suspend fun stop(): Test
}
/**
* Shared plumbing: the ids and the [TestBuilder] a collector needs in [Collector.stop], captured in
* [Collector.start] before anything that can fail.
*
* Assigned first thing on purpose. A collector whose registration throws still has to produce a
* Test saying so, and it cannot do that without a UUID source and a start timestamp so acquiring
* them is never allowed to be the step that failed.
*/
abstract class BaseCollector : Collector {
protected var ids: ProbeIds? = null
private set
private var builder: TestBuilder? = null
/** Call at the top of [Collector.start], before any platform call. */
protected fun begin(ids: ProbeIds, networkRef: String? = null) {
this.ids = ids
builder = TestBuilder(type, tier, ids, networkRef = networkRef)
}
/**
* Builds this collector's Test, or when [begin] never ran, so the collector was stopped
* without ever being started a `skipped` one saying exactly that. It reports rather than
* throws for the same reason probes do: the caller is assembling a document, and an exception
* there costs every other collector's data too.
*/
protected fun build(
status: TestStatus,
evidence: kotlinx.serialization.json.JsonObject? = null,
metrics: kotlinx.serialization.json.JsonObject? = null,
error: TestError? = null,
params: kotlinx.serialization.json.JsonObject? = null,
): Test = builder?.build(status, evidence, metrics, error, params)
?: Test(
id = "00000000-0000-7000-8000-000000000000", type = type, tier = tier,
startedMonoNs = 0, endedMonoNs = 0, status = TestStatus.SKIPPED,
error = TestError("not_started", "the collector was stopped before it was started"),
)
}
@@ -0,0 +1,74 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.probe
import android.content.Context
import android.net.ConnectivityManager
import android.net.NetworkCapabilities
import android.system.ErrnoException
import android.system.OsConstants
import app.echo_lot.measurement.Constraints
import app.echo_lot.measurement.Transport
import java.net.DatagramSocket
/**
* Detects what will prevent this run from measuring (measurement-schema.md §3 `constraints`),
* before any probe runs and independently of all of them.
*
* The known case: while a VPN holds the default route, Android refuses `Network.bindSocket()` on
* the underlying networks (EPERM) so apps cannot leak around the tunnel. Every per-network test
* then silently measures the tunnel or nothing, and the run comes out shaped exactly like a clean
* run of a healthy network. Detecting that here one throwaway bind per network is what lets
* the document say "these networks went unmeasured" instead of leaving the reader to infer it
* from a pattern of `attempted: false` scattered across the tests.
*/
object ConstraintDetector {
fun detect(ctx: Context, entries: List<NetworkInventory.Entry>): Constraints {
val cm = ctx.getSystemService(Context.CONNECTIVITY_SERVICE) as ConnectivityManager
// "A VPN holds the default route" is judged from the ACTIVE network, not from a VPN
// network merely existing in the list: a tunnel that was just disconnected lingers in
// allNetworks while it tears down, and counting it kept the app claiming "measured
// through a VPN" after the VPN was gone.
val vpnActive = runCatching {
cm.getNetworkCapabilities(cm.activeNetwork)
?.hasTransport(NetworkCapabilities.TRANSPORT_VPN) == true
}.getOrDefault(false)
val refused = ArrayList<String>()
for (e in entries) {
// The tunnel itself stays bindable — it is the underlying networks the OS walls off.
if (e.model.transport == Transport.VPN) continue
// Networks an app may never bind (carrier IMS/VoLTE, MMS, XCAP — no INTERNET, no
// NOT_RESTRICTED) refuse with EPERM permanently, VPN or no VPN. Counting them as a
// constraint is what made a phone with VoLTE report "measurement blocked" forever
// and forced every run INCONCLUSIVE. They are recorded in networks[] as
// app_usable:false; they are not something a run failed to do.
if (e.model.appUsable == false) continue
val err = try {
DatagramSocket().use { s -> e.handle.bindSocket(s) }
null
} catch (t: Throwable) {
t
}
// Only the OS *refusing* counts as blocked (EPERM: the VPN wall). A network that
// happens to die mid-snapshot fails its bind too, but with a different errno, and
// calling that "per-network probing blocked" would flip a whole healthy run to
// INCONCLUSIVE over one network going away — the probes already record
// attempted:false for that case.
if (err != null && isPermissionRefusal(err)) refused.add(e.model.id)
}
return Constraints(
vpnActive = vpnActive,
perNetworkBlocked = refused.isNotEmpty(),
unmeasuredNetworks = refused,
)
}
private fun isPermissionRefusal(t: Throwable): Boolean =
generateSequence(t) { it.cause }.any {
(it is ErrnoException && it.errno == OsConstants.EPERM) ||
(it.message?.contains("EPERM") == true)
}
}
@@ -0,0 +1,360 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.probe
/**
* Wire-format decoders for the passive discovery collectors (SSDP, LLMNR, NetBIOS-NS, WS-Discovery).
*
* No Android imports on purpose: every byte these see was broadcast by an unidentified device on
* someone else's network, so they are the part of the collectors that most needs to be exercised
* against malformed input and that is only cheap to do if it runs on a plain JVM. See
* DiscoveryParsersTest.
*
* The universal contract here is **return null, never throw**. A collector that dies on one
* malformed datagram loses the whole window's inventory, and a device that emits garbage is a
* device we still want counted. Every decoder therefore bounds-checks by hand rather than relying
* on an exception, and treats "this is not the protocol I parse" and "this is the protocol but it
* is broken" as the same answer: nothing to record.
*/
// ---- shared -------------------------------------------------------------------------------
/**
* Makes a decoded string safe to embed in a JSON document.
*
* Names arrive as arbitrary bytes. Control characters would survive JSON encoding as escapes and
* turn up in a terminal that interprets them, and an unbounded length is a memory cost decided by
* whoever is shouting on the segment so both are capped here rather than at each call site.
*/
internal fun sanitizeText(s: String, max: Int = 255): String {
val sb = StringBuilder(minOf(s.length, max))
for (c in s) {
if (sb.length >= max) break
sb.append(if (c.isISOControl() || c == '') '?' else c)
}
return sb.toString()
}
private fun u8(b: ByteArray, i: Int) = b[i].toInt() and 0xFF
private fun u16(b: ByteArray, i: Int) = (u8(b, i) shl 8) or u8(b, i + 1)
// ---- SSDP ---------------------------------------------------------------------------------
enum class SsdpKind { ALIVE, BYEBYE, UPDATE, RESPONSE, SEARCH, OTHER }
/**
* One SSDP message. [target] is NT for announcements and ST for search responses they name the
* same thing (which service is being talked about) from the two sides of the conversation, so they
* collapse into one field.
*/
data class SsdpMessage(
val kind: SsdpKind,
val target: String?,
val usn: String?,
val serverBanner: String?,
val location: String?,
)
object SsdpParser {
/**
* SSDP is HTTP-shaped but is not HTTP: there is no framing, no content length, and vendors
* disagree about line endings. Parsing it as "a start line plus colon-separated headers, be
* liberal about the rest" is the whole job — running it through an HTTP client would reject
* messages that real devices send and that we want to count.
*
* ISO-8859-1 decoding because the headers are byte-oriented and this mapping is total: no byte
* sequence can fail to decode, so a device with a Latin-1 model name in its SERVER banner is
* recorded rather than replaced by question marks.
*/
fun parse(bytes: ByteArray, len: Int): SsdpMessage? {
if (len <= 0 || len > bytes.size) return null
val text = String(bytes, 0, len, Charsets.ISO_8859_1)
val lines = text.split('\n')
val start = lines.firstOrNull()?.trim().orEmpty()
if (start.isEmpty()) return null
val headers = HashMap<String, String>()
for (i in 1 until lines.size) {
val line = lines[i].trimEnd('\r')
if (line.isBlank()) break
val c = line.indexOf(':')
if (c <= 0) continue
val key = line.substring(0, c).trim().uppercase()
// First occurrence wins: a duplicated header is a device bug, and taking the first
// matches what every SSDP implementation in the wild does.
if (key !in headers) headers[key] = sanitizeText(line.substring(c + 1).trim(), 512)
}
val kind = when {
start.startsWith("NOTIFY", ignoreCase = true) -> when (headers["NTS"]?.lowercase()) {
"ssdp:alive" -> SsdpKind.ALIVE
"ssdp:byebye" -> SsdpKind.BYEBYE
"ssdp:update" -> SsdpKind.UPDATE
else -> SsdpKind.OTHER
}
start.startsWith("M-SEARCH", ignoreCase = true) -> SsdpKind.SEARCH
start.startsWith("HTTP/", ignoreCase = true) -> SsdpKind.RESPONSE
else -> return null // not SSDP at all
}
return SsdpMessage(
kind = kind,
target = headers["NT"] ?: headers["ST"],
usn = headers["USN"],
serverBanner = headers["SERVER"],
location = headers["LOCATION"],
)
}
/**
* The make/model inside a SERVER banner, as a hint.
*
* A banner reads `Linux/4.4 UPnP/1.0 Synology/DSM-7.3` or `FRITZ!Box 7590 UPnP/1.0`: the
* interesting token is whichever one is not the OS and not the protocol version, and there is
* no grammar that says which. Dropping the known-boilerplate tokens and keeping the rest is
* therefore a heuristic and is reported as such [SsdpMessage.serverBanner] is kept verbatim
* beside it so nobody has to trust this to read the evidence.
*/
fun productHint(serverBanner: String?): String? {
val banner = serverBanner?.trim().orEmpty()
if (banner.isEmpty()) return null
val kept = banner.split(' ', '\t')
.map { it.trim() }
.filter { it.isNotEmpty() && it.substringBefore('/').lowercase() !in BOILERPLATE }
.distinct()
return kept.joinToString(" ").takeIf { it.isNotEmpty() }
}
private val BOILERPLATE = setOf(
"upnp", "http", "dlnadoc", "linux", "unix", "windows", "darwin", "posix", "sdk",
"upnp-device-host", "microsoft-windows", "mono.upnp", "webos", "android",
)
}
// ---- DNS-format questions (LLMNR, and the shape NetBIOS-NS borrows) -------------------------
/** One question from a DNS-format packet. [isQuery] separates "who is asking" from "who answered". */
data class DnsQuestion(
val id: Int,
val name: String,
val qtype: Int,
val qclass: Int,
val isQuery: Boolean,
val opcode: Int,
)
object LlmnrParser {
/**
* Decodes the first question of an LLMNR packet, which is DNS wire format with a different
* transport.
*
* Name compression is rejected rather than followed. RFC 4795 forbids it in LLMNR, so a pointer
* here is either a broken sender or someone hoping the parser will chase it and a pointer
* loop is the classic way to hang a DNS decoder. Refusing costs nothing real and removes the
* only unbounded loop this decoder could have had.
*/
fun parse(bytes: ByteArray, len: Int): DnsQuestion? {
if (len < HEADER + 5 || len > bytes.size) return null
val flags = u16(bytes, 2)
if (u16(bytes, 4) < 1) return null // no question section
val sb = StringBuilder()
var i = HEADER
var labels = 0
while (true) {
if (i >= len) return null
val l = u8(bytes, i)
if (l == 0) { i++; break }
if (l and 0xC0 != 0) return null // compression pointer / reserved length
i++
if (i + l > len) return null
if (++labels > MAX_LABELS || sb.length + l > MAX_NAME) return null
if (sb.isNotEmpty()) sb.append('.')
sb.append(String(bytes, i, l, Charsets.UTF_8))
i += l
}
if (sb.isEmpty()) return null
if (i + 4 > len) return null
return DnsQuestion(
id = u16(bytes, 0),
name = sanitizeText(sb.toString()),
qtype = u16(bytes, i),
qclass = u16(bytes, i + 2),
isQuery = (flags and 0x8000) == 0,
opcode = (flags shr 11) and 0x0F,
)
}
/** The record types worth naming in evidence; anything else is reported as its number. */
fun qtypeName(qtype: Int): String = when (qtype) {
1 -> "A"
28 -> "AAAA"
12 -> "PTR"
33 -> "SRV"
255 -> "ANY"
else -> qtype.toString()
}
private const val HEADER = 12
private const val MAX_LABELS = 64
private const val MAX_NAME = 255
}
// ---- NetBIOS name service ------------------------------------------------------------------
/**
* A decoded NetBIOS name-service question.
*
* [suffix] is the sixteenth byte of the name and is the part an engineer reads first: it says what
* the announcement is *for* (a workstation, a file server, a browser election) rather than who is
* making it.
*/
data class NetbiosName(
val name: String,
val suffix: Int,
val role: String,
val isResponse: Boolean,
val opcode: Int,
)
object NetbiosParser {
/**
* Decodes the question name from an NBNS packet (RFC 1002 §4.2).
*
* The header is DNS-shaped, but the name is not: NetBIOS first-level encoding splits each of
* the 16 name bytes into two nibbles and adds 'A' to each, so a 16-byte name is always exactly
* 32 characters drawn from A-P. That fixed shape is also the validity check anything outside
* A-P means this is not an NBNS name, and it is cheaper and safer to reject the packet than to
* guess at what a half-decodable name was supposed to say.
*/
fun parse(bytes: ByteArray, len: Int): NetbiosName? {
if (len < HEADER + 34 || len > bytes.size) return null
if (u16(bytes, 4) < 1) return null // no question section
if (u8(bytes, HEADER) != ENCODED_LEN) return null
val raw = ByteArray(16)
for (j in 0 until 16) {
val hi = u8(bytes, HEADER + 1 + j * 2) - 'A'.code
val lo = u8(bytes, HEADER + 2 + j * 2) - 'A'.code
if (hi !in 0..15 || lo !in 0..15) return null
raw[j] = ((hi shl 4) or lo).toByte()
}
// The first 15 bytes are the name, space-padded; the 16th is the suffix.
val padded = String(raw, 0, 15, Charsets.ISO_8859_1)
val name = sanitizeText(padded.trimEnd { it == ' ' || it.isISOControl() }, 15)
if (name.isEmpty()) return null
val suffix = raw[15].toInt() and 0xFF
val flags = u16(bytes, 2)
return NetbiosName(
name = name,
suffix = suffix,
role = roleOf(suffix),
isResponse = (flags and 0x8000) != 0,
opcode = (flags shr 11) and 0x0F,
)
}
/** RFC 1001 §15 / the Microsoft suffix assignments people actually see on a LAN. */
fun roleOf(suffix: Int): String = when (suffix) {
0x00 -> "workstation"
0x03 -> "messenger"
0x1B -> "domain_master_browser"
0x1C -> "domain_controllers"
0x1D -> "master_browser"
0x1E -> "browser_elections"
0x20 -> "file_server"
else -> "suffix_0x%02X".format(suffix)
}
/** NBNS opcodes: what the sender is doing, not just that it is talking. */
fun opcodeName(opcode: Int): String = when (opcode) {
0 -> "query"
5 -> "registration"
6 -> "release"
7 -> "wack"
8 -> "refresh"
else -> "opcode_$opcode"
}
private const val HEADER = 12
private const val ENCODED_LEN = 32
}
// ---- WS-Discovery --------------------------------------------------------------------------
/** The four things a WS-Discovery datagram is worth reading for. */
data class WsdMessage(
val action: String?,
val deviceUuid: String?,
val types: String?,
val xaddrs: String?,
)
object WsdParser {
/**
* Pulls four leaf values out of SOAP-over-UDP by targeted matching, deliberately **without an
* XML parser**.
*
* The input is unauthenticated, unsolicited, and written by whatever is on the segment. Handing
* that to a real XML parser buys namespace correctness and pays for it with the whole XML
* attack surface entity expansion (a 700-byte datagram that allocates gigabytes), DTDs that
* fetch external resources, and nesting deep enough to exhaust the stack all inside a probe
* whose contract is that it never throws. None of that surface is needed to read four leaf
* elements out of a message we are only ever going to count and quote.
*
* So: cap the text, then match `<[prefix:]Tag ...>value<` with a character class that cannot
* backtrack. The worst case is a value we fail to extract, which is recorded as an absence.
*/
fun parse(bytes: ByteArray, len: Int): WsdMessage? {
if (len <= 0 || len > bytes.size) return null
val text = String(bytes, 0, minOf(len, MAX_TEXT), Charsets.UTF_8)
if (!text.contains("Envelope", ignoreCase = true)) return null
// The action tail is the message type — Hello, Bye, Probe, ProbeMatches, ResolveMatches.
val action = leaf(text, "Action")?.substringAfterLast('/')?.takeIf { it.isNotEmpty() }
// The device's stable identity is an EndpointReference/Address holding a urn:uuid. Other
// Address elements exist (wsa:To names the discovery group), so the uuid shape picks the
// right one rather than the first one.
val uuid = leaves(text, "Address").firstOrNull { it.contains("uuid:", ignoreCase = true) }
val msg = WsdMessage(
action = action?.let { sanitizeText(it, 64) },
deviceUuid = uuid?.let { sanitizeText(it, 128) },
types = leaf(text, "Types")?.let { sanitizeText(it, 200) },
xaddrs = leaf(text, "XAddrs")?.let { sanitizeText(it, 300) },
)
val empty = msg.action == null && msg.deviceUuid == null &&
msg.types == null && msg.xaddrs == null
return if (empty) null else msg
}
private fun leaf(xml: String, tag: String): String? = leaves(xml, tag).firstOrNull()
private fun leaves(xml: String, tag: String): List<String> {
val re = TAGS[tag] ?: return emptyList()
return re.findAll(xml).map { it.groupValues[1].trim() }.filter { it.isNotEmpty() }.toList()
}
/**
* `[^<]{0,512}` rather than a lazy `.*?`: it can only ever match forward, so there is no input
* that makes this regex expensive. WS-Discovery leaves hold no markup, so nothing is lost.
*
* Built once, eagerly, because the alternative is a memoizing map touched from the capture
* thread a data race for the sake of four Regex allocations.
*/
private val TAGS: Map<String, Regex> =
listOf("Action", "Address", "Types", "XAddrs").associateWith { tag ->
Regex(
"""<(?:[A-Za-z0-9_.\-]{1,32}:)?$tag\b[^>]{0,256}>([^<]{0,512})<""",
RegexOption.IGNORE_CASE,
)
}
private const val MAX_TEXT = 16 * 1024
}
@@ -0,0 +1,227 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.probe
import android.content.Context
import app.echo_lot.measurement.Test
import app.echo_lot.measurement.TestStatus
import app.echo_lot.measurement.TestType
import app.echo_lot.measurement.Tier
import kotlinx.coroutines.Dispatchers
import kotlinx.coroutines.withContext
import kotlinx.serialization.json.JsonObject
import kotlinx.serialization.json.buildJsonObject
import kotlinx.serialization.json.put
import java.net.DatagramPacket
import java.net.DatagramSocket
import java.net.InetAddress
import java.net.InetSocketAddress
import java.util.Random
/**
* Asks the network's own DNS servers directly, then asks Android to resolve the same name, and
* compares.
*
* The comparison is the point. A name that fails to resolve looks the same to a user whatever the
* cause, but the causes want opposite responses: if the server does not answer, the network is
* broken and the router is the thing to look at; if the server answers a raw query while the
* platform still cannot resolve, the device's own resolver has wedged and toggling wifi fixes it in
* seconds. Nothing else on a phone will tell you which of those you have.
*
* This is deliberately not a general DNS test no recursion checks, no DNSSEC, no rewriting
* detection; [DnsCanaryProbe] covers interception. This one answers a single question: is the
* resolver on this device doing its job.
*/
class DnsResolverProbe(
private val entries: List<NetworkInventory.Entry>,
/** Resolved directly rather than through any cache; any name with a stable answer will do. */
private val probeName: String = "one.one.one.one",
) : Probe {
override val type = TestType.DNS_RESOLVER
override val tier = Tier.APP
// A query per server with a 3s ceiling, plus one getaddrinfo that may sit out its own timeout.
override val estimatedMs = 6_000L
override suspend fun run(ctx: Context, ids: ProbeIds): Test = withContext(Dispatchers.IO) {
val b = TestBuilder(type, tier, ids)
val perNetwork = LinkedHashMap<String, Pair<String, JsonObject>>()
for (e in entries) {
val servers = e.model.link.dns?.servers.orEmpty()
if (servers.isEmpty()) continue
val label = "${e.model.transport.name.lowercase()}:${e.model.id}"
// Directly: does the configured server answer at all?
var direct: Boolean? = null
var directDetail = "no server answered"
for (s in servers) {
val r = try {
queryDirect(s, probeName)
} catch (e: DnsRefused) {
// Distinguished deliberately: a server that replies with a failure is a
// working server saying no, which points at the network rather than here.
direct = false
directDetail = "$s ${e.why}"
break
}
if (r != null) {
direct = true
directDetail = "$s answered in ${r}ms"
break
}
direct = false
directDetail = "$s did not answer"
}
// The search domains the network handed out, asked about separately.
//
// A resolver appends these to a lookup, so a search domain the server will not answer
// for stalls every name a client asks about — and it fails as silence, which is
// indistinguishable from packet loss, so clients retry rather than moving on. Asking
// about a name that cannot exist is deliberate: the answer wanted here is NXDOMAIN,
// and what matters is only whether anything comes back at all.
val searchDomains = e.model.link.dns?.searchDomains.orEmpty()
var searchAnswered: Boolean? = null
var searchDetail = ""
for (d in searchDomains) {
val nonce = "echolot-probe-" + java.util.UUID.randomUUID().toString().take(8)
val answered = servers.any { srv ->
runCatching { queryDirect(srv, "$nonce.$d") != null }
.getOrElse { it is DnsRefused } // a refusal is still an answer
}
if (!answered) {
searchAnswered = false
searchDetail = "$d is not answered at all — queries under it vanish"
break
}
searchAnswered = true
searchDetail = "$d answers"
}
// Through the platform: what an app actually gets.
val viaSystem = runCatching {
e.handle.getAllByName(probeName).isNotEmpty()
}.getOrElse { false }
perNetwork[label] = e.model.id to buildJsonObject {
put("network_ref", e.model.id)
put("servers", servers.joinToString(","))
direct?.let { put("direct_answer", it) }
put("direct_detail", directDetail)
if (searchDomains.isNotEmpty()) {
put("search_domains", searchDomains.joinToString(","))
searchAnswered?.let { put("search_answered", it) }
put("search_detail", searchDetail)
}
put("system_resolves", viaSystem)
// Named here rather than left for a finding to infer, because the pairing is the
// whole observation and splitting it across two places invites reading one alone.
put(
"verdict",
when {
// Ordered by which component is at fault, most specific first. A search
// domain that swallows queries explains a failure that would otherwise be
// blamed on the device, so it has to be tested before that conclusion.
searchAnswered == false -> "search domain swallows queries"
viaSystem -> "resolver working"
direct == true -> "server answers, device resolver does not"
direct == false -> "server does not answer"
else -> "not determined"
},
)
}
}
if (perNetwork.isEmpty()) {
return@withContext b.build(
TestStatus.SKIPPED,
evidence = buildJsonObject { put("reason", "no network advertised a DNS server") },
)
}
val evidence = buildJsonObject {
put("name", probeName)
for ((label, v) in perNetwork) put(label, v.second)
}
// OK means the measurement ran, not that DNS is healthy — the finding says that.
b.build(TestStatus.OK, evidence = evidence)
}
/**
* Sends one A query straight to [server] over UDP. Returns the round trip in ms, or null.
*
* Hand-rolled rather than via any resolver API on purpose: the entire point is to bypass the
* component under suspicion. Anything that goes through the platform resolver would inherit
* exactly the fault this is trying to detect.
*/
private fun queryDirect(server: String, name: String): Long? {
return try {
queryDirectOrThrow(server, name)
} catch (e: DnsRefused) {
throw e
} catch (t: Throwable) {
null
}
}
private fun queryDirectOrThrow(server: String, name: String): Long? = run {
val id = Random().nextInt(0xFFFF)
val query = buildQuery(id, name)
DatagramSocket().use { sock ->
sock.soTimeout = 3000
val addr = InetAddress.getByName(server) // a literal from DHCP; no lookup happens
val t0 = System.nanoTime()
sock.send(DatagramPacket(query, query.size, InetSocketAddress(addr, 53)))
val buf = ByteArray(512)
val reply = DatagramPacket(buf, buf.size)
sock.receive(reply)
val ms = (System.nanoTime() - t0) / 1_000_000
// A reply is not an answer. Counting any packet as success would let a REFUSED or
// SERVFAIL — both perfectly well-formed responses — be reported as "the server
// answers", and this probe's whole output is the claim that the server is fine and
// the device is not. That would be an accusation pointed at the wrong component,
// stated with confidence.
val replyId = ((buf[0].toInt() and 0xFF) shl 8) or (buf[1].toInt() and 0xFF)
val rcode = if (reply.length >= 4) buf[3].toInt() and 0x0F else -1
val answers = if (reply.length >= 8) {
((buf[6].toInt() and 0xFF) shl 8) or (buf[7].toInt() and 0xFF)
} else 0
when {
replyId != id || reply.length < 12 -> null
rcode != 0 -> throw DnsRefused(rcodeName(rcode))
answers == 0 -> throw DnsRefused("answered with no records")
else -> ms
}
}
}
/** The server replied, but with a failure — which is a network fault, not a device one. */
private class DnsRefused(val why: String) : Exception(why)
private fun rcodeName(rcode: Int): String = when (rcode) {
1 -> "rejected the query as malformed"
2 -> "reported its own failure (SERVFAIL)"
3 -> "said the name does not exist (NXDOMAIN)"
4 -> "does not implement this query"
5 -> "refused the query (REFUSED)"
else -> "returned rcode $rcode"
}
/** A minimal DNS query: one question, class IN, type A, recursion desired. */
private fun buildQuery(id: Int, name: String): ByteArray {
val labels = name.split('.').filter { it.isNotEmpty() }
val out = ArrayList<Byte>(32)
out.add((id shr 8).toByte()); out.add(id.toByte())
out.add(0x01); out.add(0x00) // recursion desired
out.add(0x00); out.add(0x01) // one question
repeat(6) { out.add(0x00) } // no answers, authority or additional
for (l in labels) {
out.add(l.length.toByte())
for (c in l.toByteArray(Charsets.US_ASCII)) out.add(c)
}
out.add(0x00) // root label
out.add(0x00); out.add(0x01) // type A
out.add(0x00); out.add(0x01) // class IN
return out.toByteArray()
}
}
@@ -0,0 +1,113 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.probe
import android.net.Network
import android.system.Os
import android.system.OsConstants
import android.system.StructTimeval
import java.io.FileDescriptor
import java.net.InetAddress
import java.nio.ByteBuffer
import java.util.Locale
/**
* One echo exchange over the unprivileged ICMP datagram socket.
*
* Extracted from [IcmpProbe] when the long-run [PingSeriesCollector] needed the same exchange a
* few hundred times instead of once. The two differ only in how often they call this; a second
* copy of the checksum and the sent/not-sent bookkeeping would only be a second place for them to
* drift apart.
*
* Works without root because Android ships an open `ping_group_range` validated on hardware by
* the prober, and the reason this probe family exists at app tier at all.
*/
internal object IcmpEcho {
/**
* One attempt's result.
*
* [attempted] separates "we sent an echo request and heard nothing" from "we never got as far
* as sending one". Both leave [ok] false, and collapsing them is how a probe ends up asserting
* something about a network it never touched: binding to a non-default network can fail with
* EPERM, and reporting that as ICMP silence blames the network for the app's own inability to
* use the interface. For a series it is also the difference between a lost packet and a socket
* that was never usable one is loss, the other is not.
*/
data class Result(
val ok: Boolean,
val attempted: Boolean,
val detail: String,
val rttMs: Double?,
)
fun ping(
network: Network?,
target: String,
v6: Boolean,
timeoutMs: Int,
seq: Int = 1,
): Result {
var fd: FileDescriptor? = null
var sent = false
return try {
val proto = if (v6) OsConstants.IPPROTO_ICMPV6 else OsConstants.IPPROTO_ICMP
val family = if (v6) OsConstants.AF_INET6 else OsConstants.AF_INET
fd = Os.socket(family, OsConstants.SOCK_DGRAM, proto)
// Everything up to and including sendto is setup. A failure here means the test did
// not run on this network — not that the network stayed silent.
network?.bindSocket(fd)
Os.setsockoptTimeval(
fd, OsConstants.SOL_SOCKET, OsConstants.SO_RCVTIMEO,
StructTimeval.fromMillis(timeoutMs.toLong()),
)
val addr = network?.getByName(target) ?: InetAddress.getByName(target)
val ident = (Os.getpid() and 0xFFFF)
val packet = buildEchoRequest(v6, ident.toShort(), seq.toShort())
val t0 = System.nanoTime()
Os.sendto(fd, packet, 0, packet.size, 0, addr, 0)
sent = true
val buf = ByteBuffer.allocate(1500)
val received = Os.recvfrom(fd, buf, 0, null)
val rttMs = (System.nanoTime() - t0) / 1_000_000.0
val replyType = if (received > 0) buf.get(0).toInt() and 0xFF else -1
val ok = replyType == (if (v6) 129 else 0)
Result(
ok, true,
"reply type=$replyType rtt_ms=${"%.1f".format(Locale.ROOT, rttMs)} bytes=$received",
if (ok) rttMs else null,
)
} catch (e: Throwable) {
// A timeout after a successful send is a real "no reply"; anything before it is not.
Result(false, sent, "error: ${e.message ?: e.javaClass.simpleName}", null)
} finally {
fd?.let { runCatching { Os.close(it) } }
}
}
private fun buildEchoRequest(v6: Boolean, ident: Short, seq: Short): ByteArray {
val type = if (v6) 128 else 8
val payload = "echolot".toByteArray()
val pkt = ByteBuffer.allocate(8 + payload.size)
pkt.put(type.toByte()); pkt.put(0); pkt.putShort(0)
pkt.putShort(ident); pkt.putShort(seq); pkt.put(payload)
val bytes = pkt.array()
// The v6 checksum is computed by the kernel over a pseudo-header the socket owns; filling
// it in here would be wrong, not merely redundant.
if (!v6) {
val cs = checksum(bytes)
bytes[2] = (cs.toInt() shr 8).toByte(); bytes[3] = (cs.toInt() and 0xFF).toByte()
}
return bytes
}
private fun checksum(b: ByteArray): Short {
var sum = 0; var i = 0
while (i < b.size - 1) { sum += ((b[i].toInt() and 0xFF) shl 8) or (b[i + 1].toInt() and 0xFF); i += 2 }
if (i < b.size) sum += (b[i].toInt() and 0xFF) shl 8
while (sum shr 16 != 0) sum = (sum and 0xFFFF) + (sum shr 16)
return sum.inv().toShort()
}
}
@@ -5,9 +5,6 @@ package app.echo_lot.probe
import android.content.Context import android.content.Context
import android.net.Network import android.net.Network
import android.system.Os
import android.system.OsConstants
import android.system.StructTimeval
import app.echo_lot.measurement.Test import app.echo_lot.measurement.Test
import app.echo_lot.measurement.TestStatus import app.echo_lot.measurement.TestStatus
import app.echo_lot.measurement.TestType import app.echo_lot.measurement.TestType
@@ -17,10 +14,6 @@ import kotlinx.coroutines.withContext
import kotlinx.serialization.json.JsonObject import kotlinx.serialization.json.JsonObject
import kotlinx.serialization.json.buildJsonObject import kotlinx.serialization.json.buildJsonObject
import kotlinx.serialization.json.put import kotlinx.serialization.json.put
import java.io.FileDescriptor
import java.net.InetAddress
import java.nio.ByteBuffer
import java.util.Locale
/** /**
* icmp.ping4 / icmp.ping6 via the unprivileged ICMP datagram socket, per active network * icmp.ping4 / icmp.ping6 via the unprivileged ICMP datagram socket, per active network
@@ -40,24 +33,37 @@ class IcmpProbe(
override suspend fun run(ctx: Context, ids: ProbeIds): Test = withContext(Dispatchers.IO) { override suspend fun run(ctx: Context, ids: ProbeIds): Test = withContext(Dispatchers.IO) {
val b = TestBuilder(type, tier, ids) val b = TestBuilder(type, tier, ids)
val perNetwork = LinkedHashMap<String, String>() val perNetwork = LinkedHashMap<String, Pair<String?, IcmpEcho.Result>>()
var anyOk = false var anyOk = false
val rtts = ArrayList<Double>() val rtts = ArrayList<Double>()
// Default network first, then each active network explicitly. // Default network first, then each active network explicitly.
attempt(null).let { (ok, detail, rtt) -> attempt(null).let { a ->
perNetwork["default"] = detail; if (ok) { anyOk = true; rtt?.let(rtts::add) } perNetwork["default"] = null to a
if (a.ok) { anyOk = true; a.rttMs?.let(rtts::add) }
} }
for (e in entries) { for (e in entries) {
val label = "${e.model.transport.name.lowercase()}:${e.model.id}" val label = "${e.model.transport.name.lowercase()}:${e.model.id}"
val (ok, detail, rtt) = attempt(e.handle) val a = attempt(e.handle)
perNetwork[label] = detail perNetwork[label] = e.model.id to a
if (ok) { anyOk = true; rtt?.let(rtts::add) } if (a.ok) { anyOk = true; a.rttMs?.let(rtts::add) }
} }
// Per-network results are recorded structurally, not just as prose. The aggregate status
// can only say "some network answered"; a finding needs to know *which* network failed,
// and recovering that by parsing a human-readable detail string would be a trap waiting to
// spring the first time the wording changes.
val evidence: JsonObject = buildJsonObject { val evidence: JsonObject = buildJsonObject {
put("target", target) put("target", target)
for ((k, v) in perNetwork) put(k, v) for ((label, r) in perNetwork) {
val (netId, a) = r
put(label, buildJsonObject {
netId?.let { put("network_ref", it) }
put("ok", a.ok)
put("attempted", a.attempted)
put("detail", a.detail)
})
}
} }
val metrics: JsonObject = buildJsonObject { val metrics: JsonObject = buildJsonObject {
put("networks_ok", rtts.size) put("networks_ok", rtts.size)
@@ -69,56 +75,9 @@ class IcmpProbe(
b.build(status, evidence = evidence, metrics = metrics) b.build(status, evidence = evidence, metrics = metrics)
} }
private data class Attempt(val ok: Boolean, val detail: String, val rttMs: Double?) /** The echo exchange itself lives in [IcmpEcho], shared with the long-run ping series. */
private fun attempt(network: Network?): IcmpEcho.Result =
private fun attempt(network: Network?): Attempt { IcmpEcho.ping(network, target, v6, timeoutMs = 3000)
var fd: FileDescriptor? = null
return try {
val proto = if (v6) OsConstants.IPPROTO_ICMPV6 else OsConstants.IPPROTO_ICMP
val family = if (v6) OsConstants.AF_INET6 else OsConstants.AF_INET
fd = Os.socket(family, OsConstants.SOCK_DGRAM, proto)
network?.bindSocket(fd)
Os.setsockoptTimeval(fd, OsConstants.SOL_SOCKET, OsConstants.SO_RCVTIMEO, StructTimeval.fromMillis(3000))
val addr = network?.getByName(target) ?: InetAddress.getByName(target)
val ident = (Os.getpid() and 0xFFFF)
val packet = buildEchoRequest(v6, ident.toShort(), 1)
val t0 = System.nanoTime()
Os.sendto(fd, packet, 0, packet.size, 0, addr, 0)
val buf = ByteBuffer.allocate(1500)
val received = Os.recvfrom(fd, buf, 0, null)
val rttMs = (System.nanoTime() - t0) / 1_000_000.0
val replyType = if (received > 0) buf.get(0).toInt() and 0xFF else -1
val ok = replyType == (if (v6) 129 else 0)
Attempt(ok, "reply type=$replyType rtt_ms=${"%.1f".format(Locale.ROOT, rttMs)} bytes=$received", if (ok) rttMs else null)
} catch (e: Throwable) {
Attempt(false, "error: ${e.message ?: e.javaClass.simpleName}", null)
} finally {
fd?.let { runCatching { Os.close(it) } }
}
}
private fun buildEchoRequest(v6: Boolean, ident: Short, seq: Short): ByteArray {
val type = if (v6) 128 else 8
val payload = "echolot".toByteArray()
val pkt = ByteBuffer.allocate(8 + payload.size)
pkt.put(type.toByte()); pkt.put(0); pkt.putShort(0)
pkt.putShort(ident); pkt.putShort(seq); pkt.put(payload)
val bytes = pkt.array()
if (!v6) {
val cs = checksum(bytes)
bytes[2] = (cs.toInt() shr 8).toByte(); bytes[3] = (cs.toInt() and 0xFF).toByte()
}
return bytes
}
private fun checksum(b: ByteArray): Short {
var sum = 0; var i = 0
while (i < b.size - 1) { sum += ((b[i].toInt() and 0xFF) shl 8) or (b[i + 1].toInt() and 0xFF); i += 2 }
if (i < b.size) sum += (b[i].toInt() and 0xFF) shl 8
while (sum shr 16 != 0) sum = (sum and 0xFFFF) + (sum shr 16)
return sum.inv().toShort()
}
private companion object { private companion object {
fun round1(v: Double) = Math.round(v * 10.0) / 10.0 fun round1(v: Double) = Math.round(v * 10.0) / 10.0
@@ -0,0 +1,124 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.probe
import android.content.Context
import app.echo_lot.measurement.Test
import app.echo_lot.measurement.TestError
import app.echo_lot.measurement.TestType
import app.echo_lot.measurement.Tier
import kotlinx.serialization.json.JsonObject
import kotlinx.serialization.json.buildJsonObject
import kotlinx.serialization.json.put
/**
* local.llmnr_inventory who is resolving names with LLMNR on this segment, and what for.
*
* The measurement is twofold, and the second half is the one people underestimate:
*
* 1. **Which hostnames are being looked up.** An LLMNR query is a device saying out loud, to
* everyone, the name of something it wants to reach. Over a window that is a map of who talks
* to whom and, when the names are things like `wpad` or a server that no longer exists, a map
* of what is failing to resolve through DNS and falling back.
* 2. **That LLMNR is in use at all.** LLMNR (and its NetBIOS sibling) is a name-resolution
* fallback that trusts whoever answers first, which is the mechanism behind the standard
* credential-relay attack on Windows networks. Its mere presence on a segment is a finding
* independent of any individual query which is why this earns a test id rather than folding
* into the mDNS inventory.
*
* Purely passive: queries are broadcast to the group, so listening is the entire measurement, and
* answering or querying would make this device a participant in exactly the trust relationship the
* measurement is about.
*/
class LlmnrCollector : BaseCollector() {
override val type = TestType.LOCAL_LLMNR_INVENTORY
override val tier = Tier.APP
private val capture = MulticastCapture(
label = "llmnr",
port = PORT,
group4 = GROUP4,
// ff02::1:3 is the IPv6 LLMNR group. Joined where IPv6 exists; a v4-only network simply
// records an empty joined_v6 rather than an error, because that is not one.
group6 = GROUP6,
)
private var startedAtMonoNs = 0L
override suspend fun start(ctx: Context, ids: ProbeIds) {
begin(ids)
startedAtMonoNs = ids.monoNs()
capture.acquireLock(ctx)
capture.start(ids)
}
override suspend fun stop(): Test {
val packets = capture.stop()
val table = DiscoveryTable()
var queries = 0
var responses = 0
var undecodable = 0
for (p in packets) {
val q = LlmnrParser.parse(p.data, p.data.size)
if (q == null) {
undecodable++
continue
}
if (q.isQuery) queries++ else responses++
// Responses are normally unicast back to the querier, so what lands here is
// overwhelmingly queries — but a response that does reach the group is still a device
// claiming a name, which is worth the same row.
table.observe(p.sourceIp, "${q.name}/${q.qtype}", p.atMonoNs) {
buildJsonObject {
put("query_name", q.name)
put("qtype", LlmnrParser.qtypeName(q.qtype))
put("kind", if (q.isQuery) "query" else "response")
}
}
}
val evidence: JsonObject = buildJsonObject {
put("capture", capture.statusJson())
put("llmnr_queries", table.toJson())
put("evidence_truncated", capture.truncated || table.overflowed)
}
val metrics: JsonObject = buildJsonObject {
put("distinct_sources", table.distinctSources)
put("distinct_names", table.size)
put("packets", capture.packetsSeen)
put("queries", queries)
put("responses", responses)
put("undecodable_packets", undecodable)
// The headline: whether this protocol is live here at all. A boolean rather than an
// inference from a count, so a reader (or a future finding rule) never has to decide
// what "zero packets" meant — the capture's own status says whether zero is trustworthy.
put("llmnr_in_use", table.distinctSources > 0)
put("truncated", capture.truncated || table.overflowed)
}
val saw = !table.isEmpty
return build(
capture.outcome(saw),
evidence = evidence,
metrics = metrics,
params = params(),
error = capture.reason(saw)?.let { TestError("listen_incomplete", it) },
)
}
private fun params(): JsonObject = buildJsonObject {
put("mode", "passive")
put("group_v4", "$GROUP4:$PORT")
put("group_v6", "[$GROUP6]:$PORT")
put("started_mono_ns", startedAtMonoNs)
}
private companion object {
const val PORT = 5355
const val GROUP4 = "224.0.0.252"
const val GROUP6 = "ff02::1:3"
}
}
@@ -0,0 +1,115 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.probe
import android.content.Context
import android.net.nsd.NsdManager
import android.net.nsd.NsdServiceInfo
import android.net.wifi.WifiManager
import app.echo_lot.measurement.Test
import app.echo_lot.measurement.TestStatus
import app.echo_lot.measurement.TestType
import app.echo_lot.measurement.Tier
import kotlinx.coroutines.Dispatchers
import kotlinx.coroutines.delay
import kotlinx.coroutines.withContext
import kotlinx.serialization.json.buildJsonObject
import kotlinx.serialization.json.put
import kotlinx.serialization.json.putJsonObject
import java.util.Collections
/**
* local.mdns_inventory what answers mDNS on this network (MulticastLock + NSD discovery).
* The service inventory doubles as the VLAN-leakage detector: a chromecast answering on the
* guest wifi is a segmentation fault made visible. Folded from the prober, validated on both
* known devices (4 services each).
*
* Two hardware-bought lessons are load-bearing here:
* - The `_services._dns-sd._udp.` meta-query returned 0 on BOTH devices while concrete types
* found live services NsdManager's meta-query support is unreliable across builds, so the
* concrete types are the measurement and the meta-query result is itself evidence.
* - 4 s of listening missed services that 10 s catches; mDNS answers straggle.
*
* That last lesson is why [listenMs] is a parameter. 10 s is the short-mode default because it is
* the shortest window that was not demonstrably lossy; a long run hands it the whole measurement
* window, since the curve does not stop at ten seconds devices announce on their own schedule,
* and a printer that is asleep answers when something else wakes it.
*/
class MdnsInventoryProbe(private val listenMs: Long = 10_000) : Probe {
override val type = TestType.LOCAL_MDNS_INVENTORY
override val tier = Tier.APP
override val estimatedMs = listenMs + 500
/** Meta-query + common concrete types (HTTP covers HA/printers/NAS; googlecast is ubiquitous). */
private val queries = listOf(
"meta" to "_services._dns-sd._udp.",
"http" to "_http._tcp.",
"googlecast" to "_googlecast._tcp.",
)
private class Recorder : NsdManager.DiscoveryListener {
val names: MutableList<String> = Collections.synchronizedList(mutableListOf())
@Volatile var started = false
@Volatile var startFailCode: Int? = null
override fun onStartDiscoveryFailed(t: String?, code: Int) { startFailCode = code }
override fun onStopDiscoveryFailed(t: String?, code: Int) {}
override fun onDiscoveryStarted(t: String?) { started = true }
override fun onDiscoveryStopped(t: String?) {}
override fun onServiceFound(s: NsdServiceInfo?) { s?.serviceName?.let { names.add(it) } }
override fun onServiceLost(s: NsdServiceInfo?) {}
}
override suspend fun run(ctx: Context, ids: ProbeIds): Test = withContext(Dispatchers.IO) {
val b = TestBuilder(type, tier, ids)
val wifi = ctx.getSystemService(WifiManager::class.java)
val lock = wifi?.createMulticastLock("echolot")?.apply {
setReferenceCounted(false)
runCatching { acquire() }
}
val nsd = ctx.getSystemService(NsdManager::class.java)
?: return@withContext b.build(
TestStatus.UNSUPPORTED,
evidence = buildJsonObject { put("reason", "NsdManager unavailable") },
)
val recorders = queries.map { (label, type) ->
val r = Recorder()
runCatching { nsd.discoverServices(type, NsdManager.PROTOCOL_DNS_SD, r) }
.onFailure { r.startFailCode = -1 }
Triple(label, type, r)
}
try {
delay(listenMs)
var total = 0
var anyStarted = false
val evidence = buildJsonObject {
put("multicast_lock", lock?.isHeld == true)
for ((label, type, r) in recorders) {
runCatching { nsd.stopServiceDiscovery(r) }
anyStarted = anyStarted || r.started
val names = r.names.distinct()
total += names.size
putJsonObject(label) {
put("query", type)
put("started", r.started)
r.startFailCode?.let { put("start_fail_code", it) }
put("found", names.size)
if (names.isNotEmpty()) put("names", names.joinToString(", ").take(300))
}
}
}
val metrics = buildJsonObject { put("services_found", total) }
// How long it listened belongs in params: "4 services" means something different after
// ten seconds than after five minutes, and the number alone cannot say which it was.
val params = buildJsonObject { put("listen_ms", listenMs) }
// Zero services on a started discovery is a legitimate result (an empty or properly
// isolated network), not a failure — only discovery refusing to start is one.
b.build(if (anyStarted) TestStatus.OK else TestStatus.FAILED,
evidence = evidence, metrics = metrics, params = params)
} finally {
recorders.forEach { (_, _, r) -> runCatching { nsd.stopServiceDiscovery(r) } }
runCatching { lock?.release() }
}
}
}
@@ -0,0 +1,420 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.probe
import android.content.Context
import android.net.wifi.WifiManager
import app.echo_lot.measurement.TestStatus
import kotlinx.coroutines.CoroutineScope
import kotlinx.coroutines.Dispatchers
import kotlinx.coroutines.Job
import kotlinx.coroutines.SupervisorJob
import kotlinx.coroutines.launch
import kotlinx.coroutines.withContext
import kotlinx.serialization.json.JsonObject
import kotlinx.serialization.json.buildJsonObject
import kotlinx.serialization.json.put
import kotlinx.serialization.json.putJsonArray
import kotlinx.serialization.json.add
import java.net.DatagramPacket
import java.net.Inet4Address
import java.net.Inet6Address
import java.net.InetAddress
import java.net.InetSocketAddress
import java.net.MulticastSocket
import java.net.NetworkInterface
import java.net.SocketAddress
import java.net.SocketTimeoutException
/**
* The listening half of every passive-discovery collector: hold a multicast lock, bind a
* well-known UDP port, join a group on every interface that will take it, and retain the datagrams
* that arrive until the window closes.
*
* Four things here are the difference between a working listener and one that silently reports an
* empty network, and each of them has cost somebody a day:
*
* - **The MulticastLock.** Android's wifi stack filters out frames not addressed to this device
* unless an app holds one. Without it every collector built on this returns "nothing on this
* network" on a network full of chatter — the single most common reason a listener like this
* appears to work and measures nothing. Whether it was actually held is therefore reported, not
* assumed: an unheld lock plus silence is not a clean result and [outcome] refuses to call it one.
* - **SO_REUSEADDR before bind.** The well-known discovery ports are shared by construction the
* system's own mDNS/SSDP responders and any other app doing this are already there so binding
* exclusively fails on exactly the networks worth measuring.
* - **Joining per interface.** The group must be joined on the interface the traffic arrives on,
* which is not necessarily the default route: a phone with wifi plus a VPN plus cellular has
* three, and the LAN chatter is on the one that is not carrying the default route.
* - **A bound buffer, bounded per source too.** A chatty network must not decide how much memory
* this uses, and one device announcing every two seconds must not be able to fill the buffer and
* hide the twenty quieter devices behind it. Both caps set [truncated] rather than being silent.
*
* Nothing here throws. Every failure lands in [failure] (the capture is dead) or [degraded] (it is
* running but cannot see everything), which the owning collector turns into a recorded status.
*/
internal class MulticastCapture(
/** Names the MulticastLock, so a `dumpsys wifi` during a run says which listener holds what. */
private val label: String,
private val port: Int,
private val group4: String? = null,
private val group6: String? = null,
private val maxPackets: Int = 400,
private val maxBytesRetained: Int = 128 * 1024,
private val maxPerSource: Int = 24,
/**
* Whether to fall back to an ephemeral port when the well-known one cannot be bound.
*
* Only useful for a protocol with an active half: replies to our own searches come back to
* whatever port we sent from, so an ephemeral socket still collects those but it can never
* see the unsolicited announcements, which is why it is [degraded] and not normal operation.
*/
private val allowEphemeralFallback: Boolean = false,
) {
/** One retained datagram. Not a data class: the payload is a ByteArray, and structural equality
* over it would be both wrong and expensive. */
class Packet(val sourceIp: String, val atMonoNs: Long, val data: ByteArray)
/** Set when the capture could not be brought up at all; the collector reports `unsupported`. */
var failure: String? = null
private set
/** Set when the capture is running but blind to part of what it exists to see. */
var degraded: String? = null
private set
var lockHeld: Boolean = false
private set
var boundPort: Int = 0
private set
var packetsSeen: Int = 0
private set
var truncated: Boolean = false
private set
val joined4 = ArrayList<String>()
val joined6 = ArrayList<String>()
private val joinErrors = ArrayList<String>()
/** Whether this protocol needs a group join at all — NetBIOS is broadcast, not multicast. */
private val expectsGroup = group4 != null || group6 != null
private val packets = ArrayList<Packet>()
private val perSource = HashMap<String, Int>()
private var retainedBytes = 0
@Volatile private var running = false
private var socket: MulticastSocket? = null
private var lock: WifiManager.MulticastLock? = null
private var scope: CoroutineScope? = null
/**
* Brings the listener up and returns whether it is capturing. Setup is a handful of syscalls,
* so it finishes in milliseconds and the [Collector.start] promptness contract holds; only the
* receive loop is handed to a background scope.
*/
suspend fun start(ids: ProbeIds): Boolean = withContext(Dispatchers.IO) {
val s = bind() ?: return@withContext false
socket = s
running = true
// A degraded (ephemeral-port) socket is deliberately not joined to the group: it would be
// joining for a port nothing sends to, and the resulting "joined wlan0" in the evidence
// would claim a capability the capture does not have.
if (expectsGroup && degraded == null) joinGroups(s)
scope = CoroutineScope(SupervisorJob() + Dispatchers.IO).also { it.launch { pump(s, ids) } }
true
}
/** Acquires the multicast lock. Separate from [start] because it needs a Context and the rest
* does not, and because whether it succeeded is itself evidence. */
fun acquireLock(ctx: Context) {
val wifi = runCatching { ctx.applicationContext.getSystemService(WifiManager::class.java) }
.getOrNull()
lock = runCatching {
wifi?.createMulticastLock("echolot-$label")?.apply {
setReferenceCounted(false)
acquire()
}
}.getOrNull()
lockHeld = runCatching { lock?.isHeld == true }.getOrDefault(false)
}
/** Sends from the capture socket, so replies land back in this capture rather than on a second
* socket nobody is reading. Returns whether the datagram left the device. */
fun send(payload: ByteArray, host: String, toPort: Int): Boolean = runCatching {
val s = socket ?: return false
s.send(DatagramPacket(payload, payload.size, InetSocketAddress(host, toPort)))
true
}.getOrDefault(false)
/** Stops listening and hands back what was retained. Safe after a failed [start], and never
* throws a cancelled run still owes the caller the packets it did see. */
fun stop(): List<Packet> {
running = false
// Closed before the coroutine is cancelled: a blocking receive() does not notice
// cancellation, and closing the socket is what makes it return.
runCatching { socket?.close() }
socket = null
scope?.coroutineContext?.get(Job)?.cancel()
scope = null
runCatching { lock?.release() }
lock = null
// lockHeld deliberately survives the release: it records whether the capture *could* hear
// multicast while it ran, which is what [outcome] needs to decide whether silence is a
// fact about the network. Clearing it here would make every quiet network report that the
// lock was missing.
return synchronized(packets) { packets.toList() }
}
// ---- status ------------------------------------------------------------------------------
/**
* The status this capture's Test deserves, given whether anything was decoded.
*
* The interesting case is the last one. Silence on a network is a legitimate and useful
* finding but only when we know the listener could have heard something. Without the
* multicast lock, or without a single successful group join, "nothing was seen" describes this
* app and not the network, and reporting it as `ok` would be the collector lying by omission.
*/
fun outcome(sawAnything: Boolean): TestStatus = when {
failure != null -> TestStatus.UNSUPPORTED
degraded != null -> TestStatus.PARTIAL
sawAnything -> TestStatus.OK
expectsGroup && joined4.isEmpty() && joined6.isEmpty() -> TestStatus.PARTIAL
!lockHeld -> TestStatus.PARTIAL
else -> TestStatus.OK
}
/** Why [outcome] was not OK, in words, or null when it was. */
fun reason(sawAnything: Boolean): String? = when {
failure != null -> failure
degraded != null -> degraded
sawAnything -> null
expectsGroup && joined4.isEmpty() && joined6.isEmpty() ->
"no interface accepted the group join, so silence here says nothing about the network"
!lockHeld ->
"the wifi multicast lock was not held, so silence here says nothing about the network"
else -> null
}
/** The capture's own facts, for the collector's evidence. Every collector reports these
* identically so that "saw nothing" can always be told apart from "could not listen". */
fun statusJson(): JsonObject = buildJsonObject {
put("multicast_lock", lockHeld)
put("bound_port", boundPort)
putJsonArray("joined_v4") { for (n in joined4) add(n) }
putJsonArray("joined_v6") { for (n in joined6) add(n) }
if (joinErrors.isNotEmpty()) {
put("join_errors", joinErrors.take(6).joinToString("; ").take(400))
}
degraded?.let { put("degraded", it) }
failure?.let { put("failure", it) }
put("packets_seen", packetsSeen)
put("packets_retained", synchronized(packets) { packets.size })
put("evidence_truncated", truncated)
}
// ---- internals ---------------------------------------------------------------------------
private fun bind(): MulticastSocket? {
// Unbound first, so SO_REUSEADDR is set *before* bind — setting it afterwards has no
// effect, and these ports are always already in use by something.
runCatching {
val s = MulticastSocket(null as SocketAddress?)
s.reuseAddress = true
s.bind(InetSocketAddress(port))
s.soTimeout = SO_TIMEOUT_MS
boundPort = port
return s
}.onFailure { first ->
val why = describe(first)
if (!allowEphemeralFallback) {
// Ports below 1024 are privileged on Android as on any Linux, so this is the
// expected outcome for NetBIOS and is a finding rather than a bug — see
// NetbiosCollector.
failure = "could not bind UDP $port: $why"
return null
}
runCatching {
val s = MulticastSocket(null as SocketAddress?)
s.reuseAddress = true
s.bind(InetSocketAddress(0))
s.soTimeout = SO_TIMEOUT_MS
boundPort = s.localPort
degraded = "UDP $port could not be bound ($why); listening on an ephemeral port " +
"instead, so only replies to our own searches are visible and unsolicited " +
"announcements are not"
return s
}.onFailure { failure = "could not bind UDP $port ($why) or any ephemeral port" }
}
return null
}
private fun joinGroups(s: MulticastSocket) {
val ifaces = runCatching { NetworkInterface.getNetworkInterfaces()?.toList() }
.getOrNull().orEmpty()
.filter {
runCatching { it.isUp && !it.isLoopback && it.supportsMulticast() }
.getOrDefault(false)
}
val g4 = group4?.let { runCatching { InetAddress.getByName(it) }.getOrNull() }
val g6 = group6?.let { runCatching { InetAddress.getByName(it) }.getOrNull() }
for (ni in ifaces) {
val addrs = runCatching { ni.inetAddresses.toList() }.getOrNull().orEmpty()
if (g4 != null && addrs.any { it is Inet4Address }) {
join(s, g4, ni)?.let { joined4.add(it) }
}
// A network with no IPv6 address is not a failure to report as an error — it is a
// v4-only network, which is most of them. Only interfaces that could have joined are
// asked to.
if (g6 != null && addrs.any { it is Inet6Address }) {
join(s, g6, ni)?.let { joined6.add(it) }
}
}
// Last resort: let the kernel pick the interface. Some vendor builds refuse the explicit
// form on the very interface that carries the traffic, and a default-interface join is
// better than no listener at all.
if (joined4.isEmpty() && joined6.isEmpty()) {
@Suppress("DEPRECATION")
(g4 ?: g6)?.let { g ->
runCatching { s.joinGroup(g) }
.onSuccess { (if (g is Inet4Address) joined4 else joined6).add("(default)") }
.onFailure { joinErrors.add("default: ${describe(it)}") }
}
}
}
private fun join(s: MulticastSocket, group: InetAddress, ni: NetworkInterface): String? =
runCatching {
s.joinGroup(InetSocketAddress(group, boundPort), ni)
ni.name
}.onFailure {
joinErrors.add("${ni.name}/${if (group is Inet4Address) "v4" else "v6"}: ${describe(it)}")
}.getOrNull()
private fun pump(s: MulticastSocket, ids: ProbeIds) {
val buf = ByteArray(READ_BUFFER)
while (running) {
val p = DatagramPacket(buf, buf.size)
try {
s.receive(p)
} catch (t: SocketTimeoutException) {
// The timeout exists only so this loop notices `running` going false if the socket
// close somehow does not wake it. Nothing to record.
continue
} catch (t: Throwable) {
// The normal exit: stop() closed the socket underneath us. It is also what a
// vanishing interface looks like, and neither is worth a status of its own — the
// packets gathered so far are still the measurement.
return
}
packetsSeen++
val ip = p.address?.hostAddress ?: continue
val len = p.length
if (len <= 0) continue
synchronized(packets) {
val fromThis = perSource[ip] ?: 0
if (packets.size >= maxPackets ||
retainedBytes + len > maxBytesRetained ||
fromThis >= maxPerSource
) {
truncated = true
} else {
packets.add(Packet(ip, ids.monoNs(), p.data.copyOf(len)))
perSource[ip] = fromThis + 1
retainedBytes += len
}
}
}
}
private companion object {
const val SO_TIMEOUT_MS = 1_000
/** Larger than any discovery datagram anyone sends; oversized ones are truncated by the
* kernel, which the parsers survive by design. */
const val READ_BUFFER = 4_096
fun describe(t: Throwable): String =
(t.message ?: t.javaClass.simpleName).take(160)
}
}
/**
* The other half of what every discovery collector does: fold a stream of decoded messages into a
* deduplicated inventory of *things*, not packets.
*
* Deduplication is by source **and** identity, never by identity alone. Two devices announcing the
* same UPnP service type are two devices, and collapsing them would turn the one measurement worth
* having (how many things are on this segment) into a count of protocols. Conversely one device
* re-announcing every thirty seconds must not appear ten times that is what [count] is for, and
* the repetition rate is itself readable from count over the window length.
*
* First-seen is kept on the monotonic clock per the two-clock rule: it answers "was this device
* here from the start, or did it appear four minutes in", which is exactly the question a long run
* exists to answer and the one a wall-clock stamp cannot be trusted for.
*/
internal class DiscoveryTable(private val maxEntries: Int = MAX_ENTRIES) {
private class Row(val firstSeenMonoNs: Long, val fields: JsonObject) {
var count = 0
}
private val rows = LinkedHashMap<Pair<String, String>, Row>()
private val sources = HashSet<String>()
/** True when a network was busy enough that entries had to be dropped — reported, never hidden. */
var overflowed = false
private set
/**
* [fields] is a lambda so the JSON for a repeat sighting is never built: on a chatty segment
* the overwhelming majority of packets are the same device saying the same thing again.
*/
fun observe(sourceIp: String, identity: String, atMonoNs: Long, fields: () -> JsonObject) {
sources.add(sourceIp)
val key = sourceIp to identity
val existing = rows[key]
if (existing != null) {
existing.count++
return
}
if (rows.size >= maxEntries) {
overflowed = true
return
}
val built = buildJsonObject {
put("source_ip", sourceIp)
for ((k, v) in fields()) put(k, v)
}
rows[key] = Row(atMonoNs, built).also { it.count = 1 }
}
val distinctSources: Int get() = sources.size
val size: Int get() = rows.size
val isEmpty: Boolean get() = rows.isEmpty()
fun toJson(): kotlinx.serialization.json.JsonArray = kotlinx.serialization.json.JsonArray(
rows.values.map { r ->
JsonObject(
r.fields + mapOf(
"first_seen_mono_ns" to kotlinx.serialization.json.JsonPrimitive(r.firstSeenMonoNs),
"count" to kotlinx.serialization.json.JsonPrimitive(r.count),
)
)
}
)
private companion object {
/** Generous enough for any real segment, small enough that a broadcast storm cannot make
* one test dominate the document. */
const val MAX_ENTRIES = 200
}
}
@@ -0,0 +1,131 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.probe
import android.content.Context
import app.echo_lot.measurement.Test
import app.echo_lot.measurement.TestError
import app.echo_lot.measurement.TestType
import app.echo_lot.measurement.Tier
import kotlinx.serialization.json.JsonObject
import kotlinx.serialization.json.buildJsonObject
import kotlinx.serialization.json.put
/**
* local.netbios_inventory NetBIOS name-service chatter (UDP 137) on this segment.
*
* NetBIOS name registration and query traffic is broadcast, not multicast, so there is no group to
* join: a socket bound to 137 sees it by being on the segment. Every packet carries a
* first-level-encoded name plus a suffix byte saying what the name is *for* a workstation, a file
* server, a master browser which makes a passive window over it a Windows-side inventory of the
* LAN, and, like LLMNR, a security observation in its own right: NBT-NS is the other half of the
* classic name-resolution spoofing surface.
*
* **Expect this to report `unsupported` on the app tier, and read that as a result rather than a
* bug.** 137 is below 1024, and Linux Android included reserves those ports for privileged
* processes; an unprivileged app UID cannot bind one. So the honest measurement here is usually
* "this tier cannot observe NetBIOS on this device", recorded with the exact bind error, rather
* than a silent absence that reads as a quiet network. The observation belongs to the Shizuku tier,
* whose shell UID can bind it; the collector is written now so that the decoder, the evidence shape
* and the registry id are settled and tested by the time that lands.
*
* An active alternative exists send an NBSTAT query from an ephemeral port and read the unicast
* replies and is deliberately not taken: it is a host sweep, which is scanning rather than
* measuring, and it would change what the run does to the network it is observing.
*/
class NetbiosCollector : BaseCollector() {
override val type = TestType.LOCAL_NETBIOS_INVENTORY
override val tier = Tier.APP
private val capture = MulticastCapture(
label = "netbios",
port = PORT,
// No group: NBT-NS is subnet broadcast. The multicast lock is still taken, because
// Android's wifi firmware filters broadcast as well as multicast under power save.
group4 = null,
group6 = null,
// Registration bursts repeat the same name several times a second; the per-source cap is
// what keeps one noisy Windows box from filling the buffer.
maxPerSource = 12,
)
private var startedAtMonoNs = 0L
override suspend fun start(ctx: Context, ids: ProbeIds) {
begin(ids)
startedAtMonoNs = ids.monoNs()
capture.acquireLock(ctx)
capture.start(ids)
}
override suspend fun stop(): Test {
val packets = capture.stop()
val table = DiscoveryTable()
var queries = 0
var registrations = 0
var undecodable = 0
for (p in packets) {
val n = NetbiosParser.parse(p.data, p.data.size)
if (n == null) {
undecodable++
continue
}
when (n.opcode) {
0 -> queries++
5, 8 -> registrations++
}
// Identity is name + suffix, not the name alone: one host registers the same name
// several times with different suffixes, and those are different facts about it.
table.observe(p.sourceIp, n.name + "#" + "%02X".format(n.suffix), p.atMonoNs) {
buildJsonObject {
put("netbios_name", n.name)
put("suffix", "0x%02X".format(n.suffix))
put("role", n.role)
put("operation", NetbiosParser.opcodeName(n.opcode))
put("kind", if (n.isResponse) "response" else "request")
}
}
}
val evidence: JsonObject = buildJsonObject {
put("capture", capture.statusJson())
put("netbios_names", table.toJson())
put("evidence_truncated", capture.truncated || table.overflowed)
}
val metrics: JsonObject = buildJsonObject {
put("distinct_sources", table.distinctSources)
put("distinct_names", table.size)
put("packets", capture.packetsSeen)
put("name_queries", queries)
put("name_registrations", registrations)
put("undecodable_packets", undecodable)
put("netbios_in_use", table.distinctSources > 0)
put("truncated", capture.truncated || table.overflowed)
}
val saw = !table.isEmpty
return build(
capture.outcome(saw),
evidence = evidence,
metrics = metrics,
params = params(),
error = capture.reason(saw)?.let {
TestError(if (capture.failure != null) "port_unavailable" else "listen_incomplete", it)
},
)
}
private fun params(): JsonObject = buildJsonObject {
put("mode", "passive")
put("port", PORT)
put("transport", "udp broadcast")
put("started_mono_ns", startedAtMonoNs)
}
private companion object {
const val PORT = 137
}
}
@@ -0,0 +1,274 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.probe
import android.content.Context
import android.net.ConnectivityManager
import android.net.LinkProperties
import android.net.Network
import android.net.NetworkCapabilities
import android.net.NetworkRequest
import app.echo_lot.measurement.NetworkChange
import app.echo_lot.measurement.NetworkChanges
import app.echo_lot.measurement.Test
import app.echo_lot.measurement.TestError
import app.echo_lot.measurement.TestStatus
import app.echo_lot.measurement.TestType
import app.echo_lot.measurement.Tier
import kotlinx.serialization.json.JsonObject
import kotlinx.serialization.json.buildJsonObject
import kotlinx.serialization.json.put
import kotlinx.serialization.json.putJsonArray
import kotlinx.serialization.json.addJsonObject
import java.util.Collections
/**
* Watches every network for the whole run and fills `networks[].changes[]` (measurement-schema §4).
*
* That array has been in the schema since the first draft and has never been populated, because
* nothing in a battery of one-shot probes is in a position to fill it. It is the single most
* valuable thing a long run adds: a wifi link that drops and returns mid-window is invisible to
* every probe the ones before and after the gap both succeed and it is exactly the fault people
* open a network diagnostic to chase.
*
* Reported as [TestType.LINK_IP_MONITOR] at app tier. The registry lists that type as Shizuku's
* (`ip monitor`), and this is deliberately the same observation from a tier that does not need it:
* link state as it changes over time. A document may therefore carry two `link.ip_monitor` tests,
* told apart by `tier` which is what `tier` is for.
*/
class NetworkChangeCollector(
private val entries: List<NetworkInventory.Entry>,
) : BaseCollector() {
override val type = TestType.LINK_IP_MONITOR
override val tier = Tier.APP
private data class Change(
val atMonoNs: Long,
val kind: String,
val event: String,
val iface: String,
val detail: String,
)
private val changes: MutableList<Change> = Collections.synchronizedList(mutableListOf())
private var cm: ConnectivityManager? = null
private var callback: ConnectivityManager.NetworkCallback? = null
private var registerError: String? = null
/**
* Last seen state per network, so only real changes are recorded.
*
* Both capability and link-property callbacks fire constantly on a live device signal
* strength alone re-delivers capabilities every few seconds and a five-minute window of that
* would bury the four events that matter under several hundred that do not. The first callback
* after a network appears is the baseline, not a change.
*/
private val lastCaps = HashMap<String, String>()
private val lastLink = HashMap<String, String>()
/** Interface name per network handle, remembered because `onLost` can no longer look it up. */
private val ifaceOf = HashMap<String, String>()
/**
* The networks that were already up when the window opened.
*
* `registerNetworkCallback` replays `onAvailable` for every matching network the instant it is
* registered, so without this every run would open with three "the wifi appeared" events that
* describe the registration and not the network. A link that drops and returns comes back as a
* new handle, which is not in this set, so real re-appearances are still recorded.
*/
private val seeded = HashSet<String>()
override suspend fun start(ctx: Context, ids: ProbeIds) {
begin(ids)
val manager = ctx.getSystemService(ConnectivityManager::class.java)
if (manager == null) {
registerError = "ConnectivityManager unavailable"
return
}
cm = manager
// Seed the interface names from the snapshot the run already took, so a network that is
// lost without ever having delivered a callback here is still attributable.
for (e in entries) {
e.model.iface?.let { ifaceOf[key(e.handle)] = it }
seeded.add(key(e.handle))
}
val cb = object : ConnectivityManager.NetworkCallback() {
override fun onAvailable(network: Network) {
val iface = resolveIface(network)
if (key(network) in seeded) return
record(ids, NetworkChanges.GAINED, "available", iface, "network became available")
}
override fun onLost(network: Network) {
val iface = ifaceOf[key(network)] ?: "(unknown)"
record(ids, NetworkChanges.LOST, "lost", iface, "network went away")
// Dropped so a returning link re-baselines instead of reporting every property it
// ever had as a change the moment it comes back.
lastCaps.remove(key(network)); lastLink.remove(key(network))
}
override fun onCapabilitiesChanged(network: Network, caps: NetworkCapabilities) {
val iface = resolveIface(network)
val print = capsFingerprint(caps)
val previous = lastCaps.put(key(network), print)
if (previous == null || previous == print) return
record(ids, NetworkChanges.LINK_CHANGED, "capabilities", iface, "$previous$print")
}
override fun onLinkPropertiesChanged(network: Network, lp: LinkProperties) {
lp.interfaceName?.let { ifaceOf[key(network)] = it }
val print = linkFingerprint(lp)
val previous = lastLink.put(key(network), print)
if (previous == null || previous == print) return
record(ids, NetworkChanges.LINK_CHANGED, "link_properties", lp.interfaceName ?: "(unknown)",
describeLinkDelta(previous, print))
}
override fun onLosing(network: Network, maxMsToLive: Int) {
val iface = ifaceOf[key(network)] ?: "(unknown)"
record(ids, NetworkChanges.LINK_CHANGED, "losing", iface, "about to be torn down in ${maxMsToLive} ms")
}
}
callback = cb
// clearCapabilities(), or the default request only matches INTERNET + NOT_RESTRICTED and
// the carrier's IMS/MMS networks — and, more importantly, a network in the middle of
// failing validation — never appear. The transports are named explicitly so this does not
// also follow whatever internal networks a vendor keeps in the list.
val request = NetworkRequest.Builder()
.clearCapabilities()
.addTransportType(NetworkCapabilities.TRANSPORT_WIFI)
.addTransportType(NetworkCapabilities.TRANSPORT_CELLULAR)
.addTransportType(NetworkCapabilities.TRANSPORT_ETHERNET)
.addTransportType(NetworkCapabilities.TRANSPORT_VPN)
.build()
runCatching { manager.registerNetworkCallback(request, cb) }
.onFailure { registerError = it.message ?: it.javaClass.simpleName; callback = null }
}
override suspend fun stop(): Test {
callback?.let { cb -> runCatching { cm?.unregisterNetworkCallback(cb) } }
callback = null
val snapshot = synchronized(changes) { changes.toList() }
if (registerError != null) {
return build(
TestStatus.UNSUPPORTED,
error = TestError("callback_unavailable", registerError),
)
}
val evidence: JsonObject = buildJsonObject {
putJsonArray("changes") {
for (c in snapshot) addJsonObject {
put("at_mono_ns", c.atMonoNs)
put("kind", c.kind)
put("event", c.event)
put("interface", c.iface)
put("detail", c.detail)
}
}
}
val perIface = snapshot.groupBy { it.iface }
val metrics: JsonObject = buildJsonObject {
put("changes_total", snapshot.size)
put("networks_lost", snapshot.count { it.kind == NetworkChanges.LOST })
put("networks_gained", snapshot.count { it.event == "available" })
put("link_changes", snapshot.count { it.kind == NetworkChanges.LINK_CHANGED })
// The number a reader actually wants: how many times a link went away and came back.
// Counted by the same function the flapping finding uses, so the metric and the
// finding can never tell different stories about one window.
put("flap_cycles", perIface.values.sumOf { NetworkChanges.flapCycles(it.map { c -> c.kind }) })
}
// Zero changes over the window is a real, useful result — a stable network — so it is OK
// rather than a failure. The window that produced it is what makes that mean anything, and
// it is recorded in run.mode plus the sibling collectors' params.
return build(TestStatus.OK, evidence = evidence, metrics = metrics)
}
/**
* The changes belonging to each network in `networks[]`, keyed by its model id.
*
* Matched by interface name rather than by Android's `Network` handle, because a link that
* drops and returns comes back as a *different* handle with the same interface and the
* flapping case is precisely the one this must not lose. Changes on an interface that was not
* in the run's initial snapshot stay in the test evidence but have no `networks[]` entry to
* hang from.
*/
fun changesByNetwork(): Map<String, List<NetworkChange>> {
val byIface = entries.mapNotNull { e -> e.model.iface?.let { it to e.model.id } }.toMap()
val out = LinkedHashMap<String, MutableList<NetworkChange>>()
for (c in synchronized(changes) { changes.toList() }) {
val id = byIface[c.iface] ?: continue
out.getOrPut(id) { mutableListOf() }.add(
NetworkChange(
atMonoNs = c.atMonoNs,
kind = c.kind,
detail = buildJsonObject {
put("event", c.event)
put("interface", c.iface)
put("detail", c.detail)
},
)
)
}
return out
}
private fun record(ids: ProbeIds, kind: String, event: String, iface: String, detail: String) {
changes.add(Change(ids.monoNs(), kind, event, iface, detail))
}
private fun resolveIface(network: Network): String {
val known = ifaceOf[key(network)]
if (known != null) return known
val name = runCatching { cm?.getLinkProperties(network)?.interfaceName }.getOrNull()
if (name != null) ifaceOf[key(network)] = name
return name ?: "(unknown)"
}
private companion object {
/** Android's own network id, stable for the life of one Network object. */
private fun key(n: Network): String = n.toString()
/**
* Only the capabilities whose change means something diagnostically.
*
* Bandwidth estimates and signal strength are deliberately excluded: they change every few
* seconds on a moving device, and including them turns a change log into a sampling log.
* VALIDATED and CAPTIVE_PORTAL are the two that matter most they are the moment Android
* decides a network does or does not carry the internet.
*/
private fun capsFingerprint(c: NetworkCapabilities): String = buildString {
fun flag(name: String, cap: Int) {
if (runCatching { c.hasCapability(cap) }.getOrDefault(false)) append(name).append(' ')
}
flag("internet", NetworkCapabilities.NET_CAPABILITY_INTERNET)
flag("validated", NetworkCapabilities.NET_CAPABILITY_VALIDATED)
flag("captive", NetworkCapabilities.NET_CAPABILITY_CAPTIVE_PORTAL)
flag("not-metered", NetworkCapabilities.NET_CAPABILITY_NOT_METERED)
flag("not-suspended", NET_CAPABILITY_NOT_SUSPENDED)
flag("not-restricted", NetworkCapabilities.NET_CAPABILITY_NOT_RESTRICTED)
}.trim().ifEmpty { "(none)" }
/** NetworkCapabilities.NET_CAPABILITY_NOT_SUSPENDED, API 28+ (@SystemApi constant). */
private const val NET_CAPABILITY_NOT_SUSPENDED = 21
private fun linkFingerprint(lp: LinkProperties): String {
val addrs = lp.linkAddresses.map { it.toString() }.sorted().joinToString(",")
val routes = lp.routes.map { it.toString() }.sorted().joinToString(",")
val dns = lp.dnsServers.mapNotNull { it.hostAddress }.sorted().joinToString(",")
return "mtu=${lp.mtu}|addr=$addrs|route=$routes|dns=$dns"
}
/** Names which part of the link changed, so the detail is readable without a diff tool. */
private fun describeLinkDelta(before: String, after: String): String {
val b = before.split('|'); val a = after.split('|')
val changed = b.indices.filter { it < a.size && b[it] != a[it] }
.map { a[it].substringBefore('=') }
return if (changed.isEmpty()) "changed" else "changed: ${changed.joinToString(", ")}"
}
}
}
@@ -39,6 +39,46 @@ object NetworkInventory {
return out return out
} }
/**
* Android's own verdict on the network, read straight from the capabilities it already has.
*
* PARTIAL_CONNECTIVITY only exists from API 28 and CAPTIVE_PORTAL from 23, so both are read
* defensively: an older platform that cannot answer should leave the field null rather than
* assert a false.
*/
private fun systemVerdict(caps: NetworkCapabilities): app.echo_lot.measurement.SystemVerdict =
app.echo_lot.measurement.SystemVerdict(
validated = caps.hasCapability(NetworkCapabilities.NET_CAPABILITY_VALIDATED),
captivePortal = runCatching {
caps.hasCapability(NetworkCapabilities.NET_CAPABILITY_CAPTIVE_PORTAL)
}.getOrNull(),
// NET_CAPABILITY_PARTIAL_CONNECTIVITY is @SystemApi, so the constant is not in the
// public SDK even though the platform sets it from API 28. The number is stable —
// changing it would break every system app that reads it — but this is a value the
// SDK does not promise us, so it is asked for defensively and reported as unknown
// rather than as false if anything about it is not as expected.
partialConnectivity = runCatching {
caps.hasCapability(NET_CAPABILITY_PARTIAL_CONNECTIVITY)
}.getOrNull(),
)
/** @SystemApi NetworkCapabilities.NET_CAPABILITY_PARTIAL_CONNECTIVITY, API 28+. */
private const val NET_CAPABILITY_PARTIAL_CONNECTIVITY = 24
/**
* Can an ordinary app send on this network?
*
* The carrier's special-purpose networks (IMS/VoLTE, MMS, XCAP) sit in `allNetworks` next to
* the real ones, carrying `IMS`/`MMS` but neither `INTERNET` nor `NOT_RESTRICTED`. Binding
* to them fails with EPERM forever, because it needs `CONNECTIVITY_USE_RESTRICTED_NETWORKS`
* signature-level, unobtainable for this app. Asking the capabilities up front is what
* separates "the OS will never let us measure this" from "something is blocking us", which
* a bind attempt alone cannot distinguish and which the constraint logic must not confuse.
*/
private fun appUsable(caps: NetworkCapabilities): Boolean =
caps.hasCapability(NetworkCapabilities.NET_CAPABILITY_INTERNET) &&
caps.hasCapability(NetworkCapabilities.NET_CAPABILITY_NOT_RESTRICTED)
private fun toModel(id: String, caps: NetworkCapabilities, lp: LinkProperties): MNetwork { private fun toModel(id: String, caps: NetworkCapabilities, lp: LinkProperties): MNetwork {
val transport = when { val transport = when {
caps.hasTransport(NetworkCapabilities.TRANSPORT_WIFI) -> Transport.WIFI caps.hasTransport(NetworkCapabilities.TRANSPORT_WIFI) -> Transport.WIFI
@@ -72,6 +112,8 @@ object NetworkInventory {
return MNetwork( return MNetwork(
id = id, transport = transport, iface = lp.interfaceName, id = id, transport = transport, iface = lp.interfaceName,
link = Link(mtu = lp.mtu.takeIf { it > 0 }, addresses = addresses, routes = routes, dns = dns), link = Link(mtu = lp.mtu.takeIf { it > 0 }, addresses = addresses, routes = routes, dns = dns),
systemVerdict = systemVerdict(caps),
appUsable = appUsable(caps),
) )
} }
} }
@@ -0,0 +1,70 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.probe
import android.system.Os
import java.io.FileDescriptor
/**
* Linux socket-option ABI numbers that android.system.OsConstants does NOT reliably expose.
* Stable across Android's supported ABIs at the IP/IPv6 protocol levels, which is why they can
* be hardcoded: if setsockoptInt with one of these succeeds, the kernel accepted the option; if
* it throws ErrnoException, it did not. Either outcome is data. Do not "fix" these to
* OsConstants names they don't exist there (validated in the prober; see its OsAbi.kt).
*
* Measured fact worth keeping: `Os.getsockoptInt` is absent on both known devices (OnePlus 15
* A16, Lenovo TB330FU A15), so path-MTU values must be read from the errqueue (`ee_info`), never
* from getsockopt(IP_MTU).
*/
object OsAbi {
// IP level
const val IP_TTL = 2
const val IP_MTU_DISCOVER = 10
const val IP_MTU = 14
const val IP_RECVERR = 11
const val IP_PMTUDISC_DO = 2 // set DF, honor PMTU
const val IP_PMTUDISC_PROBE = 3 // set DF, ignore PMTU (for probing)
// IPv6 level
const val IPV6_MTU_DISCOVER = 23
const val IPV6_MTU = 24
const val IPV6_RECVERR = 25
const val IPV6_UNICAST_HOPS = 16
const val IPV6_PMTUDISC_PROBE = 3
// recv flags — not in OsConstants on any current API level
const val MSG_ERRQUEUE = 0x2000
const val MSG_DONTWAIT = 0x40
// struct sock_extended_err (uapi/linux/errqueue.h), fixed layout on all Android ABIs:
// u32 ee_errno; u8 ee_origin; u8 ee_type; u8 ee_code; u8 ee_pad; u32 ee_info; u32 ee_data;
// followed directly by the offender sockaddr (SO_EE_OFFENDER).
const val SOCK_EE_SIZE = 16
const val SO_EE_ORIGIN_ICMP = 2
const val ICMP_TIME_EXCEEDED = 11
const val ICMP_DEST_UNREACH = 3
/** Try setsockoptInt; return null on success, or the errno name on failure. */
fun trySetIntOpt(fd: FileDescriptor, level: Int, opt: Int, value: Int): String? =
try {
Os.setsockoptInt(fd, level, opt, value)
null
} catch (e: Throwable) {
e.message ?: e.javaClass.simpleName
}
/**
* getsockoptInt is not part of the stable public Os surface on every API level, so it is
* reached via reflection; callers must treat failure as "unreadable", not as an error.
*/
fun tryGetIntOpt(fd: FileDescriptor, level: Int, opt: Int): Result<Int> = runCatching {
val m = Os::class.java.getMethod(
"getsockoptInt",
FileDescriptor::class.java,
Int::class.javaPrimitiveType,
Int::class.javaPrimitiveType,
)
m.invoke(null, fd, level, opt) as Int
}
}
@@ -0,0 +1,158 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.probe
import android.content.Context
import app.echo_lot.measurement.Test
import app.echo_lot.measurement.TestStatus
import app.echo_lot.measurement.TestType
import app.echo_lot.measurement.Tier
import kotlinx.coroutines.CoroutineScope
import kotlinx.coroutines.Dispatchers
import kotlinx.coroutines.Job
import kotlinx.coroutines.SupervisorJob
import kotlinx.coroutines.delay
import kotlinx.coroutines.isActive
import kotlinx.coroutines.launch
import kotlinx.serialization.json.JsonNull
import kotlinx.serialization.json.JsonObject
import kotlinx.serialization.json.JsonPrimitive
import kotlinx.serialization.json.buildJsonObject
import kotlinx.serialization.json.put
import kotlinx.serialization.json.putJsonArray
import kotlinx.serialization.json.add
import java.util.Collections
/**
* icmp.ping4 sampled across the window loss and jitter over minutes instead of one packet.
*
* The battery's [IcmpProbe] answers "does this network reply at all", which one echo can settle.
* It cannot answer "how often does it not", and that is the complaint people actually have:
* 2 % loss is invisible to a single ping and ruins a video call. A series over five minutes also
* catches loss that comes in bursts, which an average taken over ten packets in one second cannot
* distinguish from a clean link.
*
* Emitted as its own `icmp.ping4` test alongside the battery's. `params` carries the window and
* the interval precisely so the two are never mistaken for each other a reader seeing 300 sent
* packets in one and 1 in the other must be able to tell which is which without guessing.
*
* Sustained loss here deliberately emits **no finding**. Every loss code in the registry is about
* the server path `connectivity.udp_loss` and its directional siblings all say "UDP", and they
* mean the probe protocol's traffic, whose direction the server can attest to. ICMP echo to a
* public address is a different measurement with a different set of benign explanations (rate
* limiting at the target is the obvious one), and borrowing a code that claims otherwise would put
* two unrelated things under one dashboard entry the exact failure the registry exists to
* prevent. The metrics say what was seen; a code for it can be added when it has been defined.
*/
class PingSeriesCollector(
private val target: String = "1.1.1.1",
private val intervalMs: Long = 2_000,
/** Deliberately below [intervalMs]: a reply that arrives after the next probe was due is lost
* for any practical purpose, and waiting for it would make the series drift out of cadence. */
private val timeoutMs: Int = 1_500,
) : BaseCollector() {
override val type = TestType.ICMP_PING4
override val tier = Tier.APP
private val txMonoNs: MutableList<Long> = Collections.synchronizedList(mutableListOf())
private val rttMs: MutableList<Double?> = Collections.synchronizedList(mutableListOf())
private var notSent = 0
private var scope: CoroutineScope? = null
private var startedAtMonoNs = 0L
override suspend fun start(ctx: Context, ids: ProbeIds) {
begin(ids)
startedAtMonoNs = ids.monoNs()
// The default network, and only the default network: this measures what the device's own
// traffic experiences over the window. Per-network binding is the battery's job, and doing
// it here would multiply the packet rate by the number of interfaces for no new answer.
val s = CoroutineScope(SupervisorJob() + Dispatchers.IO).also { scope = it }
s.launch {
var seq = 1
while (isActive) {
val t0 = ids.monoNs()
val r = IcmpEcho.ping(null, target, v6 = false, timeoutMs = timeoutMs, seq = seq)
if (r.attempted) {
txMonoNs.add(t0)
rttMs.add(r.rttMs)
} else {
// Never left the device — a socket or bind failure is not packet loss, and
// counting it as loss would blame the network for the app's own trouble.
notSent++
}
seq = (seq + 1) and 0xFFFF
delay(intervalMs)
}
}
}
override suspend fun stop(): Test {
scope?.coroutineContext?.get(Job)?.cancel()
scope = null
val tx = synchronized(txMonoNs) { txMonoNs.toList() }
val rtt = synchronized(rttMs) { rttMs.toList() }
val received = rtt.filterNotNull()
val evidence: JsonObject = buildJsonObject {
putJsonArray("seq") { for (i in tx.indices) add(i) }
putJsonArray("t_tx_ns") { for (v in tx) add(v) }
// null at an index is a lost probe, per the §6.2 columnar convention.
putJsonArray("rtt_ms") {
for (v in rtt) add(v?.let { JsonPrimitive(round1(it)) } ?: JsonNull)
}
}
val metrics: JsonObject = buildJsonObject {
put("sent", tx.size)
put("received", received.size)
put("not_sent", notSent)
if (tx.isNotEmpty()) {
put("loss_pct", round1((tx.size - received.size) * 100.0 / tx.size))
}
if (received.isNotEmpty()) {
put("rtt_ms_min", round1(received.min()))
put("rtt_ms_avg", round1(received.average()))
put("rtt_ms_max", round1(received.max()))
put("jitter_ms", round1(meanDeviation(received)))
}
}
val status = when {
tx.isEmpty() -> TestStatus.UNSUPPORTED
received.isEmpty() -> TestStatus.FAILED
received.size < tx.size -> TestStatus.PARTIAL
else -> TestStatus.OK
}
return build(status, evidence = evidence, metrics = metrics, params = params())
}
private fun params(): JsonObject = buildJsonObject {
// What separates this from the battery's single ping, and what a reader needs to reproduce
// it. Without these two numbers "300 packets, 2 % loss" is a rate nobody can interpret.
put("mode", "series")
put("target", target)
put("interval_ms", intervalMs)
put("timeout_ms", timeoutMs)
put("started_mono_ns", startedAtMonoNs)
put("network", "default")
}
private companion object {
fun round1(v: Double) = Math.round(v * 10.0) / 10.0
/**
* Mean deviation between consecutive round trips jitter as a stream experiences it.
*
* Not the spread around the average: a link that alternates 20 ms / 200 ms and one that
* drifts slowly from 20 ms to 200 ms have the same standard deviation, and only the first
* one breaks a call.
*/
fun meanDeviation(values: List<Double>): Double {
if (values.size < 2) return 0.0
var sum = 0.0
for (i in 1 until values.size) sum += kotlin.math.abs(values[i] - values[i - 1])
return sum / (values.size - 1)
}
}
}
@@ -55,9 +55,10 @@ class TestBuilder(
evidence: JsonObject? = null, evidence: JsonObject? = null,
metrics: JsonObject? = null, metrics: JsonObject? = null,
error: TestError? = null, error: TestError? = null,
params: JsonObject? = null,
): Test = Test( ): Test = Test(
id = id, type = type, networkRef = networkRef, sessionRef = sessionRef, tier = tier, id = id, type = type, networkRef = networkRef, sessionRef = sessionRef, tier = tier,
startedMonoNs = startedMonoNs, endedMonoNs = ids.monoNs(), startedMonoNs = startedMonoNs, endedMonoNs = ids.monoNs(),
status = status, error = error, evidence = evidence, metrics = metrics, status = status, error = error, params = params, evidence = evidence, metrics = metrics,
) )
} }
@@ -0,0 +1,69 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.probe
import android.content.Context
import app.echo_lot.measurement.Test
import app.echo_lot.measurement.TestError
import app.echo_lot.measurement.TestStatus
import kotlinx.coroutines.CoroutineScope
import kotlinx.coroutines.Deferred
import kotlinx.coroutines.Dispatchers
import kotlinx.coroutines.Job
import kotlinx.coroutines.SupervisorJob
import kotlinx.coroutines.async
import kotlinx.coroutines.withTimeoutOrNull
/**
* Runs an ordinary [Probe] beside the battery instead of inside it.
*
* Some probes are already listeners with a fixed window [MdnsInventoryProbe] does nothing but
* wait for answers and in a long run their window should be the run's window. Running them in
* the sequential battery would then stall every probe behind them for five minutes, which is a
* scheduling problem and not a measurement one, so the fix is to move them rather than to shorten
* them.
*
* Durations are deliberately *not* fed back into the estimate learning: a probe that listens for
* the whole window would teach the short-mode progress bar that mDNS discovery takes five minutes.
*/
class ProbeCollector(private val probe: Probe) : BaseCollector() {
override val type get() = probe.type
override val tier get() = probe.tier
private var scope: CoroutineScope? = null
private var running: Deferred<Test>? = null
override suspend fun start(ctx: Context, ids: ProbeIds) {
begin(ids)
val s = CoroutineScope(SupervisorJob() + Dispatchers.IO).also { scope = it }
running = s.async { probe.run(ctx, ids) }
}
override suspend fun stop(): Test {
val job = running
running = null
val finished = job?.let {
// A short grace, not a long one. A probe timed to the window has already finished by
// the time this is called, so the normal path returns instantly; the grace only covers
// it being slightly late. It is deliberately kept to a second and a half because the
// other caller is the Cancel button, where every millisecond spent waiting for a
// listener that will not finish is a millisecond the user watches nothing happen.
withTimeoutOrNull(GRACE_MS) { runCatching { it.await() }.getOrNull() }
}
scope?.coroutineContext?.get(Job)?.cancel()
scope = null
return finished ?: build(
TestStatus.PARTIAL,
error = TestError(
"window_closed",
"the run's window ended before this listener finished",
),
)
}
private companion object {
const val GRACE_MS = 1_500L
}
}
@@ -0,0 +1,181 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.probe
import android.content.Context
import app.echo_lot.measurement.Test
import app.echo_lot.measurement.TestError
import app.echo_lot.measurement.TestType
import app.echo_lot.measurement.Tier
import kotlinx.coroutines.CoroutineScope
import kotlinx.coroutines.Dispatchers
import kotlinx.coroutines.Job
import kotlinx.coroutines.SupervisorJob
import kotlinx.coroutines.delay
import kotlinx.coroutines.isActive
import kotlinx.coroutines.launch
import kotlinx.serialization.json.JsonObject
import kotlinx.serialization.json.buildJsonObject
import kotlinx.serialization.json.put
/**
* local.ssdp_inventory what UPnP/SSDP devices are on this segment, gathered over the whole window.
*
* **Both halves are needed, and neither is sufficient.** Passive listening catches ssdp:alive and
* ssdp:byebye announcements, which is the only way to see a device that ignores searches (plenty
* do, deliberately) and the only way to see one leave. But announcements are periodic and sparse
* a device re-announces on its own cache-control interval, commonly 30 minutes so a five-minute
* passive window silently misses most of the segment. An M-SEARCH provokes an immediate reply from
* everything that is listening, which is most things, and catches the quiet ones. Running only the
* active half is what [RouterIdentityProbe] already does in the battery, and it is why that probe
* cannot tell you that a device disappeared halfway through the run.
*
* The searches are paced, not flooded: [maxSearches] of them spread [searchIntervalMs] apart. A
* repeat catches devices that joined the network after the run started or were asleep at t=0, while
* staying orders of magnitude below a rate that would itself perturb the network being measured
* the tool must not become the fault it is looking for. `upnp:rootdevice` rather than `ssdp:all`
* for the same reason: one reply per device instead of one per service.
*
* The searches go out from the capture socket bound to 1900, so unicast replies land in the same
* capture as the multicast announcements rather than needing a second socket nobody is reading.
*/
class SsdpCollector(
private val searchIntervalMs: Long = 60_000,
private val maxSearches: Int = 5,
) : BaseCollector() {
override val type = TestType.LOCAL_SSDP_INVENTORY
override val tier = Tier.APP
private val capture = MulticastCapture(
label = "ssdp",
port = PORT,
group4 = GROUP4,
// The IPv6 link-local SSDP group. Free to join where IPv6 exists and skipped where it does
// not, so a v4-only network costs nothing and a v6-only device is not invisible.
group6 = GROUP6,
// SSDP is the chattiest of the four; a device announcing every service it hosts can emit a
// dozen NOTIFYs per cycle, so the per-source cap does most of the work here.
maxPackets = 600,
// The only one of the four with an active half worth keeping when 1900 is taken.
allowEphemeralFallback = true,
)
private var scope: CoroutineScope? = null
private var searchesSent = 0
private var startedAtMonoNs = 0L
override suspend fun start(ctx: Context, ids: ProbeIds) {
begin(ids)
startedAtMonoNs = ids.monoNs()
capture.acquireLock(ctx)
if (!capture.start(ids)) return
val s = CoroutineScope(SupervisorJob() + Dispatchers.IO).also { scope = it }
s.launch {
while (isActive && searchesSent < maxSearches) {
if (capture.send(MSEARCH, GROUP4, PORT)) searchesSent++ else break
delay(searchIntervalMs)
}
}
}
override suspend fun stop(): Test {
scope?.coroutineContext?.get(Job)?.cancel()
scope = null
val packets = capture.stop()
val table = DiscoveryTable()
var alive = 0
var byebye = 0
var responses = 0
var undecodable = 0
for (p in packets) {
val msg = SsdpParser.parse(p.data, p.data.size)
if (msg == null) {
undecodable++
continue
}
when (msg.kind) {
SsdpKind.ALIVE -> alive++
SsdpKind.BYEBYE -> byebye++
SsdpKind.RESPONSE -> responses++
// Our own M-SEARCH comes back to us through the group; counting it as a device
// would inventory the phone doing the measuring.
SsdpKind.SEARCH -> continue
else -> Unit
}
// USN is the device+service identity SSDP itself uses; NT/ST is the fallback for
// devices that omit it, and the source IP already separates two devices offering the
// same service type.
val identity = msg.usn ?: msg.target ?: "(unidentified)"
table.observe(p.sourceIp, identity, p.atMonoNs) {
buildJsonObject {
put("usn", identity)
msg.target?.let { put("target", it) }
msg.serverBanner?.let { put("server_banner", it) }
SsdpParser.productHint(msg.serverBanner)?.let { put("product_hint", it) }
msg.location?.let { put("location", it) }
put("kind", msg.kind.name.lowercase())
}
}
}
val evidence: JsonObject = buildJsonObject {
put("capture", capture.statusJson())
put("ssdp_devices", table.toJson())
// Two truncation sources, reported apart: the capture dropping datagrams and the
// inventory dropping distinct entries mean different things about the network.
put("evidence_truncated", capture.truncated || table.overflowed)
}
val metrics: JsonObject = buildJsonObject {
put("distinct_sources", table.distinctSources)
put("distinct_advertisements", table.size)
put("packets", capture.packetsSeen)
put("alive", alive)
put("byebye", byebye)
put("search_responses", responses)
put("undecodable_packets", undecodable)
put("searches_sent", searchesSent)
put("truncated", capture.truncated || table.overflowed)
}
val status = capture.outcome(sawAnything = !table.isEmpty)
val reason = capture.reason(sawAnything = !table.isEmpty)
return build(
status,
evidence = evidence,
metrics = metrics,
params = params(),
error = reason?.let { TestError("listen_incomplete", it) },
)
}
private fun params(): JsonObject = buildJsonObject {
put("mode", "passive+msearch")
put("group_v4", "$GROUP4:$PORT")
put("group_v6", "[$GROUP6]:$PORT")
put("search_target", SEARCH_TARGET)
put("search_interval_ms", searchIntervalMs)
put("max_searches", maxSearches)
put("started_mono_ns", startedAtMonoNs)
}
private companion object {
const val PORT = 1900
const val GROUP4 = "239.255.255.250"
const val GROUP6 = "ff02::c"
const val SEARCH_TARGET = "upnp:rootdevice"
/** MX is the maximum random delay a responder waits, in seconds. 3 spreads the replies
* enough that a segment full of devices does not answer in one burst we then drop. */
val MSEARCH: ByteArray = (
"M-SEARCH * HTTP/1.1\r\n" +
"HOST: $GROUP4:$PORT\r\n" +
"MAN: \"ssdp:discover\"\r\n" +
"MX: 3\r\n" +
"ST: $SEARCH_TARGET\r\n\r\n"
).toByteArray(Charsets.ISO_8859_1)
}
}
@@ -60,6 +60,15 @@ class StunProbe(
override suspend fun run(ctx: Context, ids: ProbeIds): Test = withContext(Dispatchers.IO) { override suspend fun run(ctx: Context, ids: ProbeIds): Test = withContext(Dispatchers.IO) {
val b = TestBuilder(type, tier, ids) val b = TestBuilder(type, tier, ids)
// Without a server there is nothing to ask. Skipped rather than failed: "the STUN test
// failed" reads as a finding about the network, when the truth is that this device is
// not enrolled anywhere and no packet was ever sent.
if (serverHost.isBlank()) {
return@withContext b.build(
TestStatus.SKIPPED,
evidence = buildJsonObject { put("reason", "no server configured to ask") },
)
}
DatagramSocket().use { sock -> DatagramSocket().use { sock ->
sock.soTimeout = 3000 sock.soTimeout = 3000
val localPort = sock.localPort val localPort = sock.localPort
@@ -0,0 +1,255 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.probe
import android.content.Context
import android.system.Os
import android.system.OsConstants
import app.echo_lot.measurement.Flow
import app.echo_lot.measurement.Hop
import app.echo_lot.measurement.HopProbe
import app.echo_lot.measurement.Test
import app.echo_lot.measurement.TestError
import app.echo_lot.measurement.TestStatus
import app.echo_lot.measurement.TestType
import app.echo_lot.measurement.Tier
import app.echo_lot.measurement.TracerouteEvidence
import app.echo_lot.measurement.toEvidence
import kotlinx.coroutines.Dispatchers
import kotlinx.coroutines.delay
import kotlinx.coroutines.withContext
import kotlinx.serialization.json.buildJsonObject
import kotlinx.serialization.json.put
import java.io.FileDescriptor
import java.net.InetAddress
import java.net.InetSocketAddress
import java.nio.ByteBuffer
import java.nio.ByteOrder
/**
* traceroute.udp4 UDP traceroute reading ICMP time-exceeded off the socket error queue via
* Os.recvmsg(MSG_ERRQUEUE): no root, no raw socket, no native code. Folded from the prober,
* which validated real hop addresses on both known devices (6 hops on the OnePlus 15, 5 on the
* Lenovo) and thereby retired the planned C-over-JNI errqueue shim.
*
* StructMsghdr/StructCmsghdr/recvmsg are reached via reflection (repo convention for uncertain
* OS paths): present since roughly API 34, absent before, and the probe must run and report
* on both. An absent API is UNSUPPORTED with the reason, never a crash.
*/
class TracerouteProbe(
private val targetHost: String = "1.1.1.1",
private val maxHops: Int = 6,
) : Probe {
override val type = TestType.TRACEROUTE_UDP4
override val tier = Tier.APP
// Validated wall clock is ~250 ms on a healthy path; the ceiling is maxHops silent hops at
// 900 ms each, which only a blackholing path produces.
override val estimatedMs = 1_500L
private companion object {
const val BASE_PORT = 33434
/** ICMP errors take one RTT to surface on the errqueue; poll briefly, never block. */
const val HOP_DEADLINE_NS = 900_000_000L
const val POLL_INTERVAL_MS = 40L
}
override suspend fun run(ctx: Context, ids: ProbeIds): Test = withContext(Dispatchers.IO) {
val b = TestBuilder(type, tier, ids)
val params = buildJsonObject {
put("target", targetHost); put("max_hops", maxHops); put("base_port", BASE_PORT)
}
val api = ErrqueueApi.resolve()
?: return@withContext b.build(
TestStatus.UNSUPPORTED,
params = params,
error = TestError(
"no_recvmsg",
"StructMsghdr/Os.recvmsg not on this API level — errqueue unreadable",
),
)
var fd: FileDescriptor? = null
try {
fd = Os.socket(OsConstants.AF_INET, OsConstants.SOCK_DGRAM, OsConstants.IPPROTO_UDP)
OsAbi.trySetIntOpt(fd, OsConstants.IPPROTO_IP, OsAbi.IP_RECVERR, 1)?.let {
return@withContext b.build(
TestStatus.UNSUPPORTED,
params = params,
error = TestError("ip_recverr_rejected", it),
)
}
val target = InetAddress.getByName(targetHost)
val hops = ArrayList<Hop>(maxHops)
var hopsSeen = 0
var reachedTarget = false
var srcPort = 0
for (ttl in 1..maxHops) {
OsAbi.trySetIntOpt(fd, OsConstants.IPPROTO_IP, OsAbi.IP_TTL, ttl)
val t0 = System.nanoTime()
val sent = runCatching {
Os.sendto(fd, ByteArray(32), 0, 32, 0, target, BASE_PORT + ttl)
}
if (sent.isFailure) {
hops.add(Hop(ttl, listOf(HopProbe(icmp = "sendto failed: " +
(sent.exceptionOrNull()?.message ?: "?")))))
continue
}
if (srcPort == 0) {
// Only readable after the implicit bind the first send performs.
srcPort = runCatching {
(Os.getsockname(fd) as? InetSocketAddress)?.port ?: 0
}.getOrDefault(0)
}
var hop: ErrqueueApi.ErrEvent? = null
val deadline = System.nanoTime() + HOP_DEADLINE_NS
while (hop == null && System.nanoTime() < deadline) {
hop = api.pollErrqueue(fd)
if (hop == null) delay(POLL_INTERVAL_MS)
}
val rttNs = System.nanoTime() - t0
when {
hop == null -> hops.add(Hop(ttl, listOf(HopProbe()))) // silent hop: all null
hop.parseError != null ->
hops.add(Hop(ttl, listOf(HopProbe(icmp = "unparsed: ${hop.parseError}"))))
else -> {
hopsSeen++
hops.add(Hop(ttl, listOf(HopProbe(
replyFrom = hop.offender,
rttNs = rttNs,
icmp = when (hop.icmpType) {
OsAbi.ICMP_TIME_EXCEEDED -> "time_exceeded"
OsAbi.ICMP_DEST_UNREACH -> "dest_unreachable"
else -> "type_${hop.icmpType}"
},
))))
if (hop.icmpType == OsAbi.ICMP_DEST_UNREACH) reachedTarget = true
}
}
if (reachedTarget) break
}
// dst_port varies per TTL (classic traceroute, and what was validated on hardware),
// so this flow is explicitly NOT fixed-tuple; base_port is in params.
val evidence = TracerouteEvidence(
flow = Flow(srcPort = srcPort, dstPort = BASE_PORT, fixedTuple = false),
hops = hops,
).toEvidence()
val metrics = buildJsonObject {
put("hops_seen", hopsSeen)
put("reached_target", reachedTarget)
}
val status = when {
hopsSeen > 0 -> TestStatus.OK
// API present, sends succeeded, nothing surfaced: a fact about this path or
// kernel, not proof the mechanism is missing.
else -> TestStatus.PARTIAL
}
b.build(status, params = params, evidence = evidence, metrics = metrics)
} catch (e: Throwable) {
b.build(
TestStatus.FAILED,
params = params,
error = TestError("uncaught", e.message ?: e.javaClass.simpleName),
)
} finally {
fd?.let { runCatching { Os.close(it) } }
}
}
}
/**
* Reflection facade over android.system.{StructMsghdr, StructCmsghdr, Os.recvmsg}.
* Resolved once; null if any piece is missing on this API level.
*/
internal class ErrqueueApi private constructor(
private val msghdrCtor: java.lang.reflect.Constructor<*>,
private val recvmsg: java.lang.reflect.Method,
private val cmsgLevel: java.lang.reflect.Field,
private val cmsgType: java.lang.reflect.Field,
private val cmsgData: java.lang.reflect.Field,
private val msgControl: java.lang.reflect.Field,
) {
class ErrEvent(
val offender: String?,
val icmpType: Int,
val origin: Int,
val parseError: String? = null,
)
/** One non-blocking MSG_ERRQUEUE read; null when the queue is empty. */
fun pollErrqueue(fd: FileDescriptor): ErrEvent? {
return try {
val iov = arrayOf(ByteBuffer.allocate(512))
// (SocketAddress msg_name, ByteBuffer[] msg_iov, StructCmsghdr[] msg_control, flags)
val msghdr = msghdrCtor.newInstance(
InetSocketAddress(0), iov, null, 0,
)
recvmsg.invoke(null, fd, msghdr, OsAbi.MSG_ERRQUEUE or OsAbi.MSG_DONTWAIT)
val control = msgControl.get(msghdr) as? Array<*>
?: return ErrEvent(null, -1, -1, "msg_control empty after recvmsg")
for (cmsg in control.filterNotNull()) {
val level = cmsgLevel.getInt(cmsg)
val type = cmsgType.getInt(cmsg)
if (level == OsConstants.IPPROTO_IP && type == OsAbi.IP_RECVERR) {
return parseSockExtendedErr(cmsgData.get(cmsg))
}
}
ErrEvent(null, -1, -1, "no IP_RECVERR cmsg among ${control.size}")
} catch (e: Throwable) {
// The single most load-bearing line: reflection wraps errno in
// InvocationTargetException, and EAGAIN there means "queue empty", not failure.
val cause = (e as? java.lang.reflect.InvocationTargetException)?.cause ?: e
val msg = cause.message ?: cause.javaClass.simpleName
if ("EAGAIN" in msg || "EWOULDBLOCK" in msg) null
else ErrEvent(null, -1, -1, msg)
}
}
/** cmsg_data = struct sock_extended_err + offender sockaddr_in (see OsAbi). */
private fun parseSockExtendedErr(data: Any?): ErrEvent {
val bytes: ByteArray = when (data) {
is ByteArray -> data
is ByteBuffer -> ByteArray(data.remaining()).also { data.duplicate().get(it) }
else -> return ErrEvent(null, -1, -1, "cmsg_data is ${data?.javaClass?.name}")
}
if (bytes.size < OsAbi.SOCK_EE_SIZE) {
return ErrEvent(null, -1, -1, "cmsg_data too short: ${bytes.size}")
}
val origin = bytes[4].toInt() and 0xFF
val icmpType = bytes[5].toInt() and 0xFF
// SO_EE_OFFENDER: sockaddr_in directly after the fixed struct; family is in native
// byte order, sin_addr at offset +4 within the sockaddr.
val offender = if (bytes.size >= OsAbi.SOCK_EE_SIZE + 8) {
val family = ByteBuffer.wrap(bytes, OsAbi.SOCK_EE_SIZE, 2)
.order(ByteOrder.nativeOrder()).short.toInt()
if (family == OsConstants.AF_INET) {
val a = bytes.copyOfRange(OsAbi.SOCK_EE_SIZE + 4, OsAbi.SOCK_EE_SIZE + 8)
InetAddress.getByAddress(a).hostAddress
} else null
} else null
return ErrEvent(offender, icmpType, origin)
}
companion object {
fun resolve(): ErrqueueApi? = runCatching {
val msghdrCls = Class.forName("android.system.StructMsghdr")
val cmsghdrCls = Class.forName("android.system.StructCmsghdr")
ErrqueueApi(
// Picked by shape, not by position: the 4-arg form is
// (SocketAddress, ByteBuffer[], StructCmsghdr[], int) on every level that has it.
msghdrCtor = msghdrCls.constructors.first { it.parameterCount == 4 },
recvmsg = Os::class.java.getMethod(
"recvmsg", FileDescriptor::class.java, msghdrCls, Int::class.javaPrimitiveType,
),
cmsgLevel = cmsghdrCls.getField("cmsg_level"),
cmsgType = cmsghdrCls.getField("cmsg_type"),
cmsgData = cmsghdrCls.getField("cmsg_data"),
msgControl = msghdrCls.getField("msg_control"),
)
}.getOrNull()
}
}
@@ -0,0 +1,131 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.probe
import android.content.Context
import app.echo_lot.measurement.Test
import app.echo_lot.measurement.TestStatus
import app.echo_lot.measurement.TestType
import app.echo_lot.measurement.Tier
import app.echo_lot.measurement.Network as MNetwork
import kotlinx.coroutines.Dispatchers
import kotlinx.coroutines.withContext
import kotlinx.serialization.json.JsonObject
import kotlinx.serialization.json.buildJsonObject
import kotlinx.serialization.json.put
import java.net.Inet6Address
import java.net.InetSocketAddress
import java.util.Locale
/**
* v6.brokenness does IPv6 actually carry traffic, asked with a real TCP connection.
*
* This exists to corroborate (or refute) the ICMPv6 silence that icmp.ping6 observes. ICMPv6 echo
* is widely filtered on networks where IPv6 works fine, so silence alone cannot distinguish
* "IPv6 is broken" from "ping is filtered" a phone that reported v6.broken while happily
* loading IPv6-only sites is what proved the point. A TCP connect over IPv6 to the configured
* server settles it: if it succeeds, IPv6 works and the ICMP silence is filtering; if it fails
* too, on a network that advertises IPv6, the brokenness claim finally has evidence behind it.
*
* Only networks that claim to offer IPv6 (a global address or a v6 default route) are attempted:
* connecting over v6 on an IPv4-only network fails by design, and recording that as evidence
* would manufacture the exact false positive this probe exists to kill.
*/
class V6ConnectProbe(
private val entries: List<NetworkInventory.Entry>,
private val serverHost: String,
private val port: Int = 443,
) : Probe {
override val type = TestType.V6_BROKENNESS
override val tier = Tier.APP
// One 3s connect timeout per v6-provisioned network, at most.
override val estimatedMs = 4_000L
override suspend fun run(ctx: Context, ids: ProbeIds): Test = withContext(Dispatchers.IO) {
val b = TestBuilder(type, tier, ids)
// Same rule as the STUN and canary probes: with no server there is no target, and
// borrowing someone else's infrastructure to get one is not this app's call to make.
if (serverHost.isBlank()) {
return@withContext b.build(
TestStatus.SKIPPED,
evidence = buildJsonObject { put("reason", "no server configured to connect to") },
)
}
val candidates = entries.filter { ipv6Provisioned(it.model) }
if (candidates.isEmpty()) {
return@withContext b.build(
TestStatus.SKIPPED,
evidence = buildJsonObject {
put("reason", "no active network claims to offer IPv6")
},
)
}
var okCount = 0
var attemptedCount = 0
val evidence: JsonObject = buildJsonObject {
put("target", "$serverHost:$port")
for (e in candidates) {
val label = "${e.model.transport.name.lowercase()}:${e.model.id}"
val a = attempt(e)
if (a.attempted) attemptedCount++
if (a.ok) okCount++
put(label, buildJsonObject {
put("network_ref", e.model.id)
put("ok", a.ok)
put("attempted", a.attempted)
put("detail", a.detail)
})
}
}
val status = when {
attemptedCount == 0 -> TestStatus.SKIPPED // resolution/binding never got that far
okCount == attemptedCount -> TestStatus.OK
okCount > 0 -> TestStatus.PARTIAL
else -> TestStatus.FAILED
}
b.build(status, evidence = evidence)
}
/** Same attempted/ok separation as IcmpProbe: a connect we never sent proves nothing. */
private data class Attempt(val ok: Boolean, val attempted: Boolean, val detail: String)
private fun attempt(e: NetworkInventory.Entry): Attempt {
// Resolved through this network's own resolver; a v6 address obtained over another
// network would still be connected to over this one, which is what matters.
val addr = runCatching {
e.handle.getAllByName(serverHost).filterIsInstance<Inet6Address>().firstOrNull()
}.getOrNull()
?: return Attempt(false, false, "no AAAA answer for $serverHost via this network")
// createSocket() binds to the network at creation; failing here means the app could not
// use the interface at all (e.g. EPERM under a VPN) — nothing was sent, nothing is known.
val socket = try {
e.handle.socketFactory.createSocket()
} catch (t: Throwable) {
return Attempt(false, false, "socket unavailable: ${t.message ?: t.javaClass.simpleName}")
}
return try {
val t0 = System.nanoTime()
socket.connect(InetSocketAddress(addr, port), 3000)
val rttMs = (System.nanoTime() - t0) / 1_000_000.0
Attempt(true, true, "connected to [${addr.hostAddress}]:$port " +
"rtt_ms=${"%.1f".format(Locale.ROOT, rttMs)}")
} catch (t: Throwable) {
// A refused connection would still prove the path forwards IPv6, but against our own
// server's 443 the realistic failures are timeout and unreachable — both silence.
Attempt(false, true, "error: ${t.message ?: t.javaClass.simpleName}")
} finally {
runCatching { socket.close() }
}
}
/** The network claims IPv6: a global (non-link-local) address or a v6 default route. */
private fun ipv6Provisioned(n: MNetwork): Boolean =
n.link.addresses.any { a ->
a.addr.contains(':') &&
!a.addr.startsWith("fe80", ignoreCase = true) &&
!a.addr.startsWith("::1")
} || n.link.routes.any { it.dst == "::/0" }
}
@@ -0,0 +1,165 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.probe
import android.content.Context
import android.net.wifi.WifiManager
import app.echo_lot.measurement.Test
import app.echo_lot.measurement.TestError
import app.echo_lot.measurement.TestStatus
import app.echo_lot.measurement.TestType
import app.echo_lot.measurement.Tier
import app.echo_lot.measurement.Transport
import kotlinx.coroutines.CoroutineScope
import kotlinx.coroutines.Dispatchers
import kotlinx.coroutines.Job
import kotlinx.coroutines.SupervisorJob
import kotlinx.coroutines.delay
import kotlinx.coroutines.isActive
import kotlinx.coroutines.launch
import kotlinx.serialization.json.JsonObject
import kotlinx.serialization.json.buildJsonObject
import kotlinx.serialization.json.put
import kotlinx.serialization.json.putJsonArray
import kotlinx.serialization.json.add
import java.util.Collections
/**
* wifi.signal_log RSSI, link speed and frequency sampled across the whole window.
*
* One reading of the signal strength says almost nothing: -67 dBm is fine, and -67 dBm that was
* -45 dBm ninety seconds ago is somebody walking away from the AP, or an AP whose power is being
* managed, or a band steer about to happen. The series is the measurement; the snapshot in
* `networks[].wifi` is only its first sample.
*
* Evidence is columnar (measurement-schema §6.2 conventions): parallel arrays keep a five-minute
* log at 2 s intervals in a few kB.
*/
class WifiSignalCollector(
private val entries: List<NetworkInventory.Entry>,
private val intervalMs: Long = 2_000,
) : BaseCollector() {
override val type = TestType.WIFI_SIGNAL_LOG
override val tier = Tier.APP
private val atMonoNs: MutableList<Long> = Collections.synchronizedList(mutableListOf())
private val rssi: MutableList<Int> = Collections.synchronizedList(mutableListOf())
private val speed: MutableList<Int> = Collections.synchronizedList(mutableListOf())
private val freq: MutableList<Int> = Collections.synchronizedList(mutableListOf())
/** Only the count of distinct BSSIDs leaves this class a roam is the fact worth reporting,
* and the addresses themselves are neighbours' hardware identifiers. */
private val bssids = Collections.synchronizedSet(HashSet<String>())
private var scope: CoroutineScope? = null
private var unsupported: String? = null
private var startedAtMonoNs = 0L
override suspend fun start(ctx: Context, ids: ProbeIds) {
// network_ref up front: this samples the wifi link, and a signal log with nothing to
// attach it to is a series of numbers about an unnamed thing.
val wifiNet = entries.firstOrNull { it.model.transport == Transport.WIFI }
begin(ids, networkRef = wifiNet?.model?.id)
startedAtMonoNs = ids.monoNs()
val wifi = ctx.applicationContext.getSystemService(WifiManager::class.java)
if (wifi == null) {
unsupported = "WifiManager unavailable"
return
}
if (wifiNet == null) {
unsupported = "no wifi network is connected"
return
}
// Own scope, not the caller's: the run job is cancelled the instant the user taps Cancel,
// and the samples taken up to that point are exactly what a cancelled long run still owes
// them. stop() ends this scope.
val s = CoroutineScope(SupervisorJob() + Dispatchers.IO).also { scope = it }
s.launch {
while (isActive) {
sample(ids, wifi)
delay(intervalMs)
}
}
}
@Suppress("DEPRECATION")
private fun sample(ids: ProbeIds, wifi: WifiManager) {
// WifiManager.getConnectionInfo is deprecated in favour of the NetworkCallback's
// TransportInfo, which delivers a WifiInfo only when the capabilities change — i.e. at the
// platform's cadence, not ours, and with no way to ask for a sample. For a fixed-interval
// log the deprecated call is the one that answers the question, and it still works.
val info = runCatching { wifi.connectionInfo }.getOrNull() ?: return
val r = info.rssi
// -127 and 0 are the "no reading" sentinels; recording them would drag every average down
// and invent a signal cliff that never happened.
if (r == 0 || r <= -127) return
atMonoNs.add(ids.monoNs())
rssi.add(r)
speed.add(info.linkSpeed)
freq.add(runCatching { info.frequency }.getOrDefault(0))
runCatching { info.bssid }.getOrNull()
?.takeIf { it.isNotBlank() && it != "02:00:00:00:00:00" }
?.let { bssids.add(it) }
}
override suspend fun stop(): Test {
scope?.coroutineContext?.get(Job)?.cancel()
scope = null
unsupported?.let {
return build(
TestStatus.UNSUPPORTED,
params = params(),
error = TestError("no_wifi", it),
)
}
val t = synchronized(atMonoNs) { atMonoNs.toList() }
val r = synchronized(rssi) { rssi.toList() }
val sp = synchronized(speed) { speed.toList() }
val f = synchronized(freq) { freq.toList() }
val evidence: JsonObject = buildJsonObject {
putJsonArray("at_mono_ns") { for (v in t) add(v) }
putJsonArray("rssi_dbm") { for (v in r) add(v) }
putJsonArray("link_speed_mbps") { for (v in sp) add(v) }
putJsonArray("frequency_mhz") { for (v in f) add(v) }
}
val metrics: JsonObject = buildJsonObject {
put("samples", r.size)
if (r.isNotEmpty()) {
put("rssi_dbm_min", r.min())
put("rssi_dbm_avg", round1(r.average()))
put("rssi_dbm_max", r.max())
put("rssi_dbm_range", r.max() - r.min())
}
sp.filter { it > 0 }.let { valid ->
if (valid.isNotEmpty()) {
put("link_speed_mbps_min", valid.min())
put("link_speed_mbps_avg", round1(valid.average()))
put("link_speed_mbps_max", valid.max())
}
}
// Distinct BSSIDs minus the one we started on: how often the phone changed AP without
// the network ever going down — invisible to any one-shot probe, and a common cause of
// "the call drops when I walk into the kitchen".
put("roams", (bssids.size - 1).coerceAtLeast(0))
}
return build(
if (r.isEmpty()) TestStatus.PARTIAL else TestStatus.OK,
evidence = evidence, metrics = metrics, params = params(),
)
}
private fun params(): JsonObject = buildJsonObject {
put("interval_ms", intervalMs)
put("started_mono_ns", startedAtMonoNs)
put("source", "WifiManager.connectionInfo")
}
private companion object {
fun round1(v: Double) = Math.round(v * 10.0) / 10.0
}
}
@@ -0,0 +1,133 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.probe
import android.content.Context
import app.echo_lot.measurement.Test
import app.echo_lot.measurement.TestError
import app.echo_lot.measurement.TestType
import app.echo_lot.measurement.Tier
import kotlinx.serialization.json.JsonObject
import kotlinx.serialization.json.buildJsonObject
import kotlinx.serialization.json.put
/**
* local.wsd_inventory WS-Discovery (SOAP-over-UDP, 3702) Hello / Bye / Probe / ProbeMatches.
*
* This is the protocol Windows and modern printers use to find each other, and it inventories a
* class of device the other three miss: network printers, scanners and IP cameras announce here and
* frequently nowhere else. A `Hello` is a device arriving, a `Bye` is one leaving, and a `Probe`
* from a workstation names what it is hunting for so a window over 3702 shows both the equipment
* on the segment and which machines are looking for it.
*
* Passive only. WS-Discovery's active half is a Probe multicast, which would make this app a
* participant announcing itself to every device on the segment; SSDP's M-SEARCH is a single small
* request that devices expect constantly, whereas a WSD Probe from an unknown host is the kind of
* thing that shows up in someone's security log. Listening costs the network nothing.
*
* Payloads are decoded by [WsdParser], which does targeted extraction rather than XML parsing
* see its documentation for why that is the right call for unauthenticated broadcast input.
*/
class WsdCollector : BaseCollector() {
override val type = TestType.LOCAL_WSD_INVENTORY
override val tier = Tier.APP
private val capture = MulticastCapture(
label = "wsd",
port = PORT,
group4 = GROUP4,
group6 = GROUP6,
// SOAP envelopes are an order of magnitude larger than the other three protocols'
// datagrams, so the byte ceiling binds before the packet count does. Both are set
// explicitly rather than left to the default, which was chosen for 200-byte packets.
maxPackets = 250,
maxBytesRetained = 192 * 1024,
)
private var startedAtMonoNs = 0L
override suspend fun start(ctx: Context, ids: ProbeIds) {
begin(ids)
startedAtMonoNs = ids.monoNs()
capture.acquireLock(ctx)
capture.start(ids)
}
override suspend fun stop(): Test {
val packets = capture.stop()
val table = DiscoveryTable()
var hello = 0
var bye = 0
var probes = 0
var matches = 0
var undecodable = 0
for (p in packets) {
val m = WsdParser.parse(p.data, p.data.size)
if (m == null) {
undecodable++
continue
}
when (m.action?.lowercase()) {
"hello" -> hello++
"bye" -> bye++
"probe" -> probes++
"probematches", "resolvematches" -> matches++
}
// The device UUID is WS-Discovery's own stable identity and survives address changes,
// so it is the identity where present; a Probe carries none (it names what it wants,
// not who it is), and there the action plus the types is what distinguishes one
// observation from a repeat of it.
val identity = m.deviceUuid ?: (m.action.orEmpty() + "/" + m.types.orEmpty())
table.observe(p.sourceIp, identity.ifEmpty { "(unidentified)" }, p.atMonoNs) {
buildJsonObject {
m.deviceUuid?.let { put("device_uuid", it) }
m.action?.let { put("action", it) }
m.types?.let { put("wsd_types", it) }
m.xaddrs?.let { put("wsd_xaddrs", it) }
}
}
}
val evidence: JsonObject = buildJsonObject {
put("capture", capture.statusJson())
put("wsd_devices", table.toJson())
put("evidence_truncated", capture.truncated || table.overflowed)
}
val metrics: JsonObject = buildJsonObject {
put("distinct_sources", table.distinctSources)
put("distinct_devices", table.size)
put("packets", capture.packetsSeen)
put("hello", hello)
put("bye", bye)
put("probes", probes)
put("probe_matches", matches)
put("undecodable_packets", undecodable)
put("truncated", capture.truncated || table.overflowed)
}
val saw = !table.isEmpty
return build(
capture.outcome(saw),
evidence = evidence,
metrics = metrics,
params = params(),
error = capture.reason(saw)?.let { TestError("listen_incomplete", it) },
)
}
private fun params(): JsonObject = buildJsonObject {
put("mode", "passive")
put("group_v4", "$GROUP4:$PORT")
put("group_v6", "[$GROUP6]:$PORT")
put("started_mono_ns", startedAtMonoNs)
}
private companion object {
const val PORT = 3702
const val GROUP4 = "239.255.255.250"
const val GROUP6 = "ff02::c"
}
}
@@ -0,0 +1,386 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.probe
import kotlin.test.Test
import kotlin.test.assertEquals
import kotlin.test.assertFalse
import kotlin.test.assertNotNull
import kotlin.test.assertNull
import kotlin.test.assertTrue
/**
* The decoders in DiscoveryParsers.kt, against payloads shaped like the ones real devices emit and
* against the ones a broken or hostile device emits.
*
* The malformed cases are the point. These four decoders are the only place in the app where bytes
* from an unidentified third party on the local segment are interpreted; they run inside a
* collector whose contract is that it never throws, and every one of them is reachable by anyone
* who can put a frame on the wire. So each protocol is fed truncation, junk, and the specific abuse
* its format invites a DNS compression pointer, a NetBIOS name outside the A-P alphabet, an XML
* entity bomb and the assertion is always the same pair: no exception, and no invented data.
*/
class DiscoveryParsersTest {
private fun bytes(s: String) = s.toByteArray(Charsets.ISO_8859_1)
// ---- SSDP --------------------------------------------------------------------------------
private val notifyAlive = bytes(
"NOTIFY * HTTP/1.1\r\n" +
"HOST: 239.255.255.250:1900\r\n" +
"CACHE-CONTROL: max-age=1800\r\n" +
"LOCATION: http://192.168.1.44:8060/\r\n" +
"NT: upnp:rootdevice\r\n" +
"NTS: ssdp:alive\r\n" +
"SERVER: Roku/12.5.5 UPnP/1.0 Roku/12.5.5\r\n" +
"USN: uuid:roku:ecp:YH00E1234567::upnp:rootdevice\r\n\r\n"
)
private val searchResponse = bytes(
"HTTP/1.1 200 OK\r\n" +
"CACHE-CONTROL: max-age=1800\r\n" +
"EXT:\r\n" +
"LOCATION: http://192.168.1.1:49000/rootDesc.xml\r\n" +
"SERVER: FRITZ!Box 7590 UPnP/1.0 AVM FRITZ!Box 7590 154.07.57\r\n" +
"ST: upnp:rootdevice\r\n" +
"USN: uuid:75802409-bccb-40e7-8e6c-c0ff33445566::upnp:rootdevice\r\n\r\n"
)
@Test
fun ssdpAliveAnnouncementYieldsIdentityAndLocation() {
val m = assertNotNull(SsdpParser.parse(notifyAlive, notifyAlive.size))
assertEquals(SsdpKind.ALIVE, m.kind)
assertEquals("upnp:rootdevice", m.target)
assertEquals("uuid:roku:ecp:YH00E1234567::upnp:rootdevice", m.usn)
assertEquals("http://192.168.1.44:8060/", m.location)
assertEquals("Roku/12.5.5 UPnP/1.0 Roku/12.5.5", m.serverBanner)
}
@Test
fun ssdpByebyeIsDistinguishedFromAlive() {
val byebye = bytes(
"NOTIFY * HTTP/1.1\r\nHOST: 239.255.255.250:1900\r\n" +
"NT: urn:schemas-upnp-org:device:MediaRenderer:1\r\nNTS: ssdp:byebye\r\n" +
"USN: uuid:aabbccdd::urn:schemas-upnp-org:device:MediaRenderer:1\r\n\r\n"
)
val m = assertNotNull(SsdpParser.parse(byebye, byebye.size))
assertEquals(SsdpKind.BYEBYE, m.kind)
assertEquals("urn:schemas-upnp-org:device:MediaRenderer:1", m.target)
assertNull(m.location)
}
@Test
fun ssdpSearchResponseReadsStAsTheTarget() {
val m = assertNotNull(SsdpParser.parse(searchResponse, searchResponse.size))
assertEquals(SsdpKind.RESPONSE, m.kind)
assertEquals("upnp:rootdevice", m.target)
assertEquals("http://192.168.1.1:49000/rootDesc.xml", m.location)
}
/** Our own M-SEARCH comes back through the group; it must be recognisable so the collector
* does not inventory the phone doing the measuring. */
@Test
fun ssdpOwnSearchIsRecognisable() {
val search = bytes(
"M-SEARCH * HTTP/1.1\r\nHOST: 239.255.255.250:1900\r\n" +
"MAN: \"ssdp:discover\"\r\nMX: 3\r\nST: upnp:rootdevice\r\n\r\n"
)
assertEquals(SsdpKind.SEARCH, assertNotNull(SsdpParser.parse(search, search.size)).kind)
}
@Test
fun ssdpProductHintDropsBoilerplateAndKeepsTheModel() {
assertEquals("Roku/12.5.5", SsdpParser.productHint("Roku/12.5.5 UPnP/1.0 Roku/12.5.5"))
assertEquals("Synology/DSM-7.3", SsdpParser.productHint("Linux/4.4 UPnP/1.0 Synology/DSM-7.3"))
val fritz = assertNotNull(SsdpParser.productHint("FRITZ!Box 7590 UPnP/1.0 AVM FRITZ!Box 7590"))
assertTrue(fritz.contains("FRITZ!Box"))
assertFalse(fritz.contains("UPnP"), "protocol boilerplate leaked into the model hint")
assertNull(SsdpParser.productHint(null))
assertNull(SsdpParser.productHint(" "))
// A banner that is nothing but boilerplate reveals no model and must say so, rather than
// returning an empty string that reads as a name nobody could see.
assertNull(SsdpParser.productHint("Linux/4.4 UPnP/1.0"))
}
@Test
fun ssdpMalformedIsRejectedWithoutThrowing() {
// Truncated mid-header: the headers that did arrive are still usable, and the missing NTS
// makes the kind unknown rather than making the packet a lie.
val cut = notifyAlive.copyOf(70)
assertEquals(SsdpKind.OTHER, assertNotNull(SsdpParser.parse(cut, cut.size)).kind)
assertNull(SsdpParser.parse(ByteArray(0), 0))
val blank = bytes("\r\n\r\n")
assertNull(SsdpParser.parse(blank, blank.size))
val http = bytes("GET / HTTP/1.0\r\n\r\n")
assertNull(SsdpParser.parse(http, http.size), "not an SSDP verb")
// Arbitrary binary, including the high bytes ISO-8859-1 must not choke on.
val junk = ByteArray(256) { it.toByte() }
assertNull(SsdpParser.parse(junk, junk.size))
// A declared length longer than the buffer must be refused, not read past.
assertNull(SsdpParser.parse(notifyAlive, notifyAlive.size + 100))
// Header lines with no colon are skipped rather than fatal.
val noColon = bytes("NOTIFY * HTTP/1.1\r\ngarbage line\r\nNTS: ssdp:alive\r\n\r\n")
assertEquals(SsdpKind.ALIVE, assertNotNull(SsdpParser.parse(noColon, noColon.size)).kind)
}
// ---- LLMNR -------------------------------------------------------------------------------
/** Builds a DNS-format packet: header, one question, nothing else. */
private fun dnsQuery(
name: String,
qtype: Int = 1,
id: Int = 0x1234,
flags: Int = 0x0000,
qdcount: Int = 1,
): ByteArray {
val out = ArrayList<Byte>()
fun u16(v: Int) { out.add((v shr 8).toByte()); out.add((v and 0xFF).toByte()) }
u16(id); u16(flags); u16(qdcount); u16(0); u16(0); u16(0)
for (label in name.split('.')) {
val b = label.toByteArray(Charsets.UTF_8)
out.add(b.size.toByte())
b.forEach { out.add(it) }
}
out.add(0)
u16(qtype); u16(1)
return out.toByteArray()
}
@Test
fun llmnrQueryYieldsTheNameAWorkstationIsHuntingFor() {
val p = dnsQuery("wpad")
val q = assertNotNull(LlmnrParser.parse(p, p.size))
assertEquals("wpad", q.name)
assertEquals(1, q.qtype)
assertEquals("A", LlmnrParser.qtypeName(q.qtype))
assertTrue(q.isQuery)
assertEquals(0, q.opcode)
}
@Test
fun llmnrMultiLabelNamesAndAaaaSurviveIntact() {
val p = dnsQuery("nas-backup.local", qtype = 28)
val q = assertNotNull(LlmnrParser.parse(p, p.size))
assertEquals("nas-backup.local", q.name)
assertEquals("AAAA", LlmnrParser.qtypeName(q.qtype))
}
@Test
fun llmnrResponsesAreSeparatedFromQueries() {
val p = dnsQuery("DESKTOP-A1B2C3", flags = 0x8000)
assertFalse(assertNotNull(LlmnrParser.parse(p, p.size)).isQuery)
}
@Test
fun llmnrMalformedIsRejectedWithoutThrowing() {
val good = dnsQuery("printer")
assertNull(LlmnrParser.parse(ByteArray(0), 0))
assertNull(LlmnrParser.parse(good, 8), "a header-length prefix is not a question")
assertNull(LlmnrParser.parse(good, good.size + 50), "declared length past the buffer")
val noQuestion = dnsQuery("x", qdcount = 0)
assertNull(LlmnrParser.parse(noQuestion, noQuestion.size))
// A label length that runs off the end of the datagram — the classic truncation.
val overrun = good.copyOf(good.size - 6)
assertNull(LlmnrParser.parse(overrun, overrun.size))
// A compression pointer: legal DNS, forbidden in LLMNR, and the shape that makes a naive
// decoder loop forever. It must be refused rather than followed.
val pointer = good.copyOf(20)
pointer[12] = 0xC0.toByte()
pointer[13] = 0x0C
assertNull(LlmnrParser.parse(pointer, pointer.size))
// Random bytes behind a plausible header: whatever comes back, it is not an exception.
val junk = ByteArray(64) { (it * 37).toByte() }
junk[4] = 0; junk[5] = 1
LlmnrParser.parse(junk, junk.size)
}
// ---- NetBIOS -----------------------------------------------------------------------------
/** First-level encoding, the same transform the decoder has to undo. */
private fun nbnsPacket(name: String, suffix: Int, flags: Int = 0x2810): ByteArray {
val raw = ByteArray(16) { ' '.code.toByte() }
name.forEachIndexed { i, c -> if (i < 15) raw[i] = c.code.toByte() }
raw[15] = suffix.toByte()
val out = ArrayList<Byte>()
fun u16(v: Int) { out.add((v shr 8).toByte()); out.add((v and 0xFF).toByte()) }
u16(0x8001); u16(flags); u16(1); u16(0); u16(0); u16(1)
out.add(32)
for (b in raw) {
val v = b.toInt() and 0xFF
out.add(('A'.code + (v shr 4)).toByte())
out.add(('A'.code + (v and 0x0F)).toByte())
}
out.add(0)
u16(0x0020); u16(0x0001)
return out.toByteArray()
}
@Test
fun netbiosNameDecodesWithItsSuffixAndRole() {
val p = nbnsPacket("DESKTOP-A1B2C3", 0x20)
val n = assertNotNull(NetbiosParser.parse(p, p.size))
assertEquals("DESKTOP-A1B2C3", n.name)
assertEquals(0x20, n.suffix)
assertEquals("file_server", n.role)
assertFalse(n.isResponse)
assertEquals("registration", NetbiosParser.opcodeName(n.opcode))
}
@Test
fun netbiosPaddingIsStrippedAndSuffixesAreNamed() {
val p = nbnsPacket("WORKGROUP", 0x1E)
val n = assertNotNull(NetbiosParser.parse(p, p.size))
assertEquals("WORKGROUP", n.name, "the 15-byte space padding leaked into the name")
assertEquals("browser_elections", n.role)
assertEquals("workstation", NetbiosParser.roleOf(0x00))
assertEquals("suffix_0xAB", NetbiosParser.roleOf(0xAB))
}
@Test
fun netbiosResponsesAreSeparatedFromRequests() {
val p = nbnsPacket("FILESRV", 0x20, flags = 0x8500)
assertTrue(assertNotNull(NetbiosParser.parse(p, p.size)).isResponse)
}
@Test
fun netbiosMalformedIsRejectedWithoutThrowing() {
val good = nbnsPacket("PRINTER", 0x00)
assertNull(NetbiosParser.parse(ByteArray(0), 0))
assertNull(NetbiosParser.parse(good, 20), "truncated before the encoded name ends")
assertNull(NetbiosParser.parse(good, good.size + 40), "declared length past the buffer")
// A character outside A-P cannot be half of an encoded byte. Guessing at the rest would
// fabricate a hostname, so the whole packet is refused.
val badAlphabet = good.copyOf()
badAlphabet[15] = 'Z'.code.toByte()
assertNull(NetbiosParser.parse(badAlphabet, badAlphabet.size))
// The length byte must be exactly 32; anything else is a different protocol on this port.
val badLen = good.copyOf()
badLen[12] = 16
assertNull(NetbiosParser.parse(badLen, badLen.size))
// A name that decodes to nothing but padding is not a name.
val blank = nbnsPacket("", 0x00)
assertNull(NetbiosParser.parse(blank, blank.size))
val junk = ByteArray(80) { (it * 13).toByte() }
NetbiosParser.parse(junk, junk.size)
}
@Test
fun netbiosControlCharactersInANameAreNeutralised() {
// Legal first-level encoding, illegal content: these bytes decode cleanly and would
// otherwise reach a JSON document, and from there somebody's terminal.
val p = nbnsPacket("A\u0001B\u0002C", 0x00)
assertEquals("A?B?C", assertNotNull(NetbiosParser.parse(p, p.size)).name)
}
// ---- WS-Discovery ------------------------------------------------------------------------
private val wsdHello = bytes(
"""<?xml version="1.0" encoding="utf-8"?>
<soap:Envelope xmlns:soap="http://www.w3.org/2003/05/soap-envelope"
xmlns:wsa="http://schemas.xmlsoap.org/ws/2004/08/addressing"
xmlns:wsd="http://schemas.xmlsoap.org/ws/2005/04/discovery"
xmlns:wsdp="http://schemas.xmlsoap.org/ws/2006/02/devprof">
<soap:Header>
<wsa:To>urn:schemas-xmlsoap-org:ws:2005:04:discovery</wsa:To>
<wsa:Action>http://schemas.xmlsoap.org/ws/2005/04/discovery/Hello</wsa:Action>
<wsa:MessageID>urn:uuid:0a7e6d1b-0000-4000-8000-000000000001</wsa:MessageID>
</soap:Header>
<soap:Body>
<wsd:Hello>
<wsa:EndpointReference>
<wsa:Address>urn:uuid:9f8e7d6c-1111-4222-8333-444455556666</wsa:Address>
</wsa:EndpointReference>
<wsd:Types>wsdp:Device pub:Computer</wsd:Types>
<wsd:XAddrs>http://192.168.1.77:5357/8f2c-4b1a/</wsd:XAddrs>
<wsd:MetadataVersion>1</wsd:MetadataVersion>
</wsd:Hello>
</soap:Body>
</soap:Envelope>"""
)
@Test
fun wsdHelloYieldsActionUuidTypesAndXaddrs() {
val m = assertNotNull(WsdParser.parse(wsdHello, wsdHello.size))
assertEquals("Hello", m.action)
assertEquals(
"urn:uuid:9f8e7d6c-1111-4222-8333-444455556666", m.deviceUuid,
"the EndpointReference address is the device identity, not the header MessageID",
)
assertEquals("wsdp:Device pub:Computer", m.types)
assertEquals("http://192.168.1.77:5357/8f2c-4b1a/", m.xaddrs)
}
@Test
fun wsdProbeCarriesNoDeviceIdentityAndIsStillRecorded() {
val probe = bytes(
"""<soap:Envelope xmlns:soap="http://www.w3.org/2003/05/soap-envelope"
xmlns:wsa="http://schemas.xmlsoap.org/ws/2004/08/addressing"
xmlns:wsd="http://schemas.xmlsoap.org/ws/2005/04/discovery">
<soap:Header>
<wsa:Action>http://schemas.xmlsoap.org/ws/2005/04/discovery/Probe</wsa:Action>
</soap:Header>
<soap:Body><wsd:Probe><wsd:Types>wsdp:Device</wsd:Types></wsd:Probe></soap:Body>
</soap:Envelope>"""
)
val m = assertNotNull(WsdParser.parse(probe, probe.size))
assertEquals("Probe", m.action)
assertNull(m.deviceUuid)
assertEquals("wsdp:Device", m.types)
}
@Test
fun wsdMalformedIsRejectedWithoutThrowing() {
assertNull(WsdParser.parse(ByteArray(0), 0))
assertNull(WsdParser.parse(wsdHello, wsdHello.size + 100), "declared length past the buffer")
val prose = bytes("hello world")
assertNull(WsdParser.parse(prose, prose.size), "not a SOAP envelope")
// An envelope with no leaf worth reading is an absence, not an empty device.
val bare = bytes("<soap:Envelope></soap:Envelope>")
assertNull(WsdParser.parse(bare, bare.size))
// Cut mid-element: whatever was complete is extracted, the rest is simply absent.
val cut = wsdHello.copyOf(wsdHello.size / 2)
WsdParser.parse(cut, cut.size)?.let { assertNull(it.xaddrs, "an unterminated element was invented") }
val junk = ByteArray(512) { (it * 7).toByte() }
assertNull(WsdParser.parse(junk, junk.size))
}
@Test
fun wsdHostileXmlCostsNothingBecauseNothingParsesItAsXml() {
// A billion-laughs entity bomb. Against a real XML parser this expands to gigabytes; here
// the entities are never resolved, which is the entire reason for not using one.
val bomb = bytes(
"<!DOCTYPE lolz [<!ENTITY lol \"lol\">" +
(0..8).joinToString("") { i ->
val prev = if (i == 0) "" else (i - 1).toString()
"<!ENTITY lol$i \"&lol$prev;&lol$prev;\">"
} +
"]><soap:Envelope><wsa:Action>x/Bye</wsa:Action><body>&lol8;</body></soap:Envelope>"
)
assertEquals("Bye", assertNotNull(WsdParser.parse(bomb, bomb.size)).action)
// Nesting deep enough to blow a recursive-descent parser's stack.
val deep = bytes("<soap:Envelope>" + "<a>".repeat(20_000) + "</soap:Envelope>")
WsdParser.parse(deep, deep.size)
// A payload far larger than any real datagram, pinning that the text cap is applied before
// the matching rather than after it.
val huge = bytes("<soap:Envelope><wsa:Action>x/Hello</wsa:Action>" + "z".repeat(200_000))
assertEquals("Hello", assertNotNull(WsdParser.parse(huge, huge.size)).action)
}
}
@@ -1,92 +0,0 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.protocol
import kotlinx.serialization.json.Json
import java.net.URL
import javax.net.ssl.HttpsURLConnection
/**
* The control-plane client (probe-protocol.md §2): enrollment, profile, sessions over
* SPKI-pinned HTTPS. Uses HttpsURLConnection (available since Android API 1, unlike
* java.net.http.HttpClient which needs API 34) with a pin-based SSLSocketFactory and hostname
* verification DISABLED: trust is the SPKI pin, never the certificate name (self-signed servers
* with no SAN are first-class). Blocking; the Android layer wraps calls in coroutines.
*
* @param controlUrl e.g. "https://fmr-1.echo-lot.app:8443"
* @param pins the `pin-sha256` value(s) from the enrollment QR (base64, no prefix)
*/
class ControlClient(private val controlUrl: String, pins: Set<String>) {
private val json = Json { ignoreUnknownKeys = true }
private val socketFactory = Pinning.sslContext(pins).socketFactory
private fun open(path: String, method: String, credential: String?): HttpsURLConnection {
val conn = URL(controlUrl.trimEnd('/') + path).openConnection() as HttpsURLConnection
conn.sslSocketFactory = socketFactory
conn.setHostnameVerifier { _, _ -> true } // pin is the trust, not the name
conn.requestMethod = method
conn.connectTimeout = 10_000
conn.readTimeout = 10_000
credential?.let { conn.setRequestProperty("Authorization", "Bearer $it") }
return conn
}
private fun body(conn: HttpsURLConnection): String {
val stream = if (conn.responseCode in 200..299) conn.inputStream else conn.errorStream
return stream?.bufferedReader()?.use { it.readText() } ?: ""
}
private fun writeJson(conn: HttpsURLConnection, payload: String) {
conn.doOutput = true
conn.setRequestProperty("Content-Type", "application/json")
conn.outputStream.use { it.write(payload.toByteArray()) }
}
// Minimal JSON string literal (the only bodies we send are one short field).
private fun jstr(s: String): String {
val sb = StringBuilder("\"")
for (c in s) when (c) {
'"' -> sb.append("\\\"")
'\\' -> sb.append("\\\\")
'\n' -> sb.append("\\n")
'\r' -> sb.append("\\r")
'\t' -> sb.append("\\t")
else -> sb.append(c)
}
return sb.append('"').toString()
}
/** Redeem a single-use enrollment token for a device credential (§2.1). */
fun enroll(token: String, name: String? = null): EnrollResponse {
val conn = open("/v1/enroll", "POST", null)
conn.setRequestProperty("Authorization", "Bearer $token")
writeJson(conn, if (name != null) """{"name":${jstr(name)}}""" else "{}")
check(conn.responseCode == 201) { "enroll failed: ${conn.responseCode} ${body(conn)}" }
return json.decodeFromString(EnrollResponse.serializer(), body(conn))
}
fun profile(credential: String): Profile {
val conn = open("/v1/profile", "GET", credential)
check(conn.responseCode == 200) { "profile failed: ${conn.responseCode} ${body(conn)}" }
return json.decodeFromString(Profile.serializer(), body(conn))
}
fun createSession(credential: String, target: String): SessionResponse {
val conn = open("/v1/sessions", "POST", credential)
writeJson(conn, """{"target":${jstr(target)}}""")
check(conn.responseCode == 201) { "session failed: ${conn.responseCode} ${body(conn)}" }
return json.decodeFromString(SessionResponse.serializer(), body(conn))
}
fun observations(credential: String, sessionId: String): String {
val conn = open("/v1/sessions/$sessionId/observations", "GET", credential)
check(conn.responseCode == 200) { "observations failed: ${conn.responseCode}" }
return body(conn)
}
fun deleteSession(credential: String, sessionId: String) {
open("/v1/sessions/$sessionId", "DELETE", credential).responseCode
}
}
@@ -1,53 +0,0 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.protocol
import javax.crypto.Mac
import javax.crypto.spec.SecretKeySpec
/**
* The protocol crypto primitives, matching the server exactly (probe-protocol.md §2.4/§3.1):
* HMAC-SHA256 for the data-plane gate, and HKDF-SHA256 for the session key
* `HKDF(ikm = device_credential, salt = key_salt, info = "echolot-v1/" + session_id)`.
* JDK-only (javax.crypto) no third-party crypto.
*/
object Crypto {
fun hmacSha256(key: ByteArray, data: ByteArray): ByteArray =
Mac.getInstance("HmacSHA256").run {
init(SecretKeySpec(key, "HmacSHA256"))
doFinal(data)
}
/** First 4 bytes of HMAC-SHA256 — the wire anti-abuse gate (spec §3.1). */
fun hmac32(key: ByteArray, data: ByteArray): ByteArray = hmacSha256(key, data).copyOf(4)
/**
* HKDF-SHA256 (RFC 5869) extract-then-expand. The JDK exposes no HKDF, so it is built from
* HMAC small and standard.
*/
fun hkdfSha256(ikm: ByteArray, salt: ByteArray, info: ByteArray, length: Int): ByteArray {
val prk = hmacSha256(if (salt.isEmpty()) ByteArray(32) else salt, ikm) // extract
val out = ByteArray(length)
var t = ByteArray(0)
var pos = 0
var counter = 1
while (pos < length) {
val mac = Mac.getInstance("HmacSHA256").apply { init(SecretKeySpec(prk, "HmacSHA256")) }
mac.update(t)
mac.update(info)
mac.update(counter.toByte())
t = mac.doFinal()
val n = minOf(t.size, length - pos)
t.copyInto(out, pos, 0, n)
pos += n
counter++
}
return out
}
/** Derives the 32-byte session key for a session (spec §2.4). */
fun sessionKey(credential: String, keySalt: ByteArray, sessionId: String): ByteArray =
hkdfSha256(credential.toByteArray(), keySalt, "echolot-v1/$sessionId".toByteArray(), 32)
}
@@ -1,64 +0,0 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.protocol
import kotlinx.serialization.SerialName
import kotlinx.serialization.Serializable
import kotlinx.serialization.json.JsonElement
/** Control-plane JSON shapes (probe-protocol.md §2). Only fields the client uses are modeled;
* unknown fields are ignored by the lenient Json in [ControlClient]. */
@Serializable
data class EnrollResponse(
@SerialName("device_id") val deviceId: String,
val credential: String,
)
@Serializable
data class Target(
val id: String,
val ip4: String? = null,
val ip6: String? = null,
@SerialName("udp_port") val udpPort: Int = 0,
@SerialName("tcp_port") val tcpPort: Int = 0,
@SerialName("stun_port") val stunPort: Int = 0,
)
@Serializable
data class SelfTest(
@SerialName("mtu_ok") val mtuOk: Boolean? = null,
@SerialName("sysctl_ok") val sysctlOk: Boolean? = null,
)
@Serializable
data class Profile(
@SerialName("profile_version") val profileVersion: Int = 0,
val name: String = "",
@SerialName("server_version") val serverVersion: String = "",
val capabilities: List<String> = emptyList(),
val targets: List<Target> = emptyList(),
@SerialName("canary_zone") val canaryZone: String = "",
@SerialName("server_selftest") val serverSelftest: SelfTest? = null,
val pins: List<String> = emptyList(),
) {
fun supports(capability: String) = capability in capabilities
}
@Serializable
data class SessionResponse(
@SerialName("session_id") val sessionId: String,
@SerialName("key_salt") val keySalt: String, // base64
val epoch: String,
@SerialName("expires_s") val expiresS: Int,
)
/** Observations bundle (§6). Kept as raw JSON where the shape is still evolving server-side. */
@Serializable
data class Observations(
val udp: JsonElement? = null,
val tcp: JsonElement? = null,
@SerialName("connect_back") val connectBack: JsonElement? = null,
@SerialName("dns_canary") val dnsCanary: JsonElement? = null,
)
@@ -1,39 +0,0 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.protocol
import java.security.MessageDigest
import java.security.cert.X509Certificate
import javax.net.ssl.SSLContext
import javax.net.ssl.X509TrustManager
/**
* SPKI-pinned trust (probe-protocol.md §1): the client trusts the server ONLY against the
* `pin-sha256` from enrollment CA validation is not required and self-signed is first-class.
* The pin is base64(SHA-256(SubjectPublicKeyInfo)), RFC 7469.
*/
object Pinning {
fun spkiPin(cert: X509Certificate): String {
val spki = cert.publicKey.encoded // DER SubjectPublicKeyInfo
val digest = MessageDigest.getInstance("SHA-256").digest(spki)
return java.util.Base64.getEncoder().encodeToString(digest)
}
/** An SSLContext that accepts a chain iff its leaf SPKI matches one of the expected pins. */
fun sslContext(expectedPins: Set<String>): SSLContext {
val tm = object : X509TrustManager {
override fun checkServerTrusted(chain: Array<out X509Certificate>, authType: String) {
val leaf = chain.firstOrNull() ?: throw java.security.cert.CertificateException("empty chain")
val pin = spkiPin(leaf)
if (pin !in expectedPins) {
throw java.security.cert.CertificateException("SPKI pin mismatch: got $pin")
}
}
override fun checkClientTrusted(chain: Array<out X509Certificate>, authType: String) = Unit
override fun getAcceptedIssuers(): Array<X509Certificate> = emptyArray()
}
return SSLContext.getInstance("TLS").apply { init(null, arrayOf(tm), null) }
}
}
@@ -1,77 +0,0 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.protocol
import java.net.DatagramPacket
import java.net.DatagramSocket
import java.net.InetSocketAddress
import java.util.Base64
/**
* A data-plane session (probe-protocol.md §3): derives the session key, then sends signed ELT1
* packets to the server's UDP endpoint and reads back verified responses. One session one
* server target. Blocking; the caller owns threading.
*/
class ProbeSession(
private val credential: String,
private val session: SessionResponse,
private val serverHost: String,
private val serverUdpPort: Int,
) : AutoCloseable {
private val key: ByteArray =
Crypto.sessionKey(credential, Base64.getDecoder().decode(session.keySalt), session.sessionId)
private val prefix: ByteArray = Wire.wirePrefix(session.sessionId)
private val epochNanos = System.nanoTime()
private val socket = DatagramSocket().apply { soTimeout = 3000 }
private val server = InetSocketAddress(serverHost, serverUdpPort)
private var seq = 0
private fun nowNs() = System.nanoTime() - epochNanos
/**
* One ECHO round trip. Returns RTT in ms and the server's observation, or null on loss.
*
* The response is capped at the request size (§3.4 anti-amplification) and the observation
* block is 40 bytes, so the request must be at least header+40 = 72 bytes for the full
* observation to fit hence the 40 default padding. Smaller requests still measure RTT.
*/
fun echo(paddingBytes: Int = 40): EchoResult? {
val t0 = System.nanoTime()
val pkt = Wire.build(Wire.TYPE_ECHO_REQ, prefix, ++seq, nowNs(), key, ByteArray(paddingBytes))
socket.send(DatagramPacket(pkt, pkt.size, server))
val resp = receive(Wire.TYPE_ECHO_RESP) ?: return null
val rttMs = (System.nanoTime() - t0) / 1_000_000.0
return EchoResult(rttMs, Observation.parse(resp.payload))
}
/** One MTU probe of [totalSize] bytes (DF is set by the OS on the socket where supported).
* Returns the size the server acknowledged receiving, or null if the probe was lost. */
fun mtuProbe(totalSize: Int): Int? {
val payloadLen = (totalSize - Wire.HEADER_SIZE).coerceAtLeast(0)
val pkt = Wire.build(Wire.TYPE_MTU_PROBE, prefix, ++seq, nowNs(), key, ByteArray(payloadLen))
socket.send(DatagramPacket(pkt, pkt.size, server))
val resp = receive(Wire.TYPE_MTU_ACK) ?: return null
if (resp.payload.size < 4) return null
return ((resp.payload[0].toInt() and 0xFF) shl 24) or
((resp.payload[1].toInt() and 0xFF) shl 16) or
((resp.payload[2].toInt() and 0xFF) shl 8) or
(resp.payload[3].toInt() and 0xFF)
}
private fun receive(wantType: Int): Wire.Packet? {
val buf = ByteArray(2048)
return try {
val dp = DatagramPacket(buf, buf.size)
socket.receive(dp)
Wire.parseVerified(buf, dp.length, key)?.takeIf { it.type == wantType }
} catch (e: java.net.SocketTimeoutException) {
null
}
}
override fun close() = socket.close()
data class EchoResult(val rttMs: Double, val observation: Observation?)
}
@@ -1,115 +0,0 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.protocol
import java.nio.ByteBuffer
import java.nio.ByteOrder
/**
* The binary UDP probe protocol wire format (probe-protocol.md §3.1): a fixed 32-byte header
* plus payload, HMAC-gated. Mirrors the Go server's dataplane package byte-for-byte.
*
* ```
* 0 4 magic "ELT1" 8 8 session_prefix (first 8 bytes of session id)
* 4 1 type 16 4 seq
* 5 1 flags 20 8 t_ns (sender clock, ns since session epoch)
* 6 2 payload_len 28 4 hmac32(session_key, header[0..28] || payload)
* ```
*/
object Wire {
const val HEADER_SIZE = 32
val MAGIC = byteArrayOf('E'.code.toByte(), 'L'.code.toByte(), 'T'.code.toByte(), '1'.code.toByte())
const val TYPE_ECHO_REQ: Int = 0x01
const val TYPE_ECHO_RESP: Int = 0x02
const val TYPE_TIMESYNC_REQ: Int = 0x07
const val TYPE_TIMESYNC_RSP: Int = 0x08
const val TYPE_MTU_PROBE: Int = 0x09
const val TYPE_MTU_ACK: Int = 0x0A
const val TYPE_DELAYED_ECHO: Int = 0x0B
/** The 8-byte on-the-wire prefix = first 16 hex chars of the session id, decoded. */
fun wirePrefix(sessionId: String): ByteArray {
require(sessionId.length >= 16) { "session id too short" }
val p = ByteArray(8)
for (i in 0 until 8) {
p[i] = ((hex(sessionId[i * 2]) shl 4) or hex(sessionId[i * 2 + 1])).toByte()
}
return p
}
private fun hex(c: Char): Int = when (c) {
in '0'..'9' -> c - '0'
in 'a'..'f' -> c - 'a' + 10
in 'A'..'F' -> c - 'A' + 10
else -> 0
}
/** Builds a signed packet ready to send. */
fun build(
type: Int, sessionPrefix: ByteArray, seq: Int, tNs: Long, key: ByteArray,
payload: ByteArray = ByteArray(0),
): ByteArray {
val buf = ByteBuffer.allocate(HEADER_SIZE + payload.size).order(ByteOrder.BIG_ENDIAN)
buf.put(MAGIC)
buf.put(type.toByte())
buf.put(0) // flags
buf.putShort(payload.size.toShort())
buf.put(sessionPrefix, 0, 8)
buf.putInt(seq)
buf.putLong(tNs)
buf.position(28) // leave hmac slot; fill after
buf.putInt(0)
buf.put(payload)
val bytes = buf.array()
// HMAC over header[0..28] || payload (the hmac slot itself excluded).
val mac = Crypto.hmacSha256(key, concat(bytes, 0, 28, bytes, HEADER_SIZE, payload.size))
mac.copyInto(bytes, 28, 0, 4)
return bytes
}
/** A parsed, HMAC-verified inbound packet. */
data class Packet(val type: Int, val seq: Int, val tNs: Long, val payload: ByteArray)
/** Parses and verifies an inbound datagram; null if malformed or the HMAC fails. */
fun parseVerified(data: ByteArray, len: Int, key: ByteArray): Packet? {
if (len < HEADER_SIZE) return null
for (i in MAGIC.indices) if (data[i] != MAGIC[i]) return null
val bb = ByteBuffer.wrap(data, 0, len).order(ByteOrder.BIG_ENDIAN)
val type = bb.get(4).toInt() and 0xFF
val payloadLen = bb.getShort(6).toInt() and 0xFFFF
if (HEADER_SIZE + payloadLen > len) return null
val expect = Crypto.hmacSha256(key, concat(data, 0, 28, data, HEADER_SIZE, payloadLen))
for (i in 0 until 4) if (expect[i] != data[28 + i]) return null
val seq = bb.getInt(16)
val tNs = bb.getLong(20)
val payload = data.copyOfRange(HEADER_SIZE, HEADER_SIZE + payloadLen)
return Packet(type, seq, tNs, payload)
}
private fun concat(a: ByteArray, aOff: Int, aLen: Int, b: ByteArray, bOff: Int, bLen: Int): ByteArray {
val out = ByteArray(aLen + bLen)
a.copyInto(out, 0, aOff, aOff + aLen)
b.copyInto(out, aLen, bOff, bOff + bLen)
return out
}
}
/** Server observation block appended to ECHO_RESP (spec §3.3), fixed 40 bytes. */
data class Observation(
val tRxNs: Long, val tTxNs: Long, val observedPort: Int, val receivedSize: Int,
) {
companion object {
fun parse(payload: ByteArray): Observation? {
if (payload.size < 40) return null
val bb = ByteBuffer.wrap(payload).order(ByteOrder.BIG_ENDIAN)
return Observation(
tRxNs = bb.getLong(0),
tTxNs = bb.getLong(8),
observedPort = bb.getShort(32).toInt() and 0xFFFF,
receivedSize = bb.getInt(36),
)
}
}
}
@@ -1,76 +0,0 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.protocol
import kotlin.test.Test
import kotlin.test.assertEquals
import kotlin.test.assertNull
import kotlin.test.assertNotNull
import kotlin.test.assertTrue
class CryptoWireTest {
@Test
fun hkdfMatchesRfc5869Vector() {
// RFC 5869 Appendix A.1 (SHA-256).
val ikm = ByteArray(22) { 0x0b }
val salt = byteArrayOf(0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12)
val info = byteArrayOf(
0xf0.toByte(), 0xf1.toByte(), 0xf2.toByte(), 0xf3.toByte(), 0xf4.toByte(),
0xf5.toByte(), 0xf6.toByte(), 0xf7.toByte(), 0xf8.toByte(), 0xf9.toByte(),
)
val okm = Crypto.hkdfSha256(ikm, salt, info, 42)
val expect = "3cb25f25faacd57a90434f64d0362f2a" +
"2d2d0a90cf1a5a4c5db02d56ecc4c5bf" +
"34007208d5b887185865"
assertEquals(expect, okm.joinToString("") { "%02x".format(it) })
}
@Test
fun wirePrefixDecodesHex() {
val prefix = Wire.wirePrefix("805a43f8395ae08ace7a14803766cb11")
assertEquals("805a43f8395ae08a", prefix.joinToString("") { "%02x".format(it) })
}
@Test
fun buildThenParseRoundTripsAndVerifies() {
val key = ByteArray(32) { it.toByte() }
val prefix = ByteArray(8) { (it + 1).toByte() }
val payload = "hello-echolot".toByteArray()
val pkt = Wire.build(Wire.TYPE_ECHO_REQ, prefix, 7, 123_456L, key, payload)
assertEquals(Wire.HEADER_SIZE + payload.size, pkt.size)
val parsed = Wire.parseVerified(pkt, pkt.size, key)
assertNotNull(parsed)
assertEquals(Wire.TYPE_ECHO_REQ, parsed.type)
assertEquals(7, parsed.seq)
assertEquals(123_456L, parsed.tNs)
assertEquals("hello-echolot", String(parsed.payload))
}
@Test
fun tamperedHmacIsRejected() {
val key = ByteArray(32) { it.toByte() }
val pkt = Wire.build(Wire.TYPE_ECHO_REQ, ByteArray(8), 1, 0, key, ByteArray(4))
pkt[pkt.size - 1] = (pkt[pkt.size - 1].toInt() xor 0xFF).toByte() // flip a payload byte
assertNull(Wire.parseVerified(pkt, pkt.size, key))
}
@Test
fun wrongKeyIsRejected() {
val pkt = Wire.build(Wire.TYPE_ECHO_REQ, ByteArray(8), 1, 0, ByteArray(32) { 1 }, ByteArray(0))
assertNull(Wire.parseVerified(pkt, pkt.size, ByteArray(32) { 2 }))
}
@Test
fun observationParses() {
// 40-byte block: t_rx, t_tx, 16-byte addr, port, ttl/dscp, size.
val b = ByteArray(40)
b[33] = 0x1F // port low byte = 8191... set port bytes 32..33
b[32] = 0x00
val obs = Observation.parse(b)
assertNotNull(obs)
assertTrue(obs.observedPort in 0..65535)
}
}
@@ -1,66 +0,0 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.protocol
import kotlin.test.Test
import kotlin.test.assertNotNull
import kotlin.test.assertTrue
/**
* End-to-end test of the Kotlin client against a REAL running server. It self-skips unless the
* environment provides a live target, so it never breaks CI (no network / no server):
*
* ECHOLOT_LIVE_URL = https://fmr-1.echo-lot.app:8443
* ECHOLOT_LIVE_PIN = <base64 pin-sha256>
* ECHOLOT_LIVE_CRED = <device credential from an enrollment>
* ECHOLOT_LIVE_UDP = fmr-1.echo-lot.app:8442
* ECHOLOT_LIVE_TARGET = fmr (profile target id)
*
* The harness (test-fmr.sh) mints a token over SSH, enrolls via the public control plane, and
* exports these proving the client talks to the deployed server over the wire.
*/
class LiveServerTest {
private val url = System.getenv("ECHOLOT_LIVE_URL")
private val pin = System.getenv("ECHOLOT_LIVE_PIN")
private val cred = System.getenv("ECHOLOT_LIVE_CRED")
private val udp = System.getenv("ECHOLOT_LIVE_UDP")
private val target = System.getenv("ECHOLOT_LIVE_TARGET") ?: "fmr"
@Test
fun fullFlowAgainstLiveServer() {
if (url == null || pin == null || cred == null || udp == null) {
println("LiveServerTest skipped (no ECHOLOT_LIVE_* env)")
return
}
val control = ControlClient(url, setOf(pin))
val profile = control.profile(cred)
println("profile: name=${profile.name} v=${profile.serverVersion} caps=${profile.capabilities}")
assertTrue(profile.supports("udp-probe"), "server must offer udp-probe")
val session = control.createSession(cred, target)
println("session: ${session.sessionId} expires=${session.expiresS}s")
val (host, port) = udp.split(":").let { it[0] to it[1].toInt() }
ProbeSession(cred, session, host, port).use { ps ->
// ECHO: verified response + observation with our observed port.
val echo = ps.echo(paddingBytes = 64) // ≥40 so the observation block fits (§3.4)
assertNotNull(echo, "no verified ECHO_RESP from live server")
println("echo rtt=${"%.1f".format(echo.rttMs)}ms observedPort=${echo.observation?.observedPort} size=${echo.observation?.receivedSize}")
assertNotNull(echo.observation, "ECHO_RESP missing observation block")
// MTU probe: server acks the size it received.
val acked = ps.mtuProbe(1400)
assertNotNull(acked, "no MTU_ACK from live server")
println("mtu probe 1400 -> server received $acked bytes")
assertTrue(acked!! in 1300..1500, "acked size implausible: $acked")
}
val obs = control.observations(cred, session.sessionId)
println("observations bytes: ${obs.length}")
assertTrue(obs.contains("packets_seen"), "observations should report packets_seen")
control.deleteSession(cred, session.sessionId)
}
}
@@ -0,0 +1,202 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.protocol
/**
* Version compatibility, the client side of the same question the server asks about us.
*
* Versions are SemVer, but what is really being checked is whether the peer speaks a wire protocol
* and schema this build understands; the version is a proxy, and it only works because the
* breaking axis is bumped when the contract changes. Bounds therefore sit at breaking boundaries
* rather than at individual releases a server patch release must never make the app refuse to
* talk to it.
*
* Mirrors `server/internal/compat`. Two implementations of one rule is a duplication worth
* accepting: each side must be able to state and enforce its own limits without asking the other,
* which is the entire point of a compatibility check.
*/
data class SemVer(
val major: Int,
val minor: Int,
val patch: Int,
val pre: String = "",
) : Comparable<SemVer> {
override fun compareTo(other: SemVer): Int {
(major - other.major).let { if (it != 0) return it.coerceIn(-1, 1) }
(minor - other.minor).let { if (it != 0) return it.coerceIn(-1, 1) }
(patch - other.patch).let { if (it != 0) return it.coerceIn(-1, 1) }
// A pre-release sorts below the same version without one (SemVer §11).
return when {
pre == other.pre -> 0
pre.isEmpty() -> 1
other.pre.isEmpty() -> -1
else -> pre.compareTo(other.pre).coerceIn(-1, 1)
}
}
override fun toString(): String = "$major.$minor.$patch" + if (pre.isEmpty()) "" else "-$pre"
/**
* The first version that may break compatibility with this one. Below 1.0.0 the minor is the
* breaking axis (SemVer §4), so 0.4.2's next break is 0.5.0 treating it as 1.0.0 would let
* this build accept a peer it cannot actually talk to.
*/
fun nextBreaking(): SemVer =
if (major == 0) SemVer(0, minor + 1, 0) else SemVer(major + 1, 0, 0)
companion object {
/** Accepts "1.2.3", "v1.2.3" and namespaced tags like "server-v1.2.3". Null if unusable. */
fun parse(raw: String?): SemVer? {
var s = raw?.trim().orEmpty()
if (s.isEmpty()) return null
// Strip a tag prefix ending in "v", guarded on the prefix having no digits so a
// pre-release identifier containing a "v" is left alone.
val v = s.lastIndexOf('v')
if (v >= 0 && v + 1 < s.length && s[v + 1].isDigit() && s.take(v).none { it.isDigit() }) {
s = s.substring(v + 1)
}
s = s.substringBefore('+')
val pre = s.substringAfter('-', "")
val core = s.substringBefore('-')
val parts = core.split(".")
if (parts.size != 3) return null
val nums = parts.map { it.toIntOrNull() ?: return null }
if (nums.any { it < 0 }) return null
return SemVer(nums[0], nums[1], nums[2], pre)
}
}
}
/** [min, max): minimum inclusive, maximum exclusive. A null max is unbounded. */
data class VersionRange(val min: SemVer, val max: SemVer? = null) {
operator fun contains(v: SemVer): Boolean = v >= min && (max == null || v < max)
override fun toString(): String = ">= $min" + (max?.let { ", < $it" } ?: "")
companion object {
fun of(min: String, max: String?): VersionRange {
val lo = requireNotNull(SemVer.parse(min)) { "bad minimum version: $min" }
val hi = max?.takeIf { it.isNotBlank() }?.let {
requireNotNull(SemVer.parse(it)) { "bad maximum version: $it" }
}
require(hi == null || hi > lo) { "maximum $max is not above minimum $min" }
return VersionRange(lo, hi)
}
}
}
/**
* What this build of the app requires of a server, and what it tells servers about itself.
*
* Two separate questions, deliberately not conflated:
*
* - **Protocol version** can these builds talk at all? This is the correctness axis, and its
* breaking boundary is enforced strictly.
* - **Release-version window** should they, per policy? A coarse safety net over the peer's
* SemVer, with bounds at breaking boundaries so a patch release never strands anyone. Both
* sides publish their own, and the operator can tighten the server's.
*
* `MIN_SERVER` is 0.4.2 for a concrete reason, not caution: below it a multi-homed server sent
* granted traffic from an address the session never used, so every downstream packet was dropped
* in transit and reported as 100 % downstream loss. A confidently wrong measurement is worse than
* a refused one, so talking to those builds is not something to allow "just in case".
*/
object Compat {
/** Header the app sets on every control-plane request. */
const val APP_VERSION_HEADER = "X-Echolot-App-Version"
/** The wire contract (probe-protocol.md) this build implements. */
const val PROTOCOL_VERSION = "1.0.0"
const val MIN_SERVER = "0.4.2"
/** Exclusive. The next breaking series is refused until this app is taught about it. */
const val MAX_SERVER = "1.0.0"
val serverRange: VersionRange = VersionRange.of(MIN_SERVER, MAX_SERVER)
enum class Verdict { OK, PROTOCOL_MISMATCH, SERVER_TOO_OLD, SERVER_TOO_NEW, APP_REFUSED, UNKNOWN }
/**
* The result of checking a server, with a message written for the person holding the phone.
* [usable] is what callers branch on; [message] is what they show.
*/
data class Result(val verdict: Verdict, val message: String?) {
val usable: Boolean get() = verdict == Verdict.OK || verdict == Verdict.UNKNOWN
}
/**
* Checks a profile both ways: is the server within our range, and are we within the server's.
*
* Asking both is the point of advertising the window in the profile. Discovering that the
* server will refuse us only when a measurement fails halfway through is a much worse
* experience than being told before the run starts.
*/
fun check(profile: Profile, appVersion: String): Result {
// The protocol version is the axis that decides whether these two builds *can* talk;
// the release-version window below is the operator's policy about whether they *may*.
// Checking the real thing first means a mismatch is reported as what it is.
val ours = SemVer.parse(PROTOCOL_VERSION)!!
val theirs = SemVer.parse(profile.compat.protocolVersion)
if (theirs != null && theirs >= ours.nextBreaking()) {
return Result(
Verdict.PROTOCOL_MISMATCH,
"This server speaks probe protocol ${profile.compat.protocolVersion}; this app " +
"speaks $PROTOCOL_VERSION and does not understand that revision. Update the app.",
)
}
if (theirs != null && ours >= theirs.nextBreaking()) {
return Result(
Verdict.PROTOCOL_MISMATCH,
"This server speaks probe protocol ${profile.compat.protocolVersion}, which this " +
"app ($PROTOCOL_VERSION) has moved past. Update the server.",
)
}
val server = SemVer.parse(profile.serverVersion)
?: return Result(
Verdict.UNKNOWN,
"Server did not report a usable version (\"${profile.serverVersion}\") — " +
"continuing without a compatibility check.",
)
if (server < serverRange.min) {
return Result(
Verdict.SERVER_TOO_OLD,
"This server runs $server; Echolot needs $serverRange. " +
"Measurements against older servers can be wrong rather than merely missing, " +
"so update the server.",
)
}
if (serverRange.max != null && server >= serverRange.max) {
return Result(
Verdict.SERVER_TOO_NEW,
"This server runs $server, which is newer than this app understands " +
"($serverRange). Update the app.",
)
}
// And the server's own view of us.
val app = SemVer.parse(appVersion)
val serverWantsMin = SemVer.parse(profile.compat.appMin)
val serverWantsMax = SemVer.parse(profile.compat.appMax)
if (app != null && serverWantsMin != null) {
if (app < serverWantsMin) {
return Result(
Verdict.APP_REFUSED,
"This server only accepts Echolot $serverWantsMin or newer; this app is " +
"$app. Update the app.",
)
}
if (serverWantsMax != null && app >= serverWantsMax) {
return Result(
Verdict.APP_REFUSED,
"This server refuses Echolot $serverWantsMax and newer; this app is $app. " +
"Use an older app, or a server that has caught up.",
)
}
}
return Result(Verdict.OK, null)
}
}
@@ -4,9 +4,21 @@
package app.echo_lot.protocol package app.echo_lot.protocol
import kotlinx.serialization.json.Json import kotlinx.serialization.json.Json
import kotlinx.serialization.json.jsonObject
import kotlinx.serialization.json.jsonPrimitive
import java.net.URL import java.net.URL
import javax.net.ssl.HttpsURLConnection import javax.net.ssl.HttpsURLConnection
/** The server declined to store this run, per the operator's policy. Not a transport failure. */
class UploadRefused(message: String) : Exception(message)
/**
* The server refused this app's version. Distinct from a transport failure and from an auth
* failure: nothing about the request was wrong, the two builds simply do not go together, and the
* message says which versions do.
*/
class VersionRefused(message: String) : Exception(message)
/** /**
* The control-plane client (probe-protocol.md §2): enrollment, profile, sessions over * The control-plane client (probe-protocol.md §2): enrollment, profile, sessions over
* SPKI-pinned HTTPS. Uses HttpsURLConnection (available since Android API 1, unlike * SPKI-pinned HTTPS. Uses HttpsURLConnection (available since Android API 1, unlike
@@ -16,26 +28,106 @@ import javax.net.ssl.HttpsURLConnection
* *
* @param controlUrl e.g. "https://fmr-1.echo-lot.app:8443" * @param controlUrl e.g. "https://fmr-1.echo-lot.app:8443"
* @param pins the `pin-sha256` value(s) from the enrollment QR (base64, no prefix) * @param pins the `pin-sha256` value(s) from the enrollment QR (base64, no prefix)
* @param appVersion this build's SemVer, sent on every request so the server can refuse a build
* it cannot serve *before* a measurement half-runs (BuildConfig.VERSION_NAME).
*/ */
class ControlClient(private val controlUrl: String, pins: Set<String>) { class ControlClient(
private val controlUrl: String,
pins: Set<String>,
private val appVersion: String = "",
/**
* Addresses to fall back to when the server's name will not resolve, learned from its profile.
*
* A measurement tool that cannot report from a broken network is useless exactly when it
* matters, and a wedged resolver is one of the faults it is built to find it should not also
* be the thing that stops the finding being delivered.
*
* Safe because the pin is the trust and the name is not part of it: the server presents the
* same certificate whether it was reached by name or by address, and a wrong address fails the
* pin like anything else would.
*/
private val fallbackAddrs: List<String> = emptyList(),
) {
private val json = Json { ignoreUnknownKeys = true } private val json = Json { ignoreUnknownKeys = true }
private val socketFactory = Pinning.sslContext(pins).socketFactory private val socketFactory = Pinning.sslContext(pins).socketFactory
/**
* The base URL to use, substituting a cached address only when the name genuinely fails.
*
* Resolved once per client and only on failure, so a working network pays nothing and never
* silently drifts onto an address that may be stale.
*/
private val base: String by lazy { resolveBase() }
private fun resolveBase(): String {
if (fallbackAddrs.isEmpty()) return controlUrl
val uri = runCatching { java.net.URI(controlUrl) }.getOrNull() ?: return controlUrl
val host = uri.host ?: return controlUrl
if (runCatching { java.net.InetAddress.getByName(host) }.isSuccess) return controlUrl
val port = if (uri.port > 0) uri.port else 443
for (ip in fallbackAddrs) {
// Checked rather than assumed: on a v4-only network a v6 address would otherwise be
// chosen and fail slowly, which is the wrong answer delivered late.
val reachable = runCatching {
java.net.Socket().use { sock ->
sock.connect(java.net.InetSocketAddress(ip, port), 4000)
true
}
}.getOrDefault(false)
if (reachable) {
val literal = if (ip.contains(':')) "[$ip]" else ip
return uri.scheme + "://" + literal + ":" + port
}
}
return controlUrl
}
private fun open(path: String, method: String, credential: String?): HttpsURLConnection { private fun open(path: String, method: String, credential: String?): HttpsURLConnection {
val conn = URL(controlUrl.trimEnd('/') + path).openConnection() as HttpsURLConnection val conn = URL(base.trimEnd('/') + path).openConnection() as HttpsURLConnection
conn.sslSocketFactory = socketFactory conn.sslSocketFactory = socketFactory
conn.setHostnameVerifier { _, _ -> true } // pin is the trust, not the name conn.setHostnameVerifier { _, _ -> true } // pin is the trust, not the name
conn.requestMethod = method conn.requestMethod = method
conn.connectTimeout = 10_000 conn.connectTimeout = 10_000
conn.readTimeout = 10_000 conn.readTimeout = 10_000
credential?.let { conn.setRequestProperty("Authorization", "Bearer $it") } credential?.let { conn.setRequestProperty("Authorization", "Bearer $it") }
if (appVersion.isNotBlank()) conn.setRequestProperty(Compat.APP_VERSION_HEADER, appVersion)
return conn return conn
} }
/**
* 426 is the server saying "your version, not your request". Raised as a distinct exception
* from every call site so callers never report it as a network error the whole value of the
* check is that the failure is legible.
*/
private fun checkVersion(conn: HttpsURLConnection, body: String) {
if (conn.responseCode == 426) throw VersionRefused(extractError(body) ?: body.take(200))
}
/**
* Pulls the "error" string out of a JSON body.
*
* Parsed rather than pattern-matched: an encoder may legitimately escape characters in the
* message (Go escapes ">" by default), and a regex hands the user "needs \u003e= 0.2.0".
* The parser knows how to undo every escape; a regex would have to be taught each one.
*/
private fun extractError(body: String): String? = runCatching {
json.parseToJsonElement(body).jsonObject["error"]?.jsonPrimitive?.content
}.getOrNull()
/**
* Reads the response body, and turns a 426 into [VersionRefused] first.
*
* Every call goes through here, so the version check cannot be forgotten at a new call site
* the alternative (a check per method) is exactly the kind of thing that gets missed once and
* then reports "upload failed: 426" to a user for a year.
*/
private fun body(conn: HttpsURLConnection): String { private fun body(conn: HttpsURLConnection): String {
val stream = if (conn.responseCode in 200..299) conn.inputStream else conn.errorStream val stream = if (conn.responseCode in 200..299) conn.inputStream else conn.errorStream
return stream?.bufferedReader()?.use { it.readText() } ?: "" val text = stream?.bufferedReader()?.use { it.readText() } ?: ""
checkVersion(conn, text)
return text
} }
private fun writeJson(conn: HttpsURLConnection, payload: String) { private fun writeJson(conn: HttpsURLConnection, payload: String) {
@@ -63,27 +155,141 @@ class ControlClient(private val controlUrl: String, pins: Set<String>) {
val conn = open("/v1/enroll", "POST", null) val conn = open("/v1/enroll", "POST", null)
conn.setRequestProperty("Authorization", "Bearer $token") conn.setRequestProperty("Authorization", "Bearer $token")
writeJson(conn, if (name != null) """{"name":${jstr(name)}}""" else "{}") writeJson(conn, if (name != null) """{"name":${jstr(name)}}""" else "{}")
check(conn.responseCode == 201) { "enroll failed: ${conn.responseCode} ${body(conn)}" } val text = body(conn) // reads and raises VersionRefused on 426
return json.decodeFromString(EnrollResponse.serializer(), body(conn)) check(conn.responseCode == 201) { "enroll failed: ${conn.responseCode} $text" }
return json.decodeFromString(EnrollResponse.serializer(), text)
} }
fun profile(credential: String): Profile { fun profile(credential: String): Profile {
val conn = open("/v1/profile", "GET", credential) val conn = open("/v1/profile", "GET", credential)
check(conn.responseCode == 200) { "profile failed: ${conn.responseCode} ${body(conn)}" } val text = body(conn)
return json.decodeFromString(Profile.serializer(), body(conn)) check(conn.responseCode == 200) { "profile failed: ${conn.responseCode} $text" }
return json.decodeFromString(Profile.serializer(), text)
} }
fun createSession(credential: String, target: String): SessionResponse { fun createSession(credential: String, target: String): SessionResponse {
val conn = open("/v1/sessions", "POST", credential) val conn = open("/v1/sessions", "POST", credential)
writeJson(conn, """{"target":${jstr(target)}}""") writeJson(conn, """{"target":${jstr(target)}}""")
check(conn.responseCode == 201) { "session failed: ${conn.responseCode} ${body(conn)}" } val text = body(conn)
return json.decodeFromString(SessionResponse.serializer(), body(conn)) check(conn.responseCode == 201) { "session failed: ${conn.responseCode} $text" }
return json.decodeFromString(SessionResponse.serializer(), text)
}
/**
* Requests a §5 action. The server creates an asymmetric grant for the granted ones
* (downtrain / big_send) and starts sending toward the session's observed data-plane source,
* so the caller must already have sent at least one ECHO. Returns the raw JSON reply.
*/
fun action(credential: String, sessionId: String, bodyJson: String): String {
val conn = open("/v1/sessions/$sessionId/actions", "POST", credential)
writeJson(conn, bodyJson)
val body = body(conn)
check(conn.responseCode in 200..299) { "action failed: ${conn.responseCode} $body" }
return body
}
/**
* Uploads one measurement document. The body is sent exactly as given whatever the
* anonymizer produced is what the server stores, so what the user was shown is what left
* the device. Returns the server's index entry as raw JSON.
*
* A refusal is not an error condition to retry: 403 means the operator's policy says no
* (uploads off, accounts required, or not anonymized enough), so it is surfaced as
* [UploadRefused] for the caller to show rather than swallow.
*/
fun uploadRun(credential: String, documentJson: String): String {
val conn = open("/v1/runs", "POST", credential)
writeJson(conn, documentJson)
val body = body(conn)
when (conn.responseCode) {
in 200..299 -> return body
403 -> throw UploadRefused(extractError(body) ?: body.take(200))
413 -> throw UploadRefused("run is larger than this server accepts: $body")
else -> error("upload failed: ${conn.responseCode} $body")
}
}
/**
* Reports where adbd's wireless-debug listener can be reached on this device's LAN.
*
* Dev scaffolding, not a measurement: it exists because mDNS does not cross subnets, so a
* developer working from a different network cannot discover the port that rotates every few
* minutes. A device sitting on the test LAN can see it and say so. Deliberately not part of
* the probe protocol's capability set the server documents it as a dev relay.
*/
fun reportAdbEndpoint(
credential: String,
host: String,
port: Int,
deviceName: String? = null,
note: String? = null,
): String {
val conn = open("/v1/devtools/adb-endpoint", "POST", credential)
val fields = buildString {
append("""{"host":${jstr(host)},"port":$port""")
deviceName?.let { append(""","device_name":${jstr(it)}""") }
note?.let { append(""","note":${jstr(it)}""") }
append("}")
}
writeJson(conn, fields)
val body = body(conn)
check(conn.responseCode in 200..299) {
"adb endpoint report failed: ${conn.responseCode} $body"
}
return body
}
/** Lists this device's runs stored on the server. */
fun listRuns(credential: String): String {
val conn = open("/v1/runs", "GET", credential)
val text = body(conn)
check(conn.responseCode == 200) { "list runs failed: ${conn.responseCode} $text" }
return text
}
fun getRun(credential: String, runId: String): String {
val conn = open("/v1/runs/$runId", "GET", credential)
val text = body(conn)
check(conn.responseCode == 200) { "get run failed: ${conn.responseCode} $text" }
return text
}
fun deleteRun(credential: String, runId: String) {
open("/v1/runs/$runId", "DELETE", credential).responseCode
}
/**
* Ties this device to the person the ID token identifies.
*
* The device credential proves *which device*, the token proves *which person*; the server
* requires both. Returns the raw JSON reply (account id and display name).
*/
fun linkAccount(credential: String, idToken: String): String {
val conn = open("/v1/account/link", "POST", credential)
writeJson(conn, """{"id_token":${jstr(idToken)}}""")
val text = body(conn)
check(conn.responseCode in 200..299) { "sign-in failed: ${conn.responseCode} $text" }
return text
}
/** Signs out on this device. The device stays enrolled. */
fun unlinkAccount(credential: String) {
open("/v1/account/link", "DELETE", credential).responseCode
}
/** Whether anyone is signed in on this device, and who. */
fun accountStatus(credential: String): String {
val conn = open("/v1/account", "GET", credential)
val text = body(conn)
check(conn.responseCode == 200) { "account status failed: ${conn.responseCode} $text" }
return text
} }
fun observations(credential: String, sessionId: String): String { fun observations(credential: String, sessionId: String): String {
val conn = open("/v1/sessions/$sessionId/observations", "GET", credential) val conn = open("/v1/sessions/$sessionId/observations", "GET", credential)
check(conn.responseCode == 200) { "observations failed: ${conn.responseCode}" } val text = body(conn)
return body(conn) check(conn.responseCode == 200) { "observations failed: ${conn.responseCode} $text" }
return text
} }
fun deleteSession(credential: String, sessionId: String) { fun deleteSession(credential: String, sessionId: String) {
@@ -0,0 +1,161 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.protocol
import java.net.URLDecoder
import java.net.URLEncoder
/**
* The enrollment bootstrap of probe-protocol.md §2.1.
*
* ```
* echolot://enroll?v=1&u=<control-URL, urlencoded>&p=pin-sha256:<b64 SPKI hash>&t=<token>
* ```
*
* One string carries everything a device needs to start trusting a server: where it is, which key
* to pin, and a single-use token proving the operator meant to admit this device. That is the
* whole point it is why enrollment can be a paste or a QR scan rather than three fields typed
* from a screenshot, which is what people actually do wrong.
*
* **The link is a secret.** It contains a bearer token; anyone who sees it before the device does
* can enroll instead. Tokens are single-use and short-lived precisely so a leaked link is a
* bounded problem, but it should be treated like a password while it is live.
*/
data class EnrollmentLink(
/** e.g. "https://fmr-1.echo-lot.app:8443" */
val controlUrl: String,
/** Base64 SPKI hash, without the "pin-sha256:" prefix — the form [ControlClient] wants. */
val pin: String,
val token: String,
) {
/** Rebuilds the URI. Round-trips with [parse]; used for tests and for sharing a link on. */
fun toUri(): String = buildString {
append("echolot://enroll?v=1")
append("&u=").append(enc(controlUrl))
append("&p=").append(enc(PIN_PREFIX + pin))
append("&t=").append(enc(token))
}
/**
* Redeems the token and returns a usable server configuration.
*
* The pin is applied to the very request that redeems the token, so a link pointing at an
* impostor fails at the TLS handshake rather than after handing it a token. That ordering is
* the reason the pin travels in the link at all.
*/
fun redeem(deviceName: String? = null, appVersion: String = ""): Enrolled {
// The link may name the server's public address rather than its control endpoint, so that
// a person is handed a name they recognise. Ask where to actually connect.
//
// Only the address comes from here. The pin still comes from the link, because a pin
// fetched over an ordinary TLS connection would be worth exactly what the certificate
// authorities are worth — and pinning exists to survive one the operator does not
// control, such as a root injected by corporate device management. An intercepted
// discovery can therefore send this device to the wrong host, where the pin will not
// match: an outage, not a compromise.
val endpoint = discover(controlUrl) ?: controlUrl
val client = ControlClient(endpoint, setOf(pin), appVersion)
val response = client.enroll(token, deviceName)
val profile = client.profile(response.credential)
return Enrolled(
controlUrl = endpoint,
publicUrl = controlUrl,
pin = pin,
credential = response.credential,
deviceId = response.deviceId,
profile = profile,
)
}
/**
* Asks a server where its control plane lives. Null when it does not say, or cannot be asked.
*
* Deliberately forgiving: a server that predates this, or one whose link already names the
* control endpoint directly, simply answers nothing and the link's own URL is used. Enrollment
* must not start failing because an optional lookup did.
*/
private fun discover(publicUrl: String): String? = runCatching {
val conn = (java.net.URL(publicUrl.trimEnd('/') + "/v1/discover").openConnection()
as java.net.HttpURLConnection).apply {
connectTimeout = 8_000
readTimeout = 8_000
setRequestProperty("Accept", "application/json")
}
if (conn.responseCode !in 200..299) return null
val body = conn.inputStream.bufferedReader().use { it.readText() }
kotlinx.serialization.json.Json { ignoreUnknownKeys = true }
.parseToJsonElement(body)
.let { (it as kotlinx.serialization.json.JsonObject)["control_url"] }
?.let { (it as kotlinx.serialization.json.JsonPrimitive).content }
?.takeIf { it.isNotBlank() }
}.getOrNull()
companion object {
const val SCHEME = "echolot"
const val HOST = "enroll"
private const val PIN_PREFIX = "pin-sha256:"
/**
* Parses a bootstrap link. Returns null for anything that is not one a malformed link
* must not be half-applied, because a half-configured server is a confusing failure much
* later rather than an obvious one now.
*/
fun parse(raw: String?): EnrollmentLink? {
val s = raw?.trim() ?: return null
val scheme = s.substringBefore("://", "")
if (!scheme.equals(SCHEME, ignoreCase = true)) return null
val rest = s.substringAfter("://")
val host = rest.substringBefore('?').trim('/')
if (!host.equals(HOST, ignoreCase = true)) return null
val params = HashMap<String, String>()
for (pair in rest.substringAfter('?', "").split('&')) {
if (pair.isEmpty()) continue
val k = pair.substringBefore('=')
val v = pair.substringAfter('=', "")
params[k] = dec(v)
}
// v is the link format, not the protocol. Unknown versions are refused rather than
// guessed at: the fields could mean anything.
val version = params["v"] ?: "1"
if (version != "1") return null
val url = params["u"]?.trim().orEmpty()
val pinRaw = params["p"]?.trim().orEmpty()
val token = params["t"]?.trim().orEmpty()
if (url.isEmpty() || pinRaw.isEmpty() || token.isEmpty()) return null
if (!url.startsWith("https://", ignoreCase = true)) return null
// A "+" in a query string decodes to a space, so a link whose base64 pin was pasted
// in unencoded arrives with spaces where "+" belonged — and a pin that is wrong by
// one character does not fail loudly, it just never matches, which surfaces much
// later as an inexplicable TLS error. Base64 has no spaces, so putting them back is
// unambiguous and cannot damage a correctly-encoded pin.
val pin = pinRaw.removePrefix(PIN_PREFIX).replace(' ', '+')
if (pin.isEmpty()) return null
return EnrollmentLink(controlUrl = url.trimEnd('/'), pin = pin, token = token)
}
private fun enc(s: String) = URLEncoder.encode(s, "UTF-8")
private fun dec(s: String) = runCatching { URLDecoder.decode(s, "UTF-8") }.getOrDefault(s)
}
}
/** A server this device is now enrolled with, ready to be stored in settings. */
data class Enrolled(
/** Where this device connects: the endpoint whose certificate the pin matches. */
val controlUrl: String,
/**
* The address a person was given, kept for display.
*
* Shown instead of [controlUrl] because the endpoint is plumbing it exists to select a
* certificate while this is the name the operator handed out and would recognise.
*/
val publicUrl: String,
val pin: String,
val credential: String,
val deviceId: String,
val profile: Profile,
)
@@ -13,14 +13,31 @@ import kotlinx.serialization.json.JsonElement
@Serializable @Serializable
data class EnrollResponse( data class EnrollResponse(
@SerialName("device_id") val deviceId: String, @SerialName("device_id") val deviceId: String,
val credential: String, /** The spec's name (§2.1). */
) @SerialName("device_credential") val deviceCredential: String? = null,
/** What the first server implementation shipped. Read for older servers; do not emit. */
@SerialName("credential") val legacyCredential: String? = null,
) {
/** Whichever field the server used. */
val credential: String
get() = deviceCredential ?: legacyCredential
?: error("enroll response carried no credential")
}
@Serializable @Serializable
data class Target( data class Target(
val id: String, val id: String,
val ip4: String? = null, val ip4: String? = null,
val ip6: String? = null, val ip6: String? = null,
/**
* The second address, which RFC 5780 behaviour discovery redirects to.
*
* Worth surfacing rather than treating as an implementation detail: a report that says "the
* server did not answer" means something different depending on which of its addresses was
* asked, and an operator reading one needs to be able to tell.
*/
@SerialName("ip4_alt") val ip4Alt: String? = null,
@SerialName("ip6_alt") val ip6Alt: String? = null,
@SerialName("udp_port") val udpPort: Int = 0, @SerialName("udp_port") val udpPort: Int = 0,
@SerialName("tcp_port") val tcpPort: Int = 0, @SerialName("tcp_port") val tcpPort: Int = 0,
@SerialName("stun_port") val stunPort: Int = 0, @SerialName("stun_port") val stunPort: Int = 0,
@@ -32,6 +49,59 @@ data class SelfTest(
@SerialName("sysctl_ok") val sysctlOk: Boolean? = null, @SerialName("sysctl_ok") val sysctlOk: Boolean? = null,
) )
/**
* The operator's upload rules, advertised in the profile so the app can present the choice
* honestly greyed out with a reason when the server refuses, and pre-set to the server's
* minimum anonymization when it accepts instead of discovering the policy by being rejected.
*/
@Serializable
data class UploadPolicy(
val mode: String = "off",
@SerialName("max_bytes") val maxBytes: Long = 0,
@SerialName("retention_days") val retentionDays: Int = 0,
@SerialName("max_runs_per_device") val maxRunsPerDevice: Int = 0,
@SerialName("min_anonymization") val minAnonymization: String = "full",
val reason: String? = null,
) {
val accepted: Boolean get() = mode == "anonymous" || mode == "account"
/** Why uploads are unavailable, in words a user can act on. */
fun refusalReason(): String? = when (mode) {
"off" -> reason ?: "This server does not accept uploaded runs."
"account" -> "This server only accepts uploads from signed-in accounts."
else -> null
}
}
/** The server's declaration of what it speaks and which app versions it will serve. */
@Serializable
data class CompatInfo(
@SerialName("protocol_version") val protocolVersion: String = "",
@SerialName("schema_version") val schemaVersion: String = "",
@SerialName("app_min") val appMin: String = "",
/** Exclusive; empty means the server sets no upper bound. */
@SerialName("app_max") val appMax: String = "",
)
/**
* How to sign in to this server's identity provider, advertised so the app can offer the button
* only when there is something behind it and drive the flow without anyone typing an issuer URL.
*/
@Serializable
data class AuthInfo(
val enabled: Boolean = false,
val issuer: String = "",
@SerialName("client_id") val clientId: String = "",
val flow: String = "",
@SerialName("redirect_uri") val redirectUri: String = "",
val scopes: String = "openid profile email",
@SerialName("authorization_endpoint") val authorizationEndpoint: String = "",
@SerialName("token_endpoint") val tokenEndpoint: String = "",
@SerialName("end_session_endpoint") val endSessionEndpoint: String = "",
/** Present when the server has an issuer configured but could not reach it. */
@SerialName("discovery_error") val discoveryError: String? = null,
)
@Serializable @Serializable
data class Profile( data class Profile(
@SerialName("profile_version") val profileVersion: Int = 0, @SerialName("profile_version") val profileVersion: Int = 0,
@@ -42,6 +112,9 @@ data class Profile(
@SerialName("canary_zone") val canaryZone: String = "", @SerialName("canary_zone") val canaryZone: String = "",
@SerialName("server_selftest") val serverSelftest: SelfTest? = null, @SerialName("server_selftest") val serverSelftest: SelfTest? = null,
val pins: List<String> = emptyList(), val pins: List<String> = emptyList(),
val uploads: UploadPolicy = UploadPolicy(),
val compat: CompatInfo = CompatInfo(),
val auth: AuthInfo = AuthInfo(),
) { ) {
fun supports(capability: String) = capability in capabilities fun supports(capability: String) = capability in capabilities
} }
@@ -0,0 +1,145 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.protocol
import kotlinx.serialization.json.Json
import kotlinx.serialization.json.jsonObject
import kotlinx.serialization.json.jsonPrimitive
import java.io.IOException
import java.net.HttpURLConnection
import java.net.URL
import java.net.URLEncoder
import java.security.MessageDigest
import java.security.SecureRandom
import java.util.Base64
/**
* Sign-in for the app: authorization code with PKCE (RFC 7636).
*
* The app is a *public* client it ships to devices, so any secret compiled into it can be read
* out of the APK with `unzip` and `strings`. PKCE is what replaces the client secret, and it
* defends a specific attack that matters here more than most places: the redirect comes back
* through a custom URI scheme, and on Android *any* app may register `echolot://`. A malicious one
* could intercept the callback and take the authorization code. Because the code can only be
* exchanged by presenting the verifier which never left this process and cannot be derived from
* the challenge that did a stolen code is worth nothing.
*
* Nothing from the IdP is kept afterwards. The ID token is used once, to prove to the server who
* is signing in, and then discarded: the device credential is what authenticates every later
* request. So there are no access tokens to store, no refresh tokens to rotate, and no token
* lifetime for the app to manage.
*/
object OidcLogin {
/** A started sign-in. [verifier] and [state] must survive until the callback returns. */
data class Pending(val authorizationUrl: String, val verifier: String, val state: String)
/**
* Builds the authorization URL and the secrets that must be held until the callback.
*
* Everything comes from the server's profile rather than being compiled in, so pointing the
* app at a different server with a different IdP is configuration, not a rebuild.
*/
fun begin(auth: AuthInfo, random: SecureRandom = SecureRandom()): Pending {
require(auth.enabled && auth.authorizationEndpoint.isNotBlank()) {
"this server has no identity provider configured"
}
val verifier = randomUrlSafe(random)
val state = randomUrlSafe(random)
val challenge = b64(MessageDigest.getInstance("SHA-256").digest(verifier.toByteArray()))
val q = buildString {
append("response_type=code")
append("&client_id=").append(enc(auth.clientId))
append("&redirect_uri=").append(enc(auth.redirectUri))
append("&scope=").append(enc(auth.scopes))
append("&state=").append(enc(state))
append("&code_challenge=").append(enc(challenge))
append("&code_challenge_method=S256")
}
val sep = if (auth.authorizationEndpoint.contains('?')) "&" else "?"
return Pending(auth.authorizationEndpoint + sep + q, verifier, state)
}
/** What came back on the `echolot://auth` redirect. */
data class Callback(val code: String?, val state: String?, val error: String?)
/** Parses the redirect URI the browser handed back to the app. */
fun parseCallback(uri: String): Callback {
val q = uri.substringAfter('?', "")
var code: String? = null
var state: String? = null
var error: String? = null
for (pair in q.split('&')) {
val k = pair.substringBefore('=')
val v = dec(pair.substringAfter('=', ""))
when (k) {
"code" -> code = v
"state" -> state = v
"error" -> error = v
"error_description" -> if (error != null) error = "$error: $v"
}
}
return Callback(code, state, error)
}
/** The sign-in failed in a way worth showing someone, rather than a transport error. */
class LoginFailed(message: String) : Exception(message)
/**
* Exchanges the code for an ID token.
*
* The state is compared before anything else happens. A callback whose state does not match
* the one this process generated did not come from a flow this process started which is
* precisely how an attacker gets a victim to complete *their* login so it is refused before
* the code is spent.
*/
fun complete(auth: AuthInfo, pending: Pending, callbackUri: String): String {
val cb = parseCallback(callbackUri)
if (cb.error != null) throw LoginFailed(cb.error)
if (cb.state.isNullOrEmpty() || cb.state != pending.state) {
throw LoginFailed("this sign-in did not start on this device — start again")
}
val code = cb.code ?: throw LoginFailed("the identity provider returned no authorization code")
val body = buildString {
append("grant_type=authorization_code")
append("&code=").append(enc(code))
append("&redirect_uri=").append(enc(auth.redirectUri))
append("&client_id=").append(enc(auth.clientId))
append("&code_verifier=").append(enc(pending.verifier))
}
val conn = (URL(auth.tokenEndpoint).openConnection() as HttpURLConnection).apply {
requestMethod = "POST"
doOutput = true
connectTimeout = 15_000
readTimeout = 15_000
setRequestProperty("Content-Type", "application/x-www-form-urlencoded")
setRequestProperty("Accept", "application/json")
}
conn.outputStream.use { it.write(body.toByteArray()) }
val text = try {
val stream = if (conn.responseCode in 200..299) conn.inputStream else conn.errorStream
stream?.bufferedReader()?.use { it.readText() } ?: ""
} catch (e: IOException) {
throw LoginFailed("could not reach the identity provider: ${e.message}")
}
if (conn.responseCode !in 200..299) {
throw LoginFailed("the identity provider refused the sign-in (${conn.responseCode})")
}
val idToken = runCatching {
Json.parseToJsonElement(text).jsonObject["id_token"]?.jsonPrimitive?.content
}.getOrNull()
return idToken?.takeIf { it.isNotBlank() }
?: throw LoginFailed("the identity provider returned no id_token")
}
private fun randomUrlSafe(random: SecureRandom): String =
ByteArray(32).also(random::nextBytes).let(::b64)
private fun b64(b: ByteArray): String = Base64.getUrlEncoder().withoutPadding().encodeToString(b)
private fun enc(s: String): String = URLEncoder.encode(s, "UTF-8")
private fun dec(s: String): String =
runCatching { java.net.URLDecoder.decode(s, "UTF-8") }.getOrDefault(s)
}

Some files were not shown because too many files have changed in this diff Show More