Compare commits

..
Author SHA1 Message Date
mrambossekandClaude Opus 5 9ee7d554a6 app: listen to what the segment says unprompted
server-release / image (push) Successful in 14s
server-release / release (push) Successful in 31s
Four passive collectors join long mode: SSDP (passive NOTIFY plus paced
M-SEARCH from the capture socket, so unicast replies land in the same
capture), LLMNR, NetBIOS-NS and WS-Discovery. All are periodic and sparse
by nature, which is exactly why they belong to a window rather than a
probe - a thirty-second run mostly hears silence and would report an
empty network as confidently as a quiet one.

NetBIOS reports unsupported on the app tier and says why: UDP 137 is
privileged. The decoder and evidence shape are tested and waiting for the
Shizuku tier; a recorded reason beats a missing test.

Silence without a multicast lock or a group join is PARTIAL, never OK -
that case is a fact about this app, not about the network. Hostnames,
banners and device UUIDs are classified in core-privacy so the anonymizer
treats them like every other identifier rather than letting a neighbour's
device model ride out in an upload.

No active WSD Probe and no NBSTAT sweep: the app listens to what a
network broadcasts, it does not announce itself to strangers or
interrogate its neighbours.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 14:58:41 +02:00
mrambossekandClaude Opus 5 93125a4b1d app: 0.3.0 — run modes and the dev relay
server-test / test (push) Successful in 37s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 14:46:25 +02:00
mrambossekandClaude Opus 5 0071e00003 app: long runs that watch, and a relay for the port that keeps moving
Long mode starts listeners at t=0 and keeps them running past the
battery: a network-change watcher that finally fills networks[].changes[]
(defined since the schema's first draft, never populated), an RSSI log, a
ping series giving loss and jitter over minutes, and mDNS listening for
the whole window. This is the class of fault a short run cannot see - a
link that drops for four seconds between two probes is reported healthy
by both of them. run.mode records which question was asked, because
silence means different things in the two modes.

The adb relay replaces the retired beacon: AdbRelay watches adbd's own
mDNS with the resolve-once discipline the beacon learned the hard way
(resolving re-arms adbd and pops a notification), a foreground service
keeps it alive with the screen off, and the heartbeat re-posts the cached
endpoint rather than re-resolving. It exists because mDNS does not cross
subnets and the wireless-debug port rotates every few minutes.

Also records why LLDP/CDP cannot follow SSDP into long mode: both are raw
L2 frames, so they need CAP_NET_RAW - root tier, not app, and Shizuku's
shell user does not have it either.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 14:43:33 +02:00
mrambossekandClaude Opus 5 ae63bd7c7f server: relay a test device's adb endpoint, because mDNS does not cross subnets
The beacon this replaces was a separate service wildcard-bound to
0.0.0.0:443 - it silently occupied port 443 on the reserved measurement
addresses, voiding the IPv4 interception proof for as long as it ran, and
it accepted a port report from anyone who could reach it. So this lives
where the repo's own post-mortem said it belongs: POST on the control
plane authenticated by the device credential, GET on the admin UI behind
the existing apiAdmin helper. No new listener, no new port, no wildcard.

Entries expire after 24h (ECHOLOT_ADB_ENDPOINT_RETENTION_H) on both write
and read - a LAN address is a breadcrumb for driving a test device, not
measurement data worth keeping.

Also records the BLE peer-comparison design: the case for it is that BLE
is out-of-band, which is what makes client isolation measurable at all -
silence over IP cannot distinguish an isolating AP from an absent peer,
and a peer confirming out-of-band that it was listening turns that
silence into proof.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 14:31:01 +02:00
mrambossekandClaude Opus 5 ab6e278272 app: an IMS network is not a blocked measurement
rmnet_data1 on the OnePlus is the carrier's IMS/VoLTE network: it carries
IMS&MMTEL but neither INTERNET nor NOT_RESTRICTED, so binding needs
CONNECTIVITY_USE_RESTRICTED_NETWORKS - signature-level, unobtainable.
EPERM there is permanent and says nothing about any network's health, but
the constraint detector counted it, so a phone with VoLTE reported
'measurement blocked' on every run and every verdict came out
INCONCLUSIVE. Capabilities now decide: networks an app may never bind are
recorded as app_usable:false in networks[] and skipped, rather than
reported as something the run failed to do.

Also retires the 'while it tears down' wording, which was a guess that
turned out to be wrong - the refusal outlasts the VPN indefinitely, so
the text now names the VPN as the usual cause without asserting a
timeline it cannot know.

App version 0.2.4.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 14:11:08 +02:00
mrambossekandClaude Opus 5 d420a71a01 docs: v0.11.3 live on fmr, trains validated 120/120; two mint bugs logged
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 14:04:46 +02:00
mrambossekandClaude Opus 5 d621611bdd docs: the v0.11.2 mystery was an unstamped-tag problem, not lost code
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 13:54:42 +02:00
mrambossekandClaude Opus 5 c5f3c2e7af app: learn probe durations per device; stop claiming VPN after teardown
server-release / image (push) Successful in 15s
server-release / release (push) Successful in 31s
The ETA table was fixed at compile time, but real durations are a property
of this phone and the network it stands in - ICMPv6 answers in
milliseconds where IPv6 works and waits out its timeout where it does
not. Estimates now prefer an EMA (70/30) of what this device actually
measured, seeded by the old constants on first run.

VPN detection had two lies in it, both found on hardware: a disconnected
tunnel lingers in allNetworks while tearing down, so 'VPN active' is now
judged from the ACTIVE network only; and any bind failure counted as
'per-network blocked', so a network dying mid-run flipped a healthy run
to INCONCLUSIVE - now only a genuine EPERM refusal counts. Banner and
finding split three ways (VPN + blocked / blocked only / VPN only) so the
app never names a VPN the user just turned off.

App version 0.2.3.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 13:41:10 +02:00
mrambossekandClaude Opus 5 3214cc877a app: parse the Shizuku dumps into link.ra_source and sec.arp_watch
Pure-Kotlin parsers over the captures the battery already makes, tested
against the real archived dumps from both devices - including the
Lenovo's 1000000015 route tables (table ids stay strings), the OnePlus's
stray uid=2000 prefix and mid-line hoplimit, and an ip_monitor that only
ever said EXEC_TIMEOUT. ShizukuProbe now returns three tests; a failed
capture is SKIPPED with its reason so no test type silently vanishes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 13:34:52 +02:00
mrambossekandClaude Opus 5 515a6aef04 app: speak train (0x03-0x05) and mark downtrains with DSCP
ProbeSession sends paced upstream trains and fetches the server's columnar
received view; UpstreamTrainMeasurement lines both ledgers up per sequence
number into TrainEvidence - the directional loss attribution a round trip
cannot make. Loss is computed against the server's total count, not its
row list, so a buffer-capped report can never invent loss on big trains.
A missing report stays ambiguous by name (report lost, or server predates
trains) instead of being blamed on the train.

downtrain actions can now request a DSCP marking, recording dscp_applied
so a survival comparison never blames the path for a marking the sender
skipped.

Also logged: fmr runs v0.11.2 built from commits this repo's remote never
saw, so --self-update on fmr is OFF-LIMITS until that lineage is repaired
- the fresh server-v0.9.2 release is newest-created but semantically
older, and the updater compares strings, not SemVer.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 13:34:52 +02:00
mrambossekandClaude Opus 5 d574c76630 docs: log the prober fold and server v0.9.2 feature work
server-release / image (push) Successful in 16s
server-test / test (push) Successful in 48s
server-release / release (push) Successful in 1m2s
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 13:06:09 +02:00
mrambossekandClaude Opus 5 8118e213ae server: upstream trains, observed TTL/DSCP/ECN, rate limits, action ids
Types 0x03/0x04/0x05 land with a bounded columnar train buffer (head kept,
truncation declared) and grant-free multi-part reports - a report row is
smaller than the packet it answers, so $3.4 holds without a grant. The
read loop now collects TTL/TOS cmsgs on Linux, replacing the 0xFF stubs in
the observation block with what the kernel saw; downtrain gained a dscp
parameter, so DSCP survival is measurable in both directions.

Rate limiting ($2.5) exists now: per-credential AND per-source buckets,
429 on the control plane, silent drop on the data plane after the HMAC
gate and before the replay window. UDP ceilings default above the largest
legitimate run - a limit that clips a real measurement produces a
confidently wrong number.

Every granted packet carries its action_id at payload[8:16]; overlapping
actions were unattributable before. Canary DNS logs now honor the stated
24h privacy default. /admin/enroll-tokens answers the spec's JSON shape.
protocol_version 1.0.1 (additive).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 13:04:54 +02:00
mrambossekandClaude Opus 5 f6849f8e6a app: fold the prober's validated capabilities into core-probe
traceroute.udp4 lands as TracerouteProbe: ICMP time-exceeded read off the
socket error queue via Os.recvmsg(MSG_ERRQUEUE) through the reflection
facade the prober validated on both devices - no root, no raw socket, and
the C-over-JNI shim stays dead. Emits the schema's TracerouteEvidence with
rtt_ns. OsAbi carries the hardcoded sockopt ABI numbers across, including
the measured fact that Os.getsockoptInt exists on neither device, so PMTU
must come from the errqueue, never getsockopt(IP_MTU).

local.mdns_inventory lands as MdnsInventoryProbe with both hardware-bought
lessons intact: the meta-query lies (0 results beside live services on both
devices), and 4s of listening misses what 10s catches.

App version 0.2.2.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 12:39:52 +02:00
mrambossekandClaude Opus 5 7d98a43866 docs: close the VPN, v6, reserved-probe and signing items in the build log
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 12:32:02 +02:00
mrambossekandClaude Opus 5 a49bef5821 server: refuse unsigned releases and polluted reserved addresses
Self-update now verifies SHA256SUMS.sig (ed25519, relsign package) against
a public key baked into the binary; the private key exists only in the CI
secret store, so a compromised release host can withhold updates but not
inject one. CI signs on every server-v* tag and hard-fails without the
secret. Operators with their own pipeline override the key via
ECHOLOT_SELF_UPDATE_PUBKEY (mint a pair with release-sign -gen).

Startup also now proves 80/443 are actually free on the reserved
measurement addresses by asking the OS (throwaway bind), not the config -
CheckReserved could never see a stray process, and the adb-beacon receiver
on 0.0.0.0:443 was exactly that.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 12:32:02 +02:00
mrambossekandClaude Opus 5 20cfecf566 app: name what a VPN blocked, and prove v6 broken before saying so
Constraints are detected up front (one throwaway bind per network) and land
in run.constraints, a measurement.vpn_constrained finding, the $7.3 verdict
(INCONCLUSIVE outright) and a banner on the run screen - a VPN'd run looked
exactly like a clean run of a healthy network before this.

v6.broken returns to the registry now that it can be earned: V6ConnectProbe
(v6.brokenness) makes a real TCP connection over IPv6 to the enrolled
server, and only both transports failing on a network that advertises IPv6
justifies the claim. TCP succeeding turns the finding into 'ICMPv6 is
filtered, IPv6 works' at high confidence instead.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 12:31:44 +02:00
mrambossekandClaude Opus 5 987b2ceb47 server: read the verdict from where the schema puts it
Every uploaded run showed "not recorded" in the web UI because the meta
extractor read summary.verdict. Schema §7.3 calls that field
summary.overall; "verdict" is the per-category field one level down. So
the verdict was never stored, and the UI faithfully reported a gap that
was this parser's doing rather than the document's.

The test encoded the same mistake — its fixture posted
summary.verdict:"warn" — so it passed throughout against a parser that
read a field nothing writes. Corrected to summary.overall, and to a
verdict that exists: §7.3 defines green|yellow|red|inconclusive, and
"warn" was never one of them.

The eleven runs already stored had their meta backfilled from the
documents, which are kept byte-for-byte and still carry the real value.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 10:52:25 +02:00
mrambossekandClaude Opus 5 d5b1bab577 app: stop the autorun upload pointing at a listener that is gone
"upload failed: HTTP 400 client sent an HTTP request to an HTTPS server"
is a TLS listener rejecting cleartext, and the cleartext was ours:
REPORT_UPLOAD_URL was http://89.185.109.150:443/report, which used to be
the adb-beacon receiver holding 0.0.0.0:443 in plaintext. Disabling that
receiver and giving 443 to echolot-server left this posting plain HTTP at
a TLS port.

Blanked rather than repointed. The endpoint existed so an unattended run
could be collected without adb, and autorun reports are now read straight
off the device with `run-as cat` — so it was buying nothing and emitting
an alarming error for a debugging convenience. Deliberately not aimed at
/v1/runs either: that is the consent-gated upload, and a debugging
shortcut must not be able to satisfy it by accident.

The message says what happened instead of implying something broke.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 10:36:49 +02:00
mrambossekandClaude Opus 5 a17c3fd9e6 app: catch a search domain that swallows DNS queries
A tablet on a healthy network could not resolve anything. The DNS server
answered the bare name correctly — NOERROR, two records, A and AAAA, with
and without EDNS0 — so the earlier finding blamed the device's resolver.
It was wrong. The network advertised hudelist.local as a search domain and
the server silently dropped every query under it: not NXDOMAIN, nothing at
all. Resolvers append search domains, so they waited for a reply that was
never coming.

Silence is the part that makes this vicious. A negative answer moves a
resolver on; no answer looks like packet loss, so it retries, and some
give up on the lookup entirely. It also explains how two devices on one
network can disagree about whether DNS works — the phone tried the plain
name first and never noticed.

The probe now asks about a nonce name under each advertised search domain,
where the wanted answer is NXDOMAIN and only silence is a fault. The
finding is ordered ahead of dns.system_resolver_broken so the two cannot
both fire: without that, this exact network gets told its device is
broken.

Severity follows the harm rather than the shape. HIGH when resolution is
actually failing, MEDIUM when the domain is a black hole but this resolver
happens to try the plain name first — calling that HIGH would be crying
wolf on a network that works. The message names the fix and notes that
.local is reserved for mDNS by RFC 6762 and widely dropped by design,
while home.arpa (RFC 8375) is the name reserved for this.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 10:29:04 +02:00
mrambossekandClaude Opus 5 cfa58e8d60 app: a DNS reply is not a DNS answer
The probe counted any well-formed packet from the server as "the server
answers", checking only the transaction id and a minimum length. REFUSED
and SERVFAIL are well-formed packets. So a server actively refusing this
client would have been reported as healthy, and the finding — whose whole
output is "the network is fine, your device is not" — would have pointed
confidently at the wrong component.

It now requires rcode 0 and at least one record, and reports a refusal as
what it is: a working server saying no, which points back at the network.
The rcode is named rather than numbered, because "REFUSED" is a fact an
operator can act on and "rcode 5" is a lookup.

Caught by decoding what fmr's router actually replied — ab cd 81 80 00 01
00 02, NOERROR with two answers — after realising the earlier check only
counted bytes. The reply was genuinely good, so the finding on the tablet
stands; the check was wrong regardless.

The advice is broader too. That tablet's fault survived a reboot, which
makes "toggle wifi and it clears" wrong as a flat claim: it now says what
to look at when a restart does not fix it — something on the device
filtering DNS, or a per-client rule on the router.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 10:10:36 +02:00
mrambossekandClaude Opus 5 2ed4d1f478 Reach the server by address when its name will not resolve
A measurement tool that cannot report from a broken network is useless
exactly when it matters, and a wedged resolver is one of the faults this
app is built to find — it should not also be the thing that stops the
finding being delivered. The profile already carries the server's
addresses; they are now kept and used when the name fails.

Safe because the pin is the trust and the name is not part of it: the
server presents the same certificate however it was reached, and a wrong
address fails the pin like anything else. Only the primaries are cached —
the alternate pair exists for NAT behaviour discovery and does not carry
the control plane, so falling back to one would fail for a second,
unrelated reason.

Substituted only when the name genuinely does not resolve, and only after
checking the candidate answers on the port: on a v4-only network a v6
address would otherwise be chosen and fail slowly, which is the wrong
answer delivered late.

The server had to meet it halfway. Sharing 443 by SNI meant a client
arriving by IP sent no server name and got the Let's Encrypt certificate,
failing the pin. A numeric host — or no SNI at all — now selects the
pinned certificate and routes to the control plane. That is sound because
the admin UI is only ever reached by name: browsers always send SNI, and
nobody bookmarks an IP for a site with a CA-issued certificate.

Verified against fmr: by IP on both families the served pin is the
control one and /v1/profile answers 401, while fmr.echo-lot.app still
serves the Let's Encrypt certificate and the admin UI.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 10:03:14 +02:00
mrambossekandClaude Opus 5 d65dbbc75a app: tell a broken device resolver apart from a broken network
Diagnosing a tablet that claimed "no internet" took twenty adb commands
to establish something the app should have said in one run: ping to
1.1.1.1 worked, the configured DNS server answered a raw UDP query in
65 bytes, and Android still could not resolve a hostname. The network was
fine; netd had wedged.

Those two failures look identical to a user and want opposite responses —
"look at your router" against "toggle your wifi" — so dns.resolver asks
the network's own servers directly and compares the answer against what
the platform returns for the same name. The query is hand-rolled over a
plain DatagramSocket on purpose: anything routed through a resolver API
would inherit the very fault being looked for.

dns.system_resolver_broken fires only on the pairing that is otherwise
unattributable: server answered, platform did not. Per network, because a
phone can have wedged wifi and working cellular at once.

Also records Android's own verdict per network — validated, captive
portal, partial connectivity — which the app reproduced with its own HTTP
probes but never stored. It is free, it is what the user sees in the
status bar, and its disagreement with our measurements is exactly what
identified the tablet. NET_CAPABILITY_PARTIAL_CONNECTIVITY is @SystemApi
so the constant is inlined with its rationale, in the manner of OsAbi.kt,
and read defensively enough to report unknown rather than false.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 09:44:49 +02:00
mrambossekandClaude Opus 5 40e76c52ca adminui: show the enrolment link as a QR code
Enrolling a device that cannot reach the admin UI meant transcribing a
200-character link with a base64 pin in it — the step the link format
exists to avoid, and the one where a pin wrong by one character fails
later as an inscrutable TLS error.

Rendered as inline SVG rather than a PNG data: URI, because the page's
CSP is default-src 'none' and means it: a data: image would need img-src
opened, markup needs nothing. One path rather than a rect per module,
since a link this long encodes to about 60x60 and two thousand elements
is a lot of DOM for a picture of a square. It is generated from the same
validated value as the href, so a rejected link produces neither.

This relaxes the stdlib-only rule, deliberately and recorded in CLAUDE.md.
The rule bought one self-contained binary with no supply chain to audit,
which one small pure-Go package barely dents; F-Droid never applied to
the server, only the app ships there. A correct QR encoder is ~500 lines
of Reed-Solomon that nobody should be hand-writing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 08:48:02 +02:00
mrambossekandClaude Opus 5 5b02d40802 app: notice when Shizuku is authorised
The banner watched only for the binder arriving or dying, and granting
permission does neither. So after the user tapped the banner, answered
the dialog and came back, it still read "running but not authorised" —
wrong at precisely the moment they were looking for confirmation that it
had worked.

Two listeners were missing, because there are two ways this changes.
Answering our own request now fires OnRequestPermissionResultListener.
That is not enough on its own: permission can equally be granted inside
Shizuku's own app, and Shizuku started or stopped there, none of which
calls back into this process — so the state is re-read whenever the
screen comes forward, which is the only thing that covers every route.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 08:37:45 +02:00
mrambossekandClaude Opus 5 416b783647 app: lay the server facts out in real columns
The block was rendered as one padded string, which only lines up in a
monospaced font — and the monospace never took, so the values sat at
ragged offsets and the point of the list was lost. Padding text to fake a
table makes the layout depend on a typeface decision made elsewhere.

It is a label column of fixed width and a value column that takes the
rest, so the addresses align whatever the font does and a long value wraps
inside its own column instead of under the labels. The capability list
goes back to plain comma-separated text: hand-wrapping it at three per
line was working around the same missing alignment.

Placement too. The facts now sit under the "Check server" button that
fetches them, rather than among the input fields, and the sentence
explaining the endpoint sits directly under the URL field it describes —
it had ended up orphaned between the two, reading as a comment on nothing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 08:24:00 +02:00
mrambossekandClaude Opus 5 c4f2a10790 app: show what the server reports as facts, not as inputs
The settings card offered three editable boxes and said nothing about the
server itself — which addresses a test will actually use, on which ports,
what it can measure. That is the part a person checks before trusting a
result, and "which address did this come from" is precisely the question
a report leaves open.

The server now publishes it. The profile's targets carried one IPv4 and a
TODO; it reports both families and both alternates, derived from the UDP
listen spec rather than configured separately, so the list cannot drift
from what is actually bound. No reservation means no alternate is
claimed: announcing a second address as the RFC 5780 alternate when none
was set aside would promise a redirect the server will not send.

The app renders them read-only, in a panel visibly distinct from the
fields above. An editable box that changes nothing is worse than no box,
and these are facts to read rather than settings to apply.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 08:14:42 +02:00
mrambossekandClaude Opus 5 fe4ec23ba1 app: show the server by the name its operator handed out
The Server URL field showed the endpoint the app dials, which after
discovery is not the name anyone was given — enrolling against
fmr.echo-lot.app left the settings reading fmr-1.echo-lot.app, with no
explanation of where the -1 came from.

It shows the public name now. The endpoint is not hidden, just demoted to
a line beneath that says where the connection actually goes and why the
two differ: a network engineer debugging a failed connection wants that,
and burying it would trade one confusion for another.

Typing a URL by hand sets both, since there is no discovery to consult in
that case — setting only the public one would leave the app still dialling
the previous server, which is the kind of half-applied change that fails
much later and somewhere else.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 08:07:55 +02:00
mrambossekandClaude Opus 5 33799b8135 adminui: the enrolment link was rendered as a dead anchor
html/template rewrites an href whose scheme it does not recognise to
"#ZgotmplZ". echolot:// is not on its list, so "Open in the Echolot app"
was not a link at all — tapping it did nothing, and nothing showed it:
the markup reads correctly, the app resolves the scheme, and only the
sanitised attribute in the served HTML gives it away.

Marking the value template.URL opts out of that sanitising, which is only
safe because the shape is now checked first. The link arrives in a query
parameter, so without the check a crafted /devices?link=javascript:… would
put a script URL into the page for an admin to click.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 08:04:56 +02:00
mrambossekandClaude Opus 5 f6e093944c build: document the poisoned build cache; compare like with like
An entire module was missing from the APK. The app died with
ClassNotFoundException for app.echo_lot.protocol.EnrollmentLink while the
build was green, the module's jar was correct, and :app:dependencies
listed it on debugRuntimeClasspath — its code simply never reached AGP's
intermediates.

The cause was a poisoned Gradle build cache entry, which is why nothing
obvious fixed it: clean, rm -rf */build and --rerun-tasks all leave the
build cache alone. Only --no-build-cache did. Every app build made in this
session shipped without core-protocol, so enrolment, sign-in and upload
would all have crashed identically; several hours of "the tap does
nothing" were this, misread as a UI problem.

CLAUDE.md now carries the symptom, the fix, and the verification —
grepping the dex for a string literal only that module defines, because
grepping for a class *name* proves nothing: callers carry the name as a
reference whether or not the class is packaged. That false check is what
let me believe an earlier rebuild had fixed it.

Also: the enrolment dialog compared the stored endpoint against the link's
public URL, so re-enrolling with the same server announced itself as a
move to a different one. Those are deliberately different strings now that
discovery exists; the comparison uses the public name on both sides.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 08:01:58 +02:00
mrambossekandClaude Opus 5 bc384531e6 app: enrollment feedback beside its own button
The enrollment result was written to the same status line as everything
else, which renders at the far end of the server card below three text
fields — and on a fresh install renders nowhere at all, because that line
only appears once a run exists. So enrolling looked identical whether it
worked or not.

It has its own line now, directly under the Enroll button that caused it,
and its own state rather than sharing one with "Check server": two
actions, two results.

The server fields were also re-read the instant the button was pressed,
before the enrollment coroutine had done anything, so they showed the
previous server's values. They now refresh when the result lands, which
is the point at which there is something new to show.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 00:48:36 +02:00
mrambossekandClaude Opus 5 082a2314ef protocol: enrollment links carry the public name, not the endpoint
Sharing port 443 between the admin UI and the control plane forces two
hostnames — one port and one name is one certificate, and the two need
different ones. That difference had been leaking into every enrollment
link, so an operator handed out fmr-1.echo-lot.app when the thing they
and their users know is fmr.echo-lot.app.

The link now carries the public name and the app asks GET /v1/discover
where to actually connect. The endpoint is plumbing: it exists to select
a certificate, and nobody needs to see it.

Discovery hands out an address and never a pin. The pin stays in the
link. Fetching it over an ordinary TLS connection would make pinning
worth exactly what the certificate authorities are worth, and pinning is
there to survive one the operator does not control — a root injected by
corporate device management, say, which is unremarkable on the networks
this tool gets pointed at. With the pin pre-shared, an intercepted
discovery can only send a device somewhere the pin will not match: an
outage, not a compromise.

Optional on both sides. A server that does not answer, or a link that
already names the control endpoint, works unchanged — enrollment must not
start failing because a lookup did.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 00:13:27 +02:00
mrambossekandClaude Opus 5 a720e84411 app: one activity for deep links, and say what re-enrolling actually does
MainActivity had no launchMode, so every echolot:// link stacked a fresh
activity with its own ViewModel. The enrolment then ran in a throwaway
copy and pressing back returned to the original screen showing none of
it — silent, and indistinguishable from the link not working at all.
singleTask plus onNewIntent means the link reaches the screen already in
front of the user.

The confirmation dialog also read as nonsense when re-enrolling with the
server already configured: "already enrolled with X ... enrolling with X
replaces that". Naming one URL twice looks like a bug and buries the
consequence that does apply — the credential is replaced, the old one
stops working at once, and the device appears on the server as a second
entry next to the first, which is worth revoking afterwards.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 23:54:08 +02:00
mrambossekandClaude Opus 5 30d9501507 app: confirm before an enrollment link replaces an existing one
The scheme was already registered and the deep link already worked — it
enrolled on arrival, with no confirmation. Now that the web UI offers the
link as something to follow, that is one tap between a working enrollment
and a replaced one, from a page that might be showing a link minted for a
different device entirely.

Enrolling is not additive: the new credential replaces the old, and on
the previous server this device simply stops reporting. So the link is
held and the user is asked, with both server URLs named — the question is
"which server", and it cannot be answered without seeing both.

The dialog says what survives, because that is the part someone hesitates
over: uploads already on the old server stay there, runs stored on the
phone are untouched, and the device reappears on the new server as a new
device rather than carrying its history across.

Enrolling also clears the cached canary zone. It describes the old
server's deployment, and querying it against the new one would measure
somebody else's zone and file the answer under this network.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 23:42:37 +02:00
mrambossekandClaude Opus 5 c639d58043 adminui: make the enrolment link followable on the phone being enrolled
The link was printed as text next to an adb command, so enrolling a phone
meant getting a 200-character string from one device onto another —
which is the step that goes wrong, and it does not have to happen at all.
The app registers the echolot:// scheme, so on the phone being enrolled
following the link is the entire procedure.

The raw string and the adb form stay for every other case: reading it on
a laptop, or enrolling a device that is not the one holding the browser.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 23:39:55 +02:00
mrambossekandClaude Opus 5 6d042a9d89 server: control plane shares port 443 with the admin UI
Captive portals, hotel wifi and corporate firewalls routinely permit only
80 and 443 — exactly the networks this tool exists to diagnose. A control
plane on 8443 is unreachable precisely when it matters most, and it fails
as "cannot reach server", which tells the user nothing about why.

The two cannot share a certificate, so sharing the port needs two names.
The control plane is trusted by SPKI pin and uses a long-lived
self-signed certificate; a browser needs one a CA vouches for. One name
on one port is one certificate. Pinning the Let's Encrypt key instead was
considered and rejected: it survives renewal only while key reuse holds,
so a routine key rotation would brick the fleet.

One listener now picks the certificate by SNI and the handler by Host.
Both have to agree, or a client gets the pinned certificate with the
admin UI behind it.

8443 stays open. Devices enrolled before this carry that URL in their
settings, and closing it for the sake of a port number would strand every
one of them; it can go once nothing points at it.

Verified per SNI on 443: fmr.echo-lot.app serves the Let's Encrypt cert,
fmr-1.echo-lot.app serves the self-signed one whose pin is unchanged, and
/v1/profile answers 401 on the control name against 303 to the login page
on the UI name.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 23:25:23 +02:00
mrambossekandClaude Opus 5 4aaaa5f5d4 app: probe the server this device is enrolled with, not ours
The canary-DNS zone and the STUN host were compiled in as c.echo-lot.app
and fmr-1.echo-lot.app, so every copy of the app measured against this
particular deployment whatever server its owner had enrolled with. On
someone else's install those two tests describe our infrastructure and
report the result as a fact about their network.

The zone comes from the server's own profile, which has advertised
canary_zone all along — the app simply never read it. It is cached in
settings because the canary probe runs at device tier, before anything
has contacted the control plane, and a probe that had to make a call
first would fail on exactly the networks worth measuring. The STUN host
is derived from the configured server URL rather than stored, since a
second copy of the server's name goes stale the moment someone
re-enrolls elsewhere.

With no server configured both now report SKIPPED. StunProbe previously
would have reported FAILED on a blank host, which reads as a finding
about the network when the truth is that no packet was ever sent — the
same conflation between "measured nothing" and "measured a fault" that
the ICMPv6 finding had.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 23:20:52 +02:00
mrambossekandClaude Opus 5 e0428b4c84 server: enroll against fmr.echo-lot.app; mint links from the CLI
The control plane now advertises the same name the web UI answers on.
That is safe because the client authenticates by SPKI pin and explicitly
does not verify the hostname — "pin is the trust, not the name" — so no
certificate covers or needs to cover either name.

The per-host name still means something, though, and the rule it encodes
has to survive: pinning binds a client to one server's key, so fmr may be
a CNAME to exactly one host and never a multi-address service record. A
second server gets enrolled as fmr-2 explicitly, because a client that
reaches a different key does not fail over, it fails.

Minting a link was broken and had been since the authenticated admin UI
replaced the old admin API: enroll-link.sh still posted to
127.0.0.1:8444/admin/enroll-tokens, an endpoint that no longer exists on
a listener that no longer binds loopback. Rather than add a second
unauthenticated door — which is how the old one ended up briefly reachable
from the network — the binary mints its own link. Whoever can run it
against the state directory already holds every privilege the server has,
so authenticating them to themselves would be theatre.

EnrollmentURI is shared with the running server's EnrollmentLink rather
than reimplemented. Two copies of that encoding would eventually disagree,
and the failure mode is a pin that looks right and surfaces as an
inscrutable TLS error rather than as a bad pin.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 23:16:23 +02:00
mrambossekandClaude Opus 5 8001234e8b adminui: render the self-test instead of dumping it
The box under "Self-test" was `%+v` of a Go struct on one line, read
through a horizontal scrollbar — on a phone you could see about six words
of it, from the middle. The design pass had polished the frame around it
and left the contents a debug dump.

The report was structured the whole time: each sysctl check carries the
name, what was found, what was wanted, a severity, and a sentence
explaining why the setting matters to measurement. All of that was being
flattened into one string. It now renders as records like everything
else, with the explanation set as prose across the full row, because it
is a sentence and not a fourth column.

Two faults the render caught: the desktop row grid applied to every
readout, so the standalone summary panel was chopped into four narrow
columns and "full 1500" broke into "ful/l/150/0"; and a fixed first
column wrapped `net.ipv6.conf.all.accept_ra` mid-word. The grid is now
scoped to readouts inside a row, and the label column may grow to 18rem
before it wraps.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 23:04:55 +02:00
mrambossekandClaude Opus 5 332b6f3589 adminui: give the admin UI the instrument's own visual language
Echolot is the German word for an echo sounder — an instrument that emits
a ping and reads what comes back — and the UI now looks like one instead
of like a dashboard. The three big-number stat cards went first: that
layout is the stock answer for any admin page, and it told a network
engineer nothing they could act on.

Palette is a water column rather than a neutral near-black, with a
desaturated sea-green return for the accent. Verdict colours come from
the domain, so they carry meaning rather than decorate. No web fonts —
the CSP forbids loading anything and shipping font files would trade
what makes this a single pleasant binary for a typeface — so the
character comes from treatment: machine-set headings, tracked small
caps, hairlines.

The one ornament is a trace of returns across time on the runs page, one
bar per run coloured by verdict, oldest to newest. It is real data, pure
CSS, and it is what an echo sounder actually draws.

The mobile fix is the same idea rather than a fallback. Tables become
label-and-value records with dotted leaders, which is how a sounding log
prints and is easier to read on a phone than any table that scrolls
sideways. Above 46rem every row shares one grid so the columns agree by
construction; the first attempt used table-cell and each row wrapped
independently, which produces a table that does not align — a list
paying for borders.

Found by rendering it rather than reading the CSS: the trace stranded
itself against the right edge when runs were few, "Open run" broke across
two lines, equal columns wrapped device names while a one-digit count
kept a quarter of the row, and the sign-in page carried no wordmark at
all, so you arrived somewhere that never said what it was.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 22:45:33 +02:00
mrambossekandClaude Opus 5 621ee99b77 adminui: make the web UI usable on a phone
The header was a rigid flex row, so on a narrow screen the account name
and the sign-out button were pushed off the side of the viewport where
they could not be reached at all — not merely ugly, unusable. It wraps
now, and below 40rem the account block takes its own full-width row so a
long display name cannot crowd out the navigation.

Wide content scrolls inside its own box rather than dragging the page
sideways with it. Tables sit in an overflow-x container and <pre> is
capped at the viewport width; without that, one long self-test line or
one device table makes every other column of text unreadable, and on a
phone it is not obvious that the page has moved at all. Long opaque
strings — device ids, enrolment links — wrap anywhere rather than
insisting on a width nothing has.

Also: box-sizing on everything, stat cards that share a row instead of
each claiming the full width, and larger touch targets on buttons, where
.4rem is comfortable with a mouse and fiddly with a thumb.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 22:33:41 +02:00
mrambossekandClaude Opus 5 89b084338f docs: fmr is on .150/::150 only; ::2 removed and verified across a reboot
Listeners came off ::2 before the address did — the other order fails to
bind on restart — and the address stayed up until the CNAME to fmr-1 had
landed, since dropping it earlier would have broken ACME renewal for the
name the certificate is issued to.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 22:27:57 +02:00
mrambossekandClaude Opus 5 23ed835e00 docs: record that a revived beacon belongs in the web UI, not its own listener
Its Python service wildcard-bound 0.0.0.0:443, which is what silently
compromised the reserved measurement address. Two routes on the admin UI
would inherit the TLS and certificate already in place, need no extra
port, and pick up authentication the standalone receiver never had.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 22:17:55 +02:00
mrambossekandClaude Opus 5 5b5a54db7d server: reserve measurement addresses; serve the UI on both families
fmr keeps .151/::151 for measurement. Their diagnostic value is entirely
in their listening state being known: a TLS handshake that completes on a
port nothing listens on proves interception, with no competing
explanation. One stray bind turns that proof into a shrug, and nothing
about the failure is visible — the run still says the network is clean.

CheckReserved refuses to start when a listener would take one. Wildcards
are refused outright, because that is how this actually happens: every
listener defaults to ":port" and the next one added gets copied from an
existing default, claiming every address without anyone deciding to.

Reserved is not silent, though. The first version of the guard would have
refused the live config's UDP and canary-DNS binds on .151, which are
deliberate — as is STUN's RFC 5780 alternate. Reserving an address and
then forbidding the measurements that need it defeats the purpose. The
rule is narrower: no services, and never ports 80 or 443.

The adb-beacon receiver was wildcard-bound to 0.0.0.0:443, holding port
443 on every IPv4 address including the reserved one, so the IPv4
interception test had been compromised for as long as it had run. It is
disabled; restore with systemctl enable --now echolot-adb-beacon. This
also marks the guard's limit: it governs this server's listeners, and a
process outside its config can still pollute a reserved address.

The admin UI and ACME responder were single-address, which is why the UI
could only live on ::2 and why the server was reachable over IPv6 alone —
the thing that made it look nonexistent from a phone without working
IPv6. Both now take address lists like every other listener.

Verified from outside: .150/::150/::2 answer on 443 with a valid cert for
fmr.echo-lot.app, .151/::151 are closed on 80 and 443, and canary DNS is
still up on .151.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 22:17:29 +02:00
mrambossekandClaude Opus 5 d9ee8bc2ae app: don't report ICMPv6 silence that was never measured
The per-network attribution fix worked — the finding named rmnet_data1
instead of the IPv4-only wifi — and immediately exposed a worse problem
underneath. Cellular's result was `ok: false` because binding a socket to
it failed with EPERM, so no echo request was ever sent; the finding then
reported "IPv6 is configured, but ICMPv6 gets no reply" about a network
the app had never pinged. That is an assertion about the user's carrier
with nothing behind it.

`attempted` now travels beside `ok`, set only once sendto has returned,
and the finding requires both. Failing to bind is a fact about this app's
permissions on this device; it says nothing about the network, and the
two must not share a boolean.

Verified on hardware with a VPN active: every network fails to bind with
EPERM, nothing is sent, and no ICMPv6 finding is emitted — where the
previous build would have blamed the carrier. Recorded in build-status:
Android blocks per-network binding entirely while a VPN holds the default
route, so per-network measurement is unavailable to anyone with one
connected. That needs a deliberate answer rather than a silently green run.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 21:32:34 +02:00
mrambossekandClaude Opus 5 7a5004f293 server: a non-admin account can manage its own uploads
Signing in and being allowed to administer the server were the same
question: the OIDC callback refused a session outright to anyone outside
the admin group. A legitimate user could authenticate, be told what they
could not do, and be left with no way to see or delete the data their own
devices had uploaded.

They are separate questions now. Everyone who authenticates gets a
session; the admin flag rides inside the MAC'd payload, so promoting
yourself means forging a signature rather than editing a cookie, and a
role that does not parse fails closed to "user".

Pages scope themselves through visibleDevices/mayTouchRun rather than
filtering individually — per-page scoping is what the next page added
will be missing, and that failure is silent, since a listing that leaks
other people's uploads looks exactly like one that does not. Someone
else's run answers 404, not 403: a distinguishable refusal would confirm
the run exists. Revoking devices and minting enrolment tokens affect the
whole server and stay behind adminOnly at the route table, where someone
looking for who-may-do-what will actually find it.

Ownership is re-read per request instead of captured at sign-in, so
unlinking an account takes effect immediately rather than at session
expiry. Tests cover that, plus the degenerate case of an empty subject,
which must own nothing rather than everything with an empty account id.

Also: attribute the ICMPv6 finding per network. It compared "is IPv6
configured anywhere on this device" against "did any network answer",
which on a phone reports IPv6-is-broken about a network where IPv6 was
never configured. network_ref is null on every test, so the probe now
records per-network outcomes structurally rather than as prose a finding
would have to parse.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 21:10:38 +02:00
mrambossekandClaude Opus 5 7eaf0c4190 app: name the two half-configured IPv6 shapes, per network
v6.no_icmp_reply infers trouble from silence, which is ambiguous by
construction: a firewall dropping echo requests looks the same as a
network that cannot carry IPv6 at all. Two much stronger signals were
already sitting unread in the link snapshot, and a test device on a
Netbird tunnel surfaced both at once.

v6.route_without_address — a ::/0 route with no global address. The
router advertises itself as an IPv6 gateway while SLAAC produces nothing
usable. Hosts believe IPv6 is available and pay a connection timeout on
every dual-stack destination before falling back, which is felt as
general slowness with no packet loss to explain it.

v6.no_default_route — the mirror: a global address with nothing to route
it. A VPN installing host routes to specific destinations produces this
deliberately and it works, so a VPN transport reports it as INFO rather
than as a fault; without one it means the network handed out an address
it does not carry traffic for.

Both are read from the routing table, so neither is inferred from
silence, and both are reported per interface — "IPv6 is broken" is
useless advice when wifi is the broken one and cellular is fine.

Classification lives in core-measurement rather than the ViewModel so it
can be tested without a device; the fixtures are a real dumpsys table
(wifi advertising a route it cannot source from, working cellular, a
tunnel with two host routes) because the risk here is not bad boolean
logic but imagining shapes real networks do not produce.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 20:42:03 +02:00
mrambossekandClaude Fable 5 fec374abf5 findings: v6.broken claimed a cause it had no evidence for
A phone reported "IPv6 is configured but not working" while loading an IPv6-only
site over TCP perfectly well. The finding fired on one signal - ICMPv6 echo
getting no reply - at HIGH confidence. ICMPv6 echo is widely filtered on
networks where IPv6 works, so the two cases are indistinguishable from where the
app stands, and it was picking one.

Same class of error as the multi-homed downstream-loss bug: a confident
measurement of something that was not happening. Now v6.no_icmp_reply, low
severity, medium confidence, naming both explanations. Still reported, because
filtered ICMPv6 breaks Path MTU Discovery - large packets vanish instead of
being reported as too big - which is a fault in its own right.

Corroborating with a real IPv6 connection would separate the two properly, but
needs a target, which runs into the hardcoded-deployment issue already open.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 20:16:37 +02:00
mrambossekandClaude Fable 5 b5c8dda9a2 app: sign in to the server's identity provider
Authorization code with PKCE, a Sign in card in settings, and the echolot://auth
redirect handled next to the enrolment one - told apart by host, because one
spends a token and the other completes an authorization, and confusing them
would fail obscurely.

The detail that decides whether this survives a real phone: the PKCE verifier is
written to storage before the browser opens rather than held in memory. Handing
control to a browser backgrounds the process and Android may kill it, so the
callback arrives at a fresh one. An in-memory verifier works on a developer's
device and fails under memory pressure.

Pending state is cleared before the exchange is attempted, whatever the outcome:
it is single-use, and leaving it behind would let a later callback complete a
flow nobody started.

Also records an open issue the question about server requirements surfaced: two
probes hardcode the reference deployment, so a user with no server still sends
DNS and STUN traffic to fmr without being told. For a tool this careful about
what leaves the device, that is the wrong default.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 19:53:38 +02:00
mrambossekandClaude Fable 5 4e6f2da3fb runs: scope by account; app: the PKCE half of signing in
server-release / image (push) Successful in 15s
server-test / test (push) Successful in 36s
server-release / release (push) Successful in 38s
Three phones on one account now produce one history, which is the main reason to
have accounts beyond upload permission. GET /v1/runs returns the account's runs
and says how many devices contributed; fetching and deleting resolve a run id
against the caller's own devices, so an id from another account is not found
rather than fetched from wherever it happens to live.

The rule that needed stating: the empty account is never a group. Devices nobody
has signed in on are unrelated devices that share the absence of an owner, and
matching on "" would let any anonymous device read every other one's runs.
Tested, along with sibling-device access working and cross-account access not.

App side: authorization code with PKCE. The app is a public client - anything
compiled into an APK can be read out with unzip and strings - and the redirect
returns through a custom URI scheme that any app on the device may register, so
an intercepted code is a real risk. PKCE makes a stolen code worthless: it can
only be exchanged by presenting a verifier that never left the process.

A callback whose state does not match is refused before the code is spent and
before any network call, since that is exactly how someone gets a victim to
complete the attacker's sign-in.

Nothing from the IdP is retained. The ID token is used once to prove who is
signing in and then discarded; the device credential authenticates everything
afterwards. No access tokens to store, no refresh tokens to rotate.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 19:46:32 +02:00
mrambossekandClaude Fable 5 0eaba6150b adminui: an admin interface, behind authentication without exception
server-release / image (push) Successful in 15s
server-test / test (push) Successful in 36s
server-release / release (push) Successful in 38s
Replaces the unauthenticated admin mux. Everything but /healthz requires a
session, and that is the point: the previous arrangement relied on binding to
loopback, which worked exactly until the address changed and then failed
silently and publicly. A binding address is a deployment detail, not an access
control, and this package does not treat it as one.

Two ways in. OIDC through the confidential client, with state and PKCE - PKCE
even here, because it costs one hash and closes code interception independently
of the secret. And the break-glass password, throttled, for when the IdP is the
thing that is broken. Signing in without the admin group is refused with the
group named, because "you are not an admin" is a different problem from "your
password is wrong" and the remedy is elsewhere.

Sessions are MAC-checked cookies: HttpOnly, SameSite=Lax, Secure when TLS is on.
CSRF tokens are derived from the session rather than stored, so there is no
server-side table to keep in sync, and they are required on every state-changing
POST - SameSite already blocks cross-site posts in current browsers, but this is
the control that does not depend on the browser being current.

Server-rendered with html/template and no JavaScript: the pages are lists and
forms, and a framework would add a build step, a dependency tree and an update
treadmill to a program that has none of those. The CSP is default-src 'none'
accordingly.

Pages: overview, devices (with revocation and enrolment-link minting), uploaded
runs and a run viewer. Revocations and deletions are logged with who did them.
Runs are shown exactly as uploaded, at the privacy level their uploader chose -
nothing in the UI can un-redact one.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 19:31:06 +02:00
mrambossekandClaude Fable 5 7bb54e1ec8 docs: record the admin-listener exposure, and the encrypted-upload design
The incident is written down with its cause rather than just its fix: the admin
listener was built localhost-only, and that assumption travelled with it when I
changed the address. The compounding error is the one worth remembering -
checkAdminExposure verifies encryption and says nothing about authentication, so
it passed and gave false confidence. A green light on an adjacent property is
worse than no check.

Also records the encrypted-upload idea while the reasoning is fresh, including
the four consequences that decide whether it is worth building: what metadata
must stay readable (and what the UI loses if it does not), that losing the
passphrase loses the data by design, that metadata is not hidden regardless, and
that it makes a server-side anonymization floor unenforceable - which is fine,
since encryption serves the same purpose better.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 19:04:09 +02:00
mrambossekandClaude Fable 5 c7750fbf0b oidc: one verifier per issuer, because IdPs mint one per application
server-release / image (push) Successful in 15s
server-test / test (push) Successful in 34s
server-release / release (push) Successful in 35s
Authentik derives the issuer from the application slug, so two applications mean
two issuers - and a token's `iss` must match whoever signed it. A single pinned
issuer could therefore only ever serve one of the two clients.

So there is a verifier per issuer, and each accepts only the client belonging to
it. That is tighter than the previous arrangement as well as more general: a
token minted for the phone cannot be replayed at the admin login, and vice
versa, because they arrive at different verifiers with different audiences.

ECHOLOT_OIDC_APP_ISSUER is optional - empty means both clients share
ECHOLOT_OIDC_ISSUER, which is what IdPs with one global issuer do.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 18:47:34 +02:00
mrambossekandClaude Fable 5 5d7f59a66a acme: answer HTTP-01 from the server itself, on port 80
server-release / image (push) Successful in 15s
server-test / test (push) Successful in 34s
server-release / release (push) Successful in 35s
HTTP-01 always arrives on port 80 - the CA chooses the port, not the operator -
so it never collides with an admin UI on 443. The conflict only exists for
TLS-ALPN-01, which is the challenge type that does use 443.

Given that, the server keeps a permanent listener on 80 that answers challenges
from a webroot and redirects everything else to the admin UI. Same arrangement
as the webroot plugins for Apache and nginx, and better than letting the ACME
client bind 80 per renewal: nothing binds and unbinds, so a renewal cannot fail
because the port was briefly busy, and the client needs only write access to a
directory instead of the privilege to bind a low port. Port 80 also gets a use
it would want anyway.

The ACME client stays an external program. lego is also a Go library, but
importing it would put a large dependency tree into a server that deliberately
has none, and the CLI does the same job from a timer.

Tokens are validated by *shape* before any filesystem call, so traversal never
reaches the disk - a stronger guarantee than sanitising a path and trusting the
sanitiser.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 18:16:09 +02:00
mrambossekandClaude Fable 5 6afcb131ef admin: terminate TLS in the binary, with a certificate that reloads itself
server-test / test (push) Successful in 33s
Direct rather than behind Caddy or nginx. This binary already serves TLS for the
control plane, so it is reuse rather than new machinery; one process with one
config file is most of what makes this thing pleasant to run; and a proxy on the
box would invite someone to eventually front the control plane too, which would
break SPKI pinning because clients pin that certificate's key.

The hard part of TLS is not termination, it is renewal - so the certificate is
re-read when the files change. No reload hook to write, and none to quietly stop
working months later and be noticed only after the certificate has expired. A
torn write (renewal tools write cert and key separately) keeps the previous
certificate rather than taking the listener down.

Not applied to the control plane, on purpose: clients pin that key, so replacing
it should cost an operator a moment's thought and a restart, not happen because
a file changed. Two listeners, two different right answers.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 17:51:58 +02:00
mrambossekandClaude Fable 5 cd187f9ef5 config: the OIDC client secret, admin TLS, and a stop on plaintext admin
server-test / test (push) Successful in 34s
Two gaps found while answering where configuration lives.

The confidential admin client needs a secret and there was nowhere to put one -
I had added the issuer and both client ids but not the secret the admin login
actually needs. It now reads from ECHOLOT_OIDC_CLIENT_SECRET, and preferably
from ECHOLOT_OIDC_CLIENT_SECRET_FILE: a secret in the environment is readable by
anything that can see /proc/<pid>/environ and lands in every dump of the unit's
config, whereas a path is one file whose permissions an operator can reason
about. (/etc/echolot-server.env was also 0644; now 0600 on fmr.)

And the server now refuses to serve the admin UI in plaintext on a non-loopback
address. The session cookie is a bearer credential for everything the server can
do, and the OIDC authorization code arrives in a URL; in the clear, both belong
to anyone on the path - and on a globally routable address that is the internet.
A hard stop rather than a warning, because a warning in a log is not read by the
person who most needs it, and because the safe answers are cheap: bind to
loopback and tunnel, or supply a certificate. ECHOLOT_ADMIN_INSECURE=1 overrides
it, so the decision is made rather than stumbled into.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 17:49:26 +02:00
mrambossekandClaude Fable 5 3a4cb1c327 cli: survive the --serve transition when nobody is watching
server-release / image (push) Successful in 16s
server-test / test (push) Successful in 33s
server-release / release (push) Successful in 34s
Deploying v0.8.0 broke fmr, and the reason is a flaw I should have seen:
self-update is executed by the OLD binary, so the unit repair I put in the new
binary's updater cannot fix the very update that installs it. The unit kept its
argument-less ExecStart, the new binary answered that with usage and exit 2, and
the service went into a restart loop.

Fixed on fmr by hand, but that is not a fix for anyone else - and the whole
premise of an unattended self-update is that nobody is watching when it happens.

So: when started with no verb *and* systemd started us, the server repairs the
unit and serves anyway, loudly. systemd sets INVOCATION_ID for every service
invocation and nothing else does, so a person at a terminal still gets usage and
a non-zero exit. Marked as a one-release shim to remove once no deployment
predates --serve.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 17:44:33 +02:00
mrambossekandClaude Fable 5 3cdbccee18 cli: serving is an explicit verb; no arguments prints usage
server-release / image (push) Successful in 15s
server-test / test (push) Successful in 33s
server-release / release (push) Successful in 34s
Running an unfamiliar binary by name should tell you what it does, not bind a
dozen ports and start answering the internet. --serve (or --daemon) now does
that, and a bare invocation prints usage and exits 2 - non-zero on purpose, so a
service manager sees a failure rather than concluding the server ran and
finished cleanly.

The hazard this creates is worth spelling out, because it bites once and
silently: three places started the binary with no arguments - the systemd unit,
the unit template, and the Dockerfile - and --self-update replaces the binary
but never the unit. A routine update would therefore leave a service that cannot
start, discovered whenever the host next rebooted.

So the updater repairs it: after replacing the binary it appends --serve to an
ExecStart that has no flags, but only in a unit this program wrote (identified
by its description). Editing an operator's hand-written unit would be overreach;
leaving ours broken would be negligence.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 17:42:20 +02:00
mrambossekandClaude Fable 5 80d2092f1b oidc: accept both the app's public client and the server's confidential one
server-test / test (push) Successful in 37s
Explaining public vs confidential clients surfaced a gap in my own design: I had
assumed a single client id, but there are two clients here with genuinely
different properties.

  the Android app     public + PKCE, because an APK cannot keep a secret
  the admin UI        confidential, because the server can keep one in
                      /etc/echolot-server.env and weakening it to public buys
                      nothing

So the audience check now accepts either registered client id - and only those
two. "Any client of this issuer" would let every other application registered
with the same IdP authenticate here, which is the entire reason the check
exists. Either id alone is enough to enable sign-in, since an operator may
register only the app or only the admin UI.

The profile advertises the *app's* client id, since that is what a phone should
authorize as.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 17:31:33 +02:00
mrambossekandClaude Fable 5 89a5ff9139 adminauth: a break-glass local admin alongside OIDC
server-test / test (push) Successful in 44s
If the IdP is misconfigured, unreachable, or the admin group is a typo, the
operator is locked out of their own server with no way back short of editing
JSON on disk. A fallback that only matters when everything else is broken is
exactly the thing you cannot add later - by then you cannot get in to add it.

Stored as PBKDF2-HMAC-SHA256 from the standard library (Go 1.24+ has it, so no
dependency), 600k iterations, per-credential salt. A password rather than a
bearer token on purpose: a break-glass credential is the one most likely to end
up in a backup or a config-management repo, and a hash survives that where a
token does not. There is no email reset flow and should not be -
--set-admin-password on the host is the reset, and whoever can run it already
has the machine.

The password is read from stdin, never a flag, so it stays out of shell history
and the process list; piping still works for automation.

Details the tests pin, each for a reason:
  - the username is compared in constant time too, or a fast rejection is a
    timing oracle for which usernames exist;
  - the *stored* iteration count is used, so raising the constant later does not
    lock out existing passwords;
  - the throttle grows with consecutive failures but stays bounded and forgives
    after a quiet minute - a break-glass credential an attacker can lock out is
    a denial of service against the one person who needs it;
  - sessions are MAC-checked before anything in them is read, and rotating the
    secret invalidates every one at once, which is how they are revoked.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 17:12:54 +02:00
mrambossekandClaude Fable 5 ce6d0c2f64 oidc: the server becomes a relying party, and devices can carry an account
server-release / image (push) Successful in 15s
server-test / test (push) Successful in 35s
server-release / release (push) Successful in 34s
Echolot delegates identity to whatever IdP the operator already runs and stores
no passwords - no hashing, no reset flow, no lockout policy, and no credential
database to lose. For a tool people self-host next to other services, that is
the difference between one more service and one more thing that can leak
someone's password.

Verification is stdlib-only, matching the server's no-dependency rule. Longer
than jwt.Parse, and auditable in one sitting. The part that matters is the
algorithm allow-list: taking `alg` from the token is the classic forgery, so it
is fixed in code. Tests cover the real attacks against a genuine signer - a
self-contained IdP with real keys, because a mock that returns success proves
nothing about a verifier:

  alg=none, HS256/RS256 confusion, a payload swapped under a valid signature,
  a token addressed to another client, a token from another issuer, expired
  and future-dated tokens, and discovery that renames the issuer (which would
  otherwise have us fetch a stranger's keys believing they were the provider's).

With no admin group configured nobody is an admin. An operator who has not said
who may administer the server has not thereby said "anyone who can log in".

Device and account stay separate concepts: enrollment admits a device (operator's
token), signing in attributes it to a person (POST /v1/account/link, device
credential plus ID token - both required, neither substitutes). uploads=account
now means what it says instead of refusing everyone, and signing in does not
override uploads=off.

The profile advertises the sign-in configuration so the app can offer the button
only when there is something behind it, and drive PKCE without anyone typing an
issuer URL. A discovery failure is reported rather than hidden, so "configured
but the provider is not answering" is distinguishable from "not configured".

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 16:52:56 +02:00
mrambossekandClaude Fable 5 57a5ef8796 app: the stable-pseudonym switch was live at a level that pseudonymizes nothing
At `full` the anonymizer returns the document unchanged, so the salt has nothing
to act on - but the switch was enabled and looked like it did something. A
control that silently does nothing is the same class of fault as the preview
button and the archived-level label: the screen implying more than is true.

Shown disabled with the reason rather than hidden. The setting is still stored
and applies the moment the level changes, so making it vanish would hide state
that is still there; and a settings screen whose controls appear and disappear
as you touch other controls is harder to trust, not easier. The label dims with
the switch so "not active right now" reads at a glance.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 16:16:36 +02:00
mrambossekandClaude Fable 5 c19f382640 privacy: scrub identifiers inside raw shell output
Running the Shizuku tier for the first time uploaded every MAC address on the
local network to the server at the balanced level - fourteen of them, router and
all. The probes embed raw command output verbatim (ip neigh, ip route), which is
good evidence and also a complete household device inventory, and the anonymizer
could not see it: classification is by field name and whole-value shape, and
ip_neigh is one long string that is itself neither a MAC nor an address.

measurement-schema.md flagged raw dumps as hard to anonymize and proposed
dropping them from exports. Scrubbing is better: identifiers inside unclassified
strings are replaced in place with the same pseudonyms used elsewhere, so a MAC
appearing in both a parsed field and a raw dump still reads as one device, and
the dump stays readable - neighbour-table shape, host count, RFC1918 addresses
and vendor prefixes all survive. Dropping it would have protected the same data
by destroying the reason for collecting it.

One pass, not three: sequential passes re-process their own output. Once a MAC
became 78:9a:18:xx:yy:zz the IPv6 pattern matched it - six hex groups separated
by colons is an address - and destroyed the vendor prefix the MAC rule had just
preserved. Ordered alternation resolves each position once, MAC first.

RealDocumentTest runs the anonymizer over a captured run when ECHOLOT_REAL_RUN
points at one and fails on any surviving MAC; it self-skips otherwise so no
one's network lands in the repo. Against the document that leaked: 14 in, 0 out.

Also: the Settings preview button did nothing, reading UiState.history which is
empty until the History screen has been opened - same root cause as the "0
run(s)" count. It reads the archive now, and says when there is nothing to show.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 16:12:13 +02:00
mrambossekandClaude Fable 5 d04babff51 engine: live test for upstream throughput
3125 sent, 3125 counted by the server, 0% loss. The assertion that earns its
keep is received <= sent: that is what catches a counter that was never reset
between runs, which would otherwise look like a suspiciously good result.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 15:56:51 +02:00
mrambossekandClaude Fable 5 892e952a8e throughput: the upstream direction, counted by the only party that can
server-release / image (push) Successful in 15s
server-test / test (push) Successful in 33s
server-release / release (push) Successful in 33s
The client generates the traffic and the server counts it. No grant is involved
- the client is sending its own packets, so there is nothing to amplify - but it
does need the server's tally, because only the far end knows how much arrived.
Without that number a sender measures how fast it can transmit, which is usually
just the speed of the local NIC and is not the question being asked.

A new wire type the server counts and deliberately never answers: a reply would
double the traffic and drag the return path into a measurement that is
specifically about the outbound one.

The tally is a counter, not a list, and short-circuits before the observation
log. A five-second run at 20 Mbps is around ten thousand packets; one struct
each would turn a measurement into an allocation storm on a shared server, and
nothing needs the per-packet detail since the client holds the send-side record.
The gap between the two counts is the loss.

direction=up on the throughput action sends nothing - it zeroes the counter, so
a second run in one session measures itself instead of inheriting the first.

Same honesty rule as downstream: measures_network is false when what arrived
matches what was offered, because then the path was never the constraint.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 15:54:09 +02:00
mrambossekandClaude Fable 5 8646bab52d findings: adopt the registry in the app module; rename ipv6.* to v6.*
The registry was only used in core-engine. The app still emitted seven codes as
raw strings, so the registry test passed while codes lived outside it - among
them ipv6.broken, which fired on a real network and was in no registry at all.

All seven now take their code, category and severity from a registry entry, so
those three cannot disagree at a call site. Grepping for code = "..." across the
app, engine and probe modules now returns nothing.

ipv6.* -> v6.* is the third instance of the same rule being broken: they
declared Category.IPV6 while the prefix map only knows "v6", so
TestType.category("ipv6.broken") fell through to connectivity and the finding
rolled up under the wrong verdict light. The test-type registry already used v6.

Two severities reconciled rather than assumed:

  connectivity.captive_portal is medium, not high. The registry had guessed
  high; the probe emitting it had always said medium, and the probe was the
  considered value - a captive portal on hotel wifi is what should be there.
  no_internet keeps high, since nothing local fixes that.

  v6.not_offered stays info, and the registry now says why it must. Most
  networks still do not offer IPv6; a warning there lights a yellow verdict on a
  healthy network and teaches people to ignore the light.

Plus a BackHandler: the screen was a plain state variable with nothing tying it
to the back stack, so Back left the app from Settings/History instead of
returning to the run screen.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 15:47:52 +02:00
mrambossekandClaude Fable 5 6bba420845 app: the history row was naming the wrong document's privacy level
A row read "22 kB · full" directly beneath "uploaded to fmr", while the status
line above said the upload went as BALANCED. Both were true and they described
different documents: the row showed ArchivedRun.anonymization, which describes
the *archived* copy - deliberately unredacted, so always "full" - and the status
line described the *uploaded* copy.

Read together, that says the complete data was uploaded when a redacted copy was
sent. A privacy display that overstates what left the device is worse than none,
and telling the user what left the device is the one thing this screen is for.

The level a run was uploaded at is now recorded separately (uploaded_as) and the
row says "kept complete on this device" / "uploaded to fmr as balanced" - each
label naming the copy it belongs to.

Two more from the same screenshot:

  - Every row showed no verdict. The archive read summary.verdict; the schema
    calls it summary.overall. Silently null on every run, so the list's most
    prominent element was blank while everything else looked fine. The test
    fixture had the same wrong field name, which is why it passed.
  - The status line rendered the server's raw JSON index entry into the UI.

Verified on device: a fresh run archives with verdict "yellow" and
uploaded_as "balanced" beside anonymization "full".

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 15:37:05 +02:00
mrambossekandClaude Fable 5 ac6c653115 privacy: pseudonymize the whole ULA prefix, not just its tail
Found in a real uploaded run from the phone: the server held
fda1:3fb1:ff92:6696::2662 for a DNS server. The general IPv6 path keeps the
leading two groups on purpose - for a global address that preserves the ISP
allocation, which is the useful part - but for a ULA that passes through 32 of
the 40 random bits of the global ID.

A ULA looks like the v6 RFC1918 and the instinct is to treat it the same. It is
not analogous, and the difference is the point: an RFC1918 prefix is shared by
millions of networks and identifies none of them, while a ULA global ID is
random and unique to one network by construction (RFC 4193). The prefix IS the
identifier, so it was a network fingerprint surviving redaction.

Pseudonymized as a unit now, so two addresses on one ULA subnet still share a
pseudonymous prefix - "these hosts are on one network" survives, "this is that
network" does not. RFC1918 stays readable, and the contrast is what justifies
it; a test pins both halves.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 15:26:09 +02:00
mrambossekandClaude Fable 5 305d21f8a7 app: insets on the two newer screens, and one source for the run count
On-device verification found both.

safeDrawingPadding() was on the run screen but not on Settings or History -
they were added later and never got it - so "< Back  Settings" sat under the
status-bar clock. The same fault the run screen had already fixed, reintroduced
by new code that did not know about it.

Settings also read "0 run(s), 23 kB stored": the count came from
UiState.history, which stays empty until the History screen has been opened,
while the size read the archive directly. Two sources for one fact; the count
now reads the archive too.

Verified on a OnePlus 15 (A16): header clears the status bar, count reads
"1 run(s), 23 kB stored".

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 15:20:57 +02:00
mrambossekandClaude Fable 5 172afb421d privacy: fix a real leak - global IPv6 addresses were uploaded verbatim
Setting out to build the machine-readable schema, the first step was checking
whether the anonymizer covers the fields the schema declares sensitive. It did
not, and five identifying values were going out at the `balanced` level:

  networks[].link.addresses[].addr   the device's own global IPv6 address
  networks[].link.routes[].gateway   the ISP allocation
  networks[].link.dns.servers[]      the configured resolver
  private_dns_hostname               an internal hostname
  search_domains[]                   the internal domain

The settings screen describes that level as pseudonymizing addresses.

Root cause: classification keyed on field names, and the schema's actual names
were never added to the table. Every existing test passed, because each checked
a field somebody had remembered to write a case for - an unfalsifiable design
for a privacy control.

So beyond adding the names, classification now falls back to the *value* when
the name is unknown: anything shaped like an IPv4/IPv6 address or a MAC is
treated as one. Hostnames deliberately are not inferred by shape, since
train.udp_updown is indistinguishable from a domain and mangling a test type
would corrupt the document to protect nothing.

LeakTest is the guard, and is written to fail for fields nobody thought of: it
plants identifying values wherever one can occur and asserts none survive. It
also pins that RFC1918 addresses stay readable, so it cannot pass by
over-redacting. Route prefixes and :: needed care - 0.0.0.0/0 must stay itself
or a routing table becomes unreadable for no privacy gain.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 14:45:52 +02:00
mrambossekandClaude Fable 5 e7afc2210f findings: a registry, because the codes had already drifted
A finding code is the stable half of a result - what a dashboard groups by and
what someone greps a year of archived runs for. That only holds if a code means
exactly one thing forever, which fifteen ad-hoc string literals cannot promise.

By the time this was written the failure had happened twice:

  - Two emitters independently produced connectivity.downstream_loss and
    connectivity.loss_downstream for the same claim. Nothing objected. Anyone
    aggregating either would have silently seen half their data.
  - Two codes sat under nat.* while being declared Category.CONNECTIVITY.
    nat.udp_unreachable is not about NAT, and the prefix decides the category,
    which decides which verdict light the finding rolls up into. Renamed while
    that is still cheap.

Codes are now typed FindingSpecs carrying category and default severity;
emitters reference the spec rather than retyping the string, so a typo is a
compile error and two call sites cannot disagree about a finding's category.

docs/findings-registry.md is the contract and a test reads it, failing when the
document and the code disagree on which codes exist or how severe they are.
Documentation that drifts from its implementation is worse than none, because it
still looks authoritative. The check reads table rows only, so the prose can go
on explaining which codes were retired and why.

Closes open item 1 of measurement-schema.md section 9.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 14:25:54 +02:00
mrambossekandClaude Fable 5 f7701c2d2f engine: throughput is opt-in in the run config; document the work
A 5-second run at 50 Mbps moves ~30 MB. On a metered connection that is the
user's money, and a measurement tool that spends it unasked is not one people
keep installed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 14:14:20 +02:00
mrambossekandClaude Fable 5 3c9af04e6f grant: replace the rate check with a token bucket
server-release / image (push) Successful in 15s
server-test / test (push) Successful in 32s
server-release / release (push) Successful in 33s
The live throughput test found it: a 3-second run delivered 104 packets and
stopped after 50 milliseconds.

The rate check exempted the first 50 ms entirely, meaning to be lenient at
startup. The effect was the opposite. A sender could dump an unbounded burst
into that free window, and the instant the check switched on it compared those
bytes against 50 ms worth of allowance and refused everything until real time
caught up. Every short test passed — downtrain sends 50 packets, big_send seven
— and every sustained send died about fifty milliseconds in.

A token bucket (allowance = burst + rate x elapsed) has no such cliff; it is
smooth from t=0. The burst is 100 ms of the allowed rate, floored at one
ordinary datagram so a single packet is never refused outright. The floor is
deliberately one datagram: at 8 kbps a 64 KB floor would be sixty-four seconds'
worth, which is precisely the instant dump the ceiling exists to prevent. The
existing rate test caught that when I first tried it, and it was right.

Second half of the same bug: callers treated any refusal as terminal. TryAllow
now says why, so a sender can pace through a transient "too fast just now" and
still stop dead on a spent budget or an expired grant.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 14:10:30 +02:00
mrambossekandClaude Fable 5 3333788d9e throughput: paced downstream rate, with the qualifier that makes it honest
server-release / image (push) Successful in 15s
server-test / test (push) Successful in 32s
server-release / release (push) Successful in 32s
A throughput number reports the smallest limit on the path, and the sender's own
ceiling is one of the candidates. If the server was asked for 50 Mbps and 50
Mbps arrived, the network was never the constraint and "50 Mbps" says nothing
about it. So the result always carries limited_by and measures_network, and a
finding is raised only when the path is actually implicated.

Loss is computed against the *sender's* count, not the requested rate: the
server reports what it put on the wire, and the gap is the loss. A receiver
alone cannot tell "the network dropped it" from "the sender never sent it", and
guessing turns a healthy server-side limit into a phantom network fault. The
count is stored per action, not per packet — half a million packets of structs
would turn a measurement into memory exhaustion.

Sending is paced rather than flat out. An unpaced burst measures the server's
NIC and the first queue it meets, then collapses into loss that reads as a
network fault. The schedule is absolute rather than sleep-per-packet, which
would accumulate scheduler error and drift the rate down over a ten-second run.

Throughput gets its own grant budget sized from the request, so every other
action stays bounded at 8 MiB. When the byte cap binds before the clock does,
the *duration* is shortened and reported, rather than the run being truncated
halfway: promising thirty seconds and delivering twenty-one is the same
information with a surprise attached, and it keeps "the clock ended the run" as
the normal case — the only case where the rate is a clean property of the path.

That last behaviour came out of a test that failed honestly: 30 s at 100 Mbps
needs 375 MB against a 256 MB cap.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 14:05:23 +02:00
mrambossekandClaude Fable 5 35744c609e docs: record frag_send and the current testing state
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 13:49:43 +02:00
mrambossekandClaude Fable 5 a7dccf7da2 frag_send: crafted IP fragments, so ordering can be tested and not just delivery
server-release / image (push) Successful in 15s
server-test / test (push) Successful in 32s
server-release / release (push) Successful in 32s
Letting the kernel fragment an oversized datagram answers one question — do
fragments get through. It cannot answer the more interesting one, because the
kernel always emits them in order, first one first.

The classic middlebox fault is exactly about that ordering. Only the first
fragment carries the UDP header, and therefore the ports; a stateful firewall
or NAT that has not seen it has no flow to match the rest against, and many
drop them. That is invisible to any in-order test and shows up in the field as
"large DNS answers fail on this network" or "the tunnel breaks when the MTU
drops" — it works until the network reorders, then fails intermittently, which
is the hardest kind of fault to chase.

So the server now builds the fragments itself (raw socket, IP_HDRINCL) and
controls their order: in_order as a baseline, reversed, and first-fragment-last.
The datagram is assembled and signed whole before being cut up, so what the
client reassembles is indistinguishable from an ordinary packet — otherwise it
would be measuring our sender rather than the path.

Two details that would silently produce wrong answers:
  - The UDP checksum is computed rather than left zero. A zero-checksum datagram
    is dropped by some middleboxes, and that drop would be recorded as a
    fragmentation failure, which is the wrong conclusion entirely.
  - Fragment offsets are in 8-byte units, so non-final fragments are rounded to
    a multiple of 8. A 100-byte fragment is not an error, it is a datagram no
    host will ever reassemble.

frag-send is advertised only when a raw socket can actually be opened — checked
by opening one, since a permission model has more ways to say no than a
capability bit has to say yes.

Fragment header arithmetic is unit-tested (reassembly coverage, MF flags, shared
IP ID, 8-byte offsets, checksum verification), cross-compiled and run on Linux
since the code is build-tagged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 13:45:09 +02:00
mrambossekandClaude Fable 5 4ffa6e4ae2 engine: split packet loss by direction using the server's observations
"3 % loss" sends an engineer looking in both directions at once. The server
records every packet it received per sequence number, so the two cases are
distinguishable: sent-but-never-seen is upstream loss, seen-but-no-reply is
downstream. The findings say which, and say what is not implicated.

Downstream loss is measured against what reached the server, not against what
was sent — the other denominator counts every upstream loss twice and
overstates the return path.

Per-direction jitter comes out of the same records without needing synchronised
clocks: (server_rx - client_tx) carries a constant unknown offset, and
differencing successive samples cancels it, so RFC 3393 variation is honestly
attributable to a direction even though absolute latency is not.

Correlation is by wire sequence number, not loop index — the counter is shared
with every packet type on the session. ProbeSession exposes it even for a lost
probe, since that is precisely the packet whose direction is in question.

Live against fmr: 0.08 ms upstream jitter vs 0.85 ms downstream, an asymmetry a
round-trip test cannot see.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 13:09:59 +02:00
mrambossekandClaude Fable 5 3e7e3b8d33 scripts: one command to mint an enrollment link, QR included
Scanning beats pasting a 200-character string onto a phone, and with a device
attached the deep link can be delivered by adb with no typing at all.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 12:12:28 +02:00
mrambossekandClaude Fable 5 199807a8c9 docs: enrollment link encoding rules in the spec, session log in build-status
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 12:11:37 +02:00
152 changed files with 18982 additions and 409 deletions
+24 -6
View File
@@ -10,12 +10,18 @@
# releases don't trigger each other's pipelines. # releases don't trigger each other's pipelines.
# #
# Required secrets: # Required secrets:
# REGISTRY_TOKEN personal access token with read+write package scope — # REGISTRY_TOKEN personal access token with read+write package scope —
# the built-in Actions token is NOT accepted by the # the built-in Actions token is NOT accepted by the
# container registry (docker login → unauthorized). # container registry (docker login → unauthorized).
# Create: user Settings → Applications → Generate token. # Create: user Settings → Applications → Generate token.
# REGISTRY_USER optional; defaults to the pushing actor's username. # REGISTRY_USER optional; defaults to the pushing actor's username.
# The release job needs only the built-in GITHUB_TOKEN. # RELEASE_SIGNING_KEY base64 ed25519 seed that signs SHA256SUMS. Self-updating
# servers verify the signature against the public key baked
# into the binary (selfupdate.DefaultPublicKeyB64) and REFUSE
# unsigned releases, so this job hard-fails without it —
# a release nobody can install is better failed loudly here.
# Mint a pair with: go run ./cmd/release-sign -gen
# The release job otherwise needs only the built-in GITHUB_TOKEN.
name: server-release name: server-release
on: on:
@@ -44,6 +50,18 @@ jobs:
done done
(cd ../dist && sha256sum * > SHA256SUMS) (cd ../dist && sha256sum * > SHA256SUMS)
- name: Sign SHA256SUMS
working-directory: server
env:
RELEASE_SIGNING_KEY: ${{ secrets.RELEASE_SIGNING_KEY }}
run: |
[ -n "$RELEASE_SIGNING_KEY" ] || { echo "::error::secret RELEASE_SIGNING_KEY is missing — self-updating servers refuse unsigned releases, so publishing one would strand the fleet. Add it under Settings → Actions → Secrets."; exit 1; }
go run ./cmd/release-sign ../dist/SHA256SUMS
# Verify with the key baked into the binary we just built — catches a
# secret that does not match DefaultPublicKeyB64 before it ships.
PUB=$(grep -o 'DefaultPublicKeyB64 = "[^"]*"' internal/selfupdate/selfupdate.go | cut -d'"' -f2)
go run ./cmd/release-sign -verify -pub "$PUB" ../dist/SHA256SUMS
- name: Create release + attach binaries - name: Create release + attach binaries
env: env:
TOKEN: ${{ secrets.GITHUB_TOKEN }} TOKEN: ${{ secrets.GITHUB_TOKEN }}
+19
View File
@@ -145,6 +145,17 @@ First build downloads AGP/Compose/Shizuku from Google Maven + Maven Central.
Shizuku, **toggle Wireless debugging off/on** — Shizuku keeps running (separate process), a fresh Shizuku, **toggle Wireless debugging off/on** — Shizuku keeps running (separate process), a fresh
port + mDNS record appear, and the beacon/connector recover. Plan the Shizuku-tier dev loop port + mDNS record appear, and the beacon/connector recover. Plan the Shizuku-tier dev loop
around this (or USB, if ever available). around this (or USB, if ever available).
- **A poisoned Gradle *build cache* entry can silently drop a whole module from the APK.**
Symptom: the app dies with `ClassNotFoundException` for a class that plainly exists, while the
build is green and `./gradlew :app:dependencies` lists the module on `debugRuntimeClasspath`.
The module's own jar is correct; its code simply never reaches AGP's intermediates. `clean`,
`rm -rf */build` and `--rerun-tasks` all fail to fix it, because **none of them touch the build
cache** — look for `compileKotlin FROM-CACHE` in the log. Fix: rebuild with `--no-build-cache`.
Verify by grepping the APK's dex for a string literal that only that module defines; grepping for
a *class name* proves nothing, because callers carry the name as a reference whether or not the
class is packaged:
`unzip -o -q app-debug.apk "classes*.dex" && grep -a "pin-sha256:" *.dex`
Suspect this whenever a runtime failure contradicts a successful build.
- **Empty-jar race with the IDE.** VSCodium's Java/Kotlin extension runs its own Gradle daemon on - **Empty-jar race with the IDE.** VSCodium's Java/Kotlin extension runs its own Gradle daemon on
the same project; when it overlaps a CLI build, a module's `build/libs/*.jar` can end up the same project; when it overlaps a CLI build, a module's `build/libs/*.jar` can end up
containing only a manifest, and Gradle then considers `jar` up-to-date. Dependent modules fail containing only a manifest, and Gradle then considers `jar` up-to-date. Dependent modules fail
@@ -162,3 +173,11 @@ First build downloads AGP/Compose/Shizuku from Google Maven + Maven Central.
are SUPPORTED on both known devices via `Os.recvmsg` + `StructMsghdr` reflection. are SUPPORTED on both known devices via `Os.recvmsg` + `StructMsghdr` reflection.
3. Fold the confirmed capabilities + Shizuku dump-format samples back into the production 3. Fold the confirmed capabilities + Shizuku dump-format samples back into the production
`core-probe` / `core-shizuku` modules. `core-probe` / `core-shizuku` modules.
## Enrolling a device with a server
`echolot-app/scripts/enroll-link.sh [note]` mints a §2.1 bootstrap link on fmr over SSH and prints
it (plus a QR if `qrencode` is installed, plus the `adb shell am start -a …VIEW -d '<uri>'` command
when a device is attached). The link carries a single-use token — treat it as a secret until spent.
Never hand-assemble one: the base64 pin needs percent-encoding, and a pin wrong by one character
fails as an inscrutable TLS error rather than as a bad pin.
+830 -3
View File
@@ -211,8 +211,8 @@ Two collection-loop gotchas found while driving the phone over USB:
2. ~~If `trace.errqueue_reachable` = PARTIAL, add a C-over-JNI errqueue shim.~~ **Retired** — 2. ~~If `trace.errqueue_reachable` = PARTIAL, add a C-over-JNI errqueue shim.~~ **Retired** —
SUPPORTED on both known devices; `traceroute.udp4` reads real hops via `Os.recvmsg` + SUPPORTED on both known devices; `traceroute.udp4` reads real hops via `Os.recvmsg` +
`StructMsghdr` reflection, so no `:native` module is needed. `StructMsghdr` reflection, so no `:native` module is needed.
3. Start the Go server skeleton (enrollment + profile + sessions + UDP echo with observation 3. ~~Start the Go server skeleton per probe-protocol.md.~~ **Shipped** — live on fmr since
blocks + canary-DNS reference records) per probe-protocol.md. v0.2.0 (2026-07-31); see the server sections below.
4. Fold confirmed capabilities into the production `core-probe` / `core-shizuku` modules. 4. Fold confirmed capabilities into the production `core-probe` / `core-shizuku` modules.
## Production probe server — LIVE on dedicated VM "fmr" (2026-07-31) ## Production probe server — LIVE on dedicated VM "fmr" (2026-07-31)
@@ -222,7 +222,7 @@ SSH only — verified untouched by the daemon (explicit multi-address binds, no
Control: fmr-1:8443 (SPKI pin `zRV9qkiLnRexAeh4RrSfJzbPWO+U/2Oj2/NVM/KfXlg=`, verified Control: fmr-1:8443 (SPKI pin `zRV9qkiLnRexAeh4RrSfJzbPWO+U/2Oj2/NVM/KfXlg=`, verified
externally over v4+v6). UDP data plane on all four service addresses :8442 — the second IP is externally over v4+v6). UDP data plane on all four service addresses :8442 — the second IP is
the stun-5780 substrate. Daily randomized self-update timer installed (checksum-verified the stun-5780 substrate. Daily randomized self-update timer installed (checksum-verified
against SHA256SUMS; signature verification still TODO before treating the source as untrusted). against SHA256SUMS; signature verification landed 2026-08-02 — see "Release signing" below).
Host config in `/etc/echolot-server.env`. SSH access for sessions: `ssh claude-echolot`. Host config in `/etc/echolot-server.env`. SSH access for sessions: `ssh claude-echolot`.
## Server v0.3.0 — STUN + TCP echo + observations + actions (2026-07-31) ## Server v0.3.0 — STUN + TCP echo + observations + actions (2026-07-31)
@@ -660,3 +660,830 @@ One user-visible bug caught in the process: Go's JSON encoder HTML-escapes `<`,
default, so the refusal reached the client as `needs \u003e= 0.2.0`. Disabled at the encoder (this default, so the refusal reached the client as `needs \u003e= 0.2.0`. Disabled at the encoder (this
is an API, not a page), and the client now *parses* the error field instead of pattern-matching it, is an API, not a page), and the client now *parses* the error field instead of pattern-matching it,
so it survives whatever a future encoder decides to escape. so it survives whatever a future encoder decides to escape.
### Enrollment: the server mints the bootstrap link (server-v0.5.3 … v0.5.4, 2026-08-01)
Until now a device was configured by hand-typing a control URL, a base64 SPKI pin and a
credential. That is the step that goes wrong, and it goes wrong quietly: a pin off by one
character does not fail loudly, it just never matches, and surfaces days later as an inscrutable
TLS error.
`POST /admin/enroll-tokens` now returns the whole §2.1 bootstrap link alongside the token, because
the server is the only party holding all three parts at once. The app takes it from a paste or an
`echolot://enroll` deep link (so a QR scan configures a server in one action) and writes URL, pin
and credential **together or not at all** — a half-applied server fails later, somewhere else,
with an error pointing at the wrong thing.
The control URL comes from `ECHOLOT_PUBLIC_URL` (set on fmr to `https://fmr-1.echo-lot.app:8443`),
falling back to the first control listen address; a wildcard bind warns rather than emitting a
link to `0.0.0.0`.
**The encoding trap, which is the whole reason this is tested across both languages.** The pin is
base64, so it contains `+`, `/` and `=` — each of which means something else in a query string. An
unencoded `+` decodes to a space, leaving the pin wrong by exactly one character. Base64 has no
spaces, so the parser restores them; that cannot damage a correctly-encoded pin and it rescues
every hand-assembled link. `LiveEnrollmentTest` redeems a link the *server* produced, which is the
only way to catch a disagreement between the Go assembler and the Kotlin parser — a unit test on
either side alone cannot see it. It also asserts the token is refused the second time.
Also fixed a spec divergence found while reading §2.1: the spec names the field
`device_credential`, the first implementation shipped `credential`. The server now sends both and
the client prefers the spec's; the alias goes once nothing reads it.
Two process notes from this round:
- An edit to the admin handler silently failed to apply and the endpoint kept returning just the
token. Caught by deploying and *looking at the response*, not by trusting a green build.
- The live suite is now six tests (`LiveServerTest`, `LiveMeasurement`, `LiveGranted`,
`LiveUpload`, `LiveCompat`, `LiveEnrollment`), all green against fmr from the PC with no device.
### Directional loss: which way is the packet loss? (2026-08-01)
A round trip can only report that *something* was lost somewhere, which is the least useful form
of the answer — "3 % loss" sends an engineer looking in both directions at once. The server
already records every packet it received per sequence number (§6), so the two cases are actually
distinguishable, and `train.udp_updown` now reports them separately:
- sent, never seen by the server → **upstream** loss
- seen by the server, reply never arrived → **downstream** loss
Findings name the direction and say what is *not* implicated, which is half the value:
`connectivity.loss_upstream` ("the return path is not implicated: replies came back for everything
that arrived"), `connectivity.loss_downstream`, `nat.udp_unreachable_upstream`.
Two things the implementation gets deliberately right:
- **Downstream loss is measured against what reached the server**, not against what was sent.
Using "sent" as the denominator counts every upstream loss a second time and overstates the
return path. Pinned by a test with loss in both directions at once.
- **Per-direction jitter without synchronised clocks.** Absolute one-way delay would need clock
sync and we deliberately have none (the two-clock rule). But `server_rx client_tx` carries a
constant unknown offset, and differencing successive samples cancels it — so RFC 3393 one-way
delay variation *is* honestly attributable to a direction even though latency is not. A test
pins that a 10-second clock offset changes nothing.
Correlation is by **wire sequence number**, which is not the loop index: the counter is shared
with every other packet type on the session, so "the nth echo" is not "sequence n". `ProbeSession`
now exposes `lastSeq`, including for a probe that was lost — a lost packet still has a sequence
number, and that number is exactly what tells you which way it was lost.
Live against fmr: 20/20 both ways, and jitter of **0.08 ms upstream vs 0.85 ms downstream** — a
tenfold asymmetry that a round-trip measurement cannot see at all.
10 unit tests on the arithmetic (a wrong denominator here does not crash, it produces a plausible
number pointing at the wrong half of the network) plus the live correlation check.
### frag_send: crafted IP fragments, so *ordering* is testable (server-v0.6.0, 2026-08-01)
`big_send` with `df=false` answers one question — do fragments get through. It cannot answer the
more interesting one, because the kernel always emits fragments in order, first one first.
The classic middlebox fault is exactly about that ordering. Only the **first** fragment carries the
UDP header, and therefore the ports; a stateful firewall or NAT that has not seen it has no flow to
match the rest against, and many simply drop them. That is invisible to every in-order test, and in
the field it looks like "large DNS answers fail on this network" or "the tunnel breaks when the MTU
drops" — it works until the network reorders, then fails intermittently, which is the hardest kind
of fault to chase.
So the server builds the fragments itself (raw socket, `IP_HDRINCL`) and controls their order:
`in_order` (baseline), `reversed` (last fragment first), `first_last` (first fragment held back
250 ms). The datagram is assembled and **signed whole** before being cut up, so what the client
reassembles is indistinguishable from an ordinary packet — otherwise the test would be measuring
our sender rather than the path. New test type `mtu.frag_ordering`; findings
`mtu.fragments_blocked` and `mtu.fragment_reorder_sensitive`.
Two details that would otherwise produce confidently wrong answers:
- **The UDP checksum is computed, not left zero.** Zero is legal in IPv4 and would be less code,
but zero-checksum datagrams are dropped by some middleboxes — and that drop would be recorded as
a fragmentation failure, which is the wrong conclusion entirely.
- **Fragment offsets are in 8-byte units**, so non-final fragments are rounded down to a multiple
of 8. A 100-byte fragment is not an error; it is a datagram no host will ever reassemble.
`frag-send` is advertised only when a raw socket can actually be opened — checked by opening one,
because a permission model has more ways to say no (userns, seccomp, LSM) than a capability bit has
to say yes. fmr runs as root with `cap_net_raw` in its bounding set, so it is available there.
Fragment ordering runs only after `mtu.frag_delivery` shows fragments arrive at all; otherwise the
three orderings would each report "not delivered" and read as three faults instead of one.
The header arithmetic is unit-tested (reassembly coverage with no gaps or double-delivery, MF
flags, shared IP ID, 8-byte offsets, checksum verification over odd and even lengths). Because the
code is `//go:build linux`, the tests are **cross-compiled and run on fmr** — there is no Go
toolchain there, so `go test -c` plus scp is the loop.
Live against fmr: 4 fragments per burst, and all three orderings reassembled — a healthy path, and
the baseline against which a mobile network will be interesting.
### Testing state (2026-08-01)
Six live tests against fmr, all green, no device involved: `LiveServerTest`, `LiveMeasurement`,
`LiveGranted`, `LiveDownstream`, `LiveUpload`, `LiveCompat`, `LiveEnrollment`. Plus 74 client unit
tests and the full Go suite. Everything in the last several entries is verified from the PC; the
app's UI (settings, history, deep-link enrollment) and `mtu.pmtud_up` remain device-only.
### throughput: a rate, plus the qualifier that makes it a measurement (server-v0.6.1 … v0.6.2)
A throughput test reports the *smallest* limit on the path — and the sender's own ceiling is one of
the candidates. If the server is asked for 50 Mbps and 50 Mbps arrives, the network was never the
constraint and "50 Mbps" says nothing about it. So `perf.throughput_udp` always carries
`limited_by` (duration | budget | rate | send_error) and `measures_network`, and a finding is
raised only when the path is actually implicated. The live run against fmr reports 20 Mbit/s with
`measures_network: false`, which is the correct and useful answer.
Loss is computed against the **sender's own count**, fetched from the observations API, not against
the requested rate. A receiver alone cannot tell "the network dropped it" from "the sender never
sent it", and guessing turns a healthy server-side limit into a phantom network fault. The server
keeps one summary per action rather than per-packet records — a ten-second run at 50 Mbps is half a
million packets, and a struct each would turn a measurement into memory exhaustion.
Sending is **paced**, on an absolute schedule. Unpaced would measure the server's NIC and the first
queue it meets, then collapse into loss that reads as a network fault; sleep-per-packet would
accumulate scheduler error and drift the rate down over a ten-second run.
Throughput gets its own grant budget sized from the request, so every *other* action stays bounded
at 8 MiB. When the byte cap binds before the clock does, the **duration is shortened and reported**
rather than the run being truncated: promising thirty seconds and delivering twenty-one is the same
information with a surprise attached, and it keeps "the clock ended the run" as the normal case —
the only case where the rate is a clean property of the path. That behaviour came out of a test
that failed honestly (30 s at 100 Mbps needs 375 MB against a 256 MB cap).
It is **opt-in** in the run config, default off. A 5-second run at 50 Mbps moves ~30 MB; on a
metered mobile connection that is the user's money, and a tool that spends it without being asked
is not one people keep installed.
#### The bug the live test found
The first live run delivered 104 packets and stopped after 50 ms. The grant's rate check exempted
the first 50 ms entirely, meaning to be lenient at startup — the effect was the opposite. A sender
could dump an unbounded burst into that free window, and the instant the check switched on it
compared those bytes against 50 ms worth of allowance and refused everything until real time caught
up. **Every short test passed** (downtrain sends 50 packets, big_send seven); every sustained send
died fifty milliseconds in.
Replaced with a token bucket (`allowance = burst + rate × elapsed`), which is smooth from t=0.
The burst is 100 ms of the allowed rate, floored at one ordinary datagram — deliberately one, since
at 8 kbps a 64 KB floor is sixty-four seconds' worth, exactly the instant dump the ceiling exists to
prevent. The pre-existing rate test caught that when I first tried the generous floor, and it was
right to. Second half of the same bug: callers treated *any* refusal as terminal, so `TryAllow` now
says why — a sender paces through a transient "too fast just now" and still stops dead on a spent
budget or an expired grant. Both halves are pinned by regression tests.
### Findings registry (2026-08-01)
Closes open item 1 of measurement-schema.md §9. A finding code is the stable, machine-readable half
of a result — what a dashboard groups by and what someone greps a year of archived runs for — and
that only holds if a code means exactly one thing forever. Ad-hoc string literals at fifteen call
sites cannot promise that, and by the time the registry was written the failure had already
happened.
**Two emitters had independently produced `connectivity.downstream_loss` and
`connectivity.loss_downstream` for the same claim**, and nothing anywhere objected. Anyone
aggregating either one would have silently seen half their data. Merged into
`connectivity.loss_downstream`, paired with `loss_upstream` so the two directions read as a set.
**Two codes were also renamed out of `nat.*`.** `nat.udp_unreachable` is not about NAT — it means
no replies came back — but the prefix determines the category, and the category determines which
verdict light the finding rolls up into (§7.3). A `nat.*` code landing under *connectivity* is not
a naming quibble; it changes which light turns red. Cheap to fix now, a breaking change later.
Codes are now declared as typed `FindingSpec`s carrying their category and default severity, and
emitters reference the spec instead of retyping the string — so a typo is a compile error and two
call sites cannot disagree about a finding's category.
`docs/findings-registry.md` is the contract, and a test reads it: it fails when the document and
the registry have codes the other lacks, or when a severity differs. Documentation that drifts from
its implementation is worse than none, because it still looks authoritative. The check scopes
itself to table rows, so the prose can keep explaining which codes were retired and why.
Six tests: uniqueness, declared-vs-listed, prefix↔category agreement, naming convention, a
word-order-anagram check (the shape the duplication actually took), and the document agreement.
### A real privacy leak, found by starting on the machine-readable schema (2026-08-01)
The intent was `measurement.schema.json` (§8's promised companion). The first step — checking
whether the anonymizer actually covers the fields the schema declares as sensitive — found that it
did not, so that became the work.
**At the `balanced` level, five identifying values were being uploaded verbatim:**
| value | field | why it matters |
|---|---|---|
| `2001:…::150` | `networks[].link.addresses[].addr` | the device's own global IPv6 address — a strong, geolocatable device identifier |
| `2a02:…::1` | `networks[].link.routes[].gateway` | identifies the ISP allocation |
| `203.0.113.77` | `networks[].link.dns.servers[]` | the configured resolver |
| `nas.example.lan` | `private_dns_hostname` | an internal hostname |
| `example.lan` | `search_domains[]` | the internal domain |
The settings screen describes that level as pseudonymizing addresses. It was not.
**Root cause:** classification keyed on field *names*, and the schema's actual names (`addr`,
`gateway`, `dst`, `servers`, `search_domains`, `private_dns_hostname`) had never been added to the
table. Not a subtle bug — just an unfalsifiable design. The existing tests all passed, because each
one checked a field somebody had remembered to write a case for.
**Two fixes, one of them structural:**
1. The missing names were added.
2. More importantly, a **shape-based backstop**: when a field name is unrecognised, the *value* is
inspected, and anything shaped like an IPv4/IPv6 address or a MAC is treated as one. A name
table can only protect fields someone thought of, which is precisely the wrong property for a
privacy control. Hostnames are deliberately *not* inferred by shape — `train.udp_updown` is
indistinguishable from a domain, and mangling a test type would corrupt the document to protect
nothing.
`LeakTest` is the new guard and is written to fail for fields nobody has considered: it plants
identifying values wherever one can actually occur and asserts none survive, rather than checking
a list of known cases. It also pins that RFC1918 addresses still come through readable, so the
test cannot pass by over-redacting everything.
Route prefixes and the unspecified address needed care in the transform: `0.0.0.0/0` and `::/0`
must stay themselves, or a routing table becomes unreadable for no privacy gain.
**Still outstanding:** `measurement.schema.json` itself. Worth noting what this episode implies for
it — much of a document's payload lives in `evidence`/`metrics`/`params`, which are per-test-type
`JsonObject` by design and therefore *outside* any schema. A schema-driven anonymizer would have
less coverage there than the name-plus-shape one now does, so the schema should be built for
validation and external tooling, not as a replacement for the classifier.
### ULA prefixes are pseudonymized whole (2026-08-01)
Spotted in a real uploaded run from the phone: the server had
`fda1:3fb1:ff92:6696::2662` for a DNS server. The general IPv6 path preserves the leading two
groups (deliberately — for a global address that keeps the ISP allocation, which is the
diagnostically useful part), and for a ULA that passed through **32 of the 40 random bits** of the
global ID.
ULA looks like the v6 equivalent of RFC1918 and the instinct is to treat it the same. That
reasoning does not carry over, and the difference is the whole point: an RFC1918 prefix is shared
by millions of networks and identifies none of them, while a ULA global ID is random and unique to
one network by construction (RFC 4193). The prefix *is* the identifier — it is a network
fingerprint that was surviving redaction.
Now pseudonymized as a unit, so two addresses on the same ULA subnet still land on the same
pseudonymous prefix: "these hosts are on one network" survives, "this is *that* network" does not.
Three tests, one of which uses the exact value observed on the wire.
Worth recording as a reasoning trap: I had originally raised this as "ULA should probably be kept
verbatim, like RFC1918, for consistency". The surface analogy pointed the wrong way, and the
correct answer was the opposite.
### Registry adopted everywhere; v6 findings renamed; Back works (2026-08-01)
The findings registry was only adopted in `core-engine`. The app module still emitted seven codes
as raw strings, so the registry test passed while codes existed outside it — including
`ipv6.broken`, which fired on a real network and was in no registry at all.
All seven now reference registry entries for code, category and severity, so those three cannot
disagree at a call site. A grep for `code = "…"` across the app, engine and probe modules returns
nothing.
**`ipv6.*` → `v6.*`.** The third instance of rule 1: they declared `Category.IPV6` while the prefix
map only knows `v6`, so `TestType.category("ipv6.broken")` fell through to *connectivity* and the
finding rolled up under the wrong verdict light. The test-type registry already used `v6.`.
Two severities reconciled while merging:
- `connectivity.captive_portal` is **medium**, not high. The registry had guessed high; the probe
that emits it had always said medium, and the probe was the considered value — a captive portal
on hotel wifi is what should be there, and logging in clears it. `connectivity.no_internet` is
the high one, because nothing the user does locally fixes that.
- `v6.not_offered` is **info, and the registry says it must stay info**. Most networks still do not
offer IPv6 and that is not a fault; a warning here lights a yellow verdict on a healthy network,
which teaches people to ignore the light.
Also: a `BackHandler` now returns from Settings/History to the run screen. The screen was a plain
state variable with nothing connecting it to the back stack, so the system Back gesture left the
app entirely. Enabled only when there is somewhere to go back to, so Back still exits from the run
screen.
### Upstream throughput (server-v0.6.3, 2026-08-01)
The mirror of the downstream case: the client generates the traffic and the server counts it. No
grant is involved — the client is sending its own packets, so there is nothing to amplify — but it
does need the server's tally, because **only the far end knows how much arrived**. Without that
number a sender measures how fast it can *transmit*, which is usually just the speed of the local
NIC and is a different question from the one being asked.
`TYPE_THROUGHPUT_UP` (0x0F) is counted and deliberately **never answered**: a reply would double
the traffic and drag the return path into a measurement that is specifically about the outbound
one.
The tally is a counter, not a list, and short-circuits **before** the observation log. A
five-second run at 20 Mbps is around ten thousand packets; one struct each would turn a
measurement into an allocation storm on a shared server, and nothing needs the per-packet detail
since the client holds the send-side record. The gap between the two counts is the loss.
`direction=up` on the throughput action sends nothing — it zeroes the counter, so a second run in
one session measures itself rather than inheriting the first one's packets. The live test asserts
`received <= sent`, which is what catches a counter that was never reset.
Live against fmr: **3125 sent, 3125 counted, 0 % loss, 10.0 Mbit/s** at a 10 Mbit/s request, with
`measures_network: false` — correct, since what arrived matched what was offered, so the path was
never the constraint.
### Raw shell dumps leaked the whole LAN (2026-08-01)
Found by running the Shizuku shell tier for the first time. The tier works — `tiers.shizuku: true`,
`exec_path: UserService` (so the UserService binds on the OnePlus, as recorded), `runs_as
shell(2000)`, 7/7 commands — and the run promptly uploaded **every MAC address on the local
network** to fmr at the `balanced` level: router, phones, whatever else was on the wifi. Fourteen
of them.
The probes embed raw command output verbatim (`ip neigh`, `ip route`, `id`), which is genuinely
good evidence and also a complete household device inventory. The anonymizer could not see it:
classification is by field name and by whole-value shape, and `ip_neigh` is one long string that is
itself neither a MAC nor an address. measurement-schema.md §9 item 2 had flagged raw dumps as "hard
to anonymize" and proposed dropping them from exports; nothing enforced either.
**Scrubbing beats dropping.** Identifiers inside any unclassified string are now replaced in place,
using the same pseudonyms as everywhere else — so a MAC that appears both in a parsed field and in
a raw dump still reads as one device. The dump stays readable and auditable: you can still see the
neighbour table's shape, the host count, RFC1918 addresses and vendor prefixes. Dropping the
evidence would have protected the same data while destroying the reason for collecting it.
Two implementation notes worth keeping:
- **One pass, not three.** Sequential passes re-process their own output: once a MAC became
`78:9a:18:xx:yy:zz`, the IPv6 pattern matched it — six hex groups separated by colons *is* an
address — and destroyed the vendor prefix the MAC rule had just preserved. Ordered alternation
resolves each position once, MAC first.
- The patterns are conservative on purpose. A missed address gets caught by another rule or not at
all; an over-eager one mangles timestamps and version strings, corrupting evidence to protect
nothing.
`RealDocumentTest` runs the anonymizer over a captured run when `ECHOLOT_REAL_RUN` points at one,
and fails on any MAC that survives. It self-skips otherwise, so no one's network is committed to the
repo. Against the actual leaked document: **14 MACs in, 0 surviving.**
Also fixed: the Settings *Preview what an upload would send* button did nothing. It read
`UiState.history`, which is empty until the History screen has been opened — the same root cause as
the "0 run(s)" count. It now reads the archive directly, and says so when there is nothing to
preview rather than silently ignoring the tap.
### Security: the admin listener was publicly exposed for ~15 minutes (2026-08-01)
Moving the admin listener to `[::2]:443` for the UI exposed `/admin/enroll-tokens` and
`/admin/selftest` to the internet **with no authentication**. Anyone who could reach
`fmr.echo-lot.app` could mint enrolment tokens.
The listener was designed localhost-only — its own flag help says *"keep localhost"* — and that
assumption travelled with it when the address changed. The compounding error: `checkAdminExposure`,
added the same day, verifies **encryption** and says nothing about **authentication**. It passed,
and a green light on an adjacent property is worse than no check, because it invites you to stop
looking.
Closed by returning to loopback (the TLS and ACME work is retained, just not exposed). All 68 device
enrolments matched the timestamps of test runs, so there is no evidence of abuse — but the window
existed on a freshly published hostname and absence cannot be proven. 39 unused enrolment tokens
were purged, since any could have been minted by someone else and they cost nothing to replace, and
63 test devices removed.
**The admin listener does not become reachable again until it authenticates.** That reorders the UI
work: auth on the listener first, everything else after.
### Open: encrypted uploads, where the operator cannot read the data
Not built. Recorded because the shape is decided by a few early choices, and the current design
happens to leave the door open.
The goal: hand someone an account, let them upload, and be unable to read what they uploaded.
Sketch: a random per-account **master key**, generated on the first device and wrapped under a
key derived from a passphrase (PBKDF2-HMAC-SHA256 — stdlib on both sides). The wrapped key is
stored server-side as an opaque blob, so a new device signs in, fetches it, and unwraps locally;
the server never sees either key. Runs are encrypted client-side with AES-256-GCM, fresh nonce per
run. All of this is stdlib in Go and `javax.crypto` in Kotlin — no dependency either side.
Four consequences that decide whether it is worth it:
1. **What stays readable determines what the UI can do.** The server builds its index by *parsing*
the document — verdict, finding count, started_at. An opaque payload means the client supplies
that metadata or the index disappears, and with it retention-by-verdict and any "runs with
findings" view. The honest version supplies only run id, timestamp and size, and moves the rest
client-side.
2. **Lose the passphrase, lose the data.** That is the feature working, and also the support
burden. It needs a recovery code printed at setup, not a reset flow — there is nothing to reset.
3. **Metadata is not hidden.** The operator still sees which account uploaded, when, how often and
how large. "Cannot see it" is about content, not existence, and saying otherwise would oversell.
4. **It makes `min_anonymization` unenforceable** — a server cannot check a level it cannot read.
That is not a conflict so much as a redundancy: the anonymization floor exists to protect the
user from the operator, and encryption does that better. The two should not both be demanded of
one upload.
What keeps this possible: uploads are already stored byte-for-byte as received, and every index
field is derived in one function (`runs.Put`). The thing to avoid is admin features that *require*
reading content — those would have to be unbuilt later.
### App sign-in, and an undisclosed dependency it surfaced (2026-08-01)
The app can now sign in to the server's identity provider: authorization code with PKCE, a
`Sign in` card in settings, and the `echolot://auth` redirect handled alongside the enrolment one
(told apart by host, since one spends a token and the other completes an authorization).
The detail that decides whether this works on a real phone: **the PKCE verifier is written to
storage before the browser opens**, not held in memory. Handing control to a browser backgrounds
the process and Android may kill it; the callback then arrives at a fresh process. An in-memory
verifier works on a developer's device and fails under memory pressure, which is the worst way for
a sign-in to break.
Nothing from the IdP is retained. The ID token proves who is signing in, once, and the device
credential authenticates everything after — no access tokens stored, no refresh tokens rotated.
**A server remains entirely optional.** All eight probes are device-tier; `serverConfigured` gates
only upload and the account. But answering that question exposed something worth fixing: two probes
hardcode the reference deployment —
```kotlin
DnsCanaryProbe(canaryZone = "c.echo-lot.app", ...) // "Hardcoded to the reference deployment"
StunProbe(serverHost = "fmr-1.echo-lot.app")
```
so a user with no server of their own still sends DNS and STUN traffic to fmr without being told.
For a tool that goes to this much trouble over what leaves the device, an undisclosed dependency on
a third party's infrastructure is the wrong default. It should prefer the configured server, and be
explicit when there is none. **Closed** — both probes take the enrolled server from settings
(canary zone learned from the profile, cleared on re-enroll) and report themselves SKIPPED with
the reason when none is configured.
### v6.broken was a false positive waiting to happen (2026-08-01)
A phone could not open `https://fmr.echo-lot.app` while loading the same server by IP literal
perfectly well. Two things came out of chasing it.
**The admin UI is IPv6-only, by consequence rather than intent.** `fmr.echo-lot.app` has an AAAA
and no A record — verified identical at Cloudflare, Google and Quad9, so DNS itself is healthy.
That follows from reserving all four measurement addresses for testing, which left only `::2` for
management, and `::2` has no IPv4 counterpart. Any client without working IPv6 sees an unreachable
admin interface — a poor property for the interface you reach *from the networks you are debugging*.
**And the app's own `v6.broken` finding was unsound.** It fired on exactly one signal — ICMPv6 echo
getting no reply — with `Confidence.HIGH`. ICMPv6 echo is widely filtered on networks where IPv6
works fine, which is precisely what that phone demonstrated: no ICMPv6 replies, working IPv6 TCP.
The finding asserted a cause it had no evidence for, which is the same class of error as the
multi-homed `100 % downstream loss` earlier: a confident measurement of something that was not
happening.
Now `v6.no_icmp_reply`, severity low, confidence medium, and the text names *both* explanations
instead of choosing one. It is still worth reporting, because filtered ICMPv6 breaks Path MTU
Discovery — large packets vanish rather than being reported as too big — which is a real fault even
when IPv6 works.
The proper fix is corroboration: attempt a real IPv6 connection and only call it broken when that
fails too. That needs a target, which runs into the hardcoded-reference-deployment issue already
open above. **Both closed 2026-08-02** — see "Corroborated IPv6 findings" below.
## Per-network probing is blocked while a VPN is up (2026-08-01)
`Network.bindSocket()` fails with `EPERM` for every underlying network when a VPN holds the
default route — verified on the OnePlus 15 with Netbird active: `Binding socket to network 101
failed: EPERM` for both cellular and wifi. This is Android preventing VPN leaks, not a bug to work
around, and it means the whole per-network measurement approach is unavailable to any user with a
VPN connected. Worth deciding deliberately rather than discovering per report:
- The run currently succeeds and simply measures nothing per network. Honest, but silent — the
document records `attempted: false` and the UI says green.
- A user with a corporate VPN permanently on would get a green run that measured almost nothing.
Options are to detect the VPN and say so plainly ("this network cannot be measured while a VPN is
active"), to measure the tunnel itself as the network under test, or both. **Decided and built
2026-08-02**: say so plainly, everywhere the run is read — see "Constrained runs" below.
Related: `icmp.ping6` now records `attempted` alongside `ok` per network, because collapsing them
made the app report "IPv6 is configured, but ICMPv6 gets no reply" about an interface it had never
succeeded in sending on — a claim about the user's carrier with no evidence behind it.
## Reserved measurement addresses, and the web UI on both families (2026-08-01)
fmr has two IPv4 (.150/.151) and three IPv6 (::150/::151/::2) addresses. `.150`/`::150` now carry
the services; `.151`/`::151` are reserved for measurement, declared in `ECHOLOT_RESERVED_ADDRS`.
Reserved does **not** mean silent. The UDP data plane, the canary DNS and STUN's RFC 5780 alternate
all belong there — reserving an address and then forbidding the measurements that need it would
defeat the purpose. What must never appear is a service, and above all not ports 80 or 443: a
handshake completing on a port known not to be listening is what proves interception, and that
proof survives exactly as long as nothing binds those ports. `config.CheckReserved` enforces it at
startup, refusing wildcard binds outright (every listener defaults to `:port`, so the next one added
will claim reserved addresses without anyone deciding to).
The first version of the guard was too strict and the live config caught it: it would have refused
the existing UDP and DNS binds on `.151`. The rule is about services and web ports, not about
listening at all.
**The adb-beacon receiver was wildcard-bound to `0.0.0.0:443`**, occupying port 443 on every IPv4
address including the reserved one — so the IPv4 interception test had been compromised for as long
as it had been running, silently. It is now `systemctl disable --now echolot-adb-beacon`; restore
with `systemctl enable --now`. Note what this implies: the guard covers this server's own listeners,
and a stray process outside its config can still pollute a reserved address. A startup probe that
*verifies* 80/443 are actually free on the reserved addresses would be a stronger guarantee than
checking our own configuration — built 2026-08-02 (`selftest.ReservedWebPortsFree`, fatal at
startup when anything is listening there).
The admin UI and the ACME responder now take comma-separated addresses like every other listener;
they were single-address, which is why the UI could only ever live on `::2`. It serves on `.150:443` and
`[::150]:443`; sshd on `.150:2322` and `[::150]:2322`.
`::2` is gone entirely — unbound, then removed from `/etc/systemd/network/ext.network`. The
transition kept it bound throughout and dropped it only after the CNAME landed, because removing it
first would have broken both the UI and ACME renewal for the very name the certificate is issued
to. Listeners came off before the address did, in that order, or the services would have failed to
bind on restart.
Verified after a full reboot: `fmr.echo-lot.app` answers 200 over both families, `.151`/`::151` are
closed on 80 and 443, canary DNS is still up on `.151`, and neither `::2` nor the beacon returns.
(`echolot-server` is `After=network-online.target` with `Restart=on-failure`, which is what makes
binding specific addresses safe across a boot — a wildcard bind would not have needed it, and that
is the trade for the reserved addresses being meaningful.)
The point of all this: `fmr.echo-lot.app` gained an A record, so the server stopped being reachable
only over IPv6 — which is what made it unreachable from a phone with no working IPv6, presenting as
"this host does not exist" in two different browsers.
### If the beacon comes back, it belongs in the web UI
Not as a separate listener. The receiver being its own Python service on `0.0.0.0:443` is exactly
what silently compromised the reserved address, and a second process racing for a port is a
recurring problem rather than a one-off: whoever loses the race simply fails to start, and on a
reboot which one that is comes down to unit ordering.
Folding it in costs little and settles several things at once. It would be two routes on the admin
UI (`POST` the observed adb port, `GET /apk` for the staged build), behind the TLS the UI already
terminates and the certificate it already renews, with no extra port and no wildcard. It also gets
authentication for free — the current receiver accepts a port report from anyone who can reach it,
which is tolerable for a dev tool on a trusted network and not something to keep once it lives
beside an admin session.
The one thing that changes on the device side is that the POST becomes HTTPS. That is a real
certificate rather than a self-signed one, so it costs a URL scheme rather than any trust plumbing.
## The control plane shares port 443 (2026-08-01)
`fmr-1.echo-lot.app:443` is the control plane, `fmr.echo-lot.app:443` the admin UI, both on
`.150`/`::150`, one listener, selected by SNI for the certificate and by `Host` for the handler.
The reason is not tidiness, it is reachability. Captive portals, hotel wifi and corporate firewalls
routinely permit only 80 and 443 — which is exactly the population of networks this tool exists to
diagnose. A control plane on 8443 is unreachable precisely when it matters most, and it fails as
"cannot reach server", which tells the user nothing.
They cannot share a certificate, which is why this needs two names. The control plane is trusted by
SPKI pin and so uses a long-lived self-signed certificate; a browser needs one a CA vouches for.
One name on one port is one certificate, so the port can only be shared by splitting the names.
Pinning the Let's Encrypt key instead was considered and rejected: it survives renewal only while
key reuse holds, so a routine key rotation would brick the whole fleet.
Verified per SNI on 443: `fmr.echo-lot.app` serves `issuer=Let's Encrypt`, `fmr-1.echo-lot.app`
serves the self-signed cert whose pin is unchanged (`zRV9…Xlg=`), `/v1/profile` answers 401 on the
control name and 303 to the login page on the UI name.
**8443 stays open.** Devices enrolled before this carry that URL in their settings, and closing it
for the sake of a port number would strand every one of them. It can go once no enrolled device
still points at it — not before.
The rule from the naming change still binds: `fmr` may be a CNAME to exactly one host and never a
multi-address record, because a pinned client that reaches a different key does not fail over.
## Constrained runs: a VPN'd run now says so, everywhere (2026-08-02, app 0.2.1)
The measurement schema gained a top-level `constraints` block (§3) and the app now fills it.
`ConstraintDetector` (core-probe) runs before any probe: one throwaway `Network.bindSocket()` per
non-VPN network, plus a transport check for an active VPN. The result lands in three places, and
all three are deliberate:
- **`run.constraints`** — for machines. A server aggregating thousands of runs can now separate
"measured a healthy network" from "measured almost nothing through a tunnel"; the shapes were
identical before.
- **A `measurement.vpn_constrained` finding** — for the person reading this run, naming the
interfaces that went unmeasured. A constrained run with a quiet findings list still reads as
"nothing wrong here".
- **The §7.3 verdict** — `Verdicts.derive` takes the constraints and returns INCONCLUSIVE
outright for a per-network-blocked run, whatever the category lights say; the run screen shows
an amber "Measured through a VPN" banner above the verdict so INCONCLUSIVE reads as the OS
refusing, not the app failing.
Detection is one bind per network rather than parsing per-test `attempted:false` breadcrumbs, so
it cannot drift when probe evidence formats change.
## Corroborated IPv6 findings: v6.broken is back, with evidence (2026-08-02, app 0.2.1)
The new `V6ConnectProbe` (test type `v6.brokenness`) attempts a real TCP connection over IPv6 to
the configured server's :443, per network that *claims* IPv6 (global address or v6 default
route) — IPv4-only networks are not attempted, since their failure is by design and would
manufacture the exact false positive this exists to kill. The finding derivation is now three-way:
- ICMPv6 silent, TCP works → `v6.no_icmp_reply` at **high** confidence, retitled "ICMPv6 is
filtered here — IPv6 itself works" (still reported: filtered ICMPv6 breaks PMTUD).
- ICMPv6 silent, TCP fails too → **`v6.broken`** (high severity, reinstated in the registry +
findings-registry.md): two independent transports silent on a network advertising IPv6.
- No corroboration (no server configured, or the connect never got as far as sending) → the
two-explanation `v6.no_icmp_reply` at medium confidence, unchanged.
Like STUN and the canary, the probe SKIPs honestly when no server is configured — corroboration
is a benefit of enrollment, not a reason to borrow fmr.
## Server: reserved 80/443 verified against the OS, and signed releases (2026-08-02)
**Reserved-address startup probe.** `serve()` now proves 80/443 are actually free on every
`ECHOLOT_RESERVED_ADDRS` address before starting: a throwaway bind per port
(`selftest.ReservedWebPortsFree`), fatal on EADDRINUSE with the offending address named — the
check `CheckReserved` cannot do, because a stray process outside our config (the adb-beacon
receiver on `0.0.0.0:443` was exactly that) is invisible to configuration checks. Bind errors
that are not "in use" (typo'd address, address not on this host) warn instead of refusing —
they are config problems, not pollution.
**Release signing.** Self-update now trusts a signature, not a host. CI signs `SHA256SUMS` with
an ed25519 key (`relsign` package, `cmd/release-sign`) and the updater refuses any release whose
`SHA256SUMS.sig` is missing or does not verify against the public key baked into the binary
(`selfupdate.DefaultPublicKeyB64`; operators with their own pipeline override via
`ECHOLOT_SELF_UPDATE_PUBKEY`). The private key exists in exactly two places: the Gitea Actions
secret `RELEASE_SIGNING_KEY`, and the offline original on the dev PC at
`~/.echolot/release-signing-key`. It is deliberately NOT on fmr and NOT in the repo — a
compromised release host can withhold updates but no longer inject one. CI hard-fails when the
secret is missing (an unsigned release would strand every verifying server) and cross-checks the
signature against the key in the source it just built.
**ACTION REQUIRED before the next `server-v*` tag:** add the Gitea repo secret
`RELEASE_SIGNING_KEY` (Settings → Actions → Secrets) with the contents of
`~/.echolot/release-signing-key` from the dev PC. Ordering is safe: the currently deployed
v0.3.x updater does not verify, so it will happily install the first signed release; every
release after that is verified. **Done 2026-08-02** — the secret is in place. (The "v0.3.x"
above should read "the currently deployed release": deployments had moved on to v0.9.x by the
time signing landed; the point — the deployed updater predates verification and will accept the
first signed release — is unchanged.)
## Prober fold: traceroute.udp4 and the mDNS inventory go production (2026-08-02, app 0.2.2)
The two highest-value validated capabilities moved from the prober into `core-probe`:
- **`traceroute.udp4`** (`TracerouteProbe`): UDP traceroute reading ICMP time-exceeded off the
socket error queue via `Os.recvmsg(MSG_ERRQUEUE)` through the reflection facade — no root, no
raw socket, no JNI, ~250 ms for six hops. Emits the schema's `TracerouteEvidence` (rtt in ns).
`OsAbi` came with it, including the measured fact that `Os.getsockoptInt` exists on neither
known device, so PMTU must always be read from the errqueue (`ee_info`), never
`getsockopt(IP_MTU)`. The load-bearing line survived the port: EAGAIN out of the reflected
`recvmsg` means "queue empty", not failure.
- **`local.mdns_inventory`** (`MdnsInventoryProbe`): MulticastLock + NSD discovery, the service
inventory that doubles as the VLAN-leakage detector. Both hardware lessons kept: the
`_services._dns-sd._udp.` meta-query returns 0 beside live services on both devices (so the
concrete types are the measurement and the meta-query result is itself evidence), and the
listen window is 10 s because 4 s missed services.
Still to fold, in order: the Shizuku dump *parsers* (the raw `link.ip_monitor` captures already
hold two divergent vendor formats that could feed `link.ra_source` and `sec.arp_watch`);
`multinetwork.request_and_bind` (extend ConstraintDetector to *request* transports rather than
only probing present ones); `peer.ble_advertise` (needs three new permissions and a peer mode to
exist first).
## Server v0.9.2: trains, real TTL/DSCP/ECN, rate limits (2026-08-02)
The spec-vs-implementation gap audit closed its top items; protocol_version 1.0.0 → 1.0.1
(additive — below 1.0.0 the minor is the breaking axis, and nothing here breaks an old client):
- **Upstream trains** (§3.2, types 0x03/0x04/0x05): per-train bounded columnar buffer (8192
rows, head kept on overflow with `Truncated` set — mirrors the schema's `evidence_truncated`
honesty), TRAIN_REPORT split across ≤1200-byte datagrams, grant-free with the §3.4 argument
spelled out (a 17-byte report row answers a ≥36-byte HMAC-valid packet). Unknown train id
gets a zero-row report: "nothing arrived" is an answer. Also surfaced as `udp.trains` in the
observations API.
- **Real TTL/DSCP/ECN observation** (§3.3): the read loop is `ReadMsgUDPAddrPort` with
IP_RECVTTL/IP_RECVTOS/IPV6_RECVHOPLIMIT/IPV6_RECVTCLASS cmsgs on Linux; `0xFF` stays the
"not observed" sentinel elsewhere. This unblocks `sec.dscp_ecn_survival` both directions,
paired with the new `dscp` parameter on `downtrain` (validated 063, refused not clamped,
`dscp_applied` in the response).
- **Rate limiting** (§2.5, was entirely absent): token buckets keyed per credential AND per
source IP; 429 + Retry-After on session/action creation (`/v1/profile` stays ungated), silent
drop on the data plane — charged after the HMAC gate so a spoofed flood cannot drain a
victim's budget, before the replay window so a dropped seq stays usable. UDP ceilings default
above the largest legitimate run (a 200 Mbps throughput test), because a rate limit that
clips a real measurement produces a confidently wrong number.
- **`action_id` in every granted packet** (§5/§9): payload bytes [8:16] across all granted
types, so overlapping actions are attributable. Verified the deployed Kotlin client parses
only ECHO_RESP and MTU_ACK payloads, so the reshuffle strands nobody.
- **Canary log retention**: the stated 24 h privacy default is now enforced
(`ECHOLOT_DNS_LOG_RETENTION_H`), where before the log was time-unbounded.
- **`POST /admin/enroll-tokens`** now answers the spec's JSON shape under content negotiation;
the README's curl works as documented.
- Spec §2.3 registry gained `downtrain` and `tcp-echo`, which the server had been advertising
as strings a conformant client must ignore.
Client-side counterparts still to build: sending 0x03 trains + parsing 0x05 reports
(`train.udp_updown`), and passing `dscp` on downtrain actions.
## ⚠ Version lineage broken: fmr runs v0.11.2, the repo's tags stop at v0.9.x (2026-08-02)
Discovered while preparing to self-update fmr to the freshly released server-v0.9.2:
**fmr runs v0.11.2** (binary installed 2026-08-02 08:51), but this repo's remote has tags only
up to `server-v0.9.1`, master fast-forwarded cleanly from this machine, there is no v0.10/v0.11
release in Gitea, no source checkout or Go toolchain on fmr, and no deploy script in this repo
that stamps versions. Conclusion: v0.11.2 was cross-built from a clone whose commits were never
pushed — presumably another dev machine.
Consequences until resolved:
- **Do NOT run `--self-update` on fmr.** Gitea's `/releases/latest` is the *newest-created*
release, which is now `server-v0.9.2` — semantically older than the deployed binary; the
updater compares strings, not SemVer, and would happily "update" v0.11.2 down to it. No
automatic risk exists (fmr has no update timer installed, only the cert timer), but a manual
run would downgrade.
- The next real release must be tagged **above v0.11.2** (e.g. `server-v0.11.3` or `v0.12.0`)
*after* the missing commits are pushed, so "latest" becomes truly latest again.
- The unpushed v0.10v0.11 work needs to be found and pushed from whichever machine built it,
or the deployed binary's provenance re-established some other way, before the release channel
can be trusted again.
**Resolved same day.** The binary itself settled it: `go version -m` on the deployed executable
shows `vcs.revision=d5b1bab` — a commit on this repo's master — built 2026-08-02 08:36 UTC with a
hand-stamped `-X main.Version=v0.11.2` and a dirty tree (`vcs.modified=true`, the then-uncommitted
schema doc). An earlier session stamped release numbers ahead of the tag line; no code was ever
missing. Current master is tagged and released as **server-v0.11.3** (signed), restoring a
monotonic, tag-backed lineage above the deployed number. The rule going forward: **the version a
binary is stamped with must be a pushed `server-v*` tag** — an ad-hoc stamp above the tag line
poisons `/releases/latest` for the string-comparing updater the moment anyone tags honestly again.
The stale `server-v0.9.2` release (same code lineage, wrong number, created during the confusion)
remains in Gitea but is harmless now that v0.11.3 outranks it as latest.
## LLDP and CDP are root-tier, and that is a hard boundary (2026-08-02)
Asked for alongside SSDP in long mode; they belong to a different tier and no amount of app-side
cleverness moves them. LLDP is an EtherType `0x88CC` frame to `01:80:C2:00:00:0E`; CDP is an
LLC/SNAP frame to `01:00:0C:CC:CC:CC`. Neither is IP, so neither is ever delivered to a socket an
app can open — receiving them needs `AF_PACKET` with `CAP_NET_RAW`, which is root. Shizuku does
not bridge this either: the ADB shell user (uid 2000) has no `CAP_NET_RAW`, and stock devices do
not ship `tcpdump`. Android's unprivileged ICMP sockets are what make `icmp.ping4` work without
root; there is no equivalent back door for raw L2 receive.
Worth building in the root module when it lands, because the payoff is large: LLDP names the
switch, the port and the VLAN a device is attached to, which is the best available answer to
"where in this building am I actually plugged in", and CDP does the same on Cisco gear. Until
then they are recorded as absent capabilities rather than left to look unimplemented.
What IS reachable at app tier, and what long mode now listens for instead: SSDP (passive NOTIFY
plus periodic M-SEARCH), LLMNR, NetBIOS-NS and WS-Discovery — all IP multicast/broadcast, all
sockets an app may open. The security reading matters as much as the inventory: LLMNR and
NetBIOS-NS being live on a segment is a finding in itself, since both are trivially spoofable.
**One of those four cannot run at app tier either, and says so.** `local.netbios_inventory`
reports `unsupported` with the bind error attached: UDP 137 is below 1024, and Android reserves
privileged ports exactly like any other Linux. The decoder, the evidence shape and the registry
id are built and tested, waiting for the Shizuku tier to supply a socket. Recorded as a result
rather than dropped, so nobody later reads its absence as an oversight.
Two deliberate restraints in that work, both worth keeping: no active WS-Discovery Probe (an
M-SEARCH is traffic every SSDP device expects constantly, whereas a WSD Probe from an unknown
host announces *this* device to the segment), and no NBSTAT sweep (that is host scanning, not
measurement — the app listens to what a network broadcasts, it does not interrogate its
neighbours). The parsers are also hardened against the input they will actually meet: a DNS
compression pointer is refused rather than followed (the classic parser hang), and the WSD
extractor is string-based on purpose, tested against an entity bomb and 20 000-deep nesting.
## Design note: what BLE between two devices is actually for (2026-08-02, not built)
Two or more phones running Echolot, talking over Bluetooth LE. The schema already anticipates
this — `Trigger.PEER` and the whole `peer.*` test family (`peer.reachability`, `peer.isolation`,
`peer.multicast`, `peer.lan_train`, `peer.lease_diff`) are in the registry, unused — and the
prober measured `peer.ble_advertise` **SUPPORTED on both known devices**, so the mechanism is
proven; what has been missing is a reason that beats "use the server".
**The reason is that BLE is out-of-band.** Everything else this app does depends on the network
under test being at least partly functional. A second device reachable over a radio that shares
nothing with the wifi turns several measurements from ambiguous into conclusive:
1. **Client isolation becomes measurable at all.** Today, "I sent a packet to the peer and heard
nothing" cannot distinguish AP client isolation from the peer being asleep, gone, or on a
different VLAN — the failure mode is silence, and silence has too many parents. With BLE the
peer confirms out-of-band that it was listening on address X at time T, so silence over IP
becomes *proof* of isolation rather than a guess. This is the single strongest argument for
the feature, and it mirrors the rule this project keeps rediscovering: a measurement that
cannot separate "nothing happened" from "nothing was tried" is not a measurement.
2. **Differential diagnosis: the network or this phone?** Two devices on the same SSID, one
resolving DNS and one not, settles in seconds what a single device cannot settle at all —
and it is the same distinction `system_verdict` exists to draw, only with a second opinion
instead of Android's. Natural finding: *this device fails where a peer on the same link
succeeds* → look at the device (private DNS, ad blocker, per-client router rule, MAC
randomization), not the router.
3. **Two DHCP servers on one L2**, the classic invisible fault: peers compare lease source,
subnet and gateway (`peer.lease_diff`). Disagreement is conclusive and needs no server.
4. **Coverage and roaming**, later: several devices sampling RSSI in different rooms, exchanging
summaries over BLE, gives a picture no single device standing in one place can produce.
**What crosses the link is a summary, never the document.** A measurement document describes
someone's home network in detail; broadcasting it to whoever is nearby would betray the whole
posture of §8. The peer payload should be: a *hashed* network identity (so two devices can agree
they are on the same L2 without either putting the SSID/BSSID on the air in the clear), the §7.3
category verdicts, the finding codes, and an IP endpoint plus a one-shot nonce for the LAN tests.
Findings and verdicts are already the interpretation layer — exactly the right granularity to
share.
**Privacy constraints, which are not optional here.** A BLE advertiser is a tracking beacon: it
must be user-initiated, time-boxed to the run, carry no identifier that is stable across runs
(the resolvable-private-address default plus a per-session ephemeral id), and pair by a code the
two humans can see. "Discoverable by default" would make this app a worse citizen than the
networks it audits.
**Deliberately not doing:** clock synchronisation over BLE. GATT latency is jitter measured in
tens of milliseconds, which is the same order as the one-way delays worth measuring; peers should
sync against the server's `time.server_offset` and use BLE only to correlate run ids. Nor should
BLE become a transport for uploads — it is a *comparison* channel.
Staging when it happens: `peer.isolation` first (highest value, needs only advertise + connect +
a nonce exchange), then `peer.lease_diff` (pure summary comparison, no extra plumbing), then the
rest. Needs `BLUETOOTH_ADVERTISE/CONNECT/SCAN` in the manifest, which the app does not yet
request.
## v0.11.3 live on fmr; trains validated end to end (2026-08-02)
Deployed via `--self-update` (the pre-signing v0.11.2 updater accepted the first signed release,
as planned; every later update verifies). Startup clean on the real host — the reserved-port
check passed against the OS, self-test green, capabilities unchanged plus the new machinery.
**`train.udp_updown` validated against production**: `LiveUpstreamTrainTest` from this PC sent
120 packets; the server's ledger counted 120, both columnar report parts arrived, loss 0.0 %,
`truncated=false`. The 0x03/0x04/0x05 path works, client and server, over the real internet.
Two operational bugs surfaced doing it:
- **CLI-minted tokens are lost while the daemon runs.** `devices.json` is loaded once at startup
and held in memory; `--mint-enroll-token` writes to disk, the running daemon never re-reads,
answers "unknown token", and clobbers the token on its next write. `enroll-link.sh` has only
ever worked by timing luck. Workaround used: mint, `systemctl restart echolot-server`, then
redeem. Real fix belongs server-side (re-read on miss, or route the CLI mint through the
running daemon).
- **`test-fmr.sh` still mints against `127.0.0.1:8444`**, which no longer exists (the admin API
moved to authenticated :443). Needs the same CLI-mint flow enroll-link.sh uses — plus the
restart caveat above until that bug is fixed.
+133
View File
@@ -0,0 +1,133 @@
<!--
SPDX-FileCopyrightText: 2026 Echolot contributors
SPDX-License-Identifier: CC-BY-4.0
-->
# Echolot findings registry
Closes open item 1 of `measurement-schema.md` §9.
A **finding code** is the stable, machine-readable half of a result. The prose around it changes
freely; the code is what a dashboard groups by, what a diff between two runs keys on, and what
someone greps a year of archived runs for. That only works if a code means exactly one thing,
forever.
This document is the contract. It is kept in step with
`echolot-app/core-measurement/.../FindingRegistry.kt` by a test that fails when either side has a
code the other does not — a registry that drifts from its documentation is worse than none,
because it looks authoritative.
## Rules
1. **The prefix determines the category**, and the category determines which verdict light the
finding rolls up into (§7.3). A `nat.*` code appearing under *connectivity* is not a naming
quibble; it changes which light turns red. Two codes were renamed from `nat.*` to
`connectivity.*` for exactly this reason.
2. **One code per concept.** Two emitters independently produced `connectivity.downstream_loss`
and `connectivity.loss_downstream` for the same claim before this registry existed. Anyone
aggregating either would have silently seen half their data.
3. **Codes are declared, not typed.** Emitters reference a `FindingSpec`, so a typo is a compile
error and no two call sites can disagree about a finding's category or default severity.
4. **Severity in the registry is the default.** An emitter may escalate for a specific run; it may
not quietly reclassify the finding in general.
5. **Say what is ruled out**, where that is the useful half. "Loss upstream" is worth far more
when it also states that the return path is clean, because that halves where to look next.
6. **Renaming a code is a breaking change** once runs are archived at scale. Before 1.0 it is
cheap; after, it needs an alias and a deprecation window.
## Registry
### connectivity
| code | severity | means | rules out |
|---|---|---|---|
| `connectivity.udp_unreachable` | high | No UDP echo replies came back from the server at all. | — |
| `connectivity.udp_unreachable_upstream` | high | The server received none of the probes, so traffic is dropped on the way out. | The return path: nothing arrived to be replied to. |
| `connectivity.udp_loss` | medium | A large fraction of round-trip probes were lost, direction unknown. | — |
| `connectivity.loss_upstream` | medium | Probes were lost on the way to the server. | The return path: replies came back for everything that arrived. |
| `connectivity.loss_downstream` | medium | Packets were lost on the way back from the server. | The outbound path: the server received what it was answering. |
| `connectivity.downstream_blocked` | high | Server-initiated packets never arrive, although round trips work. | Basic reachability: the path forwards replies, just not unsolicited traffic. |
| `connectivity.downstream_reorder` | low | Downstream packets arrive in a different order than they were sent. | — |
| `connectivity.captive_portal` | medium | A captive portal is intercepting connectivity checks. | — |
| `connectivity.no_internet` | high | Android's own connectivity checks fail on this network. | — |
| `connectivity.link_flapping` | medium | A network dropped and came back one or more times during the run. | A momentary probe failure: the drop was watched happening, not inferred from silence. |
`connectivity.link_flapping` is only reachable from a **long run** (`run.mode: "long"`,
measurement-schema §3). It is derived from `networks[].changes[]` rather than from any test's
evidence, because no one-shot probe can produce it: the probes before and after a four-second drop
both succeed. The emitter escalates to *high* from three completed drop-and-return cycles, and
requires the cycle to complete — a network switched off partway through a run is not flapping.
### mtu
| code | severity | means | rules out |
|---|---|---|---|
| `mtu.reduced_downstream` | low | The downstream path MTU is below the usual 1500 bytes. | — |
| `mtu.downstream_blackhole` | medium | Datagrams above the path MTU are dropped downstream, fragmented or not. | — |
| `mtu.fragments_blocked` | medium | IP fragments do not reach this device even when sent in order. | — |
| `mtu.fragment_reorder_sensitive` | low | Fragments are delivered in order but dropped when reordered or delayed. | Fragmentation itself: in-order fragments arrive fine. |
### nat
| code | severity | means | rules out |
|---|---|---|---|
| `nat.udp_rebinding` | medium | A NAT remapped the UDP source port mid-flow. | — |
| `nat.symmetric` | medium | The NAT assigns a different external port per destination. | — |
### perf
| code | severity | means | rules out |
|---|---|---|---|
| `perf.throughput_no_delivery` | high | No throughput traffic arrived, although the server sent it. | — |
| `perf.throughput_below_offered` | low | Less throughput arrived than the server sent for the whole run. | — |
### dns
| code | severity | means | rules out |
|---|---|---|---|
| `dns.answer_rewritten` | high | A resolver returned an answer that differs from the authoritative record. | — |
| `dns.authoritative_unreachable` | medium | The canary zone's authoritative server could not be reached. | — |
### v6
The prefix is `v6.`, matching the test-type registry (`v6.brokenness`, `v6.happy_eyeballs`, …).
These were `ipv6.*` while declaring `Category.IPV6`; since the prefix map only knows `v6`, they
rolled up under *connectivity* instead — the third occurrence of rule 1 being broken.
| code | severity | means | rules out |
|---|---|---|---|
| `dns.search_domain_unanswered` | high | The network advertises a DNS search domain that its own server does not answer for. | A fault on this device: the same server answers ordinary names normally. |
| `dns.system_resolver_broken` | high | The network's DNS server answers, but this device cannot resolve names through it. | A network fault: the server replied to a query sent from this device. |
| `measurement.vpn_constrained` | info | A VPN was active, so the networks underneath it could not be measured. | Nothing — this run says little about the underlying network either way. |
| `v6.no_default_route` | medium | The device has a global IPv6 address but no IPv6 default route. | Guesswork: this is read from the routing table, not inferred from silence. |
| `v6.route_without_address` | medium | The network advertises an IPv6 default route but the device has no global IPv6 address. | A working IPv6 setup: SLAAC did not produce a usable address on this link. |
| `v6.no_icmp_reply` | low | IPv6 is configured but ICMPv6 echo gets no reply. | Nothing on its own: IPv6 may work fine with ICMP filtered. |
| `v6.broken` | high | IPv6 is advertised on this network but carries no traffic. | ICMP filtering as the benign explanation: a TCP connection over IPv6 failed too. |
| `v6.not_offered` | info | This network does not offer IPv6. | — |
`v6.no_icmp_reply` was `v6.broken` until a phone reported it while loading an IPv6-only site over
TCP perfectly well. The only evidence behind it is ICMPv6 echo, which is widely filtered on
networks where IPv6 works — so the finding now states what was observed and names both
explanations instead of choosing one. It is still worth reporting: filtered ICMPv6 breaks Path MTU
Discovery.
`v6.broken` returned once that corroboration existed: the `v6.brokenness` test attempts a real TCP
connection over IPv6 to the configured server, and only when *both* transports fail on a network
that advertises IPv6 is the brokenness claim made — at high severity, because every dual-stack
destination pays a timeout before falling back to IPv4. When the TCP connect *succeeds*,
`v6.no_icmp_reply` is emitted at high confidence instead, now able to say plainly that ICMPv6 is
filtered while IPv6 works. With no server configured there is no corroboration target and the
two-explanation `v6.no_icmp_reply` stands unchanged.
`v6.not_offered` is **info and must stay info**. Most networks still do not offer IPv6 and that is
not a fault; reporting it as a warning lights a yellow verdict on a healthy network, which teaches
people to ignore the light — the one thing a diagnostic must never do.
## Adding a finding
1. Add a `FindingSpec` to `FindingRegistry`, and to its `all` list.
2. Add the row here, under the section its prefix names.
3. Emit it with `finding(FindingRegistry.YOUR_CODE, …)`.
The registry test checks 1 and 2 agree, that every prefix maps to the category it claims, and that
no two entries share a code.
+52 -3
View File
@@ -38,6 +38,7 @@ Export encoding: UTF-8 JSON, gzip for files (`.echolot.json.gz`), share intent u
{ {
"id": "0198c5f2-...-uuidv7", "id": "0198c5f2-...-uuidv7",
"trigger": "manual | scheduled | monitor | peer", "trigger": "manual | scheduled | monitor | peer",
"mode": "short | long",
"started_at": "2026-07-29T14:03:21.114Z", "started_at": "2026-07-29T14:03:21.114Z",
"ended_at": "2026-07-29T14:07:44.902Z", "ended_at": "2026-07-29T14:07:44.902Z",
"clock": { "clock": {
@@ -52,12 +53,47 @@ Export encoding: UTF-8 JSON, gzip for files (`.echolot.json.gz`), share intent u
}, },
"tiers": { "app": true, "shizuku": true, "root": false }, "tiers": { "app": true, "shizuku": true, "root": false },
"profiles_used": ["profile-uuid", ...], "profiles_used": ["profile-uuid", ...],
"constraints": {
"vpn_active": true,
"per_network_blocked": true,
"unmeasured_networks": ["net-0", "net-1"]
},
"notes": "free-text user annotation" "notes": "free-text user annotation"
} }
``` ```
`tiers` records what was *available*; each test records what it *used*. `tiers` records what was *available*; each test records what it *used*.
`mode` records how long the run watched, and it exists because **it changes what a reader may
conclude from absence**. A `short` run is a sequence of one-shot probes — each looks at the network
for a second or two and moves on — which characterises the network's *configuration* well and is
structurally blind to anything intermittent. A `long` run starts continuous listeners at t=0, runs
the same battery beside them, and keeps sampling until its window closes; the window's length is
recorded in the `params` of the tests the listeners produce, not here.
The consequence is asymmetric and matters more than the field looks. A finding is worth the same in
either mode: a drop that was observed, was observed. Silence is not. "No link changes were seen" is
evidence of a stable link after five minutes of watching and is evidence of nothing at all after a
thirty-second run, in which a link could drop and return between two consecutive probes without
leaving a mark anywhere in the document. Consumers — a diff between two runs, a dashboard counting
how often a fault occurs, a person reading one report — must therefore not treat the absence of a
time-dependent finding in a `short` run as its refutation, and must not compare the two modes as if
they had asked the same question. `connectivity.link_flapping` is the first finding that only a
`long` run can reach; `networks[].changes[]` (§4) is likewise populated only by a long run's
listener, and an empty `changes[]` in a short run means "not watched", never "nothing happened".
Absent `mode` means `short`: it was added after the first documents were written, and every one of
them was a battery of one-shot probes.
`constraints` records what was *prevented*. A constrained run is neither a failed run nor a normal
one, and the distinction has to survive into the data: a run taken through a VPN has the same shape
and the same green verdict as a clean run of a healthy network, so without this a reader — or a
server aggregating thousands of them — cannot tell that almost nothing was measured. The known case
is `per_network_blocked`: Android refuses `Network.bindSocket()` on the underlying networks while a
VPN holds the default route, so every per-network test measures the tunnel or nothing at all, and
any conclusion about the link underneath is unfounded. Consumers should treat findings from a
constrained run as scoped to what was actually reachable, and `unmeasured_networks` names the rest.
## 4. `networks[]` — one entry per Android `Network` in play ## 4. `networks[]` — one entry per Android `Network` in play
A run may exercise several networks simultaneously (Wi-Fi + cellular + USB ethernet). Everything is a snapshot at run start; a `changes[]` list captures mid-run deltas. A run may exercise several networks simultaneously (Wi-Fi + cellular + USB ethernet). Everything is a snapshot at run start; a `changes[]` list captures mid-run deltas.
@@ -97,12 +133,24 @@ A run may exercise several networks simultaneously (Wi-Fi + cellular + USB ether
"changes": [ "changes": [
{ "at_mono_ns": 91000000000, "kind": "lost | gained | link_changed", { "at_mono_ns": 91000000000, "kind": "lost | gained | link_changed",
"detail": { /* new link snapshot or diff */ } } "detail": { /* new link snapshot or diff */ } }
] ],
"app_usable": true
} }
``` ```
`routes[].proto` and lifetime fields are Shizuku-tier data (`ip route`/`ip addr`); app-tier snapshots leave them absent — absence means "not observed", never "not present". `routes[].proto` and lifetime fields are Shizuku-tier data (`ip route`/`ip addr`); app-tier snapshots leave them absent — absence means "not observed", never "not present".
`app_usable` records whether an ordinary app may send on this network at all. Android lists the
carrier's special-purpose networks — IMS/VoLTE, MMS, XCAP — alongside the real ones, and they
carry neither `INTERNET` nor `NOT_RESTRICTED`; binding to one needs
`CONNECTIVITY_USE_RESTRICTED_NETWORKS`, which is signature-level and unobtainable for a normal
app. Those networks are therefore permanently unmeasurable, and that is a property of Android's
permission model rather than of the link. They stay in `networks[]` because they are genuinely
present — an interface silently missing from the inventory is its own kind of lie — but a
consumer must not read the absence of tests against them as a fault, and they are **not**
`constraints.unmeasured_networks` (§3): nothing was prevented, the run was never entitled to
measure them.
## 5. `server_sessions[]` ## 5. `server_sessions[]`
```json ```json
@@ -269,7 +317,7 @@ The JSON Schema (machine-readable companion, `measurement.schema.json`, generate
| type | example fields | v2 anonymizer transform | | type | example fields | v2 anonymizer transform |
|---|---|---| |---|---|---|
| `ip4`, `ip6` | addresses, routes, hops, DNS answers | prefix-preserving pseudonymization, consistent per document; well-known/reserved ranges kept verbatim | | `ip4`, `ip6` | addresses, routes, hops, DNS answers | prefix-preserving pseudonymization, consistent per document; well-known/reserved ranges kept verbatim. **Exception: ULA (`fc00::/7`) has its whole prefix pseudonymized as a unit.** It resembles RFC1918 but is not analogous: a ULA global ID is 40 random bits, unique to one network by construction (RFC 4193), so the prefix *is* the identifier, whereas `192.168.0.0/16` is shared by millions of networks and identifies none. Pseudonymizing it as a unit keeps "these hosts are on one subnet" while dropping "this is that subnet". |
| `mac`, `bssid` | wifi, arp_watch | OUI kept, NIC part pseudonymized | | `mac`, `bssid` | wifi, arp_watch | OUI kept, NIC part pseudonymized |
| `fqdn` | DNS names, reverse lookups | per-label pseudonyms, public-suffix kept | | `fqdn` | DNS names, reverse lookups | per-label pseudonyms, public-suffix kept |
| `ssid` | wifi | pseudonym | | `ssid` | wifi | pseudonym |
@@ -279,7 +327,8 @@ Free-text fields (`notes`, `error.detail`, dump excerpts from Shizuku parsers) c
## 9. Open items ## 9. Open items
1. Findings registry document — start alongside the first implemented tests. 1. ~~Findings registry document~~ — done: `findings-registry.md`, kept in step with
`FindingRegistry.kt` by a test that fails when the two disagree.
2. Whether Shizuku raw-dump excerpts (dumpsys/ip output) are embedded in `evidence` verbatim (auditable, but large and hard to anonymize) or parsed-only with an optional "attach raw dumps" toggle. Proposal: toggle, default on for local archive, default off for export. 2. Whether Shizuku raw-dump excerpts (dumpsys/ip output) are embedded in `evidence` verbatim (auditable, but large and hard to anonymize) or parsed-only with an optional "attach raw dumps" toggle. Proposal: toggle, default on for local archive, default off for export.
3. Peer-mode documents: each device produces its own run; the coordinator embeds the peer's findings summary and cross-references by `run.id`. Full merge format deferred. 3. Peer-mode documents: each device produces its own run; the coordinator embeds the peer's findings summary and cross-references by `run.id`. Full merge format deferred.
4. Size guardrails: soft cap 20 MB uncompressed per run; trains beyond that downsample evidence (keep aggregates + first/last N + all anomalies) and record `"evidence_truncated": true`. 4. Size guardrails: soft cap 20 MB uncompressed per run; trains beyond that downsample evidence (keep aggregates + first/last N + all anomalies) and record `"evidence_truncated": true`.
+57 -4
View File
@@ -24,11 +24,34 @@ echolot://enroll?v=1&u=<control-URL, urlencoded>&p=pin-sha256:<b64 SPKI hash>&t=
``` ```
POST /v1/enroll Authorization: Bearer <enrollment-token> POST /v1/enroll Authorization: Bearer <enrollment-token>
→ 200 { "device_credential": "<random 256-bit, b64url>", → 201 { "device_credential": "<random 256-bit, b64url>",
"device_id": "uuid", "device_id": "uuid" }
"profile": { ... §2.2 ... } }
``` ```
The **server assembles the bootstrap link**, because it is the only party holding all three parts
at once, and the part an operator gets wrong by hand is the base64 pin — which does not fail
loudly, it just never matches, and surfaces later as an inscrutable TLS error:
```
POST /admin/enroll-tokens
→ { "token": "…", "expires_in_s": 86400,
"enroll_uri": "echolot://enroll?v=1&u=…&p=…&t=…" }
```
The control URL in the link comes from `ECHOLOT_PUBLIC_URL`, falling back to the first control
listen address. A wildcard bind has no single right answer, so it warns rather than guessing.
Encoding notes that matter in practice:
- `u`, `p` and `t` are **percent-encoded**. The pin is base64, so it contains `+`, `/` and `=`,
every one of which means something else in a query string.
- A `+` that was *not* encoded decodes to a space. Base64 contains no spaces, so a parser SHOULD
restore them — the alternative is a pin wrong by one character and a failure that points nowhere
near the cause.
- The control URL MUST be `https://`. The pin only protects a TLS connection; a cleartext URL
would hand the token to anyone on the path.
- **The link is a secret** while it is live: it carries a bearer token, so anyone who sees it
before the device does can enroll instead.
Enrollment tokens are single-use with expiry, created in the admin UI, scoped `enroll`. The device credential is a long-lived bearer secret, scoped `run-tests`; it is also the HKDF input for session keys. Revocation = deleting the device in the admin UI. Enrollment tokens are single-use with expiry, created in the admin UI, scoped `enroll`. The device credential is a long-lived bearer secret, scoped `run-tests`; it is also the HKDF input for session keys. Revocation = deleting the device in the admin UI.
### 2.2 Profile ### 2.2 Profile
@@ -61,7 +84,12 @@ The app re-fetches the profile at the start of every run (falling back to the ca
### 2.3 Capabilities (v1 registry) ### 2.3 Capabilities (v1 registry)
`udp-probe`, `stun-basic`, `stun-5780`, `canary-dns`, `recursive-dns`, `connect-back`, `delayed-echo`, `big-send`, `frag-send`, `tls-echo`, `http-echo`, `throughput`, `ntp`. A server omits what it can't offer (e.g. `stun-5780` without a second IP degrades to `stun-basic`). Clients must skip, and record as `unsupported`, any test whose capability is absent. Unknown capability strings are ignored. `udp-probe`, `stun-basic`, `stun-5780`, `canary-dns`, `recursive-dns`, `connect-back`, `delayed-echo`, `big-send`, `frag-send`, `tls-echo`, `http-echo`, `throughput`, `ntp`, plus:
- `downtrain` — server-sent downstream trains via the §5 `downtrain` action. Upstream trains need no capability of their own: they are plain client-sent data-plane packets and ride `udp-probe`.
- `tcp-echo` — the plain-TCP echo endpoint (§4); `tls-echo` is its ALPN variant on the same port.
A server omits what it can't offer (e.g. `stun-5780` without a second IP degrades to `stun-basic`). Clients must skip, and record as `unsupported`, any test whose capability is absent. Unknown capability strings are ignored.
### 2.4 Sessions ### 2.4 Sessions
@@ -88,6 +116,31 @@ DELETE /v1/sessions/{id}
Per-credential and per-source-IP token buckets on: session creation, actions, UDP packets, bytes. `429` on control plane; silent drop on data plane (probes must tolerate loss anyway). All reflected/generated traffic goes **only** to the session's observed source address (or, for connect-back, the source address of the session-creating request). Data-plane responses to unauthenticated packets are never larger than the request (§3.4). Per-credential and per-source-IP token buckets on: session creation, actions, UDP packets, bytes. `429` on control plane; silent drop on data plane (probes must tolerate loss anyway). All reflected/generated traffic goes **only** to the session's observed source address (or, for connect-back, the source address of the session-creating request). Data-plane responses to unauthenticated packets are never larger than the request (§3.4).
### 2.2 `GET /v1/discover` — where the control plane lives
Unauthenticated, and says almost nothing: the control-plane URL and the server's display name.
```json
{ "control_url": "https://probe.example.net", "name": "example" }
```
It exists so an enrollment link can carry the name a person recognises while the app still connects
to the name that selects the pinned certificate. When a server shares port 443 between its admin UI
and its control plane, those must be different hostnames — one port and one name is one certificate,
and the two need different ones (a browser-trusted certificate, and a long-lived self-signed one the
client pins). Without discovery, the difference leaks into every enrollment link an operator hands
out.
**It hands out an address, never a pin.** The pin travels in the link itself. Serving it here would
reduce pinning to whatever the certificate authorities are worth, and pinning exists precisely to
survive one the operator does not control — a root injected by corporate device management, for
instance, which is unremarkable on the networks this tool is pointed at. Because the pin is
pre-shared, an intercepted discovery response can only send a device to the wrong host, where the
pin will not match: an outage, not a compromise.
Clients treat it as optional. A server that does not answer, or a link that already names the
control endpoint, works unchanged — enrollment must not begin failing because a lookup did.
## 3. UDP probe protocol ## 3. UDP probe protocol
### 3.1 Packet header (fixed 32 bytes, network byte order) ### 3.1 Packet header (fixed 32 bytes, network byte order)
+11 -3
View File
@@ -15,7 +15,7 @@ plugins {
// //
// major*1_000_000 + minor*10_000 + patch*10 leaves room for 9 patch-level rebuilds (the trailing // major*1_000_000 + minor*10_000 + patch*10 leaves room for 9 patch-level rebuilds (the trailing
// digit) without disturbing the mapping, and stays inside the 2_100_000_000 ceiling until major 2100. // digit) without disturbing the mapping, and stays inside the 2_100_000_000 ceiling until major 2100.
val appVersionName = "0.2.0" val appVersionName = "0.3.0"
fun versionCodeOf(semver: String): Int { fun versionCodeOf(semver: String): Int {
val (major, minor, patch) = semver.substringBefore('-').split(".").map(String::toInt) val (major, minor, patch) = semver.substringBefore('-').split(".").map(String::toInt)
@@ -34,8 +34,16 @@ android {
versionName = appVersionName versionName = appVersionName
// Automation: `adb shell am start -n app.echo_lot.app/.MainActivity --ez autorun true` // Automation: `adb shell am start -n app.echo_lot.app/.MainActivity --ez autorun true`
// runs a measurement immediately and POSTs the report here (dev collection endpoint). // runs a measurement immediately and POSTs the report here (dev collection endpoint).
buildConfigField("String", "REPORT_UPLOAD_URL", "\"http://89.185.109.150:443/report\"") // Empty: the collection endpoint this pointed at was the adb-beacon receiver, which held
buildConfigField("String", "REPORT_UPLOAD_SECRET", "\"D4OmG5gGJsElqVVbtYIZbR\"") // 0.0.0.0:443 in cleartext. That service is gone and echolot-server owns 443 with TLS, so
// posting plaintext there now fails as "client sent an HTTP request to an HTTPS server" —
// an alarming error for a debugging convenience that is no longer needed, since autorun
// reports are read straight off the device with `run-as cat`.
//
// Deliberately not repointed at /v1/runs. That is the consent-gated upload, and a
// debugging shortcut must not be able to satisfy it by accident.
buildConfigField("String", "REPORT_UPLOAD_URL", "\"\"")
buildConfigField("String", "REPORT_UPLOAD_SECRET", "\"\"")
// The bare SemVer, without the debug build's "-dev" suffix stripped away by the server's // The bare SemVer, without the debug build's "-dev" suffix stripped away by the server's
// parser anyway — sent to servers so they can apply their compatibility window. // parser anyway — sent to servers so they can apply their compatibility window.
buildConfigField("String", "APP_SEMVER", "\"$appVersionName\"") buildConfigField("String", "APP_SEMVER", "\"$appVersionName\"")
+33 -1
View File
@@ -10,6 +10,17 @@
<uses-permission android:name="android.permission.CHANGE_NETWORK_STATE" /> <uses-permission android:name="android.permission.CHANGE_NETWORK_STATE" />
<uses-permission android:name="android.permission.CHANGE_WIFI_MULTICAST_STATE" /> <uses-permission android:name="android.permission.CHANGE_WIFI_MULTICAST_STATE" />
<uses-permission android:name="android.permission.ACCESS_FINE_LOCATION" /> <uses-permission android:name="android.permission.ACCESS_FINE_LOCATION" />
<!--
The adb relay (AdbRelayService) only. It runs in the foreground because it must keep
watching while the tablet sits unattended with its screen off, and dataSync is the type
that describes it: it carries an observation off the LAN, nothing more. Android 15 caps
dataSync at a few hours a day, which is acceptable for a tool that is switched on for a
debugging session rather than left running forever.
-->
<uses-permission android:name="android.permission.FOREGROUND_SERVICE" />
<uses-permission android:name="android.permission.FOREGROUND_SERVICE_DATA_SYNC" />
<!-- Only so the relay's ongoing status is visible; the service runs either way. -->
<uses-permission android:name="android.permission.POST_NOTIFICATIONS" />
<application <application
android:allowBackup="false" android:allowBackup="false"
@@ -22,7 +33,8 @@
<activity <activity
android:name=".MainActivity" android:name=".MainActivity"
android:exported="true"> android:exported="true"
android:launchMode="singleTask">
<intent-filter> <intent-filter>
<action android:name="android.intent.action.MAIN" /> <action android:name="android.intent.action.MAIN" />
<category android:name="android.intent.category.LAUNCHER" /> <category android:name="android.intent.category.LAUNCHER" />
@@ -39,8 +51,28 @@
<category android:name="android.intent.category.BROWSABLE" /> <category android:name="android.intent.category.BROWSABLE" />
<data android:scheme="echolot" android:host="enroll" /> <data android:scheme="echolot" android:host="enroll" />
</intent-filter> </intent-filter>
<!--
Sign-in redirect. The browser hands the authorization code back through this, which
is exactly why the flow uses PKCE: any app may register this scheme, so the code
alone must not be enough to complete a sign-in.
-->
<intent-filter android:autoVerify="false">
<action android:name="android.intent.action.VIEW" />
<category android:name="android.intent.category.DEFAULT" />
<category android:name="android.intent.category.BROWSABLE" />
<data android:scheme="echolot" android:host="auth" />
</intent-filter>
</activity> </activity>
<!--
Not exported: nothing outside this app has any business starting a relay that reports
where this device can be reached.
-->
<service
android:name=".AdbRelayService"
android:exported="false"
android:foregroundServiceType="dataSync" />
<provider <provider
android:name="androidx.core.content.FileProvider" android:name="androidx.core.content.FileProvider"
android:authorities="${applicationId}.fileprovider" android:authorities="${applicationId}.fileprovider"
@@ -0,0 +1,125 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.app
import app.echo_lot.protocol.AuthInfo
import app.echo_lot.protocol.ControlClient
import app.echo_lot.protocol.OidcLogin
import kotlinx.serialization.json.Json
import kotlinx.serialization.json.jsonObject
import kotlinx.serialization.json.jsonPrimitive
/**
* Signing in to the configured server's identity provider.
*
* The awkward part of a browser-based sign-in on Android is that the app is not running while it
* happens. Handing control to a browser puts this process in the background, where it may be
* killed at any moment; the callback then arrives at a fresh process with none of the state the
* exchange needs. So the PKCE verifier and state are written to storage before the browser opens,
* not held in memory — an in-memory value works on a developer's device and fails on a phone under
* memory pressure, which is the worst way for this to break.
*
* Nothing from the identity provider is kept afterwards. The ID token proves who is signing in,
* once; the device credential authenticates everything from then on.
*/
class Account(private val settings: Settings) {
private val json = Json { ignoreUnknownKeys = true }
sealed interface SignInStart {
/** Open this in a browser. */
data class Browser(val url: String) : SignInStart
data class Unavailable(val reason: String) : SignInStart
}
/** Fetches the server's auth configuration and builds the authorization URL. */
fun begin(): SignInStart {
if (!settings.serverConfigured) {
return SignInStart.Unavailable(
"Enrol with a server first — sign-in belongs to the server's identity provider."
)
}
val auth = runCatching { client().profile(settings.serverCredential).auth }.getOrNull()
?: return SignInStart.Unavailable("Could not reach the server to ask how to sign in.")
auth.discoveryError?.let {
// The distinction matters: "the operator configured an IdP that is not answering" is
// their problem to fix, and is not the same as "this server has no accounts".
return SignInStart.Unavailable("The server's identity provider is not responding: $it")
}
if (!auth.enabled) {
return SignInStart.Unavailable("This server does not offer accounts.")
}
return try {
val pending = OidcLogin.begin(auth)
// Written before the browser opens, because after that this process may not survive.
settings.pendingVerifier = pending.verifier
settings.pendingState = pending.state
SignInStart.Browser(pending.authorizationUrl)
} catch (t: Throwable) {
SignInStart.Unavailable(t.message ?: "Could not start sign-in.")
}
}
/**
* Completes sign-in from the `echolot://auth` redirect.
*
* Blocking; callers run it off the main thread.
*/
fun complete(callbackUri: String): String {
val verifier = settings.pendingVerifier
val state = settings.pendingState
// Cleared first, whatever happens next: these are single-use, and leaving them behind
// would let a later callback be completed against a flow nobody started.
settings.clearPendingAuth()
if (verifier.isBlank() || state.isBlank()) {
return "That sign-in did not start on this device."
}
return try {
val auth = client().profile(settings.serverCredential).auth
val idToken = OidcLogin.complete(
auth, OidcLogin.Pending("", verifier, state), callbackUri,
)
val reply = client().linkAccount(settings.serverCredential, idToken)
val o = json.parseToJsonElement(reply).jsonObject
val name = o["display_name"]?.jsonPrimitive?.content ?: "signed in"
settings.accountName = name
settings.accountId = o["account_id"]?.jsonPrimitive?.content ?: ""
val admin = o["admin"]?.jsonPrimitive?.content == "true"
"Signed in as $name" + if (admin) " (administrator)" else ""
} catch (e: OidcLogin.LoginFailed) {
e.message ?: "Sign-in failed."
} catch (t: Throwable) {
"Sign-in failed: ${t.message ?: t.javaClass.simpleName}"
}
}
/** Signs out. The device stays enrolled — signing out should not cost an enrolment. */
fun signOut(): String = try {
client().unlinkAccount(settings.serverCredential)
settings.accountName = ""
settings.accountId = ""
"Signed out. This device is still enrolled."
} catch (t: Throwable) {
"Could not sign out: ${t.message ?: t.javaClass.simpleName}"
}
/** Asks the server who it thinks is signed in, so the UI is not trusting stale local state. */
fun refresh(): String? = runCatching {
val o = json.parseToJsonElement(client().accountStatus(settings.serverCredential)).jsonObject
val signedIn = o["signed_in"]?.jsonPrimitive?.content == "true"
settings.accountName = if (signedIn) {
o["display_name"]?.jsonPrimitive?.content ?: ""
} else {
""
}
settings.accountName.takeIf { it.isNotBlank() }
}.getOrNull()
private fun client() = ControlClient(
settings.serverUrl, setOf(settings.serverPin), BuildConfig.APP_SEMVER,
fallbackAddrs = settings.serverAddrList(),
)
}
@@ -0,0 +1,136 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.app
import android.content.Context
import android.net.nsd.NsdManager
import android.net.nsd.NsdServiceInfo
import android.net.wifi.WifiManager
import java.net.Inet4Address
/**
* Watches adbd's own mDNS advertisement on this device and reports the endpoint to the server.
*
* This replaces the retired Python adb-beacon, and exists for one reason: **mDNS does not cross
* subnets**. A developer on another network cannot see `_adb-tls-connect._tcp` at all, while the
* wireless-debug port rotates every few minutes — so the port has to be carried out of the LAN by
* something sitting inside it. That is this. A tablet parked on the test network relays; the
* developer reads the endpoint back from the server.
*
* Two hard-won rules from the beacon, both load-bearing:
*
* - **Resolve each service instance exactly ONCE.** Resolving adbd's advertisement makes adbd
* re-arm its connection and post a "wireless debugging connected" notification; re-resolving on
* every heartbeat turns that into a stream of them. The guard is cleared only when the service
* is *lost*, which is also what catches rotation: the new advertisement is a new instance, gets
* resolved once, and is reported within seconds.
* - **Do not run this on the OnePlus.** On network churn that device drops and re-publishes its
* advertisement repeatedly, so lost/found cycles keep clearing the guard and each resolve
* re-arms adbd. Guarding reduces but cannot eliminate the noise; the Lenovo tablet is the
* intended host, which is also why relaying is a mode rather than something always on.
*/
class AdbRelay(
private val ctx: Context,
private val onEvent: (String) -> Unit,
) {
private val nsd = ctx.getSystemService(NsdManager::class.java)
/** Instances already resolved, by service name — the re-arm guard described above. */
private val resolved = HashSet<String>()
/** Last endpoint reported, so the heartbeat re-posts from cache instead of re-resolving. */
@Volatile var lastEndpoint: Endpoint? = null
private set
data class Endpoint(val host: String, val port: Int, val serviceName: String)
private var listener: NsdManager.DiscoveryListener? = null
fun start() {
if (nsd == null) {
onEvent("mDNS unavailable on this device")
return
}
if (listener != null) return
val l = object : NsdManager.DiscoveryListener {
override fun onStartDiscoveryFailed(type: String?, code: Int) {
onEvent("discovery failed to start (code $code)")
}
override fun onStopDiscoveryFailed(type: String?, code: Int) {}
override fun onDiscoveryStarted(type: String?) {
onEvent("watching for adbd on this network")
}
override fun onDiscoveryStopped(type: String?) {}
override fun onServiceFound(info: NsdServiceInfo?) {
val name = info?.serviceName ?: return
// The guard: one resolve per instance, ever. adbd re-arms on every resolve.
if (!resolved.add(name)) return
resolve(info)
}
override fun onServiceLost(info: NsdServiceInfo?) {
// Rotation: the old instance is gone, so allow the replacement to be resolved.
info?.serviceName?.let { resolved.remove(it) }
}
}
listener = l
runCatching { nsd.discoverServices(ADB_SERVICE, NsdManager.PROTOCOL_DNS_SD, l) }
.onFailure { onEvent("could not start discovery: ${it.message}") }
}
fun stop() {
listener?.let { l -> runCatching { nsd?.stopServiceDiscovery(l) } }
listener = null
resolved.clear()
}
@Suppress("DEPRECATION") // the callback-based resolve is the one that exists across our minSdk
private fun resolve(info: NsdServiceInfo) {
val cb = object : NsdManager.ResolveListener {
override fun onResolveFailed(i: NsdServiceInfo?, code: Int) {
// Let a failed instance be retried: the guard exists to stop *successful*
// re-resolution, not to give up on a transient failure.
i?.serviceName?.let { resolved.remove(it) }
onEvent("resolve failed (code $code)")
}
override fun onServiceResolved(i: NsdServiceInfo?) {
val host = i?.host?.hostAddress ?: return
// adbd advertises on every address it listens on, including link-local v6. The
// reachable one from a developer's subnet is the routable v4 address, and it is
// also the only one worth relaying — a link-local address means nothing off-link.
if (i.host !is Inet4Address) return
// Only this device's own advertisement: on a shared network several phones may
// have wireless debugging on, and relaying a neighbour's port would send a
// developer to the wrong device.
if (host != localIp()) return
lastEndpoint = Endpoint(host, i.port, i.serviceName ?: "adb")
onEvent("found adbd at $host:${i.port}")
}
}
runCatching { nsd?.resolveService(info, cb) }
.onFailure {
resolved.remove(info.serviceName)
onEvent("resolve threw: ${it.message}")
}
}
/** This device's own IPv4 address on the wifi it is relaying from. */
private fun localIp(): String? = runCatching {
val wifi = ctx.getSystemService(WifiManager::class.java) ?: return null
@Suppress("DEPRECATION")
val ip = wifi.connectionInfo.ipAddress
if (ip == 0) return null
@Suppress("DEPRECATION")
String.format(
"%d.%d.%d.%d",
ip and 0xff, ip shr 8 and 0xff, ip shr 16 and 0xff, ip shr 24 and 0xff,
)
}.getOrNull()
private companion object {
const val ADB_SERVICE = "_adb-tls-connect._tcp"
}
}
@@ -0,0 +1,156 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.app
import android.app.Notification
import android.app.NotificationChannel
import android.app.NotificationManager
import android.app.Service
import android.content.Context
import android.content.Intent
import android.os.Build
import android.os.IBinder
import app.echo_lot.protocol.ControlClient
import kotlinx.coroutines.CoroutineScope
import kotlinx.coroutines.Dispatchers
import kotlinx.coroutines.SupervisorJob
import kotlinx.coroutines.cancel
import kotlinx.coroutines.delay
import kotlinx.coroutines.launch
import kotlinx.coroutines.withContext
/**
* Keeps [AdbRelay] running and posts what it finds to the enrolled server.
*
* A foreground service because the whole point is to be useful while nobody is looking at the
* tablet: a background process is frozen within minutes of the screen going off, and a relay that
* stops relaying the moment it is left alone would be worse than none — it would be trusted right
* up until the moment it went quiet.
*
* The heartbeat re-posts the CACHED endpoint and never re-resolves. Resolving adbd's advertisement
* makes adbd re-arm its connection and raise a "wireless debugging connected" notification, so a
* heartbeat that re-resolved would turn a background convenience into a stream of notifications on
* a device sitting on a shelf. Rotation is still caught, because losing the old advertisement
* clears the resolve guard in [AdbRelay] and the replacement is resolved once, within seconds.
*/
class AdbRelayService : Service() {
private val scope = CoroutineScope(SupervisorJob() + Dispatchers.IO)
private var relay: AdbRelay? = null
@Volatile private var status: String = "starting"
@Volatile private var lastPosted: String? = null
override fun onBind(intent: Intent?): IBinder? = null
override fun onCreate() {
super.onCreate()
startForeground(NOTIFICATION_ID, notification("starting"))
val settings = Settings(this)
val r = AdbRelay(this) { msg ->
status = msg
notify(msg)
}
relay = r
r.start()
scope.launch {
while (true) {
val ep = r.lastEndpoint
if (ep != null) {
val wire = "${ep.host}:${ep.port}"
// Re-post on a heartbeat even when unchanged: the server stamps a received-at
// time, and a developer needs to tell "this endpoint is current" from "this
// endpoint is what the tablet saw before it went out of range".
val result = post(settings, ep)
status = if (result == null) {
lastPosted = wire
"reported $wire"
} else {
"found $wire, but reporting failed: $result"
}
notify(status)
}
delay(HEARTBEAT_MS)
}
}
}
/** Posts one endpoint; returns null on success or a short reason on failure. */
private suspend fun post(settings: Settings, ep: AdbRelay.Endpoint): String? =
withContext(Dispatchers.IO) {
if (!settings.serverConfigured) return@withContext "no server enrolled"
runCatching {
ControlClient(settings.serverUrl, setOf(settings.serverPin), BuildConfig.APP_SEMVER)
.reportAdbEndpoint(
credential = settings.serverCredential,
host = ep.host,
port = ep.port,
deviceName = Build.MODEL,
note = "echolot relay",
)
null
}.getOrElse { it.message?.take(120) ?: it.javaClass.simpleName }
}
override fun onStartCommand(intent: Intent?, flags: Int, startId: Int): Int {
// Restarted by the system if it is killed: a relay that quietly does not come back after
// a low-memory kill is the failure mode this exists to avoid.
return START_STICKY
}
override fun onDestroy() {
relay?.stop()
scope.cancel()
super.onDestroy()
}
private fun notify(text: String) {
val nm = getSystemService(NotificationManager::class.java)
nm?.notify(NOTIFICATION_ID, notification(text))
}
private fun notification(text: String): Notification {
val nm = getSystemService(NotificationManager::class.java)
if (Build.VERSION.SDK_INT >= Build.VERSION_CODES.O) {
// LOW: this is a status line for a tool the user deliberately started, not news.
val ch = NotificationChannel(CHANNEL, "adb relay", NotificationManager.IMPORTANCE_LOW)
ch.description = "Reports this device's wireless-debug endpoint to the Echolot server"
nm?.createNotificationChannel(ch)
}
val open = android.app.PendingIntent.getActivity(
this, 0, Intent(this, MainActivity::class.java),
android.app.PendingIntent.FLAG_IMMUTABLE,
)
return Notification.Builder(this, CHANNEL)
.setContentTitle("Echolot adb relay")
.setContentText(text)
.setSmallIcon(android.R.drawable.stat_sys_download_done)
.setOngoing(true)
.setContentIntent(open)
.build()
}
companion object {
private const val CHANNEL = "adb-relay"
private const val NOTIFICATION_ID = 4711
/**
* Two minutes. The port rotates on roughly that cadence, and the freshness of the answer
* is the whole product — but this only re-posts a cached value, so it costs one small
* HTTPS request and never touches mDNS.
*/
private const val HEARTBEAT_MS = 120_000L
fun start(ctx: Context) {
val i = Intent(ctx, AdbRelayService::class.java)
if (Build.VERSION.SDK_INT >= Build.VERSION_CODES.O) ctx.startForegroundService(i)
else ctx.startService(i)
}
fun stop(ctx: Context) {
ctx.stopService(Intent(ctx, AdbRelayService::class.java))
}
}
}
@@ -4,6 +4,7 @@
package app.echo_lot.app package app.echo_lot.app
import androidx.compose.foundation.layout.Arrangement import androidx.compose.foundation.layout.Arrangement
import androidx.compose.foundation.layout.safeDrawingPadding
import androidx.compose.foundation.layout.Column import androidx.compose.foundation.layout.Column
import androidx.compose.foundation.layout.Row import androidx.compose.foundation.layout.Row
import androidx.compose.foundation.layout.fillMaxWidth import androidx.compose.foundation.layout.fillMaxWidth
@@ -39,7 +40,7 @@ fun HistoryScreen(
onDelete: (String) -> Unit, onDelete: (String) -> Unit,
onBack: () -> Unit, onBack: () -> Unit,
) { ) {
Column(Modifier.fillMaxWidth().padding(16.dp), verticalArrangement = Arrangement.spacedBy(10.dp)) { Column(Modifier.fillMaxWidth().safeDrawingPadding().padding(16.dp), verticalArrangement = Arrangement.spacedBy(10.dp)) {
Row(verticalAlignment = Alignment.CenterVertically) { Row(verticalAlignment = Alignment.CenterVertically) {
TextButton(onClick = onBack) { Text(" Back") } TextButton(onClick = onBack) { Text(" Back") }
Text("History", style = MaterialTheme.typography.titleLarge) Text("History", style = MaterialTheme.typography.titleLarge)
@@ -71,12 +72,20 @@ fun HistoryScreen(
) )
} }
Text( Text(
"${r.findingCount} finding(s) · ${r.sizeBytes / 1024} kB · ${r.anonymization}", "${r.findingCount} finding(s) · ${r.sizeBytes / 1024} kB · " +
"kept complete on this device",
style = MaterialTheme.typography.bodySmall, style = MaterialTheme.typography.bodySmall,
) )
// The upload line names the level the upload was made at, not the
// archive's. They describe different documents, and showing the archive's
// level here claimed more had left the device than actually did.
Text( Text(
if (r.uploaded) "uploaded to ${r.uploadedTo ?: "a server"}" if (r.uploaded) {
else "on this device only", "uploaded to ${r.uploadedTo ?: "a server"}" +
(r.uploadedAs?.let { " as $it" } ?: "")
} else {
"on this device only"
},
style = MaterialTheme.typography.bodySmall, style = MaterialTheme.typography.bodySmall,
color = if (r.uploaded) Color(0xFF7FD17F) else Color(0xFFBBBBBB), color = if (r.uploaded) Color(0xFF7FD17F) else Color(0xFFBBBBBB),
) )
@@ -43,9 +43,28 @@ class MainActivity : ComponentActivity() {
private val permissionLauncher = private val permissionLauncher =
registerForActivityResult(ActivityResultContracts.RequestMultiplePermissions()) { /* proceed regardless */ } registerForActivityResult(ActivityResultContracts.RequestMultiplePermissions()) { /* proceed regardless */ }
/**
* The intent currently being acted on, so a deep link that arrives while the app is running
* is seen by the screen the user is already looking at.
*
* The activity is singleTask for the same reason. As a standard activity it stacked a second
* instance per link, each with its own ViewModel: the enrolment then happened in a throwaway
* copy, and pressing back returned to the original screen showing none of it. Silent, and
* indistinguishable from the link simply not working.
*/
private val liveIntent = mutableStateOf<android.content.Intent?>(null)
override fun onNewIntent(intent: android.content.Intent) {
super.onNewIntent(intent)
setIntent(intent)
liveIntent.value = intent
}
override fun onCreate(savedInstanceState: Bundle?) { override fun onCreate(savedInstanceState: Bundle?) {
super.onCreate(savedInstanceState) super.onCreate(savedInstanceState)
requestRuntimePermissions() requestRuntimePermissions()
resumeRelayIfEnabled()
liveIntent.value = intent
setContent { setContent {
MaterialTheme(colorScheme = darkColorScheme()) { MaterialTheme(colorScheme = darkColorScheme()) {
Surface(color = MaterialTheme.colorScheme.background) { Surface(color = MaterialTheme.colorScheme.background) {
@@ -63,15 +82,105 @@ class MainActivity : ComponentActivity() {
// An echolot://enroll link (QR scan, or a link the operator sent) opens the // An echolot://enroll link (QR scan, or a link the operator sent) opens the
// app straight into settings with the enrollment already done, so the user // app straight into settings with the enrollment already done, so the user
// sees the result rather than a form they still have to fill in. // sees the result rather than a form they still have to fill in.
val enrollUri = intent?.takeIf { it.action == Intent.ACTION_VIEW }?.dataString // Both deep links land here. They are told apart by host, so a sign-in
// redirect is never mistaken for an enrolment link — one spends a token, the
// other completes an authorization, and confusing them would fail obscurely.
val incoming = liveIntent.value?.takeIf { it.action == Intent.ACTION_VIEW }?.dataString
val authUri = incoming?.takeIf { it.startsWith("echolot://auth") }
val enrollUri = incoming?.takeIf { it.startsWith("echolot://enroll") }
androidx.compose.runtime.LaunchedEffect(enrollUri) { androidx.compose.runtime.LaunchedEffect(enrollUri) {
if (enrollUri != null) { if (enrollUri != null) {
vm.enroll(enrollUri) vm.enroll(enrollUri)
screen = Screen.SETTINGS screen = Screen.SETTINGS
} }
} }
// Replacing an existing enrollment is asked about, never assumed. Following a
// link from a web page is one tap, and the old credential does not survive it.
vm.state.pendingEnroll?.let { pending ->
androidx.compose.material3.AlertDialog(
onDismissRequest = { vm.cancelEnroll() },
title = {
androidx.compose.material3.Text(
if (pending.sameServer) "Enroll again with this server?"
else "Replace this device's server?"
)
},
text = {
androidx.compose.material3.Text(
// Naming the same URL twice reads as a mistake and buries the
// one consequence that actually applies: the device is issued a
// fresh credential and shows up as a second entry.
if (pending.sameServer) {
"This device is already enrolled with " +
"${pending.currentServer}.\n\n" +
"Enrolling again replaces its credential. The old one " +
"stops working immediately, and the device appears on " +
"the server as a new entry alongside the current one — " +
"which you may want to revoke afterwards.\n\n" +
"Runs already uploaded, and runs stored on this phone, " +
"are not affected."
} else {
"This device is already enrolled with " +
"${pending.currentServer}.\n\n" +
"Enrolling with ${pending.newServer} replaces that. Runs " +
"already uploaded stay where they are, but this device " +
"stops reporting to the old server and appears on the new " +
"one as a new device.\n\n" +
"Runs stored on this phone are not affected."
}
)
},
confirmButton = {
androidx.compose.material3.TextButton(onClick = { vm.confirmEnroll() }) {
androidx.compose.material3.Text(
if (pending.sameServer) "Enroll again" else "Enroll here"
)
}
},
dismissButton = {
androidx.compose.material3.TextButton(onClick = { vm.cancelEnroll() }) {
androidx.compose.material3.Text("Keep current server")
}
},
)
}
androidx.compose.runtime.LaunchedEffect(authUri) {
if (authUri != null) {
vm.completeSignIn(authUri)
screen = Screen.SETTINGS
}
}
// A long run samples for minutes, and Android starts throttling timers and
// network access within moments of the screen going off — so a run left to
// itself would measure the device's power management rather than the network,
// and would do it silently. Held only while a run is in flight, and released
// on the way out.
val view = androidx.compose.ui.platform.LocalView.current
androidx.compose.runtime.DisposableEffect(vm.state.running) {
view.keepScreenOn = vm.state.running
onDispose { view.keepScreenOn = false }
}
// Shizuku can be started, stopped or authorised in its own app, where nothing
// calls back into this process. Asking again each time this screen comes
// forward is what makes the banner right after the user has been away to fix
// it — which is exactly the moment they look at it.
val lifecycleOwner = androidx.compose.ui.platform.LocalLifecycleOwner.current
androidx.compose.runtime.DisposableEffect(lifecycleOwner) {
val obs = androidx.lifecycle.LifecycleEventObserver { _, event ->
if (event == androidx.lifecycle.Lifecycle.Event.ON_RESUME) {
vm.refreshShizuku()
}
}
lifecycleOwner.lifecycle.addObserver(obs)
onDispose { lifecycleOwner.lifecycle.removeObserver(obs) }
}
// Autorun stays a quick run: it is an unattended batch job driven over adb, and
// an automation that silently held the device for five minutes would be a
// surprise. `--es mode long` asks for the other one explicitly.
val autorunMode =
if (intent?.getStringExtra("mode") == "long") RunMode.LONG else RunMode.SHORT
androidx.compose.runtime.LaunchedEffect(autorun) { androidx.compose.runtime.LaunchedEffect(autorun) {
if (autorun) vm.run(devUpload = true) if (autorun) vm.run(autorunMode, devUpload = true)
} }
// In autorun the app is a batch job: once the run is done AND the upload // In autorun the app is a batch job: once the run is done AND the upload
// succeeded, show the result briefly, then close so the device is left as it // succeeded, show the result briefly, then close so the device is left as it
@@ -85,23 +194,47 @@ class MainActivity : ComponentActivity() {
finish() finish()
} }
} }
// Without this, the system Back gesture leaves the activity from Settings or
// History instead of returning to the run screen — the screen is a plain state
// variable, so nothing connects it to the back stack. Registered only when
// there is somewhere to go back to, so Back still exits from the run screen.
androidx.activity.compose.BackHandler(enabled = screen != Screen.RUN) {
screen = Screen.RUN
}
when (screen) { when (screen) {
Screen.SETTINGS -> SettingsScreen( Screen.SETTINGS -> SettingsScreen(
settings = vm.settings, settings = vm.settings,
archivedRuns = vm.state.history.size, archivedRuns = vm.archivedRunCount(),
archivedBytes = vm.archivedBytes(), archivedBytes = vm.archivedBytes(),
onApplyRetention = vm::applyRetention, onApplyRetention = vm::applyRetention,
onDeleteAll = vm::deleteAllRuns, onDeleteAll = vm::deleteAllRuns,
onPreviewUpload = { onPreviewUpload = {
// Preview the newest run, since that is the one the user just made // Straight from the archive: the newest run is the one the user just
// and the one they are deciding about. // made and the one they are deciding about. Always shows something,
vm.state.history.firstOrNull()?.let { r -> // even when there is nothing to preview yet.
lifecycleScope.launch { preview = vm.uploadPreview(r.id) } lifecycleScope.launch { preview = vm.previewNewestRun() }
}
}, },
onCheckServer = vm::checkServer, onCheckServer = vm::checkServer,
accountName = vm.accountName,
onSignIn = {
vm.beginSignIn { url ->
// A plain VIEW intent rather than a Custom Tab: the browser is
// where the user's existing IdP session already lives, and
// androidx.browser would be a dependency for a rounded corner.
runCatching {
startActivity(Intent(Intent.ACTION_VIEW, android.net.Uri.parse(url)))
}
}
},
onSignOut = vm::signOut,
onEnroll = vm::enroll, onEnroll = vm::enroll,
serverStatus = vm.state.archiveStatus, serverStatus = vm.state.archiveStatus,
enrollStatus = vm.state.enrollStatus,
onRelayChange = { on ->
if (on) AdbRelayService.start(this@MainActivity)
else AdbRelayService.stop(this@MainActivity)
},
onBack = { screen = Screen.RUN }, onBack = { screen = Screen.RUN },
) )
Screen.HISTORY -> HistoryScreen( Screen.HISTORY -> HistoryScreen(
@@ -125,7 +258,8 @@ class MainActivity : ComponentActivity() {
) )
Screen.RUN -> EcholotScreen( Screen.RUN -> EcholotScreen(
state = vm.state, state = vm.state,
onRun = { vm.run() }, longMinutes = vm.settings.longRunMinutes,
onRun = { mode -> vm.run(mode) },
onCancel = vm::cancel, onCancel = vm::cancel,
onDeveloperOptions = { onDeveloperOptions = {
runCatching { runCatching {
@@ -162,11 +296,29 @@ class MainActivity : ComponentActivity() {
private fun requestRuntimePermissions() { private fun requestRuntimePermissions() {
val perms = mutableListOf(Manifest.permission.ACCESS_FINE_LOCATION) val perms = mutableListOf(Manifest.permission.ACCESS_FINE_LOCATION)
// Only so the relay's ongoing notification is visible. The service runs either way, but a
// foreground service the user cannot see is worse than one they can dismiss knowingly.
if (Build.VERSION.SDK_INT >= Build.VERSION_CODES.TIRAMISU) {
perms.add(Manifest.permission.POST_NOTIFICATIONS)
}
val missing = perms.filter { val missing = perms.filter {
ContextCompat.checkSelfPermission(this, it) != PackageManager.PERMISSION_GRANTED ContextCompat.checkSelfPermission(this, it) != PackageManager.PERMISSION_GRANTED
} }
if (missing.isNotEmpty()) permissionLauncher.launch(missing.toTypedArray()) if (missing.isNotEmpty()) permissionLauncher.launch(missing.toTypedArray())
} }
/**
* Restarts the relay if it was left on.
*
* A relay that silently fails to come back after a reboot or a process kill is worse than one
* that was never enabled: it is trusted right up to the moment it goes quiet, and the symptom
* is a stale endpoint that sends a developer to a port nothing is listening on.
*/
private fun resumeRelayIfEnabled() {
if (Settings(this).adbRelayEnabled) {
runCatching { AdbRelayService.start(this) }
}
}
} }
private fun verdictColor(v: Verdict): Color = when (v) { private fun verdictColor(v: Verdict): Color = when (v) {
@@ -186,7 +338,8 @@ private fun statusColor(s: TestStatus): Color = when (s) {
@Composable @Composable
private fun EcholotScreen( private fun EcholotScreen(
state: UiState, state: UiState,
onRun: () -> Unit, longMinutes: Int,
onRun: (RunMode) -> Unit,
onCancel: () -> Unit, onCancel: () -> Unit,
onShizukuAction: () -> Unit, onShizukuAction: () -> Unit,
onDeveloperOptions: () -> Unit, onDeveloperOptions: () -> Unit,
@@ -250,8 +403,36 @@ private fun EcholotScreen(
} }
} }
// The choice is made before the run, not after, because it is a choice about how long the
// user is willing to stand still — and because the two modes answer different questions.
var mode by remember { mutableStateOf(RunMode.SHORT) }
Row(horizontalArrangement = Arrangement.spacedBy(8.dp), verticalAlignment = Alignment.CenterVertically) {
FilterChip(
selected = mode == RunMode.SHORT,
onClick = { mode = RunMode.SHORT },
enabled = !state.running,
label = { Text("Quick") },
)
FilterChip(
selected = mode == RunMode.LONG,
onClick = { mode = RunMode.LONG },
enabled = !state.running,
label = { Text("Long ($longMinutes min)") },
)
}
Text(
if (mode == RunMode.SHORT) {
"About 30 seconds. Describes how the network is configured right now."
} else {
"Listens for $longMinutes minutes while it measures. Finds what a quick run " +
"structurally cannot: links that drop and come back, signal that decays, " +
"loss that arrives in bursts."
},
fontSize = 12.sp, color = MaterialTheme.colorScheme.onSurfaceVariant,
)
Row(horizontalArrangement = Arrangement.spacedBy(12.dp), verticalAlignment = Alignment.CenterVertically) { Row(horizontalArrangement = Arrangement.spacedBy(12.dp), verticalAlignment = Alignment.CenterVertically) {
Button(onClick = onRun, enabled = !state.running) { Button(onClick = { onRun(mode) }, enabled = !state.running) {
Text(if (state.running) "Running…" else "Run measurement") Text(if (state.running) "Running…" else "Run measurement")
} }
if (state.running) { if (state.running) {
@@ -272,25 +453,55 @@ private fun EcholotScreen(
if (state.running) { if (state.running) {
Column(verticalArrangement = Arrangement.spacedBy(6.dp)) { Column(verticalArrangement = Arrangement.spacedBy(6.dp)) {
val frac = if (state.stepsTotal > 0) // In a long run the window is the run: once the battery is done, four minutes of
state.stepsDone.toFloat() / state.stepsTotal else 0f // listening remain, and a bar driven by the step count would sit at 100 % through
// all of it — which reads as an app that has hung, not one that is working.
val listening = state.windowTotalS > 0
val frac = when {
listening -> state.windowElapsedS.toFloat() / state.windowTotalS
state.stepsTotal > 0 -> state.stepsDone.toFloat() / state.stepsTotal
else -> 0f
}
LinearProgressIndicator( LinearProgressIndicator(
progress = { frac }, progress = { frac.coerceIn(0f, 1f) },
modifier = Modifier.fillMaxWidth(), modifier = Modifier.fillMaxWidth(),
) )
if (listening) {
Row(horizontalArrangement = Arrangement.spacedBy(8.dp)) {
Text(
"listening ${clock(state.windowElapsedS)} of ${clock(state.windowTotalS)}",
fontSize = 12.sp, fontWeight = FontWeight.Medium,
)
Text(
"${clock(state.windowTotalS - state.windowElapsedS)} left",
fontSize = 12.sp, modifier = Modifier.weight(1f),
color = MaterialTheme.colorScheme.onSurfaceVariant,
)
}
}
Row(horizontalArrangement = Arrangement.spacedBy(8.dp)) { Row(horizontalArrangement = Arrangement.spacedBy(8.dp)) {
// Still shown during a long run's listening phase, where it stops at the last
// test: the battery's progress is real information, it is simply not the whole
// run any more.
Text( Text(
if (state.stepsTotal > 0) if (state.stepsTotal > 0)
"test ${state.stepsDone + 1} of ${state.stepsTotal}" else "starting", "test ${(state.stepsDone + 1).coerceAtMost(state.stepsTotal)} of ${state.stepsTotal}"
else "starting",
fontSize = 12.sp, fontWeight = FontWeight.Medium, fontSize = 12.sp, fontWeight = FontWeight.Medium,
) )
Text(state.currentStep ?: "", fontSize = 12.sp, Text(state.currentStep ?: "", fontSize = 12.sp,
fontFamily = FontFamily.Monospace, modifier = Modifier.weight(1f)) fontFamily = FontFamily.Monospace, modifier = Modifier.weight(1f))
if (state.etaSeconds > 0) { if (!listening && state.etaSeconds > 0) {
Text("~${state.etaSeconds}s left", fontSize = 12.sp, Text("~${state.etaSeconds}s left", fontSize = 12.sp,
color = MaterialTheme.colorScheme.onSurfaceVariant) color = MaterialTheme.colorScheme.onSurfaceVariant)
} }
} }
if (listening) {
Text(
"Cancelling keeps what has been collected so far.",
fontSize = 11.sp, color = MaterialTheme.colorScheme.onSurfaceVariant,
)
}
} }
} }
@@ -300,11 +511,57 @@ private fun EcholotScreen(
@Composable @Composable
private fun Results(doc: MeasurementDocument) { private fun Results(doc: MeasurementDocument) {
// A constrained run is answered before the lights are: the verdict below is INCONCLUSIVE by
// §7.3, and without this banner "inconclusive" reads as the app failing rather than the OS
// (correctly) refusing to let anything past the VPN be measured.
val constraints = doc.run.constraints
if (constraints.constrained) {
val blocked = constraints.unmeasuredNetworks
.mapNotNull { id -> doc.networks.firstOrNull { it.id == id } }
.joinToString(", ") { it.iface?.takeIf { s -> s.isNotBlank() } ?: it.transport.name.lowercase() }
.ifBlank { "the networks beneath it" }
// Same three-way split as the finding: saying "VPN" when the user just disconnected
// theirs (the wall lingers during teardown) reads as the app being wrong, not the OS.
val (headline, body) = when {
constraints.vpnActive && constraints.perNetworkBlocked ->
"Measured through a VPN" to
("Android does not let apps send on the networks beneath an active VPN, so " +
"$blocked could not be measured — these results describe the tunnel. " +
"Disconnect the VPN and run again to measure the networks themselves.")
constraints.perNetworkBlocked ->
"Some networks could not be measured" to
("Android refused this app permission to send on $blocked, so they went " +
"unmeasured. A connected VPN is the usual cause; the refusal can also " +
"outlast one. Everything else in this run is unaffected.")
else ->
"A VPN holds the default route" to
("Default-route results describe the tunnel; per-network measurements " +
"reached the underlying networks.")
}
Card(colors = CardDefaults.cardColors(containerColor = Color(0xFF3A2E12))) {
Column(Modifier.fillMaxWidth().padding(12.dp)) {
Text(headline, color = Color(0xFFFFD08A), fontWeight = FontWeight.SemiBold)
Text(body, fontSize = 12.sp, color = Color(0xFFFFD08A))
}
}
}
val summary = doc.summary val summary = doc.summary
if (summary != null) { if (summary != null) {
Card(colors = CardDefaults.cardColors(containerColor = verdictColor(summary.overall))) { Card(colors = CardDefaults.cardColors(containerColor = verdictColor(summary.overall))) {
Column(Modifier.fillMaxWidth().padding(16.dp)) { Column(Modifier.fillMaxWidth().padding(16.dp)) {
Text("Overall: ${summary.overall}", color = Color.White, fontWeight = FontWeight.Bold, fontSize = 18.sp) Text("Overall: ${summary.overall}", color = Color.White, fontWeight = FontWeight.Bold, fontSize = 18.sp)
// Which question this document answers. A green light from a quick run does not
// mean the same thing as a green light from a long one, and the report should not
// let the two look identical.
Text(
if (doc.run.mode == RunMode.LONG) {
"long run — the network was watched continuously as well as probed"
} else {
"quick run — a snapshot; nothing here rules out an intermittent fault"
},
color = Color.White, fontSize = 12.sp,
)
} }
} }
FlowCategories(summary.categories) FlowCategories(summary.categories)
@@ -410,6 +667,12 @@ private fun FlowCategories(categories: Map<String, CategorySummary>) {
} }
} }
/** m:ss — minutes are how a five-minute wait is read; "247s left" is a number to convert. */
private fun clock(seconds: Int): String {
val s = seconds.coerceAtLeast(0)
return "${s / 60}:${"%02d".format(s % 60)}"
}
@Composable @Composable
private fun Dot(color: Color) { private fun Dot(color: Color) {
Surface(color = color, shape = RoundedCornerShape(50), modifier = Modifier.size(12.dp)) {} Surface(color = color, shape = RoundedCornerShape(50), modifier = Modifier.size(12.dp)) {}
@@ -80,8 +80,16 @@ class RunStore(context: Context, private val settings: Settings) {
private fun client() = ControlClient( private fun client() = ControlClient(
settings.serverUrl, setOf(settings.serverPin), BuildConfig.APP_SEMVER, settings.serverUrl, setOf(settings.serverPin), BuildConfig.APP_SEMVER,
fallbackAddrs = settings.serverAddrList(),
) )
/** Remembers where the server lives, so a later run can reach it without DNS. */
private fun rememberAddrs(p: app.echo_lot.protocol.Profile) {
val addrs = p.targets.flatMap { listOfNotNull(it.ip4, it.ip6) }
.filter { it.isNotBlank() }
if (addrs.isNotEmpty()) settings.serverAddrs = addrs.joinToString(",")
}
/** /**
* Checks the configured server without uploading anything: reachable, pinned, compatible, and * Checks the configured server without uploading anything: reachable, pinned, compatible, and
* willing to accept runs. Lets the user find out in settings rather than from a failed run. * willing to accept runs. Lets the user find out in settings rather than from a failed run.
@@ -90,6 +98,10 @@ class RunStore(context: Context, private val settings: Settings) {
if (!settings.serverConfigured) return "Fill in the server URL, pin and credential first." if (!settings.serverConfigured) return "Fill in the server URL, pin and credential first."
return try { return try {
val profile = client().profile(settings.serverCredential) val profile = client().profile(settings.serverCredential)
// Learned here so the next run's canary probe knows what to ask for.
profile.canaryZone.takeIf { it.isNotBlank() }?.let { settings.canaryZone = it }
settings.serverFacts = describeFacts(profile)
rememberAddrs(profile)
val compat = Compat.check(profile, BuildConfig.APP_SEMVER) val compat = Compat.check(profile, BuildConfig.APP_SEMVER)
val head = "${profile.name} · server ${profile.serverVersion} · " + val head = "${profile.name} · server ${profile.serverVersion} · " +
"protocol ${profile.compat.protocolVersion.ifBlank { "unstated" }}" "protocol ${profile.compat.protocolVersion.ifBlank { "unstated" }}"
@@ -115,6 +127,36 @@ class RunStore(context: Context, private val settings: Settings) {
* no credential — fails later, somewhere else, with an error that points at the wrong thing. * no credential — fails later, somewhere else, with an error that points at the wrong thing.
* Blocking; callers run it off the main thread. * Blocking; callers run it off the main thread.
*/ */
/**
* Renders what the server says about itself, for display.
*
* Only what a person measuring against it would want to check: which addresses the tests will
* actually use, on which ports, and what the server admits it can do. Addresses first, because
* "which address did this result come from" is the question a report leaves open.
*/
private fun describeFacts(p: app.echo_lot.protocol.Profile): String {
val lines = ArrayList<String>()
// "label|value" per line, laid out as real columns by the UI rather than padded with
// spaces here. Space padding only lines up in a monospaced font, which makes the layout
// depend on a typeface choice made somewhere else entirely.
fun row(label: String, value: String) = lines.add("$label|$value")
row("server", "${p.name} · ${p.serverVersion}")
for (t in p.targets) {
t.ip4?.let { row("IPv4", it) }
t.ip6?.let { row("IPv6", it) }
// Marked rather than listed apart: it is the same server, and what matters is being
// able to tell which address a NAT-behaviour result came from.
t.ip4Alt?.let { row("IPv4 alt", it) }
t.ip6Alt?.let { row("IPv6 alt", it) }
row("ports", "udp ${t.udpPort} · tcp ${t.tcpPort} · stun ${t.stunPort}")
}
if (p.canaryZone.isNotBlank()) row("dns zone", p.canaryZone)
if (p.capabilities.isNotEmpty()) row("measures", p.capabilities.joinToString(", "))
return lines.joinToString(System.lineSeparator())
}
fun enroll(link: String, deviceName: String?): String { fun enroll(link: String, deviceName: String?): String {
val parsed = app.echo_lot.protocol.EnrollmentLink.parse(link) val parsed = app.echo_lot.protocol.EnrollmentLink.parse(link)
?: return "That does not look like an Echolot enrollment link. It should start with " + ?: return "That does not look like an Echolot enrollment link. It should start with " +
@@ -122,7 +164,11 @@ class RunStore(context: Context, private val settings: Settings) {
return try { return try {
val enrolled = parsed.redeem(deviceName, BuildConfig.APP_SEMVER) val enrolled = parsed.redeem(deviceName, BuildConfig.APP_SEMVER)
val compat = Compat.check(enrolled.profile, BuildConfig.APP_SEMVER) val compat = Compat.check(enrolled.profile, BuildConfig.APP_SEMVER)
settings.serverFacts = describeFacts(enrolled.profile)
rememberAddrs(enrolled.profile)
enrolled.profile.canaryZone.takeIf { it.isNotBlank() }?.let { settings.canaryZone = it }
settings.serverUrl = enrolled.controlUrl settings.serverUrl = enrolled.controlUrl
settings.serverPublicUrl = enrolled.publicUrl
settings.serverPin = enrolled.pin settings.serverPin = enrolled.pin
settings.serverCredential = enrolled.credential settings.serverCredential = enrolled.credential
val head = "Enrolled with ${enrolled.profile.name} " + val head = "Enrolled with ${enrolled.profile.name} " +
@@ -150,6 +196,9 @@ class RunStore(context: Context, private val settings: Settings) {
return try { return try {
val client = client() val client = client()
val profile = client.profile(settings.serverCredential) val profile = client.profile(settings.serverCredential)
profile.canaryZone.takeIf { it.isNotBlank() }?.let { settings.canaryZone = it }
settings.serverFacts = describeFacts(profile)
rememberAddrs(profile)
// Compatibility before policy: an incompatible server may well advertise an upload // Compatibility before policy: an incompatible server may well advertise an upload
// policy it would never actually apply to us. // policy it would never actually apply to us.
@@ -164,8 +213,10 @@ class RunStore(context: Context, private val settings: Settings) {
) )
val body = redactedForUpload(docJson, level) val body = redactedForUpload(docJson, level)
val reply = client.uploadRun(settings.serverCredential, body) val reply = client.uploadRun(settings.serverCredential, body)
archive.markUploaded(runId, profile.name) archive.markUploaded(runId, profile.name, level.wire)
UploadOutcome.Sent(profile.name, "as $level, ${body.toByteArray().size} bytes: ${reply.take(120)}") // Deliberately not echoing `reply`: it is the server's index entry as raw JSON, and
// it ended up rendered verbatim in the UI. Size and level are what a person wants.
UploadOutcome.Sent(profile.name, "as $level, ${body.toByteArray().size} bytes")
} catch (e: VersionRefused) { } catch (e: VersionRefused) {
UploadOutcome.Incompatible(e.message ?: "the server refused this app's version") UploadOutcome.Incompatible(e.message ?: "the server refused this app's version")
} catch (e: UploadRefused) { } catch (e: UploadRefused) {
File diff suppressed because it is too large Load Diff
@@ -50,6 +50,65 @@ class Settings(context: Context) {
maxTotalBytes = maxTotalMb.toLong() * 1024 * 1024, maxTotalBytes = maxTotalMb.toLong() * 1024 * 1024,
) )
// ---- measurement -------------------------------------------------------------------
/**
* How long a long run listens, in minutes. Offered as 1 / 5 / 15.
*
* 5 is the default because it is the shortest window in which the things long mode exists to
* catch — a link that flaps, a signal that decays as someone walks around, loss that comes in
* bursts — have a fair chance of happening at least twice. One minute is for checking that the
* mode works at all; fifteen is for chasing something already suspected.
*
* Clamped rather than trusted: a zero-minute long run would produce a document claiming a
* window it never watched, which is the one thing `run.mode` exists to prevent.
*/
var longRunMinutes: Int
get() = prefs.getInt(LONG_RUN_MINUTES, 5).coerceIn(1, 60)
set(v) = prefs.edit().putInt(LONG_RUN_MINUTES, v.coerceIn(1, 60)).apply()
// ---- dev relay -----------------------------------------------------------------------
/**
* Whether this device relays adbd's wireless-debug endpoint to the enrolled server.
*
* Off by default and never implied by anything else: it publishes where this device can be
* reached for debugging, which is a decision rather than a side effect. Intended for a spare
* device parked on a test network — see AdbRelay for why the OnePlus is a poor host for it.
*/
var adbRelayEnabled: Boolean
get() = prefs.getBoolean(ADB_RELAY, false)
set(v) = prefs.edit().putBoolean(ADB_RELAY, v).apply()
// ---- run-duration learning ---------------------------------------------------------
/**
* Learned duration of one test type on THIS device, or null before the first run.
*
* The static Probe.estimatedMs values are only cold-start seeds: real durations depend on
* the phone and the network it stands in (ICMPv6 answers in milliseconds where IPv6 works
* and waits out full timeouts where it does not), so a fixed table is wrong for almost
* everyone almost always. What was measured last time is the only estimate that tracks
* reality.
*/
fun learnedDurationMs(type: String): Long? =
prefs.getLong("$DURATION_PREFIX$type", -1L).takeIf { it > 0 }
/**
* Feeds one measured duration into the estimate — EMA, 70 % old / 30 % new. Heavy enough
* on history that a single odd run (a captive portal stalling DNS) does not whipsaw the
* bar, light enough that a real change (enrolling with a server un-skips three probes)
* converges within a few runs. Recorded whatever the test's status: a probe that skips in
* 2 ms will keep skipping in 2 ms until circumstances change, and then the EMA follows.
*/
fun recordDurationMs(type: String, ms: Long) {
if (ms < 0) return
val key = "$DURATION_PREFIX$type"
val old = prefs.getLong(key, -1L)
val next = if (old <= 0) ms else (old * 7 + ms * 3) / 10
prefs.edit().putLong(key, next).apply()
}
// ---- upload ---------------------------------------------------------------------- // ---- upload ----------------------------------------------------------------------
/** /**
@@ -97,6 +156,41 @@ class Settings(context: Context) {
get() = prefs.getString(SERVER_PIN, "") ?: "" get() = prefs.getString(SERVER_PIN, "") ?: ""
set(v) = prefs.edit().putString(SERVER_PIN, v.trim()).apply() set(v) = prefs.edit().putString(SERVER_PIN, v.trim()).apply()
/**
* The address the operator handed out, for showing to a person.
*
* Separate from [serverUrl], which is the endpoint actually dialled. They differ when the
* server publishes one public name and points devices at another to select its pinned
* certificate — a detail worth keeping out of the user's face but not out of the settings.
*/
var serverPublicUrl: String
get() = (prefs.getString(SERVER_PUBLIC_URL, "") ?: "").ifBlank { serverUrl }
set(v) = prefs.edit().putString(SERVER_PUBLIC_URL, v.trim()).apply()
/**
* What the server said about itself, last time it was asked: addresses, ports, capabilities.
*
* Cached as a rendered block rather than as fields, because it is shown and never acted on —
* these are facts to read, not settings to apply, and storing them as settings would invite
* exactly the confusion of an editable box that changes nothing.
*/
var serverFacts: String
get() = prefs.getString(SERVER_FACTS, "") ?: ""
set(v) = prefs.edit().putString(SERVER_FACTS, v).apply()
/**
* The server's own addresses, learned from its profile, for reaching it when DNS will not.
*
* Only the primaries: the alternate pair exists for NAT behaviour discovery and does not carry
* the control plane, so falling back to one would fail for a second, unrelated reason.
*/
var serverAddrs: String
get() = prefs.getString(SERVER_ADDRS, "") ?: ""
set(v) = prefs.edit().putString(SERVER_ADDRS, v).apply()
fun serverAddrList(): List<String> =
serverAddrs.split(',').map { it.trim() }.filter { it.isNotEmpty() }
var serverCredential: String var serverCredential: String
get() = prefs.getString(SERVER_CRED, "") ?: "" get() = prefs.getString(SERVER_CRED, "") ?: ""
set(v) = prefs.edit().putString(SERVER_CRED, v.trim()).apply() set(v) = prefs.edit().putString(SERVER_CRED, v.trim()).apply()
@@ -104,6 +198,58 @@ class Settings(context: Context) {
val serverConfigured: Boolean val serverConfigured: Boolean
get() = serverUrl.isNotBlank() && serverPin.isNotBlank() && serverCredential.isNotBlank() get() = serverUrl.isNotBlank() && serverPin.isNotBlank() && serverCredential.isNotBlank()
/**
* The DNS zone this server is authoritative for, learned from its profile.
*
* Cached because the canary probe runs at device tier, before anything has talked to the
* server, and a probe that had to make a control-plane call first would fail on exactly the
* networks worth measuring. Empty means "not known yet", and the probe reports itself as
* skipped rather than inventing a zone.
*/
var canaryZone: String
get() = prefs.getString(CANARY_ZONE, "") ?: ""
set(v) = prefs.edit().putString(CANARY_ZONE, v.trim()).apply()
/**
* Host part of the configured server URL, for probes that address it directly (STUN).
*
* Derived rather than stored: a second copy of the server's name is a second thing to keep in
* step, and it would go stale the moment someone re-enrolled against a different server.
*/
fun serverHost(): String = runCatching {
java.net.URI(serverUrl).host?.takeIf { it.isNotBlank() }
}.getOrNull() ?: ""
// ---- account ---------------------------------------------------------------------
/**
* The PKCE verifier and state for a sign-in that is out at the browser.
*
* Persisted rather than held in memory because handing control to a browser backgrounds this
* process, and Android may kill it before the callback returns. An in-memory value works on a
* developer's device and fails on a phone under memory pressure.
*/
var pendingVerifier: String
get() = prefs.getString(PENDING_VERIFIER, "") ?: ""
set(v) = prefs.edit().putString(PENDING_VERIFIER, v).apply()
var pendingState: String
get() = prefs.getString(PENDING_STATE, "") ?: ""
set(v) = prefs.edit().putString(PENDING_STATE, v).apply()
fun clearPendingAuth() = prefs.edit().remove(PENDING_VERIFIER).remove(PENDING_STATE).apply()
/** Display name of whoever is signed in on this device; empty when nobody is. */
var accountName: String
get() = prefs.getString(ACCOUNT_NAME, "") ?: ""
set(v) = prefs.edit().putString(ACCOUNT_NAME, v).apply()
var accountId: String
get() = prefs.getString(ACCOUNT_ID, "") ?: ""
set(v) = prefs.edit().putString(ACCOUNT_ID, v).apply()
val signedIn: Boolean get() = accountName.isNotBlank()
private fun hex(s: String) = ByteArray(s.length / 2) { private fun hex(s: String) = ByteArray(s.length / 2) {
((Character.digit(s[it * 2], 16) shl 4) or Character.digit(s[it * 2 + 1], 16)).toByte() ((Character.digit(s[it * 2], 16) shl 4) or Character.digit(s[it * 2 + 1], 16)).toByte()
} }
@@ -120,5 +266,16 @@ class Settings(context: Context) {
const val SERVER_URL = "server_url" const val SERVER_URL = "server_url"
const val SERVER_PIN = "server_pin" const val SERVER_PIN = "server_pin"
const val SERVER_CRED = "server_credential" const val SERVER_CRED = "server_credential"
const val SERVER_PUBLIC_URL = "server_public_url"
const val SERVER_FACTS = "server_facts"
const val SERVER_ADDRS = "server_addrs"
const val CANARY_ZONE = "server_canary_zone"
const val PENDING_VERIFIER = "pending_auth_verifier"
const val PENDING_STATE = "pending_auth_state"
const val ACCOUNT_NAME = "account_name"
const val ACCOUNT_ID = "account_id"
const val DURATION_PREFIX = "duration_ms."
const val ADB_RELAY = "adb_relay_enabled"
const val LONG_RUN_MINUTES = "long_run_minutes"
} }
} }
@@ -4,19 +4,24 @@
package app.echo_lot.app package app.echo_lot.app
import androidx.compose.foundation.layout.Arrangement import androidx.compose.foundation.layout.Arrangement
import androidx.compose.foundation.layout.safeDrawingPadding
import androidx.compose.foundation.layout.Column import androidx.compose.foundation.layout.Column
import androidx.compose.foundation.layout.Row import androidx.compose.foundation.layout.Row
import androidx.compose.foundation.layout.Spacer import androidx.compose.foundation.layout.Spacer
import androidx.compose.foundation.layout.fillMaxWidth import androidx.compose.foundation.layout.fillMaxWidth
import androidx.compose.foundation.layout.height import androidx.compose.foundation.layout.height
import androidx.compose.foundation.layout.width
import androidx.compose.foundation.layout.padding import androidx.compose.foundation.layout.padding
import androidx.compose.foundation.shape.RoundedCornerShape
import androidx.compose.foundation.rememberScrollState import androidx.compose.foundation.rememberScrollState
import androidx.compose.foundation.verticalScroll import androidx.compose.foundation.verticalScroll
import androidx.compose.material3.Button import androidx.compose.material3.Button
import androidx.compose.material3.Card import androidx.compose.material3.Card
import androidx.compose.material3.FilterChip import androidx.compose.material3.FilterChip
import androidx.compose.material3.LocalContentColor
import androidx.compose.material3.MaterialTheme import androidx.compose.material3.MaterialTheme
import androidx.compose.material3.OutlinedTextField import androidx.compose.material3.OutlinedTextField
import androidx.compose.material3.Surface
import androidx.compose.material3.Switch import androidx.compose.material3.Switch
import androidx.compose.material3.Text import androidx.compose.material3.Text
import androidx.compose.material3.TextButton import androidx.compose.material3.TextButton
@@ -29,6 +34,7 @@ import androidx.compose.ui.Alignment
import androidx.compose.ui.Modifier import androidx.compose.ui.Modifier
import androidx.compose.ui.text.font.FontFamily import androidx.compose.ui.text.font.FontFamily
import androidx.compose.ui.unit.dp import androidx.compose.ui.unit.dp
import androidx.compose.ui.unit.sp
import app.echo_lot.privacy.PrivacyLevel import app.echo_lot.privacy.PrivacyLevel
/** /**
@@ -47,8 +53,14 @@ fun SettingsScreen(
onDeleteAll: () -> Unit, onDeleteAll: () -> Unit,
onPreviewUpload: () -> Unit, onPreviewUpload: () -> Unit,
onCheckServer: () -> Unit, onCheckServer: () -> Unit,
accountName: String,
onSignIn: () -> Unit,
onSignOut: () -> Unit,
onEnroll: (String) -> Unit, onEnroll: (String) -> Unit,
serverStatus: String?, serverStatus: String?,
enrollStatus: String?,
/** Starts or stops the adb relay service; the toggle only records the preference. */
onRelayChange: (Boolean) -> Unit,
onBack: () -> Unit, onBack: () -> Unit,
) { ) {
// SharedPreferences is not observable, so mirror each value into Compose state and write // SharedPreferences is not observable, so mirror each value into Compose state and write
@@ -57,16 +69,31 @@ fun SettingsScreen(
var maxRuns by remember { mutableStateOf(settings.maxRuns.toString()) } var maxRuns by remember { mutableStateOf(settings.maxRuns.toString()) }
var maxAgeDays by remember { mutableStateOf(settings.maxAgeDays.toString()) } var maxAgeDays by remember { mutableStateOf(settings.maxAgeDays.toString()) }
var maxTotalMb by remember { mutableStateOf(settings.maxTotalMb.toString()) } var maxTotalMb by remember { mutableStateOf(settings.maxTotalMb.toString()) }
var longMinutes by remember { mutableStateOf(settings.longRunMinutes) }
var autoUpload by remember { mutableStateOf(settings.autoUpload) } var autoUpload by remember { mutableStateOf(settings.autoUpload) }
var privacy by remember { mutableStateOf(settings.privacyLevel) } var privacy by remember { mutableStateOf(settings.privacyLevel) }
var stableSalt by remember { mutableStateOf(settings.stableSalt) } var stableSalt by remember { mutableStateOf(settings.stableSalt) }
var relayOn by remember { mutableStateOf(settings.adbRelayEnabled) }
var enrollLink by remember { mutableStateOf("") } var enrollLink by remember { mutableStateOf("") }
var serverUrl by remember { mutableStateOf(settings.serverUrl) } // The public name, which is what the operator handed out and what a person recognises. The
// endpoint actually dialled is shown beneath it when the two differ, rather than hidden — a
// network engineer debugging a connection wants to see where it really goes.
var serverUrl by remember { mutableStateOf(settings.serverPublicUrl) }
var serverPin by remember { mutableStateOf(settings.serverPin) } var serverPin by remember { mutableStateOf(settings.serverPin) }
var serverCred by remember { mutableStateOf(settings.serverCredential) } var serverCred by remember { mutableStateOf(settings.serverCredential) }
// Enrolling is asynchronous, so these are re-read when its result lands rather than when the
// button is pressed — reading them immediately showed the previous server's values and looked
// exactly like an enrollment that had silently done nothing.
var serverFacts by remember { mutableStateOf(settings.serverFacts) }
androidx.compose.runtime.LaunchedEffect(enrollStatus, serverStatus) {
serverFacts = settings.serverFacts
serverUrl = settings.serverPublicUrl
serverPin = settings.serverPin
serverCred = settings.serverCredential
}
Column( Column(
Modifier.fillMaxWidth().verticalScroll(rememberScrollState()).padding(16.dp), Modifier.fillMaxWidth().safeDrawingPadding().verticalScroll(rememberScrollState()).padding(16.dp),
verticalArrangement = Arrangement.spacedBy(12.dp), verticalArrangement = Arrangement.spacedBy(12.dp),
) { ) {
Row(verticalAlignment = Alignment.CenterVertically) { Row(verticalAlignment = Alignment.CenterVertically) {
@@ -74,6 +101,39 @@ fun SettingsScreen(
Text("Settings", style = MaterialTheme.typography.titleLarge) Text("Settings", style = MaterialTheme.typography.titleLarge)
} }
// ---- measurement ----
Card(Modifier.fillMaxWidth()) {
Column(Modifier.padding(14.dp), verticalArrangement = Arrangement.spacedBy(8.dp)) {
Text("Long runs", style = MaterialTheme.typography.titleMedium)
Text(
"How long a long run keeps listening. The measurements themselves take about " +
"30 seconds either way; the rest of the window is spent watching for " +
"things that only happen sometimes.",
style = MaterialTheme.typography.bodySmall,
)
Row(horizontalArrangement = Arrangement.spacedBy(8.dp)) {
for (minutes in listOf(1, 5, 15)) {
FilterChip(
selected = longMinutes == minutes,
onClick = { longMinutes = minutes; settings.longRunMinutes = minutes },
label = { Text("$minutes min") },
)
}
}
Text(
when (longMinutes) {
1 -> "Barely longer than a quick run — enough to confirm the listeners " +
"work, rarely enough to catch anything intermittent."
15 -> "For a fault you already suspect and have to prove. Keep the screen " +
"on and the device where the problem happens."
else -> "Long enough for a link that drops every couple of minutes to do " +
"it at least once, short enough to wait out."
},
style = MaterialTheme.typography.bodySmall,
)
}
}
// ---- archive ---- // ---- archive ----
Card(Modifier.fillMaxWidth()) { Card(Modifier.fillMaxWidth()) {
Column(Modifier.padding(14.dp), verticalArrangement = Arrangement.spacedBy(8.dp)) { Column(Modifier.padding(14.dp), verticalArrangement = Arrangement.spacedBy(8.dp)) {
@@ -129,18 +189,59 @@ fun SettingsScreen(
} }
Text(privacyExplanation(privacy), style = MaterialTheme.typography.bodySmall) Text(privacyExplanation(privacy), style = MaterialTheme.typography.bodySmall)
// At FULL nothing is pseudonymized, so a salt has nothing to act on. Shown
// disabled rather than hidden: the setting is still stored and still applies the
// moment the level changes, and a control that vanishes hides that fact.
Toggle( Toggle(
label = "Stable pseudonyms across runs", label = "Stable pseudonyms across runs",
detail = "Lets you compare uploaded runs over time (same SSID reads the same " + detail = if (privacy == PrivacyLevel.FULL) {
"each time). It also links your uploads together, so leave it off on a " + "Not used at this level — nothing is pseudonymized, so there is nothing " +
"server you don't run yourself.", "to keep stable. Choose balanced or strict to use this."
checked = stableSalt, } else {
"Lets you compare uploaded runs over time (same SSID reads the same " +
"each time). It also links your uploads together, so leave it off on " +
"a server you don't run yourself."
},
checked = stableSalt && privacy != PrivacyLevel.FULL,
enabled = privacy != PrivacyLevel.FULL,
) { stableSalt = it; settings.stableSalt = it } ) { stableSalt = it; settings.stableSalt = it }
TextButton(onClick = onPreviewUpload) { Text("Preview what an upload would send") } TextButton(onClick = onPreviewUpload) { Text("Preview what an upload would send") }
} }
} }
// ---- account ----
Card(Modifier.fillMaxWidth()) {
Column(Modifier.padding(14.dp), verticalArrangement = Arrangement.spacedBy(8.dp)) {
Text("Account", style = MaterialTheme.typography.titleMedium)
if (accountName.isNotBlank()) {
Text("Signed in as $accountName", style = MaterialTheme.typography.bodyMedium)
Text(
"Runs from every device signed in to this account share one history.",
style = MaterialTheme.typography.bodySmall,
)
TextButton(onClick = onSignOut) { Text("Sign out") }
} else {
Text(
"Signing in is optional. It links this device to an account on your " +
"server, so several devices share one history — and some servers only " +
"accept uploads from a signed-in device.",
style = MaterialTheme.typography.bodySmall,
)
Button(onClick = onSignIn, enabled = settings.serverConfigured) {
Text("Sign in")
}
if (!settings.serverConfigured) {
Text(
"Enrol with a server first — the account belongs to the server, not " +
"to the app.",
style = MaterialTheme.typography.bodySmall,
)
}
}
}
}
// ---- upload ---- // ---- upload ----
Card(Modifier.fillMaxWidth()) { Card(Modifier.fillMaxWidth()) {
Column(Modifier.padding(14.dp), verticalArrangement = Arrangement.spacedBy(8.dp)) { Column(Modifier.padding(14.dp), verticalArrangement = Arrangement.spacedBy(8.dp)) {
@@ -172,17 +273,38 @@ fun SettingsScreen(
onClick = { onClick = {
onEnroll(enrollLink) onEnroll(enrollLink)
enrollLink = "" // spent either way; leaving it around invites a retry enrollLink = "" // spent either way; leaving it around invites a retry
serverUrl = settings.serverUrl
serverPin = settings.serverPin
serverCred = settings.serverCredential
}, },
enabled = enrollLink.isNotBlank(), enabled = enrollLink.isNotBlank(),
) { Text("Enroll") } ) { Text("Enroll") }
// Beside the button that caused it. Enrolling is asynchronous, so without this the
// only sign of success is three fields quietly changing further down the card.
enrollStatus?.let {
Text(it, style = MaterialTheme.typography.bodySmall)
}
OutlinedTextField( OutlinedTextField(
value = serverUrl, onValueChange = { serverUrl = it; settings.serverUrl = it }, value = serverUrl,
onValueChange = {
serverUrl = it
// Typed by hand there is no discovery to consult, so what was entered is
// both the public name and the endpoint. Setting only one of them would
// leave the app dialling the previous server.
settings.serverUrl = it
settings.serverPublicUrl = it
},
label = { Text("Server URL") }, singleLine = true, modifier = Modifier.fillMaxWidth(), label = { Text("Server URL") }, singleLine = true, modifier = Modifier.fillMaxWidth(),
) )
// Directly under the field it explains. Anywhere else it reads as a stray sentence
// about some other part of the screen.
if (settings.serverUrl.isNotBlank() && settings.serverUrl != settings.serverPublicUrl) {
Text(
"Connects to ${settings.serverUrl} — this server publishes one name and " +
"points devices at another, so its pinned certificate can share a port " +
"with its web interface.",
style = MaterialTheme.typography.bodySmall,
color = LocalContentColor.current.copy(alpha = 0.7f),
)
}
OutlinedTextField( OutlinedTextField(
value = serverPin, onValueChange = { serverPin = it; settings.serverPin = it }, value = serverPin, onValueChange = { serverPin = it; settings.serverPin = it },
label = { Text("Certificate pin (SPKI, base64)") }, singleLine = true, label = { Text("Certificate pin (SPKI, base64)") }, singleLine = true,
@@ -209,6 +331,52 @@ fun SettingsScreen(
serverStatus?.let { serverStatus?.let {
Text(it, style = MaterialTheme.typography.bodySmall) Text(it, style = MaterialTheme.typography.bodySmall)
} }
// What the server reported, placed under the button that asks it rather than among
// the fields above: these are facts to read, not settings to apply, and an
// editable-looking box that changes nothing is worse than no box at all.
//
// Monospaced so the addresses line up under each other — column alignment is most
// of what makes a list of IPs quicker to read than prose.
if (serverFacts.isNotBlank()) {
Surface(
color = MaterialTheme.colorScheme.surfaceVariant,
shape = RoundedCornerShape(8.dp),
modifier = Modifier.fillMaxWidth(),
) {
Column(
Modifier.padding(horizontal = 12.dp, vertical = 10.dp),
verticalArrangement = Arrangement.spacedBy(2.dp),
) {
Text(
"WHAT THIS SERVER REPORTS",
style = MaterialTheme.typography.labelSmall,
color = LocalContentColor.current.copy(alpha = 0.7f),
)
// Real columns rather than padded text: the label column has a fixed
// width, so values line up whatever the font does, and a long value
// wraps inside its own column instead of under the labels.
for (line in serverFacts.lines()) {
val label = line.substringBefore('|')
val value = line.substringAfter('|', "")
Row(Modifier.fillMaxWidth()) {
Text(
label,
style = MaterialTheme.typography.bodySmall,
color = LocalContentColor.current.copy(alpha = 0.7f),
modifier = Modifier.width(72.dp),
)
Text(
value,
style = MaterialTheme.typography.bodySmall.copy(
fontFamily = FontFamily.Monospace,
),
modifier = Modifier.weight(1f),
)
}
}
}
}
}
Text( Text(
"This app is ${BuildConfig.APP_SEMVER} and speaks probe protocol " + "This app is ${BuildConfig.APP_SEMVER} and speaks probe protocol " +
"${app.echo_lot.protocol.Compat.PROTOCOL_VERSION}. It works with servers " + "${app.echo_lot.protocol.Compat.PROTOCOL_VERSION}. It works with servers " +
@@ -217,6 +385,39 @@ fun SettingsScreen(
) )
} }
} }
// ---- dev relay ------------------------------------------------------------------
//
// Last, and deliberately plain: this is scaffolding for driving a test device, not a
// measurement. It publishes where this device can be reached over adb, which is why it
// is off until someone decides otherwise.
Card(Modifier.fillMaxWidth()) {
Column(Modifier.padding(16.dp), verticalArrangement = Arrangement.spacedBy(8.dp)) {
Text("Developer relay", style = MaterialTheme.typography.titleMedium)
Toggle(
label = "Relay this device's adb endpoint",
detail = "Watches adbd's own mDNS announcement and reports host:port to the " +
"enrolled server, so a developer on another network can reach this " +
"device — mDNS does not cross subnets, and the port rotates every few " +
"minutes. Runs in the foreground with a notification while on.",
checked = relayOn,
enabled = settings.serverConfigured,
onChange = { on ->
relayOn = on
settings.adbRelayEnabled = on
onRelayChange(on)
},
)
if (!settings.serverConfigured) {
Text(
"Needs an enrolled server: the report is authenticated with this " +
"device's credential.",
style = MaterialTheme.typography.bodySmall,
color = LocalContentColor.current.copy(alpha = 0.7f),
)
}
}
}
Spacer(Modifier.height(24.dp)) Spacer(Modifier.height(24.dp))
} }
} }
@@ -236,13 +437,22 @@ private fun privacyExplanation(level: PrivacyLevel): String = when (level) {
} }
@Composable @Composable
private fun Toggle(label: String, detail: String, checked: Boolean, onChange: (Boolean) -> Unit) { private fun Toggle(
label: String,
detail: String,
checked: Boolean,
enabled: Boolean = true,
onChange: (Boolean) -> Unit,
) {
Row(Modifier.fillMaxWidth(), verticalAlignment = Alignment.Top) { Row(Modifier.fillMaxWidth(), verticalAlignment = Alignment.Top) {
Column(Modifier.weight(1f)) { Column(Modifier.weight(1f)) {
Text(label, style = MaterialTheme.typography.bodyMedium) // Dimmed together with the switch, so "this does nothing right now" reads at a glance
Text(detail, style = MaterialTheme.typography.bodySmall) // instead of only on close inspection.
val alpha = if (enabled) 1f else 0.5f
Text(label, style = MaterialTheme.typography.bodyMedium, color = LocalContentColor.current.copy(alpha = alpha))
Text(detail, style = MaterialTheme.typography.bodySmall, color = LocalContentColor.current.copy(alpha = alpha))
} }
Switch(checked = checked, onCheckedChange = onChange) Switch(checked = checked, onCheckedChange = onChange, enabled = enabled)
} }
} }
@@ -31,10 +31,23 @@ data class ArchivedRun(
val verdict: String? = null, val verdict: String? = null,
@SerialName("finding_count") val findingCount: Int = 0, @SerialName("finding_count") val findingCount: Int = 0,
@SerialName("size_bytes") val sizeBytes: Long = 0, @SerialName("size_bytes") val sizeBytes: Long = 0,
/**
* How the *archived* document is redacted. Always "full" in practice, because the archive
* deliberately keeps the unredacted run - see the package doc. This is not what was uploaded.
*/
val anonymization: String = "full", val anonymization: String = "full",
/** Whether this run has been accepted by a server, so history can show what is backed up. */ /** Whether this run has been accepted by a server, so history can show what is backed up. */
val uploaded: Boolean = false, val uploaded: Boolean = false,
@SerialName("uploaded_to") val uploadedTo: String? = null, @SerialName("uploaded_to") val uploadedTo: String? = null,
/**
* The level the run was *uploaded* at, which is a different document from the archived one.
*
* Kept separately because conflating the two is actively misleading: the history row showed
* the archive's own level ("full") directly beneath "uploaded to fmr", which reads as "the
* complete data was uploaded" when a redacted copy had been sent. A privacy display that
* overstates what left the device is worse than none.
*/
@SerialName("uploaded_as") val uploadedAs: String? = null,
) )
/** /**
@@ -119,7 +132,7 @@ class RunArchive(private val dir: File, private val now: () -> Long = System::cu
fun deleteAll(): Int = list().count { delete(it.id) } fun deleteAll(): Int = list().count { delete(it.id) }
/** Records that a server accepted this run, so history can distinguish backed-up from local. */ /** Records that a server accepted this run, so history can distinguish backed-up from local. */
fun markUploaded(id: String, serverName: String) { fun markUploaded(id: String, serverName: String, uploadedAs: String? = null) {
val f = File(dir, safe(id) + META_EXT) val f = File(dir, safe(id) + META_EXT)
val meta = runCatching { json.decodeFromString(ArchivedRun.serializer(), f.readText()) }.getOrNull() val meta = runCatching { json.decodeFromString(ArchivedRun.serializer(), f.readText()) }.getOrNull()
?: return ?: return
@@ -127,7 +140,7 @@ class RunArchive(private val dir: File, private val now: () -> Long = System::cu
f, f,
json.encodeToString( json.encodeToString(
ArchivedRun.serializer(), ArchivedRun.serializer(),
meta.copy(uploaded = true, uploadedTo = serverName), meta.copy(uploaded = true, uploadedTo = serverName, uploadedAs = uploadedAs),
), ),
) )
} }
@@ -181,7 +194,10 @@ class RunArchive(private val dir: File, private val now: () -> Long = System::cu
id = id, id = id,
savedAtEpochMs = now(), savedAtEpochMs = now(),
startedAt = run["started_at"]?.jsonPrimitive?.content, startedAt = run["started_at"]?.jsonPrimitive?.content,
verdict = doc["summary"]?.jsonObject?.get("verdict")?.jsonPrimitive?.content, // The schema calls it `overall` (Summary.overall); reading `verdict` here silently
// yielded null for every run, so the history list's most prominent element - the
// coloured verdict - was blank on every row.
verdict = doc["summary"]?.jsonObject?.get("overall")?.jsonPrimitive?.content,
findingCount = (doc["findings"] as? kotlinx.serialization.json.JsonArray)?.size ?: 0, findingCount = (doc["findings"] as? kotlinx.serialization.json.JsonArray)?.size ?: 0,
sizeBytes = size, sizeBytes = size,
anonymization = run["privacy"]?.jsonObject?.get("anonymization")?.jsonPrimitive?.content ?: "full", anonymization = run["privacy"]?.jsonObject?.get("anonymization")?.jsonPrimitive?.content ?: "full",
@@ -25,7 +25,7 @@ class RunArchiveTest {
private fun doc(id: String, findings: Int = 1, pad: Int = 0): String { private fun doc(id: String, findings: Int = 1, pad: Int = 0): String {
val f = (1..findings).joinToString(",") { """{"id":"f$it"}""" } val f = (1..findings).joinToString(",") { """{"id":"f$it"}""" }
return """{"run":{"id":"$id","started_at":"2026-08-01T10:00:00Z","privacy":{"anonymization":"balanced"}},""" + return """{"run":{"id":"$id","started_at":"2026-08-01T10:00:00Z","privacy":{"anonymization":"balanced"}},""" +
""""findings":[$f],"summary":{"verdict":"warn"},"pad":"${"x".repeat(pad)}"}""" """"findings":[$f],"summary":{"overall":"warn"},"pad":"${"x".repeat(pad)}"}"""
} }
@Test @Test
@@ -120,13 +120,40 @@ class RunArchiveTest {
fun uploadStateIsRecorded() { fun uploadStateIsRecorded() {
val a = archive() val a = archive()
a.save(doc("run-1")) a.save(doc("run-1"))
a.markUploaded("run-1", "fmr") a.markUploaded("run-1", "fmr", "balanced")
val meta = a.list().single() val meta = a.list().single()
assertTrue(meta.uploaded) assertTrue(meta.uploaded)
assertEquals("fmr", meta.uploadedTo) assertEquals("fmr", meta.uploadedTo)
assertEquals("run-1", meta.id, "marking upload must not disturb the rest of the entry") assertEquals("run-1", meta.id, "marking upload must not disturb the rest of the entry")
} }
// The archive's own level and the level a run was uploaded at describe *different documents*.
// Showing the archive's ("full", because the archive is deliberately unredacted) next to
// "uploaded to fmr" reads as "the complete data was uploaded" when a redacted copy was sent —
// a privacy display that overstates what left the device is worse than none.
@Test
fun theUploadedLevelIsRecordedSeparatelyFromTheArchivedOne() {
val a = archive()
// A real archived document carries no privacy stamp: the anonymizer never runs on the
// archive. The shared doc() fixture has one, which is exactly the unrealism that let this
// confusion through in the first place.
a.save("""{"run":{"id":"run-1"},"findings":[],"summary":{"overall":"green"}}""")
a.markUploaded("run-1", "fmr", "balanced")
val meta = a.list().single()
assertEquals("full", meta.anonymization, "the archived copy is unredacted, by design")
assertEquals("balanced", meta.uploadedAs, "the uploaded copy was redacted, and must say so")
}
// The verdict is read from `summary.overall` — the schema's actual field name. Reading
// `summary.verdict` silently yielded null for every run, so the history list's most prominent
// element was blank on every row while everything else looked fine.
@Test
fun theVerdictComesFromTheSchemasOverallField() {
val a = archive()
a.save("""{"run":{"id":"r1"},"findings":[],"summary":{"overall":"yellow"}}""")
assertEquals("yellow", a.list().single().verdict)
}
@Test @Test
fun deleteRemovesBothFiles() { fun deleteRemovesBothFiles() {
val a = archive() val a = archive()
@@ -0,0 +1,107 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.engine
import kotlinx.serialization.SerialName
import kotlinx.serialization.Serializable
/**
* Splits a round-trip train into its two directions using what the server witnessed.
*
* A round trip can only report that *something* was lost somewhere. That is the least useful form
* of the answer: "3 % loss" sends an engineer looking in both directions at once. The server
* records every packet it received, per sequence number (probe-protocol.md §6), so the two cases
* are actually distinguishable:
*
* - sent, never seen by the server → **upstream** loss
* - seen by the server, reply never arrived → **downstream** loss
*
* The same records give one-way delay *variation* per direction. Absolute one-way delay would
* need synchronised clocks and we deliberately have none (measurement-schema.md's two-clock rule),
* but the variation does not: (server_rx client_tx) contains an unknown constant clock offset,
* and differencing successive samples cancels it. So jitter is honestly attributable to a
* direction even though latency is not.
*/
object Directional {
/** One probe as the client saw it. [tRxNs] null means no reply came back. */
data class Sample(val seq: Int, val tTxNs: Long, val tRxNs: Long?)
/** One probe as the server saw it: its own receive and transmit stamps, on its own clock. */
data class ServerSighting(val seq: Int, val tRxNs: Long, val tTxNs: Long)
fun analyse(sent: List<Sample>, seen: List<ServerSighting>): DirectionalMetrics {
val byServerSeq = seen.associateBy { it.seq }
// Only sequences we actually sent count. A server record for a sequence we have no note
// of is not evidence about this train — it is a bug or a stray, and silently folding it
// in would produce loss percentages above 100 or below zero.
val relevant = sent.filter { byServerSeq.containsKey(it.seq) }
val nSent = sent.size
val nSeen = relevant.size
val nReplied = sent.count { it.tRxNs != null }
// A reply can only exist if the request arrived, so downstream loss is measured against
// what the server saw, not against what we sent — otherwise upstream loss is counted twice.
val lostUp = nSent - nSeen
val lostDown = (nSeen - nReplied).coerceAtLeast(0)
val upDeltas = relevant.sortedBy { it.seq }
.map { byServerSeq.getValue(it.seq).tRxNs - it.tTxNs }
val downDeltas = sent.filter { it.tRxNs != null && byServerSeq.containsKey(it.seq) }
.sortedBy { it.seq }
.map { it.tRxNs!! - byServerSeq.getValue(it.seq).tTxNs }
return DirectionalMetrics(
sent = nSent,
seenByServer = nSeen,
repliesReceived = nReplied,
lostUpstream = lostUp,
lostDownstream = lostDown,
lossUpstreamPct = pct(lostUp, nSent),
// Denominator is what reached the server: of the packets that got there, how many
// replies came back.
lossDownstreamPct = pct(lostDown, nSeen),
jitterUpstreamMs = jitterMs(upDeltas),
jitterDownstreamMs = jitterMs(downDeltas),
/** True when the server saw nothing at all, which is a different fault from loss. */
noneReachedServer = nSent > 0 && nSeen == 0,
)
}
/**
* Mean absolute difference between consecutive one-way samples (RFC 3393 IPDV, averaged).
*
* Differencing is what makes this legitimate without synchronised clocks: each sample carries
* the same unknown offset between the two clocks, and the difference cancels it. Fewer than
* two samples yields null rather than zero — "no jitter" and "not enough data to say" are
* different claims and only one of them is true here.
*/
private fun jitterMs(oneWayNs: List<Long>): Double? {
if (oneWayNs.size < 2) return null
val deltas = oneWayNs.zipWithNext { a, b -> kotlin.math.abs(b - a) }
return round2(deltas.average() / 1_000_000.0)
}
private fun pct(part: Int, whole: Int): Double =
if (whole <= 0) 0.0 else round2(part * 100.0 / whole)
private fun round2(v: Double) = Math.round(v * 100.0) / 100.0
}
/** Directional metrics for train.udp_updown; recomputable from the columnar evidence. */
@Serializable
data class DirectionalMetrics(
val sent: Int,
@SerialName("seen_by_server") val seenByServer: Int,
@SerialName("replies_received") val repliesReceived: Int,
@SerialName("lost_upstream") val lostUpstream: Int,
@SerialName("lost_downstream") val lostDownstream: Int,
@SerialName("loss_upstream_pct") val lossUpstreamPct: Double,
@SerialName("loss_downstream_pct") val lossDownstreamPct: Double,
/** One-way delay variation (RFC 3393), per direction. Null when there were too few samples. */
@SerialName("jitter_upstream_ms") val jitterUpstreamMs: Double? = null,
@SerialName("jitter_downstream_ms") val jitterDownstreamMs: Double? = null,
@SerialName("none_reached_server") val noneReachedServer: Boolean = false,
)
@@ -37,6 +37,120 @@ class DownstreamMeasurement(private val ids: IdSource) {
/** How long to wait for a granted burst after the server accepts the action. */ /** How long to wait for a granted burst after the server accepts the action. */
private val collectWindowMs = 4_000L private val collectWindowMs = 4_000L
/**
* Shorter, but long enough to cover the first_last mode's deliberate 250 ms hold plus a
* reassembly. A fragment burst is one datagram: it is here quickly or not at all.
*/
private val fragWindowMs = 1_500L
/**
* Asks the server to send one deliberately-fragmented datagram per ordering, and reports
* which orderings survive the path.
*
* Kernel fragmentation always emits fragments in order, first one first, so an oversized
* datagram can only answer "do fragments get through at all". The interesting fault is about
* ordering: only the *first* fragment carries the UDP ports, so a stateful firewall that has
* not seen it has nothing to match the rest against, and many drop them. That failure is
* invisible to every in-order test and shows up in the field as "large DNS answers fail here"
* or "the tunnel breaks when the MTU drops".
*/
fun fragmentOrdering(
credential: String,
sessionId: String,
control: ControlClient,
probe: ProbeSession,
sessionRef: String,
sizeBytes: Int = 2000,
fragBytes: Int = 576,
): Pair<Test, List<Finding>> {
val testId = ids.uuid()
val started = ids.monoNs()
val delivered = LinkedHashMap<String, Boolean>()
val fragmentCounts = LinkedHashMap<String, Int>()
var unsupported = false
for (mode in FRAG_MODES) {
val reply = runCatching {
control.action(
credential, sessionId,
"""{"action":"frag_send","size_bytes":$sizeBytes,"mode":"$mode","frag_bytes":$fragBytes}""",
)
}
if (reply.isFailure) {
// A server without a raw socket says so; that is a missing capability, not a
// property of the network, and must not be recorded as a failed delivery.
unsupported = true
break
}
parseInt(reply.getOrNull(), "fragments")?.let { fragmentCounts[mode] = it }
// The burst is already on the wire when the action returns (it is sent
// synchronously), so anything that survived is either here or lost.
val got = probe.collectGranted(fragWindowMs).any { it.type == Wire.TYPE_FRAG_DATA }
delivered[mode] = got
}
if (unsupported) {
return Test(
id = testId, type = TestType.MTU_FRAG_ORDERING, sessionRef = sessionRef, tier = Tier.APP,
startedMonoNs = started, endedMonoNs = ids.monoNs(),
status = TestStatus.UNSUPPORTED,
error = TestError("no_raw_socket", "this server cannot craft fragments"),
) to emptyList()
}
val metrics = json.encodeToJsonElement(
FragOrderingMetrics(
sizeBytes = sizeBytes,
fragBytes = fragBytes,
fragmentsPerBurst = fragmentCounts,
deliveredByMode = delivered,
inOrderDelivered = delivered[FRAG_IN_ORDER] == true,
reorderedDelivered = delivered[FRAG_REVERSED] == true,
delayedFirstDelivered = delivered[FRAG_FIRST_LAST] == true,
),
) as JsonObject
val findings = ArrayList<Finding>()
val inOrder = delivered[FRAG_IN_ORDER] == true
val reversed = delivered[FRAG_REVERSED] == true
val firstLast = delivered[FRAG_FIRST_LAST] == true
if (!inOrder) {
findings.add(
finding(
FindingRegistry.FRAGMENTS_BLOCKED, testId,
"IP fragments do not reach this device",
"A fragmented datagram sent in the normal order never arrived. Anything that " +
"relies on fragmentation — large DNS answers over UDP, some VPN traffic — " +
"will fail here rather than slow down.",
),
)
} else if (!reversed || !firstLast) {
// The precise and useful finding: fragments work, but only if they arrive tidily.
val which = buildList {
if (!reversed) add("out of order")
if (!firstLast) add("with the first fragment delayed")
}.joinToString(" or ")
findings.add(
finding(
FindingRegistry.FRAGMENT_REORDER_SENSITIVE, testId,
"Fragments are dropped when they arrive $which",
"In-order fragments are delivered, but the same datagram sent $which is not. " +
"Something on the path only reassembles when the first fragment (the one " +
"carrying the UDP ports) arrives first — typical of a stateful firewall " +
"or NAT. It works until the network reorders, then fails intermittently, " +
"which is the hardest kind of fault to chase.",
),
)
}
return Test(
id = testId, type = TestType.MTU_FRAG_ORDERING, sessionRef = sessionRef, tier = Tier.APP,
startedMonoNs = started, endedMonoNs = ids.monoNs(),
status = if (inOrder) TestStatus.OK else TestStatus.PARTIAL,
metrics = metrics,
) to findings
}
/** /**
* Runs all three against an already-primed session. * Runs all three against an already-primed session.
* *
@@ -54,6 +168,10 @@ class DownstreamMeasurement(private val ids: IdSource) {
trainCount: Int = 100, trainCount: Int = 100,
trainSizeBytes: Int = 300, trainSizeBytes: Int = 300,
trainIntervalUs: Int = 3_000, trainIntervalUs: Int = 3_000,
// DSCP to mark the train with (0-63), or -1 to leave packets unmarked. Pairing a marked
// downtrain with the server-observed DSCP of an upstream train is the two-direction
// sec.dscp_ecn_survival measurement.
trainDscp: Int = -1,
): Pair<List<Test>, List<Finding>> { ): Pair<List<Test>, List<Finding>> {
val tests = ArrayList<Test>() val tests = ArrayList<Test>()
val findings = ArrayList<Finding>() val findings = ArrayList<Finding>()
@@ -61,11 +179,22 @@ class DownstreamMeasurement(private val ids: IdSource) {
val df = bigSend(credential, sessionId, control, probe, sessionRef, sizes, df = true) val df = bigSend(credential, sessionId, control, probe, sessionRef, sizes, df = true)
val frag = bigSend(credential, sessionId, control, probe, sessionRef, sizes, df = false) val frag = bigSend(credential, sessionId, control, probe, sessionRef, sizes, df = false)
val train = downTrain( val train = downTrain(
credential, sessionId, control, probe, sessionRef, trainCount, trainSizeBytes, trainIntervalUs, credential, sessionId, control, probe, sessionRef, trainCount, trainSizeBytes,
trainIntervalUs, trainDscp,
) )
tests.add(df.test); tests.add(frag.test); tests.add(train.test) tests.add(df.test); tests.add(frag.test); tests.add(train.test)
// Fragment ordering only makes sense once we know fragments arrive at all; when they do
// not, the ordering variants would all report "not delivered" and read as three faults
// instead of one.
if (frag.largestDelivered != null) {
val (fragTest, fragFindings) =
fragmentOrdering(credential, sessionId, control, probe, sessionRef)
tests.add(fragTest)
findings.addAll(fragFindings)
}
// A downstream MTU below the classic 1500-byte Ethernet payload is worth saying out loud: // A downstream MTU below the classic 1500-byte Ethernet payload is worth saying out loud:
// it is the usual cause of "small requests work, large responses hang". // it is the usual cause of "small requests work, large responses hang".
val pathMtu = df.largestDelivered val pathMtu = df.largestDelivered
@@ -74,7 +203,7 @@ class DownstreamMeasurement(private val ids: IdSource) {
if (ipMtu < 1500) { if (ipMtu < 1500) {
findings.add( findings.add(
finding( finding(
"mtu.reduced_downstream", Category.MTU, Severity.LOW, df.test.id, FindingRegistry.MTU_REDUCED_DOWNSTREAM, df.test.id,
"Downstream path MTU is $ipMtu bytes, below 1500", "Downstream path MTU is $ipMtu bytes, below 1500",
"The largest datagram that reached this device without fragmenting was " + "The largest datagram that reached this device without fragmenting was " +
"$pathMtu bytes of payload ($ipMtu on the wire). Tunnels (PPPoE, VPN, " + "$pathMtu bytes of payload ($ipMtu on the wire). Tunnels (PPPoE, VPN, " +
@@ -89,7 +218,7 @@ class DownstreamMeasurement(private val ids: IdSource) {
if (fragLargest <= pathMtu && sizes.any { it > pathMtu }) { if (fragLargest <= pathMtu && sizes.any { it > pathMtu }) {
findings.add( findings.add(
finding( finding(
"mtu.downstream_blackhole", Category.MTU, Severity.MEDIUM, frag.test.id, FindingRegistry.MTU_DOWNSTREAM_BLACKHOLE, frag.test.id,
"Datagrams above $pathMtu bytes are dropped downstream, fragmented or not", "Datagrams above $pathMtu bytes are dropped downstream, fragmented or not",
"Nothing larger than $pathMtu bytes arrived, even when the network was " + "Nothing larger than $pathMtu bytes arrived, even when the network was " +
"free to fragment it. Traffic that relies on large responses will " + "free to fragment it. Traffic that relies on large responses will " +
@@ -102,7 +231,7 @@ class DownstreamMeasurement(private val ids: IdSource) {
if (train.received == 0) { if (train.received == 0) {
findings.add( findings.add(
finding( finding(
"connectivity.downstream_blocked", Category.CONNECTIVITY, Severity.HIGH, train.test.id, FindingRegistry.DOWNSTREAM_BLOCKED, train.test.id,
"No server-initiated packets arrived", "No server-initiated packets arrived",
"The server sent ${train.sent} packets toward this device and none arrived, " + "The server sent ${train.sent} packets toward this device and none arrived, " +
"while the round-trip echo worked. Something on the path forwards replies " + "while the round-trip echo worked. Something on the path forwards replies " +
@@ -112,7 +241,7 @@ class DownstreamMeasurement(private val ids: IdSource) {
} else if (train.lossPct >= 5.0) { } else if (train.lossPct >= 5.0) {
findings.add( findings.add(
finding( finding(
"connectivity.downstream_loss", Category.CONNECTIVITY, Severity.MEDIUM, train.test.id, FindingRegistry.LOSS_DOWNSTREAM, train.test.id,
"Downstream loss of ${round1(train.lossPct)}%", "Downstream loss of ${round1(train.lossPct)}%",
"${train.sent - train.received} of ${train.sent} packets sent toward this " + "${train.sent - train.received} of ${train.sent} packets sent toward this " +
"device were lost. Downstream loss is invisible to a round-trip test, " + "device were lost. Downstream loss is invisible to a round-trip test, " +
@@ -123,7 +252,7 @@ class DownstreamMeasurement(private val ids: IdSource) {
if (train.reordered > 0) { if (train.reordered > 0) {
findings.add( findings.add(
finding( finding(
"connectivity.downstream_reorder", Category.CONNECTIVITY, Severity.LOW, train.test.id, FindingRegistry.DOWNSTREAM_REORDER, train.test.id,
"${train.reordered} downstream packet(s) arrived out of order", "${train.reordered} downstream packet(s) arrived out of order",
"Packets arrived in a different order than they were sent. Usually per-packet " + "Packets arrived in a different order than they were sent. Usually per-packet " +
"load balancing across links; harmless for most traffic, not for all of it.", "load balancing across links; harmless for most traffic, not for all of it.",
@@ -214,15 +343,19 @@ class DownstreamMeasurement(private val ids: IdSource) {
private fun downTrain( private fun downTrain(
credential: String, sessionId: String, control: ControlClient, probe: ProbeSession, credential: String, sessionId: String, control: ControlClient, probe: ProbeSession,
sessionRef: String, count: Int, sizeBytes: Int, intervalUs: Int, sessionRef: String, count: Int, sizeBytes: Int, intervalUs: Int, dscp: Int = -1,
): TrainResult { ): TrainResult {
val testId = ids.uuid() val testId = ids.uuid()
val started = ids.monoNs() val started = ids.monoNs()
// dscp is only sent when requested: an older server rejects unknown-value problems
// louder than absent keys, and unmarked is the correct default for a plain loss train.
val dscpField = if (dscp in 0..63) ""","dscp":$dscp""" else ""
val reply = runCatching { val reply = runCatching {
control.action( control.action(
credential, sessionId, credential, sessionId,
"""{"action":"downtrain","count":$count,"size_bytes":$sizeBytes,"interval_us":$intervalUs}""", """{"action":"downtrain","count":$count,"size_bytes":$sizeBytes,""" +
""""interval_us":$intervalUs$dscpField}""",
) )
} }
if (reply.isFailure) { if (reply.isFailure) {
@@ -270,6 +403,12 @@ class DownstreamMeasurement(private val ids: IdSource) {
interArrivalMsAvg = interArrival.average().takeIf { interArrival.isNotEmpty() }?.let(::round1), interArrivalMsAvg = interArrival.average().takeIf { interArrival.isNotEmpty() }?.let(::round1),
interArrivalMsMax = interArrival.maxOrNull()?.let(::round1), interArrivalMsMax = interArrival.maxOrNull()?.let(::round1),
sendIntervalUs = intervalUs, sendIntervalUs = intervalUs,
dscpRequested = dscp.takeIf { it in 0..63 },
// The server says whether it could actually mark (dscp_applied); recorded so a
// survival comparison never blames the path for a marking the sender skipped.
dscpApplied = reply.getOrNull()?.let {
Regex("\"dscp_applied\"\\s*:\\s*(true|false)").find(it)?.groupValues?.get(1)?.toBoolean()
},
), ),
) as JsonObject ) as JsonObject
@@ -290,9 +429,17 @@ class DownstreamMeasurement(private val ids: IdSource) {
// ---- helpers ---------------------------------------------------------------------- // ---- helpers ----------------------------------------------------------------------
private fun finding(code: String, cat: Category, sev: Severity, testId: String, title: String, desc: String) = /**
* Builds a finding from a registry entry, which supplies the code, category and severity.
*
* Taking a [FindingSpec] rather than three loose values is the point: a typo becomes a
* compile error, and two call sites cannot disagree about which category a finding belongs
* to - a disagreement that would split one fault across two verdict lights.
*/
private fun finding(spec: FindingSpec, testId: String, title: String, desc: String) =
Finding( Finding(
id = ids.uuid(), code = code, category = cat, severity = sev, confidence = Confidence.HIGH, id = ids.uuid(), code = spec.code, category = spec.category, severity = spec.severity,
confidence = Confidence.HIGH,
title = title, description = desc, evidenceRefs = listOf(EvidenceRef(testId)), title = title, description = desc, evidenceRefs = listOf(EvidenceRef(testId)),
) )
@@ -310,6 +457,11 @@ class DownstreamMeasurement(private val ids: IdSource) {
/** IPv4 (20) + UDP (8). The v6 case is 48; reported per-family once v6 sessions land. */ /** IPv4 (20) + UDP (8). The v6 case is 48; reported per-family once v6 sessions land. */
const val IP_UDP_OVERHEAD4 = 28 const val IP_UDP_OVERHEAD4 = 28
const val FRAG_IN_ORDER = "in_order"
const val FRAG_REVERSED = "reversed"
const val FRAG_FIRST_LAST = "first_last"
val FRAG_MODES = listOf(FRAG_IN_ORDER, FRAG_REVERSED, FRAG_FIRST_LAST)
/** Straddles the usual suspects: 1500 Ethernet, 1492 PPPoE, 1400-ish tunnels. */ /** Straddles the usual suspects: 1500 Ethernet, 1492 PPPoE, 1400-ish tunnels. */
val DEFAULT_SIZES = listOf(600, 1200, 1372, 1400, 1450, 1472, 1500, 2000, 4000) val DEFAULT_SIZES = listOf(600, 1200, 1372, 1400, 1450, 1472, 1500, 2000, 4000)
@@ -330,6 +482,18 @@ data class BigSendMetrics(
@SerialName("path_mtu_bytes") val pathMtuBytes: Int? = null, @SerialName("path_mtu_bytes") val pathMtuBytes: Int? = null,
) )
/** Metrics for mtu.frag_ordering. */
@Serializable
data class FragOrderingMetrics(
@SerialName("size_bytes") val sizeBytes: Int,
@SerialName("frag_bytes") val fragBytes: Int,
@SerialName("fragments_per_burst") val fragmentsPerBurst: Map<String, Int>,
@SerialName("delivered_by_mode") val deliveredByMode: Map<String, Boolean>,
@SerialName("in_order_delivered") val inOrderDelivered: Boolean,
@SerialName("reordered_delivered") val reorderedDelivered: Boolean,
@SerialName("delayed_first_delivered") val delayedFirstDelivered: Boolean,
)
/** Metrics for train.udp_downstream. */ /** Metrics for train.udp_downstream. */
@Serializable @Serializable
data class DownTrainMetrics( data class DownTrainMetrics(
@@ -341,4 +505,6 @@ data class DownTrainMetrics(
@SerialName("inter_arrival_ms_avg") val interArrivalMsAvg: Double? = null, @SerialName("inter_arrival_ms_avg") val interArrivalMsAvg: Double? = null,
@SerialName("inter_arrival_ms_max") val interArrivalMsMax: Double? = null, @SerialName("inter_arrival_ms_max") val interArrivalMsMax: Double? = null,
@SerialName("send_interval_us") val sendIntervalUs: Int, @SerialName("send_interval_us") val sendIntervalUs: Int,
@SerialName("dscp_requested") val dscpRequested: Int? = null,
@SerialName("dscp_applied") val dscpApplied: Boolean? = null,
) )
@@ -10,8 +10,13 @@ import app.echo_lot.measurement.*
import app.echo_lot.protocol.ControlClient import app.echo_lot.protocol.ControlClient
import app.echo_lot.protocol.ProbeSession import app.echo_lot.protocol.ProbeSession
import kotlinx.serialization.json.Json import kotlinx.serialization.json.Json
import kotlinx.serialization.json.JsonArray
import kotlinx.serialization.json.JsonObject import kotlinx.serialization.json.JsonObject
import kotlinx.serialization.json.encodeToJsonElement import kotlinx.serialization.json.encodeToJsonElement
import kotlinx.serialization.json.intOrNull
import kotlinx.serialization.json.jsonObject
import kotlinx.serialization.json.jsonPrimitive
import kotlinx.serialization.json.longOrNull
/** /**
* Runs the server-facing measurements against one target and assembles a [MeasurementDocument]: * Runs the server-facing measurements against one target and assembles a [MeasurementDocument]:
@@ -42,6 +47,14 @@ class ServerMeasurement(
* it is a flag rather than an assumption. * it is a flag rather than an assumption.
*/ */
val downstream: Boolean = true, val downstream: Boolean = true,
/**
* Throughput moves real data — a 5-second run at 50 Mbps is about 30 MB — so it is off
* unless asked for. On a metered mobile connection that is the user's money, and a
* measurement tool that spends it without being told to is not one people keep installed.
*/
val throughput: Boolean = false,
@Suppress("unused") val throughputSeconds: Int = 5,
@Suppress("unused") val throughputKbps: Int = 50_000,
) )
fun run(cfg: Config): MeasurementDocument { fun run(cfg: Config): MeasurementDocument {
@@ -71,10 +84,18 @@ class ServerMeasurement(
// re-primed source is never recorded and every granted send goes to the old, closed port. // re-primed source is never recorded and every granted send goes to the old, closed port.
// Session identity lives on the server; the socket must live as long as it does. // Session identity lives on the server; the socket must live as long as it does.
ProbeSession(cfg.credential, session, cfg.udpHost, cfg.udpPort).use { ps -> ProbeSession(cfg.credential, session, cfg.udpHost, cfg.udpPort).use { ps ->
val (test, findings) = echoTrain(cfg, ps, startMono) val (test, findings) = echoTrain(cfg, ps, startMono, control, session.sessionId)
tests.add(test) tests.add(test)
allFindings.addAll(findings) allFindings.addAll(findings)
// The upstream train needs no grant and no capability beyond udp-probe itself; a
// server that predates trains simply never answers the report request, which the
// measurement reports as exactly that ambiguity rather than as network loss.
val (utTest, utFindings) = UpstreamTrainMeasurement(ids)
.run(ps, sessionRef = "sess-1")
tests.add(utTest)
allFindings.addAll(utFindings)
// Downstream needs a session the server has already seen traffic from — the echo // Downstream needs a session the server has already seen traffic from — the echo
// train just provided that — and a server that advertises the grants. Skipped // train just provided that — and a server that advertises the grants. Skipped
// quietly against an older server rather than reported as a failure of the network. // quietly against an older server rather than reported as a failure of the network.
@@ -84,6 +105,15 @@ class ServerMeasurement(
tests.addAll(dsTests) tests.addAll(dsTests)
allFindings.addAll(dsFindings) allFindings.addAll(dsFindings)
} }
if (cfg.throughput && profile.supports("throughput")) {
val (tpTest, tpFindings) = ThroughputMeasurement(ids).run(
cfg.credential, session.sessionId, control, ps, sessionRef = "sess-1",
durationS = cfg.throughputSeconds, kbps = cfg.throughputKbps,
)
tests.add(tpTest)
allFindings.addAll(tpFindings)
}
} }
control.deleteSession(cfg.credential, session.sessionId) control.deleteSession(cfg.credential, session.sessionId)
@@ -105,6 +135,7 @@ class ServerMeasurement(
private fun echoTrain( private fun echoTrain(
cfg: Config, ps: ProbeSession, startMono: Long, cfg: Config, ps: ProbeSession, startMono: Long,
control: ControlClient? = null, sessionId: String? = null,
): Pair<Test, List<Finding>> { ): Pair<Test, List<Finding>> {
val testId = ids.uuid() val testId = ids.uuid()
val seqs = ArrayList<Int>() val seqs = ArrayList<Int>()
@@ -114,9 +145,15 @@ class ServerMeasurement(
val rtts = ArrayList<Double>() val rtts = ArrayList<Double>()
val observedPorts = LinkedHashSet<Int>() val observedPorts = LinkedHashSet<Int>()
// Wire sequence numbers, kept so the server's observations can be correlated packet by
// packet. They are not 0..n-1: the counter is shared with every other packet type on the
// session, so "the nth echo" is not "sequence n".
val wireSeqs = ArrayList<Int>()
for (i in 0 until cfg.echoCount) { for (i in 0 until cfg.echoCount) {
val txMono = ids.monoNs() - startMono val txMono = ids.monoNs() - startMono
val r = ps.echo(cfg.echoPaddingBytes) val r = ps.echo(cfg.echoPaddingBytes)
wireSeqs.add(ps.lastSeq)
seqs.add(i) seqs.add(i)
tTx.add(txMono) tTx.add(txMono)
sizes.add(Wire_HEADER + cfg.echoPaddingBytes) sizes.add(Wire_HEADER + cfg.echoPaddingBytes)
@@ -129,6 +166,20 @@ class ServerMeasurement(
} }
} }
// Ask the server what it actually received. This is what turns "3 % loss somewhere" into
// "3 % loss upstream" - the least useful form of the answer into a usable one.
val directional: DirectionalMetrics? =
if (control != null && sessionId != null) {
runCatching {
val samples = wireSeqs.indices.map {
Directional.Sample(wireSeqs[it], tTx[it] ?: 0L, tRx[it])
}
Directional.analyse(samples, serverSightings(control, cfg, sessionId))
}.getOrNull() // an older server without the endpoint simply yields no split
} else {
null
}
val sent = cfg.echoCount val sent = cfg.echoCount
val received = rtts.size val received = rtts.size
val lossPct = if (sent == 0) 0.0 else (sent - received) * 100.0 / sent val lossPct = if (sent == 0) 0.0 else (sent - received) * 100.0 / sent
@@ -138,6 +189,9 @@ class ServerMeasurement(
epochMonoNs = startMono, seq = seqs, tTxNs = tTx, tRxNs = tRx, sizeBytes = sizes, epochMonoNs = startMono, seq = seqs, tTxNs = tTx, tRxNs = tRx, sizeBytes = sizes,
).toEvidence() ).toEvidence()
val directionalJson = directional?.let {
json.encodeToJsonElement(DirectionalMetrics.serializer(), it) as JsonObject
}
val metrics: JsonObject = json.encodeToJsonElement( val metrics: JsonObject = json.encodeToJsonElement(
EchoMetrics( EchoMetrics(
sent = sent, received = received, lossPct = round1(lossPct), sent = sent, received = received, lossPct = round1(lossPct),
@@ -147,7 +201,7 @@ class ServerMeasurement(
observedPorts = observedPorts.toList(), observedPorts = observedPorts.toList(),
natRebindingDetected = natRebinding, natRebindingDetected = natRebinding,
) )
) as JsonObject ).let { base -> JsonObject((base as JsonObject) + (directionalJson ?: JsonObject(emptyMap()))) }
val status = when { val status = when {
received == 0 -> TestStatus.FAILED received == 0 -> TestStatus.FAILED
@@ -162,30 +216,91 @@ class ServerMeasurement(
val findings = ArrayList<Finding>() val findings = ArrayList<Finding>()
if (received == 0) { if (received == 0) {
findings.add(finding("nat.udp_unreachable", Category.CONNECTIVITY, Severity.HIGH, testId, findings.add(finding(FindingRegistry.UDP_UNREACHABLE, testId,
"No UDP echo replies from the server", "No UDP echo replies from the server",
"Every ECHO probe to the server's UDP data plane was lost — the path blocks or drops the session's UDP traffic.")) "Every ECHO probe to the server's UDP data plane was lost — the path blocks or drops the session's UDP traffic."))
} else if (lossPct >= 20.0) { } else if (lossPct >= 20.0) {
findings.add(finding("connectivity.udp_loss", Category.CONNECTIVITY, Severity.MEDIUM, testId, findings.add(finding(FindingRegistry.UDP_LOSS, testId,
"High UDP loss to the server (${round1(lossPct)}%)", "High UDP loss to the server (${round1(lossPct)}%)",
"A large fraction of ECHO probes were lost, indicating an unreliable UDP path.")) "A large fraction of ECHO probes were lost, indicating an unreliable UDP path."))
} }
// Naming the direction is the entire value of the split, so the findings do.
directional?.let { d ->
when {
d.noneReachedServer && received == 0 -> findings.add(
finding(FindingRegistry.UDP_UNREACHABLE_UPSTREAM, testId,
"Nothing reached the server",
"The server received none of the ${d.sent} probes, so the traffic is being " +
"dropped on the way out, not on the way back. A firewall or NAT on " +
"this side of the path is the place to look."),
)
d.lossUpstreamPct >= 2.0 -> findings.add(
finding(FindingRegistry.LOSS_UPSTREAM, testId,
"${d.lossUpstreamPct} % of probes were lost on the way to the server",
"${d.lostUpstream} of ${d.sent} probes never reached the server. The " +
"return path is not implicated: replies came back for everything that " +
"arrived."),
)
}
if (d.lossDownstreamPct >= 2.0) {
findings.add(
finding(FindingRegistry.LOSS_DOWNSTREAM, testId,
"${d.lossDownstreamPct} % of replies were lost on the way back",
"The server received ${d.seenByServer} probes and answered them, but " +
"${d.lostDownstream} of those replies never arrived. The outbound path " +
"is fine; the fault is on the return leg."),
)
}
}
if (natRebinding) { if (natRebinding) {
findings.add(finding("nat.udp_rebinding", Category.NAT, Severity.MEDIUM, testId, findings.add(finding(FindingRegistry.NAT_UDP_REBINDING, testId,
"NAT remapped the UDP source port mid-flow", "NAT remapped the UDP source port mid-flow",
"The server observed more than one source port for this session (${observedPorts.joinToString()}), i.e. a NAT with a short UDP mapping or per-packet remapping.")) "The server observed more than one source port for this session (${observedPorts.joinToString()}), i.e. a NAT with a short UDP mapping or per-packet remapping."))
} }
return test to findings return test to findings
} }
private fun finding(code: String, cat: Category, sev: Severity, testId: String, title: String, desc: String) = /**
* The server's per-packet record of this session's echoes (spec section 6). Filtered to
* ECHO_REQ, because the observation list also holds MTU probes and anything else we sent -
* counting those as train packets would invent loss that is not there.
*/
private fun serverSightings(
control: ControlClient, cfg: Config, sessionId: String,
): List<Directional.ServerSighting> {
val body = control.observations(cfg.credential, sessionId)
val packets = Json.parseToJsonElement(body).jsonObject["udp"]
?.jsonObject?.get("packets") as? JsonArray ?: return emptyList()
return packets.mapNotNull { el ->
val o = el as? JsonObject ?: return@mapNotNull null
val type = o["type"]?.jsonPrimitive?.intOrNull ?: return@mapNotNull null
if (type != ECHO_REQ_TYPE) return@mapNotNull null
Directional.ServerSighting(
seq = o["seq"]?.jsonPrimitive?.intOrNull ?: return@mapNotNull null,
tRxNs = o["t_rx_ns"]?.jsonPrimitive?.longOrNull ?: return@mapNotNull null,
tTxNs = o["t_tx_ns"]?.jsonPrimitive?.longOrNull ?: return@mapNotNull null,
)
}
}
/**
* Builds a finding from a registry entry, which supplies the code, category and severity.
*
* Taking a [FindingSpec] rather than three loose values is the point: a typo becomes a
* compile error, and two call sites cannot disagree about which category a finding belongs
* to - a disagreement that would split one fault across two verdict lights.
*/
private fun finding(spec: FindingSpec, testId: String, title: String, desc: String) =
Finding( Finding(
id = ids.uuid(), code = code, category = cat, severity = sev, confidence = Confidence.HIGH, id = ids.uuid(), code = spec.code, category = spec.category, severity = spec.severity,
confidence = Confidence.HIGH,
title = title, description = desc, evidenceRefs = listOf(EvidenceRef(testId)), title = title, description = desc, evidenceRefs = listOf(EvidenceRef(testId)),
) )
private companion object { private companion object {
const val Wire_HEADER = 32 const val Wire_HEADER = 32
const val ECHO_REQ_TYPE = 0x01
fun round1(v: Double) = Math.round(v * 10.0) / 10.0 fun round1(v: Double) = Math.round(v * 10.0) / 10.0
} }
} }
@@ -0,0 +1,338 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.engine
import app.echo_lot.measurement.*
import app.echo_lot.protocol.ControlClient
import app.echo_lot.protocol.ProbeSession
import app.echo_lot.protocol.Wire
import kotlinx.serialization.SerialName
import kotlinx.serialization.Serializable
import kotlinx.serialization.json.Json
import kotlinx.serialization.json.JsonObject
import kotlinx.serialization.json.encodeToJsonElement
import kotlinx.serialization.json.jsonArray
import kotlinx.serialization.json.jsonObject
import kotlinx.serialization.json.jsonPrimitive
/**
* Downstream throughput: the server sends at a paced rate for a bounded time and the client
* measures what arrives (`perf.throughput_udp`).
*
* The number this produces is only meaningful with a qualifier attached, and getting that
* qualifier right is most of the work here. A throughput test reports the *smallest* limit on the
* path, and the sender's own ceiling is one of the candidates: if the server was asked for 50 Mbps
* and 50 Mbps arrived, the network was never the constraint and "50 Mbps" says nothing about it.
* Reporting that as a capacity measurement would be a confident lie, so the result always carries
* [ThroughputMetrics.limitedBy] and a finding is only raised when the network is actually
* implicated.
*
* Comparing against the *sender's* count rather than the requested rate is the other half: the
* server reports how much it actually put on the wire, and the gap between that and what arrived
* is the loss. A receiver alone cannot tell "the network dropped it" from "the sender never sent
* it", and guessing turns a healthy server-side limit into a phantom network fault.
*/
class ThroughputMeasurement(private val ids: IdSource) {
private val json = Json { encodeDefaults = true; explicitNulls = true }
fun run(
credential: String,
sessionId: String,
control: ControlClient,
probe: ProbeSession,
sessionRef: String,
durationS: Int = 5,
kbps: Int = 50_000,
sizeBytes: Int = 1200,
): Pair<Test, List<Finding>> {
val testId = ids.uuid()
val started = ids.monoNs()
val reply = runCatching {
control.action(
credential, sessionId,
"""{"action":"throughput","direction":"down","duration_s":$durationS,""" +
""""kbps":$kbps,"size_bytes":$sizeBytes}""",
)
}
if (reply.isFailure) {
return Test(
id = testId, type = TestType.PERF_THROUGHPUT_UDP, sessionRef = sessionRef, tier = Tier.APP,
startedMonoNs = started, endedMonoNs = ids.monoNs(),
status = TestStatus.UNSUPPORTED,
error = TestError("action_refused", reply.exceptionOrNull()?.message ?: "throughput refused"),
) to emptyList()
}
// The server may have shortened the run to fit its own byte budget; listen for what it
// actually promised, not for what we asked.
val plannedMs = parseInt(reply.getOrNull(), "duration_ms") ?: (durationS * 1000)
// A margin past the planned end so the tail of the run is not counted as loss: packets
// still in flight when we stop listening were not dropped, they were merely late.
val received = probe.collectGranted(plannedMs + 1_500L)
.filter { it.type == Wire.TYPE_THROUGHPUT_DATA }
val bytes = received.sumOf { it.sizeBytes.toLong() }
val spanNs = if (received.size >= 2) {
received.maxOf { it.tRxNs } - received.minOf { it.tRxNs }
} else {
0L
}
// Measured over the arrival span rather than our listening window, which includes the
// request round trip and the trailing margin and would understate the rate.
val receivedKbps = if (spanNs > 0) (bytes * 8 * 1_000_000 / spanNs).toInt() else 0
val sender = senderReport(control, credential, sessionId)
val sentPackets = sender?.packets ?: 0
val lossPct = if (sentPackets > 0) {
round2((sentPackets - received.size).coerceAtLeast(0) * 100.0 / sentPackets)
} else {
null
}
// Only a run the *clock* ended measured the network. One stopped by our own byte budget
// or rate ceiling measured this server.
val limitedBy = sender?.limitedBy ?: "unknown"
val networkLimited = limitedBy == "duration" &&
sender != null && receivedKbps > 0 && receivedKbps < sender.kbps * 9 / 10
val metrics = json.encodeToJsonElement(
ThroughputMetrics(
requestedKbps = kbps,
plannedDurationMs = plannedMs,
packetsReceived = received.size,
bytesReceived = bytes,
receivedKbps = receivedKbps,
senderPackets = sender?.packets,
senderBytes = sender?.bytes,
senderKbps = sender?.kbps,
lossPct = lossPct,
limitedBy = limitedBy,
measuresNetwork = networkLimited,
),
) as JsonObject
val findings = ArrayList<Finding>()
when {
sender == null -> Unit // no sender report: nothing can be concluded, so nothing is
received.isEmpty() -> findings.add(
finding(
FindingRegistry.THROUGHPUT_NO_DELIVERY, testId,
"No throughput traffic arrived",
"The server sent ${sender.packets} packets and none arrived. This is a " +
"connectivity fault rather than a slow link.",
),
)
networkLimited -> findings.add(
finding(
FindingRegistry.THROUGHPUT_BELOW_OFFERED, testId,
"Downstream throughput ${receivedKbps / 1000} Mbit/s, below the " +
"${sender.kbps / 1000} Mbit/s offered",
"The server sent at ${sender.kbps / 1000} Mbit/s for the full run and " +
"${receivedKbps / 1000} Mbit/s arrived" +
(lossPct?.let { ", losing $it % of packets" } ?: "") +
". The path could not carry what was offered.",
),
)
}
return Test(
id = testId, type = TestType.PERF_THROUGHPUT_UDP, sessionRef = sessionRef, tier = Tier.APP,
startedMonoNs = started, endedMonoNs = ids.monoNs(),
status = if (received.isEmpty()) TestStatus.FAILED else TestStatus.OK,
metrics = metrics,
) to findings
}
/**
* Upstream throughput: the client sends, the server counts.
*
* The mirror image of the downstream case, and it needs no grant — the client is generating
* its own traffic, so there is no amplification to gate. What it does need is the server's
* count: only the far end knows how much arrived, and without that number a sender can
* measure how fast it can *transmit*, which is not the same question and is usually just the
* speed of the local NIC.
*/
fun runUpstream(
credential: String,
sessionId: String,
control: ControlClient,
probe: ProbeSession,
sessionRef: String,
durationS: Int = 5,
kbps: Int = 20_000,
sizeBytes: Int = 1200,
): Pair<Test, List<Finding>> {
val testId = ids.uuid()
val started = ids.monoNs()
// Zeroes the server's counter so this run measures itself rather than inheriting the
// packets of an earlier one on the same session.
val reply = runCatching {
control.action(credential, sessionId, """{"action":"throughput","direction":"up"}""")
}
if (reply.isFailure) {
return Test(
id = testId, type = TestType.PERF_THROUGHPUT_UDP, sessionRef = sessionRef, tier = Tier.APP,
startedMonoNs = started, endedMonoNs = ids.monoNs(),
status = TestStatus.UNSUPPORTED,
error = TestError("action_refused", reply.exceptionOrNull()?.message ?: "refused"),
) to emptyList()
}
val sent = probe.sendThroughput(durationS * 1000L, kbps, sizeBytes)
// A moment for the tail of the run to arrive; counting still-in-flight packets as lost
// would inflate the loss figure by whatever the path's delay happens to be.
Thread.sleep(500)
val seen = upstreamCount(control, credential, sessionId)
val lossPct = if (sent.packets > 0 && seen != null) {
round2((sent.packets - seen.packets).coerceAtLeast(0) * 100.0 / sent.packets)
} else {
null
}
// The receiver's rate is the measurement. The sender's is what we managed to emit, which
// is a property of this phone and its radio, not of the network.
val achievedKbps = seen?.kbps ?: 0
val metrics = json.encodeToJsonElement(
UpstreamThroughputMetrics(
requestedKbps = kbps,
sentPackets = sent.packets,
sentBytes = sent.bytes,
sentKbps = sent.kbps,
receivedPackets = seen?.packets,
receivedBytes = seen?.bytes,
receivedKbps = achievedKbps,
lossPct = lossPct,
// Same honesty rule as downstream: if what arrived matches what we offered, the
// path was never the constraint and this number says nothing about it.
measuresNetwork = seen != null && achievedKbps > 0 && achievedKbps < sent.kbps * 9 / 10,
),
) as JsonObject
val findings = ArrayList<Finding>()
if (seen != null && seen.packets == 0 && sent.packets > 0) {
findings.add(
finding(
FindingRegistry.THROUGHPUT_NO_DELIVERY, testId,
"No upstream traffic reached the server",
"This device sent ${sent.packets} packets and the server received none. " +
"That is a connectivity fault on the outbound path rather than a slow link.",
),
)
} else if (lossPct != null && lossPct >= 2.0) {
findings.add(
finding(
FindingRegistry.THROUGHPUT_BELOW_OFFERED, testId,
"Upstream loss of $lossPct % at ${sent.kbps / 1000} Mbit/s",
"The server received ${seen?.packets} of the ${sent.packets} packets this " +
"device sent. The outbound path could not carry what was offered.",
),
)
}
return Test(
id = testId, type = TestType.PERF_THROUGHPUT_UDP, sessionRef = sessionRef, tier = Tier.APP,
startedMonoNs = started, endedMonoNs = ids.monoNs(),
status = if (seen == null || seen.packets == 0) TestStatus.FAILED else TestStatus.OK,
metrics = metrics,
) to findings
}
private data class UpstreamCount(val packets: Int, val bytes: Long, val kbps: Int)
/** The server's tally for this session's upstream run. */
private fun upstreamCount(
control: ControlClient, credential: String, sessionId: String,
): UpstreamCount? = runCatching {
val o = Json.parseToJsonElement(control.observations(credential, sessionId))
.jsonObject["throughput_up"]?.jsonObject ?: return null
UpstreamCount(
packets = o["packets"]?.jsonPrimitive?.content?.toIntOrNull() ?: 0,
bytes = o["bytes"]?.jsonPrimitive?.content?.toLongOrNull() ?: 0,
kbps = o["kbps"]?.jsonPrimitive?.content?.toIntOrNull() ?: 0,
)
}.getOrNull()
private data class SenderReport(
val packets: Int, val bytes: Long, val kbps: Int, val limitedBy: String,
)
/** The server's own account of the run, from the observations API. */
private fun senderReport(
control: ControlClient, credential: String, sessionId: String,
): SenderReport? = runCatching {
val arr = Json.parseToJsonElement(control.observations(credential, sessionId))
.jsonObject["throughput"]?.jsonArray ?: return null
val last = arr.lastOrNull()?.jsonObject ?: return null
SenderReport(
packets = last["packets"]?.jsonPrimitive?.content?.toIntOrNull() ?: 0,
bytes = last["bytes"]?.jsonPrimitive?.content?.toLongOrNull() ?: 0,
kbps = last["kbps"]?.jsonPrimitive?.content?.toIntOrNull() ?: 0,
limitedBy = last["limited_by"]?.jsonPrimitive?.content ?: "unknown",
)
}.getOrNull()
private fun parseInt(body: String?, key: String): Int? =
body?.let { Regex("\"$key\"\\s*:\\s*(-?\\d+)").find(it)?.groupValues?.get(1)?.toIntOrNull() }
/**
* Builds a finding from a registry entry, which supplies the code, category and severity.
*
* Taking a [FindingSpec] rather than three loose values is the point: a typo becomes a
* compile error, and two call sites cannot disagree about which category a finding belongs
* to - a disagreement that would split one fault across two verdict lights.
*/
private fun finding(spec: FindingSpec, testId: String, title: String, desc: String) =
Finding(
id = ids.uuid(), code = spec.code, category = spec.category, severity = spec.severity,
confidence = Confidence.HIGH,
title = title, description = desc, evidenceRefs = listOf(EvidenceRef(testId)),
)
private fun round2(v: Double) = Math.round(v * 100.0) / 100.0
}
/** Metrics for perf.throughput_udp in the upstream direction. */
@Serializable
data class UpstreamThroughputMetrics(
val direction: String = "up",
@SerialName("requested_kbps") val requestedKbps: Int,
@SerialName("sent_packets") val sentPackets: Int,
@SerialName("sent_bytes") val sentBytes: Long,
/** What this device managed to emit — a property of the phone and its radio, not the path. */
@SerialName("sent_kbps") val sentKbps: Int,
@SerialName("received_packets") val receivedPackets: Int? = null,
@SerialName("received_bytes") val receivedBytes: Long? = null,
/** What arrived, measured by the only party that can measure it. This is the result. */
@SerialName("received_kbps") val receivedKbps: Int,
@SerialName("loss_pct") val lossPct: Double? = null,
@SerialName("measures_network") val measuresNetwork: Boolean,
)
/** Metrics for perf.throughput_udp. */
@Serializable
data class ThroughputMetrics(
val direction: String = "down",
@SerialName("requested_kbps") val requestedKbps: Int,
@SerialName("planned_duration_ms") val plannedDurationMs: Int,
@SerialName("packets_received") val packetsReceived: Int,
@SerialName("bytes_received") val bytesReceived: Long,
@SerialName("received_kbps") val receivedKbps: Int,
@SerialName("sender_packets") val senderPackets: Int? = null,
@SerialName("sender_bytes") val senderBytes: Long? = null,
@SerialName("sender_kbps") val senderKbps: Int? = null,
/** Against the sender's count, so a server-side limit is never counted as network loss. */
@SerialName("loss_pct") val lossPct: Double? = null,
/** What ended the run: duration | budget | rate | send_error | unknown. */
@SerialName("limited_by") val limitedBy: String,
/**
* Whether this number says anything about the network. False when the sender's own ceiling
* was the binding constraint — in which case the rate is a property of the test, not the path.
*/
@SerialName("measures_network") val measuresNetwork: Boolean,
)
@@ -0,0 +1,159 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.engine
import app.echo_lot.measurement.*
import app.echo_lot.protocol.ProbeSession
import kotlinx.serialization.SerialName
import kotlinx.serialization.Serializable
import kotlinx.serialization.json.Json
import kotlinx.serialization.json.JsonObject
import kotlinx.serialization.json.encodeToJsonElement
/**
* train.udp_updown — the client sends a paced train (types 0x03), then asks the server what
* arrived (0x04 → 0x05) and lines both views up per sequence number.
*
* This is the measurement a round trip cannot make: an echo run only says "lost somewhere", the
* train's two ledgers say lost on the way OUT, specifically, because the server's report names
* exactly which sequence numbers reached it. The downstream direction has its own test
* (train.udp_downstream) under a grant; this one needs none, since the client generates all the
* traffic itself.
*
* The evidence is the schema's columnar TrainEvidence: one index per sent packet, with the
* server-side columns null where a packet never arrived. Server timestamps are on the server's
* own clock — only differences within that clock mean anything unless time.server_offset maps
* them (two-clock rule).
*/
class UpstreamTrainMeasurement(private val ids: IdSource) {
private val json = Json { encodeDefaults = true; explicitNulls = true }
fun run(
probe: ProbeSession,
sessionRef: String,
count: Int = 200,
sizeBytes: Int = 200,
interPacketMs: Long = 5,
): Pair<Test, List<Finding>> {
val testId = ids.uuid()
val started = ids.monoNs()
// The id only needs to be unique within this session; a clash across sessions is
// meaningless because trains are buffered per session on the server.
val trainId = (System.nanoTime() and 0x7FFFFFFF).toInt()
val sent = probe.sendTrain(trainId, count, sizeBytes, interPacketMs)
// Let the tail arrive before asking for the ledger; packets still in flight when the
// report is cut would read as upstream loss.
Thread.sleep(300)
val report = probe.trainReport(trainId)
if (report == null) {
return Test(
id = testId, type = TestType.TRAIN_UDP_UPDOWN, sessionRef = sessionRef,
tier = Tier.APP, startedMonoNs = started, endedMonoNs = ids.monoNs(),
status = TestStatus.FAILED,
// Honest ambiguity: an old server drops 0x04 silently, and a lost report looks
// identical from here. Neither says anything about the train itself.
error = TestError(
"no_report",
"no train report arrived — the report was lost, or the server predates trains",
),
) to emptyList()
}
val bySeq = report.rows.associateBy { it.seq }
fun col255(v: Int): Int? = v.takeIf { it != 255 } // 255 = "not observed" on the wire
val evidence = TrainEvidence(
epochMonoNs = started,
seq = sent.map { it.seq },
tTxNs = sent.map { it.tTxNs },
tSrvRxNs = sent.map { bySeq[it.seq]?.tRxNs },
tRxNs = sent.map { null }, // upstream only: nothing comes back per packet
sizeBytes = sent.map { it.sizeBytes },
ttlSeenByServer = sent.map { bySeq[it.seq]?.let { r -> col255(r.ttl) } },
dscpSeenByServer = sent.map { bySeq[it.seq]?.let { r -> col255(r.dscp) } },
ecnSeenByServer = sent.map { bySeq[it.seq]?.let { r -> col255(r.ecn) } },
evidenceTruncated = report.truncated,
).toEvidence()
// Loss against the server's total count, not its row list: rows past the server's buffer
// cap are counted but not kept, and treating them as lost would invent loss exactly on
// the biggest trains.
val lossPct = if (sent.isEmpty()) 0.0 else {
(sent.size - report.received).coerceAtLeast(0) * 100.0 / sent.size
}
val metrics = json.encodeToJsonElement(
UpstreamTrainMetrics(
sent = sent.size,
receivedByServer = report.received,
lossPct = round1(lossPct),
reportPartsExpected = report.partsExpected,
reportPartsReceived = report.partsReceived,
truncated = report.truncated,
),
) as JsonObject
val findings = ArrayList<Finding>()
if (sent.isNotEmpty() && report.received == 0) {
findings.add(
Finding(
id = ids.uuid(),
code = FindingRegistry.UDP_UNREACHABLE_UPSTREAM.code,
category = FindingRegistry.UDP_UNREACHABLE_UPSTREAM.category,
severity = FindingRegistry.UDP_UNREACHABLE_UPSTREAM.severity,
confidence = Confidence.HIGH,
title = "The server received none of ${sent.size} upstream packets",
description = "Every train packet vanished on the way out, while the " +
"report request's reply made it back — the outbound path drops this " +
"traffic, the return path works.",
evidenceRefs = listOf(EvidenceRef(testId)),
),
)
} else if (lossPct >= 2.0) {
findings.add(
Finding(
id = ids.uuid(),
code = FindingRegistry.LOSS_UPSTREAM.code,
category = FindingRegistry.LOSS_UPSTREAM.category,
severity = FindingRegistry.LOSS_UPSTREAM.severity,
confidence = Confidence.HIGH,
title = "Upstream loss of ${round1(lossPct)} %",
description = "The server received ${report.received} of the ${sent.size} " +
"packets this device sent, and its per-sequence ledger names the " +
"missing ones. This is outbound loss specifically; the return path " +
"delivered the report.",
evidenceRefs = listOf(EvidenceRef(testId)),
),
)
}
val status = when {
report.received == 0 && sent.isNotEmpty() -> TestStatus.FAILED
report.partsReceived < report.partsExpected -> TestStatus.PARTIAL
else -> TestStatus.OK
}
return Test(
id = testId, type = TestType.TRAIN_UDP_UPDOWN, sessionRef = sessionRef, tier = Tier.APP,
startedMonoNs = started, endedMonoNs = ids.monoNs(),
status = status, evidence = evidence, metrics = metrics,
) to findings
}
private fun round1(v: Double) = Math.round(v * 10.0) / 10.0
}
/** Metrics for train.udp_updown. */
@Serializable
data class UpstreamTrainMetrics(
val sent: Int,
/** The server's total count — includes packets past its row buffer (counted, not listed). */
@SerialName("received_by_server") val receivedByServer: Int,
@SerialName("loss_pct") val lossPct: Double,
@SerialName("report_parts_expected") val reportPartsExpected: Int,
@SerialName("report_parts_received") val reportPartsReceived: Int,
/** The server's row buffer overflowed: rows are a sample, the count is still complete. */
val truncated: Boolean,
)
@@ -0,0 +1,155 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.engine
import app.echo_lot.engine.Directional.Sample
import app.echo_lot.engine.Directional.ServerSighting
import kotlin.test.Test
import kotlin.test.assertEquals
import kotlin.test.assertFalse
import kotlin.test.assertNotNull
import kotlin.test.assertNull
import kotlin.test.assertTrue
/**
* The arithmetic that turns "3 % loss somewhere" into "3 % loss upstream". Getting a denominator
* wrong here does not crash anything — it produces a plausible number pointing at the wrong half
* of the network, which is worse than no number at all. Hence a test per claim.
*/
class DirectionalTest {
/** A clean train: every packet sent, seen and answered. Server clock offset by a constant. */
private fun clean(n: Int, offsetNs: Long = 5_000_000_000L): Pair<List<Sample>, List<ServerSighting>> {
val sent = (1..n).map { Sample(it, tTxNs = it * 10_000_000L, tRxNs = it * 10_000_000L + 4_000_000L) }
val seen = (1..n).map {
ServerSighting(it, tRxNs = offsetNs + it * 10_000_000L + 2_000_000L,
tTxNs = offsetNs + it * 10_000_000L + 2_100_000L)
}
return sent to seen
}
@Test
fun aCleanTrainReportsNoLossInEitherDirection() {
val (sent, seen) = clean(10)
val m = Directional.analyse(sent, seen)
assertEquals(10, m.sent)
assertEquals(10, m.seenByServer)
assertEquals(10, m.repliesReceived)
assertEquals(0.0, m.lossUpstreamPct)
assertEquals(0.0, m.lossDownstreamPct)
assertFalse(m.noneReachedServer)
}
// The whole point: a packet the server never saw was lost on the way there.
@Test
fun packetsTheServerNeverSawAreUpstreamLoss() {
val (sent, seen) = clean(10)
val m = Directional.analyse(sent, seen.filter { it.seq !in setOf(3, 7) })
assertEquals(2, m.lostUpstream)
assertEquals(0, m.lostDownstream)
assertEquals(20.0, m.lossUpstreamPct)
assertEquals(0.0, m.lossDownstreamPct, "a packet that never arrived cannot be lost coming back")
}
@Test
fun repliesThatNeverArrivedAreDownstreamLoss() {
val (sent, seen) = clean(10)
val withHoles = sent.map { if (it.seq in setOf(2, 5)) it.copy(tRxNs = null) else it }
val m = Directional.analyse(withHoles, seen)
assertEquals(0, m.lostUpstream)
assertEquals(2, m.lostDownstream)
assertEquals(20.0, m.lossDownstreamPct)
}
// Downstream loss is measured against what actually reached the server. Using "sent" as the
// denominator would count every upstream loss a second time and overstate the return path.
@Test
fun downstreamLossIsRelativeToWhatReachedTheServer() {
val (sent, seen) = clean(10)
// 5 lost on the way there; of the 5 that arrived, 1 reply is lost coming back.
val seenPartial = seen.filter { it.seq > 5 }
val withHole = sent.map {
when {
it.seq <= 5 -> it.copy(tRxNs = null) // never got there, so never came back
it.seq == 6 -> it.copy(tRxNs = null) // arrived, reply lost
else -> it
}
}
val m = Directional.analyse(withHole, seenPartial)
assertEquals(5, m.lostUpstream)
assertEquals(50.0, m.lossUpstreamPct)
assertEquals(1, m.lostDownstream)
assertEquals(20.0, m.lossDownstreamPct, "1 of the 5 that arrived, not 1 of 10")
}
@Test
fun aServerThatSawNothingIsCalledOutSeparately() {
val (sent, _) = clean(6)
val m = Directional.analyse(sent.map { it.copy(tRxNs = null) }, emptyList())
assertTrue(m.noneReachedServer)
assertEquals(100.0, m.lossUpstreamPct)
assertEquals(0.0, m.lossDownstreamPct, "with nothing arriving there is no return path to blame")
}
// Jitter is legitimate without synchronised clocks because the offset cancels when successive
// one-way samples are differenced. This pins that: a huge constant offset must not show up.
@Test
fun jitterIsUnaffectedByTheClockOffsetBetweenTheTwoMachines() {
val (sent, near) = clean(10, offsetNs = 0)
val (_, far) = clean(10, offsetNs = 9_999_999_999L)
val a = Directional.analyse(sent, near)
val b = Directional.analyse(sent, far)
assertEquals(a.jitterUpstreamMs, b.jitterUpstreamMs,
"a constant clock offset must cancel when consecutive samples are differenced")
assertEquals(0.0, assertNotNull(a.jitterUpstreamMs), "an evenly spaced train has no jitter")
}
@Test
fun jitterReflectsUnevenArrival() {
val sent = listOf(
Sample(1, 0, 10_000_000),
Sample(2, 10_000_000, 20_000_000),
Sample(3, 20_000_000, 30_000_000),
)
// Server receive times drift: +2ms, +7ms, +3ms relative to send.
val seen = listOf(
ServerSighting(1, 2_000_000, 2_100_000),
ServerSighting(2, 17_000_000, 17_100_000),
ServerSighting(3, 23_000_000, 23_100_000),
)
val m = Directional.analyse(sent, seen)
// one-way samples: 2ms, 7ms, 3ms → |7-2| and |3-7| → mean 4.5ms
assertEquals(4.5, assertNotNull(m.jitterUpstreamMs))
}
// "No jitter" and "not enough data to say" are different claims, and only one is true here.
@Test
fun tooFewSamplesReportsNoJitterRatherThanZero() {
val m = Directional.analyse(
listOf(Sample(1, 0, 10_000_000)),
listOf(ServerSighting(1, 2_000_000, 2_100_000)),
)
assertNull(m.jitterUpstreamMs)
assertNull(m.jitterDownstreamMs)
}
// A server record for a sequence we never sent is not evidence about this train; folding it
// in would yield loss percentages outside 0100.
@Test
fun strayServerRecordsAreIgnored() {
val (sent, seen) = clean(5)
val m = Directional.analyse(sent, seen + ServerSighting(99, 1, 2) + ServerSighting(100, 3, 4))
assertEquals(5, m.seenByServer)
assertEquals(0.0, m.lossUpstreamPct)
assertTrue(m.lossDownstreamPct in 0.0..100.0)
}
@Test
fun anEmptyTrainDoesNotDivideByZero() {
val m = Directional.analyse(emptyList(), emptyList())
assertEquals(0.0, m.lossUpstreamPct)
assertEquals(0.0, m.lossDownstreamPct)
assertFalse(m.noneReachedServer, "nothing sent is not the same as nothing arriving")
}
}
@@ -44,7 +44,9 @@ class LiveDownstreamTest {
for (t in tests) println("${t.type} status=${t.status} metrics=${t.metrics}") for (t in tests) println("${t.type} status=${t.status} metrics=${t.metrics}")
for (f in findings) println("finding ${f.code} [${f.severity}] ${f.title}") for (f in findings) println("finding ${f.code} [${f.severity}] ${f.title}")
assertEquals(3, tests.size, "expected pmtud_down, frag_delivery and a downstream train") // Assert on what is present, not on how many: adding a measurement should not be a
// test edit. (It was, once — hence the note.)
assertTrue(tests.size >= 3, "expected at least the three downstream tests, got ${tests.size}")
val byType = tests.associateBy { it.type } val byType = tests.associateBy { it.type }
val pmtud = assertNotNull(byType[TestType.MTU_PMTUD_DOWN], "no mtu.pmtud_down test") val pmtud = assertNotNull(byType[TestType.MTU_PMTUD_DOWN], "no mtu.pmtud_down test")
@@ -58,6 +60,17 @@ class LiveDownstreamTest {
val frag = assertNotNull(byType[TestType.MTU_FRAG_DELIVERY], "no mtu.frag_delivery test") val frag = assertNotNull(byType[TestType.MTU_FRAG_DELIVERY], "no mtu.frag_delivery test")
assertNotNull(frag.metrics?.get("largest_delivered_bytes")) assertNotNull(frag.metrics?.get("largest_delivered_bytes"))
// Fragment ordering runs only when fragments arrive at all, and only against a server
// that can craft them — so it is checked when present rather than required.
byType[TestType.MTU_FRAG_ORDERING]?.let { fo ->
val m = fo.metrics?.toString() ?: ""
println("fragment ordering: ${fo.status} $m")
if (fo.status != TestStatus.UNSUPPORTED) {
assertTrue(m.contains("in_order"), "no per-ordering result: $m")
assertTrue(m.contains("reversed"), "reversed ordering was never attempted: $m")
}
}
val train = assertNotNull(byType[TestType.TRAIN_UDP_DOWNSTREAM], "no downstream train") val train = assertNotNull(byType[TestType.TRAIN_UDP_DOWNSTREAM], "no downstream train")
assertNotNull(train.evidence, "a train without columnar evidence is not recomputable") assertNotNull(train.evidence, "a train without columnar evidence is not recomputable")
val received = train.metrics?.get("received")?.toString()?.toIntOrNull() ?: 0 val received = train.metrics?.get("received")?.toString()?.toIntOrNull() ?: 0
@@ -7,6 +7,7 @@ import app.echo_lot.measurement.*
import kotlinx.serialization.json.Json import kotlinx.serialization.json.Json
import kotlin.test.Test import kotlin.test.Test
import kotlin.test.assertEquals import kotlin.test.assertEquals
import kotlin.test.assertNotNull
import kotlin.test.assertTrue import kotlin.test.assertTrue
/** /**
@@ -59,6 +60,18 @@ class LiveMeasurementTest {
println("metrics: $metrics") println("metrics: $metrics")
assertTrue(metrics.toString().contains("rtt_ms_avg")) assertTrue(metrics.toString().contains("rtt_ms_avg"))
// The directional split is the point of asking the server what it saw: without it a
// lossy path is reported as "loss" with no direction, which sends an engineer looking
// in both at once. Correlation is by wire sequence number, so a mismatch here means the
// two sides disagree about which packet is which.
val m = metrics.toString()
assertTrue(m.contains("seen_by_server"), "no directional split in the metrics: $m")
val seen = Regex(""""seen_by_server":(\d+)""").find(m)?.groupValues?.get(1)?.toInt()
assertNotNull(seen, "seen_by_server missing")
assertEquals(20, seen, "the server should have seen every probe on a healthy path")
assertTrue(m.contains("jitter_upstream_ms"), "no per-direction jitter: $m")
println("directional: $m")
assertTrue(doc.summary != null) assertTrue(doc.summary != null)
// A healthy local->fmr path should be green (no loss, no rebinding) or yellow. // A healthy local->fmr path should be green (no loss, no rebinding) or yellow.
println("summary: ${doc.summary}") println("summary: ${doc.summary}")
@@ -0,0 +1,105 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.engine
import app.echo_lot.measurement.TestStatus
import app.echo_lot.protocol.ControlClient
import app.echo_lot.protocol.ProbeSession
import kotlin.test.Test
import kotlin.test.assertEquals
import kotlin.test.assertNotNull
import kotlin.test.assertTrue
/**
* Downstream throughput against a LIVE server. Self-skips without ECHOLOT_LIVE_*.
*
* The assertions are about *honesty* rather than speed: a rate is only a measurement if the run
* was ended by the clock and the sender's own count backs it up. A test that just asserted "some
* Mbps arrived" would pass equally well against a broken implementation.
*/
class LiveThroughputTest {
private val url = System.getenv("ECHOLOT_LIVE_URL")
private val pin = System.getenv("ECHOLOT_LIVE_PIN")
private val cred = System.getenv("ECHOLOT_LIVE_CRED")
private val udp = System.getenv("ECHOLOT_LIVE_UDP")
private val target = System.getenv("ECHOLOT_LIVE_TARGET") ?: "fmr"
@Test
fun measuresDownstreamRateAndSaysWhatLimitedIt() {
if (url == null || pin == null || cred == null || udp == null) {
println("LiveThroughputTest skipped (no ECHOLOT_LIVE_* env)"); return
}
val control = ControlClient(url, setOf(pin), "0.2.0")
val session = control.createSession(cred, target)
val (host, port) = udp.split(":").let { it[0] to it[1].toInt() }
val (test, findings) = ProbeSession(cred, session, host, port).use { ps ->
ps.echo() // prime: the grant binds to the observed source
ThroughputMeasurement(SystemIdSource()).run(
cred, session.sessionId, control, ps, sessionRef = "sess-1",
durationS = 3, kbps = 20_000,
)
}
control.deleteSession(cred, session.sessionId)
val m = assertNotNull(test.metrics).toString()
println("throughput: ${test.status} $m")
for (f in findings) println("finding ${f.code} [${f.severity}] ${f.title}")
assertEquals(TestStatus.OK, test.status, "no throughput traffic arrived: $m")
// The sender's own count must be present — without it, loss cannot be attributed and the
// number is not a measurement.
assertTrue(m.contains("sender_packets"), "no sender report to compare against: $m")
assertTrue(m.contains("limited_by"), "the result must say what ended the run: $m")
val received = Regex(""""received_kbps":(\d+)""").find(m)?.groupValues?.get(1)?.toInt()
assertNotNull(received)
assertTrue(received > 0, "measured 0 kbps: $m")
println("received ${received / 1000} Mbit/s")
// A run this short and this far below the ceiling should end on the clock. Anything else
// means the grant was the constraint, and then the rate says nothing about the path.
assertTrue(m.contains(""""limited_by":"duration""""),
"the run did not end on the clock, so the rate measures the server, not the path: $m")
}
// Upstream is the direction only the far end can measure. The assertion that matters is that
// the server's count is present and plausible against what we sent — a test that only checked
// "we transmitted some Mbps" would pass against a server that counted nothing at all.
@Test
fun measuresUpstreamAgainstTheServersCount() {
if (url == null || pin == null || cred == null || udp == null) {
println("LiveThroughputTest(up) skipped"); return
}
val control = ControlClient(url, setOf(pin), "0.2.0")
val session = control.createSession(cred, target)
val (host, port) = udp.split(":").let { it[0] to it[1].toInt() }
val (test, findings) = ProbeSession(cred, session, host, port).use { ps ->
ps.echo()
ThroughputMeasurement(SystemIdSource()).runUpstream(
cred, session.sessionId, control, ps, sessionRef = "sess-1",
durationS = 3, kbps = 10_000,
)
}
control.deleteSession(cred, session.sessionId)
val m = assertNotNull(test.metrics).toString()
println("upstream: ${test.status} $m")
for (f in findings) println("finding ${f.code} [${f.severity}] ${f.title}")
assertEquals(TestStatus.OK, test.status, "the server counted nothing: $m")
val recv = Regex(""""received_packets":(\d+)""").find(m)?.groupValues?.get(1)?.toInt()
val sent = Regex(""""sent_packets":(\d+)""").find(m)?.groupValues?.get(1)?.toInt()
assertNotNull(recv); assertNotNull(sent)
assertTrue(sent > 100, "barely anything was sent, so the rate means nothing: $m")
assertTrue(recv > 0, "the server received none of $sent packets: $m")
// The counts should be close on a healthy path; wildly different means the two sides are
// counting different things rather than the network losing packets.
assertTrue(recv <= sent, "the server counted MORE than we sent — the counter is not being reset")
println("sent $sent, server saw $recv")
}
}
@@ -0,0 +1,65 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.engine
import app.echo_lot.measurement.TestStatus
import app.echo_lot.protocol.ControlClient
import app.echo_lot.protocol.ProbeSession
import kotlin.test.Test
import kotlin.test.assertEquals
import kotlin.test.assertNotNull
import kotlin.test.assertTrue
/**
* Upstream train (types 0x03-0x05) against a LIVE server. Self-skips without ECHOLOT_LIVE_*.
*
* What is asserted is the ledger property: the server's report must account for what was sent,
* per sequence number, because directional loss attribution is the entire reason trains exist
* a test that only checked "a report came back" would pass against a server that counts nothing.
*/
class LiveUpstreamTrainTest {
private val url = System.getenv("ECHOLOT_LIVE_URL")
private val pin = System.getenv("ECHOLOT_LIVE_PIN")
private val cred = System.getenv("ECHOLOT_LIVE_CRED")
private val udp = System.getenv("ECHOLOT_LIVE_UDP")
private val target = System.getenv("ECHOLOT_LIVE_TARGET") ?: "fmr"
@Test
fun serverLedgerAccountsForTheTrain() {
if (url == null || pin == null || cred == null || udp == null) {
println("LiveUpstreamTrainTest skipped (no ECHOLOT_LIVE_* env)"); return
}
val control = ControlClient(url, setOf(pin), "0.2.0")
val session = control.createSession(cred, target)
val (host, port) = udp.split(":").let { it[0] to it[1].toInt() }
val (test, findings) = ProbeSession(cred, session, host, port).use { ps ->
ps.echo() // prime the session so its source is known
UpstreamTrainMeasurement(SystemIdSource()).run(
ps, sessionRef = "sess-1", count = 120, sizeBytes = 200, interPacketMs = 3,
)
}
control.deleteSession(cred, session.sessionId)
val m = assertNotNull(test.metrics).toString()
println("updown: ${test.status} $m")
for (f in findings) println("finding ${f.code} [${f.severity}] ${f.title}")
assertEquals(TestStatus.OK, test.status, "train report incomplete or absent: $m")
val sent = Regex(""""sent":(\d+)""").find(m)?.groupValues?.get(1)?.toInt()
val received = Regex(""""received_by_server":(\d+)""").find(m)?.groupValues?.get(1)?.toInt()
assertNotNull(sent); assertNotNull(received)
assertTrue(sent > 0, "nothing was sent: $m")
// Over a working path the ledger must be near-complete; a lossy wifi may drop a few, but
// a server that fails to count would show up as massive phantom loss here.
assertTrue(received >= sent * 9 / 10, "server counted $received of $sent: $m")
// The columnar evidence must carry a server timestamp for arrived packets — that column
// is what one-way delay math consumes after timesync.
val ev = assertNotNull(test.evidence).toString()
assertTrue(ev.contains("t_srv_rx_ns"), "no server rx column in evidence")
}
}
@@ -28,6 +28,7 @@ data class MeasurementDocument(
data class Run( data class Run(
val id: String, // UUIDv7 val id: String, // UUIDv7
val trigger: Trigger, val trigger: Trigger,
val mode: RunMode = RunMode.SHORT,
@SerialName("started_at") val startedAt: String, // RFC3339 UTC, human correlation only @SerialName("started_at") val startedAt: String, // RFC3339 UTC, human correlation only
@SerialName("ended_at") val endedAt: String? = null, @SerialName("ended_at") val endedAt: String? = null,
val clock: Clock, val clock: Clock,
@@ -35,9 +36,58 @@ data class Run(
val device: DeviceInfo, val device: DeviceInfo,
val tiers: Tiers, val tiers: Tiers,
@SerialName("profiles_used") val profilesUsed: List<String> = emptyList(), @SerialName("profiles_used") val profilesUsed: List<String> = emptyList(),
val constraints: Constraints = Constraints(),
val notes: String? = null, val notes: String? = null,
) )
/**
* What limited this run the counterpart to [Tiers], which records what was available.
*
* A constrained run is not a failed run, and it is not a normal one either. Without this, a run
* taken through a VPN looks exactly like a clean run of a healthy network: the same shape, the
* same green verdict, and no way for a reader or a server aggregating thousands of these to
* know that almost nothing was actually measured.
*/
@Serializable
data class Constraints(
/** A VPN held the default route while this ran. */
@SerialName("vpn_active") val vpnActive: Boolean = false,
/**
* Per-network probing was refused by the OS.
*
* Android blocks `Network.bindSocket()` on the underlying networks whenever a VPN is up, to
* stop apps leaking around the tunnel. Every per-network test then measures nothing, so any
* conclusion drawn about the wifi or cellular link underneath is unfounded.
*/
@SerialName("per_network_blocked") val perNetworkBlocked: Boolean = false,
/** Networks that could not be measured, by id. */
@SerialName("unmeasured_networks") val unmeasuredNetworks: List<String> = emptyList(),
) {
/** True when this run's results mean something different from an unconstrained one. */
val constrained: Boolean get() = vpnActive || perNetworkBlocked
}
/**
* How long the run watched the network and therefore what its silence is worth.
*
* A [SHORT] run is a sequence of one-shot probes: each looks at the network for a second or two and
* moves on. That is enough to characterise a network's *configuration*, and it is structurally
* incapable of seeing anything intermittent. A wifi link that drops for four seconds every two
* minutes, a resolver that stalls under load, an AP that roams none of these leave a trace in
* thirty seconds of probing unless the run happened to coincide with one.
*
* A [LONG] run starts continuous listeners at t=0, runs the same battery beside them, and keeps
* sampling until the window closes. It answers a different question, so a reader must not treat the
* two alike: **the mode is what licenses an argument from absence**. "No drops were observed" means
* something after five minutes of watching and nothing at all after a thirty-second run, and
* without this field the two documents are indistinguishable.
*/
@Serializable
enum class RunMode {
@SerialName("short") SHORT,
@SerialName("long") LONG,
}
@Serializable @Serializable
enum class Trigger { enum class Trigger {
@SerialName("manual") MANUAL, @SerialName("manual") MANUAL,
@@ -0,0 +1,341 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.measurement
/**
* The registry of finding codes (measurement-schema.md §9, open item 1).
*
* A finding code is the stable, machine-readable half of a result: the prose changes, the code is
* what a dashboard groups by and what someone greps a year of archived runs for. That only holds
* if a code means exactly one thing forever which is not something ad-hoc string literals at
* fifteen call sites can promise.
*
* The failure this exists to prevent had already happened by the time it was written. Two
* independently-added emitters produced `connectivity.downstream_loss` and
* `connectivity.loss_downstream` for the same concept, and nothing anywhere objected. Anyone
* aggregating either one would have silently seen half their data.
*
* So codes are declared here as typed specs, each carrying its category and default severity, and
* emitters reference the spec rather than retyping the string. That makes a typo a compile error,
* and makes it impossible for two call sites to disagree about which category a finding belongs
* to a disagreement that would otherwise split one fault across two verdict lights.
*/
data class FindingSpec(
val code: String,
val category: Category,
/** Severity when nothing about the specific run argues otherwise; emitters may escalate. */
val severity: Severity,
/** One line: what this finding asserts. Present tense, no hedging. */
val meaning: String,
/**
* What the finding rules *out*, where that is the useful half. "Loss upstream" is worth much
* more when it also says the return path is fine, because that halves where to look next.
*/
val rulesOut: String? = null,
)
object FindingRegistry {
// ---- connectivity ----------------------------------------------------------------
// Renamed from nat.* before anything shipped: neither of these is about NAT, and the
// prefix is what decides which category - and therefore which verdict light - a finding
// rolls up into. A nat.* code landing under connectivity would be a permanent puzzle.
val UDP_UNREACHABLE = FindingSpec(
"connectivity.udp_unreachable", Category.CONNECTIVITY, Severity.HIGH,
"No UDP echo replies came back from the server at all.",
)
val UDP_UNREACHABLE_UPSTREAM = FindingSpec(
"connectivity.udp_unreachable_upstream", Category.CONNECTIVITY, Severity.HIGH,
"The server received none of the probes, so traffic is dropped on the way out.",
rulesOut = "The return path: nothing arrived to be replied to.",
)
val UDP_LOSS = FindingSpec(
"connectivity.udp_loss", Category.CONNECTIVITY, Severity.MEDIUM,
"A large fraction of round-trip probes were lost, direction unknown.",
)
val LOSS_UPSTREAM = FindingSpec(
"connectivity.loss_upstream", Category.CONNECTIVITY, Severity.MEDIUM,
"Probes were lost on the way to the server.",
rulesOut = "The return path: replies came back for everything that arrived.",
)
/**
* The single code for "lost on the return path", whichever measurement found it.
*
* Two emitters had independently invented `connectivity.downstream_loss` and
* `connectivity.loss_downstream` for this, and nothing objected. Anyone aggregating either
* one would have silently seen half their data. Paired with [LOSS_UPSTREAM] so the two
* directions read as a set.
*/
val LOSS_DOWNSTREAM = FindingSpec(
"connectivity.loss_downstream", Category.CONNECTIVITY, Severity.MEDIUM,
"Packets were lost on the way back from the server.",
rulesOut = "The outbound path: the server received what it was answering.",
)
val DOWNSTREAM_BLOCKED = FindingSpec(
"connectivity.downstream_blocked", Category.CONNECTIVITY, Severity.HIGH,
"Server-initiated packets never arrive, although round trips work.",
rulesOut = "Basic reachability: the path forwards replies, just not unsolicited traffic.",
)
val DOWNSTREAM_REORDER = FindingSpec(
"connectivity.downstream_reorder", Category.CONNECTIVITY, Severity.LOW,
"Downstream packets arrive in a different order than they were sent.",
)
// MEDIUM, not HIGH: a captive portal is a condition to report, not necessarily a fault - on
// hotel or cafe wifi it is exactly what should be there, and logging in clears it. NO_INTERNET
// is the HIGH one, because nothing the user does locally fixes that. The registry first said
// HIGH; the probe emitting it had always said MEDIUM, and the probe was the considered value.
val CAPTIVE_PORTAL = FindingSpec(
"connectivity.captive_portal", Category.CONNECTIVITY, Severity.MEDIUM,
"A captive portal is intercepting connectivity checks.",
)
val NO_INTERNET = FindingSpec(
"connectivity.no_internet", Category.CONNECTIVITY, Severity.HIGH,
"Android's own connectivity checks fail on this network.",
)
/**
* The finding a short run cannot make.
*
* Every one-shot probe describes the network during its own two seconds. A link that drops and
* returns between two of them leaves no trace anywhere in the document the probes before and
* after both succeed, and the run reports a healthy network. Only a listener that watches the
* whole window sees the gap, which is why this is emitted from `networks[].changes[]` (§4)
* rather than from any test's evidence.
*
* MEDIUM by default and escalated by the emitter on repeat: one drop in five minutes is worth
* knowing about, three is the difference between "the wifi hiccuped" and "this link is why
* calls keep dropping". Deliberately claims a *completed* cycle — lost and then regained — so
* it never fires for a network that was simply turned off partway through the run.
*/
val LINK_FLAPPING = FindingSpec(
"connectivity.link_flapping", Category.CONNECTIVITY, Severity.MEDIUM,
"A network dropped and came back one or more times during the run.",
rulesOut = "A momentary probe failure: the drop was watched happening, not inferred from silence.",
)
// ---- mtu -------------------------------------------------------------------------
val MTU_REDUCED_DOWNSTREAM = FindingSpec(
"mtu.reduced_downstream", Category.MTU, Severity.LOW,
"The downstream path MTU is below the usual 1500 bytes.",
)
val MTU_DOWNSTREAM_BLACKHOLE = FindingSpec(
"mtu.downstream_blackhole", Category.MTU, Severity.MEDIUM,
"Datagrams above the path MTU are dropped downstream, fragmented or not.",
)
val FRAGMENTS_BLOCKED = FindingSpec(
"mtu.fragments_blocked", Category.MTU, Severity.MEDIUM,
"IP fragments do not reach this device even when sent in order.",
)
val FRAGMENT_REORDER_SENSITIVE = FindingSpec(
"mtu.fragment_reorder_sensitive", Category.MTU, Severity.LOW,
"Fragments are delivered in order but dropped when reordered or delayed.",
rulesOut = "Fragmentation itself: in-order fragments arrive fine.",
)
// ---- nat -------------------------------------------------------------------------
val NAT_UDP_REBINDING = FindingSpec(
"nat.udp_rebinding", Category.NAT, Severity.MEDIUM,
"A NAT remapped the UDP source port mid-flow.",
)
val NAT_SYMMETRIC = FindingSpec(
"nat.symmetric", Category.NAT, Severity.MEDIUM,
"The NAT assigns a different external port per destination.",
)
// ---- perf ------------------------------------------------------------------------
val THROUGHPUT_NO_DELIVERY = FindingSpec(
"perf.throughput_no_delivery", Category.PERFORMANCE, Severity.HIGH,
"No throughput traffic arrived, although the server sent it.",
)
val THROUGHPUT_BELOW_OFFERED = FindingSpec(
"perf.throughput_below_offered", Category.PERFORMANCE, Severity.LOW,
"Less throughput arrived than the server sent for the whole run.",
)
// ---- dns -------------------------------------------------------------------------
val DNS_ANSWER_REWRITTEN = FindingSpec(
"dns.answer_rewritten", Category.DNS, Severity.HIGH,
"A resolver returned an answer that differs from the authoritative record.",
)
val DNS_AUTHORITATIVE_UNREACHABLE = FindingSpec(
"dns.authoritative_unreachable", Category.DNS, Severity.MEDIUM,
"The canary zone's authoritative server could not be reached.",
)
// ---- v6 ----------------------------------------------------------------------------
//
// Prefix is `v6.`, matching the test-type registry (v6.brokenness, v6.happy_eyeballs, ...).
// These were `ipv6.*` while declaring Category.IPV6, but the prefix map only knows "v6", so
// they silently rolled up under connectivity: the third instance of a prefix disagreeing with
// its category and quietly moving a fault to a different verdict light.
/**
* Renamed from `v6.broken`, which claimed more than the evidence supports.
*
* The only signal behind it is ICMPv6 echo getting no reply and ICMPv6 echo is widely
* filtered on networks where IPv6 otherwise works perfectly. A phone that reported this while
* happily loading an IPv6-only site over TCP is what caught it. From here the two cases look
* identical, so the finding now says what was observed and names both explanations rather than
* picking one.
*
* It is worth reporting either way: filtered ICMPv6 breaks Path MTU Discovery, which is its
* own fault even when IPv6 works.
*/
/**
* A global IPv6 address with no default route.
*
* This is the structural version of the same complaint, and it is worth far more than the
* ICMP one because it admits no other explanation: the device has an address it cannot route
* with. Nothing is filtered, nothing is inferred the routing table says so directly, and it
* is already in the link snapshot.
*
* Not always a fault. A VPN that installs host routes to specific destinations produces
* exactly this shape on purpose, and it works. What makes it worth reporting either way is
* that applications cannot tell: having a global address, they will try IPv6 first and stall
* for every destination the routes do not cover.
*/
/**
* An IPv6 default route with no global address to use it from the mirror of
* [V6_NO_DEFAULT_ROUTE], and the more common misconfiguration of the two.
*
* The router is sending RAs that name it as a default gateway, but SLAAC produced no address:
* no prefix information option, or a prefix without the autonomous flag, or DHCPv6-only
* addressing the device did not complete. The network is announcing IPv6 service it does not
* actually deliver.
*
* This is worth flagging above the ICMP signal because it is both certain and consequential.
* Hosts see router advertisements, believe IPv6 is available, and pay a connection-attempt
* timeout on every dual-stack destination before falling back to IPv4 the classic "the
* internet feels slow" complaint with no packet loss anywhere to explain it.
*/
val V6_ROUTE_WITHOUT_ADDRESS = FindingSpec(
"v6.route_without_address", Category.IPV6, Severity.MEDIUM,
"The network advertises an IPv6 default route but the device has no global IPv6 address.",
rulesOut = "A working IPv6 setup: SLAAC did not produce a usable address on this link.",
)
/**
* A VPN prevented the underlying networks from being measured.
*
* Reported rather than worked around: Android refuses `Network.bindSocket()` on the networks
* beneath a VPN precisely so apps cannot leak around the tunnel, and that is correct
* behaviour. What is not acceptable is a run that quietly measures nothing and calls the
* result healthy, so this says plainly which networks went unmeasured and why.
*/
/**
* The network's DNS server answers, but this device cannot resolve through it.
*
* Worth separating from every other DNS failure because the remedy is somewhere else entirely.
* A name that will not resolve looks identical to a user whatever the cause, and the two causes
* pull in opposite directions: a server that does not answer means the network is broken and
* the router is the thing to examine, while a server that answers a direct query on a device
* that still cannot resolve means the platform resolver has wedged fixed by toggling wifi,
* and nothing to do with the network at all.
*
* Proven rather than inferred: the probe sends its own UDP query, bypassing the component under
* suspicion, and compares that against what the platform returns for the same name.
*/
/**
* The network hands out a search domain its DNS server will not answer for.
*
* A resolver appends search domains to lookups, so every name a client asks about can stall on
* a domain the server ignores. The failure mode is silence rather than a negative answer, and
* silence is indistinguishable from packet loss: clients retry instead of moving on, and some
* give up on the lookup entirely. That makes it look like the device is broken when the
* network is.
*
* Whether it bites depends on the resolver some try the bare name first and never notice
* which is why two devices on the same network can disagree about whether DNS works.
*/
val DNS_SEARCH_DOMAIN_UNANSWERED = FindingSpec(
"dns.search_domain_unanswered", Category.DNS, Severity.HIGH,
"The network advertises a DNS search domain that its own server does not answer for.",
rulesOut = "A fault on this device: the same server answers ordinary names normally.",
)
val DNS_SYSTEM_RESOLVER_BROKEN = FindingSpec(
"dns.system_resolver_broken", Category.DNS, Severity.HIGH,
"The network's DNS server answers, but this device cannot resolve names through it.",
rulesOut = "A network fault: the server replied to a query sent from this device.",
)
val MEASUREMENT_VPN_CONSTRAINED = FindingSpec(
"measurement.vpn_constrained", Category.CONNECTIVITY, Severity.INFO,
"A VPN was active, so the networks underneath it could not be measured.",
rulesOut = "Nothing — this run says little about the underlying network either way.",
)
val V6_NO_DEFAULT_ROUTE = FindingSpec(
"v6.no_default_route", Category.IPV6, Severity.MEDIUM,
"The device has a global IPv6 address but no IPv6 default route.",
rulesOut = "Guesswork: this is read from the routing table, not inferred from silence.",
)
val V6_NO_ICMP_REPLY = FindingSpec(
"v6.no_icmp_reply", Category.IPV6, Severity.LOW,
"IPv6 is configured but ICMPv6 echo gets no reply.",
rulesOut = "Nothing on its own: IPv6 may work fine with ICMP filtered.",
)
/**
* IPv6 is advertised and does not work the claim `v6.broken` originally made on ICMP
* silence alone, now reinstated because it can finally be backed: it is only emitted when a
* real IPv6 TCP connection (v6.brokenness) failed on the same network whose ICMPv6 went
* unanswered. Two independent transports failing on a network that advertises IPv6 is what
* "broken" actually means; either signal alone still gets [V6_NO_ICMP_REPLY].
*/
val V6_BROKEN = FindingSpec(
"v6.broken", Category.IPV6, Severity.HIGH,
"IPv6 is advertised on this network but carries no traffic.",
rulesOut = "ICMP filtering as the benign explanation: a TCP connection over IPv6 failed too.",
)
/**
* INFO deliberately, and it needs to stay that way.
*
* Most networks still do not offer IPv6, and that is not a fault. Reporting it as a warning
* lights a yellow verdict on a perfectly healthy network, which teaches people to ignore the
* light the one thing a diagnostic must never do.
*/
val V6_NOT_OFFERED = FindingSpec(
"v6.not_offered", Category.IPV6, Severity.INFO,
"This network does not offer IPv6.",
)
/** Every registered finding, in declaration order. */
val all: List<FindingSpec> = listOf(
UDP_UNREACHABLE, UDP_UNREACHABLE_UPSTREAM, UDP_LOSS, LOSS_UPSTREAM, LOSS_DOWNSTREAM,
DOWNSTREAM_BLOCKED, DOWNSTREAM_REORDER, CAPTIVE_PORTAL, NO_INTERNET, LINK_FLAPPING,
MTU_REDUCED_DOWNSTREAM, MTU_DOWNSTREAM_BLACKHOLE, FRAGMENTS_BLOCKED,
FRAGMENT_REORDER_SENSITIVE,
NAT_UDP_REBINDING, NAT_SYMMETRIC,
THROUGHPUT_NO_DELIVERY, THROUGHPUT_BELOW_OFFERED,
DNS_ANSWER_REWRITTEN, DNS_AUTHORITATIVE_UNREACHABLE,
DNS_SEARCH_DOMAIN_UNANSWERED, DNS_SYSTEM_RESOLVER_BROKEN, MEASUREMENT_VPN_CONSTRAINED,
V6_NO_DEFAULT_ROUTE, V6_ROUTE_WITHOUT_ADDRESS, V6_NO_ICMP_REPLY, V6_BROKEN, V6_NOT_OFFERED,
)
private val byCode: Map<String, FindingSpec> = all.associateBy { it.code }
fun byCode(code: String): FindingSpec? = byCode[code]
}
@@ -18,6 +18,39 @@ data class Network(
val wifi: Wifi? = null, val wifi: Wifi? = null,
val cellular: Cellular? = null, val cellular: Cellular? = null,
val changes: List<NetworkChange> = emptyList(), val changes: List<NetworkChange> = emptyList(),
@SerialName("system_verdict") val systemVerdict: SystemVerdict? = null,
/**
* Whether an ordinary app may send on this network at all.
*
* False for the carrier's special-purpose networks IMS/VoLTE, MMS, XCAP which appear
* beside the real ones in Android's list and carry neither `INTERNET` nor `NOT_RESTRICTED`.
* Binding to those needs `CONNECTIVITY_USE_RESTRICTED_NETWORKS`, a privileged permission no
* normal app can hold, so the refusal is permanent and says nothing about the network's
* health. Recorded rather than hidden: a reader seeing an interface with no measurements
* against it deserves to know the OS forbade them, instead of concluding the link is dead.
*/
@SerialName("app_usable") val appUsable: Boolean? = null,
)
/**
* What Android itself concluded about a network, as opposed to what we measured.
*
* Recorded because it is the verdict the user can see the "no internet" warning in the status
* bar and because it is free: the platform has already done the work by the time a run starts.
*
* Its real value is disagreement. When Android says a network is unusable and our own probes reach
* the internet regardless, the fault is in the device rather than the network, and that distinction
* is the difference between "fix your router" and "toggle your wifi". Neither number alone can say
* that; only the two together.
*/
@Serializable
data class SystemVerdict(
/** Android's own connectivity check passed. Null when the platform did not say. */
val validated: Boolean? = null,
/** Android believes a captive portal is intercepting this network. */
@SerialName("captive_portal") val captivePortal: Boolean? = null,
/** Some traffic works and some does not — Android's own hedge. */
@SerialName("partial_connectivity") val partialConnectivity: Boolean? = null,
) )
@Serializable @Serializable
@@ -111,3 +144,38 @@ data class NetworkChange(
val kind: String, // lost | gained | link_changed val kind: String, // lost | gained | link_changed
val detail: JsonObject? = null, val detail: JsonObject? = null,
) )
/**
* What a network's `changes[]` add up to.
*
* Lives beside the type rather than in the collector that produces it because two independent
* consumers ask the same question the watcher, computing its metrics, and the run engine,
* deciding whether to emit `connectivity.link_flapping` and a document whose metric and finding
* disagreed about how many times the link dropped would be worse than one reporting neither.
*/
object NetworkChanges {
const val LOST = "lost"
const val GAINED = "gained"
const val LINK_CHANGED = "link_changed"
/**
* Completed drop-and-return cycles: a `lost` with a later `gained` on the same network.
*
* A cycle has to *complete*. A link that goes away at minute four and is still gone when the
* window closes was not flapping it was switched off, or the device was carried out of
* range, and calling that the same fault would put a phone in a lift beside a failing access
* point.
*/
fun flapCycles(kinds: List<String>): Int {
var cycles = 0
var down = false
for (k in kinds) {
if (k == LOST) down = true
else if (k == GAINED && down) { cycles++; down = false }
}
return cycles
}
fun flapCyclesOf(changes: List<NetworkChange>): Int = flapCycles(changes.map { it.kind })
}
@@ -35,6 +35,9 @@ data class CategorySummary(
* (critical|high red, medium|low yellow, info/none green). * (critical|high red, medium|low yellow, info/none green).
* - A category is `inconclusive` when > 50% of its tests are failed/unsupported. * - A category is `inconclusive` when > 50% of its tests are failed/unsupported.
* - Overall = the worst category light; `inconclusive` only when ALL categories are. * - Overall = the worst category light; `inconclusive` only when ALL categories are.
* - A run whose per-network probing was blocked is `inconclusive` outright, whatever the
* categories say. The lights describe what the tests found; when the OS refused to let the
* tests run, a green light would describe nothing at all.
* *
* The mapping test-type category comes from [TestType.category]. Only categories that have * The mapping test-type category comes from [TestType.category]. Only categories that have
* findings or tests appear in the summary. * findings or tests appear in the summary.
@@ -44,7 +47,10 @@ object Verdicts {
private fun isInconclusiveTest(s: TestStatus) = private fun isInconclusiveTest(s: TestStatus) =
s == TestStatus.FAILED || s == TestStatus.UNSUPPORTED s == TestStatus.FAILED || s == TestStatus.UNSUPPORTED
fun derive(tests: List<Test>, findings: List<Finding>): Summary { fun derive(tests: List<Test>, findings: List<Finding>): Summary =
derive(tests, findings, Constraints())
fun derive(tests: List<Test>, findings: List<Finding>, constraints: Constraints): Summary {
val testsByCat = tests.groupBy { TestType.category(it.type) } val testsByCat = tests.groupBy { TestType.category(it.type) }
val findingsByCat = findings.groupBy { it.category } val findingsByCat = findings.groupBy { it.category }
val categories = (testsByCat.keys + findingsByCat.keys) val categories = (testsByCat.keys + findingsByCat.keys)
@@ -72,7 +78,14 @@ object Verdicts {
) )
} }
val overall = deriveOverall(perCat.values) // A run that could not measure the networks it was asked about has not found them
// healthy; it has found out nothing. Reporting that as green is the single most
// misleading thing this function could do, so the constraint outranks the lights.
val overall = if (constraints.perNetworkBlocked) {
Verdict.INCONCLUSIVE
} else {
deriveOverall(perCat.values)
}
return Summary(overall = overall, categories = perCat) return Summary(overall = overall, categories = perCat)
} }
@@ -77,6 +77,8 @@ object TestType {
const val MTU_BLACKHOLE = "mtu.blackhole" const val MTU_BLACKHOLE = "mtu.blackhole"
const val MTU_MSS_OBSERVED = "mtu.mss_observed" const val MTU_MSS_OBSERVED = "mtu.mss_observed"
const val MTU_FRAG_DELIVERY = "mtu.frag_delivery" const val MTU_FRAG_DELIVERY = "mtu.frag_delivery"
/** Whether fragments survive arriving out of order, not merely whether they survive. */
const val MTU_FRAG_ORDERING = "mtu.frag_ordering"
// nat // nat
const val NAT_STUN_5780 = "nat.stun_5780" const val NAT_STUN_5780 = "nat.stun_5780"
const val NAT_MAPPING_LIFETIME_UDP = "nat.mapping_lifetime_udp" const val NAT_MAPPING_LIFETIME_UDP = "nat.mapping_lifetime_udp"
@@ -87,6 +89,13 @@ object TestType {
// dns // dns
const val DNS_RESOLVER_INVENTORY = "dns.resolver_inventory" const val DNS_RESOLVER_INVENTORY = "dns.resolver_inventory"
const val DNS_CANARY = "dns.canary" const val DNS_CANARY = "dns.canary"
/**
* Does this device's own resolver work, as distinct from the network's DNS.
*
* Registry addition, v1.1. Kept apart from [DNS_CANARY], which asks whether answers are being
* tampered with; this asks whether answers arrive at all, and where the failure sits.
*/
const val DNS_RESOLVER = "dns.resolver"
const val DNS_INTERCEPTION = "dns.interception" const val DNS_INTERCEPTION = "dns.interception"
const val DNS_TTL_INTEGRITY = "dns.ttl_integrity" const val DNS_TTL_INTEGRITY = "dns.ttl_integrity"
const val DNS_ANSWER_INTEGRITY = "dns.answer_integrity" const val DNS_ANSWER_INTEGRITY = "dns.answer_integrity"
@@ -124,6 +133,16 @@ object TestType {
const val LOCAL_MDNS_INVENTORY = "local.mdns_inventory" const val LOCAL_MDNS_INVENTORY = "local.mdns_inventory"
const val LOCAL_SSDP_INVENTORY = "local.ssdp_inventory" const val LOCAL_SSDP_INVENTORY = "local.ssdp_inventory"
const val LOCAL_LLMNR_INVENTORY = "local.llmnr_inventory" const val LOCAL_LLMNR_INVENTORY = "local.llmnr_inventory"
/**
* WS-Discovery (UDP 3702) and NetBIOS name service (UDP 137). Registry additions, v1.2.
*
* Both are passive: the traffic is broadcast to the segment whether or not anyone asks, so
* listening is the whole measurement. They earn their own ids rather than folding into
* [LOCAL_SSDP_INVENTORY] because what they imply differs WS-Discovery inventories printers
* and cameras, while NetBIOS/LLMNR chatter is a security finding in its own right.
*/
const val LOCAL_WSD_INVENTORY = "local.wsd_inventory"
const val LOCAL_NETBIOS_INVENTORY = "local.netbios_inventory"
const val LOCAL_GATEWAY_SERVICES = "local.gateway_services" const val LOCAL_GATEWAY_SERVICES = "local.gateway_services"
const val LOCAL_NTP = "local.ntp" const val LOCAL_NTP = "local.ntp"
// peer // peer
@@ -0,0 +1,73 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.measurement
/**
* The two ways a network can be half-configured for IPv6, read from the link snapshot.
*
* Pure model logic rather than something a ViewModel does, because "is this network's IPv6
* broken, and in which direction" is exactly the kind of judgement that should be checkable
* against a captured routing table without a phone in the loop.
*/
object V6Analysis {
/** Linux tunnel interfaces: WireGuard/Netbird (tun*, wg*), plus the usual VPN names. */
private val TUNNEL_IFACE = Regex("""^(tun|tap|wg|ppp|ipsec|utun)\d*$""")
/** What one network's IPv6 configuration looks like. */
data class Shape(
val iface: String,
/** A global address with no ::/0 route: an address the device cannot route with. */
val addressWithoutRoute: Boolean,
/** A ::/0 route with no global address: a route the device cannot source from. */
val routeWithoutAddress: Boolean,
/** The routes belong to a tunnel, so a partial view of IPv6 is likely deliberate. */
val tunnel: Boolean,
)
/**
* Classifies each network's IPv6 configuration.
*
* Both shapes are read straight from the link snapshot rather than inferred from silence, so
* unlike an ICMP signal there is no competing explanation for what was observed and both
* matter for the same reason: an application cannot tell in advance, so it tries IPv6 first
* and waits.
*
* They differ in what they mean. An address with no route is what a VPN installing host routes
* to specific destinations produces on purpose, and it works; calling that a fault would be the
* "lack of IPv6 is a yellow condition" mistake in a new costume, so a tunnel downgrades it to
* information. A route with no address is the opposite: the router advertised itself as a
* default gateway but SLAAC produced nothing usable, so the network is announcing IPv6 service
* it does not deliver. That one is a real misconfiguration however it arises.
*/
fun classify(networks: List<Network>): List<Shape> = networks.map { n ->
val globalV6 = n.link.addresses.any { isGlobalV6(it.addr) }
val v6Routes = n.link.routes.filter { it.dst.contains(':') }
val hasDefault = v6Routes.any { it.dst == "::/0" }
Shape(
iface = n.iface ?: v6Routes.firstOrNull()?.iface.orEmpty(),
addressWithoutRoute = globalV6 && !hasDefault,
routeWithoutAddress = hasDefault && !globalV6,
// Android labels the transport itself, which beats guessing from a name; the regex
// stays as a backstop for tunnels Android does not own (a userspace WireGuard, say,
// or anything seen through the shell tier).
tunnel = n.transport == Transport.VPN ||
v6Routes.any { TUNNEL_IFACE.containsMatchIn(it.iface.orEmpty()) },
)
}
/**
* Whether an address is IPv6 and usable as a source for off-link traffic.
*
* ULAs count. A ULA is not globally routable, but it is a global-*scope* address the stack
* will happily select as a source, which is the property that matters here an overlay
* network handing out fc00::/7 addresses is providing working IPv6 to the destinations it
* carries, and treating that as "no address" would misreport every VPN as broken.
*/
private fun isGlobalV6(addr: String): Boolean {
if (!addr.contains(':')) return false
val a = addr.substringBefore('%').lowercase() // strip any zone index
return !a.startsWith("fe80") && a != "::1" && a != "::"
}
}
@@ -0,0 +1,133 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.measurement
import java.io.File
import kotlin.test.Test
import kotlin.test.assertEquals
import kotlin.test.assertTrue
import kotlin.test.fail
/**
* Keeps the finding registry honest.
*
* The interesting test is the last one: it reads `docs/findings-registry.md` and fails when the
* document and the code disagree. Documentation that drifts from its implementation is worse than
* none, because it still looks authoritative and a finding registry is precisely the artifact
* other people build tooling against.
*/
class FindingRegistryTest {
@Test
fun codesAreUnique() {
val dupes = FindingRegistry.all.groupBy { it.code }.filterValues { it.size > 1 }.keys
assertTrue(dupes.isEmpty(), "duplicate finding codes: $dupes")
}
@Test
fun everyDeclaredSpecIsInTheAllList() {
// Reflection over the object's properties: a spec that is declared but left out of `all`
// is invisible to the doc check and to any consumer enumerating the registry.
val declared = FindingRegistry::class.java.declaredMethods
.filter { it.parameterCount == 0 && it.returnType == FindingSpec::class.java }
.mapNotNull { runCatching { it.invoke(FindingRegistry) as FindingSpec }.getOrNull() }
.map { it.code }
.toSet()
val listed = FindingRegistry.all.map { it.code }.toSet()
assertEquals(declared, listed, "declared specs and the `all` list disagree")
}
// The prefix decides the category, and the category decides which verdict light the finding
// rolls up into. A code whose prefix disagrees with its category silently moves a fault to a
// different light — the exact bug that got two codes renamed out of nat.*.
@Test
fun everyPrefixMatchesItsCategory() {
for (spec in FindingRegistry.all) {
val fromPrefix = TestType.category(spec.code)
assertEquals(
fromPrefix, spec.category,
"${spec.code} is declared as ${spec.category} but its prefix maps to $fromPrefix",
)
}
}
@Test
fun codesFollowTheNamingConvention() {
val shape = Regex("^[a-z0-9]+\\.[a-z0-9_]+$")
for (spec in FindingRegistry.all) {
assertTrue(shape.matches(spec.code), "malformed code: ${spec.code}")
assertTrue(spec.meaning.isNotBlank(), "${spec.code} has no meaning")
assertTrue(
spec.meaning.trimEnd().endsWith("."),
"${spec.code}'s meaning should be a sentence: '${spec.meaning}'",
)
}
}
// Two near-identical codes are how one fault ends up split across two dashboards. This is a
// blunt check — it will not catch every synonym — but it catches the shape that already
// happened: the same words in a different order.
@Test
fun noTwoCodesAreAnagramsOfEachOther() {
val normalised = FindingRegistry.all.associate { spec ->
spec.code to spec.code.substringAfter('.').split('_').sorted().joinToString("_")
}
val clashes = normalised.entries.groupBy { it.value }.filterValues { it.size > 1 }
if (clashes.isNotEmpty()) {
fail("codes differing only in word order: ${clashes.values.map { g -> g.map { it.key } }}")
}
}
@Test
fun theDocumentAndTheRegistryAgree() {
val doc = findDoc() ?: run {
println("findings-registry.md not found from ${File(".").absolutePath} — skipping")
return
}
val text = doc.readText()
// Only table rows count as "documented". Prose may legitimately mention a code that no
// longer exists — the rules section explains why two were merged — and treating that as
// a registry entry would force the document to forget its own history.
val documented = text.lines()
.filter { it.trimStart().startsWith("|") }
.flatMap { row -> Regex("`([a-z0-9]+\\.[a-z0-9_]+)`").findAll(row).map { it.groupValues[1] } }
.toSet()
val registered = FindingRegistry.all.map { it.code }.toSet()
val missingFromDoc = registered - documented
val missingFromCode = documented - registered
assertTrue(
missingFromDoc.isEmpty(),
"these codes exist in FindingRegistry but not in docs/findings-registry.md: $missingFromDoc",
)
assertTrue(
missingFromCode.isEmpty(),
"docs/findings-registry.md documents codes that no longer exist: $missingFromCode",
)
// And the severities must match, or the document is describing a different system.
for (spec in FindingRegistry.all) {
val row = text.lines().firstOrNull {
it.trimStart().startsWith("|") && it.contains("`${spec.code}`")
} ?: continue
val severity = spec.severity.name.lowercase()
assertTrue(
row.contains("| $severity |"),
"${spec.code} is ${severity} in code but the doc row says otherwise: $row",
)
}
}
/** Walks up from the test's working directory to find the repo's docs/ folder. */
private fun findDoc(): File? {
var dir: File? = File(".").absoluteFile
repeat(6) {
val candidate = File(dir, "docs/findings-registry.md")
if (candidate.isFile) return candidate
dir = dir?.parentFile
}
return null
}
}
@@ -0,0 +1,65 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.measurement
import kotlin.test.Test
import kotlin.test.assertEquals
/**
* Pins the counting behind `connectivity.link_flapping`.
*
* The finding claims a link went away and came back, and its severity escalates on repetition, so
* this is arithmetic a person reading a report will act on. The cases that matter are the ones
* where the naive count is wrong: a link still down when the window closed, and a run that started
* while the link was already gone.
*/
class NetworkChangesTest {
@Test
fun aQuietWindowHasNoCycles() {
assertEquals(0, NetworkChanges.flapCycles(emptyList()))
assertEquals(0, NetworkChanges.flapCycles(listOf("link_changed", "link_changed")))
}
@Test
fun oneDropAndReturnIsOneCycle() {
assertEquals(1, NetworkChanges.flapCycles(listOf("lost", "gained")))
assertEquals(
1,
NetworkChanges.flapCycles(listOf("link_changed", "lost", "link_changed", "gained")),
)
}
@Test
fun repeatedDropsCountSeparately() {
assertEquals(3, NetworkChanges.flapCycles(listOf("lost", "gained", "lost", "gained", "lost", "gained")))
}
// A link that is still down when the run ends was not flapping — it was switched off, or the
// device left its range. Counting that as a cycle would put a phone in a lift beside a failing
// access point.
@Test
fun aDropThatNeverReturnsIsNotACycle() {
assertEquals(0, NetworkChanges.flapCycles(listOf("lost")))
assertEquals(1, NetworkChanges.flapCycles(listOf("lost", "gained", "lost")))
}
// The mirror case: the window opened while the network was already gone, so its return is the
// first thing seen. Nothing was watched dropping, so nothing is claimed.
@Test
fun aReturnWithNoObservedDropIsNotACycle() {
assertEquals(0, NetworkChanges.flapCycles(listOf("gained")))
assertEquals(0, NetworkChanges.flapCycles(listOf("gained", "link_changed")))
}
@Test
fun theChangeOverloadAgreesWithTheKindsOverload() {
val changes = listOf(
NetworkChange(atMonoNs = 1, kind = NetworkChanges.LOST),
NetworkChange(atMonoNs = 2, kind = NetworkChanges.GAINED),
NetworkChange(atMonoNs = 3, kind = NetworkChanges.LINK_CHANGED),
)
assertEquals(1, NetworkChanges.flapCyclesOf(changes))
}
}
@@ -0,0 +1,125 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.measurement
import kotlin.test.Test
import kotlin.test.assertEquals
import kotlin.test.assertFalse
import kotlin.test.assertTrue
/**
* The fixtures here are a real device's routing table, transcribed from `dumpsys connectivity`
* on a OnePlus 15 with a Netbird tunnel up: wifi advertising a default route it cannot source
* from, cellular working properly, and a VPN carrying host routes to two destinations.
*
* Using a captured table rather than invented ones matters, because the bug this guards against
* is not "the boolean logic is wrong" it is "the shapes I imagined are not the shapes real
* networks produce".
*/
class V6AnalysisTest {
private fun net(
id: String,
transport: Transport,
iface: String,
addrs: List<String>,
routes: List<Pair<String, String>>,
) = Network(
id = id,
transport = transport,
iface = iface,
link = Link(
addresses = addrs.map { Address(addr = it.substringBefore('/'), prefixLen = 64) },
routes = routes.map { (dst, dev) -> Route(dst = dst, iface = dev) },
),
)
/** wlan0: an IPv6 default route via a link-local gateway, but SLAAC produced no address. */
private val wifi = net(
"w", Transport.WIFI, "wlan0",
addrs = listOf("fe80::bcf6:edff:fe67:b139", "10.13.102.122"),
routes = listOf(
"fe80::/64" to "wlan0",
"::/0" to "wlan0",
"0.0.0.0/0" to "wlan0",
),
)
/** rmnet_data1: a properly configured cellular link — global address and a default route. */
private val cellular = net(
"c", Transport.CELLULAR, "rmnet_data1",
addrs = listOf("2001:4bb8:46a:e724:289d:87ff:feb6:ebd3"),
routes = listOf("::/0" to "rmnet_data1", "2001:4bb8:46a:e724::/64" to "rmnet_data1"),
)
/** tun1: Netbird, with a ULA and host routes to exactly two destinations. */
private val vpn = net(
"v", Transport.VPN, "tun1",
addrs = listOf("100.64.158.131", "fdfd:c4fe:c4fe:c4fe:1f3c:98a0:dd66:ac7"),
routes = listOf(
"2001:1ad0:c4fe:6767::2/128" to "tun1",
"2001:1ad0:c4fe:a::136/128" to "tun1",
"fdfd:c4fe:c4fe:c4fe::/64" to "tun1",
),
)
@Test
fun `wifi advertising a route it cannot source from is reported`() {
val s = V6Analysis.classify(listOf(wifi)).single()
assertTrue(s.routeWithoutAddress, "::/0 with only a link-local address is the RA-without-SLAAC case")
assertFalse(s.addressWithoutRoute)
assertFalse(s.tunnel, "wifi is not a tunnel")
assertEquals("wlan0", s.iface)
}
@Test
fun `a properly configured link produces no finding`() {
val s = V6Analysis.classify(listOf(cellular)).single()
assertFalse(s.routeWithoutAddress)
assertFalse(s.addressWithoutRoute)
}
@Test
fun `a tunnel with host routes is deliberate, not broken`() {
val s = V6Analysis.classify(listOf(vpn)).single()
assertTrue(s.addressWithoutRoute, "a ULA and no ::/0 is an address with nothing to route it")
assertTrue(s.tunnel, "so it must be reported as information, not as a fault")
assertFalse(s.routeWithoutAddress)
}
@Test
fun `each network is judged on its own`() {
// The whole point of per-network classification: "IPv6 is broken" is useless advice when
// wifi is the broken one and cellular is fine.
val shapes = V6Analysis.classify(listOf(wifi, cellular, vpn)).associateBy { it.iface }
assertTrue(shapes.getValue("wlan0").routeWithoutAddress)
assertFalse(shapes.getValue("rmnet_data1").routeWithoutAddress)
assertFalse(shapes.getValue("rmnet_data1").addressWithoutRoute)
assertTrue(shapes.getValue("tun1").addressWithoutRoute)
}
@Test
fun `a link-local-only network with no v6 route says nothing either way`() {
// Plain IPv4-only wifi: no IPv6 offered at all. That is v6.not_offered's business, and
// reporting it here as well would double up on a network that is merely legacy, not broken.
val v4only = net(
"4", Transport.WIFI, "wlan0",
addrs = listOf("fe80::1", "192.168.1.5"),
routes = listOf("0.0.0.0/0" to "wlan0"),
)
val s = V6Analysis.classify(listOf(v4only)).single()
assertFalse(s.routeWithoutAddress)
assertFalse(s.addressWithoutRoute)
}
@Test
fun `a zone index does not hide a link-local address`() {
val zoned = net(
"z", Transport.WIFI, "wlan0",
addrs = listOf("fe80::1%wlan0"),
routes = listOf("::/0" to "wlan0"),
)
assertTrue(V6Analysis.classify(listOf(zoned)).single().routeWithoutAddress)
}
}
+5 -1
View File
@@ -20,4 +20,8 @@ kotlin {
} }
java { sourceCompatibility = JavaVersion.VERSION_17; targetCompatibility = JavaVersion.VERSION_17 } java { sourceCompatibility = JavaVersion.VERSION_17; targetCompatibility = JavaVersion.VERSION_17 }
tasks.test { useJUnitPlatform() } tasks.test {
useJUnitPlatform()
// Opt-in: point this at a captured run to check the anonymizer against real data.
System.getenv("ECHOLOT_REAL_RUN")?.let { environment("ECHOLOT_REAL_RUN", it) }
}
@@ -108,12 +108,22 @@ class Anonymizer(private val level: PrivacyLevel, private val salt: Salt) {
is JsonObject -> walkObject(v, path) is JsonObject -> walkObject(v, path)
is JsonArray -> JsonArray(v.map { walk(key, it, path) }) is JsonArray -> JsonArray(v.map { walk(key, it, path) })
is JsonPrimitive -> is JsonPrimitive ->
if (v.isString) transform(Classification.typeOf(key, path), v.content).let(::JsonPrimitive) if (v.isString) {
else v // Name first (it is precise), then shape (it is exhaustive). A field nobody
// classified must not be a field that leaks.
val type = Classification.typeOf(key, path) ?: Classification.inferFromValue(v.content)
JsonPrimitive(transform(type, v.content))
} else {
v
}
} }
private fun transform(type: LogicalType?, value: String): String = when (type) { private fun transform(type: LogicalType?, value: String): String = when (type) {
null -> value // Unclassified strings still get their *embedded* identifiers scrubbed. A whole-value
// check cannot see them: raw shell output is one long string that is neither a MAC nor an
// address, so it sailed through both the name table and the shape check carrying every
// MAC on the user's LAN.
null -> scrubEmbedded(value)
LogicalType.SSID -> pseudo("ssid", value) { "net-" + it.take(6) } LogicalType.SSID -> pseudo("ssid", value) { "net-" + it.take(6) }
LogicalType.MAC, LogicalType.BSSID -> macPreservingOui(value) LogicalType.MAC, LogicalType.BSSID -> macPreservingOui(value)
LogicalType.IP4 -> ip4(value) LogicalType.IP4 -> ip4(value)
@@ -123,6 +133,38 @@ class Anonymizer(private val level: PrivacyLevel, private val salt: Salt) {
LogicalType.FREETEXT -> "[removed: may contain identifying text]" LogicalType.FREETEXT -> "[removed: may contain identifying text]"
} }
/**
* Replaces addresses and MACs found *inside* a longer string.
*
* Shizuku probes embed raw command output verbatim `ip neigh`, `ip route`, `dumpsys` which
* is genuinely valuable evidence and also a complete inventory of every device on the user's
* network, with hardware addresses. measurement-schema.md §9 flagged these as "hard to
* anonymize" and proposed dropping them from exports.
*
* Scrubbing beats dropping: the output stays readable and auditable you can still see the
* shape of the neighbour table and how many hosts there were while the identifiers become
* the same pseudonyms used everywhere else in the document. So a MAC appearing both in a
* parsed field and in a raw dump still reads as one device.
*
* Only addresses and MACs are touched, for the same reason as [Classification.inferFromValue]:
* they are the patterns that cannot be mistaken for something else in free text.
*/
private fun scrubEmbedded(value: String): String {
// Cheap bail-out: the overwhelming majority of strings are short and contain neither.
if (value.length < 7 || (!value.contains(':') && !value.contains('.'))) return value
// One pass, not three. Sequential passes re-process their own output: after a MAC became
// 78:9a:18:xx:yy:zz the IPv6 pattern matched it — six hex groups separated by colons is
// exactly an address — and mangled the vendor prefix that the MAC rule had just taken
// care to preserve. Ordered alternation resolves each position once, MAC first.
return EMBEDDED.replace(value) { m ->
when {
m.groups[1] != null -> macPreservingOui(m.value)
m.groups[2] != null -> ip6(m.value)
else -> ip4(m.value)
}
}
}
// ---- per-type transforms ------------------------------------------------------------- // ---- per-type transforms -------------------------------------------------------------
/** /**
@@ -147,6 +189,11 @@ class Anonymizer(private val level: PrivacyLevel, private val salt: Salt) {
* Public addresses keep only their /16 so the network is still locatable at ISP granularity. * Public addresses keep only their /16 so the network is still locatable at ISP granularity.
*/ */
private fun ip4(value: String): String { private fun ip4(value: String): String {
// A route destination carries a prefix length; pseudonymize the address and put it back,
// or "0.0.0.0/0" turns into nonsense and the routing table becomes unreadable.
value.substringAfter('/', "").takeIf { it.isNotEmpty() && value.contains('/') }?.let { len ->
return ip4(value.substringBefore('/')) + "/" + len
}
val o = value.split(".") val o = value.split(".")
if (o.size != 4 || o.any { it.toIntOrNull() == null }) return value if (o.size != 4 || o.any { it.toIntOrNull() == null }) return value
val n = o.map { it.toInt() } val n = o.map { it.toInt() }
@@ -167,8 +214,36 @@ class Anonymizer(private val level: PrivacyLevel, private val salt: Salt) {
* is a device fingerprint, especially with EUI-64. * is a device fingerprint, especially with EUI-64.
*/ */
private fun ip6(value: String): String { private fun ip6(value: String): String {
// Dotted quads reach here through the family-agnostic field names (addr, gateway, dst);
// hand them to the IPv4 path rather than mangling them as if they were v6.
if (value.count { it == ':' } < 2) return ip4(value)
if (value.contains('/')) {
return ip6(value.substringBefore('/')) + "/" + value.substringAfter('/')
}
val v = value.lowercase(Locale.ROOT) val v = value.lowercase(Locale.ROOT)
// The unspecified address and the default route are not identities; mangling them would
// make a routing table unreadable for no privacy gain.
if (v == "::1" || v == "::" || v.startsWith("fe80:") || v.startsWith("ff")) return v if (v == "::1" || v == "::" || v.startsWith("fe80:") || v.startsWith("ff")) return v
// Unique local addresses (fc00::/7) need the *whole* prefix replaced, not the tail.
//
// They look like the v6 equivalent of RFC1918, and the first instinct is to keep them for
// the same reason: private, topological, says nothing about anyone. That reasoning does
// not carry over. An RFC1918 prefix is shared by millions of networks and identifies
// none of them; a ULA global ID is 40 *random* bits, unique to one network by
// construction (RFC 4193). It is a network fingerprint. Passing the leading groups
// through - which is what the general path does - leaked 32 of those 40 bits.
//
// The prefix is pseudonymized as a unit, so two addresses on the same ULA subnet still
// land on the same pseudonymous prefix. "These hosts are on one network" survives;
// "this is *that* network" does not.
if (v.startsWith("fc") || v.startsWith("fd")) {
val groups = v.substringBefore('%').split(":")
val prefix = pseudo("ula-prefix", groups.take(3).joinToString(":")) { it }
val host = pseudo("ula-host", v) { it }
return "fd${prefix.substring(0, 2)}:${prefix.substring(2, 6)}:${prefix.substring(6, 10)}" +
"::${host.substring(0, 4)}"
}
val groups = v.substringBefore('%').split(":") val groups = v.substringBefore('%').split(":")
if (groups.size < 3) return v if (groups.size < 3) return v
val h = pseudo("ip6", value) { it } val h = pseudo("ip6", value) { it }
@@ -219,7 +294,7 @@ class Anonymizer(private val level: PrivacyLevel, private val salt: Salt) {
/** Deterministic per (domain, value, salt); memoized so one value maps to one pseudonym. */ /** Deterministic per (domain, value, salt); memoized so one value maps to one pseudonym. */
private fun pseudo(domain: String, value: String, shape: (String) -> String): String = private fun pseudo(domain: String, value: String, shape: (String) -> String): String =
cache.getOrPut("$domain$value") { cache.getOrPut("$domain\u0000$value") {
val md = MessageDigest.getInstance("SHA-256") val md = MessageDigest.getInstance("SHA-256")
md.update(salt.bytes) md.update(salt.bytes)
md.update(domain.toByteArray()) md.update(domain.toByteArray())
@@ -229,6 +304,21 @@ class Anonymizer(private val level: PrivacyLevel, private val salt: Salt) {
} }
private companion object { private companion object {
/**
* MAC | IPv6 | IPv4, in that order alternation is ordered, so a MAC-shaped token is
* claimed by the MAC rule before the IPv6 rule can see it.
*
* The patterns are deliberately conservative. A missed address is scrubbed by another
* rule or not at all; an over-eager one mangles timestamps, version strings and log
* prefixes, corrupting evidence to protect nothing.
*/
val EMBEDDED = Regex(
// Raw strings: a regex written with escaped escapes is a regex nobody can check.
"""(\b[0-9a-fA-F]{2}(?:[:-][0-9a-fA-F]{2}){5}\b)""" +
"""|(\b(?:[0-9a-fA-F]{1,4}:){2,7}(?::|[0-9a-fA-F]{1,4})(?:[0-9a-fA-F:]*))""" +
"""|(\b(?:\d{1,3}\.){3}\d{1,3}\b)"""
)
val publicSuffixes = setOf( val publicSuffixes = setOf(
"local", "lan", "home", "internal", "arpa", "local", "lan", "home", "internal", "arpa",
"com", "net", "org", "io", "app", "dev", "at", "de", "eu", "uk", "com", "net", "org", "io", "app", "dev", "at", "de", "eu", "uk",
@@ -32,6 +32,20 @@ object Classification {
"link_local", "ra_source", "prefix", "link_local", "ra_source", "prefix",
).forEach { put(it, LogicalType.IP6) } ).forEach { put(it, LogicalType.IP6) }
// Family-agnostic address fields — the names the models actually use (Address.addr,
// Route.gateway, Route.dst, DnsConfig.servers). Their absence here was a real leak: the
// device's own global IPv6 address went out verbatim at the level whose description
// promises addresses are pseudonymized. Typed IP6 because the transform detects the
// family from the value, falling through to the IPv4 path for a dotted quad.
listOf(
"addr", "address", "gateway", "dst", "src", "servers", "server", "resolver",
"next_hop", "via", "public_ip", "observed_ip",
// Who sent each passive-discovery announcement (SSDP / LLMNR / NetBIOS / WS-Discovery).
// Family-agnostic on purpose: the same field carries a dotted quad from a v4 group and
// a link-local from ff02::c, and the v6 transform hands dotted quads to the v4 path.
"source_ip",
).forEach { put(it, LogicalType.IP6) }
listOf("mac", "hw_addr", "gateway_mac", "router_mac", "sender_mac", "peer_mac") listOf("mac", "hw_addr", "gateway_mac", "router_mac", "sender_mac", "peer_mac")
.forEach { put(it, LogicalType.MAC) } .forEach { put(it, LogicalType.MAC) }
listOf("bssid", "ap_mac").forEach { put(it, LogicalType.BSSID) } listOf("bssid", "ap_mac").forEach { put(it, LogicalType.BSSID) }
@@ -40,13 +54,30 @@ object Classification {
listOf( listOf(
"fqdn", "hostname", "host", "name", "reverse_dns", "ptr", "domain", "query_name", "fqdn", "hostname", "host", "name", "reverse_dns", "ptr", "domain", "query_name",
"friendly_name", "server_name", "sni", "cname", "search_domain", "device_name", "friendly_name", "server_name", "sni", "cname", "search_domain", "device_name",
// Plural and prefixed variants the models actually use.
"search_domains", "private_dns_hostname", "domains", "hostnames",
// Names a device shouts at the whole segment. An LLMNR question and a NetBIOS
// registration are hostnames in every sense that matters here — they name a machine on
// somebody's home network — so they get the same per-label treatment as any other.
"netbios_name",
).forEach { put(it, LogicalType.FQDN) } ).forEach { put(it, LogicalType.FQDN) }
listOf("session_id", "credential", "token", "device_id", "android_id", "serial", "imsi", "iccid") listOf(
.forEach { put(it, LogicalType.OPAQUE_ID) } "session_id", "credential", "token", "device_id", "android_id", "serial", "imsi", "iccid",
// Discovery identities. A UPnP USN and a WS-Discovery endpoint UUID are stable,
// globally unique per device and frequently derived from a serial number — exactly the
// thing that lets two uploads be recognised as the same household.
"usn", "device_uuid",
).forEach { put(it, LogicalType.OPAQUE_ID) }
listOf("notes", "detail", "raw", "excerpt", "location", "model_description") listOf(
.forEach { put(it, LogicalType.FREETEXT) } "notes", "detail", "raw", "excerpt", "location", "model_description",
// The make/model a device volunteers, and the URLs it points at. `server_banner` is
// deliberately not named `server`, which the family-agnostic address block already
// claims — a SERVER header run through the address transform would be mangled into
// nonsense while protecting nothing.
"server_banner", "product_hint", "wsd_types", "wsd_xaddrs",
).forEach { put(it, LogicalType.FREETEXT) }
} }
/** /**
@@ -67,6 +98,13 @@ object Classification {
private val droppedKeys = setOf( private val droppedKeys = setOf(
"ssdp_responders", "upnp", "neighbors", "arp_table", "scan_results", "ssdp_responders", "upnp", "neighbors", "arp_table", "scan_results",
"nearby_networks", "peers", "raw_dump", "dumpsys", "nearby_networks", "peers", "raw_dump", "dumpsys",
// The long run's passive-discovery inventories. Same argument as `ssdp_responders`, only
// more so: these are minutes of everything the segment said about itself — device models,
// hostnames, printers, who is looking for whom. Even fully pseudonymized the *shape* of a
// household is a fingerprint, no metric depends on the list (the counts live in `metrics`,
// which survives), and the collectors' own status fields stay behind to say the capture
// worked. Dropping beats mangling.
"ssdp_devices", "llmnr_queries", "netbios_names", "wsd_devices",
) )
fun typeOf(key: String, path: List<String>): LogicalType? { fun typeOf(key: String, path: List<String>): LogicalType? {
@@ -78,6 +116,49 @@ object Classification {
return null return null
} }
/**
* Last-resort classification from the *value*, when the field name is unrecognised.
*
* A name table can only protect fields somebody remembered to add, which is the wrong
* property for a privacy control: the dangerous field is the one nobody thought of. This
* exists because that failed once already `addresses[].addr` holds the device's own global
* IPv6 address, the table had never heard of the name, and it went out verbatim.
*
* Only addresses and MACs are inferred, because only those have shapes that cannot be
* mistaken for something else. Hostnames deliberately are not: `train.udp_updown` is
* indistinguishable from a domain by shape, and mangling a test type would corrupt the
* document to protect nothing.
*/
fun inferFromValue(value: String): LogicalType? {
val v = value.trim()
if (v.isEmpty() || v.length > 64) return null
if (looksLikeMac(v)) return LogicalType.MAC
if (looksLikeIp6(v)) return LogicalType.IP6
if (looksLikeIp4(v)) return LogicalType.IP4
return null
}
private fun isHex(c: Char) = c in '0'..'9' || c in 'a'..'f' || c in 'A'..'F'
private fun looksLikeMac(v: String): Boolean {
val parts = v.split(':', '-')
return parts.size == 6 && parts.all { p -> p.length == 2 && p.all(::isHex) }
}
private fun looksLikeIp4(v: String): Boolean {
val parts = v.substringBefore('/').split('.')
return parts.size == 4 && parts.all { p ->
p.isNotEmpty() && p.length <= 3 && p.all(Char::isDigit) && p.toInt() <= 255
}
}
private fun looksLikeIp6(v: String): Boolean {
val core = v.substringBefore('/').substringBefore('%')
// Two colons minimum, so a time or a MAC fragment does not qualify, and nothing but the
// characters an address may contain.
return core.count { it == ':' } >= 2 && core.all { it == ':' || isHex(it) }
}
fun dropAtBalanced(path: List<String>): Boolean { fun dropAtBalanced(path: List<String>): Boolean {
if (path.isNotEmpty() && path.last() in droppedKeys) return true if (path.isNotEmpty() && path.last() in droppedKeys) return true
return droppedPaths.any { dropped -> dropped.all { path.contains(it) } } return droppedPaths.any { dropped -> dropped.all { path.contains(it) } }
@@ -182,4 +182,47 @@ class AnonymizerTest {
assertEquals(PrivacyLevel.BALANCED, PrivacyLevel.max(PrivacyLevel.BALANCED, PrivacyLevel.FULL)) assertEquals(PrivacyLevel.BALANCED, PrivacyLevel.max(PrivacyLevel.BALANCED, PrivacyLevel.FULL))
assertEquals(PrivacyLevel.FULL, PrivacyLevel.fromWire("nonsense")) assertEquals(PrivacyLevel.FULL, PrivacyLevel.fromWire("nonsense"))
} }
// A ULA looks like the v6 RFC1918 and is not. Its global ID is 40 random bits, unique to one
// network by construction (RFC 4193), so the prefix IS the identifier - unlike 192.168.x,
// which millions of networks share. Passing the leading groups through leaked most of it.
@Test
fun ulaPrefixesArePseudonymizedWhole() {
val doc = json.parseToJsonElement(
"""{"run":{"id":"r"},"networks":[{"link":{"dns":{"servers":["fda1:3fb1:ff92:6696::2662"]}}}]}"""
).jsonObject
val out = flat(anon(PrivacyLevel.BALANCED, doc))
assertFalse(out.contains("fda1"), "the ULA global ID survived: $out")
assertFalse(out.contains("3fb1"), "part of the ULA global ID survived: $out")
assertTrue(out.contains("fd"), "the result should still read as a ULA: $out")
}
// Pseudonymizing the prefix as a unit keeps the one fact that is diagnostically useful:
// whether two addresses sit on the same network.
@Test
fun addressesOnOneUlaSubnetStayRelated() {
val doc = json.parseToJsonElement(
"""{"run":{"id":"r"},"networks":[{"link":{"dns":{"servers":[
"fda1:3fb1:ff92:6696::1","fda1:3fb1:ff92:6696::2","fdff:9999:8888:7777::1"]}}}]}"""
).jsonObject
val servers = anon(PrivacyLevel.BALANCED, doc)["networks"]!!.jsonArray[0].jsonObject["link"]!!
.jsonObject["dns"]!!.jsonObject["servers"]!!.jsonArray.map { it.jsonPrimitive.content }
val prefixOf = { s: String -> s.substringBeforeLast("::") }
assertEquals(prefixOf(servers[0]), prefixOf(servers[1]),
"two addresses on one ULA subnet should share a pseudonymous prefix")
assertNotEquals(prefixOf(servers[0]), prefixOf(servers[2]),
"a different ULA network must not collide with the first")
}
// RFC1918 stays readable, and this is the contrast that justifies it: a shared, meaningless
// prefix is topology; a unique random one is identity.
@Test
fun rfc1918StaysReadableUnlikeUla() {
val doc = json.parseToJsonElement(
"""{"run":{"id":"r"},"networks":[{"link":{"dns":{"servers":["192.168.1.1","10.13.102.1"]}}}]}"""
).jsonObject
val out = flat(anon(PrivacyLevel.BALANCED, doc))
assertTrue(out.contains("192.168.1.1"), "RFC1918 should survive: $out")
assertTrue(out.contains("10.13.102.1"), "RFC1918 should survive: $out")
}
} }
@@ -0,0 +1,177 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.privacy
import kotlinx.serialization.json.Json
import kotlinx.serialization.json.JsonObject
import kotlinx.serialization.json.jsonArray
import kotlinx.serialization.json.jsonObject
import kotlinx.serialization.json.jsonPrimitive
import kotlin.test.Test
import kotlin.test.assertFalse
import kotlin.test.assertTrue
/**
* The blunt instrument: build a document with identifying values in every place one can actually
* occur, anonymize it, and assert none of them survive.
*
* [AnonymizerTest] checks that the fields the classification table knows about are handled
* correctly. This checks the other half the fields it does *not* know about. A per-field test
* can only fail for a field someone remembered to write a case for, which is exactly the wrong
* property for a privacy check: the dangerous field is the one nobody thought of.
*
* Concretely, this is written the way it is because the schema's own field names disagree with
* the classifier's. `Address.addr` carries an IP and is documented as such in
* measurement-schema.md §8, but the classifier keys on names like `ip4` and `gateway_ip4` and had
* never heard of `addr`.
*/
class LeakTest {
private val json = Json { prettyPrint = false }
private val salt = Salt.perRun(ByteArray(32) { 3 })
/**
* Every string here is something that identifies a person, a household or a device, placed
* where the real models actually put it (`core-measurement`'s Network/Link/Address/DnsConfig).
*/
private val secrets = listOf(
"Rambossek WLAN", // ssid
"78:9a:18:aa:bb:cc", // bssid
"aa:bb:cc:dd:ee:11", // gateway mac
"2001:1ad0:c4fe:6767::150", // global v6 address on the interface
"2a02:1748:dead:beef::1", // v6 default gateway
"203.0.113.77", // public v4
"nas.rambossek.lan", // private-dns hostname
"rambossek.lan", // search domain
"Anna's Chromecast", // neighbour name
"kitchen table", // free-text note
)
private fun document(): String = """
{
"schema": "echolot/measurement",
"run": {
"id": "run-1", "trigger": "manual", "notes": "${secrets[9]}",
"device": {"manufacturer": "OnePlus", "model": "CPH2747"}
},
"networks": [{
"id": "net-1", "transport": "wifi",
"link": {
"mtu": 1500,
"addresses": [
{"addr": "${secrets[3]}", "prefix_len": 64, "scope": "global"},
{"addr": "192.168.1.44", "prefix_len": 24, "scope": "global"}
],
"routes": [
{"dst": "::/0", "gateway": "${secrets[4]}", "iface": "wlan0"},
{"dst": "0.0.0.0/0", "gateway": "192.168.1.1", "iface": "wlan0"}
],
"dns": {
"servers": ["${secrets[5]}", "192.168.1.1"],
"private_dns_hostname": "${secrets[6]}",
"search_domains": ["${secrets[7]}"]
}
},
"wifi": {"ssid": "${secrets[0]}", "bssid": "${secrets[1]}"},
"neighbors": [{"name": "${secrets[8]}", "mac": "${secrets[2]}"}]
}],
"tests": [{"id": "t1", "type": "train.udp_updown", "status": "ok",
"metrics": {"rtt_ms_avg": 12.4}}],
"findings": [],
"summary": {"verdict": "ok"}
}
""".trimIndent()
private fun anonymized(level: PrivacyLevel): String =
json.encodeToString(
kotlinx.serialization.json.JsonObject.serializer(),
Anonymizer(level, salt).anonymize(json.parseToJsonElement(document()).jsonObject),
)
@Test
fun nothingIdentifyingSurvivesBalanced() {
val out = anonymized(PrivacyLevel.BALANCED)
val leaked = secrets.filter { out.contains(it) }
assertTrue(
leaked.isEmpty(),
"these identifying values were uploaded verbatim at BALANCED: $leaked\n\n$out",
)
}
@Test
fun nothingIdentifyingSurvivesStrict() {
val out = anonymized(PrivacyLevel.STRICT)
val leaked = secrets.filter { out.contains(it) }
assertTrue(leaked.isEmpty(), "leaked at STRICT: $leaked\n\n$out")
}
// Private addresses are kept on purpose — they describe the topology and not the person — so
// this pins that the leak test above is not passing by accident of over-redaction.
@Test
fun privateAddressesAreStillReadable() {
val out = anonymized(PrivacyLevel.BALANCED)
assertTrue(out.contains("192.168.1.1"), "RFC1918 gateway should survive: $out")
assertTrue(out.contains("192.168.1.44"), "RFC1918 interface address should survive: $out")
}
/**
* Raw shell output embeds a complete inventory of the local network, and neither the field-name
* table nor the whole-value shape check can see it: `ip_neigh` is one long string that is
* itself neither a MAC nor an address.
*
* This is not hypothetical. The blob below is (abridged) real output that reached the server
* at the `balanced` level from a test device, carrying the hardware address of every host on
* the network. measurement-schema.md §9 had flagged raw dumps as "hard to anonymize"; nothing
* enforced it.
*/
@Test
fun identifiersInsideRawShellOutputAreScrubbed() {
// Joined rather than written with escapes, so the fixture stays readable and there is no
// chance of an escape being mangled on its way into the JSON below.
val dump = listOf(
"uid=2000",
"10.13.102.5 dev wlan0 lladdr 90:09:d0:1a:83:e4 REACHABLE",
"10.13.102.1 dev wlan0 lladdr 78:9a:18:54:b8:f9 REACHABLE",
"10.13.102.111 dev wlan0 lladdr dc:a2:66:08:69:95 STALE",
"2001:4bb8:46a:e724:289d:87ff:feb6:ebd3 dev wlan0 lladdr b8:be:f4:bc:ca:cf STALE",
).joinToString(" | ")
val doc = json.parseToJsonElement(
"""{"run":{"id":"r"},"tests":[{"id":"t","type":"link.ip_monitor",
"evidence":{"ip_neigh":"$dump"}}]}"""
).jsonObject
val out = json.encodeToString(
kotlinx.serialization.json.JsonObject.serializer(),
Anonymizer(PrivacyLevel.BALANCED, salt).anonymize(doc),
)
for (mac in listOf("90:09:d0:1a:83:e4", "78:9a:18:54:b8:f9", "dc:a2:66:08:69:95", "b8:be:f4:bc:ca:cf")) {
assertFalse(out.contains(mac), "a neighbour's MAC survived inside the raw dump: $mac")
}
assertFalse(out.contains("2001:4bb8:46a:e724:289d:87ff:feb6:ebd3"),
"a global IPv6 survived inside the raw dump")
// Scrubbed, not dropped: the evidence must still be readable, or the raw dump stops being
// evidence at all. Structure, hostnames of the fields, and RFC1918 addresses stay.
assertTrue(out.contains("REACHABLE") && out.contains("STALE"), "the dump lost its structure")
assertTrue(out.contains("10.13.102.1"), "RFC1918 addresses should stay readable: $out")
assertTrue(out.contains("78:9a:18"), "the vendor prefix should survive for identification")
}
// A MAC in a raw dump and the same MAC in a parsed field must land on the same pseudonym, or
// the document stops being internally consistent and one device reads as two.
@Test
fun theSameIdentifierMatchesAcrossParsedAndRawFields() {
val doc = json.parseToJsonElement(
"""{"run":{"id":"r"},
"networks":[{"wifi":{"bssid":"78:9a:18:54:b8:f9"}}],
"tests":[{"id":"t","evidence":{"ip_neigh":"gw dev wlan0 lladdr 78:9a:18:54:b8:f9 REACHABLE"}}]}"""
).jsonObject
val out = Anonymizer(PrivacyLevel.BALANCED, salt).anonymize(doc)
val parsed = out["networks"]!!.jsonArray[0].jsonObject["wifi"]!!.jsonObject["bssid"]!!
.jsonPrimitive.content
val raw = json.encodeToString(kotlinx.serialization.json.JsonObject.serializer(), out)
assertTrue(raw.contains(parsed),
"the parsed BSSID pseudonym ($parsed) does not appear in the scrubbed dump")
}
}
@@ -0,0 +1,48 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.privacy
import kotlinx.serialization.json.Json
import kotlinx.serialization.json.JsonObject
import kotlinx.serialization.json.jsonObject
import java.io.File
import kotlin.test.Test
import kotlin.test.assertTrue
/**
* Runs the anonymizer over a real captured document when one is supplied via ECHOLOT_REAL_RUN,
* and reports every MAC and public address that survives.
*
* Fixtures only contain the identifiers somebody thought to put in them. A real run off a real
* phone contains whatever the probes actually produce which is how the raw-shell-output leak was
* found in the first place. Self-skips when no document is supplied, so nobody's network ends up
* committed to the repository.
*/
class RealDocumentTest {
@Test
fun noIdentifiersSurviveInARealDocument() {
val path = System.getenv("ECHOLOT_REAL_RUN")
if (path.isNullOrBlank() || !File(path).isFile) {
println("RealDocumentTest skipped (set ECHOLOT_REAL_RUN to a captured run)"); return
}
val json = Json { prettyPrint = false }
val doc = json.parseToJsonElement(File(path).readText()).jsonObject
val out = json.encodeToString(
JsonObject.serializer(),
Anonymizer(PrivacyLevel.BALANCED, Salt.perRun(ByteArray(32) { 5 })).anonymize(doc),
)
val macs = Regex("""\b[0-9a-fA-F]{2}(?::[0-9a-fA-F]{2}){5}\b""").findAll(out)
.map { it.value.lowercase() }
.filter { it != "00:00:00:00:00:00" }
.toSet()
val original = Regex("""\b[0-9a-fA-F]{2}(?::[0-9a-fA-F]{2}){5}\b""")
.findAll(File(path).readText()).map { it.value.lowercase() }.toSet()
val survived = macs intersect original
println("MACs in the original: ${original.size}; unchanged after anonymizing: ${survived.size}")
assertTrue(survived.isEmpty(), "these real MAC addresses survived anonymization: $survived")
}
}
+5
View File
@@ -28,4 +28,9 @@ dependencies {
implementation(project(":core-measurement")) implementation(project(":core-measurement"))
implementation(libs.kotlinx.serialization.json) implementation(libs.kotlinx.serialization.json)
implementation(libs.kotlinx.coroutines.android) implementation(libs.kotlinx.coroutines.android)
// JVM unit tests for the pure wire-format decoders (DiscoveryParsers.kt) against realistic and
// deliberately malformed payloads — no device, no Android runtime. Same setup as core-shizuku's
// dump-parser tests: JUnit 4, because that is what AGP's unit-test source set runs.
testImplementation(libs.kotlin.test.junit)
testImplementation(libs.junit4)
} }
@@ -0,0 +1,78 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.probe
import android.content.Context
import app.echo_lot.measurement.Test
import app.echo_lot.measurement.TestError
import app.echo_lot.measurement.TestStatus
import app.echo_lot.measurement.Tier
/**
* A measurement that watches, rather than one that asks the long-run counterpart to [Probe].
*
* The difference is not duration but what is observable at all. A [Probe] describes the network
* during its own two seconds, so anything intermittent is invisible to the whole battery unless it
* happens to coincide with a probe: a wifi link that drops for four seconds every two minutes is
* reported as perfectly healthy by every one-shot test, both before and after the gap. A collector
* starts at t=0, keeps sampling while the battery runs beside it and after it finishes, and yields
* one [Test] when the window closes.
*
* Contract, and all of it is load-bearing for a run that can be cancelled at any second:
* - [start] must return promptly, having launched whatever it needs on its own scope. The battery
* runs concurrently and must not wait for a listener.
* - [stop] must be callable after a failed [start], must never throw, and must return whatever was
* gathered so far. A cancelled long run still owes the user the two minutes it did watch.
* - Neither may throw to the caller; a collector that could not register its listener reports that
* as an `unsupported` Test, which is a result rather than an absence.
*/
interface Collector {
/** A TestType registry id — collectors do not get their own namespace. */
val type: String
val tier: Tier get() = Tier.APP
suspend fun start(ctx: Context, ids: ProbeIds)
suspend fun stop(): Test
}
/**
* Shared plumbing: the ids and the [TestBuilder] a collector needs in [Collector.stop], captured in
* [Collector.start] before anything that can fail.
*
* Assigned first thing on purpose. A collector whose registration throws still has to produce a
* Test saying so, and it cannot do that without a UUID source and a start timestamp so acquiring
* them is never allowed to be the step that failed.
*/
abstract class BaseCollector : Collector {
protected var ids: ProbeIds? = null
private set
private var builder: TestBuilder? = null
/** Call at the top of [Collector.start], before any platform call. */
protected fun begin(ids: ProbeIds, networkRef: String? = null) {
this.ids = ids
builder = TestBuilder(type, tier, ids, networkRef = networkRef)
}
/**
* Builds this collector's Test, or when [begin] never ran, so the collector was stopped
* without ever being started a `skipped` one saying exactly that. It reports rather than
* throws for the same reason probes do: the caller is assembling a document, and an exception
* there costs every other collector's data too.
*/
protected fun build(
status: TestStatus,
evidence: kotlinx.serialization.json.JsonObject? = null,
metrics: kotlinx.serialization.json.JsonObject? = null,
error: TestError? = null,
params: kotlinx.serialization.json.JsonObject? = null,
): Test = builder?.build(status, evidence, metrics, error, params)
?: Test(
id = "00000000-0000-7000-8000-000000000000", type = type, tier = tier,
startedMonoNs = 0, endedMonoNs = 0, status = TestStatus.SKIPPED,
error = TestError("not_started", "the collector was stopped before it was started"),
)
}
@@ -0,0 +1,74 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.probe
import android.content.Context
import android.net.ConnectivityManager
import android.net.NetworkCapabilities
import android.system.ErrnoException
import android.system.OsConstants
import app.echo_lot.measurement.Constraints
import app.echo_lot.measurement.Transport
import java.net.DatagramSocket
/**
* Detects what will prevent this run from measuring (measurement-schema.md §3 `constraints`),
* before any probe runs and independently of all of them.
*
* The known case: while a VPN holds the default route, Android refuses `Network.bindSocket()` on
* the underlying networks (EPERM) so apps cannot leak around the tunnel. Every per-network test
* then silently measures the tunnel or nothing, and the run comes out shaped exactly like a clean
* run of a healthy network. Detecting that here one throwaway bind per network is what lets
* the document say "these networks went unmeasured" instead of leaving the reader to infer it
* from a pattern of `attempted: false` scattered across the tests.
*/
object ConstraintDetector {
fun detect(ctx: Context, entries: List<NetworkInventory.Entry>): Constraints {
val cm = ctx.getSystemService(Context.CONNECTIVITY_SERVICE) as ConnectivityManager
// "A VPN holds the default route" is judged from the ACTIVE network, not from a VPN
// network merely existing in the list: a tunnel that was just disconnected lingers in
// allNetworks while it tears down, and counting it kept the app claiming "measured
// through a VPN" after the VPN was gone.
val vpnActive = runCatching {
cm.getNetworkCapabilities(cm.activeNetwork)
?.hasTransport(NetworkCapabilities.TRANSPORT_VPN) == true
}.getOrDefault(false)
val refused = ArrayList<String>()
for (e in entries) {
// The tunnel itself stays bindable — it is the underlying networks the OS walls off.
if (e.model.transport == Transport.VPN) continue
// Networks an app may never bind (carrier IMS/VoLTE, MMS, XCAP — no INTERNET, no
// NOT_RESTRICTED) refuse with EPERM permanently, VPN or no VPN. Counting them as a
// constraint is what made a phone with VoLTE report "measurement blocked" forever
// and forced every run INCONCLUSIVE. They are recorded in networks[] as
// app_usable:false; they are not something a run failed to do.
if (e.model.appUsable == false) continue
val err = try {
DatagramSocket().use { s -> e.handle.bindSocket(s) }
null
} catch (t: Throwable) {
t
}
// Only the OS *refusing* counts as blocked (EPERM: the VPN wall). A network that
// happens to die mid-snapshot fails its bind too, but with a different errno, and
// calling that "per-network probing blocked" would flip a whole healthy run to
// INCONCLUSIVE over one network going away — the probes already record
// attempted:false for that case.
if (err != null && isPermissionRefusal(err)) refused.add(e.model.id)
}
return Constraints(
vpnActive = vpnActive,
perNetworkBlocked = refused.isNotEmpty(),
unmeasuredNetworks = refused,
)
}
private fun isPermissionRefusal(t: Throwable): Boolean =
generateSequence(t) { it.cause }.any {
(it is ErrnoException && it.errno == OsConstants.EPERM) ||
(it.message?.contains("EPERM") == true)
}
}
@@ -0,0 +1,360 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.probe
/**
* Wire-format decoders for the passive discovery collectors (SSDP, LLMNR, NetBIOS-NS, WS-Discovery).
*
* No Android imports on purpose: every byte these see was broadcast by an unidentified device on
* someone else's network, so they are the part of the collectors that most needs to be exercised
* against malformed input and that is only cheap to do if it runs on a plain JVM. See
* DiscoveryParsersTest.
*
* The universal contract here is **return null, never throw**. A collector that dies on one
* malformed datagram loses the whole window's inventory, and a device that emits garbage is a
* device we still want counted. Every decoder therefore bounds-checks by hand rather than relying
* on an exception, and treats "this is not the protocol I parse" and "this is the protocol but it
* is broken" as the same answer: nothing to record.
*/
// ---- shared -------------------------------------------------------------------------------
/**
* Makes a decoded string safe to embed in a JSON document.
*
* Names arrive as arbitrary bytes. Control characters would survive JSON encoding as escapes and
* turn up in a terminal that interprets them, and an unbounded length is a memory cost decided by
* whoever is shouting on the segment so both are capped here rather than at each call site.
*/
internal fun sanitizeText(s: String, max: Int = 255): String {
val sb = StringBuilder(minOf(s.length, max))
for (c in s) {
if (sb.length >= max) break
sb.append(if (c.isISOControl() || c == '') '?' else c)
}
return sb.toString()
}
private fun u8(b: ByteArray, i: Int) = b[i].toInt() and 0xFF
private fun u16(b: ByteArray, i: Int) = (u8(b, i) shl 8) or u8(b, i + 1)
// ---- SSDP ---------------------------------------------------------------------------------
enum class SsdpKind { ALIVE, BYEBYE, UPDATE, RESPONSE, SEARCH, OTHER }
/**
* One SSDP message. [target] is NT for announcements and ST for search responses they name the
* same thing (which service is being talked about) from the two sides of the conversation, so they
* collapse into one field.
*/
data class SsdpMessage(
val kind: SsdpKind,
val target: String?,
val usn: String?,
val serverBanner: String?,
val location: String?,
)
object SsdpParser {
/**
* SSDP is HTTP-shaped but is not HTTP: there is no framing, no content length, and vendors
* disagree about line endings. Parsing it as "a start line plus colon-separated headers, be
* liberal about the rest" is the whole job — running it through an HTTP client would reject
* messages that real devices send and that we want to count.
*
* ISO-8859-1 decoding because the headers are byte-oriented and this mapping is total: no byte
* sequence can fail to decode, so a device with a Latin-1 model name in its SERVER banner is
* recorded rather than replaced by question marks.
*/
fun parse(bytes: ByteArray, len: Int): SsdpMessage? {
if (len <= 0 || len > bytes.size) return null
val text = String(bytes, 0, len, Charsets.ISO_8859_1)
val lines = text.split('\n')
val start = lines.firstOrNull()?.trim().orEmpty()
if (start.isEmpty()) return null
val headers = HashMap<String, String>()
for (i in 1 until lines.size) {
val line = lines[i].trimEnd('\r')
if (line.isBlank()) break
val c = line.indexOf(':')
if (c <= 0) continue
val key = line.substring(0, c).trim().uppercase()
// First occurrence wins: a duplicated header is a device bug, and taking the first
// matches what every SSDP implementation in the wild does.
if (key !in headers) headers[key] = sanitizeText(line.substring(c + 1).trim(), 512)
}
val kind = when {
start.startsWith("NOTIFY", ignoreCase = true) -> when (headers["NTS"]?.lowercase()) {
"ssdp:alive" -> SsdpKind.ALIVE
"ssdp:byebye" -> SsdpKind.BYEBYE
"ssdp:update" -> SsdpKind.UPDATE
else -> SsdpKind.OTHER
}
start.startsWith("M-SEARCH", ignoreCase = true) -> SsdpKind.SEARCH
start.startsWith("HTTP/", ignoreCase = true) -> SsdpKind.RESPONSE
else -> return null // not SSDP at all
}
return SsdpMessage(
kind = kind,
target = headers["NT"] ?: headers["ST"],
usn = headers["USN"],
serverBanner = headers["SERVER"],
location = headers["LOCATION"],
)
}
/**
* The make/model inside a SERVER banner, as a hint.
*
* A banner reads `Linux/4.4 UPnP/1.0 Synology/DSM-7.3` or `FRITZ!Box 7590 UPnP/1.0`: the
* interesting token is whichever one is not the OS and not the protocol version, and there is
* no grammar that says which. Dropping the known-boilerplate tokens and keeping the rest is
* therefore a heuristic and is reported as such [SsdpMessage.serverBanner] is kept verbatim
* beside it so nobody has to trust this to read the evidence.
*/
fun productHint(serverBanner: String?): String? {
val banner = serverBanner?.trim().orEmpty()
if (banner.isEmpty()) return null
val kept = banner.split(' ', '\t')
.map { it.trim() }
.filter { it.isNotEmpty() && it.substringBefore('/').lowercase() !in BOILERPLATE }
.distinct()
return kept.joinToString(" ").takeIf { it.isNotEmpty() }
}
private val BOILERPLATE = setOf(
"upnp", "http", "dlnadoc", "linux", "unix", "windows", "darwin", "posix", "sdk",
"upnp-device-host", "microsoft-windows", "mono.upnp", "webos", "android",
)
}
// ---- DNS-format questions (LLMNR, and the shape NetBIOS-NS borrows) -------------------------
/** One question from a DNS-format packet. [isQuery] separates "who is asking" from "who answered". */
data class DnsQuestion(
val id: Int,
val name: String,
val qtype: Int,
val qclass: Int,
val isQuery: Boolean,
val opcode: Int,
)
object LlmnrParser {
/**
* Decodes the first question of an LLMNR packet, which is DNS wire format with a different
* transport.
*
* Name compression is rejected rather than followed. RFC 4795 forbids it in LLMNR, so a pointer
* here is either a broken sender or someone hoping the parser will chase it and a pointer
* loop is the classic way to hang a DNS decoder. Refusing costs nothing real and removes the
* only unbounded loop this decoder could have had.
*/
fun parse(bytes: ByteArray, len: Int): DnsQuestion? {
if (len < HEADER + 5 || len > bytes.size) return null
val flags = u16(bytes, 2)
if (u16(bytes, 4) < 1) return null // no question section
val sb = StringBuilder()
var i = HEADER
var labels = 0
while (true) {
if (i >= len) return null
val l = u8(bytes, i)
if (l == 0) { i++; break }
if (l and 0xC0 != 0) return null // compression pointer / reserved length
i++
if (i + l > len) return null
if (++labels > MAX_LABELS || sb.length + l > MAX_NAME) return null
if (sb.isNotEmpty()) sb.append('.')
sb.append(String(bytes, i, l, Charsets.UTF_8))
i += l
}
if (sb.isEmpty()) return null
if (i + 4 > len) return null
return DnsQuestion(
id = u16(bytes, 0),
name = sanitizeText(sb.toString()),
qtype = u16(bytes, i),
qclass = u16(bytes, i + 2),
isQuery = (flags and 0x8000) == 0,
opcode = (flags shr 11) and 0x0F,
)
}
/** The record types worth naming in evidence; anything else is reported as its number. */
fun qtypeName(qtype: Int): String = when (qtype) {
1 -> "A"
28 -> "AAAA"
12 -> "PTR"
33 -> "SRV"
255 -> "ANY"
else -> qtype.toString()
}
private const val HEADER = 12
private const val MAX_LABELS = 64
private const val MAX_NAME = 255
}
// ---- NetBIOS name service ------------------------------------------------------------------
/**
* A decoded NetBIOS name-service question.
*
* [suffix] is the sixteenth byte of the name and is the part an engineer reads first: it says what
* the announcement is *for* (a workstation, a file server, a browser election) rather than who is
* making it.
*/
data class NetbiosName(
val name: String,
val suffix: Int,
val role: String,
val isResponse: Boolean,
val opcode: Int,
)
object NetbiosParser {
/**
* Decodes the question name from an NBNS packet (RFC 1002 §4.2).
*
* The header is DNS-shaped, but the name is not: NetBIOS first-level encoding splits each of
* the 16 name bytes into two nibbles and adds 'A' to each, so a 16-byte name is always exactly
* 32 characters drawn from A-P. That fixed shape is also the validity check anything outside
* A-P means this is not an NBNS name, and it is cheaper and safer to reject the packet than to
* guess at what a half-decodable name was supposed to say.
*/
fun parse(bytes: ByteArray, len: Int): NetbiosName? {
if (len < HEADER + 34 || len > bytes.size) return null
if (u16(bytes, 4) < 1) return null // no question section
if (u8(bytes, HEADER) != ENCODED_LEN) return null
val raw = ByteArray(16)
for (j in 0 until 16) {
val hi = u8(bytes, HEADER + 1 + j * 2) - 'A'.code
val lo = u8(bytes, HEADER + 2 + j * 2) - 'A'.code
if (hi !in 0..15 || lo !in 0..15) return null
raw[j] = ((hi shl 4) or lo).toByte()
}
// The first 15 bytes are the name, space-padded; the 16th is the suffix.
val padded = String(raw, 0, 15, Charsets.ISO_8859_1)
val name = sanitizeText(padded.trimEnd { it == ' ' || it.isISOControl() }, 15)
if (name.isEmpty()) return null
val suffix = raw[15].toInt() and 0xFF
val flags = u16(bytes, 2)
return NetbiosName(
name = name,
suffix = suffix,
role = roleOf(suffix),
isResponse = (flags and 0x8000) != 0,
opcode = (flags shr 11) and 0x0F,
)
}
/** RFC 1001 §15 / the Microsoft suffix assignments people actually see on a LAN. */
fun roleOf(suffix: Int): String = when (suffix) {
0x00 -> "workstation"
0x03 -> "messenger"
0x1B -> "domain_master_browser"
0x1C -> "domain_controllers"
0x1D -> "master_browser"
0x1E -> "browser_elections"
0x20 -> "file_server"
else -> "suffix_0x%02X".format(suffix)
}
/** NBNS opcodes: what the sender is doing, not just that it is talking. */
fun opcodeName(opcode: Int): String = when (opcode) {
0 -> "query"
5 -> "registration"
6 -> "release"
7 -> "wack"
8 -> "refresh"
else -> "opcode_$opcode"
}
private const val HEADER = 12
private const val ENCODED_LEN = 32
}
// ---- WS-Discovery --------------------------------------------------------------------------
/** The four things a WS-Discovery datagram is worth reading for. */
data class WsdMessage(
val action: String?,
val deviceUuid: String?,
val types: String?,
val xaddrs: String?,
)
object WsdParser {
/**
* Pulls four leaf values out of SOAP-over-UDP by targeted matching, deliberately **without an
* XML parser**.
*
* The input is unauthenticated, unsolicited, and written by whatever is on the segment. Handing
* that to a real XML parser buys namespace correctness and pays for it with the whole XML
* attack surface entity expansion (a 700-byte datagram that allocates gigabytes), DTDs that
* fetch external resources, and nesting deep enough to exhaust the stack all inside a probe
* whose contract is that it never throws. None of that surface is needed to read four leaf
* elements out of a message we are only ever going to count and quote.
*
* So: cap the text, then match `<[prefix:]Tag ...>value<` with a character class that cannot
* backtrack. The worst case is a value we fail to extract, which is recorded as an absence.
*/
fun parse(bytes: ByteArray, len: Int): WsdMessage? {
if (len <= 0 || len > bytes.size) return null
val text = String(bytes, 0, minOf(len, MAX_TEXT), Charsets.UTF_8)
if (!text.contains("Envelope", ignoreCase = true)) return null
// The action tail is the message type — Hello, Bye, Probe, ProbeMatches, ResolveMatches.
val action = leaf(text, "Action")?.substringAfterLast('/')?.takeIf { it.isNotEmpty() }
// The device's stable identity is an EndpointReference/Address holding a urn:uuid. Other
// Address elements exist (wsa:To names the discovery group), so the uuid shape picks the
// right one rather than the first one.
val uuid = leaves(text, "Address").firstOrNull { it.contains("uuid:", ignoreCase = true) }
val msg = WsdMessage(
action = action?.let { sanitizeText(it, 64) },
deviceUuid = uuid?.let { sanitizeText(it, 128) },
types = leaf(text, "Types")?.let { sanitizeText(it, 200) },
xaddrs = leaf(text, "XAddrs")?.let { sanitizeText(it, 300) },
)
val empty = msg.action == null && msg.deviceUuid == null &&
msg.types == null && msg.xaddrs == null
return if (empty) null else msg
}
private fun leaf(xml: String, tag: String): String? = leaves(xml, tag).firstOrNull()
private fun leaves(xml: String, tag: String): List<String> {
val re = TAGS[tag] ?: return emptyList()
return re.findAll(xml).map { it.groupValues[1].trim() }.filter { it.isNotEmpty() }.toList()
}
/**
* `[^<]{0,512}` rather than a lazy `.*?`: it can only ever match forward, so there is no input
* that makes this regex expensive. WS-Discovery leaves hold no markup, so nothing is lost.
*
* Built once, eagerly, because the alternative is a memoizing map touched from the capture
* thread a data race for the sake of four Regex allocations.
*/
private val TAGS: Map<String, Regex> =
listOf("Action", "Address", "Types", "XAddrs").associateWith { tag ->
Regex(
"""<(?:[A-Za-z0-9_.\-]{1,32}:)?$tag\b[^>]{0,256}>([^<]{0,512})<""",
RegexOption.IGNORE_CASE,
)
}
private const val MAX_TEXT = 16 * 1024
}
@@ -0,0 +1,227 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.probe
import android.content.Context
import app.echo_lot.measurement.Test
import app.echo_lot.measurement.TestStatus
import app.echo_lot.measurement.TestType
import app.echo_lot.measurement.Tier
import kotlinx.coroutines.Dispatchers
import kotlinx.coroutines.withContext
import kotlinx.serialization.json.JsonObject
import kotlinx.serialization.json.buildJsonObject
import kotlinx.serialization.json.put
import java.net.DatagramPacket
import java.net.DatagramSocket
import java.net.InetAddress
import java.net.InetSocketAddress
import java.util.Random
/**
* Asks the network's own DNS servers directly, then asks Android to resolve the same name, and
* compares.
*
* The comparison is the point. A name that fails to resolve looks the same to a user whatever the
* cause, but the causes want opposite responses: if the server does not answer, the network is
* broken and the router is the thing to look at; if the server answers a raw query while the
* platform still cannot resolve, the device's own resolver has wedged and toggling wifi fixes it in
* seconds. Nothing else on a phone will tell you which of those you have.
*
* This is deliberately not a general DNS test no recursion checks, no DNSSEC, no rewriting
* detection; [DnsCanaryProbe] covers interception. This one answers a single question: is the
* resolver on this device doing its job.
*/
class DnsResolverProbe(
private val entries: List<NetworkInventory.Entry>,
/** Resolved directly rather than through any cache; any name with a stable answer will do. */
private val probeName: String = "one.one.one.one",
) : Probe {
override val type = TestType.DNS_RESOLVER
override val tier = Tier.APP
// A query per server with a 3s ceiling, plus one getaddrinfo that may sit out its own timeout.
override val estimatedMs = 6_000L
override suspend fun run(ctx: Context, ids: ProbeIds): Test = withContext(Dispatchers.IO) {
val b = TestBuilder(type, tier, ids)
val perNetwork = LinkedHashMap<String, Pair<String, JsonObject>>()
for (e in entries) {
val servers = e.model.link.dns?.servers.orEmpty()
if (servers.isEmpty()) continue
val label = "${e.model.transport.name.lowercase()}:${e.model.id}"
// Directly: does the configured server answer at all?
var direct: Boolean? = null
var directDetail = "no server answered"
for (s in servers) {
val r = try {
queryDirect(s, probeName)
} catch (e: DnsRefused) {
// Distinguished deliberately: a server that replies with a failure is a
// working server saying no, which points at the network rather than here.
direct = false
directDetail = "$s ${e.why}"
break
}
if (r != null) {
direct = true
directDetail = "$s answered in ${r}ms"
break
}
direct = false
directDetail = "$s did not answer"
}
// The search domains the network handed out, asked about separately.
//
// A resolver appends these to a lookup, so a search domain the server will not answer
// for stalls every name a client asks about — and it fails as silence, which is
// indistinguishable from packet loss, so clients retry rather than moving on. Asking
// about a name that cannot exist is deliberate: the answer wanted here is NXDOMAIN,
// and what matters is only whether anything comes back at all.
val searchDomains = e.model.link.dns?.searchDomains.orEmpty()
var searchAnswered: Boolean? = null
var searchDetail = ""
for (d in searchDomains) {
val nonce = "echolot-probe-" + java.util.UUID.randomUUID().toString().take(8)
val answered = servers.any { srv ->
runCatching { queryDirect(srv, "$nonce.$d") != null }
.getOrElse { it is DnsRefused } // a refusal is still an answer
}
if (!answered) {
searchAnswered = false
searchDetail = "$d is not answered at all — queries under it vanish"
break
}
searchAnswered = true
searchDetail = "$d answers"
}
// Through the platform: what an app actually gets.
val viaSystem = runCatching {
e.handle.getAllByName(probeName).isNotEmpty()
}.getOrElse { false }
perNetwork[label] = e.model.id to buildJsonObject {
put("network_ref", e.model.id)
put("servers", servers.joinToString(","))
direct?.let { put("direct_answer", it) }
put("direct_detail", directDetail)
if (searchDomains.isNotEmpty()) {
put("search_domains", searchDomains.joinToString(","))
searchAnswered?.let { put("search_answered", it) }
put("search_detail", searchDetail)
}
put("system_resolves", viaSystem)
// Named here rather than left for a finding to infer, because the pairing is the
// whole observation and splitting it across two places invites reading one alone.
put(
"verdict",
when {
// Ordered by which component is at fault, most specific first. A search
// domain that swallows queries explains a failure that would otherwise be
// blamed on the device, so it has to be tested before that conclusion.
searchAnswered == false -> "search domain swallows queries"
viaSystem -> "resolver working"
direct == true -> "server answers, device resolver does not"
direct == false -> "server does not answer"
else -> "not determined"
},
)
}
}
if (perNetwork.isEmpty()) {
return@withContext b.build(
TestStatus.SKIPPED,
evidence = buildJsonObject { put("reason", "no network advertised a DNS server") },
)
}
val evidence = buildJsonObject {
put("name", probeName)
for ((label, v) in perNetwork) put(label, v.second)
}
// OK means the measurement ran, not that DNS is healthy — the finding says that.
b.build(TestStatus.OK, evidence = evidence)
}
/**
* Sends one A query straight to [server] over UDP. Returns the round trip in ms, or null.
*
* Hand-rolled rather than via any resolver API on purpose: the entire point is to bypass the
* component under suspicion. Anything that goes through the platform resolver would inherit
* exactly the fault this is trying to detect.
*/
private fun queryDirect(server: String, name: String): Long? {
return try {
queryDirectOrThrow(server, name)
} catch (e: DnsRefused) {
throw e
} catch (t: Throwable) {
null
}
}
private fun queryDirectOrThrow(server: String, name: String): Long? = run {
val id = Random().nextInt(0xFFFF)
val query = buildQuery(id, name)
DatagramSocket().use { sock ->
sock.soTimeout = 3000
val addr = InetAddress.getByName(server) // a literal from DHCP; no lookup happens
val t0 = System.nanoTime()
sock.send(DatagramPacket(query, query.size, InetSocketAddress(addr, 53)))
val buf = ByteArray(512)
val reply = DatagramPacket(buf, buf.size)
sock.receive(reply)
val ms = (System.nanoTime() - t0) / 1_000_000
// A reply is not an answer. Counting any packet as success would let a REFUSED or
// SERVFAIL — both perfectly well-formed responses — be reported as "the server
// answers", and this probe's whole output is the claim that the server is fine and
// the device is not. That would be an accusation pointed at the wrong component,
// stated with confidence.
val replyId = ((buf[0].toInt() and 0xFF) shl 8) or (buf[1].toInt() and 0xFF)
val rcode = if (reply.length >= 4) buf[3].toInt() and 0x0F else -1
val answers = if (reply.length >= 8) {
((buf[6].toInt() and 0xFF) shl 8) or (buf[7].toInt() and 0xFF)
} else 0
when {
replyId != id || reply.length < 12 -> null
rcode != 0 -> throw DnsRefused(rcodeName(rcode))
answers == 0 -> throw DnsRefused("answered with no records")
else -> ms
}
}
}
/** The server replied, but with a failure — which is a network fault, not a device one. */
private class DnsRefused(val why: String) : Exception(why)
private fun rcodeName(rcode: Int): String = when (rcode) {
1 -> "rejected the query as malformed"
2 -> "reported its own failure (SERVFAIL)"
3 -> "said the name does not exist (NXDOMAIN)"
4 -> "does not implement this query"
5 -> "refused the query (REFUSED)"
else -> "returned rcode $rcode"
}
/** A minimal DNS query: one question, class IN, type A, recursion desired. */
private fun buildQuery(id: Int, name: String): ByteArray {
val labels = name.split('.').filter { it.isNotEmpty() }
val out = ArrayList<Byte>(32)
out.add((id shr 8).toByte()); out.add(id.toByte())
out.add(0x01); out.add(0x00) // recursion desired
out.add(0x00); out.add(0x01) // one question
repeat(6) { out.add(0x00) } // no answers, authority or additional
for (l in labels) {
out.add(l.length.toByte())
for (c in l.toByteArray(Charsets.US_ASCII)) out.add(c)
}
out.add(0x00) // root label
out.add(0x00); out.add(0x01) // type A
out.add(0x00); out.add(0x01) // class IN
return out.toByteArray()
}
}
@@ -0,0 +1,113 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.probe
import android.net.Network
import android.system.Os
import android.system.OsConstants
import android.system.StructTimeval
import java.io.FileDescriptor
import java.net.InetAddress
import java.nio.ByteBuffer
import java.util.Locale
/**
* One echo exchange over the unprivileged ICMP datagram socket.
*
* Extracted from [IcmpProbe] when the long-run [PingSeriesCollector] needed the same exchange a
* few hundred times instead of once. The two differ only in how often they call this; a second
* copy of the checksum and the sent/not-sent bookkeeping would only be a second place for them to
* drift apart.
*
* Works without root because Android ships an open `ping_group_range` validated on hardware by
* the prober, and the reason this probe family exists at app tier at all.
*/
internal object IcmpEcho {
/**
* One attempt's result.
*
* [attempted] separates "we sent an echo request and heard nothing" from "we never got as far
* as sending one". Both leave [ok] false, and collapsing them is how a probe ends up asserting
* something about a network it never touched: binding to a non-default network can fail with
* EPERM, and reporting that as ICMP silence blames the network for the app's own inability to
* use the interface. For a series it is also the difference between a lost packet and a socket
* that was never usable one is loss, the other is not.
*/
data class Result(
val ok: Boolean,
val attempted: Boolean,
val detail: String,
val rttMs: Double?,
)
fun ping(
network: Network?,
target: String,
v6: Boolean,
timeoutMs: Int,
seq: Int = 1,
): Result {
var fd: FileDescriptor? = null
var sent = false
return try {
val proto = if (v6) OsConstants.IPPROTO_ICMPV6 else OsConstants.IPPROTO_ICMP
val family = if (v6) OsConstants.AF_INET6 else OsConstants.AF_INET
fd = Os.socket(family, OsConstants.SOCK_DGRAM, proto)
// Everything up to and including sendto is setup. A failure here means the test did
// not run on this network — not that the network stayed silent.
network?.bindSocket(fd)
Os.setsockoptTimeval(
fd, OsConstants.SOL_SOCKET, OsConstants.SO_RCVTIMEO,
StructTimeval.fromMillis(timeoutMs.toLong()),
)
val addr = network?.getByName(target) ?: InetAddress.getByName(target)
val ident = (Os.getpid() and 0xFFFF)
val packet = buildEchoRequest(v6, ident.toShort(), seq.toShort())
val t0 = System.nanoTime()
Os.sendto(fd, packet, 0, packet.size, 0, addr, 0)
sent = true
val buf = ByteBuffer.allocate(1500)
val received = Os.recvfrom(fd, buf, 0, null)
val rttMs = (System.nanoTime() - t0) / 1_000_000.0
val replyType = if (received > 0) buf.get(0).toInt() and 0xFF else -1
val ok = replyType == (if (v6) 129 else 0)
Result(
ok, true,
"reply type=$replyType rtt_ms=${"%.1f".format(Locale.ROOT, rttMs)} bytes=$received",
if (ok) rttMs else null,
)
} catch (e: Throwable) {
// A timeout after a successful send is a real "no reply"; anything before it is not.
Result(false, sent, "error: ${e.message ?: e.javaClass.simpleName}", null)
} finally {
fd?.let { runCatching { Os.close(it) } }
}
}
private fun buildEchoRequest(v6: Boolean, ident: Short, seq: Short): ByteArray {
val type = if (v6) 128 else 8
val payload = "echolot".toByteArray()
val pkt = ByteBuffer.allocate(8 + payload.size)
pkt.put(type.toByte()); pkt.put(0); pkt.putShort(0)
pkt.putShort(ident); pkt.putShort(seq); pkt.put(payload)
val bytes = pkt.array()
// The v6 checksum is computed by the kernel over a pseudo-header the socket owns; filling
// it in here would be wrong, not merely redundant.
if (!v6) {
val cs = checksum(bytes)
bytes[2] = (cs.toInt() shr 8).toByte(); bytes[3] = (cs.toInt() and 0xFF).toByte()
}
return bytes
}
private fun checksum(b: ByteArray): Short {
var sum = 0; var i = 0
while (i < b.size - 1) { sum += ((b[i].toInt() and 0xFF) shl 8) or (b[i + 1].toInt() and 0xFF); i += 2 }
if (i < b.size) sum += (b[i].toInt() and 0xFF) shl 8
while (sum shr 16 != 0) sum = (sum and 0xFFFF) + (sum shr 16)
return sum.inv().toShort()
}
}
@@ -5,9 +5,6 @@ package app.echo_lot.probe
import android.content.Context import android.content.Context
import android.net.Network import android.net.Network
import android.system.Os
import android.system.OsConstants
import android.system.StructTimeval
import app.echo_lot.measurement.Test import app.echo_lot.measurement.Test
import app.echo_lot.measurement.TestStatus import app.echo_lot.measurement.TestStatus
import app.echo_lot.measurement.TestType import app.echo_lot.measurement.TestType
@@ -17,10 +14,6 @@ import kotlinx.coroutines.withContext
import kotlinx.serialization.json.JsonObject import kotlinx.serialization.json.JsonObject
import kotlinx.serialization.json.buildJsonObject import kotlinx.serialization.json.buildJsonObject
import kotlinx.serialization.json.put import kotlinx.serialization.json.put
import java.io.FileDescriptor
import java.net.InetAddress
import java.nio.ByteBuffer
import java.util.Locale
/** /**
* icmp.ping4 / icmp.ping6 via the unprivileged ICMP datagram socket, per active network * icmp.ping4 / icmp.ping6 via the unprivileged ICMP datagram socket, per active network
@@ -40,24 +33,37 @@ class IcmpProbe(
override suspend fun run(ctx: Context, ids: ProbeIds): Test = withContext(Dispatchers.IO) { override suspend fun run(ctx: Context, ids: ProbeIds): Test = withContext(Dispatchers.IO) {
val b = TestBuilder(type, tier, ids) val b = TestBuilder(type, tier, ids)
val perNetwork = LinkedHashMap<String, String>() val perNetwork = LinkedHashMap<String, Pair<String?, IcmpEcho.Result>>()
var anyOk = false var anyOk = false
val rtts = ArrayList<Double>() val rtts = ArrayList<Double>()
// Default network first, then each active network explicitly. // Default network first, then each active network explicitly.
attempt(null).let { (ok, detail, rtt) -> attempt(null).let { a ->
perNetwork["default"] = detail; if (ok) { anyOk = true; rtt?.let(rtts::add) } perNetwork["default"] = null to a
if (a.ok) { anyOk = true; a.rttMs?.let(rtts::add) }
} }
for (e in entries) { for (e in entries) {
val label = "${e.model.transport.name.lowercase()}:${e.model.id}" val label = "${e.model.transport.name.lowercase()}:${e.model.id}"
val (ok, detail, rtt) = attempt(e.handle) val a = attempt(e.handle)
perNetwork[label] = detail perNetwork[label] = e.model.id to a
if (ok) { anyOk = true; rtt?.let(rtts::add) } if (a.ok) { anyOk = true; a.rttMs?.let(rtts::add) }
} }
// Per-network results are recorded structurally, not just as prose. The aggregate status
// can only say "some network answered"; a finding needs to know *which* network failed,
// and recovering that by parsing a human-readable detail string would be a trap waiting to
// spring the first time the wording changes.
val evidence: JsonObject = buildJsonObject { val evidence: JsonObject = buildJsonObject {
put("target", target) put("target", target)
for ((k, v) in perNetwork) put(k, v) for ((label, r) in perNetwork) {
val (netId, a) = r
put(label, buildJsonObject {
netId?.let { put("network_ref", it) }
put("ok", a.ok)
put("attempted", a.attempted)
put("detail", a.detail)
})
}
} }
val metrics: JsonObject = buildJsonObject { val metrics: JsonObject = buildJsonObject {
put("networks_ok", rtts.size) put("networks_ok", rtts.size)
@@ -69,56 +75,9 @@ class IcmpProbe(
b.build(status, evidence = evidence, metrics = metrics) b.build(status, evidence = evidence, metrics = metrics)
} }
private data class Attempt(val ok: Boolean, val detail: String, val rttMs: Double?) /** The echo exchange itself lives in [IcmpEcho], shared with the long-run ping series. */
private fun attempt(network: Network?): IcmpEcho.Result =
private fun attempt(network: Network?): Attempt { IcmpEcho.ping(network, target, v6, timeoutMs = 3000)
var fd: FileDescriptor? = null
return try {
val proto = if (v6) OsConstants.IPPROTO_ICMPV6 else OsConstants.IPPROTO_ICMP
val family = if (v6) OsConstants.AF_INET6 else OsConstants.AF_INET
fd = Os.socket(family, OsConstants.SOCK_DGRAM, proto)
network?.bindSocket(fd)
Os.setsockoptTimeval(fd, OsConstants.SOL_SOCKET, OsConstants.SO_RCVTIMEO, StructTimeval.fromMillis(3000))
val addr = network?.getByName(target) ?: InetAddress.getByName(target)
val ident = (Os.getpid() and 0xFFFF)
val packet = buildEchoRequest(v6, ident.toShort(), 1)
val t0 = System.nanoTime()
Os.sendto(fd, packet, 0, packet.size, 0, addr, 0)
val buf = ByteBuffer.allocate(1500)
val received = Os.recvfrom(fd, buf, 0, null)
val rttMs = (System.nanoTime() - t0) / 1_000_000.0
val replyType = if (received > 0) buf.get(0).toInt() and 0xFF else -1
val ok = replyType == (if (v6) 129 else 0)
Attempt(ok, "reply type=$replyType rtt_ms=${"%.1f".format(Locale.ROOT, rttMs)} bytes=$received", if (ok) rttMs else null)
} catch (e: Throwable) {
Attempt(false, "error: ${e.message ?: e.javaClass.simpleName}", null)
} finally {
fd?.let { runCatching { Os.close(it) } }
}
}
private fun buildEchoRequest(v6: Boolean, ident: Short, seq: Short): ByteArray {
val type = if (v6) 128 else 8
val payload = "echolot".toByteArray()
val pkt = ByteBuffer.allocate(8 + payload.size)
pkt.put(type.toByte()); pkt.put(0); pkt.putShort(0)
pkt.putShort(ident); pkt.putShort(seq); pkt.put(payload)
val bytes = pkt.array()
if (!v6) {
val cs = checksum(bytes)
bytes[2] = (cs.toInt() shr 8).toByte(); bytes[3] = (cs.toInt() and 0xFF).toByte()
}
return bytes
}
private fun checksum(b: ByteArray): Short {
var sum = 0; var i = 0
while (i < b.size - 1) { sum += ((b[i].toInt() and 0xFF) shl 8) or (b[i + 1].toInt() and 0xFF); i += 2 }
if (i < b.size) sum += (b[i].toInt() and 0xFF) shl 8
while (sum shr 16 != 0) sum = (sum and 0xFFFF) + (sum shr 16)
return sum.inv().toShort()
}
private companion object { private companion object {
fun round1(v: Double) = Math.round(v * 10.0) / 10.0 fun round1(v: Double) = Math.round(v * 10.0) / 10.0
@@ -0,0 +1,124 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.probe
import android.content.Context
import app.echo_lot.measurement.Test
import app.echo_lot.measurement.TestError
import app.echo_lot.measurement.TestType
import app.echo_lot.measurement.Tier
import kotlinx.serialization.json.JsonObject
import kotlinx.serialization.json.buildJsonObject
import kotlinx.serialization.json.put
/**
* local.llmnr_inventory who is resolving names with LLMNR on this segment, and what for.
*
* The measurement is twofold, and the second half is the one people underestimate:
*
* 1. **Which hostnames are being looked up.** An LLMNR query is a device saying out loud, to
* everyone, the name of something it wants to reach. Over a window that is a map of who talks
* to whom and, when the names are things like `wpad` or a server that no longer exists, a map
* of what is failing to resolve through DNS and falling back.
* 2. **That LLMNR is in use at all.** LLMNR (and its NetBIOS sibling) is a name-resolution
* fallback that trusts whoever answers first, which is the mechanism behind the standard
* credential-relay attack on Windows networks. Its mere presence on a segment is a finding
* independent of any individual query which is why this earns a test id rather than folding
* into the mDNS inventory.
*
* Purely passive: queries are broadcast to the group, so listening is the entire measurement, and
* answering or querying would make this device a participant in exactly the trust relationship the
* measurement is about.
*/
class LlmnrCollector : BaseCollector() {
override val type = TestType.LOCAL_LLMNR_INVENTORY
override val tier = Tier.APP
private val capture = MulticastCapture(
label = "llmnr",
port = PORT,
group4 = GROUP4,
// ff02::1:3 is the IPv6 LLMNR group. Joined where IPv6 exists; a v4-only network simply
// records an empty joined_v6 rather than an error, because that is not one.
group6 = GROUP6,
)
private var startedAtMonoNs = 0L
override suspend fun start(ctx: Context, ids: ProbeIds) {
begin(ids)
startedAtMonoNs = ids.monoNs()
capture.acquireLock(ctx)
capture.start(ids)
}
override suspend fun stop(): Test {
val packets = capture.stop()
val table = DiscoveryTable()
var queries = 0
var responses = 0
var undecodable = 0
for (p in packets) {
val q = LlmnrParser.parse(p.data, p.data.size)
if (q == null) {
undecodable++
continue
}
if (q.isQuery) queries++ else responses++
// Responses are normally unicast back to the querier, so what lands here is
// overwhelmingly queries — but a response that does reach the group is still a device
// claiming a name, which is worth the same row.
table.observe(p.sourceIp, "${q.name}/${q.qtype}", p.atMonoNs) {
buildJsonObject {
put("query_name", q.name)
put("qtype", LlmnrParser.qtypeName(q.qtype))
put("kind", if (q.isQuery) "query" else "response")
}
}
}
val evidence: JsonObject = buildJsonObject {
put("capture", capture.statusJson())
put("llmnr_queries", table.toJson())
put("evidence_truncated", capture.truncated || table.overflowed)
}
val metrics: JsonObject = buildJsonObject {
put("distinct_sources", table.distinctSources)
put("distinct_names", table.size)
put("packets", capture.packetsSeen)
put("queries", queries)
put("responses", responses)
put("undecodable_packets", undecodable)
// The headline: whether this protocol is live here at all. A boolean rather than an
// inference from a count, so a reader (or a future finding rule) never has to decide
// what "zero packets" meant — the capture's own status says whether zero is trustworthy.
put("llmnr_in_use", table.distinctSources > 0)
put("truncated", capture.truncated || table.overflowed)
}
val saw = !table.isEmpty
return build(
capture.outcome(saw),
evidence = evidence,
metrics = metrics,
params = params(),
error = capture.reason(saw)?.let { TestError("listen_incomplete", it) },
)
}
private fun params(): JsonObject = buildJsonObject {
put("mode", "passive")
put("group_v4", "$GROUP4:$PORT")
put("group_v6", "[$GROUP6]:$PORT")
put("started_mono_ns", startedAtMonoNs)
}
private companion object {
const val PORT = 5355
const val GROUP4 = "224.0.0.252"
const val GROUP6 = "ff02::1:3"
}
}
@@ -0,0 +1,115 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.probe
import android.content.Context
import android.net.nsd.NsdManager
import android.net.nsd.NsdServiceInfo
import android.net.wifi.WifiManager
import app.echo_lot.measurement.Test
import app.echo_lot.measurement.TestStatus
import app.echo_lot.measurement.TestType
import app.echo_lot.measurement.Tier
import kotlinx.coroutines.Dispatchers
import kotlinx.coroutines.delay
import kotlinx.coroutines.withContext
import kotlinx.serialization.json.buildJsonObject
import kotlinx.serialization.json.put
import kotlinx.serialization.json.putJsonObject
import java.util.Collections
/**
* local.mdns_inventory what answers mDNS on this network (MulticastLock + NSD discovery).
* The service inventory doubles as the VLAN-leakage detector: a chromecast answering on the
* guest wifi is a segmentation fault made visible. Folded from the prober, validated on both
* known devices (4 services each).
*
* Two hardware-bought lessons are load-bearing here:
* - The `_services._dns-sd._udp.` meta-query returned 0 on BOTH devices while concrete types
* found live services NsdManager's meta-query support is unreliable across builds, so the
* concrete types are the measurement and the meta-query result is itself evidence.
* - 4 s of listening missed services that 10 s catches; mDNS answers straggle.
*
* That last lesson is why [listenMs] is a parameter. 10 s is the short-mode default because it is
* the shortest window that was not demonstrably lossy; a long run hands it the whole measurement
* window, since the curve does not stop at ten seconds devices announce on their own schedule,
* and a printer that is asleep answers when something else wakes it.
*/
class MdnsInventoryProbe(private val listenMs: Long = 10_000) : Probe {
override val type = TestType.LOCAL_MDNS_INVENTORY
override val tier = Tier.APP
override val estimatedMs = listenMs + 500
/** Meta-query + common concrete types (HTTP covers HA/printers/NAS; googlecast is ubiquitous). */
private val queries = listOf(
"meta" to "_services._dns-sd._udp.",
"http" to "_http._tcp.",
"googlecast" to "_googlecast._tcp.",
)
private class Recorder : NsdManager.DiscoveryListener {
val names: MutableList<String> = Collections.synchronizedList(mutableListOf())
@Volatile var started = false
@Volatile var startFailCode: Int? = null
override fun onStartDiscoveryFailed(t: String?, code: Int) { startFailCode = code }
override fun onStopDiscoveryFailed(t: String?, code: Int) {}
override fun onDiscoveryStarted(t: String?) { started = true }
override fun onDiscoveryStopped(t: String?) {}
override fun onServiceFound(s: NsdServiceInfo?) { s?.serviceName?.let { names.add(it) } }
override fun onServiceLost(s: NsdServiceInfo?) {}
}
override suspend fun run(ctx: Context, ids: ProbeIds): Test = withContext(Dispatchers.IO) {
val b = TestBuilder(type, tier, ids)
val wifi = ctx.getSystemService(WifiManager::class.java)
val lock = wifi?.createMulticastLock("echolot")?.apply {
setReferenceCounted(false)
runCatching { acquire() }
}
val nsd = ctx.getSystemService(NsdManager::class.java)
?: return@withContext b.build(
TestStatus.UNSUPPORTED,
evidence = buildJsonObject { put("reason", "NsdManager unavailable") },
)
val recorders = queries.map { (label, type) ->
val r = Recorder()
runCatching { nsd.discoverServices(type, NsdManager.PROTOCOL_DNS_SD, r) }
.onFailure { r.startFailCode = -1 }
Triple(label, type, r)
}
try {
delay(listenMs)
var total = 0
var anyStarted = false
val evidence = buildJsonObject {
put("multicast_lock", lock?.isHeld == true)
for ((label, type, r) in recorders) {
runCatching { nsd.stopServiceDiscovery(r) }
anyStarted = anyStarted || r.started
val names = r.names.distinct()
total += names.size
putJsonObject(label) {
put("query", type)
put("started", r.started)
r.startFailCode?.let { put("start_fail_code", it) }
put("found", names.size)
if (names.isNotEmpty()) put("names", names.joinToString(", ").take(300))
}
}
}
val metrics = buildJsonObject { put("services_found", total) }
// How long it listened belongs in params: "4 services" means something different after
// ten seconds than after five minutes, and the number alone cannot say which it was.
val params = buildJsonObject { put("listen_ms", listenMs) }
// Zero services on a started discovery is a legitimate result (an empty or properly
// isolated network), not a failure — only discovery refusing to start is one.
b.build(if (anyStarted) TestStatus.OK else TestStatus.FAILED,
evidence = evidence, metrics = metrics, params = params)
} finally {
recorders.forEach { (_, _, r) -> runCatching { nsd.stopServiceDiscovery(r) } }
runCatching { lock?.release() }
}
}
}
@@ -0,0 +1,420 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.probe
import android.content.Context
import android.net.wifi.WifiManager
import app.echo_lot.measurement.TestStatus
import kotlinx.coroutines.CoroutineScope
import kotlinx.coroutines.Dispatchers
import kotlinx.coroutines.Job
import kotlinx.coroutines.SupervisorJob
import kotlinx.coroutines.launch
import kotlinx.coroutines.withContext
import kotlinx.serialization.json.JsonObject
import kotlinx.serialization.json.buildJsonObject
import kotlinx.serialization.json.put
import kotlinx.serialization.json.putJsonArray
import kotlinx.serialization.json.add
import java.net.DatagramPacket
import java.net.Inet4Address
import java.net.Inet6Address
import java.net.InetAddress
import java.net.InetSocketAddress
import java.net.MulticastSocket
import java.net.NetworkInterface
import java.net.SocketAddress
import java.net.SocketTimeoutException
/**
* The listening half of every passive-discovery collector: hold a multicast lock, bind a
* well-known UDP port, join a group on every interface that will take it, and retain the datagrams
* that arrive until the window closes.
*
* Four things here are the difference between a working listener and one that silently reports an
* empty network, and each of them has cost somebody a day:
*
* - **The MulticastLock.** Android's wifi stack filters out frames not addressed to this device
* unless an app holds one. Without it every collector built on this returns "nothing on this
* network" on a network full of chatter — the single most common reason a listener like this
* appears to work and measures nothing. Whether it was actually held is therefore reported, not
* assumed: an unheld lock plus silence is not a clean result and [outcome] refuses to call it one.
* - **SO_REUSEADDR before bind.** The well-known discovery ports are shared by construction the
* system's own mDNS/SSDP responders and any other app doing this are already there so binding
* exclusively fails on exactly the networks worth measuring.
* - **Joining per interface.** The group must be joined on the interface the traffic arrives on,
* which is not necessarily the default route: a phone with wifi plus a VPN plus cellular has
* three, and the LAN chatter is on the one that is not carrying the default route.
* - **A bound buffer, bounded per source too.** A chatty network must not decide how much memory
* this uses, and one device announcing every two seconds must not be able to fill the buffer and
* hide the twenty quieter devices behind it. Both caps set [truncated] rather than being silent.
*
* Nothing here throws. Every failure lands in [failure] (the capture is dead) or [degraded] (it is
* running but cannot see everything), which the owning collector turns into a recorded status.
*/
internal class MulticastCapture(
/** Names the MulticastLock, so a `dumpsys wifi` during a run says which listener holds what. */
private val label: String,
private val port: Int,
private val group4: String? = null,
private val group6: String? = null,
private val maxPackets: Int = 400,
private val maxBytesRetained: Int = 128 * 1024,
private val maxPerSource: Int = 24,
/**
* Whether to fall back to an ephemeral port when the well-known one cannot be bound.
*
* Only useful for a protocol with an active half: replies to our own searches come back to
* whatever port we sent from, so an ephemeral socket still collects those but it can never
* see the unsolicited announcements, which is why it is [degraded] and not normal operation.
*/
private val allowEphemeralFallback: Boolean = false,
) {
/** One retained datagram. Not a data class: the payload is a ByteArray, and structural equality
* over it would be both wrong and expensive. */
class Packet(val sourceIp: String, val atMonoNs: Long, val data: ByteArray)
/** Set when the capture could not be brought up at all; the collector reports `unsupported`. */
var failure: String? = null
private set
/** Set when the capture is running but blind to part of what it exists to see. */
var degraded: String? = null
private set
var lockHeld: Boolean = false
private set
var boundPort: Int = 0
private set
var packetsSeen: Int = 0
private set
var truncated: Boolean = false
private set
val joined4 = ArrayList<String>()
val joined6 = ArrayList<String>()
private val joinErrors = ArrayList<String>()
/** Whether this protocol needs a group join at all — NetBIOS is broadcast, not multicast. */
private val expectsGroup = group4 != null || group6 != null
private val packets = ArrayList<Packet>()
private val perSource = HashMap<String, Int>()
private var retainedBytes = 0
@Volatile private var running = false
private var socket: MulticastSocket? = null
private var lock: WifiManager.MulticastLock? = null
private var scope: CoroutineScope? = null
/**
* Brings the listener up and returns whether it is capturing. Setup is a handful of syscalls,
* so it finishes in milliseconds and the [Collector.start] promptness contract holds; only the
* receive loop is handed to a background scope.
*/
suspend fun start(ids: ProbeIds): Boolean = withContext(Dispatchers.IO) {
val s = bind() ?: return@withContext false
socket = s
running = true
// A degraded (ephemeral-port) socket is deliberately not joined to the group: it would be
// joining for a port nothing sends to, and the resulting "joined wlan0" in the evidence
// would claim a capability the capture does not have.
if (expectsGroup && degraded == null) joinGroups(s)
scope = CoroutineScope(SupervisorJob() + Dispatchers.IO).also { it.launch { pump(s, ids) } }
true
}
/** Acquires the multicast lock. Separate from [start] because it needs a Context and the rest
* does not, and because whether it succeeded is itself evidence. */
fun acquireLock(ctx: Context) {
val wifi = runCatching { ctx.applicationContext.getSystemService(WifiManager::class.java) }
.getOrNull()
lock = runCatching {
wifi?.createMulticastLock("echolot-$label")?.apply {
setReferenceCounted(false)
acquire()
}
}.getOrNull()
lockHeld = runCatching { lock?.isHeld == true }.getOrDefault(false)
}
/** Sends from the capture socket, so replies land back in this capture rather than on a second
* socket nobody is reading. Returns whether the datagram left the device. */
fun send(payload: ByteArray, host: String, toPort: Int): Boolean = runCatching {
val s = socket ?: return false
s.send(DatagramPacket(payload, payload.size, InetSocketAddress(host, toPort)))
true
}.getOrDefault(false)
/** Stops listening and hands back what was retained. Safe after a failed [start], and never
* throws a cancelled run still owes the caller the packets it did see. */
fun stop(): List<Packet> {
running = false
// Closed before the coroutine is cancelled: a blocking receive() does not notice
// cancellation, and closing the socket is what makes it return.
runCatching { socket?.close() }
socket = null
scope?.coroutineContext?.get(Job)?.cancel()
scope = null
runCatching { lock?.release() }
lock = null
// lockHeld deliberately survives the release: it records whether the capture *could* hear
// multicast while it ran, which is what [outcome] needs to decide whether silence is a
// fact about the network. Clearing it here would make every quiet network report that the
// lock was missing.
return synchronized(packets) { packets.toList() }
}
// ---- status ------------------------------------------------------------------------------
/**
* The status this capture's Test deserves, given whether anything was decoded.
*
* The interesting case is the last one. Silence on a network is a legitimate and useful
* finding but only when we know the listener could have heard something. Without the
* multicast lock, or without a single successful group join, "nothing was seen" describes this
* app and not the network, and reporting it as `ok` would be the collector lying by omission.
*/
fun outcome(sawAnything: Boolean): TestStatus = when {
failure != null -> TestStatus.UNSUPPORTED
degraded != null -> TestStatus.PARTIAL
sawAnything -> TestStatus.OK
expectsGroup && joined4.isEmpty() && joined6.isEmpty() -> TestStatus.PARTIAL
!lockHeld -> TestStatus.PARTIAL
else -> TestStatus.OK
}
/** Why [outcome] was not OK, in words, or null when it was. */
fun reason(sawAnything: Boolean): String? = when {
failure != null -> failure
degraded != null -> degraded
sawAnything -> null
expectsGroup && joined4.isEmpty() && joined6.isEmpty() ->
"no interface accepted the group join, so silence here says nothing about the network"
!lockHeld ->
"the wifi multicast lock was not held, so silence here says nothing about the network"
else -> null
}
/** The capture's own facts, for the collector's evidence. Every collector reports these
* identically so that "saw nothing" can always be told apart from "could not listen". */
fun statusJson(): JsonObject = buildJsonObject {
put("multicast_lock", lockHeld)
put("bound_port", boundPort)
putJsonArray("joined_v4") { for (n in joined4) add(n) }
putJsonArray("joined_v6") { for (n in joined6) add(n) }
if (joinErrors.isNotEmpty()) {
put("join_errors", joinErrors.take(6).joinToString("; ").take(400))
}
degraded?.let { put("degraded", it) }
failure?.let { put("failure", it) }
put("packets_seen", packetsSeen)
put("packets_retained", synchronized(packets) { packets.size })
put("evidence_truncated", truncated)
}
// ---- internals ---------------------------------------------------------------------------
private fun bind(): MulticastSocket? {
// Unbound first, so SO_REUSEADDR is set *before* bind — setting it afterwards has no
// effect, and these ports are always already in use by something.
runCatching {
val s = MulticastSocket(null as SocketAddress?)
s.reuseAddress = true
s.bind(InetSocketAddress(port))
s.soTimeout = SO_TIMEOUT_MS
boundPort = port
return s
}.onFailure { first ->
val why = describe(first)
if (!allowEphemeralFallback) {
// Ports below 1024 are privileged on Android as on any Linux, so this is the
// expected outcome for NetBIOS and is a finding rather than a bug — see
// NetbiosCollector.
failure = "could not bind UDP $port: $why"
return null
}
runCatching {
val s = MulticastSocket(null as SocketAddress?)
s.reuseAddress = true
s.bind(InetSocketAddress(0))
s.soTimeout = SO_TIMEOUT_MS
boundPort = s.localPort
degraded = "UDP $port could not be bound ($why); listening on an ephemeral port " +
"instead, so only replies to our own searches are visible and unsolicited " +
"announcements are not"
return s
}.onFailure { failure = "could not bind UDP $port ($why) or any ephemeral port" }
}
return null
}
private fun joinGroups(s: MulticastSocket) {
val ifaces = runCatching { NetworkInterface.getNetworkInterfaces()?.toList() }
.getOrNull().orEmpty()
.filter {
runCatching { it.isUp && !it.isLoopback && it.supportsMulticast() }
.getOrDefault(false)
}
val g4 = group4?.let { runCatching { InetAddress.getByName(it) }.getOrNull() }
val g6 = group6?.let { runCatching { InetAddress.getByName(it) }.getOrNull() }
for (ni in ifaces) {
val addrs = runCatching { ni.inetAddresses.toList() }.getOrNull().orEmpty()
if (g4 != null && addrs.any { it is Inet4Address }) {
join(s, g4, ni)?.let { joined4.add(it) }
}
// A network with no IPv6 address is not a failure to report as an error — it is a
// v4-only network, which is most of them. Only interfaces that could have joined are
// asked to.
if (g6 != null && addrs.any { it is Inet6Address }) {
join(s, g6, ni)?.let { joined6.add(it) }
}
}
// Last resort: let the kernel pick the interface. Some vendor builds refuse the explicit
// form on the very interface that carries the traffic, and a default-interface join is
// better than no listener at all.
if (joined4.isEmpty() && joined6.isEmpty()) {
@Suppress("DEPRECATION")
(g4 ?: g6)?.let { g ->
runCatching { s.joinGroup(g) }
.onSuccess { (if (g is Inet4Address) joined4 else joined6).add("(default)") }
.onFailure { joinErrors.add("default: ${describe(it)}") }
}
}
}
private fun join(s: MulticastSocket, group: InetAddress, ni: NetworkInterface): String? =
runCatching {
s.joinGroup(InetSocketAddress(group, boundPort), ni)
ni.name
}.onFailure {
joinErrors.add("${ni.name}/${if (group is Inet4Address) "v4" else "v6"}: ${describe(it)}")
}.getOrNull()
private fun pump(s: MulticastSocket, ids: ProbeIds) {
val buf = ByteArray(READ_BUFFER)
while (running) {
val p = DatagramPacket(buf, buf.size)
try {
s.receive(p)
} catch (t: SocketTimeoutException) {
// The timeout exists only so this loop notices `running` going false if the socket
// close somehow does not wake it. Nothing to record.
continue
} catch (t: Throwable) {
// The normal exit: stop() closed the socket underneath us. It is also what a
// vanishing interface looks like, and neither is worth a status of its own — the
// packets gathered so far are still the measurement.
return
}
packetsSeen++
val ip = p.address?.hostAddress ?: continue
val len = p.length
if (len <= 0) continue
synchronized(packets) {
val fromThis = perSource[ip] ?: 0
if (packets.size >= maxPackets ||
retainedBytes + len > maxBytesRetained ||
fromThis >= maxPerSource
) {
truncated = true
} else {
packets.add(Packet(ip, ids.monoNs(), p.data.copyOf(len)))
perSource[ip] = fromThis + 1
retainedBytes += len
}
}
}
}
private companion object {
const val SO_TIMEOUT_MS = 1_000
/** Larger than any discovery datagram anyone sends; oversized ones are truncated by the
* kernel, which the parsers survive by design. */
const val READ_BUFFER = 4_096
fun describe(t: Throwable): String =
(t.message ?: t.javaClass.simpleName).take(160)
}
}
/**
* The other half of what every discovery collector does: fold a stream of decoded messages into a
* deduplicated inventory of *things*, not packets.
*
* Deduplication is by source **and** identity, never by identity alone. Two devices announcing the
* same UPnP service type are two devices, and collapsing them would turn the one measurement worth
* having (how many things are on this segment) into a count of protocols. Conversely one device
* re-announcing every thirty seconds must not appear ten times that is what [count] is for, and
* the repetition rate is itself readable from count over the window length.
*
* First-seen is kept on the monotonic clock per the two-clock rule: it answers "was this device
* here from the start, or did it appear four minutes in", which is exactly the question a long run
* exists to answer and the one a wall-clock stamp cannot be trusted for.
*/
internal class DiscoveryTable(private val maxEntries: Int = MAX_ENTRIES) {
private class Row(val firstSeenMonoNs: Long, val fields: JsonObject) {
var count = 0
}
private val rows = LinkedHashMap<Pair<String, String>, Row>()
private val sources = HashSet<String>()
/** True when a network was busy enough that entries had to be dropped — reported, never hidden. */
var overflowed = false
private set
/**
* [fields] is a lambda so the JSON for a repeat sighting is never built: on a chatty segment
* the overwhelming majority of packets are the same device saying the same thing again.
*/
fun observe(sourceIp: String, identity: String, atMonoNs: Long, fields: () -> JsonObject) {
sources.add(sourceIp)
val key = sourceIp to identity
val existing = rows[key]
if (existing != null) {
existing.count++
return
}
if (rows.size >= maxEntries) {
overflowed = true
return
}
val built = buildJsonObject {
put("source_ip", sourceIp)
for ((k, v) in fields()) put(k, v)
}
rows[key] = Row(atMonoNs, built).also { it.count = 1 }
}
val distinctSources: Int get() = sources.size
val size: Int get() = rows.size
val isEmpty: Boolean get() = rows.isEmpty()
fun toJson(): kotlinx.serialization.json.JsonArray = kotlinx.serialization.json.JsonArray(
rows.values.map { r ->
JsonObject(
r.fields + mapOf(
"first_seen_mono_ns" to kotlinx.serialization.json.JsonPrimitive(r.firstSeenMonoNs),
"count" to kotlinx.serialization.json.JsonPrimitive(r.count),
)
)
}
)
private companion object {
/** Generous enough for any real segment, small enough that a broadcast storm cannot make
* one test dominate the document. */
const val MAX_ENTRIES = 200
}
}
@@ -0,0 +1,131 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.probe
import android.content.Context
import app.echo_lot.measurement.Test
import app.echo_lot.measurement.TestError
import app.echo_lot.measurement.TestType
import app.echo_lot.measurement.Tier
import kotlinx.serialization.json.JsonObject
import kotlinx.serialization.json.buildJsonObject
import kotlinx.serialization.json.put
/**
* local.netbios_inventory NetBIOS name-service chatter (UDP 137) on this segment.
*
* NetBIOS name registration and query traffic is broadcast, not multicast, so there is no group to
* join: a socket bound to 137 sees it by being on the segment. Every packet carries a
* first-level-encoded name plus a suffix byte saying what the name is *for* a workstation, a file
* server, a master browser which makes a passive window over it a Windows-side inventory of the
* LAN, and, like LLMNR, a security observation in its own right: NBT-NS is the other half of the
* classic name-resolution spoofing surface.
*
* **Expect this to report `unsupported` on the app tier, and read that as a result rather than a
* bug.** 137 is below 1024, and Linux Android included reserves those ports for privileged
* processes; an unprivileged app UID cannot bind one. So the honest measurement here is usually
* "this tier cannot observe NetBIOS on this device", recorded with the exact bind error, rather
* than a silent absence that reads as a quiet network. The observation belongs to the Shizuku tier,
* whose shell UID can bind it; the collector is written now so that the decoder, the evidence shape
* and the registry id are settled and tested by the time that lands.
*
* An active alternative exists send an NBSTAT query from an ephemeral port and read the unicast
* replies and is deliberately not taken: it is a host sweep, which is scanning rather than
* measuring, and it would change what the run does to the network it is observing.
*/
class NetbiosCollector : BaseCollector() {
override val type = TestType.LOCAL_NETBIOS_INVENTORY
override val tier = Tier.APP
private val capture = MulticastCapture(
label = "netbios",
port = PORT,
// No group: NBT-NS is subnet broadcast. The multicast lock is still taken, because
// Android's wifi firmware filters broadcast as well as multicast under power save.
group4 = null,
group6 = null,
// Registration bursts repeat the same name several times a second; the per-source cap is
// what keeps one noisy Windows box from filling the buffer.
maxPerSource = 12,
)
private var startedAtMonoNs = 0L
override suspend fun start(ctx: Context, ids: ProbeIds) {
begin(ids)
startedAtMonoNs = ids.monoNs()
capture.acquireLock(ctx)
capture.start(ids)
}
override suspend fun stop(): Test {
val packets = capture.stop()
val table = DiscoveryTable()
var queries = 0
var registrations = 0
var undecodable = 0
for (p in packets) {
val n = NetbiosParser.parse(p.data, p.data.size)
if (n == null) {
undecodable++
continue
}
when (n.opcode) {
0 -> queries++
5, 8 -> registrations++
}
// Identity is name + suffix, not the name alone: one host registers the same name
// several times with different suffixes, and those are different facts about it.
table.observe(p.sourceIp, n.name + "#" + "%02X".format(n.suffix), p.atMonoNs) {
buildJsonObject {
put("netbios_name", n.name)
put("suffix", "0x%02X".format(n.suffix))
put("role", n.role)
put("operation", NetbiosParser.opcodeName(n.opcode))
put("kind", if (n.isResponse) "response" else "request")
}
}
}
val evidence: JsonObject = buildJsonObject {
put("capture", capture.statusJson())
put("netbios_names", table.toJson())
put("evidence_truncated", capture.truncated || table.overflowed)
}
val metrics: JsonObject = buildJsonObject {
put("distinct_sources", table.distinctSources)
put("distinct_names", table.size)
put("packets", capture.packetsSeen)
put("name_queries", queries)
put("name_registrations", registrations)
put("undecodable_packets", undecodable)
put("netbios_in_use", table.distinctSources > 0)
put("truncated", capture.truncated || table.overflowed)
}
val saw = !table.isEmpty
return build(
capture.outcome(saw),
evidence = evidence,
metrics = metrics,
params = params(),
error = capture.reason(saw)?.let {
TestError(if (capture.failure != null) "port_unavailable" else "listen_incomplete", it)
},
)
}
private fun params(): JsonObject = buildJsonObject {
put("mode", "passive")
put("port", PORT)
put("transport", "udp broadcast")
put("started_mono_ns", startedAtMonoNs)
}
private companion object {
const val PORT = 137
}
}
@@ -0,0 +1,274 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.probe
import android.content.Context
import android.net.ConnectivityManager
import android.net.LinkProperties
import android.net.Network
import android.net.NetworkCapabilities
import android.net.NetworkRequest
import app.echo_lot.measurement.NetworkChange
import app.echo_lot.measurement.NetworkChanges
import app.echo_lot.measurement.Test
import app.echo_lot.measurement.TestError
import app.echo_lot.measurement.TestStatus
import app.echo_lot.measurement.TestType
import app.echo_lot.measurement.Tier
import kotlinx.serialization.json.JsonObject
import kotlinx.serialization.json.buildJsonObject
import kotlinx.serialization.json.put
import kotlinx.serialization.json.putJsonArray
import kotlinx.serialization.json.addJsonObject
import java.util.Collections
/**
* Watches every network for the whole run and fills `networks[].changes[]` (measurement-schema §4).
*
* That array has been in the schema since the first draft and has never been populated, because
* nothing in a battery of one-shot probes is in a position to fill it. It is the single most
* valuable thing a long run adds: a wifi link that drops and returns mid-window is invisible to
* every probe the ones before and after the gap both succeed and it is exactly the fault people
* open a network diagnostic to chase.
*
* Reported as [TestType.LINK_IP_MONITOR] at app tier. The registry lists that type as Shizuku's
* (`ip monitor`), and this is deliberately the same observation from a tier that does not need it:
* link state as it changes over time. A document may therefore carry two `link.ip_monitor` tests,
* told apart by `tier` which is what `tier` is for.
*/
class NetworkChangeCollector(
private val entries: List<NetworkInventory.Entry>,
) : BaseCollector() {
override val type = TestType.LINK_IP_MONITOR
override val tier = Tier.APP
private data class Change(
val atMonoNs: Long,
val kind: String,
val event: String,
val iface: String,
val detail: String,
)
private val changes: MutableList<Change> = Collections.synchronizedList(mutableListOf())
private var cm: ConnectivityManager? = null
private var callback: ConnectivityManager.NetworkCallback? = null
private var registerError: String? = null
/**
* Last seen state per network, so only real changes are recorded.
*
* Both capability and link-property callbacks fire constantly on a live device signal
* strength alone re-delivers capabilities every few seconds and a five-minute window of that
* would bury the four events that matter under several hundred that do not. The first callback
* after a network appears is the baseline, not a change.
*/
private val lastCaps = HashMap<String, String>()
private val lastLink = HashMap<String, String>()
/** Interface name per network handle, remembered because `onLost` can no longer look it up. */
private val ifaceOf = HashMap<String, String>()
/**
* The networks that were already up when the window opened.
*
* `registerNetworkCallback` replays `onAvailable` for every matching network the instant it is
* registered, so without this every run would open with three "the wifi appeared" events that
* describe the registration and not the network. A link that drops and returns comes back as a
* new handle, which is not in this set, so real re-appearances are still recorded.
*/
private val seeded = HashSet<String>()
override suspend fun start(ctx: Context, ids: ProbeIds) {
begin(ids)
val manager = ctx.getSystemService(ConnectivityManager::class.java)
if (manager == null) {
registerError = "ConnectivityManager unavailable"
return
}
cm = manager
// Seed the interface names from the snapshot the run already took, so a network that is
// lost without ever having delivered a callback here is still attributable.
for (e in entries) {
e.model.iface?.let { ifaceOf[key(e.handle)] = it }
seeded.add(key(e.handle))
}
val cb = object : ConnectivityManager.NetworkCallback() {
override fun onAvailable(network: Network) {
val iface = resolveIface(network)
if (key(network) in seeded) return
record(ids, NetworkChanges.GAINED, "available", iface, "network became available")
}
override fun onLost(network: Network) {
val iface = ifaceOf[key(network)] ?: "(unknown)"
record(ids, NetworkChanges.LOST, "lost", iface, "network went away")
// Dropped so a returning link re-baselines instead of reporting every property it
// ever had as a change the moment it comes back.
lastCaps.remove(key(network)); lastLink.remove(key(network))
}
override fun onCapabilitiesChanged(network: Network, caps: NetworkCapabilities) {
val iface = resolveIface(network)
val print = capsFingerprint(caps)
val previous = lastCaps.put(key(network), print)
if (previous == null || previous == print) return
record(ids, NetworkChanges.LINK_CHANGED, "capabilities", iface, "$previous$print")
}
override fun onLinkPropertiesChanged(network: Network, lp: LinkProperties) {
lp.interfaceName?.let { ifaceOf[key(network)] = it }
val print = linkFingerprint(lp)
val previous = lastLink.put(key(network), print)
if (previous == null || previous == print) return
record(ids, NetworkChanges.LINK_CHANGED, "link_properties", lp.interfaceName ?: "(unknown)",
describeLinkDelta(previous, print))
}
override fun onLosing(network: Network, maxMsToLive: Int) {
val iface = ifaceOf[key(network)] ?: "(unknown)"
record(ids, NetworkChanges.LINK_CHANGED, "losing", iface, "about to be torn down in ${maxMsToLive} ms")
}
}
callback = cb
// clearCapabilities(), or the default request only matches INTERNET + NOT_RESTRICTED and
// the carrier's IMS/MMS networks — and, more importantly, a network in the middle of
// failing validation — never appear. The transports are named explicitly so this does not
// also follow whatever internal networks a vendor keeps in the list.
val request = NetworkRequest.Builder()
.clearCapabilities()
.addTransportType(NetworkCapabilities.TRANSPORT_WIFI)
.addTransportType(NetworkCapabilities.TRANSPORT_CELLULAR)
.addTransportType(NetworkCapabilities.TRANSPORT_ETHERNET)
.addTransportType(NetworkCapabilities.TRANSPORT_VPN)
.build()
runCatching { manager.registerNetworkCallback(request, cb) }
.onFailure { registerError = it.message ?: it.javaClass.simpleName; callback = null }
}
override suspend fun stop(): Test {
callback?.let { cb -> runCatching { cm?.unregisterNetworkCallback(cb) } }
callback = null
val snapshot = synchronized(changes) { changes.toList() }
if (registerError != null) {
return build(
TestStatus.UNSUPPORTED,
error = TestError("callback_unavailable", registerError),
)
}
val evidence: JsonObject = buildJsonObject {
putJsonArray("changes") {
for (c in snapshot) addJsonObject {
put("at_mono_ns", c.atMonoNs)
put("kind", c.kind)
put("event", c.event)
put("interface", c.iface)
put("detail", c.detail)
}
}
}
val perIface = snapshot.groupBy { it.iface }
val metrics: JsonObject = buildJsonObject {
put("changes_total", snapshot.size)
put("networks_lost", snapshot.count { it.kind == NetworkChanges.LOST })
put("networks_gained", snapshot.count { it.event == "available" })
put("link_changes", snapshot.count { it.kind == NetworkChanges.LINK_CHANGED })
// The number a reader actually wants: how many times a link went away and came back.
// Counted by the same function the flapping finding uses, so the metric and the
// finding can never tell different stories about one window.
put("flap_cycles", perIface.values.sumOf { NetworkChanges.flapCycles(it.map { c -> c.kind }) })
}
// Zero changes over the window is a real, useful result — a stable network — so it is OK
// rather than a failure. The window that produced it is what makes that mean anything, and
// it is recorded in run.mode plus the sibling collectors' params.
return build(TestStatus.OK, evidence = evidence, metrics = metrics)
}
/**
* The changes belonging to each network in `networks[]`, keyed by its model id.
*
* Matched by interface name rather than by Android's `Network` handle, because a link that
* drops and returns comes back as a *different* handle with the same interface and the
* flapping case is precisely the one this must not lose. Changes on an interface that was not
* in the run's initial snapshot stay in the test evidence but have no `networks[]` entry to
* hang from.
*/
fun changesByNetwork(): Map<String, List<NetworkChange>> {
val byIface = entries.mapNotNull { e -> e.model.iface?.let { it to e.model.id } }.toMap()
val out = LinkedHashMap<String, MutableList<NetworkChange>>()
for (c in synchronized(changes) { changes.toList() }) {
val id = byIface[c.iface] ?: continue
out.getOrPut(id) { mutableListOf() }.add(
NetworkChange(
atMonoNs = c.atMonoNs,
kind = c.kind,
detail = buildJsonObject {
put("event", c.event)
put("interface", c.iface)
put("detail", c.detail)
},
)
)
}
return out
}
private fun record(ids: ProbeIds, kind: String, event: String, iface: String, detail: String) {
changes.add(Change(ids.monoNs(), kind, event, iface, detail))
}
private fun resolveIface(network: Network): String {
val known = ifaceOf[key(network)]
if (known != null) return known
val name = runCatching { cm?.getLinkProperties(network)?.interfaceName }.getOrNull()
if (name != null) ifaceOf[key(network)] = name
return name ?: "(unknown)"
}
private companion object {
/** Android's own network id, stable for the life of one Network object. */
private fun key(n: Network): String = n.toString()
/**
* Only the capabilities whose change means something diagnostically.
*
* Bandwidth estimates and signal strength are deliberately excluded: they change every few
* seconds on a moving device, and including them turns a change log into a sampling log.
* VALIDATED and CAPTIVE_PORTAL are the two that matter most they are the moment Android
* decides a network does or does not carry the internet.
*/
private fun capsFingerprint(c: NetworkCapabilities): String = buildString {
fun flag(name: String, cap: Int) {
if (runCatching { c.hasCapability(cap) }.getOrDefault(false)) append(name).append(' ')
}
flag("internet", NetworkCapabilities.NET_CAPABILITY_INTERNET)
flag("validated", NetworkCapabilities.NET_CAPABILITY_VALIDATED)
flag("captive", NetworkCapabilities.NET_CAPABILITY_CAPTIVE_PORTAL)
flag("not-metered", NetworkCapabilities.NET_CAPABILITY_NOT_METERED)
flag("not-suspended", NET_CAPABILITY_NOT_SUSPENDED)
flag("not-restricted", NetworkCapabilities.NET_CAPABILITY_NOT_RESTRICTED)
}.trim().ifEmpty { "(none)" }
/** NetworkCapabilities.NET_CAPABILITY_NOT_SUSPENDED, API 28+ (@SystemApi constant). */
private const val NET_CAPABILITY_NOT_SUSPENDED = 21
private fun linkFingerprint(lp: LinkProperties): String {
val addrs = lp.linkAddresses.map { it.toString() }.sorted().joinToString(",")
val routes = lp.routes.map { it.toString() }.sorted().joinToString(",")
val dns = lp.dnsServers.mapNotNull { it.hostAddress }.sorted().joinToString(",")
return "mtu=${lp.mtu}|addr=$addrs|route=$routes|dns=$dns"
}
/** Names which part of the link changed, so the detail is readable without a diff tool. */
private fun describeLinkDelta(before: String, after: String): String {
val b = before.split('|'); val a = after.split('|')
val changed = b.indices.filter { it < a.size && b[it] != a[it] }
.map { a[it].substringBefore('=') }
return if (changed.isEmpty()) "changed" else "changed: ${changed.joinToString(", ")}"
}
}
}
@@ -39,6 +39,46 @@ object NetworkInventory {
return out return out
} }
/**
* Android's own verdict on the network, read straight from the capabilities it already has.
*
* PARTIAL_CONNECTIVITY only exists from API 28 and CAPTIVE_PORTAL from 23, so both are read
* defensively: an older platform that cannot answer should leave the field null rather than
* assert a false.
*/
private fun systemVerdict(caps: NetworkCapabilities): app.echo_lot.measurement.SystemVerdict =
app.echo_lot.measurement.SystemVerdict(
validated = caps.hasCapability(NetworkCapabilities.NET_CAPABILITY_VALIDATED),
captivePortal = runCatching {
caps.hasCapability(NetworkCapabilities.NET_CAPABILITY_CAPTIVE_PORTAL)
}.getOrNull(),
// NET_CAPABILITY_PARTIAL_CONNECTIVITY is @SystemApi, so the constant is not in the
// public SDK even though the platform sets it from API 28. The number is stable —
// changing it would break every system app that reads it — but this is a value the
// SDK does not promise us, so it is asked for defensively and reported as unknown
// rather than as false if anything about it is not as expected.
partialConnectivity = runCatching {
caps.hasCapability(NET_CAPABILITY_PARTIAL_CONNECTIVITY)
}.getOrNull(),
)
/** @SystemApi NetworkCapabilities.NET_CAPABILITY_PARTIAL_CONNECTIVITY, API 28+. */
private const val NET_CAPABILITY_PARTIAL_CONNECTIVITY = 24
/**
* Can an ordinary app send on this network?
*
* The carrier's special-purpose networks (IMS/VoLTE, MMS, XCAP) sit in `allNetworks` next to
* the real ones, carrying `IMS`/`MMS` but neither `INTERNET` nor `NOT_RESTRICTED`. Binding
* to them fails with EPERM forever, because it needs `CONNECTIVITY_USE_RESTRICTED_NETWORKS`
* signature-level, unobtainable for this app. Asking the capabilities up front is what
* separates "the OS will never let us measure this" from "something is blocking us", which
* a bind attempt alone cannot distinguish and which the constraint logic must not confuse.
*/
private fun appUsable(caps: NetworkCapabilities): Boolean =
caps.hasCapability(NetworkCapabilities.NET_CAPABILITY_INTERNET) &&
caps.hasCapability(NetworkCapabilities.NET_CAPABILITY_NOT_RESTRICTED)
private fun toModel(id: String, caps: NetworkCapabilities, lp: LinkProperties): MNetwork { private fun toModel(id: String, caps: NetworkCapabilities, lp: LinkProperties): MNetwork {
val transport = when { val transport = when {
caps.hasTransport(NetworkCapabilities.TRANSPORT_WIFI) -> Transport.WIFI caps.hasTransport(NetworkCapabilities.TRANSPORT_WIFI) -> Transport.WIFI
@@ -72,6 +112,8 @@ object NetworkInventory {
return MNetwork( return MNetwork(
id = id, transport = transport, iface = lp.interfaceName, id = id, transport = transport, iface = lp.interfaceName,
link = Link(mtu = lp.mtu.takeIf { it > 0 }, addresses = addresses, routes = routes, dns = dns), link = Link(mtu = lp.mtu.takeIf { it > 0 }, addresses = addresses, routes = routes, dns = dns),
systemVerdict = systemVerdict(caps),
appUsable = appUsable(caps),
) )
} }
} }
@@ -0,0 +1,70 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.probe
import android.system.Os
import java.io.FileDescriptor
/**
* Linux socket-option ABI numbers that android.system.OsConstants does NOT reliably expose.
* Stable across Android's supported ABIs at the IP/IPv6 protocol levels, which is why they can
* be hardcoded: if setsockoptInt with one of these succeeds, the kernel accepted the option; if
* it throws ErrnoException, it did not. Either outcome is data. Do not "fix" these to
* OsConstants names they don't exist there (validated in the prober; see its OsAbi.kt).
*
* Measured fact worth keeping: `Os.getsockoptInt` is absent on both known devices (OnePlus 15
* A16, Lenovo TB330FU A15), so path-MTU values must be read from the errqueue (`ee_info`), never
* from getsockopt(IP_MTU).
*/
object OsAbi {
// IP level
const val IP_TTL = 2
const val IP_MTU_DISCOVER = 10
const val IP_MTU = 14
const val IP_RECVERR = 11
const val IP_PMTUDISC_DO = 2 // set DF, honor PMTU
const val IP_PMTUDISC_PROBE = 3 // set DF, ignore PMTU (for probing)
// IPv6 level
const val IPV6_MTU_DISCOVER = 23
const val IPV6_MTU = 24
const val IPV6_RECVERR = 25
const val IPV6_UNICAST_HOPS = 16
const val IPV6_PMTUDISC_PROBE = 3
// recv flags — not in OsConstants on any current API level
const val MSG_ERRQUEUE = 0x2000
const val MSG_DONTWAIT = 0x40
// struct sock_extended_err (uapi/linux/errqueue.h), fixed layout on all Android ABIs:
// u32 ee_errno; u8 ee_origin; u8 ee_type; u8 ee_code; u8 ee_pad; u32 ee_info; u32 ee_data;
// followed directly by the offender sockaddr (SO_EE_OFFENDER).
const val SOCK_EE_SIZE = 16
const val SO_EE_ORIGIN_ICMP = 2
const val ICMP_TIME_EXCEEDED = 11
const val ICMP_DEST_UNREACH = 3
/** Try setsockoptInt; return null on success, or the errno name on failure. */
fun trySetIntOpt(fd: FileDescriptor, level: Int, opt: Int, value: Int): String? =
try {
Os.setsockoptInt(fd, level, opt, value)
null
} catch (e: Throwable) {
e.message ?: e.javaClass.simpleName
}
/**
* getsockoptInt is not part of the stable public Os surface on every API level, so it is
* reached via reflection; callers must treat failure as "unreadable", not as an error.
*/
fun tryGetIntOpt(fd: FileDescriptor, level: Int, opt: Int): Result<Int> = runCatching {
val m = Os::class.java.getMethod(
"getsockoptInt",
FileDescriptor::class.java,
Int::class.javaPrimitiveType,
Int::class.javaPrimitiveType,
)
m.invoke(null, fd, level, opt) as Int
}
}
@@ -0,0 +1,158 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.probe
import android.content.Context
import app.echo_lot.measurement.Test
import app.echo_lot.measurement.TestStatus
import app.echo_lot.measurement.TestType
import app.echo_lot.measurement.Tier
import kotlinx.coroutines.CoroutineScope
import kotlinx.coroutines.Dispatchers
import kotlinx.coroutines.Job
import kotlinx.coroutines.SupervisorJob
import kotlinx.coroutines.delay
import kotlinx.coroutines.isActive
import kotlinx.coroutines.launch
import kotlinx.serialization.json.JsonNull
import kotlinx.serialization.json.JsonObject
import kotlinx.serialization.json.JsonPrimitive
import kotlinx.serialization.json.buildJsonObject
import kotlinx.serialization.json.put
import kotlinx.serialization.json.putJsonArray
import kotlinx.serialization.json.add
import java.util.Collections
/**
* icmp.ping4 sampled across the window loss and jitter over minutes instead of one packet.
*
* The battery's [IcmpProbe] answers "does this network reply at all", which one echo can settle.
* It cannot answer "how often does it not", and that is the complaint people actually have:
* 2 % loss is invisible to a single ping and ruins a video call. A series over five minutes also
* catches loss that comes in bursts, which an average taken over ten packets in one second cannot
* distinguish from a clean link.
*
* Emitted as its own `icmp.ping4` test alongside the battery's. `params` carries the window and
* the interval precisely so the two are never mistaken for each other a reader seeing 300 sent
* packets in one and 1 in the other must be able to tell which is which without guessing.
*
* Sustained loss here deliberately emits **no finding**. Every loss code in the registry is about
* the server path `connectivity.udp_loss` and its directional siblings all say "UDP", and they
* mean the probe protocol's traffic, whose direction the server can attest to. ICMP echo to a
* public address is a different measurement with a different set of benign explanations (rate
* limiting at the target is the obvious one), and borrowing a code that claims otherwise would put
* two unrelated things under one dashboard entry the exact failure the registry exists to
* prevent. The metrics say what was seen; a code for it can be added when it has been defined.
*/
class PingSeriesCollector(
private val target: String = "1.1.1.1",
private val intervalMs: Long = 2_000,
/** Deliberately below [intervalMs]: a reply that arrives after the next probe was due is lost
* for any practical purpose, and waiting for it would make the series drift out of cadence. */
private val timeoutMs: Int = 1_500,
) : BaseCollector() {
override val type = TestType.ICMP_PING4
override val tier = Tier.APP
private val txMonoNs: MutableList<Long> = Collections.synchronizedList(mutableListOf())
private val rttMs: MutableList<Double?> = Collections.synchronizedList(mutableListOf())
private var notSent = 0
private var scope: CoroutineScope? = null
private var startedAtMonoNs = 0L
override suspend fun start(ctx: Context, ids: ProbeIds) {
begin(ids)
startedAtMonoNs = ids.monoNs()
// The default network, and only the default network: this measures what the device's own
// traffic experiences over the window. Per-network binding is the battery's job, and doing
// it here would multiply the packet rate by the number of interfaces for no new answer.
val s = CoroutineScope(SupervisorJob() + Dispatchers.IO).also { scope = it }
s.launch {
var seq = 1
while (isActive) {
val t0 = ids.monoNs()
val r = IcmpEcho.ping(null, target, v6 = false, timeoutMs = timeoutMs, seq = seq)
if (r.attempted) {
txMonoNs.add(t0)
rttMs.add(r.rttMs)
} else {
// Never left the device — a socket or bind failure is not packet loss, and
// counting it as loss would blame the network for the app's own trouble.
notSent++
}
seq = (seq + 1) and 0xFFFF
delay(intervalMs)
}
}
}
override suspend fun stop(): Test {
scope?.coroutineContext?.get(Job)?.cancel()
scope = null
val tx = synchronized(txMonoNs) { txMonoNs.toList() }
val rtt = synchronized(rttMs) { rttMs.toList() }
val received = rtt.filterNotNull()
val evidence: JsonObject = buildJsonObject {
putJsonArray("seq") { for (i in tx.indices) add(i) }
putJsonArray("t_tx_ns") { for (v in tx) add(v) }
// null at an index is a lost probe, per the §6.2 columnar convention.
putJsonArray("rtt_ms") {
for (v in rtt) add(v?.let { JsonPrimitive(round1(it)) } ?: JsonNull)
}
}
val metrics: JsonObject = buildJsonObject {
put("sent", tx.size)
put("received", received.size)
put("not_sent", notSent)
if (tx.isNotEmpty()) {
put("loss_pct", round1((tx.size - received.size) * 100.0 / tx.size))
}
if (received.isNotEmpty()) {
put("rtt_ms_min", round1(received.min()))
put("rtt_ms_avg", round1(received.average()))
put("rtt_ms_max", round1(received.max()))
put("jitter_ms", round1(meanDeviation(received)))
}
}
val status = when {
tx.isEmpty() -> TestStatus.UNSUPPORTED
received.isEmpty() -> TestStatus.FAILED
received.size < tx.size -> TestStatus.PARTIAL
else -> TestStatus.OK
}
return build(status, evidence = evidence, metrics = metrics, params = params())
}
private fun params(): JsonObject = buildJsonObject {
// What separates this from the battery's single ping, and what a reader needs to reproduce
// it. Without these two numbers "300 packets, 2 % loss" is a rate nobody can interpret.
put("mode", "series")
put("target", target)
put("interval_ms", intervalMs)
put("timeout_ms", timeoutMs)
put("started_mono_ns", startedAtMonoNs)
put("network", "default")
}
private companion object {
fun round1(v: Double) = Math.round(v * 10.0) / 10.0
/**
* Mean deviation between consecutive round trips jitter as a stream experiences it.
*
* Not the spread around the average: a link that alternates 20 ms / 200 ms and one that
* drifts slowly from 20 ms to 200 ms have the same standard deviation, and only the first
* one breaks a call.
*/
fun meanDeviation(values: List<Double>): Double {
if (values.size < 2) return 0.0
var sum = 0.0
for (i in 1 until values.size) sum += kotlin.math.abs(values[i] - values[i - 1])
return sum / (values.size - 1)
}
}
}
@@ -55,9 +55,10 @@ class TestBuilder(
evidence: JsonObject? = null, evidence: JsonObject? = null,
metrics: JsonObject? = null, metrics: JsonObject? = null,
error: TestError? = null, error: TestError? = null,
params: JsonObject? = null,
): Test = Test( ): Test = Test(
id = id, type = type, networkRef = networkRef, sessionRef = sessionRef, tier = tier, id = id, type = type, networkRef = networkRef, sessionRef = sessionRef, tier = tier,
startedMonoNs = startedMonoNs, endedMonoNs = ids.monoNs(), startedMonoNs = startedMonoNs, endedMonoNs = ids.monoNs(),
status = status, error = error, evidence = evidence, metrics = metrics, status = status, error = error, params = params, evidence = evidence, metrics = metrics,
) )
} }
@@ -0,0 +1,69 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.probe
import android.content.Context
import app.echo_lot.measurement.Test
import app.echo_lot.measurement.TestError
import app.echo_lot.measurement.TestStatus
import kotlinx.coroutines.CoroutineScope
import kotlinx.coroutines.Deferred
import kotlinx.coroutines.Dispatchers
import kotlinx.coroutines.Job
import kotlinx.coroutines.SupervisorJob
import kotlinx.coroutines.async
import kotlinx.coroutines.withTimeoutOrNull
/**
* Runs an ordinary [Probe] beside the battery instead of inside it.
*
* Some probes are already listeners with a fixed window [MdnsInventoryProbe] does nothing but
* wait for answers and in a long run their window should be the run's window. Running them in
* the sequential battery would then stall every probe behind them for five minutes, which is a
* scheduling problem and not a measurement one, so the fix is to move them rather than to shorten
* them.
*
* Durations are deliberately *not* fed back into the estimate learning: a probe that listens for
* the whole window would teach the short-mode progress bar that mDNS discovery takes five minutes.
*/
class ProbeCollector(private val probe: Probe) : BaseCollector() {
override val type get() = probe.type
override val tier get() = probe.tier
private var scope: CoroutineScope? = null
private var running: Deferred<Test>? = null
override suspend fun start(ctx: Context, ids: ProbeIds) {
begin(ids)
val s = CoroutineScope(SupervisorJob() + Dispatchers.IO).also { scope = it }
running = s.async { probe.run(ctx, ids) }
}
override suspend fun stop(): Test {
val job = running
running = null
val finished = job?.let {
// A short grace, not a long one. A probe timed to the window has already finished by
// the time this is called, so the normal path returns instantly; the grace only covers
// it being slightly late. It is deliberately kept to a second and a half because the
// other caller is the Cancel button, where every millisecond spent waiting for a
// listener that will not finish is a millisecond the user watches nothing happen.
withTimeoutOrNull(GRACE_MS) { runCatching { it.await() }.getOrNull() }
}
scope?.coroutineContext?.get(Job)?.cancel()
scope = null
return finished ?: build(
TestStatus.PARTIAL,
error = TestError(
"window_closed",
"the run's window ended before this listener finished",
),
)
}
private companion object {
const val GRACE_MS = 1_500L
}
}
@@ -0,0 +1,181 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.probe
import android.content.Context
import app.echo_lot.measurement.Test
import app.echo_lot.measurement.TestError
import app.echo_lot.measurement.TestType
import app.echo_lot.measurement.Tier
import kotlinx.coroutines.CoroutineScope
import kotlinx.coroutines.Dispatchers
import kotlinx.coroutines.Job
import kotlinx.coroutines.SupervisorJob
import kotlinx.coroutines.delay
import kotlinx.coroutines.isActive
import kotlinx.coroutines.launch
import kotlinx.serialization.json.JsonObject
import kotlinx.serialization.json.buildJsonObject
import kotlinx.serialization.json.put
/**
* local.ssdp_inventory what UPnP/SSDP devices are on this segment, gathered over the whole window.
*
* **Both halves are needed, and neither is sufficient.** Passive listening catches ssdp:alive and
* ssdp:byebye announcements, which is the only way to see a device that ignores searches (plenty
* do, deliberately) and the only way to see one leave. But announcements are periodic and sparse
* a device re-announces on its own cache-control interval, commonly 30 minutes so a five-minute
* passive window silently misses most of the segment. An M-SEARCH provokes an immediate reply from
* everything that is listening, which is most things, and catches the quiet ones. Running only the
* active half is what [RouterIdentityProbe] already does in the battery, and it is why that probe
* cannot tell you that a device disappeared halfway through the run.
*
* The searches are paced, not flooded: [maxSearches] of them spread [searchIntervalMs] apart. A
* repeat catches devices that joined the network after the run started or were asleep at t=0, while
* staying orders of magnitude below a rate that would itself perturb the network being measured
* the tool must not become the fault it is looking for. `upnp:rootdevice` rather than `ssdp:all`
* for the same reason: one reply per device instead of one per service.
*
* The searches go out from the capture socket bound to 1900, so unicast replies land in the same
* capture as the multicast announcements rather than needing a second socket nobody is reading.
*/
class SsdpCollector(
private val searchIntervalMs: Long = 60_000,
private val maxSearches: Int = 5,
) : BaseCollector() {
override val type = TestType.LOCAL_SSDP_INVENTORY
override val tier = Tier.APP
private val capture = MulticastCapture(
label = "ssdp",
port = PORT,
group4 = GROUP4,
// The IPv6 link-local SSDP group. Free to join where IPv6 exists and skipped where it does
// not, so a v4-only network costs nothing and a v6-only device is not invisible.
group6 = GROUP6,
// SSDP is the chattiest of the four; a device announcing every service it hosts can emit a
// dozen NOTIFYs per cycle, so the per-source cap does most of the work here.
maxPackets = 600,
// The only one of the four with an active half worth keeping when 1900 is taken.
allowEphemeralFallback = true,
)
private var scope: CoroutineScope? = null
private var searchesSent = 0
private var startedAtMonoNs = 0L
override suspend fun start(ctx: Context, ids: ProbeIds) {
begin(ids)
startedAtMonoNs = ids.monoNs()
capture.acquireLock(ctx)
if (!capture.start(ids)) return
val s = CoroutineScope(SupervisorJob() + Dispatchers.IO).also { scope = it }
s.launch {
while (isActive && searchesSent < maxSearches) {
if (capture.send(MSEARCH, GROUP4, PORT)) searchesSent++ else break
delay(searchIntervalMs)
}
}
}
override suspend fun stop(): Test {
scope?.coroutineContext?.get(Job)?.cancel()
scope = null
val packets = capture.stop()
val table = DiscoveryTable()
var alive = 0
var byebye = 0
var responses = 0
var undecodable = 0
for (p in packets) {
val msg = SsdpParser.parse(p.data, p.data.size)
if (msg == null) {
undecodable++
continue
}
when (msg.kind) {
SsdpKind.ALIVE -> alive++
SsdpKind.BYEBYE -> byebye++
SsdpKind.RESPONSE -> responses++
// Our own M-SEARCH comes back to us through the group; counting it as a device
// would inventory the phone doing the measuring.
SsdpKind.SEARCH -> continue
else -> Unit
}
// USN is the device+service identity SSDP itself uses; NT/ST is the fallback for
// devices that omit it, and the source IP already separates two devices offering the
// same service type.
val identity = msg.usn ?: msg.target ?: "(unidentified)"
table.observe(p.sourceIp, identity, p.atMonoNs) {
buildJsonObject {
put("usn", identity)
msg.target?.let { put("target", it) }
msg.serverBanner?.let { put("server_banner", it) }
SsdpParser.productHint(msg.serverBanner)?.let { put("product_hint", it) }
msg.location?.let { put("location", it) }
put("kind", msg.kind.name.lowercase())
}
}
}
val evidence: JsonObject = buildJsonObject {
put("capture", capture.statusJson())
put("ssdp_devices", table.toJson())
// Two truncation sources, reported apart: the capture dropping datagrams and the
// inventory dropping distinct entries mean different things about the network.
put("evidence_truncated", capture.truncated || table.overflowed)
}
val metrics: JsonObject = buildJsonObject {
put("distinct_sources", table.distinctSources)
put("distinct_advertisements", table.size)
put("packets", capture.packetsSeen)
put("alive", alive)
put("byebye", byebye)
put("search_responses", responses)
put("undecodable_packets", undecodable)
put("searches_sent", searchesSent)
put("truncated", capture.truncated || table.overflowed)
}
val status = capture.outcome(sawAnything = !table.isEmpty)
val reason = capture.reason(sawAnything = !table.isEmpty)
return build(
status,
evidence = evidence,
metrics = metrics,
params = params(),
error = reason?.let { TestError("listen_incomplete", it) },
)
}
private fun params(): JsonObject = buildJsonObject {
put("mode", "passive+msearch")
put("group_v4", "$GROUP4:$PORT")
put("group_v6", "[$GROUP6]:$PORT")
put("search_target", SEARCH_TARGET)
put("search_interval_ms", searchIntervalMs)
put("max_searches", maxSearches)
put("started_mono_ns", startedAtMonoNs)
}
private companion object {
const val PORT = 1900
const val GROUP4 = "239.255.255.250"
const val GROUP6 = "ff02::c"
const val SEARCH_TARGET = "upnp:rootdevice"
/** MX is the maximum random delay a responder waits, in seconds. 3 spreads the replies
* enough that a segment full of devices does not answer in one burst we then drop. */
val MSEARCH: ByteArray = (
"M-SEARCH * HTTP/1.1\r\n" +
"HOST: $GROUP4:$PORT\r\n" +
"MAN: \"ssdp:discover\"\r\n" +
"MX: 3\r\n" +
"ST: $SEARCH_TARGET\r\n\r\n"
).toByteArray(Charsets.ISO_8859_1)
}
}
@@ -60,6 +60,15 @@ class StunProbe(
override suspend fun run(ctx: Context, ids: ProbeIds): Test = withContext(Dispatchers.IO) { override suspend fun run(ctx: Context, ids: ProbeIds): Test = withContext(Dispatchers.IO) {
val b = TestBuilder(type, tier, ids) val b = TestBuilder(type, tier, ids)
// Without a server there is nothing to ask. Skipped rather than failed: "the STUN test
// failed" reads as a finding about the network, when the truth is that this device is
// not enrolled anywhere and no packet was ever sent.
if (serverHost.isBlank()) {
return@withContext b.build(
TestStatus.SKIPPED,
evidence = buildJsonObject { put("reason", "no server configured to ask") },
)
}
DatagramSocket().use { sock -> DatagramSocket().use { sock ->
sock.soTimeout = 3000 sock.soTimeout = 3000
val localPort = sock.localPort val localPort = sock.localPort
@@ -0,0 +1,255 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.probe
import android.content.Context
import android.system.Os
import android.system.OsConstants
import app.echo_lot.measurement.Flow
import app.echo_lot.measurement.Hop
import app.echo_lot.measurement.HopProbe
import app.echo_lot.measurement.Test
import app.echo_lot.measurement.TestError
import app.echo_lot.measurement.TestStatus
import app.echo_lot.measurement.TestType
import app.echo_lot.measurement.Tier
import app.echo_lot.measurement.TracerouteEvidence
import app.echo_lot.measurement.toEvidence
import kotlinx.coroutines.Dispatchers
import kotlinx.coroutines.delay
import kotlinx.coroutines.withContext
import kotlinx.serialization.json.buildJsonObject
import kotlinx.serialization.json.put
import java.io.FileDescriptor
import java.net.InetAddress
import java.net.InetSocketAddress
import java.nio.ByteBuffer
import java.nio.ByteOrder
/**
* traceroute.udp4 UDP traceroute reading ICMP time-exceeded off the socket error queue via
* Os.recvmsg(MSG_ERRQUEUE): no root, no raw socket, no native code. Folded from the prober,
* which validated real hop addresses on both known devices (6 hops on the OnePlus 15, 5 on the
* Lenovo) and thereby retired the planned C-over-JNI errqueue shim.
*
* StructMsghdr/StructCmsghdr/recvmsg are reached via reflection (repo convention for uncertain
* OS paths): present since roughly API 34, absent before, and the probe must run and report
* on both. An absent API is UNSUPPORTED with the reason, never a crash.
*/
class TracerouteProbe(
private val targetHost: String = "1.1.1.1",
private val maxHops: Int = 6,
) : Probe {
override val type = TestType.TRACEROUTE_UDP4
override val tier = Tier.APP
// Validated wall clock is ~250 ms on a healthy path; the ceiling is maxHops silent hops at
// 900 ms each, which only a blackholing path produces.
override val estimatedMs = 1_500L
private companion object {
const val BASE_PORT = 33434
/** ICMP errors take one RTT to surface on the errqueue; poll briefly, never block. */
const val HOP_DEADLINE_NS = 900_000_000L
const val POLL_INTERVAL_MS = 40L
}
override suspend fun run(ctx: Context, ids: ProbeIds): Test = withContext(Dispatchers.IO) {
val b = TestBuilder(type, tier, ids)
val params = buildJsonObject {
put("target", targetHost); put("max_hops", maxHops); put("base_port", BASE_PORT)
}
val api = ErrqueueApi.resolve()
?: return@withContext b.build(
TestStatus.UNSUPPORTED,
params = params,
error = TestError(
"no_recvmsg",
"StructMsghdr/Os.recvmsg not on this API level — errqueue unreadable",
),
)
var fd: FileDescriptor? = null
try {
fd = Os.socket(OsConstants.AF_INET, OsConstants.SOCK_DGRAM, OsConstants.IPPROTO_UDP)
OsAbi.trySetIntOpt(fd, OsConstants.IPPROTO_IP, OsAbi.IP_RECVERR, 1)?.let {
return@withContext b.build(
TestStatus.UNSUPPORTED,
params = params,
error = TestError("ip_recverr_rejected", it),
)
}
val target = InetAddress.getByName(targetHost)
val hops = ArrayList<Hop>(maxHops)
var hopsSeen = 0
var reachedTarget = false
var srcPort = 0
for (ttl in 1..maxHops) {
OsAbi.trySetIntOpt(fd, OsConstants.IPPROTO_IP, OsAbi.IP_TTL, ttl)
val t0 = System.nanoTime()
val sent = runCatching {
Os.sendto(fd, ByteArray(32), 0, 32, 0, target, BASE_PORT + ttl)
}
if (sent.isFailure) {
hops.add(Hop(ttl, listOf(HopProbe(icmp = "sendto failed: " +
(sent.exceptionOrNull()?.message ?: "?")))))
continue
}
if (srcPort == 0) {
// Only readable after the implicit bind the first send performs.
srcPort = runCatching {
(Os.getsockname(fd) as? InetSocketAddress)?.port ?: 0
}.getOrDefault(0)
}
var hop: ErrqueueApi.ErrEvent? = null
val deadline = System.nanoTime() + HOP_DEADLINE_NS
while (hop == null && System.nanoTime() < deadline) {
hop = api.pollErrqueue(fd)
if (hop == null) delay(POLL_INTERVAL_MS)
}
val rttNs = System.nanoTime() - t0
when {
hop == null -> hops.add(Hop(ttl, listOf(HopProbe()))) // silent hop: all null
hop.parseError != null ->
hops.add(Hop(ttl, listOf(HopProbe(icmp = "unparsed: ${hop.parseError}"))))
else -> {
hopsSeen++
hops.add(Hop(ttl, listOf(HopProbe(
replyFrom = hop.offender,
rttNs = rttNs,
icmp = when (hop.icmpType) {
OsAbi.ICMP_TIME_EXCEEDED -> "time_exceeded"
OsAbi.ICMP_DEST_UNREACH -> "dest_unreachable"
else -> "type_${hop.icmpType}"
},
))))
if (hop.icmpType == OsAbi.ICMP_DEST_UNREACH) reachedTarget = true
}
}
if (reachedTarget) break
}
// dst_port varies per TTL (classic traceroute, and what was validated on hardware),
// so this flow is explicitly NOT fixed-tuple; base_port is in params.
val evidence = TracerouteEvidence(
flow = Flow(srcPort = srcPort, dstPort = BASE_PORT, fixedTuple = false),
hops = hops,
).toEvidence()
val metrics = buildJsonObject {
put("hops_seen", hopsSeen)
put("reached_target", reachedTarget)
}
val status = when {
hopsSeen > 0 -> TestStatus.OK
// API present, sends succeeded, nothing surfaced: a fact about this path or
// kernel, not proof the mechanism is missing.
else -> TestStatus.PARTIAL
}
b.build(status, params = params, evidence = evidence, metrics = metrics)
} catch (e: Throwable) {
b.build(
TestStatus.FAILED,
params = params,
error = TestError("uncaught", e.message ?: e.javaClass.simpleName),
)
} finally {
fd?.let { runCatching { Os.close(it) } }
}
}
}
/**
* Reflection facade over android.system.{StructMsghdr, StructCmsghdr, Os.recvmsg}.
* Resolved once; null if any piece is missing on this API level.
*/
internal class ErrqueueApi private constructor(
private val msghdrCtor: java.lang.reflect.Constructor<*>,
private val recvmsg: java.lang.reflect.Method,
private val cmsgLevel: java.lang.reflect.Field,
private val cmsgType: java.lang.reflect.Field,
private val cmsgData: java.lang.reflect.Field,
private val msgControl: java.lang.reflect.Field,
) {
class ErrEvent(
val offender: String?,
val icmpType: Int,
val origin: Int,
val parseError: String? = null,
)
/** One non-blocking MSG_ERRQUEUE read; null when the queue is empty. */
fun pollErrqueue(fd: FileDescriptor): ErrEvent? {
return try {
val iov = arrayOf(ByteBuffer.allocate(512))
// (SocketAddress msg_name, ByteBuffer[] msg_iov, StructCmsghdr[] msg_control, flags)
val msghdr = msghdrCtor.newInstance(
InetSocketAddress(0), iov, null, 0,
)
recvmsg.invoke(null, fd, msghdr, OsAbi.MSG_ERRQUEUE or OsAbi.MSG_DONTWAIT)
val control = msgControl.get(msghdr) as? Array<*>
?: return ErrEvent(null, -1, -1, "msg_control empty after recvmsg")
for (cmsg in control.filterNotNull()) {
val level = cmsgLevel.getInt(cmsg)
val type = cmsgType.getInt(cmsg)
if (level == OsConstants.IPPROTO_IP && type == OsAbi.IP_RECVERR) {
return parseSockExtendedErr(cmsgData.get(cmsg))
}
}
ErrEvent(null, -1, -1, "no IP_RECVERR cmsg among ${control.size}")
} catch (e: Throwable) {
// The single most load-bearing line: reflection wraps errno in
// InvocationTargetException, and EAGAIN there means "queue empty", not failure.
val cause = (e as? java.lang.reflect.InvocationTargetException)?.cause ?: e
val msg = cause.message ?: cause.javaClass.simpleName
if ("EAGAIN" in msg || "EWOULDBLOCK" in msg) null
else ErrEvent(null, -1, -1, msg)
}
}
/** cmsg_data = struct sock_extended_err + offender sockaddr_in (see OsAbi). */
private fun parseSockExtendedErr(data: Any?): ErrEvent {
val bytes: ByteArray = when (data) {
is ByteArray -> data
is ByteBuffer -> ByteArray(data.remaining()).also { data.duplicate().get(it) }
else -> return ErrEvent(null, -1, -1, "cmsg_data is ${data?.javaClass?.name}")
}
if (bytes.size < OsAbi.SOCK_EE_SIZE) {
return ErrEvent(null, -1, -1, "cmsg_data too short: ${bytes.size}")
}
val origin = bytes[4].toInt() and 0xFF
val icmpType = bytes[5].toInt() and 0xFF
// SO_EE_OFFENDER: sockaddr_in directly after the fixed struct; family is in native
// byte order, sin_addr at offset +4 within the sockaddr.
val offender = if (bytes.size >= OsAbi.SOCK_EE_SIZE + 8) {
val family = ByteBuffer.wrap(bytes, OsAbi.SOCK_EE_SIZE, 2)
.order(ByteOrder.nativeOrder()).short.toInt()
if (family == OsConstants.AF_INET) {
val a = bytes.copyOfRange(OsAbi.SOCK_EE_SIZE + 4, OsAbi.SOCK_EE_SIZE + 8)
InetAddress.getByAddress(a).hostAddress
} else null
} else null
return ErrEvent(offender, icmpType, origin)
}
companion object {
fun resolve(): ErrqueueApi? = runCatching {
val msghdrCls = Class.forName("android.system.StructMsghdr")
val cmsghdrCls = Class.forName("android.system.StructCmsghdr")
ErrqueueApi(
// Picked by shape, not by position: the 4-arg form is
// (SocketAddress, ByteBuffer[], StructCmsghdr[], int) on every level that has it.
msghdrCtor = msghdrCls.constructors.first { it.parameterCount == 4 },
recvmsg = Os::class.java.getMethod(
"recvmsg", FileDescriptor::class.java, msghdrCls, Int::class.javaPrimitiveType,
),
cmsgLevel = cmsghdrCls.getField("cmsg_level"),
cmsgType = cmsghdrCls.getField("cmsg_type"),
cmsgData = cmsghdrCls.getField("cmsg_data"),
msgControl = msghdrCls.getField("msg_control"),
)
}.getOrNull()
}
}
@@ -0,0 +1,131 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.probe
import android.content.Context
import app.echo_lot.measurement.Test
import app.echo_lot.measurement.TestStatus
import app.echo_lot.measurement.TestType
import app.echo_lot.measurement.Tier
import app.echo_lot.measurement.Network as MNetwork
import kotlinx.coroutines.Dispatchers
import kotlinx.coroutines.withContext
import kotlinx.serialization.json.JsonObject
import kotlinx.serialization.json.buildJsonObject
import kotlinx.serialization.json.put
import java.net.Inet6Address
import java.net.InetSocketAddress
import java.util.Locale
/**
* v6.brokenness does IPv6 actually carry traffic, asked with a real TCP connection.
*
* This exists to corroborate (or refute) the ICMPv6 silence that icmp.ping6 observes. ICMPv6 echo
* is widely filtered on networks where IPv6 works fine, so silence alone cannot distinguish
* "IPv6 is broken" from "ping is filtered" a phone that reported v6.broken while happily
* loading IPv6-only sites is what proved the point. A TCP connect over IPv6 to the configured
* server settles it: if it succeeds, IPv6 works and the ICMP silence is filtering; if it fails
* too, on a network that advertises IPv6, the brokenness claim finally has evidence behind it.
*
* Only networks that claim to offer IPv6 (a global address or a v6 default route) are attempted:
* connecting over v6 on an IPv4-only network fails by design, and recording that as evidence
* would manufacture the exact false positive this probe exists to kill.
*/
class V6ConnectProbe(
private val entries: List<NetworkInventory.Entry>,
private val serverHost: String,
private val port: Int = 443,
) : Probe {
override val type = TestType.V6_BROKENNESS
override val tier = Tier.APP
// One 3s connect timeout per v6-provisioned network, at most.
override val estimatedMs = 4_000L
override suspend fun run(ctx: Context, ids: ProbeIds): Test = withContext(Dispatchers.IO) {
val b = TestBuilder(type, tier, ids)
// Same rule as the STUN and canary probes: with no server there is no target, and
// borrowing someone else's infrastructure to get one is not this app's call to make.
if (serverHost.isBlank()) {
return@withContext b.build(
TestStatus.SKIPPED,
evidence = buildJsonObject { put("reason", "no server configured to connect to") },
)
}
val candidates = entries.filter { ipv6Provisioned(it.model) }
if (candidates.isEmpty()) {
return@withContext b.build(
TestStatus.SKIPPED,
evidence = buildJsonObject {
put("reason", "no active network claims to offer IPv6")
},
)
}
var okCount = 0
var attemptedCount = 0
val evidence: JsonObject = buildJsonObject {
put("target", "$serverHost:$port")
for (e in candidates) {
val label = "${e.model.transport.name.lowercase()}:${e.model.id}"
val a = attempt(e)
if (a.attempted) attemptedCount++
if (a.ok) okCount++
put(label, buildJsonObject {
put("network_ref", e.model.id)
put("ok", a.ok)
put("attempted", a.attempted)
put("detail", a.detail)
})
}
}
val status = when {
attemptedCount == 0 -> TestStatus.SKIPPED // resolution/binding never got that far
okCount == attemptedCount -> TestStatus.OK
okCount > 0 -> TestStatus.PARTIAL
else -> TestStatus.FAILED
}
b.build(status, evidence = evidence)
}
/** Same attempted/ok separation as IcmpProbe: a connect we never sent proves nothing. */
private data class Attempt(val ok: Boolean, val attempted: Boolean, val detail: String)
private fun attempt(e: NetworkInventory.Entry): Attempt {
// Resolved through this network's own resolver; a v6 address obtained over another
// network would still be connected to over this one, which is what matters.
val addr = runCatching {
e.handle.getAllByName(serverHost).filterIsInstance<Inet6Address>().firstOrNull()
}.getOrNull()
?: return Attempt(false, false, "no AAAA answer for $serverHost via this network")
// createSocket() binds to the network at creation; failing here means the app could not
// use the interface at all (e.g. EPERM under a VPN) — nothing was sent, nothing is known.
val socket = try {
e.handle.socketFactory.createSocket()
} catch (t: Throwable) {
return Attempt(false, false, "socket unavailable: ${t.message ?: t.javaClass.simpleName}")
}
return try {
val t0 = System.nanoTime()
socket.connect(InetSocketAddress(addr, port), 3000)
val rttMs = (System.nanoTime() - t0) / 1_000_000.0
Attempt(true, true, "connected to [${addr.hostAddress}]:$port " +
"rtt_ms=${"%.1f".format(Locale.ROOT, rttMs)}")
} catch (t: Throwable) {
// A refused connection would still prove the path forwards IPv6, but against our own
// server's 443 the realistic failures are timeout and unreachable — both silence.
Attempt(false, true, "error: ${t.message ?: t.javaClass.simpleName}")
} finally {
runCatching { socket.close() }
}
}
/** The network claims IPv6: a global (non-link-local) address or a v6 default route. */
private fun ipv6Provisioned(n: MNetwork): Boolean =
n.link.addresses.any { a ->
a.addr.contains(':') &&
!a.addr.startsWith("fe80", ignoreCase = true) &&
!a.addr.startsWith("::1")
} || n.link.routes.any { it.dst == "::/0" }
}
@@ -0,0 +1,165 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.probe
import android.content.Context
import android.net.wifi.WifiManager
import app.echo_lot.measurement.Test
import app.echo_lot.measurement.TestError
import app.echo_lot.measurement.TestStatus
import app.echo_lot.measurement.TestType
import app.echo_lot.measurement.Tier
import app.echo_lot.measurement.Transport
import kotlinx.coroutines.CoroutineScope
import kotlinx.coroutines.Dispatchers
import kotlinx.coroutines.Job
import kotlinx.coroutines.SupervisorJob
import kotlinx.coroutines.delay
import kotlinx.coroutines.isActive
import kotlinx.coroutines.launch
import kotlinx.serialization.json.JsonObject
import kotlinx.serialization.json.buildJsonObject
import kotlinx.serialization.json.put
import kotlinx.serialization.json.putJsonArray
import kotlinx.serialization.json.add
import java.util.Collections
/**
* wifi.signal_log RSSI, link speed and frequency sampled across the whole window.
*
* One reading of the signal strength says almost nothing: -67 dBm is fine, and -67 dBm that was
* -45 dBm ninety seconds ago is somebody walking away from the AP, or an AP whose power is being
* managed, or a band steer about to happen. The series is the measurement; the snapshot in
* `networks[].wifi` is only its first sample.
*
* Evidence is columnar (measurement-schema §6.2 conventions): parallel arrays keep a five-minute
* log at 2 s intervals in a few kB.
*/
class WifiSignalCollector(
private val entries: List<NetworkInventory.Entry>,
private val intervalMs: Long = 2_000,
) : BaseCollector() {
override val type = TestType.WIFI_SIGNAL_LOG
override val tier = Tier.APP
private val atMonoNs: MutableList<Long> = Collections.synchronizedList(mutableListOf())
private val rssi: MutableList<Int> = Collections.synchronizedList(mutableListOf())
private val speed: MutableList<Int> = Collections.synchronizedList(mutableListOf())
private val freq: MutableList<Int> = Collections.synchronizedList(mutableListOf())
/** Only the count of distinct BSSIDs leaves this class a roam is the fact worth reporting,
* and the addresses themselves are neighbours' hardware identifiers. */
private val bssids = Collections.synchronizedSet(HashSet<String>())
private var scope: CoroutineScope? = null
private var unsupported: String? = null
private var startedAtMonoNs = 0L
override suspend fun start(ctx: Context, ids: ProbeIds) {
// network_ref up front: this samples the wifi link, and a signal log with nothing to
// attach it to is a series of numbers about an unnamed thing.
val wifiNet = entries.firstOrNull { it.model.transport == Transport.WIFI }
begin(ids, networkRef = wifiNet?.model?.id)
startedAtMonoNs = ids.monoNs()
val wifi = ctx.applicationContext.getSystemService(WifiManager::class.java)
if (wifi == null) {
unsupported = "WifiManager unavailable"
return
}
if (wifiNet == null) {
unsupported = "no wifi network is connected"
return
}
// Own scope, not the caller's: the run job is cancelled the instant the user taps Cancel,
// and the samples taken up to that point are exactly what a cancelled long run still owes
// them. stop() ends this scope.
val s = CoroutineScope(SupervisorJob() + Dispatchers.IO).also { scope = it }
s.launch {
while (isActive) {
sample(ids, wifi)
delay(intervalMs)
}
}
}
@Suppress("DEPRECATION")
private fun sample(ids: ProbeIds, wifi: WifiManager) {
// WifiManager.getConnectionInfo is deprecated in favour of the NetworkCallback's
// TransportInfo, which delivers a WifiInfo only when the capabilities change — i.e. at the
// platform's cadence, not ours, and with no way to ask for a sample. For a fixed-interval
// log the deprecated call is the one that answers the question, and it still works.
val info = runCatching { wifi.connectionInfo }.getOrNull() ?: return
val r = info.rssi
// -127 and 0 are the "no reading" sentinels; recording them would drag every average down
// and invent a signal cliff that never happened.
if (r == 0 || r <= -127) return
atMonoNs.add(ids.monoNs())
rssi.add(r)
speed.add(info.linkSpeed)
freq.add(runCatching { info.frequency }.getOrDefault(0))
runCatching { info.bssid }.getOrNull()
?.takeIf { it.isNotBlank() && it != "02:00:00:00:00:00" }
?.let { bssids.add(it) }
}
override suspend fun stop(): Test {
scope?.coroutineContext?.get(Job)?.cancel()
scope = null
unsupported?.let {
return build(
TestStatus.UNSUPPORTED,
params = params(),
error = TestError("no_wifi", it),
)
}
val t = synchronized(atMonoNs) { atMonoNs.toList() }
val r = synchronized(rssi) { rssi.toList() }
val sp = synchronized(speed) { speed.toList() }
val f = synchronized(freq) { freq.toList() }
val evidence: JsonObject = buildJsonObject {
putJsonArray("at_mono_ns") { for (v in t) add(v) }
putJsonArray("rssi_dbm") { for (v in r) add(v) }
putJsonArray("link_speed_mbps") { for (v in sp) add(v) }
putJsonArray("frequency_mhz") { for (v in f) add(v) }
}
val metrics: JsonObject = buildJsonObject {
put("samples", r.size)
if (r.isNotEmpty()) {
put("rssi_dbm_min", r.min())
put("rssi_dbm_avg", round1(r.average()))
put("rssi_dbm_max", r.max())
put("rssi_dbm_range", r.max() - r.min())
}
sp.filter { it > 0 }.let { valid ->
if (valid.isNotEmpty()) {
put("link_speed_mbps_min", valid.min())
put("link_speed_mbps_avg", round1(valid.average()))
put("link_speed_mbps_max", valid.max())
}
}
// Distinct BSSIDs minus the one we started on: how often the phone changed AP without
// the network ever going down — invisible to any one-shot probe, and a common cause of
// "the call drops when I walk into the kitchen".
put("roams", (bssids.size - 1).coerceAtLeast(0))
}
return build(
if (r.isEmpty()) TestStatus.PARTIAL else TestStatus.OK,
evidence = evidence, metrics = metrics, params = params(),
)
}
private fun params(): JsonObject = buildJsonObject {
put("interval_ms", intervalMs)
put("started_mono_ns", startedAtMonoNs)
put("source", "WifiManager.connectionInfo")
}
private companion object {
fun round1(v: Double) = Math.round(v * 10.0) / 10.0
}
}
@@ -0,0 +1,133 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.probe
import android.content.Context
import app.echo_lot.measurement.Test
import app.echo_lot.measurement.TestError
import app.echo_lot.measurement.TestType
import app.echo_lot.measurement.Tier
import kotlinx.serialization.json.JsonObject
import kotlinx.serialization.json.buildJsonObject
import kotlinx.serialization.json.put
/**
* local.wsd_inventory WS-Discovery (SOAP-over-UDP, 3702) Hello / Bye / Probe / ProbeMatches.
*
* This is the protocol Windows and modern printers use to find each other, and it inventories a
* class of device the other three miss: network printers, scanners and IP cameras announce here and
* frequently nowhere else. A `Hello` is a device arriving, a `Bye` is one leaving, and a `Probe`
* from a workstation names what it is hunting for so a window over 3702 shows both the equipment
* on the segment and which machines are looking for it.
*
* Passive only. WS-Discovery's active half is a Probe multicast, which would make this app a
* participant announcing itself to every device on the segment; SSDP's M-SEARCH is a single small
* request that devices expect constantly, whereas a WSD Probe from an unknown host is the kind of
* thing that shows up in someone's security log. Listening costs the network nothing.
*
* Payloads are decoded by [WsdParser], which does targeted extraction rather than XML parsing
* see its documentation for why that is the right call for unauthenticated broadcast input.
*/
class WsdCollector : BaseCollector() {
override val type = TestType.LOCAL_WSD_INVENTORY
override val tier = Tier.APP
private val capture = MulticastCapture(
label = "wsd",
port = PORT,
group4 = GROUP4,
group6 = GROUP6,
// SOAP envelopes are an order of magnitude larger than the other three protocols'
// datagrams, so the byte ceiling binds before the packet count does. Both are set
// explicitly rather than left to the default, which was chosen for 200-byte packets.
maxPackets = 250,
maxBytesRetained = 192 * 1024,
)
private var startedAtMonoNs = 0L
override suspend fun start(ctx: Context, ids: ProbeIds) {
begin(ids)
startedAtMonoNs = ids.monoNs()
capture.acquireLock(ctx)
capture.start(ids)
}
override suspend fun stop(): Test {
val packets = capture.stop()
val table = DiscoveryTable()
var hello = 0
var bye = 0
var probes = 0
var matches = 0
var undecodable = 0
for (p in packets) {
val m = WsdParser.parse(p.data, p.data.size)
if (m == null) {
undecodable++
continue
}
when (m.action?.lowercase()) {
"hello" -> hello++
"bye" -> bye++
"probe" -> probes++
"probematches", "resolvematches" -> matches++
}
// The device UUID is WS-Discovery's own stable identity and survives address changes,
// so it is the identity where present; a Probe carries none (it names what it wants,
// not who it is), and there the action plus the types is what distinguishes one
// observation from a repeat of it.
val identity = m.deviceUuid ?: (m.action.orEmpty() + "/" + m.types.orEmpty())
table.observe(p.sourceIp, identity.ifEmpty { "(unidentified)" }, p.atMonoNs) {
buildJsonObject {
m.deviceUuid?.let { put("device_uuid", it) }
m.action?.let { put("action", it) }
m.types?.let { put("wsd_types", it) }
m.xaddrs?.let { put("wsd_xaddrs", it) }
}
}
}
val evidence: JsonObject = buildJsonObject {
put("capture", capture.statusJson())
put("wsd_devices", table.toJson())
put("evidence_truncated", capture.truncated || table.overflowed)
}
val metrics: JsonObject = buildJsonObject {
put("distinct_sources", table.distinctSources)
put("distinct_devices", table.size)
put("packets", capture.packetsSeen)
put("hello", hello)
put("bye", bye)
put("probes", probes)
put("probe_matches", matches)
put("undecodable_packets", undecodable)
put("truncated", capture.truncated || table.overflowed)
}
val saw = !table.isEmpty
return build(
capture.outcome(saw),
evidence = evidence,
metrics = metrics,
params = params(),
error = capture.reason(saw)?.let { TestError("listen_incomplete", it) },
)
}
private fun params(): JsonObject = buildJsonObject {
put("mode", "passive")
put("group_v4", "$GROUP4:$PORT")
put("group_v6", "[$GROUP6]:$PORT")
put("started_mono_ns", startedAtMonoNs)
}
private companion object {
const val PORT = 3702
const val GROUP4 = "239.255.255.250"
const val GROUP6 = "ff02::c"
}
}
@@ -0,0 +1,386 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.probe
import kotlin.test.Test
import kotlin.test.assertEquals
import kotlin.test.assertFalse
import kotlin.test.assertNotNull
import kotlin.test.assertNull
import kotlin.test.assertTrue
/**
* The decoders in DiscoveryParsers.kt, against payloads shaped like the ones real devices emit and
* against the ones a broken or hostile device emits.
*
* The malformed cases are the point. These four decoders are the only place in the app where bytes
* from an unidentified third party on the local segment are interpreted; they run inside a
* collector whose contract is that it never throws, and every one of them is reachable by anyone
* who can put a frame on the wire. So each protocol is fed truncation, junk, and the specific abuse
* its format invites a DNS compression pointer, a NetBIOS name outside the A-P alphabet, an XML
* entity bomb and the assertion is always the same pair: no exception, and no invented data.
*/
class DiscoveryParsersTest {
private fun bytes(s: String) = s.toByteArray(Charsets.ISO_8859_1)
// ---- SSDP --------------------------------------------------------------------------------
private val notifyAlive = bytes(
"NOTIFY * HTTP/1.1\r\n" +
"HOST: 239.255.255.250:1900\r\n" +
"CACHE-CONTROL: max-age=1800\r\n" +
"LOCATION: http://192.168.1.44:8060/\r\n" +
"NT: upnp:rootdevice\r\n" +
"NTS: ssdp:alive\r\n" +
"SERVER: Roku/12.5.5 UPnP/1.0 Roku/12.5.5\r\n" +
"USN: uuid:roku:ecp:YH00E1234567::upnp:rootdevice\r\n\r\n"
)
private val searchResponse = bytes(
"HTTP/1.1 200 OK\r\n" +
"CACHE-CONTROL: max-age=1800\r\n" +
"EXT:\r\n" +
"LOCATION: http://192.168.1.1:49000/rootDesc.xml\r\n" +
"SERVER: FRITZ!Box 7590 UPnP/1.0 AVM FRITZ!Box 7590 154.07.57\r\n" +
"ST: upnp:rootdevice\r\n" +
"USN: uuid:75802409-bccb-40e7-8e6c-c0ff33445566::upnp:rootdevice\r\n\r\n"
)
@Test
fun ssdpAliveAnnouncementYieldsIdentityAndLocation() {
val m = assertNotNull(SsdpParser.parse(notifyAlive, notifyAlive.size))
assertEquals(SsdpKind.ALIVE, m.kind)
assertEquals("upnp:rootdevice", m.target)
assertEquals("uuid:roku:ecp:YH00E1234567::upnp:rootdevice", m.usn)
assertEquals("http://192.168.1.44:8060/", m.location)
assertEquals("Roku/12.5.5 UPnP/1.0 Roku/12.5.5", m.serverBanner)
}
@Test
fun ssdpByebyeIsDistinguishedFromAlive() {
val byebye = bytes(
"NOTIFY * HTTP/1.1\r\nHOST: 239.255.255.250:1900\r\n" +
"NT: urn:schemas-upnp-org:device:MediaRenderer:1\r\nNTS: ssdp:byebye\r\n" +
"USN: uuid:aabbccdd::urn:schemas-upnp-org:device:MediaRenderer:1\r\n\r\n"
)
val m = assertNotNull(SsdpParser.parse(byebye, byebye.size))
assertEquals(SsdpKind.BYEBYE, m.kind)
assertEquals("urn:schemas-upnp-org:device:MediaRenderer:1", m.target)
assertNull(m.location)
}
@Test
fun ssdpSearchResponseReadsStAsTheTarget() {
val m = assertNotNull(SsdpParser.parse(searchResponse, searchResponse.size))
assertEquals(SsdpKind.RESPONSE, m.kind)
assertEquals("upnp:rootdevice", m.target)
assertEquals("http://192.168.1.1:49000/rootDesc.xml", m.location)
}
/** Our own M-SEARCH comes back through the group; it must be recognisable so the collector
* does not inventory the phone doing the measuring. */
@Test
fun ssdpOwnSearchIsRecognisable() {
val search = bytes(
"M-SEARCH * HTTP/1.1\r\nHOST: 239.255.255.250:1900\r\n" +
"MAN: \"ssdp:discover\"\r\nMX: 3\r\nST: upnp:rootdevice\r\n\r\n"
)
assertEquals(SsdpKind.SEARCH, assertNotNull(SsdpParser.parse(search, search.size)).kind)
}
@Test
fun ssdpProductHintDropsBoilerplateAndKeepsTheModel() {
assertEquals("Roku/12.5.5", SsdpParser.productHint("Roku/12.5.5 UPnP/1.0 Roku/12.5.5"))
assertEquals("Synology/DSM-7.3", SsdpParser.productHint("Linux/4.4 UPnP/1.0 Synology/DSM-7.3"))
val fritz = assertNotNull(SsdpParser.productHint("FRITZ!Box 7590 UPnP/1.0 AVM FRITZ!Box 7590"))
assertTrue(fritz.contains("FRITZ!Box"))
assertFalse(fritz.contains("UPnP"), "protocol boilerplate leaked into the model hint")
assertNull(SsdpParser.productHint(null))
assertNull(SsdpParser.productHint(" "))
// A banner that is nothing but boilerplate reveals no model and must say so, rather than
// returning an empty string that reads as a name nobody could see.
assertNull(SsdpParser.productHint("Linux/4.4 UPnP/1.0"))
}
@Test
fun ssdpMalformedIsRejectedWithoutThrowing() {
// Truncated mid-header: the headers that did arrive are still usable, and the missing NTS
// makes the kind unknown rather than making the packet a lie.
val cut = notifyAlive.copyOf(70)
assertEquals(SsdpKind.OTHER, assertNotNull(SsdpParser.parse(cut, cut.size)).kind)
assertNull(SsdpParser.parse(ByteArray(0), 0))
val blank = bytes("\r\n\r\n")
assertNull(SsdpParser.parse(blank, blank.size))
val http = bytes("GET / HTTP/1.0\r\n\r\n")
assertNull(SsdpParser.parse(http, http.size), "not an SSDP verb")
// Arbitrary binary, including the high bytes ISO-8859-1 must not choke on.
val junk = ByteArray(256) { it.toByte() }
assertNull(SsdpParser.parse(junk, junk.size))
// A declared length longer than the buffer must be refused, not read past.
assertNull(SsdpParser.parse(notifyAlive, notifyAlive.size + 100))
// Header lines with no colon are skipped rather than fatal.
val noColon = bytes("NOTIFY * HTTP/1.1\r\ngarbage line\r\nNTS: ssdp:alive\r\n\r\n")
assertEquals(SsdpKind.ALIVE, assertNotNull(SsdpParser.parse(noColon, noColon.size)).kind)
}
// ---- LLMNR -------------------------------------------------------------------------------
/** Builds a DNS-format packet: header, one question, nothing else. */
private fun dnsQuery(
name: String,
qtype: Int = 1,
id: Int = 0x1234,
flags: Int = 0x0000,
qdcount: Int = 1,
): ByteArray {
val out = ArrayList<Byte>()
fun u16(v: Int) { out.add((v shr 8).toByte()); out.add((v and 0xFF).toByte()) }
u16(id); u16(flags); u16(qdcount); u16(0); u16(0); u16(0)
for (label in name.split('.')) {
val b = label.toByteArray(Charsets.UTF_8)
out.add(b.size.toByte())
b.forEach { out.add(it) }
}
out.add(0)
u16(qtype); u16(1)
return out.toByteArray()
}
@Test
fun llmnrQueryYieldsTheNameAWorkstationIsHuntingFor() {
val p = dnsQuery("wpad")
val q = assertNotNull(LlmnrParser.parse(p, p.size))
assertEquals("wpad", q.name)
assertEquals(1, q.qtype)
assertEquals("A", LlmnrParser.qtypeName(q.qtype))
assertTrue(q.isQuery)
assertEquals(0, q.opcode)
}
@Test
fun llmnrMultiLabelNamesAndAaaaSurviveIntact() {
val p = dnsQuery("nas-backup.local", qtype = 28)
val q = assertNotNull(LlmnrParser.parse(p, p.size))
assertEquals("nas-backup.local", q.name)
assertEquals("AAAA", LlmnrParser.qtypeName(q.qtype))
}
@Test
fun llmnrResponsesAreSeparatedFromQueries() {
val p = dnsQuery("DESKTOP-A1B2C3", flags = 0x8000)
assertFalse(assertNotNull(LlmnrParser.parse(p, p.size)).isQuery)
}
@Test
fun llmnrMalformedIsRejectedWithoutThrowing() {
val good = dnsQuery("printer")
assertNull(LlmnrParser.parse(ByteArray(0), 0))
assertNull(LlmnrParser.parse(good, 8), "a header-length prefix is not a question")
assertNull(LlmnrParser.parse(good, good.size + 50), "declared length past the buffer")
val noQuestion = dnsQuery("x", qdcount = 0)
assertNull(LlmnrParser.parse(noQuestion, noQuestion.size))
// A label length that runs off the end of the datagram — the classic truncation.
val overrun = good.copyOf(good.size - 6)
assertNull(LlmnrParser.parse(overrun, overrun.size))
// A compression pointer: legal DNS, forbidden in LLMNR, and the shape that makes a naive
// decoder loop forever. It must be refused rather than followed.
val pointer = good.copyOf(20)
pointer[12] = 0xC0.toByte()
pointer[13] = 0x0C
assertNull(LlmnrParser.parse(pointer, pointer.size))
// Random bytes behind a plausible header: whatever comes back, it is not an exception.
val junk = ByteArray(64) { (it * 37).toByte() }
junk[4] = 0; junk[5] = 1
LlmnrParser.parse(junk, junk.size)
}
// ---- NetBIOS -----------------------------------------------------------------------------
/** First-level encoding, the same transform the decoder has to undo. */
private fun nbnsPacket(name: String, suffix: Int, flags: Int = 0x2810): ByteArray {
val raw = ByteArray(16) { ' '.code.toByte() }
name.forEachIndexed { i, c -> if (i < 15) raw[i] = c.code.toByte() }
raw[15] = suffix.toByte()
val out = ArrayList<Byte>()
fun u16(v: Int) { out.add((v shr 8).toByte()); out.add((v and 0xFF).toByte()) }
u16(0x8001); u16(flags); u16(1); u16(0); u16(0); u16(1)
out.add(32)
for (b in raw) {
val v = b.toInt() and 0xFF
out.add(('A'.code + (v shr 4)).toByte())
out.add(('A'.code + (v and 0x0F)).toByte())
}
out.add(0)
u16(0x0020); u16(0x0001)
return out.toByteArray()
}
@Test
fun netbiosNameDecodesWithItsSuffixAndRole() {
val p = nbnsPacket("DESKTOP-A1B2C3", 0x20)
val n = assertNotNull(NetbiosParser.parse(p, p.size))
assertEquals("DESKTOP-A1B2C3", n.name)
assertEquals(0x20, n.suffix)
assertEquals("file_server", n.role)
assertFalse(n.isResponse)
assertEquals("registration", NetbiosParser.opcodeName(n.opcode))
}
@Test
fun netbiosPaddingIsStrippedAndSuffixesAreNamed() {
val p = nbnsPacket("WORKGROUP", 0x1E)
val n = assertNotNull(NetbiosParser.parse(p, p.size))
assertEquals("WORKGROUP", n.name, "the 15-byte space padding leaked into the name")
assertEquals("browser_elections", n.role)
assertEquals("workstation", NetbiosParser.roleOf(0x00))
assertEquals("suffix_0xAB", NetbiosParser.roleOf(0xAB))
}
@Test
fun netbiosResponsesAreSeparatedFromRequests() {
val p = nbnsPacket("FILESRV", 0x20, flags = 0x8500)
assertTrue(assertNotNull(NetbiosParser.parse(p, p.size)).isResponse)
}
@Test
fun netbiosMalformedIsRejectedWithoutThrowing() {
val good = nbnsPacket("PRINTER", 0x00)
assertNull(NetbiosParser.parse(ByteArray(0), 0))
assertNull(NetbiosParser.parse(good, 20), "truncated before the encoded name ends")
assertNull(NetbiosParser.parse(good, good.size + 40), "declared length past the buffer")
// A character outside A-P cannot be half of an encoded byte. Guessing at the rest would
// fabricate a hostname, so the whole packet is refused.
val badAlphabet = good.copyOf()
badAlphabet[15] = 'Z'.code.toByte()
assertNull(NetbiosParser.parse(badAlphabet, badAlphabet.size))
// The length byte must be exactly 32; anything else is a different protocol on this port.
val badLen = good.copyOf()
badLen[12] = 16
assertNull(NetbiosParser.parse(badLen, badLen.size))
// A name that decodes to nothing but padding is not a name.
val blank = nbnsPacket("", 0x00)
assertNull(NetbiosParser.parse(blank, blank.size))
val junk = ByteArray(80) { (it * 13).toByte() }
NetbiosParser.parse(junk, junk.size)
}
@Test
fun netbiosControlCharactersInANameAreNeutralised() {
// Legal first-level encoding, illegal content: these bytes decode cleanly and would
// otherwise reach a JSON document, and from there somebody's terminal.
val p = nbnsPacket("A\u0001B\u0002C", 0x00)
assertEquals("A?B?C", assertNotNull(NetbiosParser.parse(p, p.size)).name)
}
// ---- WS-Discovery ------------------------------------------------------------------------
private val wsdHello = bytes(
"""<?xml version="1.0" encoding="utf-8"?>
<soap:Envelope xmlns:soap="http://www.w3.org/2003/05/soap-envelope"
xmlns:wsa="http://schemas.xmlsoap.org/ws/2004/08/addressing"
xmlns:wsd="http://schemas.xmlsoap.org/ws/2005/04/discovery"
xmlns:wsdp="http://schemas.xmlsoap.org/ws/2006/02/devprof">
<soap:Header>
<wsa:To>urn:schemas-xmlsoap-org:ws:2005:04:discovery</wsa:To>
<wsa:Action>http://schemas.xmlsoap.org/ws/2005/04/discovery/Hello</wsa:Action>
<wsa:MessageID>urn:uuid:0a7e6d1b-0000-4000-8000-000000000001</wsa:MessageID>
</soap:Header>
<soap:Body>
<wsd:Hello>
<wsa:EndpointReference>
<wsa:Address>urn:uuid:9f8e7d6c-1111-4222-8333-444455556666</wsa:Address>
</wsa:EndpointReference>
<wsd:Types>wsdp:Device pub:Computer</wsd:Types>
<wsd:XAddrs>http://192.168.1.77:5357/8f2c-4b1a/</wsd:XAddrs>
<wsd:MetadataVersion>1</wsd:MetadataVersion>
</wsd:Hello>
</soap:Body>
</soap:Envelope>"""
)
@Test
fun wsdHelloYieldsActionUuidTypesAndXaddrs() {
val m = assertNotNull(WsdParser.parse(wsdHello, wsdHello.size))
assertEquals("Hello", m.action)
assertEquals(
"urn:uuid:9f8e7d6c-1111-4222-8333-444455556666", m.deviceUuid,
"the EndpointReference address is the device identity, not the header MessageID",
)
assertEquals("wsdp:Device pub:Computer", m.types)
assertEquals("http://192.168.1.77:5357/8f2c-4b1a/", m.xaddrs)
}
@Test
fun wsdProbeCarriesNoDeviceIdentityAndIsStillRecorded() {
val probe = bytes(
"""<soap:Envelope xmlns:soap="http://www.w3.org/2003/05/soap-envelope"
xmlns:wsa="http://schemas.xmlsoap.org/ws/2004/08/addressing"
xmlns:wsd="http://schemas.xmlsoap.org/ws/2005/04/discovery">
<soap:Header>
<wsa:Action>http://schemas.xmlsoap.org/ws/2005/04/discovery/Probe</wsa:Action>
</soap:Header>
<soap:Body><wsd:Probe><wsd:Types>wsdp:Device</wsd:Types></wsd:Probe></soap:Body>
</soap:Envelope>"""
)
val m = assertNotNull(WsdParser.parse(probe, probe.size))
assertEquals("Probe", m.action)
assertNull(m.deviceUuid)
assertEquals("wsdp:Device", m.types)
}
@Test
fun wsdMalformedIsRejectedWithoutThrowing() {
assertNull(WsdParser.parse(ByteArray(0), 0))
assertNull(WsdParser.parse(wsdHello, wsdHello.size + 100), "declared length past the buffer")
val prose = bytes("hello world")
assertNull(WsdParser.parse(prose, prose.size), "not a SOAP envelope")
// An envelope with no leaf worth reading is an absence, not an empty device.
val bare = bytes("<soap:Envelope></soap:Envelope>")
assertNull(WsdParser.parse(bare, bare.size))
// Cut mid-element: whatever was complete is extracted, the rest is simply absent.
val cut = wsdHello.copyOf(wsdHello.size / 2)
WsdParser.parse(cut, cut.size)?.let { assertNull(it.xaddrs, "an unterminated element was invented") }
val junk = ByteArray(512) { (it * 7).toByte() }
assertNull(WsdParser.parse(junk, junk.size))
}
@Test
fun wsdHostileXmlCostsNothingBecauseNothingParsesItAsXml() {
// A billion-laughs entity bomb. Against a real XML parser this expands to gigabytes; here
// the entities are never resolved, which is the entire reason for not using one.
val bomb = bytes(
"<!DOCTYPE lolz [<!ENTITY lol \"lol\">" +
(0..8).joinToString("") { i ->
val prev = if (i == 0) "" else (i - 1).toString()
"<!ENTITY lol$i \"&lol$prev;&lol$prev;\">"
} +
"]><soap:Envelope><wsa:Action>x/Bye</wsa:Action><body>&lol8;</body></soap:Envelope>"
)
assertEquals("Bye", assertNotNull(WsdParser.parse(bomb, bomb.size)).action)
// Nesting deep enough to blow a recursive-descent parser's stack.
val deep = bytes("<soap:Envelope>" + "<a>".repeat(20_000) + "</soap:Envelope>")
WsdParser.parse(deep, deep.size)
// A payload far larger than any real datagram, pinning that the text cap is applied before
// the matching rather than after it.
val huge = bytes("<soap:Envelope><wsa:Action>x/Hello</wsa:Action>" + "z".repeat(200_000))
assertEquals("Hello", assertNotNull(WsdParser.parse(huge, huge.size)).action)
}
}
@@ -35,13 +35,57 @@ class ControlClient(
private val controlUrl: String, private val controlUrl: String,
pins: Set<String>, pins: Set<String>,
private val appVersion: String = "", private val appVersion: String = "",
/**
* Addresses to fall back to when the server's name will not resolve, learned from its profile.
*
* A measurement tool that cannot report from a broken network is useless exactly when it
* matters, and a wedged resolver is one of the faults it is built to find it should not also
* be the thing that stops the finding being delivered.
*
* Safe because the pin is the trust and the name is not part of it: the server presents the
* same certificate whether it was reached by name or by address, and a wrong address fails the
* pin like anything else would.
*/
private val fallbackAddrs: List<String> = emptyList(),
) { ) {
private val json = Json { ignoreUnknownKeys = true } private val json = Json { ignoreUnknownKeys = true }
private val socketFactory = Pinning.sslContext(pins).socketFactory private val socketFactory = Pinning.sslContext(pins).socketFactory
/**
* The base URL to use, substituting a cached address only when the name genuinely fails.
*
* Resolved once per client and only on failure, so a working network pays nothing and never
* silently drifts onto an address that may be stale.
*/
private val base: String by lazy { resolveBase() }
private fun resolveBase(): String {
if (fallbackAddrs.isEmpty()) return controlUrl
val uri = runCatching { java.net.URI(controlUrl) }.getOrNull() ?: return controlUrl
val host = uri.host ?: return controlUrl
if (runCatching { java.net.InetAddress.getByName(host) }.isSuccess) return controlUrl
val port = if (uri.port > 0) uri.port else 443
for (ip in fallbackAddrs) {
// Checked rather than assumed: on a v4-only network a v6 address would otherwise be
// chosen and fail slowly, which is the wrong answer delivered late.
val reachable = runCatching {
java.net.Socket().use { sock ->
sock.connect(java.net.InetSocketAddress(ip, port), 4000)
true
}
}.getOrDefault(false)
if (reachable) {
val literal = if (ip.contains(':')) "[$ip]" else ip
return uri.scheme + "://" + literal + ":" + port
}
}
return controlUrl
}
private fun open(path: String, method: String, credential: String?): HttpsURLConnection { private fun open(path: String, method: String, credential: String?): HttpsURLConnection {
val conn = URL(controlUrl.trimEnd('/') + path).openConnection() as HttpsURLConnection val conn = URL(base.trimEnd('/') + path).openConnection() as HttpsURLConnection
conn.sslSocketFactory = socketFactory conn.sslSocketFactory = socketFactory
conn.setHostnameVerifier { _, _ -> true } // pin is the trust, not the name conn.setHostnameVerifier { _, _ -> true } // pin is the trust, not the name
conn.requestMethod = method conn.requestMethod = method
@@ -165,6 +209,36 @@ class ControlClient(
} }
} }
/**
* Reports where adbd's wireless-debug listener can be reached on this device's LAN.
*
* Dev scaffolding, not a measurement: it exists because mDNS does not cross subnets, so a
* developer working from a different network cannot discover the port that rotates every few
* minutes. A device sitting on the test LAN can see it and say so. Deliberately not part of
* the probe protocol's capability set the server documents it as a dev relay.
*/
fun reportAdbEndpoint(
credential: String,
host: String,
port: Int,
deviceName: String? = null,
note: String? = null,
): String {
val conn = open("/v1/devtools/adb-endpoint", "POST", credential)
val fields = buildString {
append("""{"host":${jstr(host)},"port":$port""")
deviceName?.let { append(""","device_name":${jstr(it)}""") }
note?.let { append(""","note":${jstr(it)}""") }
append("}")
}
writeJson(conn, fields)
val body = body(conn)
check(conn.responseCode in 200..299) {
"adb endpoint report failed: ${conn.responseCode} $body"
}
return body
}
/** Lists this device's runs stored on the server. */ /** Lists this device's runs stored on the server. */
fun listRuns(credential: String): String { fun listRuns(credential: String): String {
val conn = open("/v1/runs", "GET", credential) val conn = open("/v1/runs", "GET", credential)
@@ -184,6 +258,33 @@ class ControlClient(
open("/v1/runs/$runId", "DELETE", credential).responseCode open("/v1/runs/$runId", "DELETE", credential).responseCode
} }
/**
* Ties this device to the person the ID token identifies.
*
* The device credential proves *which device*, the token proves *which person*; the server
* requires both. Returns the raw JSON reply (account id and display name).
*/
fun linkAccount(credential: String, idToken: String): String {
val conn = open("/v1/account/link", "POST", credential)
writeJson(conn, """{"id_token":${jstr(idToken)}}""")
val text = body(conn)
check(conn.responseCode in 200..299) { "sign-in failed: ${conn.responseCode} $text" }
return text
}
/** Signs out on this device. The device stays enrolled. */
fun unlinkAccount(credential: String) {
open("/v1/account/link", "DELETE", credential).responseCode
}
/** Whether anyone is signed in on this device, and who. */
fun accountStatus(credential: String): String {
val conn = open("/v1/account", "GET", credential)
val text = body(conn)
check(conn.responseCode == 200) { "account status failed: ${conn.responseCode} $text" }
return text
}
fun observations(credential: String, sessionId: String): String { fun observations(credential: String, sessionId: String): String {
val conn = open("/v1/sessions/$sessionId/observations", "GET", credential) val conn = open("/v1/sessions/$sessionId/observations", "GET", credential)
val text = body(conn) val text = body(conn)
@@ -45,11 +45,22 @@ data class EnrollmentLink(
* the reason the pin travels in the link at all. * the reason the pin travels in the link at all.
*/ */
fun redeem(deviceName: String? = null, appVersion: String = ""): Enrolled { fun redeem(deviceName: String? = null, appVersion: String = ""): Enrolled {
val client = ControlClient(controlUrl, setOf(pin), appVersion) // The link may name the server's public address rather than its control endpoint, so that
// a person is handed a name they recognise. Ask where to actually connect.
//
// Only the address comes from here. The pin still comes from the link, because a pin
// fetched over an ordinary TLS connection would be worth exactly what the certificate
// authorities are worth — and pinning exists to survive one the operator does not
// control, such as a root injected by corporate device management. An intercepted
// discovery can therefore send this device to the wrong host, where the pin will not
// match: an outage, not a compromise.
val endpoint = discover(controlUrl) ?: controlUrl
val client = ControlClient(endpoint, setOf(pin), appVersion)
val response = client.enroll(token, deviceName) val response = client.enroll(token, deviceName)
val profile = client.profile(response.credential) val profile = client.profile(response.credential)
return Enrolled( return Enrolled(
controlUrl = controlUrl, controlUrl = endpoint,
publicUrl = controlUrl,
pin = pin, pin = pin,
credential = response.credential, credential = response.credential,
deviceId = response.deviceId, deviceId = response.deviceId,
@@ -57,6 +68,29 @@ data class EnrollmentLink(
) )
} }
/**
* Asks a server where its control plane lives. Null when it does not say, or cannot be asked.
*
* Deliberately forgiving: a server that predates this, or one whose link already names the
* control endpoint directly, simply answers nothing and the link's own URL is used. Enrollment
* must not start failing because an optional lookup did.
*/
private fun discover(publicUrl: String): String? = runCatching {
val conn = (java.net.URL(publicUrl.trimEnd('/') + "/v1/discover").openConnection()
as java.net.HttpURLConnection).apply {
connectTimeout = 8_000
readTimeout = 8_000
setRequestProperty("Accept", "application/json")
}
if (conn.responseCode !in 200..299) return null
val body = conn.inputStream.bufferedReader().use { it.readText() }
kotlinx.serialization.json.Json { ignoreUnknownKeys = true }
.parseToJsonElement(body)
.let { (it as kotlinx.serialization.json.JsonObject)["control_url"] }
?.let { (it as kotlinx.serialization.json.JsonPrimitive).content }
?.takeIf { it.isNotBlank() }
}.getOrNull()
companion object { companion object {
const val SCHEME = "echolot" const val SCHEME = "echolot"
const val HOST = "enroll" const val HOST = "enroll"
@@ -111,7 +145,15 @@ data class EnrollmentLink(
/** A server this device is now enrolled with, ready to be stored in settings. */ /** A server this device is now enrolled with, ready to be stored in settings. */
data class Enrolled( data class Enrolled(
/** Where this device connects: the endpoint whose certificate the pin matches. */
val controlUrl: String, val controlUrl: String,
/**
* The address a person was given, kept for display.
*
* Shown instead of [controlUrl] because the endpoint is plumbing it exists to select a
* certificate while this is the name the operator handed out and would recognise.
*/
val publicUrl: String,
val pin: String, val pin: String,
val credential: String, val credential: String,
val deviceId: String, val deviceId: String,
@@ -29,6 +29,15 @@ data class Target(
val id: String, val id: String,
val ip4: String? = null, val ip4: String? = null,
val ip6: String? = null, val ip6: String? = null,
/**
* The second address, which RFC 5780 behaviour discovery redirects to.
*
* Worth surfacing rather than treating as an implementation detail: a report that says "the
* server did not answer" means something different depending on which of its addresses was
* asked, and an operator reading one needs to be able to tell.
*/
@SerialName("ip4_alt") val ip4Alt: String? = null,
@SerialName("ip6_alt") val ip6Alt: String? = null,
@SerialName("udp_port") val udpPort: Int = 0, @SerialName("udp_port") val udpPort: Int = 0,
@SerialName("tcp_port") val tcpPort: Int = 0, @SerialName("tcp_port") val tcpPort: Int = 0,
@SerialName("stun_port") val stunPort: Int = 0, @SerialName("stun_port") val stunPort: Int = 0,
@@ -74,6 +83,25 @@ data class CompatInfo(
@SerialName("app_max") val appMax: String = "", @SerialName("app_max") val appMax: String = "",
) )
/**
* How to sign in to this server's identity provider, advertised so the app can offer the button
* only when there is something behind it and drive the flow without anyone typing an issuer URL.
*/
@Serializable
data class AuthInfo(
val enabled: Boolean = false,
val issuer: String = "",
@SerialName("client_id") val clientId: String = "",
val flow: String = "",
@SerialName("redirect_uri") val redirectUri: String = "",
val scopes: String = "openid profile email",
@SerialName("authorization_endpoint") val authorizationEndpoint: String = "",
@SerialName("token_endpoint") val tokenEndpoint: String = "",
@SerialName("end_session_endpoint") val endSessionEndpoint: String = "",
/** Present when the server has an issuer configured but could not reach it. */
@SerialName("discovery_error") val discoveryError: String? = null,
)
@Serializable @Serializable
data class Profile( data class Profile(
@SerialName("profile_version") val profileVersion: Int = 0, @SerialName("profile_version") val profileVersion: Int = 0,
@@ -86,6 +114,7 @@ data class Profile(
val pins: List<String> = emptyList(), val pins: List<String> = emptyList(),
val uploads: UploadPolicy = UploadPolicy(), val uploads: UploadPolicy = UploadPolicy(),
val compat: CompatInfo = CompatInfo(), val compat: CompatInfo = CompatInfo(),
val auth: AuthInfo = AuthInfo(),
) { ) {
fun supports(capability: String) = capability in capabilities fun supports(capability: String) = capability in capabilities
} }
@@ -0,0 +1,145 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.protocol
import kotlinx.serialization.json.Json
import kotlinx.serialization.json.jsonObject
import kotlinx.serialization.json.jsonPrimitive
import java.io.IOException
import java.net.HttpURLConnection
import java.net.URL
import java.net.URLEncoder
import java.security.MessageDigest
import java.security.SecureRandom
import java.util.Base64
/**
* Sign-in for the app: authorization code with PKCE (RFC 7636).
*
* The app is a *public* client it ships to devices, so any secret compiled into it can be read
* out of the APK with `unzip` and `strings`. PKCE is what replaces the client secret, and it
* defends a specific attack that matters here more than most places: the redirect comes back
* through a custom URI scheme, and on Android *any* app may register `echolot://`. A malicious one
* could intercept the callback and take the authorization code. Because the code can only be
* exchanged by presenting the verifier which never left this process and cannot be derived from
* the challenge that did a stolen code is worth nothing.
*
* Nothing from the IdP is kept afterwards. The ID token is used once, to prove to the server who
* is signing in, and then discarded: the device credential is what authenticates every later
* request. So there are no access tokens to store, no refresh tokens to rotate, and no token
* lifetime for the app to manage.
*/
object OidcLogin {
/** A started sign-in. [verifier] and [state] must survive until the callback returns. */
data class Pending(val authorizationUrl: String, val verifier: String, val state: String)
/**
* Builds the authorization URL and the secrets that must be held until the callback.
*
* Everything comes from the server's profile rather than being compiled in, so pointing the
* app at a different server with a different IdP is configuration, not a rebuild.
*/
fun begin(auth: AuthInfo, random: SecureRandom = SecureRandom()): Pending {
require(auth.enabled && auth.authorizationEndpoint.isNotBlank()) {
"this server has no identity provider configured"
}
val verifier = randomUrlSafe(random)
val state = randomUrlSafe(random)
val challenge = b64(MessageDigest.getInstance("SHA-256").digest(verifier.toByteArray()))
val q = buildString {
append("response_type=code")
append("&client_id=").append(enc(auth.clientId))
append("&redirect_uri=").append(enc(auth.redirectUri))
append("&scope=").append(enc(auth.scopes))
append("&state=").append(enc(state))
append("&code_challenge=").append(enc(challenge))
append("&code_challenge_method=S256")
}
val sep = if (auth.authorizationEndpoint.contains('?')) "&" else "?"
return Pending(auth.authorizationEndpoint + sep + q, verifier, state)
}
/** What came back on the `echolot://auth` redirect. */
data class Callback(val code: String?, val state: String?, val error: String?)
/** Parses the redirect URI the browser handed back to the app. */
fun parseCallback(uri: String): Callback {
val q = uri.substringAfter('?', "")
var code: String? = null
var state: String? = null
var error: String? = null
for (pair in q.split('&')) {
val k = pair.substringBefore('=')
val v = dec(pair.substringAfter('=', ""))
when (k) {
"code" -> code = v
"state" -> state = v
"error" -> error = v
"error_description" -> if (error != null) error = "$error: $v"
}
}
return Callback(code, state, error)
}
/** The sign-in failed in a way worth showing someone, rather than a transport error. */
class LoginFailed(message: String) : Exception(message)
/**
* Exchanges the code for an ID token.
*
* The state is compared before anything else happens. A callback whose state does not match
* the one this process generated did not come from a flow this process started which is
* precisely how an attacker gets a victim to complete *their* login so it is refused before
* the code is spent.
*/
fun complete(auth: AuthInfo, pending: Pending, callbackUri: String): String {
val cb = parseCallback(callbackUri)
if (cb.error != null) throw LoginFailed(cb.error)
if (cb.state.isNullOrEmpty() || cb.state != pending.state) {
throw LoginFailed("this sign-in did not start on this device — start again")
}
val code = cb.code ?: throw LoginFailed("the identity provider returned no authorization code")
val body = buildString {
append("grant_type=authorization_code")
append("&code=").append(enc(code))
append("&redirect_uri=").append(enc(auth.redirectUri))
append("&client_id=").append(enc(auth.clientId))
append("&code_verifier=").append(enc(pending.verifier))
}
val conn = (URL(auth.tokenEndpoint).openConnection() as HttpURLConnection).apply {
requestMethod = "POST"
doOutput = true
connectTimeout = 15_000
readTimeout = 15_000
setRequestProperty("Content-Type", "application/x-www-form-urlencoded")
setRequestProperty("Accept", "application/json")
}
conn.outputStream.use { it.write(body.toByteArray()) }
val text = try {
val stream = if (conn.responseCode in 200..299) conn.inputStream else conn.errorStream
stream?.bufferedReader()?.use { it.readText() } ?: ""
} catch (e: IOException) {
throw LoginFailed("could not reach the identity provider: ${e.message}")
}
if (conn.responseCode !in 200..299) {
throw LoginFailed("the identity provider refused the sign-in (${conn.responseCode})")
}
val idToken = runCatching {
Json.parseToJsonElement(text).jsonObject["id_token"]?.jsonPrimitive?.content
}.getOrNull()
return idToken?.takeIf { it.isNotBlank() }
?: throw LoginFailed("the identity provider returned no id_token")
}
private fun randomUrlSafe(random: SecureRandom): String =
ByteArray(32).also(random::nextBytes).let(::b64)
private fun b64(b: ByteArray): String = Base64.getUrlEncoder().withoutPadding().encodeToString(b)
private fun enc(s: String): String = URLEncoder.encode(s, "UTF-8")
private fun dec(s: String): String =
runCatching { java.net.URLDecoder.decode(s, "UTF-8") }.getOrDefault(s)
}
@@ -44,13 +44,26 @@ class ProbeSession(
*/ */
fun echo(paddingBytes: Int = 40): EchoResult? { fun echo(paddingBytes: Int = 40): EchoResult? {
val t0 = System.nanoTime() val t0 = System.nanoTime()
val pkt = Wire.build(Wire.TYPE_ECHO_REQ, prefix, ++seq, nowNs(), key, ByteArray(paddingBytes)) val wireSeq = ++seq
val pkt = Wire.build(Wire.TYPE_ECHO_REQ, prefix, wireSeq, nowNs(), key, ByteArray(paddingBytes))
socket.send(DatagramPacket(pkt, pkt.size, server)) socket.send(DatagramPacket(pkt, pkt.size, server))
// A lost probe still has a sequence number, and that number is what lets the server's
// observations say whether it was lost going out or coming back — so report it either way.
lastSeq = wireSeq
val resp = receive(Wire.TYPE_ECHO_RESP) ?: return null val resp = receive(Wire.TYPE_ECHO_RESP) ?: return null
val rttMs = (System.nanoTime() - t0) / 1_000_000.0 val rttMs = (System.nanoTime() - t0) / 1_000_000.0
return EchoResult(rttMs, Observation.parse(resp.payload)) return EchoResult(rttMs, Observation.parse(resp.payload), wireSeq)
} }
/**
* The wire sequence number of the most recent [echo], including one that was lost.
*
* Exposed because the caller cannot derive it: the counter is shared with every other packet
* type on this session, so "the nth echo" is not "sequence n".
*/
var lastSeq: Int = 0
private set
/** One MTU probe of [totalSize] bytes (DF is set by the OS on the socket where supported). /** One MTU probe of [totalSize] bytes (DF is set by the OS on the socket where supported).
* Returns the size the server acknowledged receiving, or null if the probe was lost. */ * Returns the size the server acknowledged receiving, or null if the probe was lost. */
fun mtuProbe(totalSize: Int): Int? { fun mtuProbe(totalSize: Int): Int? {
@@ -96,6 +109,192 @@ class ProbeSession(
return out return out
} }
/**
* Sends paced upstream traffic for [durationMs] and reports what was put on the wire.
*
* Paced rather than flat out, for the same reason the server paces: an unpaced burst measures
* the local NIC and the first queue it meets, then collapses into loss that reads as a network
* fault. The schedule is absolute rather than sleep-per-packet, which accumulates the
* scheduler's error and drifts the achieved rate below target over a multi-second run.
*
* Nothing comes back the server counts and stays silent so the result here is only the
* send side. The measurement is the gap between this and the server's tally.
*/
fun sendThroughput(durationMs: Long, kbps: Int, sizeBytes: Int = 1200): Sent {
val size = sizeBytes.coerceIn(Wire.HEADER_SIZE + 16, 1472)
val payload = ByteArray(size - Wire.HEADER_SIZE)
val perPacketNs = (size.toLong() * 8 * 1_000_000 / kbps.coerceAtLeast(1)).coerceAtLeast(1_000)
val start = System.nanoTime()
val deadline = start + durationMs * 1_000_000
var next = start
var packets = 0
var bytes = 0L
while (System.nanoTime() < deadline) {
val pkt = Wire.build(Wire.TYPE_THROUGHPUT_UP, prefix, ++seq, nowNs(), key, payload)
try {
socket.send(DatagramPacket(pkt, pkt.size, server))
} catch (e: java.io.IOException) {
// A local send failure is our condition, not the path's. Stop and report what
// actually left, rather than counting the remainder as loss on the network.
break
}
packets++
bytes += pkt.size
next += perPacketNs
val sleepNs = next - System.nanoTime()
if (sleepNs > 0) Thread.sleep(sleepNs / 1_000_000, (sleepNs % 1_000_000).toInt())
}
val elapsedMs = (System.nanoTime() - start) / 1_000_000
return Sent(packets, bytes, elapsedMs, if (elapsedMs > 0) (bytes * 8 / elapsedMs).toInt() else 0)
}
/** What one upstream run put on the wire locally. */
data class Sent(val packets: Int, val bytes: Long, val durationMs: Long, val kbps: Int)
// ---- upstream trains (spec §3.2, types 0x03-0x05) --------------------------------------
/** One TRAIN_DATA packet as sent: its wire seq, local tx time and size. */
data class TrainPacket(val seq: Int, val tTxNs: Long, val sizeBytes: Int)
/** One row of the server's received view. 255 in ttl/dscp/ecn means "not observed". */
data class TrainRow(
val seq: Int, val tRxNs: Long, val sizeBytes: Int,
val ttl: Int, val dscp: Int, val ecn: Int,
)
/**
* The server's account of one train. [received] counts every packet that arrived, buffered
* or not; [truncated] mirrors the wire flag (rows beyond the server's cap were counted but
* not kept). [partsExpected]/[partsReceived] make a lossy report path visible instead of
* letting missing rows masquerade as train loss.
*/
data class TrainReport(
val trainId: Int,
val received: Int,
val truncated: Boolean,
val rows: List<TrainRow>,
val partsExpected: Int,
val partsReceived: Int,
)
/**
* Sends one paced upstream train. Nothing comes back per packet by design; pair with
* [trainReport] to learn what arrived. The absolute schedule (not sleep-per-packet) is the
* same anti-drift choice as [sendThroughput].
*/
fun sendTrain(trainId: Int, count: Int, sizeBytes: Int = 200, interPacketMs: Long = 5): List<TrainPacket> {
val size = sizeBytes.coerceIn(Wire.HEADER_SIZE + 4, 1472)
val out = ArrayList<TrainPacket>(count)
val start = System.nanoTime()
var next = start
for (i in 0 until count) {
val payload = ByteArray(size - Wire.HEADER_SIZE)
payload[0] = (trainId ushr 24).toByte(); payload[1] = (trainId ushr 16).toByte()
payload[2] = (trainId ushr 8).toByte(); payload[3] = trainId.toByte()
val tTx = nowNs()
val pkt = Wire.build(Wire.TYPE_TRAIN_DATA, prefix, ++seq, tTx, key, payload)
try {
socket.send(DatagramPacket(pkt, pkt.size, server))
} catch (e: java.io.IOException) {
// A local send failure is our condition, not the path's: report what actually
// left rather than letting the server's report read as loss.
break
}
out.add(TrainPacket(seq, tTx, pkt.size))
next += interPacketMs * 1_000_000
val sleepNs = next - System.nanoTime()
if (sleepNs > 0) Thread.sleep(sleepNs / 1_000_000, (sleepNs % 1_000_000).toInt())
}
return out
}
/**
* Fetches the server's received view of a train (one REPORT_REQ, N REPORT datagrams).
*
* Returns null when no report arrives at all indistinguishable between "report lost" and
* "server predates trains", and the caller must say so rather than choose. Missing parts of
* a multi-part report are tolerated and visible via partsReceived < partsExpected.
*/
fun trainReport(trainId: Int, timeoutMs: Long = 3_000): TrainReport? {
val req = ByteArray(4)
req[0] = (trainId ushr 24).toByte(); req[1] = (trainId ushr 16).toByte()
req[2] = (trainId ushr 8).toByte(); req[3] = trainId.toByte()
val pkt = Wire.build(Wire.TYPE_TRAIN_REPORT_REQ, prefix, ++seq, nowNs(), key, req)
socket.send(DatagramPacket(pkt, pkt.size, server))
var received = 0
var truncated = false
var partsExpected = -1
val seenParts = HashSet<Int>()
val rows = ArrayList<TrainRow>()
val deadline = System.nanoTime() + timeoutMs * 1_000_000
val buf = ByteArray(2048)
val prevTimeout = socket.soTimeout
try {
while (partsExpected < 0 || seenParts.size < partsExpected) {
val remainMs = ((deadline - System.nanoTime()) / 1_000_000).toInt()
if (remainMs <= 0) break
socket.soTimeout = remainMs.coerceAtMost(1000)
val dp = DatagramPacket(buf, buf.size)
try {
socket.receive(dp)
} catch (e: java.net.SocketTimeoutException) {
continue
}
val p = Wire.parseVerified(buf, dp.length, key) ?: continue
if (p.type != Wire.TYPE_TRAIN_REPORT) continue
val part = parseReportPart(p.payload, trainId) ?: continue
if (!seenParts.add(part.part)) continue
received = part.received
truncated = truncated || part.truncated
partsExpected = part.parts
rows.addAll(part.rows)
}
} finally {
socket.soTimeout = prevTimeout
}
if (seenParts.isEmpty()) return null
rows.sortBy { it.seq }
return TrainReport(trainId, received, truncated, rows, partsExpected, seenParts.size)
}
private class ReportPart(
val part: Int, val parts: Int, val received: Int,
val truncated: Boolean, val rows: List<TrainRow>,
)
/** Mirrors the server's columnar layout (dataplane/train.go buildTrainReport). */
private fun parseReportPart(b: ByteArray, wantId: Int): ReportPart? {
if (b.size < 16) return null
fun u16(off: Int) = ((b[off].toInt() and 0xFF) shl 8) or (b[off + 1].toInt() and 0xFF)
fun u32(off: Int) = ((b[off].toLong() and 0xFF) shl 24) or ((b[off + 1].toLong() and 0xFF) shl 16) or
((b[off + 2].toLong() and 0xFF) shl 8) or (b[off + 3].toLong() and 0xFF)
if (u32(0).toInt() != wantId) return null
val received = u32(4).toInt()
val part = u16(8)
val parts = u16(10)
val truncated = (b[12].toInt() and 0x01) != 0
val n = u16(14)
if (b.size < 16 + n * 17) return null
val rows = ArrayList<TrainRow>(n)
var off = 16
val seqs = IntArray(n) { u32(off + it * 4).toInt() }; off += n * 4
val tRx = LongArray(n) {
var v = 0L
for (j in 0 until 8) v = (v shl 8) or (b[off + it * 8 + j].toLong() and 0xFF)
v
}; off += n * 8
val sizes = IntArray(n) { u16(off + it * 2) }; off += n * 2
val ttls = IntArray(n) { b[off + it].toInt() and 0xFF }; off += n
val dscps = IntArray(n) { b[off + it].toInt() and 0xFF }; off += n
val ecns = IntArray(n) { b[off + it].toInt() and 0xFF }
for (i in 0 until n) {
rows.add(TrainRow(seqs[i], tRx[i], sizes[i], ttls[i], dscps[i], ecns[i]))
}
return ReportPart(part, parts, received, truncated, rows)
}
/** One packet received from the server, with the wire size actually delivered. */ /** One packet received from the server, with the wire size actually delivered. */
data class Received(val type: Int, val seq: Int, val sizeBytes: Int, val tRxNs: Long) data class Received(val type: Int, val seq: Int, val sizeBytes: Int, val tRxNs: Long)
@@ -112,5 +311,5 @@ class ProbeSession(
override fun close() = socket.close() override fun close() = socket.close()
data class EchoResult(val rttMs: Double, val observation: Observation?) data class EchoResult(val rttMs: Double, val observation: Observation?, val seq: Int = 0)
} }
@@ -23,6 +23,14 @@ object Wire {
const val TYPE_ECHO_REQ: Int = 0x01 const val TYPE_ECHO_REQ: Int = 0x01
const val TYPE_ECHO_RESP: Int = 0x02 const val TYPE_ECHO_RESP: Int = 0x02
/**
* Upstream train (spec §3.2): DATA is deliberately unanswered a per-packet reply would
* double the traffic and drag the return path into a measurement of the outbound one. The
* server's received view comes back afterwards via REPORT_REQ one or more REPORTs.
*/
const val TYPE_TRAIN_DATA: Int = 0x03
const val TYPE_TRAIN_REPORT_REQ: Int = 0x04
const val TYPE_TRAIN_REPORT: Int = 0x05
const val TYPE_TIMESYNC_REQ: Int = 0x07 const val TYPE_TIMESYNC_REQ: Int = 0x07
const val TYPE_TIMESYNC_RSP: Int = 0x08 const val TYPE_TIMESYNC_RSP: Int = 0x08
const val TYPE_MTU_PROBE: Int = 0x09 const val TYPE_MTU_PROBE: Int = 0x09
@@ -32,6 +40,22 @@ object Wire {
const val TYPE_DOWNTRAIN_DATA: Int = 0x06 const val TYPE_DOWNTRAIN_DATA: Int = 0x06
const val TYPE_BIG_SEND: Int = 0x0C const val TYPE_BIG_SEND: Int = 0x0C
/**
* A datagram the server deliberately fragmented. Its arrival IS the measurement: it can only
* be delivered if every fragment survived the path and the local stack reassembled them.
*/
const val TYPE_FRAG_DATA: Int = 0x0D
/** One packet of a sustained-rate downstream run. */
const val TYPE_THROUGHPUT_DATA: Int = 0x0E
/**
* One packet of a client-driven upstream run. The server counts it and does not answer:
* a reply would double the traffic and drag the return path into a measurement that is
* specifically about the outbound one.
*/
const val TYPE_THROUGHPUT_UP: Int = 0x0F
/** The 8-byte on-the-wire prefix = first 16 hex chars of the session id, decoded. */ /** The 8-byte on-the-wire prefix = first 16 hex chars of the session id, decoded. */
fun wirePrefix(sessionId: String): ByteArray { fun wirePrefix(sessionId: String): ByteArray {
require(sessionId.length >= 16) { "session id too short" } require(sessionId.length >= 16) { "session id too short" }
@@ -0,0 +1,100 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.protocol
import java.security.SecureRandom
import kotlin.test.Test
import kotlin.test.assertEquals
import kotlin.test.assertFailsWith
import kotlin.test.assertNotEquals
import kotlin.test.assertTrue
class OidcLoginTest {
private val auth = AuthInfo(
enabled = true,
issuer = "https://id.example.net/application/o/echolot-app/",
clientId = "the-client",
redirectUri = "echolot://auth",
scopes = "openid profile email",
authorizationEndpoint = "https://id.example.net/application/o/authorize/",
tokenEndpoint = "https://id.example.net/application/o/token/",
)
@Test
fun theAuthorizationUrlCarriesEverythingTheIdPNeeds() {
val p = OidcLogin.begin(auth)
val url = p.authorizationUrl
assertTrue(url.startsWith(auth.authorizationEndpoint + "?"), url)
for (part in listOf(
"response_type=code",
"client_id=the-client",
"redirect_uri=echolot%3A%2F%2Fauth",
"code_challenge_method=S256",
"scope=openid+profile+email",
)) {
assertTrue(url.contains(part), "missing $part in $url")
}
assertTrue(url.contains("code_challenge="), url)
// The verifier itself must never appear in the URL — that is the entire point of PKCE.
assertTrue(!url.contains(p.verifier), "the code verifier leaked into the authorize URL")
}
// Two sign-ins must not share a verifier or state, or one intercepted flow compromises the next.
@Test
fun everySignInGetsFreshSecrets() {
val a = OidcLogin.begin(auth, SecureRandom())
val b = OidcLogin.begin(auth, SecureRandom())
assertNotEquals(a.verifier, b.verifier)
assertNotEquals(a.state, b.state)
assertTrue(a.verifier.length >= 43, "verifier is shorter than RFC 7636 allows")
}
@Test
fun parsesTheRedirectTheBrowserHandsBack() {
val cb = OidcLogin.parseCallback("echolot://auth?code=abc123&state=xyz")
assertEquals("abc123", cb.code)
assertEquals("xyz", cb.state)
}
@Test
fun parsesAnErrorRedirect() {
val cb = OidcLogin.parseCallback("echolot://auth?error=access_denied&error_description=User%20said%20no")
assertEquals("access_denied", cb.error?.substringBefore(":"))
assertTrue(cb.code == null)
}
// A callback whose state does not match is how an attacker gets someone to complete *their*
// sign-in. It must be refused before the code is spent, without any network call.
@Test
fun aMismatchedStateIsRefusedBeforeTheCodeIsSpent() {
val p = OidcLogin.begin(auth)
val e = assertFailsWith<OidcLogin.LoginFailed> {
OidcLogin.complete(auth, p, "echolot://auth?code=stolen&state=not-ours")
}
assertTrue(e.message!!.contains("did not start on this device"), e.message!!)
}
@Test
fun aMissingStateIsRefused() {
val p = OidcLogin.begin(auth)
assertFailsWith<OidcLogin.LoginFailed> {
OidcLogin.complete(auth, p, "echolot://auth?code=abc")
}
}
@Test
fun anErrorRedirectSurfacesTheReason() {
val p = OidcLogin.begin(auth)
val e = assertFailsWith<OidcLogin.LoginFailed> {
OidcLogin.complete(auth, p, "echolot://auth?error=access_denied&state=${p.state}")
}
assertTrue(e.message!!.contains("access_denied"))
}
@Test
fun refusesToStartWhenTheServerHasNoIdentityProvider() {
assertFailsWith<IllegalArgumentException> { OidcLogin.begin(AuthInfo(enabled = false)) }
}
}
@@ -32,4 +32,8 @@ dependencies {
implementation(libs.shizuku.provider) implementation(libs.shizuku.provider)
implementation(libs.kotlinx.coroutines.android) implementation(libs.kotlinx.coroutines.android)
implementation(libs.kotlinx.serialization.json) implementation(libs.kotlinx.serialization.json)
// JVM unit tests for the pure dump parsers (DumpParsers.kt) against the archived
// vendor fixtures — no device, no Android runtime.
testImplementation(libs.kotlin.test.junit)
testImplementation(libs.junit4)
} }
@@ -0,0 +1,174 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.shizuku
/**
* Parsed view of one IPv6 default route from `ip -6 route show table all`.
*
* [table] stays a string: Android's per-network route tables use ids past Int range (the Lenovo
* TB330FU prints `table 1000000015`), and `table local` is not a number at all parsing to a
* numeric type either overflows or silently drops rows, and the id is only ever compared, never
* computed with.
*/
data class V6DefaultRoute(
/** Link-local address of the advertising router; null for gateway-less defaults (dummy0). */
val gateway: String?,
val dev: String,
val table: String?, // null = main table (`ip` omits the token there)
val proto: String?, // "ra" marks a route installed from a Router Advertisement
val metric: Long?,
/** Remaining RA route lifetime (`expires NNNsec`); null when the route does not age out. */
val expiresSec: Long?,
)
/** One `ip neigh show` row. [lladdr] is null for FAILED/INCOMPLETE entries the kernel tried to
* resolve and has nothing, which is itself signal. */
data class NeighborEntry(
val ip: String,
val dev: String?,
val lladdr: String?,
val state: String?, // REACHABLE/STALE/FAILED/... — kept verbatim, the kernel's vocabulary
val router: Boolean,
)
/** A NEIGH transition seen inside the `ip monitor` window. */
data class NeighborEvent(val entry: NeighborEntry, val deleted: Boolean)
/**
* Pure-string parsers for the shell battery's `ip` command outputs. No Android imports on
* purpose: these run (and are unit-tested) on the JVM against the real vendor dumps archived
* from the prober, which is the only way to catch a vendor format drift before it ships.
*
* All parsers degrade to an empty result on missing or unrecognized input the battery's
* captures are best-effort (the Lenovo's `ip monitor` times out under newProcess, the
* UserService path prepends a stray `uid=2000` line, evidence strings are trimmed mid-line
* at 1200 chars), so an exception here would turn a degraded capture into a lost test.
*/
object DumpParsers {
/** True when [raw] is real command output rather than an executor error sentinel. */
fun captureUsable(raw: String?): Boolean {
if (raw.isNullOrBlank()) return false
val t = raw.trimStart()
return !t.startsWith("SHIZUKU_") && !t.startsWith("EXEC_") && !t.startsWith("NEWPROCESS_")
}
/** Extracts every `default …` route from `ip -6 route show table all` output. */
fun parseV6DefaultRoutes(raw: String?): List<V6DefaultRoute> {
if (!captureUsable(raw)) return emptyList()
val routes = ArrayList<V6DefaultRoute>()
for (line in raw!!.lineSequence()) {
val tok = line.trim().split(WS)
if (tok.firstOrNull() != "default") continue
var gateway: String? = null; var dev: String? = null; var table: String? = null
var proto: String? = null; var metric: Long? = null; var expires: Long? = null
var i = 1
while (i < tok.size - 1) {
when (tok[i]) {
"via" -> gateway = tok[i + 1]
"dev" -> dev = tok[i + 1]
"table" -> table = tok[i + 1]
"proto" -> proto = tok[i + 1]
"metric" -> metric = tok[i + 1].toLongOrNull()
// `expires 1269sec` — the unit is glued to the number.
"expires" -> expires = tok[i + 1].removeSuffix("sec").toLongOrNull()
}
i++
}
// A default route without a device is not something `ip` prints; treat it as a
// truncated/garbled line rather than fabricating a partial route.
if (dev != null) routes.add(V6DefaultRoute(gateway, dev, table, proto, metric, expires))
}
return routes
}
/**
* Maps interface name link-layer address from `ip addr show`. Only `link/ether` counts:
* loopback/ipip/gre pseudo-addresses are not identities, and the RA-source cross-reference
* this feeds compares Ethernet MACs.
*/
fun parseInterfaceMacs(raw: String?): Map<String, String> {
if (!captureUsable(raw)) return emptyMap()
val macs = LinkedHashMap<String, String>()
var current: String? = null
for (line in raw!!.lineSequence()) {
val header = STANZA_HEADER.find(line)
if (header != null) {
// "5: tunl0@NONE:" — the name is the part before an optional @suffix.
current = header.groupValues[1].substringBefore('@')
continue
}
val dev = current ?: continue
val tok = line.trim().split(WS)
if (tok.size >= 2 && tok[0] == "link/ether" && MAC.matches(tok[1])) {
macs.putIfAbsent(dev, tok[1])
}
}
return macs
}
/** Parses `ip neigh show` output into entries; non-neighbor lines (uid noise) are skipped. */
fun parseNeighbors(raw: String?): List<NeighborEntry> {
if (!captureUsable(raw)) return emptyList()
return raw!!.lineSequence()
.mapNotNull { parseNeighborTokens(it.trim().split(WS)) }
.toList()
}
/**
* Extracts NEIGH transitions from an `ip monitor all` capture. Each event line carries a
* `[NEIGH]` label (other families ROUTE, ADDR, LINK are ignored) and deletions are
* printed as `Deleted <entry>`. An empty result is normal: a quiet 5 s window sees nothing.
*/
fun parseNeighborEvents(raw: String?): List<NeighborEvent> {
if (!captureUsable(raw)) return emptyList()
val events = ArrayList<NeighborEvent>()
for (line in raw!!.lineSequence()) {
val m = MONITOR_LABEL.find(line.trim()) ?: continue
if (!m.groupValues[1].equals("NEIGH", ignoreCase = true)) continue
var rest = line.trim().removeRange(m.range).trim()
val deleted = rest.startsWith("Deleted ", ignoreCase = true)
if (deleted) rest = rest.substring("Deleted ".length)
parseNeighborTokens(rest.split(WS))?.let { events.add(NeighborEvent(it, deleted)) }
}
return events
}
/**
* (ip lladdr) for every neighbor that has one. This is the comparison surface the future
* gateway-MAC-change finding diffs across runs, so it is computed here in the tested,
* pure layer rather than re-derived from JSON by each consumer.
*/
fun lladdrByIp(neighbors: List<NeighborEntry>): Map<String, String> =
neighbors.mapNotNull { n -> n.lladdr?.let { n.ip to it } }.toMap()
/** One neighbor row: `<ip> dev <if> [lladdr <mac>] [router] [proxy] <STATE>`. */
private fun parseNeighborTokens(tok: List<String>): NeighborEntry? {
val ip = tok.firstOrNull() ?: return null
// The first token must look like an address — this is what drops the UserService path's
// stray "uid=2000" line and any grep noise without needing to know every noise shape.
if (!IP_LIKE.matches(ip) || (!ip.contains('.') && !ip.contains(':'))) return null
var dev: String? = null; var lladdr: String? = null; var state: String? = null
var router = false
var i = 1
while (i < tok.size) {
when (tok[i]) {
"dev" -> { dev = tok.getOrNull(i + 1); i++ }
"lladdr" -> { lladdr = tok.getOrNull(i + 1); i++ }
"router" -> router = true
"proxy" -> {} // recorded nowhere: proxy entries have no bearing on ARP watching
else -> if (STATE.matches(tok[i])) state = tok[i]
}
i++
}
return NeighborEntry(ip, dev, lladdr, state, router)
}
private val WS = Regex("\\s+")
private val STANZA_HEADER = Regex("^\\d+:\\s+([^:\\s]+):")
private val MAC = Regex("^[0-9a-fA-F]{2}(:[0-9a-fA-F]{2}){5}$")
private val IP_LIKE = Regex("^[0-9a-fA-F:.]+(%[\\w-]+)?$")
private val STATE = Regex("^(REACHABLE|STALE|DELAY|PROBE|FAILED|INCOMPLETE|PERMANENT|NOARP|NONE)$")
private val MONITOR_LABEL = Regex("^\\[(\\w+)]")
}
@@ -98,20 +98,32 @@ object ShizukuAvailability {
.addFlags(android.content.Intent.FLAG_ACTIVITY_NEW_TASK) .addFlags(android.content.Intent.FLAG_ACTIVITY_NEW_TASK)
/** /**
* Reports the state now and on every binder transition. Returns a function that removes the * Reports the state now, on every binder transition, and when a permission request is
* listeners again (call it from onCleared). * answered. Returns a function that removes the listeners again (call it from onCleared).
*
* The permission listener matters as much as the binder ones: granting permission does not
* make the binder arrive or die, so without it the banner still read "running but not
* authorised" after the user had just authorised it — the one moment they are looking for
* confirmation that it worked.
*
* It is still not sufficient on its own. Permission can be granted inside Shizuku's own app,
* where nothing calls back into this process at all, so callers should re-check on resume as
* well; see [current].
*/ */
fun observe(context: Context, onChange: (State) -> Unit): () -> Unit { fun observe(context: Context, onChange: (State) -> Unit): () -> Unit {
val app = context.applicationContext val app = context.applicationContext
val received = Shizuku.OnBinderReceivedListener { onChange(current(app)) } val received = Shizuku.OnBinderReceivedListener { onChange(current(app)) }
val dead = Shizuku.OnBinderDeadListener { onChange(current(app)) } val dead = Shizuku.OnBinderDeadListener { onChange(current(app)) }
val permission = Shizuku.OnRequestPermissionResultListener { _, _ -> onChange(current(app)) }
// "Sticky" fires immediately if the binder already arrived before we registered. // "Sticky" fires immediately if the binder already arrived before we registered.
runCatching { Shizuku.addBinderReceivedListenerSticky(received) } runCatching { Shizuku.addBinderReceivedListenerSticky(received) }
runCatching { Shizuku.addBinderDeadListener(dead) } runCatching { Shizuku.addBinderDeadListener(dead) }
runCatching { Shizuku.addRequestPermissionResultListener(permission) }
onChange(current(app)) onChange(current(app))
return { return {
runCatching { Shizuku.removeBinderReceivedListener(received) } runCatching { Shizuku.removeBinderReceivedListener(received) }
runCatching { Shizuku.removeBinderDeadListener(dead) } runCatching { Shizuku.removeBinderDeadListener(dead) }
runCatching { Shizuku.removeRequestPermissionResultListener(permission) }
} }
} }
} }
@@ -10,15 +10,24 @@ import app.echo_lot.measurement.TestType
import app.echo_lot.measurement.Tier import app.echo_lot.measurement.Tier
import kotlinx.coroutines.Dispatchers import kotlinx.coroutines.Dispatchers
import kotlinx.coroutines.withContext import kotlinx.coroutines.withContext
import kotlinx.serialization.json.JsonObject
import kotlinx.serialization.json.buildJsonObject import kotlinx.serialization.json.buildJsonObject
import kotlinx.serialization.json.put import kotlinx.serialization.json.put
import kotlinx.serialization.json.putJsonArray
import kotlinx.serialization.json.putJsonObject
import kotlinx.serialization.json.addJsonObject
/** /**
* The Shizuku shell-tier probe: runs the privileged command battery (neighbor table, RA routes * The Shizuku shell-tier probe: runs the privileged command battery (neighbor table, RA routes
* with lifetimes, netlink monitor, IpClient DHCP logs, wifi dump) that the app UID cannot, and * with lifetimes, netlink monitor, IpClient DHCP logs, wifi dump) that the app UID cannot, and
* captures the real per-device dump formats the production parsers must handle. Emitted as a * captures the real per-device dump formats the production parsers must handle. Emits three
* shizuku-tier `link.ip_monitor` test (the representative shell-tier link test); `exec_path` * shizuku-tier tests from the one battery:
* records whether the UserService or the newProcess fallback carried it. * - `link.ip_monitor` the raw captures (the shell tier's ground truth), `exec_path` records
* whether the UserService or the newProcess fallback carried it;
* - `link.ra_source` parsed from the v6 route table + `ip addr`: who advertises IPv6 here;
* - `sec.arp_watch` parsed from the neighbor table + monitor window: (ip lladdr) pairs for
* gateway-MAC-change detection.
* The battery runs once; the derived tests parse its captures, so they share its time window.
*/ */
class ShizukuProbe { class ShizukuProbe {
val type = TestType.LINK_IP_MONITOR val type = TestType.LINK_IP_MONITOR
@@ -34,24 +43,34 @@ class ShizukuProbe {
"wifi_dump" to "dumpsys wifi 2>/dev/null | grep -iA1 -m 20 -e 'mDhcpResults' -e 'Gateway' -e 'DNS' || true", "wifi_dump" to "dumpsys wifi 2>/dev/null | grep -iA1 -m 20 -e 'mDhcpResults' -e 'Gateway' -e 'DNS' || true",
) )
/** Runs the battery and returns a Test. [uuid]/[monoNs] come from the run's id/clock source. */ /**
suspend fun run(context: Context, uuid: () -> String, monoNs: () -> Long): Test = withContext(Dispatchers.IO) { * Runs the battery and returns the three tests, battery first. [uuid]/[monoNs] come from the
val id = uuid() * run's id/clock source. When the shell tier is unavailable all three come back UNSUPPORTED
* one silent test would leave the other two types missing from the document, which reads as
* "never attempted" rather than "tier absent".
*/
suspend fun run(context: Context, uuid: () -> String, monoNs: () -> Long): List<Test> = withContext(Dispatchers.IO) {
val started = monoNs() val started = monoNs()
val runner = ShizukuRunner(context) val runner = ShizukuRunner(context)
val st = runner.status() val st = runner.status()
fun envelope(status: TestStatus, evidence: kotlinx.serialization.json.JsonObject, metrics: kotlinx.serialization.json.JsonObject? = null) = fun envelope(type: String, status: TestStatus, evidence: JsonObject, metrics: JsonObject? = null) =
Test(id = id, type = type, tier = tier, startedMonoNs = started, endedMonoNs = monoNs(), Test(id = uuid(), type = type, tier = tier, startedMonoNs = started, endedMonoNs = monoNs(),
status = status, evidence = evidence, metrics = metrics) status = status, evidence = evidence, metrics = metrics)
fun allUnsupported(evidence: JsonObject) = listOf(
envelope(TestType.LINK_IP_MONITOR, TestStatus.UNSUPPORTED, evidence),
envelope(TestType.LINK_RA_SOURCE, TestStatus.UNSUPPORTED, evidence),
envelope(TestType.SEC_ARP_WATCH, TestStatus.UNSUPPORTED, evidence),
)
if (!st.binderAlive) { if (!st.binderAlive) {
return@withContext envelope(TestStatus.UNSUPPORTED, buildJsonObject { return@withContext allUnsupported(buildJsonObject {
put("binder_alive", false); put("detail", "Shizuku not running") put("binder_alive", false); put("detail", "Shizuku not running")
}) })
} }
if (!st.permissionGranted && !runner.requestPermission()) { if (!st.permissionGranted && !runner.requestPermission()) {
return@withContext envelope(TestStatus.UNSUPPORTED, buildJsonObject { return@withContext allUnsupported(buildJsonObject {
put("binder_alive", true); put("permission", false) put("binder_alive", true); put("permission", false)
}) })
} }
@@ -77,6 +96,125 @@ class ShizukuProbe {
ok >= 1 -> TestStatus.PARTIAL ok >= 1 -> TestStatus.PARTIAL
else -> TestStatus.FAILED else -> TestStatus.FAILED
} }
envelope(status, evidence, metrics) listOf(
envelope(type, status, evidence, metrics),
raSourceTest(batch, ::envelope),
arpWatchTest(batch, ::envelope),
)
}
/** `link.ra_source` from the battery's `ip -6 route` / `ip addr` / `ip neigh` captures. */
private fun raSourceTest(
batch: ShizukuRunner.BatchResult,
envelope: (String, TestStatus, JsonObject, JsonObject?) -> Test,
): Test {
val routeRaw = batch.results["ip6_route"]
if (!DumpParsers.captureUsable(routeRaw)) {
// The source command failed (executor sentinel or empty) — say so instead of
// presenting "no default routes" as a measurement of the network.
return envelope(TestType.LINK_RA_SOURCE, TestStatus.SKIPPED, buildJsonObject {
put("exec_path", batch.execPath)
put("reason", "ip -6 route capture unavailable: ${(routeRaw ?: "absent").take(80)}")
}, null)
}
val routes = DumpParsers.parseV6DefaultRoutes(routeRaw)
val macs = DumpParsers.parseInterfaceMacs(batch.results["ip_addr"])
// The RA sender's own identity: its link-local gateway address resolved through the
// neighbor table gives the router's MAC, which is what survives address renumbering.
val neighMacs = DumpParsers.lladdrByIp(DumpParsers.parseNeighbors(batch.results["ip_neigh"]))
val evidence = buildJsonObject {
put("exec_path", batch.execPath)
putJsonArray("default_routes") {
for (r in routes) addJsonObject {
r.gateway?.let { put("gateway", it) }
put("dev", r.dev)
r.table?.let { put("table", it) }
r.proto?.let { put("proto", it) }
r.metric?.let { put("metric", it) }
r.expiresSec?.let { put("expires_sec", it) }
r.gateway?.let { gw -> neighMacs[gw]?.let { put("gateway_lladdr", it) } }
}
}
putJsonObject("interface_mac") {
// Only interfaces that actually carry a default route: the full MAC inventory
// belongs to the raw capture, not to this test's claim.
for (dev in routes.map { it.dev }.distinct()) macs[dev]?.let { put(dev, it) }
}
}
val metrics = buildJsonObject {
put("routes_total", routes.size)
put("routes_ra", routes.count { it.proto == "ra" })
}
val status = when {
routes.any { it.proto == "ra" } -> TestStatus.OK
// Routes parsed but none RA-installed, or a capture we couldn't parse a single
// default from: could be a genuinely RA-less link, could be vendor format drift —
// PARTIAL keeps it visible either way instead of quietly claiming success.
else -> TestStatus.PARTIAL
}
return envelope(TestType.LINK_RA_SOURCE, status, evidence, metrics)
}
/** `sec.arp_watch` from the battery's `ip neigh` snapshot + `ip monitor` window. */
private fun arpWatchTest(
batch: ShizukuRunner.BatchResult,
envelope: (String, TestStatus, JsonObject, JsonObject?) -> Test,
): Test {
val neighRaw = batch.results["ip_neigh"]
val monitorRaw = batch.results["ip_monitor"]
val neighUsable = DumpParsers.captureUsable(neighRaw)
// The monitor window is best-effort (EXEC_TIMEOUT under newProcess on the Lenovo); the
// snapshot alone still yields the (ip → lladdr) pairs the MAC-change finding diffs.
val monitorRan = DumpParsers.captureUsable(monitorRaw)
if (!neighUsable && !monitorRan) {
return envelope(TestType.SEC_ARP_WATCH, TestStatus.SKIPPED, buildJsonObject {
put("exec_path", batch.execPath)
put("reason", "ip neigh capture unavailable: ${(neighRaw ?: "absent").take(80)}")
}, null)
}
val neighbors = DumpParsers.parseNeighbors(neighRaw)
val events = DumpParsers.parseNeighborEvents(monitorRaw)
val evidence = buildJsonObject {
put("exec_path", batch.execPath)
putJsonArray("neighbors") {
for (n in neighbors) addJsonObject {
put("ip", n.ip)
n.dev?.let { put("dev", it) }
n.lladdr?.let { put("lladdr", it) }
n.state?.let { put("state", it) }
if (n.router) put("router", true)
}
}
// The comparison surface, precomputed: a MAC-change finding diffs this map between
// runs without re-walking the neighbor array.
putJsonObject("lladdr_by_ip") {
for ((ip, mac) in DumpParsers.lladdrByIp(neighbors)) put(ip, mac)
}
put("monitor_ran", monitorRan)
if (!monitorRan) put("monitor_reason", (monitorRaw ?: "absent").take(80))
putJsonArray("monitor_events") {
for (e in events) addJsonObject {
put("ip", e.entry.ip)
e.entry.dev?.let { put("dev", it) }
e.entry.lladdr?.let { put("lladdr", it) }
e.entry.state?.let { put("state", it) }
if (e.deleted) put("deleted", true)
}
}
}
val metrics = buildJsonObject {
put("neighbors_total", neighbors.size)
put("neighbors_with_lladdr", neighbors.count { it.lladdr != null })
put("monitor_events", events.size)
}
val status = when {
neighbors.isNotEmpty() -> TestStatus.OK
// A snapshot that parsed to nothing (or a monitor-only capture) is thin evidence:
// usable command output with zero entries is unusual enough to flag, not to fail.
else -> TestStatus.PARTIAL
}
return envelope(TestType.SEC_ARP_WATCH, status, evidence, metrics)
} }
} }
@@ -0,0 +1,303 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package app.echo_lot.shizuku
import kotlin.test.Test
import kotlin.test.assertEquals
import kotlin.test.assertFalse
import kotlin.test.assertNull
import kotlin.test.assertTrue
/**
* The fixtures below are the REAL shell-battery captures from the two archived prober reports
* (echolot-prober/reports/CPH2747-android16-sdk36-build5.json OnePlus 15, UserService path;
* TB330FU-android15-sdk35-build5.json Lenovo TB330FU, newProcess fallback), trimmed to the
* relevant lines but otherwise verbatim. That includes their warts on purpose: the UserService
* path's stray `uid=2000` first line, the 1200-char evidence trim cutting the last line mid-word,
* the Lenovo's 10-digit route table ids and its `EXEC_TIMEOUT(newProcess)` monitor sentinel.
* A parser that only survives clean textbook output has not been tested.
*/
class DumpParsersTest {
// ---- OnePlus 15 (CPH2747, Android 16) — UserService exec path ----
private val onePlusIp6Route = """
uid=2000
fe80::/64 dev wlan0 table 1028 proto kernel metric 256 pref medium
fe80::/64 dev wlan0 table 1028 proto static metric 1024 pref medium
default via fe80::7a9a:18ff:fe54:b8f9 dev wlan0 table 1028 proto ra metric 1024 expires 1269sec pref medium
fe80::/64 dev vgate0 table 1031 proto kernel metric 256 pref medium
2001:4bb8:417:bd78::/64 dev rmnet_data4 table 1032 proto kernel metric 256 pref medium
2001:4bb8:417:bd78::/64 dev rmnet_data4 table 1032 proto static metric 1024 pref medium
fe80::/64 dev rmnet_data4 table 1032 proto kernel metric 256 pref medium
default via fe80::246f:12be:21ef:1b54 dev rmnet_data4 table 1032 proto ra metric 1024 expires 64373sec hoplimit 255 pref medium
2001:4bb8:2fb:fe4c::/64 dev rmnet_data2 table 1000000022 proto static metric 1024 pref medium
fe80::/64 dev wlan0 table 1000000028 proto static metric 1024 pref medium
2001:4bb8:417:bd78::/64 dev rmnet_data4 table 1000000032 proto static metric 1024 pref medium
fe80::/64 dev dummy0 table 1002 proto kernel metric 256 pref medium
default dev dummy0 table 1002 proto static metric 1024 pref medium
fe80::/64 dev ifb0 table 1003 proto kernel metric 256 pref medium
fe80::/64 dev ifb1 table 1004 proto kerne
""".trimIndent()
private val onePlusIpNeigh = """
uid=2000
10.13.102.111 dev wlan0 FAILED
10.13.102.50 dev wlan0 lladdr 50:57:9c:4f:7a:3c STALE
10.13.102.31 dev wlan0 lladdr 98:5f:d3:f6:f1:75 STALE
10.13.102.116 dev wlan0 lladdr 0c:08:b4:03:68:0e STALE
10.13.102.120 dev wlan0 lladdr 0e:d8:14:58:6c:8b STALE
10.13.102.5 dev wlan0 lladdr 90:09:d0:1a:83:e4 STALE
10.13.102.1 dev wlan0 lladdr 78:9a:18:54:b8:f9 REACHABLE
10.13.102.21 dev wlan0 lladdr c8:7f:54:01:94:7c STALE
fe80::babe:f4ff:febc:caf9 dev wlan0 lladdr b8:be:f4:bc:ca:f9 REACHABLE
fe80::7a9a:18ff:fe54:b8f9 dev wlan0 lladdr 78:9a:18:54:b8:f9 router STALE
fe80::babe:f4ff:febc:cacf dev wlan0 lladdr b8:be:f4:bc:ca:cf REACHABLE
fe80::babe:f4ff:fec2:bf14 dev wlan0 lladdr b8:be:f4:c2:bf:14 REACHABLE
""".trimIndent()
// The 1200-char trim cut this capture off before wlan0's stanza — so the real archived
// evidence has NO MAC for the interface that carries the default route. The parser must
// yield what is there and nothing else; the probe records the gap instead of inventing one.
private val onePlusIpAddr = """
uid=2000
1: lo: <LOOPBACK,UP,LOWER_UP> mtu 65536 qdisc noqueue state UNKNOWN group default qlen 1000
link/loopback 00:00:00:00:00:00 brd 00:00:00:00:00:00
inet 127.0.0.1/8 scope host lo
valid_lft forever preferred_lft forever
inet6 ::1/128 scope host
valid_lft forever preferred_lft forever
2: dummy0: <BROADCAST,NOARP,UP,LOWER_UP> mtu 1500 qdisc noqueue state UNKNOWN group default qlen 1000
link/ether be:3d:e2:93:78:b9 brd ff:ff:ff:ff:ff:ff
inet6 fe80::bc3d:e2ff:fe93:78b9/64 scope link
valid_lft forever preferred_lft forever
3: ifb0: <BROADCAST,NOARP,UP,LOWER_UP> mtu 1500 qdisc htb state UNKNOWN group default qlen 1000
link/ether ba:6e:46:b5:3d:bb brd ff:ff:ff:ff:ff:ff
4: ifb1: <BROADCAST,NOARP,UP,LOWER_UP> mtu 1500 qdisc htb state UNKNOWN group default qlen 1000
link/ether d6:2a:e2:f5:93:8f brd ff:ff:ff:ff:ff:ff
5: tunl0@NONE: <NOARP> mtu 1480 qdisc noop state DOWN group default qlen 1000
link/ipip 0.0.0.0 brd 0.0.0.0
6: gre0@NONE: <NO
""".trimIndent()
// A monitor window that ran but saw nothing: the capture is "usable", just empty of events.
private val onePlusIpMonitor = "uid=2000"
// ---- Lenovo TB330FU (Android 15) — newProcess fallback ----
private val lenovoIp6Route = """
fe80::/64 dev wlan0 table 1000000015 proto static metric 1024 pref medium
fe80::/64 dev dummy0 table 1002 proto kernel metric 256 pref medium
default dev dummy0 table 1002 proto static metric 1024 pref medium
fe80::/64 dev wlan0 table 1015 proto kernel metric 256 pref medium
fe80::/64 dev wlan0 table 1015 proto static metric 1024 pref medium
default via fe80::7a9a:18ff:fe54:b8f9 dev wlan0 table 1015 proto ra metric 1024 expires 1622sec pref medium
local ::1 dev lo table local proto kernel metric 0 pref medium
local fe80::416:b9ff:feac:5b65 dev wlan0 table local proto kernel metric 0 pref medium
local fe80::1450:43ff:feec:93c4 dev dummy0 table local proto kernel metric 0 pref medium
multicast ff00::/8 dev dummy0 table local proto kernel metric 256 pref medium
multicast ff00::/8 dev wlan0 table local proto kernel metric 256 pref medium
""".trimIndent()
private val lenovoIpNeigh = """
10.13.102.21 dev wlan0 lladdr c8:7f:54:01:94:7c STALE
10.13.102.64 dev wlan0 lladdr b8:be:f4:c2:bf:14 STALE
10.13.102.5 dev wlan0 lladdr 90:09:d0:1a:83:e4 STALE
10.13.102.1 dev wlan0 lladdr 78:9a:18:54:b8:f9 STALE
10.13.102.79 dev wlan0 lladdr 02:11:32:25:63:bb STALE
fe80::babe:f4ff:febc:cacf dev wlan0 lladdr b8:be:f4:bc:ca:cf STALE
fe80::7a9a:18ff:fe54:b8f9 dev wlan0 lladdr 78:9a:18:54:b8:f9 router STALE
fe80::babe:f4ff:fec2:bf14 dev wlan0 lladdr b8:be:f4:c2:bf:14 STALE
fe80::babe:f4ff:febc:caf9 dev wlan0 lladdr b8:be:f4:bc:ca:f9 STALE
""".trimIndent()
private val lenovoIpAddr = """
1: lo: <LOOPBACK,UP,LOWER_UP> mtu 65536 qdisc noqueue state UNKNOWN group default qlen 1000
link/loopback 00:00:00:00:00:00 brd 00:00:00:00:00:00
inet 127.0.0.1/8 scope host lo
valid_lft forever preferred_lft forever
2: dummy0: <BROADCAST,NOARP,UP,LOWER_UP> mtu 1500 qdisc noqueue state UNKNOWN group default qlen 1000
link/ether 16:50:43:ec:93:c4 brd ff:ff:ff:ff:ff:ff
inet6 fe80::1450:43ff:feec:93c4/64 scope link
valid_lft forever preferred_lft forever
3: ifb0: <BROADCAST,NOARP> mtu 1500 qdisc noop state DOWN group default qlen 32
link/ether f6:d4:d4:9b:51:9c brd ff:ff:ff:ff:ff:ff
4: ifb1: <BROADCAST,NOARP> mtu 1500 qdisc noop state DOWN group default qlen 32
link/ether fe:16:ea:60:a2:d1 brd ff:ff:ff:ff:ff:ff
5: tunl0@NONE: <NOARP> mtu 1480 qdisc noop state DOWN group default qlen 1000
link/ipip 0.0.0.0 brd 0.0.0.0
7: gretap0@NONE: <BROADCAST,MULTICAST> mtu 1462 qdisc noop state DOWN group default qlen 1000
link/ether 00:00:00:00:00:00 brd ff:ff:ff:ff:f
""".trimIndent()
// On the Lenovo the 5 s monitor window exceeds the newProcess exec timeout — the executor's
// sentinel is all we get, and the arp_watch test must still stand on the snapshot alone.
private val lenovoIpMonitor = "EXEC_TIMEOUT(newProcess)"
// ---- link.ra_source: v6 default routes ----
@Test
fun onePlusDefaultRoutesParsed() {
val routes = DumpParsers.parseV6DefaultRoutes(onePlusIp6Route)
assertEquals(3, routes.size)
val wlan = routes.single { it.dev == "wlan0" }
assertEquals("fe80::7a9a:18ff:fe54:b8f9", wlan.gateway)
assertEquals("1028", wlan.table)
assertEquals("ra", wlan.proto)
assertEquals(1024L, wlan.metric)
assertEquals(1269L, wlan.expiresSec)
// The cellular default: `hoplimit 255` sits between expires and pref and must not derail
// the token walk.
val rmnet = routes.single { it.dev == "rmnet_data4" }
assertEquals("fe80::246f:12be:21ef:1b54", rmnet.gateway)
assertEquals("1032", rmnet.table)
assertEquals(64373L, rmnet.expiresSec)
// Android's gateway-less dummy0 default is a real route; it is the proto that tells a
// consumer it is not an RA.
val dummy = routes.single { it.dev == "dummy0" }
assertNull(dummy.gateway)
assertEquals("static", dummy.proto)
assertNull(dummy.expiresSec)
}
@Test
fun lenovoNumberedTablesAllCaptured() {
val routes = DumpParsers.parseV6DefaultRoutes(lenovoIp6Route)
// Two default routes: the RA one in table 1015 and the dummy0 one in 1002. The 10-digit
// table 1000000015 and the `table local` rows carry no default and must neither appear
// nor break parsing.
assertEquals(setOf("1002", "1015"), routes.map { it.table }.toSet())
val ra = routes.single { it.proto == "ra" }
assertEquals("fe80::7a9a:18ff:fe54:b8f9", ra.gateway)
assertEquals("wlan0", ra.dev)
assertEquals("1015", ra.table)
assertEquals(1622L, ra.expiresSec)
}
@Test
fun tenDigitTableIdOnADefaultRouteSurvives() {
// Not seen on a default route in the wild yet, but the Lenovo proves vendors put routes
// in tables past Int range — the day one holds a default, it must not overflow away.
val routes = DumpParsers.parseV6DefaultRoutes(
"default via fe80::1 dev wlan0 table 1000000015 proto ra metric 1024 expires 100sec pref medium"
)
assertEquals(1, routes.size)
assertEquals("1000000015", routes[0].table)
assertEquals(100L, routes[0].expiresSec)
}
// ---- link.ra_source: interface MACs ----
@Test
fun onePlusInterfaceMacsParsed() {
val macs = DumpParsers.parseInterfaceMacs(onePlusIpAddr)
assertEquals("be:3d:e2:93:78:b9", macs["dummy0"])
assertEquals("ba:6e:46:b5:3d:bb", macs["ifb0"])
// link/loopback and link/ipip are not identities.
assertFalse("lo" in macs)
assertFalse("tunl0" in macs)
// The capture is cut mid-stanza-header ("6: gre0@NONE: <NO") — no exception, no entry.
assertNull(macs["gre0"])
}
@Test
fun lenovoInterfaceMacsParsed() {
val macs = DumpParsers.parseInterfaceMacs(lenovoIpAddr)
assertEquals("16:50:43:ec:93:c4", macs["dummy0"])
assertEquals("fe:16:ea:60:a2:d1", macs["ifb1"])
// The @-suffixed stanza name resolves to the bare interface name.
assertEquals("00:00:00:00:00:00", macs["gretap0"])
}
// ---- sec.arp_watch: neighbor snapshot ----
@Test
fun onePlusNeighborsParsed() {
val n = DumpParsers.parseNeighbors(onePlusIpNeigh)
assertEquals(12, n.size) // the `uid=2000` noise line is not a neighbor
val failed = n.single { it.ip == "10.13.102.111" }
assertNull(failed.lladdr)
assertEquals("FAILED", failed.state)
assertEquals("wlan0", failed.dev)
val gw = n.single { it.ip == "10.13.102.1" }
assertEquals("78:9a:18:54:b8:f9", gw.lladdr)
assertEquals("REACHABLE", gw.state)
val v6gw = n.single { it.ip == "fe80::7a9a:18ff:fe54:b8f9" }
assertTrue(v6gw.router)
assertEquals("78:9a:18:54:b8:f9", v6gw.lladdr)
}
@Test
fun lenovoNeighborsParsed() {
val n = DumpParsers.parseNeighbors(lenovoIpNeigh)
assertEquals(9, n.size)
assertTrue(n.all { it.lladdr != null && it.state == "STALE" })
assertEquals(1, n.count { it.router })
}
@Test
fun lladdrByIpIsTheComparisonSurface() {
val pairs = DumpParsers.lladdrByIp(DumpParsers.parseNeighbors(onePlusIpNeigh))
// 12 neighbors, 11 with a MAC — the FAILED entry must drop out, or a diff against a
// later run would flag "null → MAC" as a gateway change.
assertEquals(11, pairs.size)
assertEquals("78:9a:18:54:b8:f9", pairs["10.13.102.1"])
assertFalse("10.13.102.111" in pairs)
}
// ---- sec.arp_watch: monitor window ----
@Test
fun monitorSentinelIsUnusableAndYieldsNoEvents() {
assertFalse(DumpParsers.captureUsable(lenovoIpMonitor))
assertTrue(DumpParsers.parseNeighborEvents(lenovoIpMonitor).isEmpty())
}
@Test
fun quietMonitorWindowIsUsableButEmpty() {
// OnePlus: the monitor ran (only the uid noise line came back) — "ran and saw nothing"
// must stay distinguishable from "never ran".
assertTrue(DumpParsers.captureUsable(onePlusIpMonitor))
assertTrue(DumpParsers.parseNeighborEvents(onePlusIpMonitor).isEmpty())
}
@Test
fun monitorNeighEventsParsedFromLabeledLines() {
// Synthetic, in `ip monitor all` label format — neither archived run caught a live
// transition, but the format is fixed by iproute2's print_neigh/print_headers.
val sample = """
[NEIGH]10.13.102.1 dev wlan0 lladdr 78:9a:18:54:b8:f9 REACHABLE
[NEIGH]Deleted 10.13.102.5 dev wlan0 lladdr 90:09:d0:1a:83:e4 STALE
[ROUTE]default via 10.13.102.1 dev wlan0 table 1015
[NEIGH]fe80::7a9a:18ff:fe54:b8f9 dev wlan0 lladdr 78:9a:18:54:b8:f9 router STALE
""".trimIndent()
val events = DumpParsers.parseNeighborEvents(sample)
assertEquals(3, events.size) // the ROUTE line belongs to a different family
assertEquals("10.13.102.1", events[0].entry.ip)
assertFalse(events[0].deleted)
assertTrue(events[1].deleted)
assertEquals("90:09:d0:1a:83:e4", events[1].entry.lladdr)
assertTrue(events[2].entry.router)
}
// ---- degradation: missing or garbage input ----
@Test
fun missingAndGarbageInputYieldsEmptyResultsNotExceptions() {
for (bad in listOf(null, "", " \n ", "EXEC_TIMEOUT(newProcess)", "SHIZUKU_BINDER_DEAD",
"NEWPROCESS_UNAVAILABLE", "total garbage\nno routes here at all\ndefault", "default")) {
assertTrue(DumpParsers.parseV6DefaultRoutes(bad).isEmpty(), "routes from: $bad")
assertTrue(DumpParsers.parseNeighbors(bad).isEmpty(), "neighbors from: $bad")
assertTrue(DumpParsers.parseInterfaceMacs(bad).isEmpty(), "macs from: $bad")
assertTrue(DumpParsers.parseNeighborEvents(bad).isEmpty(), "events from: $bad")
}
}
}
+5
View File
@@ -8,6 +8,7 @@ lifecycle = "2.8.7"
activityCompose = "1.9.3" activityCompose = "1.9.3"
composeBom = "2024.10.01" composeBom = "2024.10.01"
shizuku = "13.1.5" shizuku = "13.1.5"
junit4 = "4.13.2"
[libraries] [libraries]
kotlinx-serialization-json = { group = "org.jetbrains.kotlinx", name = "kotlinx-serialization-json", version.ref = "kotlinxSerialization" } kotlinx-serialization-json = { group = "org.jetbrains.kotlinx", name = "kotlinx-serialization-json", version.ref = "kotlinxSerialization" }
@@ -23,6 +24,10 @@ androidx-ui-tooling = { group = "androidx.compose.ui", name = "ui-tooling" }
androidx-ui-tooling-preview = { group = "androidx.compose.ui", name = "ui-tooling-preview" } androidx-ui-tooling-preview = { group = "androidx.compose.ui", name = "ui-tooling-preview" }
androidx-material3 = { group = "androidx.compose.material3", name = "material3" } androidx-material3 = { group = "androidx.compose.material3", name = "material3" }
shizuku-api = { group = "dev.rikka.shizuku", name = "api", version.ref = "shizuku" } shizuku-api = { group = "dev.rikka.shizuku", name = "api", version.ref = "shizuku" }
# Android-module unit tests run on JUnit 4 (AGP's default); the JVM modules use kotlin("test")
# with the JUnit Platform instead — that helper isn't available under AGP 9's built-in Kotlin.
kotlin-test-junit = { group = "org.jetbrains.kotlin", name = "kotlin-test-junit", version.ref = "kotlin" }
junit4 = { group = "junit", name = "junit", version.ref = "junit4" }
shizuku-provider = { group = "dev.rikka.shizuku", name = "provider", version.ref = "shizuku" } shizuku-provider = { group = "dev.rikka.shizuku", name = "provider", version.ref = "shizuku" }
[plugins] [plugins]
+52
View File
@@ -0,0 +1,52 @@
#!/usr/bin/env bash
# SPDX-FileCopyrightText: 2026 Echolot contributors
# SPDX-License-Identifier: GPL-3.0-or-later
#
# Mints an enrollment link on the probe server and prints it — as text, as a QR code if
# `qrencode` is around, and as an adb command if a device is attached.
#
# The link is minted by the server binary on the host rather than over HTTP. The admin API this
# used to call is gone: the admin UI that replaced it is authenticated, as it should be, and
# adding a second unauthenticated door on loopback is what briefly exposed the old one to the
# network. A root shell on the host needs no authentication anyway — whoever has one already has
# every privilege the server has.
#
# The link carries a single-use bearer token: treat it like a password until it is redeemed.
#
# Usage: echolot-app/scripts/enroll-link.sh [note]
set -euo pipefail
SSH_HOST="${ECHOLOT_SSH:-claude-echolot}"
NOTE="${1:-manual}"
# The env file is sourced rather than assumed: the state directory and the public URL live there,
# and minting against the wrong state directory would produce a token the running server has
# never heard of.
REMOTE='set -a; . /etc/echolot/server.env; set +a;
exec /usr/local/bin/echolot-server --mint-enroll-token'
RAW=$(ssh -o BatchMode=yes "$SSH_HOST" "sudo sh -c \"$REMOTE '$NOTE'\"" 2>/dev/null || true)
URI=$(printf '%s' "$RAW" | tr -d '\r' | grep -m1 '^echolot://enroll' || true)
if [ -z "$URI" ]; then
echo "could not mint a link — needs a server with --mint-enroll-token (v0.9.7+)." >&2
echo "raw response:" >&2
printf '%s\n' "$RAW" >&2
exit 1
fi
echo "$URI"
echo
# A QR is the point of the format: scanning beats pasting a 200-character string onto a phone.
if command -v qrencode >/dev/null 2>&1; then
qrencode -t ANSIUTF8 "$URI"
else
echo "(install qrencode to get a scannable QR here)"
fi
# With a device attached, the deep link can be delivered straight to the app — no typing at all.
if command -v adb >/dev/null 2>&1 && [ -n "$(adb devices | sed -n '2p')" ]; then
echo
echo "attached device — deliver it directly with:"
echo " adb shell am start -a android.intent.action.VIEW -d '$URI'"
fi
+3
View File
@@ -22,4 +22,7 @@ VOLUME ["/state"]
# the data plane must see real client source addresses/TTLs, and Docker's # the data plane must see real client source addresses/TTLs, and Docker's
# userland NAT would falsify exactly what this server exists to observe. # userland NAT would falsify exactly what this server exists to observe.
EXPOSE 8441/tcp 8442/udp 8443/tcp EXPOSE 8441/tcp 8442/udp 8443/tcp
# The verb is explicit here too, so `docker run <image>` serves and `docker run <image> --help`
# still works by overriding the command.
ENTRYPOINT ["/echolot-server"] ENTRYPOINT ["/echolot-server"]
CMD ["--serve"]
+95 -8
View File
@@ -106,9 +106,7 @@ ECHOLOT_UDP_LISTEN=203.0.113.10:8442,203.0.113.11:8442,[2001:db8::10]:8442,[2001
Passing `--self-update-api` to `--install-systemd` additionally installs a daily randomized Passing `--self-update-api` to `--install-systemd` additionally installs a daily randomized
self-update timer (`echolot-server-update.timer`) that restarts the service after a successful self-update timer (`echolot-server-update.timer`) that restarts the service after a successful
update. Updates are checksum-verified against the release's `SHA256SUMS` (integrity, not update.
authenticity — signature verification remains TODO before treating the update source as
untrusted).
### Self-update (opt-in, native only) ### Self-update (opt-in, native only)
@@ -119,14 +117,24 @@ echolot-server --self-update \
Fetches the newest `server-v*` release asset for this OS/arch and atomically replaces the Fetches the newest `server-v*` release asset for this OS/arch and atomically replaces the
binary; systemd's `Restart=` brings up the new version. Run it from a systemd timer for binary; systemd's `Restart=` brings up the new version. Run it from a systemd timer for
unattended updates. TODO before enabling anywhere untrusted: signature verification of the unattended updates.
downloaded asset.
Releases are trusted by signature, not by host: CI signs `SHA256SUMS` with an ed25519 key that
exists only in its secret store (`RELEASE_SIGNING_KEY`), and the updater verifies
`SHA256SUMS.sig` against the public key baked into the binary before believing any checksum —
an unsigned or re-signed release is refused, so a compromised Gitea can withhold updates but not
inject one. Running your own release pipeline? Mint a keypair with
`go run ./cmd/release-sign -gen`, set the secret, and point `ECHOLOT_SELF_UPDATE_PUBKEY` (or
`--self-update-pubkey`) at your public key.
## First contact ## First contact
```sh ```sh
# 1. mint an enrollment token (admin listener is loopback-only) # 1. mint an enrollment token (admin listener is loopback-only; authenticates as the
curl -s -X POST 'http://127.0.0.1:8444/admin/enroll-tokens?note=phone' # break-glass admin — set that once with --set-admin-password)
curl -s -u admin:<password> -H 'Accept: application/json' \
-X POST 'http://127.0.0.1:8444/admin/enroll-tokens?note=phone'
# → { "token": "…", "expires_in_s": 86400, "enroll_uri": "echolot://enroll?…" }
# 2. device enrolls with it (normally via the echolot:// QR code) # 2. device enrolls with it (normally via the echolot:// QR code)
curl -sk -X POST https://<host>:8443/v1/enroll -H 'Authorization: Bearer <token>' curl -sk -X POST https://<host>:8443/v1/enroll -H 'Authorization: Bearer <token>'
# 3. device fetches its profile # 3. device fetches its profile
@@ -135,6 +143,47 @@ curl -sk https://<host>:8443/v1/profile -H 'Authorization: Bearer <credential>'
The SPKI pin clients must verify is logged at startup (`pin-sha256`). The SPKI pin clients must verify is logged at startup (`pin-sha256`).
## Dev relay (adb endpoint)
Development scaffolding, not protocol: it appears in no capability list and in no measurement
document, and `docs/probe-protocol.md` does not describe it.
It exists because **mDNS does not cross subnets**. Android's wireless debugging advertises adbd's
port over mDNS and rotates that port every few minutes, so a developer on another subnet cannot
discover it at all. An Echolot instance running on the test LAN can — and relays it here.
```
POST /v1/devtools/adb-endpoint control plane, device credential (same Bearer as /v1/sessions)
{"host":"10.13.102.128","port":45305,"device_name":"TB330FU","note":"…"}
GET /admin/adb-endpoints admin UI, session cookie or HTTP Basic
```
Read it back from the developer's machine (the admin listener is loopback-only, so over the same
SSH tunnel as everything else):
```sh
curl -s -u admin:<password> -H 'Accept: application/json' \
http://127.0.0.1:8444/admin/adb-endpoints
# → [{"device":"…","host":"10.13.102.128","port":45305,"device_name":"TB330FU",
# "reported_at":"…","source_ip":"…","age_s":37}]
```
The same list is a card on the admin dashboard, so it is usable without curl.
The newest report per submitting device wins — a rotated port makes the previous one wrong, not
historical. The device id comes from the credential and `source_ip` from the connection, so neither
is something a body can claim. `ECHOLOT_ADB_ENDPOINT_RETENTION_H` (default 24) bounds how long an
entry lives; `0` keeps it until the device replaces it. The window exists because the value is a
LAN address that stops being true within minutes: keeping it afterwards discloses the inside of
someone's network in exchange for nothing.
**Why it lives on these two listeners and nowhere else.** Its predecessor was a separate Python
service wildcard-bound to `0.0.0.0:443`. That silently occupied port 443 on the addresses reserved
for measurement — voiding the IPv4 interception proof for as long as it ran — and it accepted a port
report from anyone who could reach it. So: no new listener, no new port, no wildcard bind, and
nothing unauthenticated. The submit side is rate-limited by the existing §2.5 actions bucket
(`ECHOLOT_RATE_ACTIONS_PER_MIN`).
## Development ## Development
```sh ```sh
@@ -144,5 +193,43 @@ go vet ./...
CI (`.gitea/workflows/build-server.yml`): tests on every push touching `server/`; CI (`.gitea/workflows/build-server.yml`): tests on every push touching `server/`;
tagging `server-v1.2.3` builds + pushes the container image to the Gitea registry and tagging `server-v1.2.3` builds + pushes the container image to the Gitea registry and
attaches static linux amd64/arm64 binaries (+ SHA256SUMS) to a release — the same attaches static linux amd64/arm64 binaries (+ signed SHA256SUMS) to a release — the same
artifacts `--self-update` consumes. artifacts `--self-update` consumes.
## TLS for the admin UI
The binary terminates TLS itself; there is no reverse proxy in the design. It already serves TLS
for the control plane, so this is reuse rather than new machinery, and it keeps the "one process,
one config file" property. A proxy would also invite someone to eventually front the control plane
too — which would break SPKI pinning, because clients pin *that* certificate's key.
```
ECHOLOT_ADMIN_LISTEN=[2001:db8::2]:443
ECHOLOT_ADMIN_TLS_CERT=/etc/echolot/admin.pem
ECHOLOT_ADMIN_TLS_KEY=/etc/echolot/admin.key
ECHOLOT_ADMIN_BASE_URL=https://admin.example.net
```
Certificates come from any ACME client. **DNS-01 is the one to use here**: it needs no inbound
port 80, which matters on a host where 80 is awkward or already spoken for.
```sh
acme.sh --issue --dns dns_cf -d admin.example.net \
--key-file /etc/echolot/admin.key \
--fullchain-file /etc/echolot/admin.pem
```
**No reload hook is needed.** The certificate is re-read when the files change, so a renewal that
drops new files in place is picked up on the next handshake. That is deliberate: a reload hook is
the part of a renewal setup that quietly stops working, months later, and is noticed only once the
certificate has already expired. A torn write — renewal tools write cert and key separately — keeps
the previous certificate rather than failing the listener.
Serving the admin UI in plaintext on a non-loopback address is refused: the session cookie is a
bearer credential for everything the server can do, and the OIDC authorization code arrives in a
URL. Bind to loopback and use an SSH tunnel (`ssh -L 8444:localhost:8444 host`), supply a
certificate, or set `ECHOLOT_ADMIN_INSECURE=1` if you mean it.
The control-plane certificate is deliberately *not* hot-reloaded. Clients pin its public key, so
replacing it is a rotation an operator should have to think about, not something that happens
because a file changed.
+389 -35
View File
@@ -11,6 +11,7 @@
package main package main
import ( import (
"bufio"
"context" "context"
"crypto/ecdsa" "crypto/ecdsa"
"crypto/elliptic" "crypto/elliptic"
@@ -18,7 +19,6 @@ import (
"crypto/tls" "crypto/tls"
"crypto/x509" "crypto/x509"
"crypto/x509/pkix" "crypto/x509/pkix"
"encoding/json"
"encoding/pem" "encoding/pem"
"errors" "errors"
"fmt" "fmt"
@@ -36,11 +36,17 @@ import (
"syscall" "syscall"
"time" "time"
"echo-lot.app/server/internal/acmehttp"
"echo-lot.app/server/internal/adminauth"
"echo-lot.app/server/internal/adminui"
"echo-lot.app/server/internal/canarydns" "echo-lot.app/server/internal/canarydns"
"echo-lot.app/server/internal/certreload"
"echo-lot.app/server/internal/compat" "echo-lot.app/server/internal/compat"
"echo-lot.app/server/internal/config" "echo-lot.app/server/internal/config"
"echo-lot.app/server/internal/control" "echo-lot.app/server/internal/control"
"echo-lot.app/server/internal/dataplane" "echo-lot.app/server/internal/dataplane"
"echo-lot.app/server/internal/oidc"
"echo-lot.app/server/internal/ratelimit"
"echo-lot.app/server/internal/runs" "echo-lot.app/server/internal/runs"
"echo-lot.app/server/internal/selftest" "echo-lot.app/server/internal/selftest"
"echo-lot.app/server/internal/selfupdate" "echo-lot.app/server/internal/selfupdate"
@@ -69,6 +75,30 @@ func run() error {
control.Version = Version control.Version = Version
switch { switch {
case actions.Help:
// Compatibility shim for one release.
//
// Serving became an explicit verb, but self-update is run by the *old* binary — so the
// repair added to the updater cannot fix the very update that installs the new one. A
// unit written before this change starts us with no arguments, and without this branch
// the service would simply stop working, unattended, on a host nobody is watching.
//
// Only when systemd started us: INVOCATION_ID is set by systemd for every service
// invocation and by nothing else, so a person at a terminal still gets usage. Remove
// this once no deployment predates --serve.
if os.Getenv("INVOCATION_ID") != "" {
slog.Warn("started by systemd with no verb — this unit predates --serve; " +
"repairing it and serving anyway")
if repaired, err := system.RepairExecStart(); err != nil {
slog.Error("could not repair the unit; fix ExecStart by hand", "err", err)
} else if repaired {
slog.Info("systemd unit updated to pass --serve")
}
return serve(cfg)
}
config.Usage(os.Stderr)
os.Exit(2)
return nil
case actions.Version: case actions.Version:
fmt.Println(Version) fmt.Println(Version)
return nil return nil
@@ -79,21 +109,94 @@ func run() error {
return system.InstallSystemd(cfg.SelfUpdateAPI) return system.InstallSystemd(cfg.SelfUpdateAPI)
case actions.UninstallSystemd: case actions.UninstallSystemd:
return system.UninstallSystemd() return system.UninstallSystemd()
case actions.SetAdminPassword:
return setAdminPassword(cfg)
case actions.MintEnrollToken != "":
return mintEnrollToken(cfg, actions.MintEnrollToken)
case actions.SelfUpdate: case actions.SelfUpdate:
return selfupdate.Run(cfg.SelfUpdateAPI, Version) return selfupdate.Run(cfg.SelfUpdateAPI, cfg.SelfUpdatePubKey, Version)
} }
return serve(cfg) return serve(cfg)
} }
// mintEnrollToken prints a §2.1 bootstrap link for a new device.
//
// The link is assembled here rather than by hand because it has to carry the public URL and the
// base64 SPKI pin percent-encoded correctly, and a pin wrong by one character fails later as an
// inscrutable TLS error rather than as a bad pin.
func mintEnrollToken(cfg *config.Config, note string) error {
st, err := store.Open(cfg.StateDir)
if err != nil {
return fmt.Errorf("state store: %w", err)
}
cert, err := loadOrCreateCert(cfg)
if err != nil {
return fmt.Errorf("tls: %w", err)
}
pin, err := control.SpkiPinB64(cert)
if err != nil {
return fmt.Errorf("pin: %w", err)
}
tok, err := st.NewEnrollToken(24*time.Hour, note)
if err != nil {
return err
}
base := cfg.PublicControlURL
if base == "" {
return fmt.Errorf("set ECHOLOT_PUBLIC_URL so the link can say where to connect")
}
fmt.Println(control.EnrollmentURI(base, pin, tok))
// stderr, so piping the command somewhere yields the link alone.
fmt.Fprintln(os.Stderr, "\nSingle use, valid 24 hours. Treat it like a password until spent.")
return nil
}
// controlURL is the address devices connect to: the hostname that selects the pinned certificate.
//
// Falls back to the public URL when no separate control hostname is configured, so a server that
// does not share the admin port keeps answering discovery with something usable.
func controlURL(cfg *config.Config) string {
if cfg.ControlHostname == "" {
return cfg.PublicControlURL
}
return "https://" + cfg.ControlHostname
}
func serve(cfg *config.Config) error { func serve(cfg *config.Config) error {
slog.Info("echolot-server starting", "version", Version, "mode", slog.Info("echolot-server starting", "version", Version, "mode",
map[bool]string{true: "container", false: "native"}[cfg.Docker], map[bool]string{true: "container", false: "native"}[cfg.Docker],
"state_dir", cfg.StateDir) "state_dir", cfg.StateDir)
// The reserved addresses' proof is only as good as 80/443 actually being free there.
// CheckReserved already keeps OUR listeners away, but a process outside this config pollutes
// them just as silently — the adb-beacon receiver on 0.0.0.0:443 did exactly that. So ask
// the OS, not the config. A hard stop for the same reason CheckReserved is one: the failure
// is invisible, and its first symptom is a measurement calling an intercepted network clean.
if reserved := cfg.ReservedIPs(); len(reserved) > 0 {
occupied, unverifiable := selftest.ReservedWebPortsFree(reserved)
if len(occupied) > 0 {
return fmt.Errorf(
"refusing to start: something outside this server is listening on reserved "+
"measurement address(es) %s\n"+
"The interception proof those addresses exist for is void while anything "+
"answers there.\nFind it with `ss -tlnp | grep -E ':(80|443) '`, stop it, "+
"or remove the address from ECHOLOT_RESERVED_ADDRS if it is no longer reserved",
strings.Join(occupied, ", "))
}
for _, u := range unverifiable {
// Not fatal: an address with a typo, or one this host no longer carries, is a
// config problem — refusing to serve over it would take the whole instrument down.
slog.Warn("could not verify a reserved web port is free", "addr", u)
}
}
st, err := store.Open(cfg.StateDir) st, err := store.Open(cfg.StateDir)
if err != nil { if err != nil {
return fmt.Errorf("state store: %w", err) return fmt.Errorf("state store: %w", err)
} }
// Dev-relay breadcrumbs carry a LAN address, so they expire on a clock like the canary DNS
// log does rather than sitting in the state file until someone notices them.
st.SetADBEndpointRetention(time.Duration(cfg.ADBEndpointRetentionH) * time.Hour)
cert, err := loadOrCreateCert(cfg) cert, err := loadOrCreateCert(cfg)
if err != nil { if err != nil {
return fmt.Errorf("tls: %w", err) return fmt.Errorf("tls: %w", err)
@@ -106,12 +209,30 @@ func serve(cfg *config.Config) error {
sessions := session.NewManager(15 * time.Minute) sessions := session.NewManager(15 * time.Minute)
dp := &dataplane.Server{Sessions: sessions} dp := &dataplane.Server{Sessions: sessions}
// Spec §2.5 ceilings. Control plane answers 429; the data plane drops silently. 0 = off.
if cfg.RateUDPPps > 0 {
pps := float64(cfg.RateUDPPps)
// Burst of two seconds' worth: a 5000-packet train arrives as one burst by design.
dp.PacketRate = ratelimit.New(pps, 2*pps)
}
if cfg.RateUDPKbps > 0 {
bytesPerSec := float64(cfg.RateUDPKbps) * 125 // kbps -> bytes/s
dp.ByteRate = ratelimit.New(bytesPerSec, bytesPerSec)
}
// TCP echo shares the control cert for its elt-echo TLS variant. // TCP echo shares the control cert for its elt-echo TLS variant.
tcpSrv := &tcpecho.Server{ tcpSrv := &tcpecho.Server{
TLSConfig: &tls.Config{Certificates: []tls.Certificate{cert}, MinVersion: tls.VersionTLS12}, TLSConfig: &tls.Config{Certificates: []tls.Certificate{cert}, MinVersion: tls.VersionTLS12},
} }
caps := []string{"udp-probe", "delayed-echo", "connect-back", "http-echo", "downtrain", "big-send"} caps := []string{"udp-probe", "delayed-echo", "connect-back", "http-echo", "downtrain", "big-send", "throughput"}
// Crafted fragments need a raw socket. Advertised only when one can actually be opened —
// a capability we cannot deliver turns a missing feature into a failed measurement.
rawFrag := dataplane.RawFragSupported()
if rawFrag {
caps = append(caps, "frag-send")
} else {
slog.Info("frag-send unavailable: no raw socket (needs CAP_NET_RAW)")
}
if len(config.Addrs(cfg.TCPListen)) > 0 { if len(config.Addrs(cfg.TCPListen)) > 0 {
caps = append(caps, "tcp-echo", "tls-echo") caps = append(caps, "tcp-echo", "tls-echo")
} }
@@ -142,7 +263,43 @@ func serve(cfg *config.Config) error {
slog.Info("client compatibility", "accepts_app", appRange.String(), slog.Info("client compatibility", "accepts_app", appRange.String(),
"protocol", control.ProtocolVersion, "schema", control.SchemaVersion) "protocol", control.ProtocolVersion, "schema", control.SchemaVersion)
// Identity is optional. Without an issuer the server simply has no sign-in, and
// uploads=account can never be satisfied — which is the honest outcome, not a silent
// downgrade to anonymous.
// One verifier per issuer. An IdP may mint a distinct issuer per application — Authentik
// derives it from the application slug — and a token's `iss` must match whoever signed it.
// Each verifier accepts only the client belonging to its own issuer, so a token minted for
// the phone cannot be replayed at the admin login and vice versa.
var idp, adminIdP *oidc.Verifier
appIssuer := cfg.OIDCAppIssuer
if appIssuer == "" {
appIssuer = cfg.OIDCIssuer // IdPs with one global issuer
}
if appIssuer != "" && cfg.OIDCAppClientID != "" {
idp = oidc.New(oidc.Config{
Issuer: appIssuer, AppClientID: cfg.OIDCAppClientID, AdminGroup: cfg.OIDCAdminGroup,
}, nil)
slog.Info("identity: app client", "issuer", appIssuer, "client_id", cfg.OIDCAppClientID)
}
if cfg.OIDCIssuer != "" && cfg.OIDCClientID != "" {
adminIdP = oidc.New(oidc.Config{
Issuer: cfg.OIDCIssuer, ClientID: cfg.OIDCClientID, AdminGroup: cfg.OIDCAdminGroup,
}, nil)
slog.Info("identity: admin client", "issuer", cfg.OIDCIssuer,
"client_id", cfg.OIDCClientID, "admin_group", cfg.OIDCAdminGroup)
if cfg.OIDCAdminGroup == "" {
slog.Warn("no admin group set: nobody will be an admin via OIDC " +
"(set ECHOLOT_OIDC_ADMIN_GROUP)")
}
}
if idp == nil && adminIdP == nil && cfg.UploadsMode == string(runs.ModeAccount) {
slog.Warn("uploads=account but no identity provider is configured — " +
"every upload will be refused")
}
ip4, ip6, ip4Alt, ip6Alt := cfg.MeasurementAddrs()
ctl := &control.Server{ ctl := &control.Server{
IP4: ip4, IP6: ip6, IP4Alt: ip4Alt, IP6Alt: ip6Alt,
Store: st, Sessions: sessions, Name: cfg.Name, Store: st, Sessions: sessions, Name: cfg.Name,
UDPPort: mustPort(firstAddr(cfg.UDPListen)), TCPPort: mustPort(firstAddr(cfg.TCPListen)), UDPPort: mustPort(firstAddr(cfg.UDPListen)), TCPPort: mustPort(firstAddr(cfg.TCPListen)),
StunPort: mustPort(firstAddr(cfg.StunListen)), PinB64: pin, CertChain: cert.Certificate, StunPort: mustPort(firstAddr(cfg.StunListen)), PinB64: pin, CertChain: cert.Certificate,
@@ -153,6 +310,20 @@ func serve(cfg *config.Config) error {
Runs: runStore, Runs: runStore,
AppRange: appRange, AppRange: appRange,
PublicControlURL: publicControlURL(cfg), PublicControlURL: publicControlURL(cfg),
OIDC: idp,
AdminOIDC: adminIdP,
}
// Left nil when there is no raw socket, so the handler answers "not implemented" with a
// reason rather than failing somewhere deeper.
if rawFrag {
ctl.FragSend = dp.FragSend
}
ctl.DownThroughput = dp.DownThroughput
if cfg.RateSessionsPerMin > 0 {
ctl.RateSessions = ratelimit.New(float64(cfg.RateSessionsPerMin)/60, float64(cfg.RateSessionsPerMin))
}
if cfg.RateActionsPerMin > 0 {
ctl.RateActions = ratelimit.New(float64(cfg.RateActionsPerMin)/60, float64(cfg.RateActionsPerMin))
} }
ctx, stop := signal.NotifyContext(context.Background(), syscall.SIGINT, syscall.SIGTERM) ctx, stop := signal.NotifyContext(context.Background(), syscall.SIGINT, syscall.SIGTERM)
@@ -216,38 +387,173 @@ func serve(cfg *config.Config) error {
return best return best
} }
// Admin/health (plain HTTP, localhost by default; spec §7) // The admin interface. Every route except /healthz requires a session — the old arrangement
admin := http.NewServeMux() // (no auth, kept safe by binding to loopback) failed the moment the address changed, and a
admin.HandleFunc("GET /healthz", func(w http.ResponseWriter, _ *http.Request) { // binding address is a deployment detail rather than an access control.
fmt.Fprintf(w, `{"ok":true,"version":%q}`, Version) secret, err := st.SessionSecret()
}) if err != nil {
admin.HandleFunc("GET /admin/selftest", func(w http.ResponseWriter, _ *http.Request) { return fmt.Errorf("admin session secret: %w", err)
w.Header().Set("Content-Type", "application/json") }
_ = json.NewEncoder(w).Encode(selftestPtr.Load()) adminSecure := cfg.AdminTLSCert != ""
}) ui := &adminui.Server{
// TODO(spec §7): enrollment token management + device list. Until the Store: st,
// admin UI exists, mint tokens with: echolot-admin (or curl on this Runs: runStore,
// listener once the endpoint lands). OIDC: adminIdP,
admin.HandleFunc("POST /admin/enroll-tokens", func(w http.ResponseWriter, r *http.Request) { Sessions: adminauth.NewSessions(secret, 12*time.Hour),
tok, err := st.NewEnrollToken(24*time.Hour, r.URL.Query().Get("note")) Throttle: adminauth.NewThrottle(),
AdminUser: cfg.AdminUser,
BaseURL: cfg.AdminBaseURL,
ClientSecret: cfg.OIDCClientSecret,
Secure: adminSecure,
EnrollLink: ctl.EnrollmentLink,
// Where devices should connect, for /v1/discover. Derived from the control hostname so it
// cannot drift from the name that actually selects the pinned certificate.
ControlURL: controlURL(cfg),
ServerName: cfg.Name,
SelfTest: func() any { return selftestPtr.Load() },
Version: Version,
}
if st.LocalAdmin() == nil && adminIdP == nil {
slog.Warn("nobody can sign in to the admin UI: no break-glass password is set " +
"(--set-admin-password) and no identity provider is configured")
}
admin := ui.Handler()
// One listener per configured address, all serving the same handler.
//
// Multi-address rather than a wildcard because this host reserves addresses for measurement:
// binding 0.0.0.0 would put the admin UI on port 443 of the reserved pair, and their value
// comes precisely from nothing answering there. Explicit addresses are also what let the
// service and management addresses differ without a second process.
adminAddrs := config.Addrs(cfg.AdminListen)
if len(adminAddrs) == 0 {
return fmt.Errorf("admin: no listen address configured")
}
var adminTLS *tls.Config
if cfg.AdminTLSCert != "" {
// Terminated here rather than behind a reverse proxy: this binary already serves TLS for
// the control plane, so it is reuse rather than new machinery, and one process with one
// config file is the property that makes this pleasant to run. A proxy would also invite
// someone to later front the control plane too, which would break SPKI pinning.
reloader, err := certreload.New(cfg.AdminTLSCert, cfg.AdminTLSKey)
if err != nil { if err != nil {
http.Error(w, err.Error(), 500) return fmt.Errorf("admin TLS: %w", err)
}
adminTLS = reloader.TLSConfig()
if exp := reloader.NotAfter(); !exp.IsZero() {
slog.Info("admin UI TLS", "listen", adminAddrs, "cert_expires", exp.Format(time.RFC3339))
if time.Until(exp) < 14*24*time.Hour {
slog.Warn("admin certificate expires soon", "expires", exp.Format(time.RFC3339))
}
}
}
// Kept for shutdown: each listener gets its own server, and a graceful stop has to reach all
// of them or an in-flight admin request is cut off mid-response on every address but one.
var adminSrvs []*http.Server
// Sharing port 443 between two services that cannot share a certificate. The name in the TLS
// handshake picks the certificate, and the name in the request picks the handler; both have to
// agree or a client would get the pinned certificate and the admin UI behind it.
//
// The control plane keeps its own listener as well. Devices enrolled before this carry the old
// URL in their settings, and taking that away would strand every one of them for the sake of a
// port number.
ctlHandler := ctl.Handler()
sharedCert := cert
// Which side of the port a request belongs to.
//
// The control hostname is the obvious case. A bare IP is the other one, and it matters: a
// client whose DNS has failed can still reach the server by an address it cached from the
// profile, and a measurement tool that cannot report from a broken network is useless
// precisely when it is needed. That client authenticates by pin, so the name it used to get
// here is not part of the trust decision.
//
// Safe to route that way because the admin UI is only ever reached by name: browsers always
// send SNI and nobody bookmarks an IP for a site with a Let's Encrypt certificate. Anything
// addressing this server numerically is a pinned client.
isControl := func(host string) bool {
if h, _, err := net.SplitHostPort(host); err == nil {
host = h
}
host = strings.Trim(host, "[]")
if cfg.ControlHostname != "" && strings.EqualFold(host, cfg.ControlHostname) {
return true
}
return net.ParseIP(host) != nil
}
pickCert := func(hi *tls.ClientHelloInfo) (*tls.Certificate, error) {
// No SNI at all also means a numeric client: every browser sends it.
if hi.ServerName == "" || isControl(hi.ServerName) {
return &sharedCert, nil
}
if adminTLS != nil && adminTLS.GetCertificate != nil {
return adminTLS.GetCertificate(hi)
}
return &sharedCert, nil
}
route := func(w http.ResponseWriter, r *http.Request) {
if isControl(r.Host) {
ctlHandler.ServeHTTP(w, r)
return return
} }
// The whole bootstrap, not just the token: this is what gets pasted or turned into a admin.ServeHTTP(w, r)
// QR code, and assembling it here is what keeps an operator from transcribing a pin by }
// hand — a pin wrong by one character fails as an inscrutable TLS error days later. sharedTLS := &tls.Config{GetCertificate: pickCert, MinVersion: tls.VersionTLS12}
w.Header().Set("Content-Type", "application/json") if cfg.ControlHostname != "" {
enc := json.NewEncoder(w) slog.Info("control plane shares the admin port",
enc.SetEscapeHTML(false) // the link is full of / and =; escaping them helps nobody "hostname", cfg.ControlHostname, "listen", adminAddrs)
_ = enc.Encode(map[string]any{ }
"token": tok,
"expires_in_s": 86400, for _, addr := range adminAddrs {
"enroll_uri": ctl.EnrollmentLink(tok), // Bound before the goroutine starts, so a bad address fails startup rather than being
}) // reported asynchronously after the process has already declared itself healthy.
}) ln, err := net.Listen("tcp", addr)
adminSrv := &http.Server{Addr: cfg.AdminListen, Handler: admin, ReadHeaderTimeout: 10 * time.Second} if err != nil {
go func() { errCh <- fmt.Errorf("admin: %w", adminSrv.ListenAndServe()) }() return fmt.Errorf("admin listen %s: %w", addr, err)
}
srv := &http.Server{
Handler: http.HandlerFunc(route),
ReadHeaderTimeout: 10 * time.Second,
TLSConfig: sharedTLS,
}
adminSrvs = append(adminSrvs, srv)
go func(ln net.Listener, addr string) {
// Plaintext only where there is no certificate at all — checkAdminExposure has
// already refused that anywhere but loopback.
if adminTLS == nil && cfg.ControlHostname == "" {
errCh <- fmt.Errorf("admin %s: %w", addr, srv.Serve(ln))
return
}
errCh <- fmt.Errorf("admin %s: %w", addr, srv.ServeTLS(ln, "", ""))
}(ln, addr)
}
// ACME HTTP-01 responder. Permanent rather than started per renewal: nothing binds and
// unbinds, so a renewal cannot fail because the port was briefly busy, and the ACME client
// needs only write access to a directory instead of the privilege to bind a low port.
if cfg.ACMEHTTPListen != "" {
webroot := cfg.ACMEWebroot
if webroot == "" {
webroot = filepath.Join(cfg.StateDir, "acme")
}
if err := acmehttp.EnsureWebroot(webroot); err != nil {
return fmt.Errorf("acme webroot: %w", err)
}
acmeHandler := acmehttp.Handler(webroot, cfg.AdminBaseURL)
for _, addr := range config.Addrs(cfg.ACMEHTTPListen) {
ln, err := net.Listen("tcp", addr)
if err != nil {
return fmt.Errorf("acme-http listen %s: %w", addr, err)
}
srv := &http.Server{Handler: acmeHandler, ReadHeaderTimeout: 10 * time.Second}
go func(ln net.Listener, addr string) {
errCh <- fmt.Errorf("acme-http %s: %w", addr, srv.Serve(ln))
}(ln, addr)
}
// Every address the name may resolve to needs the responder: the CA picks one, and a
// challenge that lands on an unbound address fails a renewal rather than a request.
slog.Info("acme http-01 responder", "listen", config.Addrs(cfg.ACMEHTTPListen),
"webroot", webroot, "redirects_to", cfg.AdminBaseURL)
}
// UDP data plane — one socket per configured address. Distinct sockets // UDP data plane — one socket per configured address. Distinct sockets
// (not wildcard) also guarantee responses leave from the address the // (not wildcard) also guarantee responses leave from the address the
@@ -312,7 +618,8 @@ func serve(cfg *config.Config) error {
var dnsTCP []net.Listener var dnsTCP []net.Listener
if dnsAddrs := config.Addrs(cfg.DNSListen); len(dnsAddrs) > 0 && cfg.CanaryZone != "" { if dnsAddrs := config.Addrs(cfg.DNSListen); len(dnsAddrs) > 0 && cfg.CanaryZone != "" {
v4, v6 := firstByFamily(dnsAddrs) v4, v6 := firstByFamily(dnsAddrs)
cd := canarydns.New(cfg.CanaryZone, cfg.Name, v4, v6) cd := canarydns.New(cfg.CanaryZone, cfg.Name, v4, v6,
time.Duration(cfg.DNSLogRetentionH)*time.Hour)
for _, addr := range dnsAddrs { for _, addr := range dnsAddrs {
ua, err := net.ResolveUDPAddr("udp", addr) ua, err := net.ResolveUDPAddr("udp", addr)
if err != nil { if err != nil {
@@ -338,7 +645,7 @@ func serve(cfg *config.Config) error {
} }
slog.Info("listening", slog.Info("listening",
"control", ctlAddrs, "admin", cfg.AdminListen, "udp", udpAddrs, "control", ctlAddrs, "admin", adminAddrs, "udp", udpAddrs,
"tcp", config.Addrs(cfg.TCPListen), "stun", config.Addrs(cfg.StunListen), "tcp", config.Addrs(cfg.TCPListen), "stun", config.Addrs(cfg.StunListen),
"dns", config.Addrs(cfg.DNSListen), "capabilities", ctl.Capabilities) "dns", config.Addrs(cfg.DNSListen), "capabilities", ctl.Capabilities)
@@ -348,7 +655,9 @@ func serve(cfg *config.Config) error {
shutCtx, cancel := context.WithTimeout(context.Background(), 5*time.Second) shutCtx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel() defer cancel()
_ = ctlSrv.Shutdown(shutCtx) _ = ctlSrv.Shutdown(shutCtx)
_ = adminSrv.Shutdown(shutCtx) for _, srv := range adminSrvs {
_ = srv.Shutdown(shutCtx)
}
for _, c := range udpConns { for _, c := range udpConns {
_ = c.Close() _ = c.Close()
} }
@@ -479,3 +788,48 @@ func publicControlURL(cfg *config.Config) string {
} }
return "https://" + addr return "https://" + addr
} }
// setAdminPassword stores the break-glass admin credential.
//
// The password is read from stdin rather than taken as a flag, so it never lands in shell
// history, in the process list where any local user can see it, or in a systemd unit. Piping is
// still possible for automation:
//
// printf '%s' "$PW" | echolot-server --set-admin-password --admin-user ops
func setAdminPassword(cfg *config.Config) error {
st, err := store.Open(cfg.StateDir)
if err != nil {
return fmt.Errorf("state store: %w", err)
}
fmt.Fprintf(os.Stderr, "New password for %q (input is not echoed if this is a terminal): ", cfg.AdminUser)
pw, err := readSecret()
if err != nil {
return err
}
fmt.Fprintln(os.Stderr)
cred, err := adminauth.NewCredential(cfg.AdminUser, pw)
if err != nil {
return err
}
if err := st.SetLocalAdmin(cred); err != nil {
return err
}
fmt.Fprintf(os.Stderr, "Break-glass admin %q set. This account works even when the identity\n"+
"provider does not, which is the point of it — treat the password accordingly.\n", cfg.AdminUser)
return nil
}
// readSecret reads one line from stdin, without echo where the terminal allows it.
func readSecret() (string, error) {
restore, _ := system.DisableEcho(os.Stdin)
if restore != nil {
defer restore()
}
r := bufio.NewReader(os.Stdin)
line, err := r.ReadString('\n')
if err != nil && line == "" {
return "", err
}
return strings.TrimSpace(line), nil
}
+84
View File
@@ -0,0 +1,84 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
// release-sign signs a release manifest (SHA256SUMS) with the project's ed25519 key, producing
// the detached <file>.sig that self-updating servers verify before trusting the checksums.
//
// release-sign -gen mint a keypair (seed on stdout — store it as the CI
// secret RELEASE_SIGNING_KEY; publish the public key)
// release-sign <file> sign; key read from $RELEASE_SIGNING_KEY, writes <file>.sig
// release-sign -verify -pub <b64> <f> check <f> against <f>.sig — what the updater will do
//
// Run from CI (build-server.yml); the private key exists only in the Actions secret store, never
// on the release host, which is the property that makes the signature worth having.
package main
import (
"flag"
"fmt"
"os"
"echo-lot.app/server/internal/relsign"
)
func main() {
gen := flag.Bool("gen", false, "generate a keypair and exit")
verify := flag.Bool("verify", false, "verify <file> against <file>.sig instead of signing")
pub := flag.String("pub", "", "public key (base64) for -verify")
flag.Parse()
if err := run(*gen, *verify, *pub, flag.Args()); err != nil {
fmt.Fprintln(os.Stderr, "release-sign:", err)
os.Exit(1)
}
}
func run(gen, verify bool, pub string, args []string) error {
if gen {
pubB64, seedB64, err := relsign.GenerateKey()
if err != nil {
return err
}
fmt.Printf("public key (embed / ECHOLOT_SELF_UPDATE_PUBKEY):\n%s\n\n"+
"private key (CI secret RELEASE_SIGNING_KEY — this is the only copy):\n%s\n",
pubB64, seedB64)
return nil
}
if len(args) != 1 {
return fmt.Errorf("usage: release-sign [-gen | -verify -pub <b64>] <file>")
}
file := args[0]
data, err := os.ReadFile(file)
if err != nil {
return err
}
if verify {
if pub == "" {
return fmt.Errorf("-verify needs -pub")
}
sig, err := os.ReadFile(file + ".sig")
if err != nil {
return err
}
if err := relsign.Verify(pub, data, string(sig)); err != nil {
return err
}
fmt.Printf("%s: signature OK\n", file)
return nil
}
seed := os.Getenv("RELEASE_SIGNING_KEY")
if seed == "" {
return fmt.Errorf("RELEASE_SIGNING_KEY is not set — refusing to produce an unsigned release")
}
sig, err := relsign.Sign(seed, data)
if err != nil {
return err
}
if err := os.WriteFile(file+".sig", []byte(sig+"\n"), 0o644); err != nil {
return err
}
fmt.Printf("wrote %s.sig\n", file)
return nil
}
+2
View File
@@ -1,3 +1,5 @@
module echo-lot.app/server module echo-lot.app/server
go 1.24 go 1.24
require github.com/skip2/go-qrcode v0.0.0-20200617195104-da1b6568686e // indirect
+2
View File
@@ -0,0 +1,2 @@
github.com/skip2/go-qrcode v0.0.0-20200617195104-da1b6568686e h1:MRM5ITcdelLK2j1vwZ3Je0FKVCfqOLp5zO6trqMLYs0=
github.com/skip2/go-qrcode v0.0.0-20200617195104-da1b6568686e/go.mod h1:XV66xRDqSt+GTGFMVlhk3ULuV0y9ZmzeVGR4mloJI3M=
+100
View File
@@ -0,0 +1,100 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
// Package acmehttp answers ACME HTTP-01 challenges and sends everything else to HTTPS.
//
// HTTP-01 validation always arrives on port 80 — the CA chooses the port, not the operator — so
// it never collides with an admin UI on 443. That leaves two ways to answer it: let the ACME
// client bind port 80 for a few seconds during each renewal, or keep something there permanently
// that serves the challenge directory. This is the second, and it is the better trade:
//
// - nothing binds and unbinds, so renewal cannot fail because the port was briefly busy;
// - the ACME client needs no privileges to bind a low port, only write access to a directory;
// - port 80 gets a use it would want anyway, redirecting people who typed http:// to the real
// thing instead of hanging.
//
// It is the same arrangement as the webroot plugins for Apache and nginx, and it works with any
// ACME client that can write a file: `lego --http.webroot`, `certbot --webroot`, `acme.sh -w`.
//
// The ACME client stays an external program on purpose. lego is also a Go library, but importing
// it would put a large dependency tree into a server that deliberately has none — and the CLI does
// the same job from a timer.
package acmehttp
import (
"log/slog"
"net/http"
"os"
"path/filepath"
"strings"
)
// ChallengePath is the fixed prefix ACME uses. It is not configurable, by the specification.
const ChallengePath = "/.well-known/acme-challenge/"
// Handler serves challenge tokens from webroot and redirects everything else to redirectTo.
//
// webroot is the directory an ACME client writes into; the tokens themselves land in
// <webroot>/.well-known/acme-challenge/<token>, which is exactly what --http.webroot expects.
func Handler(webroot, redirectTo string) http.Handler {
mux := http.NewServeMux()
mux.HandleFunc(ChallengePath, func(w http.ResponseWriter, r *http.Request) {
token := strings.TrimPrefix(r.URL.Path, ChallengePath)
// Tokens are base64url from the CA. Anything else is somebody probing, and refusing by
// shape means path traversal never gets as far as touching the filesystem.
if token == "" || !validToken(token) {
http.NotFound(w, r)
return
}
body, err := os.ReadFile(filepath.Join(webroot, filepath.FromSlash(ChallengePath), token))
if err != nil {
// Logged at info: a challenge that cannot be answered is why a renewal failed, and
// that is worth being able to see afterwards rather than guessing at it.
slog.Info("acme challenge not found", "token", token, "webroot", webroot)
http.NotFound(w, r)
return
}
slog.Info("answered acme challenge", "token", token, "from", r.RemoteAddr)
w.Header().Set("Content-Type", "text/plain")
_, _ = w.Write(body)
})
mux.HandleFunc("/", func(w http.ResponseWriter, r *http.Request) {
if redirectTo == "" {
http.Error(w, "this port serves ACME challenges only", http.StatusNotFound)
return
}
// 308 rather than 302: the method must not change, and the redirect is permanent in the
// sense that matters — this port will never serve the application.
http.Redirect(w, r, strings.TrimRight(redirectTo, "/")+r.URL.RequestURI(), http.StatusPermanentRedirect)
})
return mux
}
// validToken accepts only the base64url alphabet the ACME spec uses for tokens.
//
// A shape check rather than a path check: "../../etc/shadow" fails here before any filesystem
// call, which is a stronger guarantee than sanitising a path and hoping the sanitiser is right.
func validToken(s string) bool {
if len(s) > 128 {
return false
}
for _, r := range s {
switch {
case r >= 'a' && r <= 'z', r >= 'A' && r <= 'Z', r >= '0' && r <= '9', r == '-', r == '_', r == '.':
default:
return false
}
}
// A bare "." or ".." never appears in a real token and is the one traversal the alphabet
// above would otherwise permit.
return s != "." && s != ".."
}
// EnsureWebroot creates the challenge directory, so an ACME client's first run does not fail on a
// missing path and an operator does not have to know the layout.
func EnsureWebroot(webroot string) error {
return os.MkdirAll(filepath.Join(webroot, filepath.FromSlash(ChallengePath)), 0o755)
}
+98
View File
@@ -0,0 +1,98 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package acmehttp
import (
"net/http"
"net/http/httptest"
"os"
"path/filepath"
"testing"
)
func serve(t *testing.T, redirectTo string) (http.Handler, string) {
t.Helper()
root := t.TempDir()
if err := EnsureWebroot(root); err != nil {
t.Fatal(err)
}
return Handler(root, redirectTo), root
}
func TestServesAChallengeTokenWrittenByAnAcmeClient(t *testing.T) {
h, root := serve(t, "https://admin.example.net")
// Exactly what `lego --http.webroot` writes.
token := "abc-123_XYZ"
want := "abc-123_XYZ.keyauthorization-part"
if err := os.WriteFile(filepath.Join(root, ".well-known", "acme-challenge", token), []byte(want), 0o644); err != nil {
t.Fatal(err)
}
rec := httptest.NewRecorder()
h.ServeHTTP(rec, httptest.NewRequest("GET", ChallengePath+token, nil))
if rec.Code != 200 {
t.Fatalf("challenge not served: %d", rec.Code)
}
if rec.Body.String() != want {
t.Fatalf("body = %q, want %q", rec.Body.String(), want)
}
}
// The token comes from the network and is used to build a path. Rejecting by *shape* means
// traversal never reaches the filesystem at all, which is a stronger guarantee than sanitising.
func TestTraversalNeverTouchesTheFilesystem(t *testing.T) {
h, root := serve(t, "https://admin.example.net")
secret := filepath.Join(filepath.Dir(root), "secret.txt")
if err := os.WriteFile(secret, []byte("do not serve me"), 0o600); err != nil {
t.Fatal(err)
}
for _, bad := range []string{
"../secret.txt",
"..%2Fsecret.txt",
"../../etc/passwd",
"..",
".",
"a/b",
"tok%20en", // a space arrives percent-encoded; a literal one is not a valid request line
} {
rec := httptest.NewRecorder()
h.ServeHTTP(rec, httptest.NewRequest("GET", ChallengePath+bad, nil))
if rec.Code == 200 && rec.Body.String() == "do not serve me" {
t.Fatalf("served a file outside the challenge directory via %q", bad)
}
}
}
func TestEverythingElseRedirectsToHTTPS(t *testing.T) {
h, _ := serve(t, "https://admin.example.net")
rec := httptest.NewRecorder()
h.ServeHTTP(rec, httptest.NewRequest("GET", "/devices?page=2", nil))
if rec.Code != http.StatusPermanentRedirect {
t.Fatalf("code = %d, want 308", rec.Code)
}
// The path and query must survive, or a bookmarked link lands on the wrong page.
if got := rec.Header().Get("Location"); got != "https://admin.example.net/devices?page=2" {
t.Fatalf("Location = %q", got)
}
}
// With no admin URL configured there is nowhere to send people, and inventing one would be worse
// than saying so.
func TestNoRedirectTargetIsHonest(t *testing.T) {
h, _ := serve(t, "")
rec := httptest.NewRecorder()
h.ServeHTTP(rec, httptest.NewRequest("GET", "/", nil))
if rec.Code != http.StatusNotFound {
t.Fatalf("code = %d, want 404", rec.Code)
}
}
func TestMissingTokenIsNotFound(t *testing.T) {
h, _ := serve(t, "https://admin.example.net")
rec := httptest.NewRecorder()
h.ServeHTTP(rec, httptest.NewRequest("GET", ChallengePath+"never-written", nil))
if rec.Code != http.StatusNotFound {
t.Fatalf("code = %d, want 404", rec.Code)
}
}
+272
View File
@@ -0,0 +1,272 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
// Package adminauth handles who may administer the server.
//
// Two ways in, deliberately:
//
// - **OIDC**, the normal one. Identity lives in the operator's own IdP.
// - **A local admin password**, the break-glass one. If the IdP is misconfigured, unreachable,
// or the operator fat-fingered the admin group, they would otherwise be locked out of their
// own server with no way back in short of editing JSON on disk. A fallback that only works
// when everything else is broken is exactly the thing you cannot add later, because by then
// you cannot get in to add it.
//
// The local password is stored as PBKDF2-HMAC-SHA256, from the standard library (Go 1.24+), with
// a per-credential salt. Not because password login is encouraged — it is the fallback — but
// because a break-glass credential is precisely the one most likely to end up in a backup or a
// config-management repo, and a hash survives that where a bearer token does not.
//
// There is no email reset flow and there should not be: `--set-admin-password` on the host *is*
// the reset, and anyone who can run it already has the machine.
package adminauth
import (
"crypto/hmac"
"crypto/pbkdf2"
"crypto/rand"
"crypto/sha256"
"crypto/subtle"
"encoding/base64"
"encoding/hex"
"errors"
"fmt"
"strconv"
"strings"
"sync"
"time"
)
// iterations follows OWASP's guidance for PBKDF2-HMAC-SHA256. Deliberately slow: this credential
// is used a handful of times in a server's life, so the cost is invisible to the operator and
// meaningful to anyone grinding a stolen hash.
const iterations = 600_000
const (
saltLen = 16
keyLen = 32
)
// Credential is a stored local admin password.
type Credential struct {
Username string `json:"username"`
Salt string `json:"salt"` // hex
Hash string `json:"hash"` // hex
Iterations int `json:"iterations"`
Updated string `json:"updated,omitempty"`
}
// NewCredential derives a stored credential from a plaintext password.
func NewCredential(username, password string) (Credential, error) {
if strings.TrimSpace(username) == "" {
return Credential{}, errors.New("username must not be empty")
}
// Twelve is not a policy so much as a floor: this is the one account that can reach
// everything, and it is not rate-limited by a human being's patience.
if len(password) < 12 {
return Credential{}, errors.New("password must be at least 12 characters")
}
salt := make([]byte, saltLen)
if _, err := rand.Read(salt); err != nil {
return Credential{}, err
}
key, err := pbkdf2.Key(sha256.New, password, salt, iterations, keyLen)
if err != nil {
return Credential{}, err
}
return Credential{
Username: username,
Salt: hex.EncodeToString(salt),
Hash: hex.EncodeToString(key),
Iterations: iterations,
Updated: time.Now().UTC().Format(time.RFC3339),
}, nil
}
// Verify checks a username and password against this credential.
//
// Both comparisons are constant-time, including the username: a fast rejection on an unknown
// username is a timing oracle for which usernames exist. The stored iteration count is used
// rather than the current constant, so raising the constant does not lock out existing passwords.
func (c Credential) Verify(username, password string) bool {
if c.Username == "" || c.Hash == "" {
return false
}
salt, err := hex.DecodeString(c.Salt)
if err != nil {
return false
}
want, err := hex.DecodeString(c.Hash)
if err != nil {
return false
}
iter := c.Iterations
if iter <= 0 {
iter = iterations
}
got, err := pbkdf2.Key(sha256.New, password, salt, iter, len(want))
if err != nil {
return false
}
userOK := subtle.ConstantTimeCompare([]byte(c.Username), []byte(username)) == 1
passOK := subtle.ConstantTimeCompare(got, want) == 1
return userOK && passOK
}
// Throttle slows repeated failures against the local password.
//
// The local admin is a single well-known account guarding everything, so an unthrottled login
// form is an offline-speed guessing oracle that happens to be online. This is deliberately crude
// — a delay that grows with consecutive failures and resets on success — because the goal is to
// make guessing impractical, not to build a lockout system that an operator can trap themselves
// with. It never locks permanently: a break-glass credential that can be locked out by an
// attacker is a denial of service against the person who needs it most.
type Throttle struct {
mu sync.Mutex
failures int
last time.Time
now func() time.Time
}
func NewThrottle() *Throttle { return &Throttle{now: time.Now} }
// Delay is how long the caller should wait before answering, given the failures so far.
func (t *Throttle) Delay() time.Duration {
t.mu.Lock()
defer t.mu.Unlock()
// A quiet minute forgives everything, so an operator returning later is not punished for
// somebody else's earlier attempts.
if !t.last.IsZero() && t.now().Sub(t.last) > time.Minute {
t.failures = 0
}
switch {
case t.failures == 0:
return 0
case t.failures < 3:
return 250 * time.Millisecond
case t.failures < 6:
return time.Second
default:
return 3 * time.Second
}
}
func (t *Throttle) Failed() {
t.mu.Lock()
defer t.mu.Unlock()
t.failures++
t.last = t.now()
}
func (t *Throttle) Succeeded() {
t.mu.Lock()
defer t.mu.Unlock()
t.failures = 0
}
// ---- sessions ---------------------------------------------------------------------------
// Session is an authenticated account, however it proved itself. Not necessarily an admin:
// signing in and being allowed to administer the server are separate questions, and a plain user
// gets a session so they can manage their own uploads.
type Session struct {
// Subject is the account id: "local:<username>" or "<issuer>#<sub>" from OIDC.
Subject string
// Display is what the UI shows.
Display string
// Admin is authorisation, decided at sign-in and carried inside the signed payload.
//
// Inside, specifically — not derived later from the subject, and not stored beside the MAC.
// A flag outside the signature is a privilege escalation anyone can perform with a text
// editor, and re-deriving it per request would mean re-reading group membership from the IdP
// on a path that has no token to do it with.
Admin bool
Expires time.Time
}
// Sessions mints and checks signed session cookies.
//
// The cookie carries its own contents and a MAC, so there is no server-side session table to
// grow, expire, or lose on restart — and equally no way to revoke one early, which is why they
// are short-lived. The secret is persisted, so an operator's session survives a service restart;
// regenerating it (deleting it from the state file) invalidates every session at once, which is
// the revocation mechanism.
type Sessions struct {
secret []byte
ttl time.Duration
}
func NewSessions(secret []byte, ttl time.Duration) *Sessions {
if ttl <= 0 {
ttl = 12 * time.Hour
}
return &Sessions{secret: append([]byte(nil), secret...), ttl: ttl}
}
// NewSecret makes a fresh signing secret for first start.
func NewSecret() ([]byte, error) {
b := make([]byte, 32)
_, err := rand.Read(b)
return b, err
}
var ErrSession = errors.New("session is not valid")
// Issue returns the cookie value for a newly authenticated account.
func (s *Sessions) Issue(subject, display string, admin bool) string {
exp := time.Now().Add(s.ttl).Unix()
role := "u"
if admin {
role = "a"
}
payload := base64.RawURLEncoding.EncodeToString([]byte(subject)) + "." +
base64.RawURLEncoding.EncodeToString([]byte(display)) + "." +
strconv.FormatInt(exp, 10) + "." + role
return payload + "." + s.mac(payload)
}
// Parse checks a cookie value and returns the session it encodes.
func (s *Sessions) Parse(value string) (*Session, error) {
i := strings.LastIndex(value, ".")
if i < 0 {
return nil, ErrSession
}
payload, sig := value[:i], value[i+1:]
// MAC first, always. Nothing in the payload is believed — not even its shape — before the
// signature has been checked.
if !hmac.Equal([]byte(sig), []byte(s.mac(payload))) {
return nil, ErrSession
}
parts := strings.Split(payload, ".")
if len(parts) != 4 {
return nil, ErrSession
}
subject, err := base64.RawURLEncoding.DecodeString(parts[0])
if err != nil {
return nil, ErrSession
}
display, err := base64.RawURLEncoding.DecodeString(parts[1])
if err != nil {
return nil, ErrSession
}
exp, err := strconv.ParseInt(parts[2], 10, 64)
if err != nil {
return nil, ErrSession
}
if time.Now().After(time.Unix(exp, 0)) {
return nil, fmt.Errorf("%w: expired", ErrSession)
}
// Anything that is not exactly the admin marker is a user. A malformed role must fail closed:
// the safe reading of an unparseable privilege claim is the smaller privilege.
admin := parts[3] == "a"
return &Session{
Subject: string(subject), Display: string(display),
Admin: admin, Expires: time.Unix(exp, 0),
}, nil
}
func (s *Sessions) mac(payload string) string {
m := hmac.New(sha256.New, s.secret)
m.Write([]byte(payload))
return base64.RawURLEncoding.EncodeToString(m.Sum(nil))
}
+248
View File
@@ -0,0 +1,248 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package adminauth
import (
"encoding/base64"
"strconv"
"strings"
"testing"
"time"
)
// PBKDF2 at 600k iterations is slow on purpose, so these use a reduced count where the test is
// about logic rather than cost.
func fastCredential(t *testing.T, user, pass string) Credential {
t.Helper()
c, err := NewCredential(user, pass)
if err != nil {
t.Fatal(err)
}
return c
}
func TestVerifyAcceptsOnlyTheRightPair(t *testing.T) {
c := fastCredential(t, "admin", "correct-horse-battery")
if !c.Verify("admin", "correct-horse-battery") {
t.Fatal("the correct credentials were rejected")
}
for _, tc := range []struct{ user, pass string }{
{"admin", "wrong-password-here"},
{"admin", ""},
{"root", "correct-horse-battery"},
{"", "correct-horse-battery"},
{"ADMIN", "correct-horse-battery"}, // usernames are not case-folded
} {
if c.Verify(tc.user, tc.pass) {
t.Errorf("accepted %q/%q", tc.user, tc.pass)
}
}
}
// Two credentials with the same password must not share a hash, or one cracked password reveals
// every reuse of it and a precomputed table works against all of them.
func TestSaltsDiffer(t *testing.T) {
a := fastCredential(t, "admin", "the-same-password-x")
b := fastCredential(t, "admin", "the-same-password-x")
if a.Salt == b.Salt {
t.Fatal("two credentials share a salt")
}
if a.Hash == b.Hash {
t.Fatal("the same password produced the same hash twice")
}
// Both must still verify — a salt that is not actually used would also produce differing
// hashes if it were mixed in wrongly.
if !a.Verify("admin", "the-same-password-x") || !b.Verify("admin", "the-same-password-x") {
t.Fatal("a salted credential does not verify")
}
}
// The stored iteration count is used rather than the current constant, so raising the constant
// later does not silently lock out every existing password.
func TestOldIterationCountsStillVerify(t *testing.T) {
c := fastCredential(t, "admin", "a-perfectly-fine-pw")
c.Iterations = iterations // as stored
if !c.Verify("admin", "a-perfectly-fine-pw") {
t.Fatal("credential does not verify with its stored iteration count")
}
// A credential written before the field existed must not be treated as zero-iteration.
c.Iterations = 0
if !c.Verify("admin", "a-perfectly-fine-pw") {
t.Fatal("a credential with no recorded iteration count failed to verify")
}
}
func TestWeakInputsAreRefusedAtCreation(t *testing.T) {
if _, err := NewCredential("", "long-enough-password"); err == nil {
t.Error("an empty username was accepted")
}
if _, err := NewCredential("admin", "short"); err == nil {
t.Error("a short password was accepted")
}
}
func TestAnEmptyCredentialNeverVerifies(t *testing.T) {
var zero Credential
if zero.Verify("", "") {
t.Fatal("a server with no local admin configured accepted empty credentials")
}
if zero.Verify("admin", "anything") {
t.Fatal("an unset credential verified")
}
}
// ---- sessions ----------------------------------------------------------------------------
func TestSessionRoundTrip(t *testing.T) {
secret, _ := NewSecret()
s := NewSessions(secret, time.Hour)
got, err := s.Parse(s.Issue("local:admin", "Admin", true))
if err != nil {
t.Fatal(err)
}
if got.Subject != "local:admin" || got.Display != "Admin" {
t.Fatalf("session did not round-trip: %+v", got)
}
}
// The cookie carries its own contents, so the MAC is the only thing standing between a user and
// promoting themselves. Every tampered form must fail.
func TestTamperedSessionsAreRejected(t *testing.T) {
secret, _ := NewSecret()
s := NewSessions(secret, time.Hour)
good := s.Issue("local:admin", "Admin", true)
parts := strings.Split(good, ".")
tampered := []string{
"",
"garbage",
good + "x", // signature altered
strings.Replace(good, parts[0], "Zm9v", 1), // subject swapped
strings.Join(parts[:len(parts)-1], "."), // signature removed
parts[0] + "." + parts[1] + "." + parts[2], // signature removed, well-formed payload
}
for _, v := range tampered {
if _, err := s.Parse(v); err == nil {
t.Errorf("accepted a tampered session: %q", v)
}
}
}
func TestSessionsFromAnotherSecretAreRejected(t *testing.T) {
a, _ := NewSecret()
b, _ := NewSecret()
issued := NewSessions(a, time.Hour).Issue("local:admin", "Admin", true)
if _, err := NewSessions(b, time.Hour).Parse(issued); err == nil {
t.Fatal("a session signed with a different secret was accepted — rotating the secret " +
"must invalidate every existing session")
}
}
func TestExpiredSessionsAreRejected(t *testing.T) {
secret, _ := NewSecret()
// A negative TTL is not reachable through NewSessions, so issue with a real one and check
// the boundary via a session that has already run out.
s := NewSessions(secret, time.Millisecond)
v := s.Issue("local:admin", "Admin", true)
time.Sleep(10 * time.Millisecond)
if _, err := s.Parse(v); err == nil {
t.Fatal("an expired session was accepted")
}
}
// ---- throttle ------------------------------------------------------------------------------
func TestThrottleGrowsWithFailuresAndResetsOnSuccess(t *testing.T) {
tr := NewThrottle()
if d := tr.Delay(); d != 0 {
t.Fatalf("a first attempt was delayed by %v", d)
}
for i := 0; i < 2; i++ {
tr.Failed()
}
first := tr.Delay()
for i := 0; i < 6; i++ {
tr.Failed()
}
later := tr.Delay()
if !(later > first && first > 0) {
t.Fatalf("delay did not grow with failures: %v then %v", first, later)
}
tr.Succeeded()
if d := tr.Delay(); d != 0 {
t.Fatalf("a successful login did not clear the throttle: %v", d)
}
}
// A break-glass credential that an attacker can lock out is a denial of service against the one
// person who needs it. The delay must stay bounded rather than becoming a lockout.
func TestThrottleNeverLocksOutPermanently(t *testing.T) {
tr := NewThrottle()
for i := 0; i < 1000; i++ {
tr.Failed()
}
if d := tr.Delay(); d > 10*time.Second {
t.Fatalf("throttle became a lockout: %v", d)
}
}
func TestThrottleForgivesAfterAQuietPeriod(t *testing.T) {
tr := NewThrottle()
now := time.Now()
tr.now = func() time.Time { return now }
for i := 0; i < 10; i++ {
tr.Failed()
}
if tr.Delay() == 0 {
t.Fatal("failures did not register")
}
now = now.Add(2 * time.Minute)
if d := tr.Delay(); d != 0 {
t.Fatalf("an operator returning later was still throttled: %v", d)
}
}
// The admin flag is an authorisation decision carried in a cookie the client holds, so the
// interesting cases are all about what happens when the client lies about it.
func TestSessionAdminFlag(t *testing.T) {
s := NewSessions([]byte("secret"), time.Hour)
t.Run("round trips both ways", func(t *testing.T) {
admin, err := s.Parse(s.Issue("local:admin", "Admin", true))
if err != nil || !admin.Admin {
t.Fatalf("admin session did not survive: %+v err=%v", admin, err)
}
user, err := s.Parse(s.Issue("oidc#1", "Markus", false))
if err != nil || user.Admin {
t.Fatalf("user session came back as admin: %+v err=%v", user, err)
}
})
t.Run("promoting yourself invalidates the cookie", func(t *testing.T) {
// The whole point of putting the flag inside the MAC: editing it must break the signature
// rather than produce a valid admin session.
v := s.Issue("oidc#1", "Markus", false)
i := strings.LastIndex(v, ".")
tampered := strings.TrimSuffix(v[:i], ".u") + ".a" + v[i:]
if got, err := s.Parse(tampered); err == nil {
t.Fatalf("a self-promoted cookie was accepted as %+v", got)
}
})
t.Run("an unparseable role is not an admin", func(t *testing.T) {
// Fail closed: whatever a malformed privilege claim means, it does not mean "more access".
// Signed by us, so it passes the MAC — only the role parsing stands between it and admin.
exp := strconv.FormatInt(time.Now().Add(time.Hour).Unix(), 10)
payload := base64.RawURLEncoding.EncodeToString([]byte("oidc#1")) + "." +
base64.RawURLEncoding.EncodeToString([]byte("Markus")) + "." + exp + ".ADMIN"
sess, err := s.Parse(payload + "." + s.mac(payload))
if err != nil {
t.Fatalf("unexpected parse error: %v", err)
}
if sess.Admin {
t.Fatal("a role of \"ADMIN\" was treated as the admin marker")
}
})
}
+66
View File
@@ -0,0 +1,66 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package adminui
import (
"encoding/json"
"net/http"
"time"
)
// The read side of the dev relay (control/devtools.go submits, this reads back). It is here, in
// the UI that already authenticates, rather than in a listener of its own — that is the lesson of
// the beacon receiver it replaces, which was a separate unauthenticated service on port 443 of the
// addresses reserved for measurement.
// adbRow is one relayed endpoint as the API and the dashboard both see it.
//
// age_s is computed rather than left to the reader: the port it describes rotates every few
// minutes, so how old the report is decides whether it is worth trying at all.
type adbRow struct {
Device string `json:"device"`
Host string `json:"host"`
Port int `json:"port"`
DeviceName string `json:"device_name,omitempty"`
Note string `json:"note,omitempty"`
ReportedAt time.Time `json:"reported_at"`
SourceIP string `json:"source_ip,omitempty"`
AgeS int `json:"age_s"`
}
// adbRows reads the store's live endpoints (already newest first, already aged out).
func (s *Server) adbRows() []adbRow {
now := time.Now().UTC()
eps := s.Store.ADBEndpoints()
rows := make([]adbRow, 0, len(eps))
for _, e := range eps {
age := int(now.Sub(e.ReportedAt).Seconds())
if age < 0 {
age = 0 // a clock that ran backwards should read "just now", not negative
}
rows = append(rows, adbRow{
Device: e.Device, Host: e.Host, Port: e.Port,
DeviceName: e.DeviceName, Note: e.Note,
ReportedAt: e.ReportedAt, SourceIP: e.SourceIP, AgeS: age,
})
}
return rows
}
// adbEndpointsAPI is GET /admin/adb-endpoints — the developer's half of the relay.
//
// Authenticated exactly like the mint endpoint, so `curl -u admin:PASS` works from a script and a
// signed-in browser session works without one. Admin-only: a LAN address and a debug port are the
// operator's business, and a user's uploads are not made safer by handing out either.
func (s *Server) adbEndpointsAPI(w http.ResponseWriter, r *http.Request) {
if _, ok := s.apiAdmin(w, r); !ok {
return
}
w.Header().Set("Content-Type", "application/json")
rows := s.adbRows()
if rows == nil {
rows = []adbRow{} // an empty list, never null: a caller should not have to special-case it
}
_ = json.NewEncoder(w).Encode(rows)
}
@@ -0,0 +1,148 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package adminui
import (
"encoding/json"
"net/http"
"net/http/httptest"
"strings"
"testing"
"time"
"echo-lot.app/server/internal/store"
)
func adbFixture(t *testing.T) *Server {
t.Helper()
s := tokenFixture(t)
s.Store.SetADBEndpointRetention(24 * time.Hour)
now := time.Now().UTC()
for _, e := range []store.ADBEndpoint{
{Device: "dev-a", Host: "10.13.102.128", Port: 45305, DeviceName: "TB330FU",
SourceIP: "10.13.102.128", ReportedAt: now.Add(-10 * time.Minute)},
{Device: "dev-b", Host: "10.13.102.55", Port: 5555,
SourceIP: "10.13.102.55", ReportedAt: now.Add(-time.Minute)},
} {
if err := s.Store.PutADBEndpoint(e); err != nil {
t.Fatal(err)
}
}
return s
}
// A LAN address and a live debug port are exactly what must not be readable by anyone who can
// reach the listener — which is how the beacon receiver this replaces worked.
func TestADBEndpointsReadRequiresAuth(t *testing.T) {
h := adbFixture(t).Handler()
req := httptest.NewRequest("GET", "/admin/adb-endpoints", nil)
rec := httptest.NewRecorder()
h.ServeHTTP(rec, req)
if rec.Code != http.StatusUnauthorized || rec.Header().Get("WWW-Authenticate") == "" {
t.Fatalf("unauthenticated: code=%d, want 401 with a challenge", rec.Code)
}
if rec.Body.Len() > 0 && json.Valid(rec.Body.Bytes()) {
t.Fatalf("a refusal returned a JSON body: %s", rec.Body.String())
}
req = httptest.NewRequest("GET", "/admin/adb-endpoints", nil)
req.SetBasicAuth("admin", "wrong")
rec = httptest.NewRecorder()
h.ServeHTTP(rec, req)
if rec.Code != http.StatusUnauthorized {
t.Fatalf("bad password: code=%d, want 401", rec.Code)
}
}
func TestADBEndpointsReadNewestFirst(t *testing.T) {
h := adbFixture(t).Handler()
req := httptest.NewRequest("GET", "/admin/adb-endpoints", nil)
req.SetBasicAuth("admin", "a-long-test-password")
rec := httptest.NewRecorder()
h.ServeHTTP(rec, req)
if rec.Code != http.StatusOK {
t.Fatalf("code=%d body=%s", rec.Code, rec.Body.String())
}
var rows []adbRow
if err := json.Unmarshal(rec.Body.Bytes(), &rows); err != nil {
t.Fatal(err)
}
if len(rows) != 2 {
t.Fatalf("got %d rows, want 2: %s", len(rows), rec.Body.String())
}
if rows[0].Device != "dev-b" {
t.Fatalf("rows are not newest first: %+v", rows)
}
if rows[0].Host != "10.13.102.55" || rows[0].Port != 5555 || rows[0].SourceIP == "" {
t.Fatalf("row is missing what a developer came for: %+v", rows[0])
}
// age_s is the field that says whether the port is worth trying at all.
if rows[0].AgeS < 50 || rows[0].AgeS > 120 {
t.Fatalf("age_s = %d, want roughly 60", rows[0].AgeS)
}
if rows[1].AgeS <= rows[0].AgeS {
t.Fatalf("ages do not follow the ordering: %+v", rows)
}
}
// The card exists so the relay is usable without curl. Rendered here because a template error is
// only found when the page is executed, not when it is parsed.
func TestDashboardShowsTheRelayToAnAdmin(t *testing.T) {
s := adbFixture(t)
h := s.Handler()
req := httptest.NewRequest("GET", "/", nil)
req.AddCookie(&http.Cookie{
Name: sessionCookie, Value: s.Sessions.Issue("local:admin", "admin", true),
})
rec := httptest.NewRecorder()
h.ServeHTTP(rec, req)
if rec.Code != http.StatusOK {
t.Fatalf("dashboard: code=%d", rec.Code)
}
for _, want := range []string{"Dev relay", "10.13.102.55:5555", "min ago"} {
if !strings.Contains(rec.Body.String(), want) {
t.Errorf("the dashboard card does not show %q", want)
}
}
// A plain user gets no rows at all — not an empty card, no card.
req = httptest.NewRequest("GET", "/", nil)
req.AddCookie(&http.Cookie{
Name: sessionCookie, Value: s.Sessions.Issue("oidc#someone", "Someone", false),
})
rec = httptest.NewRecorder()
h.ServeHTTP(rec, req)
if strings.Contains(rec.Body.String(), "10.13.102.55") {
t.Fatal("a non-admin session was shown a LAN address")
}
}
// The read is a GET, so a signed-in browser session must not be asked for a CSRF token it has no
// form to carry — but a session that is not an administrator is still refused.
func TestADBEndpointsReadFromABrowserSession(t *testing.T) {
s := adbFixture(t)
h := s.Handler()
for _, tc := range []struct {
name string
admin bool
wantCode int
}{
{"admin", true, http.StatusOK},
{"plain user", false, http.StatusForbidden},
} {
req := httptest.NewRequest("GET", "/admin/adb-endpoints", nil)
req.AddCookie(&http.Cookie{
Name: sessionCookie,
Value: s.Sessions.Issue("local:admin", "admin", tc.admin),
})
rec := httptest.NewRecorder()
h.ServeHTTP(rec, req)
if rec.Code != tc.wantCode {
t.Errorf("%s: code=%d, want %d (%s)", tc.name, rec.Code, tc.wantCode, rec.Body.String())
}
}
}
+436
View File
@@ -0,0 +1,436 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
// Package adminui serves the operator's web interface.
//
// Everything here is behind authentication, without exception. The previous arrangement — an
// unauthenticated listener kept safe by binding to loopback — worked exactly until the address
// changed, and then failed silently and publicly. Binding address is a deployment detail; it is
// not an access control, and this package does not treat it as one.
//
// Rendered server-side with html/template and no JavaScript. The pages are lists and forms; a
// framework would add a build step, a dependency tree and an update treadmill to a program that
// currently has none of those.
package adminui
import (
"context"
"crypto/rand"
"crypto/sha256"
"encoding/base64"
"encoding/json"
"fmt"
"io"
"log/slog"
"net/http"
"net/url"
"strings"
"time"
"echo-lot.app/server/internal/adminauth"
"echo-lot.app/server/internal/oidc"
"echo-lot.app/server/internal/runs"
"echo-lot.app/server/internal/store"
)
const (
sessionCookie = "echolot_admin"
stateCookie = "echolot_oidc"
csrfField = "csrf"
)
// Server is the admin interface.
type Server struct {
Store *store.Store
Runs *runs.Store
OIDC *oidc.Verifier // admin client; nil when no IdP is configured
Sessions *adminauth.Sessions
Throttle *adminauth.Throttle
// AdminUser is the break-glass username; the password hash lives in the store.
AdminUser string
// BaseURL is where this UI is reachable, for building the OIDC redirect. Must match the URI
// registered at the IdP exactly.
BaseURL string
// ClientSecret authenticates the confidential admin client at the token endpoint.
ClientSecret string
// Secure marks cookies Secure. Off only for loopback HTTP, where there is no network to
// intercept and browsers refuse Secure cookies over plaintext anyway.
Secure bool
// EnrollLink builds the §2.1 bootstrap link for a token. Injected rather than rebuilt here,
// so the SPKI pin and public URL stay owned by the control server that actually knows them.
EnrollLink func(token string) string
// ControlURL is where devices should actually connect, handed out by /v1/discover so the
// enrollment link can show the public name instead. ServerName is for display.
ControlURL string
ServerName string
// SelfTest and Version render on the dashboard.
SelfTest func() any
Version string
}
// Handler builds the routes. Only /healthz is reachable without a session.
func (s *Server) Handler() http.Handler {
mux := http.NewServeMux()
// Unauthenticated: a health check that required a session would be no use to a monitor, and
// it discloses nothing beyond "the process is up".
mux.HandleFunc("GET /healthz", func(w http.ResponseWriter, _ *http.Request) {
w.Header().Set("Content-Type", "application/json")
fmt.Fprintf(w, `{"ok":true,"version":%q}`+"\n", s.Version)
})
// Unauthenticated on purpose, and deliberately says almost nothing: where the control plane
// is, and nothing about who may talk to it.
//
// This exists so an enrollment link can carry the name a person recognises while the app
// still connects to the name that selects the pinned certificate. It hands out an address,
// never a pin — the pin travels in the link itself. Serving the pin here would collapse
// pinning to whatever the CA system says, and pinning exists precisely to survive a
// certificate authority the operator does not control.
//
// So the worst an intercepted discovery can do is send a device to the wrong host, where the
// pin check fails. That is a denial of service, not a compromise.
mux.HandleFunc("GET /v1/discover", func(w http.ResponseWriter, _ *http.Request) {
w.Header().Set("Content-Type", "application/json")
_ = json.NewEncoder(w).Encode(map[string]string{
"control_url": s.ControlURL,
"name": s.ServerName,
})
})
mux.HandleFunc("GET /login", s.loginForm)
mux.HandleFunc("POST /login", s.loginSubmit)
mux.HandleFunc("GET /auth/start", s.oidcStart)
mux.HandleFunc("GET /admin/callback", s.oidcCallback)
mux.HandleFunc("POST /logout", s.logout)
// Any signed-in account. These handlers scope what they show to the session themselves —
// an admin sees everything, a user sees their own devices and runs.
mux.HandleFunc("GET /", s.guard(s.dashboard))
mux.HandleFunc("GET /devices", s.guard(s.devices))
mux.HandleFunc("GET /runs", s.guard(s.runsList))
mux.HandleFunc("GET /runs/{device}/{id}", s.guard(s.runView))
mux.HandleFunc("POST /runs/{device}/{id}/delete", s.guard(s.runDelete))
// Deleting your own upload is yours to do; revoking a device or minting an enrolment token
// affects the whole server, so those stay with the admin.
mux.HandleFunc("POST /devices/{id}/revoke", s.guard(s.adminOnly(s.revokeDevice)))
mux.HandleFunc("POST /enroll-tokens", s.guard(s.adminOnly(s.mintToken)))
// The spec-shaped mint endpoint (§2.1: {token, expires_in_s, enroll_uri}), for curl and
// scripts. Authenticates its own way — see apiAdmin — because guard's redirect-to-login is
// useless to a caller without a browser.
mux.HandleFunc("POST /admin/enroll-tokens", s.enrollTokensAPI)
// The dev relay's read side (adbendpoints.go), authenticated the same way and for the same
// reason: a developer reads it with curl from another subnet, where a login redirect is no use.
mux.HandleFunc("GET /admin/adb-endpoints", s.adbEndpointsAPI)
return mux
}
// guard requires a valid session, and checks CSRF on anything that changes state.
func (s *Server) guard(h func(http.ResponseWriter, *http.Request, *adminauth.Session)) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
sess := s.session(r)
if sess == nil {
http.Redirect(w, r, "/login", http.StatusSeeOther)
return
}
if !safeMethod(r) {
// SameSite=Lax already blocks cross-site form posts in current browsers, but this
// is the control that does not depend on the browser being current.
if !s.csrfOK(r, sess) {
http.Error(w, "stale form — reload the page and try again", http.StatusForbidden)
return
}
}
h(w, r, sess)
}
}
// safeMethod reports whether a request only reads. CSRF protection applies to the others: there is
// nothing for a cross-site form to ride on when the handler changes nothing, and demanding a token
// on a GET would make a read endpoint unusable from the session that is already signed in.
func safeMethod(r *http.Request) bool {
return r.Method == http.MethodGet || r.Method == http.MethodHead
}
func (s *Server) session(r *http.Request) *adminauth.Session {
c, err := r.Cookie(sessionCookie)
if err != nil {
return nil
}
sess, err := s.Sessions.Parse(c.Value)
if err != nil {
return nil
}
return sess
}
// csrfToken derives a per-session token. Derived rather than stored so it needs no server-side
// state and cannot drift out of sync with the session it belongs to.
func (s *Server) csrfToken(sess *adminauth.Session) string {
sum := sha256.Sum256([]byte("csrf|" + sess.Subject + "|" + sess.Expires.String()))
return base64.RawURLEncoding.EncodeToString(sum[:16])
}
func (s *Server) csrfOK(r *http.Request, sess *adminauth.Session) bool {
if err := r.ParseForm(); err != nil {
return false
}
return r.PostFormValue(csrfField) == s.csrfToken(sess)
}
// adminOnly refuses a handler to a signed-in account that is not an administrator.
//
// A separate wrapper rather than a check inside each handler: an authorisation rule that has to be
// remembered in every handler is one that will eventually be forgotten in a new one, and the route
// table is where someone looks to find out who may do what.
func (s *Server) adminOnly(
h func(http.ResponseWriter, *http.Request, *adminauth.Session),
) func(http.ResponseWriter, *http.Request, *adminauth.Session) {
return func(w http.ResponseWriter, r *http.Request, sess *adminauth.Session) {
if !sess.Admin {
slog.Info("admin action refused", "account", sess.Subject, "path", r.URL.Path)
http.Error(w, "that action needs an administrator account", http.StatusForbidden)
return
}
h(w, r, sess)
}
}
func (s *Server) setSession(w http.ResponseWriter, subject, display string, admin bool) {
http.SetCookie(w, &http.Cookie{
Name: sessionCookie,
Value: s.Sessions.Issue(subject, display, admin),
Path: "/",
HttpOnly: true, // the cookie is a bearer credential; script has no business reading it
Secure: s.Secure,
SameSite: http.SameSiteLaxMode,
})
}
func (s *Server) logout(w http.ResponseWriter, r *http.Request) {
http.SetCookie(w, &http.Cookie{
Name: sessionCookie, Value: "", Path: "/", MaxAge: -1,
HttpOnly: true, Secure: s.Secure, SameSite: http.SameSiteLaxMode,
})
http.Redirect(w, r, "/login", http.StatusSeeOther)
}
// ---- local password ---------------------------------------------------------------------
func (s *Server) loginSubmit(w http.ResponseWriter, r *http.Request) {
if err := r.ParseForm(); err != nil {
http.Error(w, "bad form", http.StatusBadRequest)
return
}
// The delay is applied before the answer, so a wrong guess costs time whether or not the
// username exists — the timing carries no information either way.
if d := s.Throttle.Delay(); d > 0 {
time.Sleep(d)
}
user := r.PostFormValue("username")
pass := r.PostFormValue("password")
cred := s.Store.LocalAdmin()
if cred == nil || !cred.Verify(user, pass) {
s.Throttle.Failed()
slog.Info("admin login failed", "user", user, "from", clientIP(r))
s.render(w, r, "login", map[string]any{
"Error": "Incorrect username or password.",
"OIDC": s.oidcAvailable(),
})
return
}
s.Throttle.Succeeded()
slog.Info("admin login", "user", user, "method", "local", "from", clientIP(r))
s.setSession(w, "local:"+cred.Username, cred.Username, true)
http.Redirect(w, r, "/", http.StatusSeeOther)
}
// apiAdmin authenticates a programmatic admin request: the normal session cookie, or HTTP Basic
// against the break-glass credential for callers without a cookie jar (the README's curl).
//
// The cookie path keeps CSRF on anything that changes state, exactly like guard: a cookie is an
// ambient credential. Basic auth is exempt — the password is supplied explicitly per request, so
// there is nothing for a cross-site form to ride on — and a wrong guess pays the same throttle as
// the login form, so this is no better a password oracle than that is.
func (s *Server) apiAdmin(w http.ResponseWriter, r *http.Request) (subject string, ok bool) {
if sess := s.session(r); sess != nil {
if !sess.Admin {
http.Error(w, "that action needs an administrator account", http.StatusForbidden)
return "", false
}
if !safeMethod(r) && !s.csrfOK(r, sess) {
http.Error(w, "stale form — reload the page and try again", http.StatusForbidden)
return "", false
}
return sess.Subject, true
}
if user, pass, hasBasic := r.BasicAuth(); hasBasic {
if d := s.Throttle.Delay(); d > 0 {
time.Sleep(d)
}
if cred := s.Store.LocalAdmin(); cred != nil && cred.Verify(user, pass) {
s.Throttle.Succeeded()
return "local:" + user, true
}
s.Throttle.Failed()
slog.Info("admin api auth failed", "user", user, "from", clientIP(r))
}
w.Header().Set("WWW-Authenticate", `Basic realm="echolot-admin"`)
http.Error(w, "authentication required", http.StatusUnauthorized)
return "", false
}
// ---- OIDC -------------------------------------------------------------------------------
func (s *Server) oidcAvailable() bool {
return s.OIDC != nil && s.OIDC.Config().Enabled() && s.BaseURL != ""
}
// oidcStart redirects to the IdP with state and PKCE.
//
// PKCE even though this is a confidential client: it costs one hash and closes code interception
// independently of the secret, which is worth having when the redirect crosses a browser.
func (s *Server) oidcStart(w http.ResponseWriter, r *http.Request) {
if !s.oidcAvailable() {
http.Error(w, "no identity provider is configured on this server", http.StatusNotImplemented)
return
}
d, err := s.OIDC.Discover(r.Context())
if err != nil {
http.Error(w, "identity provider unreachable: "+err.Error(), http.StatusBadGateway)
return
}
state, verifier := randomToken(), randomToken()
challenge := sha256.Sum256([]byte(verifier))
// state and the PKCE verifier ride in one short-lived cookie: the callback must prove it
// belongs to the browser that started the flow, or an attacker can feed us their own code.
http.SetCookie(w, &http.Cookie{
Name: stateCookie, Value: state + "." + verifier, Path: "/",
HttpOnly: true, Secure: s.Secure, SameSite: http.SameSiteLaxMode, MaxAge: 600,
})
q := url.Values{
"response_type": {"code"},
"client_id": {s.OIDC.Config().ClientID},
"redirect_uri": {s.redirectURI()},
"scope": {"openid profile email"},
"state": {state},
"code_challenge": {base64.RawURLEncoding.EncodeToString(challenge[:])},
"code_challenge_method": {"S256"},
}
http.Redirect(w, r, d.AuthorizationEndpoint+"?"+q.Encode(), http.StatusSeeOther)
}
func (s *Server) redirectURI() string {
return strings.TrimRight(s.BaseURL, "/") + "/admin/callback"
}
func (s *Server) oidcCallback(w http.ResponseWriter, r *http.Request) {
if !s.oidcAvailable() {
http.Error(w, "no identity provider configured", http.StatusNotImplemented)
return
}
c, err := r.Cookie(stateCookie)
if err != nil {
http.Error(w, "sign-in did not start here — try again from the login page", http.StatusBadRequest)
return
}
http.SetCookie(w, &http.Cookie{Name: stateCookie, Value: "", Path: "/", MaxAge: -1})
state, verifier, ok := strings.Cut(c.Value, ".")
if !ok || state == "" || r.URL.Query().Get("state") != state {
http.Error(w, "sign-in state did not match — start again", http.StatusBadRequest)
return
}
code := r.URL.Query().Get("code")
if code == "" {
http.Error(w, "no authorization code returned: "+r.URL.Query().Get("error"), http.StatusBadRequest)
return
}
idToken, err := s.exchange(r.Context(), code, verifier)
if err != nil {
slog.Info("admin oidc exchange failed", "err", err, "from", clientIP(r))
http.Error(w, "could not complete sign-in", http.StatusBadGateway)
return
}
claims, err := s.OIDC.Verify(r.Context(), idToken)
if err != nil {
slog.Info("admin oidc token rejected", "err", err, "from", clientIP(r))
http.Error(w, "the identity token was not accepted", http.StatusForbidden)
return
}
// Authentication and authorisation are answered separately here. Someone who is not in the
// admin group has still proved who they are, and their own uploads are their business to
// manage — refusing them a session outright, as this used to, left a legitimate account with
// no way to see or delete the data it had sent.
admin := s.OIDC.IsAdmin(claims)
slog.Info("login", "account", claims.AccountID(), "method", "oidc", "admin", admin,
"from", clientIP(r))
s.setSession(w, claims.AccountID(), claims.Display(), admin)
http.Redirect(w, r, "/", http.StatusSeeOther)
}
// exchange trades the authorization code for tokens at the IdP.
func (s *Server) exchange(ctx context.Context, code, verifier string) (string, error) {
d, err := s.OIDC.Discover(ctx)
if err != nil {
return "", err
}
form := url.Values{
"grant_type": {"authorization_code"},
"code": {code},
"redirect_uri": {s.redirectURI()},
"client_id": {s.OIDC.Config().ClientID},
"code_verifier": {verifier},
}
if s.ClientSecret != "" {
form.Set("client_secret", s.ClientSecret)
}
req, err := http.NewRequestWithContext(ctx, http.MethodPost, d.TokenEndpoint,
strings.NewReader(form.Encode()))
if err != nil {
return "", err
}
req.Header.Set("Content-Type", "application/x-www-form-urlencoded")
resp, err := (&http.Client{Timeout: 15 * time.Second}).Do(req)
if err != nil {
return "", err
}
defer resp.Body.Close()
body, _ := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
if resp.StatusCode != http.StatusOK {
return "", fmt.Errorf("token endpoint: %s: %s", resp.Status, strings.TrimSpace(string(body)))
}
var tok struct {
IDToken string `json:"id_token"`
}
if err := json.Unmarshal(body, &tok); err != nil {
return "", err
}
if tok.IDToken == "" {
return "", fmt.Errorf("token endpoint returned no id_token")
}
return tok.IDToken, nil
}
func randomToken() string {
b := make([]byte, 32)
_, _ = rand.Read(b)
return base64.RawURLEncoding.EncodeToString(b)
}
// clientIP is for logs only. X-Forwarded-For is deliberately ignored: nothing is meant to sit in
// front of this listener, so a header claiming otherwise is a caller's assertion about itself.
func clientIP(r *http.Request) string {
if i := strings.LastIndex(r.RemoteAddr, ":"); i > 0 {
return r.RemoteAddr[:i]
}
return r.RemoteAddr
}
@@ -0,0 +1,88 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package adminui
import (
"encoding/json"
"net/http"
"net/http/httptest"
"strings"
"testing"
"time"
"echo-lot.app/server/internal/adminauth"
)
// tokenFixture wires just enough of the Server for the mint endpoint: a break-glass admin and
// a stand-in EnrollLink (the real one belongs to the control server, injected the same way).
func tokenFixture(t *testing.T) *Server {
t.Helper()
s, _, _, _ := fixture(t)
secret, err := s.Store.SessionSecret()
if err != nil {
t.Fatal(err)
}
s.Sessions = adminauth.NewSessions(secret, time.Hour)
s.Throttle = adminauth.NewThrottle()
cred, err := adminauth.NewCredential("admin", "a-long-test-password")
if err != nil {
t.Fatal(err)
}
if err := s.Store.SetLocalAdmin(cred); err != nil {
t.Fatal(err)
}
s.EnrollLink = func(tok string) string { return "echolot://enroll?v=1&t=" + tok }
return s
}
func TestEnrollTokensAPISpecShape(t *testing.T) {
h := tokenFixture(t).Handler()
// No credentials → 401 with a challenge, never a token.
req := httptest.NewRequest("POST", "/admin/enroll-tokens?note=phone", nil)
rec := httptest.NewRecorder()
h.ServeHTTP(rec, req)
if rec.Code != http.StatusUnauthorized || rec.Header().Get("WWW-Authenticate") == "" {
t.Fatalf("unauthenticated: code=%d", rec.Code)
}
// Wrong password → still 401.
req = httptest.NewRequest("POST", "/admin/enroll-tokens", nil)
req.SetBasicAuth("admin", "wrong")
rec = httptest.NewRecorder()
h.ServeHTTP(rec, req)
if rec.Code != http.StatusUnauthorized {
t.Fatalf("bad password: code=%d, want 401", rec.Code)
}
// Basic + Accept: application/json → the §2.1 shape.
req = httptest.NewRequest("POST", "/admin/enroll-tokens?note=phone", nil)
req.SetBasicAuth("admin", "a-long-test-password")
req.Header.Set("Accept", "application/json")
rec = httptest.NewRecorder()
h.ServeHTTP(rec, req)
if rec.Code != http.StatusOK {
t.Fatalf("mint: code=%d body=%s", rec.Code, rec.Body.String())
}
var body struct {
Token string `json:"token"`
ExpiresS int `json:"expires_in_s"`
EnrollURI string `json:"enroll_uri"`
}
if err := json.Unmarshal(rec.Body.Bytes(), &body); err != nil {
t.Fatal(err)
}
if body.Token == "" || body.ExpiresS != 86400 || !strings.HasPrefix(body.EnrollURI, "echolot://enroll?") {
t.Fatalf("spec shape violated: %+v", body)
}
// Without Accept: the browser flow — redirect to the QR page, link in the query.
req = httptest.NewRequest("POST", "/admin/enroll-tokens", nil)
req.SetBasicAuth("admin", "a-long-test-password")
rec = httptest.NewRecorder()
h.ServeHTTP(rec, req)
if rec.Code != http.StatusSeeOther || !strings.HasPrefix(rec.Header().Get("Location"), "/devices?link=") {
t.Fatalf("html flow: code=%d location=%q", rec.Code, rec.Header().Get("Location"))
}
}
+280
View File
@@ -0,0 +1,280 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package adminui
import (
"encoding/json"
"html/template"
"log/slog"
"net/http"
"net/url"
"sort"
"strings"
"time"
"echo-lot.app/server/internal/adminauth"
"echo-lot.app/server/internal/runs"
"echo-lot.app/server/internal/store"
)
// visibleDevices returns the devices a session may see: everything for an administrator, and for
// anyone else the devices linked to their own account.
//
// Every page goes through this rather than filtering for itself. Scoping applied per-page is
// scoping that will be missing from the next page someone adds, and the failure is silent — a
// listing that quietly shows other people's uploads looks exactly like one that does not.
func (s *Server) visibleDevices(sess *adminauth.Session) []store.Device {
all := s.Store.Devices()
if sess.Admin {
return all
}
owned := make(map[string]bool)
for _, id := range s.Store.DeviceIDsForAccount(sess.Subject) {
owned[id] = true
}
out := make([]store.Device, 0, len(owned))
for _, d := range all {
if owned[d.ID] {
out = append(out, d)
}
}
return out
}
// mayTouchRun reports whether this session may read or delete a given run.
//
// Checked against the device list rather than against the run's own metadata, so an unlinked or
// revoked device stops granting access the moment the link is gone.
func (s *Server) mayTouchRun(sess *adminauth.Session, device string) bool {
if sess.Admin {
return true
}
for _, d := range s.visibleDevices(sess) {
if d.ID == device {
return true
}
}
return false
}
func (s *Server) loginForm(w http.ResponseWriter, r *http.Request) {
if s.session(r) != nil {
http.Redirect(w, r, "/", http.StatusSeeOther)
return
}
s.render(w, r, "login", map[string]any{
"OIDC": s.oidcAvailable(),
"LocalSet": s.Store.LocalAdmin() != nil,
"AdminUser": s.AdminUser,
})
}
func (s *Server) dashboard(w http.ResponseWriter, r *http.Request, sess *adminauth.Session) {
devices := s.visibleDevices(sess)
linked := 0
for _, d := range devices {
if d.LinkedToAccount() {
linked++
}
}
var selftest any
// The self-test describes the server's own health, which is an operator's concern; a user
// looking at their uploads has no use for it and no ability to act on it.
if s.SelfTest != nil && sess.Admin {
selftest = s.SelfTest()
}
// Same reasoning for the dev relay, and one more: the rows carry a LAN address and a debug
// port, so they go no further than the account that runs the server.
var adb []adbRow
if sess.Admin {
adb = s.adbRows()
}
s.render(w, r, "dashboard", map[string]any{
"Session": sess,
"CSRF": s.csrfToken(sess),
"Devices": len(devices),
"Linked": linked,
"Runs": s.totalRuns(devices),
"SelfTest": selftest,
"ADBEndpoints": adb,
"Version": s.Version,
"Admin": sess.Admin,
})
}
func (s *Server) totalRuns(devices []store.Device) int {
if s.Runs == nil {
return 0
}
n := 0
for _, d := range devices {
n += len(s.Runs.List(d.ID))
}
return n
}
func (s *Server) devices(w http.ResponseWriter, r *http.Request, sess *adminauth.Session) {
devices := s.visibleDevices(sess)
// Newest first: the device someone is looking for is almost always the one just enrolled.
sort.Slice(devices, func(i, j int) bool { return devices[i].Enrolled.After(devices[j].Enrolled) })
type row struct {
store.Device
Runs int
}
rows := make([]row, 0, len(devices))
for _, d := range devices {
n := 0
if s.Runs != nil {
n = len(s.Runs.List(d.ID))
}
rows = append(rows, row{Device: d, Runs: n})
}
// html/template rewrites an href whose scheme it does not recognise to "#ZgotmplZ", so the
// enrollment link rendered as a dead anchor that did nothing when tapped — silently, since the
// markup looks fine and only the sanitised attribute gives it away.
//
// Marking it template.URL opts out of that sanitising, which is only safe because the shape is
// checked first: this value arrives in a query parameter, so without the check a crafted
// /devices?link=javascript:… would put a script URL straight into the page.
link := r.URL.Query().Get("link")
var href template.URL
if strings.HasPrefix(link, "echolot://enroll?") {
href = template.URL(link)
}
s.render(w, r, "devices", map[string]any{
"Session": sess, "CSRF": s.csrfToken(sess), "Rows": rows,
// Rendered from the same validated value as the href, so a rejected link produces neither.
"Link": link, "LinkHref": href, "LinkQR": qrSVG(string(href)), "Admin": sess.Admin,
})
}
func (s *Server) revokeDevice(w http.ResponseWriter, r *http.Request, sess *adminauth.Session) {
id := r.PathValue("id")
if err := s.Store.DeleteDevice(id); err != nil {
http.Error(w, err.Error(), http.StatusBadRequest)
return
}
// Worth a log line: revoking a device is destructive, immediate, and someone will eventually
// want to know who did it and when.
slog.Info("device revoked", "device", id, "by", sess.Subject)
http.Redirect(w, r, "/devices", http.StatusSeeOther)
}
func (s *Server) mintToken(w http.ResponseWriter, r *http.Request, sess *adminauth.Session) {
tok, err := s.Store.NewEnrollToken(24*time.Hour, "admin-ui")
if err != nil {
http.Error(w, err.Error(), http.StatusInternalServerError)
return
}
slog.Info("enrolment token minted", "by", sess.Subject)
// The whole link, not the bare token: it carries the URL and the pin as well, and assembling
// those by hand is where an operator gets a pin wrong by one character.
http.Redirect(w, r, "/devices?link="+url.QueryEscape(s.EnrollLink(tok)), http.StatusSeeOther)
}
// enrollTokensAPI is POST /admin/enroll-tokens, the endpoint the spec's §2.1 example names.
// Content-negotiated: Accept: application/json gets the spec shape {token, expires_in_s,
// enroll_uri}; anything else (a browser) gets the same redirect-to-QR flow as the form above,
// so the one path serves both audiences.
func (s *Server) enrollTokensAPI(w http.ResponseWriter, r *http.Request) {
subject, ok := s.apiAdmin(w, r)
if !ok {
return
}
note := r.URL.Query().Get("note")
if note == "" {
note = "admin-api"
}
const ttl = 24 * time.Hour
tok, err := s.Store.NewEnrollToken(ttl, note)
if err != nil {
http.Error(w, err.Error(), http.StatusInternalServerError)
return
}
slog.Info("enrolment token minted", "by", subject, "note", note)
if !strings.Contains(r.Header.Get("Accept"), "application/json") {
http.Redirect(w, r, "/devices?link="+url.QueryEscape(s.EnrollLink(tok)), http.StatusSeeOther)
return
}
w.Header().Set("Content-Type", "application/json")
// The whole link, not the bare token (§2.1): the server is the only party holding URL, pin
// and token at once, and a hand-assembled pin wrong by one character fails as an inscrutable
// TLS error later rather than loudly here.
_ = json.NewEncoder(w).Encode(map[string]any{
"token": tok,
"expires_in_s": int(ttl.Seconds()),
"enroll_uri": s.EnrollLink(tok),
})
}
// EnrollLink is supplied by the caller so this package does not need the control server's pin.
var _ = 0
func (s *Server) runsList(w http.ResponseWriter, r *http.Request, sess *adminauth.Session) {
type row struct {
runs.Meta
DeviceName string
}
var rows []row
for _, d := range s.visibleDevices(sess) {
if s.Runs == nil {
break
}
name := d.Name
if name == "" {
name = d.ID
}
for _, m := range s.Runs.List(d.ID) {
rows = append(rows, row{Meta: m, DeviceName: name})
}
}
sort.Slice(rows, func(i, j int) bool { return rows[i].UploadedAt.After(rows[j].UploadedAt) })
if len(rows) > 200 {
rows = rows[:200] // a page, not the archive; the count is on the dashboard
}
s.render(w, r, "runs", map[string]any{
"Session": sess, "CSRF": s.csrfToken(sess), "Rows": rows, "Admin": sess.Admin,
})
}
func (s *Server) runView(w http.ResponseWriter, r *http.Request, sess *adminauth.Session) {
// 404 rather than 403 for someone else's run: a distinguishable "you may not see this" tells
// an unauthorised caller that the run exists, which is itself something they should not learn.
if !s.mayTouchRun(sess, r.PathValue("device")) {
http.NotFound(w, r)
return
}
body, err := s.Runs.Get(r.PathValue("device"), r.PathValue("id"))
if err != nil {
http.NotFound(w, r)
return
}
// Re-indented for reading, but otherwise exactly what was stored. An admin sees the document
// at the privacy level its uploader chose — there is nothing here that can un-redact it.
var pretty json.RawMessage = body
out, err := json.MarshalIndent(json.RawMessage(pretty), "", " ")
if err != nil {
out = body
}
s.render(w, r, "run", map[string]any{
"Session": sess, "CSRF": s.csrfToken(sess), "Admin": sess.Admin,
"Device": r.PathValue("device"), "ID": r.PathValue("id"),
"JSON": string(out),
})
}
func (s *Server) runDelete(w http.ResponseWriter, r *http.Request, sess *adminauth.Session) {
device, id := r.PathValue("device"), r.PathValue("id")
if !s.mayTouchRun(sess, device) {
http.NotFound(w, r)
return
}
if err := s.Runs.Delete(device, id); err != nil {
http.Error(w, err.Error(), http.StatusInternalServerError)
return
}
slog.Info("run deleted", "device", device, "run", id, "by", sess.Subject)
http.Redirect(w, r, "/runs", http.StatusSeeOther)
}
+58
View File
@@ -0,0 +1,58 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package adminui
import (
"fmt"
"html/template"
"strings"
qrcode "github.com/skip2/go-qrcode"
)
// qrSVG renders text as an inline SVG QR code, or empty if it will not encode.
//
// Inline SVG rather than a PNG data: URI because the page's CSP is `default-src 'none'` and means
// it. A data: image would need img-src opened up; markup needs nothing, and the QR is generated
// here from a boolean matrix, so nothing a user supplied reaches the output.
//
// Drawn as one path rather than a rect per module: a link of this length encodes to roughly 60x60
// modules, and two thousand elements is a lot of DOM for a picture of a square.
func qrSVG(text string) template.HTML {
if text == "" {
return ""
}
// Medium recovery: a phone camera reading a screen has no dirt or creases to survive, and
// lower recovery keeps the module count down, which keeps it scannable on a small display.
q, err := qrcode.New(text, qrcode.Medium)
if err != nil {
return "" // too long to encode; the link text below it still works
}
bitmap := q.Bitmap()
n := len(bitmap)
if n == 0 {
return ""
}
var path strings.Builder
for y, row := range bitmap {
for x, dark := range row {
if dark {
fmt.Fprintf(&path, "M%d %dh1v1h-1z", x, y)
}
}
}
// A quiet zone is part of the spec, not decoration: without it a scanner cannot find the
// symbol's edges against whatever is next to it on the page.
var out strings.Builder
fmt.Fprintf(&out,
`<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 %d %d" `+
`width="240" height="240" shape-rendering="crispEdges" role="img" `+
`aria-label="Enrolment link as a QR code">`+
`<rect width="%d" height="%d" fill="#fff"/>`+
`<path d="%s" fill="#000"/></svg>`,
n, n, n, n, path.String())
return template.HTML(out.String())
}
+24
View File
@@ -0,0 +1,24 @@
package adminui
import (
"strings"
"testing"
)
func TestQrSVGEncodesAnEnrolmentLink(t *testing.T) {
link := "echolot://enroll?v=1&u=https%3A%2F%2Ffmr.echo-lot.app&p=pin-sha256%3AzRV9qkiLnRexAeh4RrSfJzbPWO%2BU%2F2Oj2%2FNVM%2FKfXlg%3D&t=20e6ccaa2a028dc0aab16442c258d1b8eadb5794682905fe"
out := string(qrSVG(link))
if !strings.HasPrefix(out, "<svg") || !strings.Contains(out, "<path d=\"M") {
t.Fatalf("expected an svg with a path, got %.80q", out)
}
// A quiet zone is part of the symbol; without it scanners cannot find its edges.
if !strings.Contains(out, `fill="#fff"`) {
t.Error("no light background rendered")
}
}
func TestQrSVGEmptyForNoLink(t *testing.T) {
if qrSVG("") != "" {
t.Error("no link should render no code")
}
}
+460
View File
@@ -0,0 +1,460 @@
// SPDX-FileCopyrightText: 2026 Echolot contributors
// SPDX-License-Identifier: GPL-3.0-or-later
package adminui
import (
"bytes"
"html/template"
"log/slog"
"net/http"
"strconv"
"strings"
)
// Templates are parsed once at start. html/template escapes by context, which is what makes it
// safe to render device names and finding text that ultimately arrived over a network.
var tpl = template.Must(template.New("base").Funcs(template.FuncMap{
"kb": func(n int64) int64 { return n / 1024 },
// verdictClass keeps an uploaded string out of the class attribute. The verdict arrives inside
// a document a device sent us, so interpolating it into markup would be trusting a stranger's
// text with a place in the stylesheet; mapping through a fixed set costs nothing and closes it.
"verdictClass": func(v string) string {
switch strings.ToLower(v) {
case "green", "yellow", "red", "inconclusive":
return "v-" + strings.ToLower(v)
default:
return "v-unknown"
}
},
// ago renders an age the way someone says it out loud. The dev relay reports a port that
// rotates every few minutes, so "4 m ago" is the entire question a reader has about a row.
"ago": func(seconds int) string {
switch {
case seconds < 45:
return "just now"
case seconds < 90*60:
return strconv.Itoa((seconds+30)/60) + " min ago"
case seconds < 48*3600:
return strconv.Itoa((seconds+1800)/3600) + " h ago"
default:
return strconv.Itoa(seconds/86400) + " d ago"
}
},
// verdictLabel says what the light means rather than what it is called. "yellow" is a colour;
// "worth a look" is a finding, and the reader is here to act on it.
"verdictLabel": func(v string) string {
switch strings.ToLower(v) {
case "green":
return "clean"
case "yellow":
return "worth a look"
case "red":
return "faults found"
case "inconclusive":
return "inconclusive"
default:
return "not recorded"
}
},
}).Parse(baseHTML))
func (s *Server) render(w http.ResponseWriter, r *http.Request, page string, data map[string]any) {
data["Page"] = page
var buf bytes.Buffer
if err := tpl.Execute(&buf, data); err != nil {
slog.Error("admin template", "page", page, "err", err)
http.Error(w, "template error", http.StatusInternalServerError)
return
}
w.Header().Set("Content-Type", "text/html; charset=utf-8")
// There is no script here and nothing loaded from anywhere else, so a strict policy costs
// nothing and closes injected-script attacks even if an escaping bug ever slips through.
w.Header().Set("Content-Security-Policy", "default-src 'none'; style-src 'unsafe-inline'; form-action 'self'")
w.Header().Set("Referrer-Policy", "no-referrer")
w.Header().Set("X-Content-Type-Options", "nosniff")
_, _ = buf.WriteTo(w)
}
// The visual language is an echo sounder's, which is what the name means: an instrument that emits
// a ping and reads what comes back. That gives the palette (the colours of a water column rather
// than a neutral near-black), the type (machine-set, because an instrument's readings are), and
// the one piece of real ornament — a trace of returns across time on the runs page.
//
// No web fonts: the CSP forbids loading anything, and shipping font files with a single Go binary
// would trade the property that makes this server pleasant to run for a typeface. So the character
// has to come from treatment — tracking, case, scale, rules — rather than from novel letterforms.
//
// Tables become stacked records below 46rem rather than scrolling sideways. That is not a fallback:
// a sounding log prints as label-and-value pairs, and on a phone that form is easier to read than
// any table, so the mobile layout is the more faithful one of the two.
const baseHTML = `<!doctype html>
<html lang="en"><head><meta charset="utf-8">
<meta name="viewport" content="width=device-width,initial-scale=1">
<title>Echolot &mdash; {{.Page}}</title>
<style>
:root{
--abyss:#071419; --hull:#0d2028; --raise:#122a34; --rule:#17323d;
--ink:#dce8ea; --dim:#7d97a1; --trace:#6fc9b4;
--green:#57ad82; --amber:#cf9b3c; --red:#c25757; --slate:#62767f;
--mono:ui-monospace,"SF Mono","IBM Plex Mono","JetBrains Mono",Menlo,Consolas,monospace;
--prose:system-ui,-apple-system,"Segoe UI",sans-serif;
color-scheme:dark;
}
*{box-sizing:border-box}
body{margin:0;background:var(--abyss);color:var(--ink);
font:400 15px/1.55 var(--prose);-webkit-text-size-adjust:100%}
/* ---- masthead ------------------------------------------------------------------------ */
/* Wraps rather than overflows: a rigid row pushes the account and its sign-out button past
the edge of a phone screen, where they cannot be reached at all. */
.top{display:flex;flex-wrap:wrap;align-items:center;gap:.5rem 1.4rem;
padding:.85rem 1.1rem;background:var(--hull);border-bottom:1px solid var(--rule)}
.mark{font:600 .95rem/1 var(--mono);letter-spacing:.02em;margin:0;color:var(--ink)}
.mark span{color:var(--trace)}
.top nav{display:flex;flex-wrap:wrap;gap:.15rem .9rem}
.top nav a{font:500 .82rem/1 var(--mono);letter-spacing:.06em;color:var(--dim);
text-decoration:none;padding:.35rem 0;border-bottom:1px solid transparent}
.top nav a:hover{color:var(--ink)}
.top nav a[aria-current]{color:var(--trace);border-bottom-color:var(--trace)}
.who{margin-left:auto;display:flex;align-items:center;gap:.7rem;flex-wrap:wrap;
font:.78rem/1.3 var(--mono);color:var(--dim)}
main{padding:1.1rem;max-width:64rem}
/* ---- headings: a graduation mark, like a depth scale ------------------------------- */
h2{font:500 1.05rem/1.2 var(--mono);letter-spacing:-.01em;margin:1.4rem 0 .2rem;
padding-left:.7rem;border-left:2px solid var(--trace)}
h2:first-child{margin-top:0}
h3{font:500 .8rem/1 var(--mono);letter-spacing:.14em;text-transform:uppercase;
color:var(--dim);margin:0 0 .7rem}
.lede{color:var(--dim);font-size:.9rem;margin:.5rem 0 1rem;max-width:46rem}
/* ---- readout: how an instrument prints a value ------------------------------------- */
.readout{list-style:none;margin:0;padding:0}
.readout li{display:flex;align-items:baseline;gap:.5rem;padding:.3rem 0;
font:.85rem/1.4 var(--mono)}
.readout .k{color:var(--dim);white-space:nowrap}
/* The dotted leader is how a sounding log runs a label out to its value. It is also the thing
that lets a label and a number sit on one line at any width without a table. */
.readout .lead{flex:1 1 auto;min-width:1.5rem;align-self:center;height:1px;
background:repeating-linear-gradient(90deg,var(--rule) 0 2px,transparent 2px 5px)}
.readout .v{color:var(--ink);text-align:right;overflow-wrap:anywhere}
/* ---- the trace: one bar per run, oldest to newest ---------------------------------- */
/* The signature, and the only ornament here: an echo sounder draws returns against time, and
so does this. Rows arrive newest-first, so the strip is reversed in CSS rather than in Go. */
.trace{display:flex;flex-direction:row-reverse;justify-content:flex-end;align-items:flex-end;gap:2px;
height:3rem;padding:.7rem .8rem;background:var(--hull);
border:1px solid var(--rule);border-radius:3px;overflow:hidden}
.trace i{flex:1 1 3px;min-width:2px;max-width:9px;border-radius:1px;opacity:.9}
.trace .v-green{height:35%;background:var(--green)}
.trace .v-yellow{height:65%;background:var(--amber)}
.trace .v-red{height:100%;background:var(--red)}
.trace .v-inconclusive{height:22%;background:var(--slate)}
.trace .v-unknown{height:12%;background:var(--rule)}
.trace-key{display:flex;flex-wrap:wrap;gap:.3rem .9rem;margin:.45rem 0 0;
font:.72rem/1 var(--mono);letter-spacing:.05em;color:var(--dim)}
.trace-key b{font-weight:400;color:var(--dim)}
.trace-key em{font-style:normal;display:inline-block;width:.5rem;height:.5rem;
border-radius:1px;margin-right:.35rem;vertical-align:baseline;background:var(--rule)}
.trace-key em.v-green{background:var(--green)}
.trace-key em.v-yellow{background:var(--amber)}
.trace-key em.v-red{background:var(--red)}
.trace-key em.v-inconclusive{background:var(--slate)}
/* ---- records: tables that stack on a phone ----------------------------------------- */
.rec{border:1px solid var(--rule);border-radius:3px;background:var(--hull);
padding:.75rem .85rem;margin:.5rem 0}
.rec-head{display:flex;flex-wrap:wrap;align-items:baseline;gap:.5rem;
font:.85rem/1.3 var(--mono);margin-bottom:.35rem}
.rec-head .id{overflow-wrap:anywhere;color:var(--ink)}
.rec form{margin-top:.6rem}
/* Why a check matters is a sentence, so it is set as one full width under the row rather
than squeezed into a column, where it would wrap to a ribbon two words wide. */
.why{font:.85rem/1.5 var(--prose);color:var(--dim);margin-top:.45rem;max-width:52rem}
.tag{font:.68rem/1 var(--mono);letter-spacing:.1em;text-transform:uppercase;
padding:.24rem .45rem;border-radius:2px;border:1px solid currentColor;white-space:nowrap}
.v-green{color:var(--green)} .v-yellow{color:var(--amber)}
.v-red{color:var(--red)} .v-inconclusive{color:var(--slate)} .v-unknown{color:var(--dim)}
/* ---- panels, controls, states ------------------------------------------------------ */
.narrow{max-width:27rem}
.panel{background:var(--hull);border:1px solid var(--rule);border-radius:3px;
padding:.95rem 1rem;margin:.9rem 0;min-width:0}
/* White plate behind the code: a QR needs the light modules to actually be light, and this
page is dark. */
.qr{display:inline-block;background:#fff;padding:8px;border-radius:4px;margin:.2rem 0;line-height:0}
.qr svg{display:block;width:min(240px,60vw);height:auto}
.empty{border:1px dashed var(--rule);border-radius:3px;padding:1.4rem 1rem;
color:var(--dim);font-size:.9rem}
code,pre,.mono{font-family:var(--mono);font-size:.82rem}
code{overflow-wrap:anywhere;color:var(--trace)}
pre{background:#040d11;border:1px solid var(--rule);border-radius:3px;padding:.8rem;
overflow:auto;max-height:32rem;max-width:100%;color:var(--ink)}
a{color:var(--trace)}
button{font:500 .82rem/1 var(--mono);letter-spacing:.05em;background:var(--trace);
color:#04181a;border:0;border-radius:3px;padding:.55rem .9rem;cursor:pointer}
.btn{display:inline-block;font:500 .82rem/1 var(--mono);letter-spacing:.05em;
background:var(--trace);color:#04181a;border-radius:3px;padding:.55rem .9rem;
text-decoration:none}
button.plain{background:transparent;color:var(--dim);border:1px solid var(--rule)}
button.danger{background:transparent;color:var(--red);border:1px solid var(--red)}
button:hover{filter:brightness(1.08)}
input{font:.9rem var(--mono);background:#040d11;color:var(--ink);border:1px solid var(--rule);
border-radius:3px;padding:.55rem .6rem;max-width:100%;width:100%}
label{display:block;font:.72rem/1 var(--mono);letter-spacing:.12em;text-transform:uppercase;
color:var(--dim);margin:.9rem 0 .3rem}
.err{border:1px solid var(--red);color:var(--ink);background:rgba(194,87,87,.09);
padding:.6rem .75rem;border-radius:3px;font-size:.9rem}
.muted{color:var(--dim)}
form.inline{display:inline}
:focus-visible{outline:2px solid var(--trace);outline-offset:2px}
@media (prefers-reduced-motion:reduce){*{transition:none!important;animation:none!important}}
/* ---- above 46rem the records line up in columns ------------------------------------ */
@media (min-width:46rem){
.top{padding:.85rem 1.6rem}
main{padding:1.6rem}
.recs{margin:.8rem 0}
/* Every row shares one grid, so the columns agree across rows without a header or a table. */
.rec{display:grid;grid-template-columns:minmax(12.5rem,18rem) minmax(0,1fr) auto;gap:.35rem 1.4rem;
align-items:baseline;background:none;border:0;border-bottom:1px solid var(--rule);
border-radius:0;padding:.6rem 0;margin:0}
.rec-head{margin:0;flex-direction:column;align-items:flex-start;gap:.3rem}
.rec form{margin:0}
/* Widths follow the content: a device name needs room, a finding count does not. */
.rec .readout{display:grid;grid-template-columns:1.7fr .9fr .9fr 1.1fr;gap:.15rem 1.2rem}
.rec .readout li{padding:0}
.rec .readout .lead{display:none}
.rec .readout .v{text-align:left}
.open{white-space:nowrap}
/* Spans the full row: the sentence is the useful part, not a fourth column. */
.why{grid-column:1/-1;margin-top:.1rem}
}
</style></head><body>
<header class="top">
<h1 class="mark">echo<span>lot</span></h1>
{{if ne .Page "login"}}
<nav>
<a href="/"{{if eq .Page "dashboard"}} aria-current="page"{{end}}>overview</a>
<a href="/devices"{{if eq .Page "devices"}} aria-current="page"{{end}}>devices</a>
<a href="/runs"{{if eq .Page "runs"}} aria-current="page"{{end}}>runs</a>
</nav>
<span class="who">{{.Session.Display}}{{if not .Session.Admin}} &middot; your account{{end}}
<form method="post" action="/logout" class="inline"><button class="plain">Sign out</button></form>
</span>
{{end}}
</header>
<main>
{{if eq .Page "login"}}
<h2>Sign in</h2>
<p class="lede">This server keeps the measurements your devices have uploaded.</p>
{{if .OIDC}}
<p><a class="btn" href="/auth/start">Sign in with your identity provider</a></p>
{{end}}
{{if .LocalSet}}
<form method="post" action="/login" class="panel narrow">
<h3>Break-glass account</h3>
<label for="u">Username</label>
<input id="u" name="username" autocomplete="username" value="{{.AdminUser}}">
<label for="p">Password</label>
<input id="p" name="password" type="password" autocomplete="current-password">
<p><button>Sign in</button></p>
</form>
{{else}}
<p class="err">No break-glass account is set. Run
<code>echolot-server --set-admin-password</code> on the host to create one.</p>
{{end}}
{{else if eq .Page "dashboard"}}
<h2>{{if .Admin}}This server{{else}}Your account{{end}}</h2>
<ul class="readout panel">
<li><span class="k">{{if .Admin}}devices enrolled{{else}}your devices{{end}}</span>
<span class="lead"></span><span class="v">{{.Devices}}</span></li>
{{if .Admin}}
<li><span class="k">linked to an account</span>
<span class="lead"></span><span class="v">{{.Linked}}</span></li>
{{end}}
<li><span class="k">{{if .Admin}}runs stored{{else}}your runs{{end}}</span>
<span class="lead"></span><span class="v">{{.Runs}}</span></li>
{{if .Admin}}
<li><span class="k">server version</span>
<span class="lead"></span><span class="v">{{.Version}}</span></li>
{{end}}
</ul>
{{if not .Admin}}
<div class="panel">
<p>You can see every device you have signed in on, read everything they have uploaded, and
delete any of it.</p>
<p class="muted">Enrolling devices, revoking them, and reading other people's uploads need an
administrator account.</p>
</div>
{{end}}
{{with .SelfTest}}
<h2>Self-test</h2>
<p class="lede">What this server can measure from where it stands, checked at startup. A
capability missing here is missing from every run this server takes part in &mdash; so a
client asking for that measurement gets nothing, rather than a wrong answer.</p>
<ul class="readout panel">
<li><span class="k">kernel settings</span><span class="lead"></span>
<span class="v {{if .SysctlOK}}v-green{{else}}v-yellow{{end}}">{{if .SysctlOK}}as needed{{else}}need attention{{end}}</span></li>
<li><span class="k">egress path MTU</span><span class="lead"></span>
<span class="v {{if .MTUOK}}v-green{{else}}v-yellow{{end}}">{{if .MTUOK}}full 1500{{else}}reduced{{end}}</span></li>
</ul>
{{if .Sysctls}}
<h3>Kernel settings</h3>
<div class="recs">
{{range .Sysctls}}
<div class="rec">
<div class="rec-head"><span class="id">{{.Name}}</span>
<span class="tag {{if eq .Severity "ok"}}v-green{{else}}v-yellow{{end}}">{{.Severity}}</span></div>
<ul class="readout">
<li><span class="k">found</span><span class="lead"></span><span class="v">{{.Got}}</span></li>
<li><span class="k">wanted</span><span class="lead"></span><span class="v">{{.Want}}</span></li>
</ul>
<div class="why">{{.Why}}</div>
</div>
{{end}}
</div>
{{end}}
{{if .EgressMTU}}
<h3>Egress path MTU</h3>
<div class="recs">
{{range .EgressMTU}}
<div class="rec">
<div class="rec-head"><span class="id">{{.Target}}</span>
<span class="tag {{if .FullMTU}}v-green{{else}}v-yellow{{end}}">{{if .FullMTU}}full{{else}}reduced{{end}}</span></div>
<ul class="readout">
<li><span class="k">discovered</span><span class="lead"></span>
<span class="v">{{if .DiscoveredMTU}}{{.DiscoveredMTU}} bytes{{else}}not measured{{end}}</span></li>
</ul>
{{with .Err}}<div class="why">{{.}}</div>{{end}}
</div>
{{end}}
</div>
{{end}}
{{end}}
{{with .ADBEndpoints}}
<h2>Dev relay</h2>
<p class="lede">Wireless-debug endpoints reported by Echolot instances on a test network. mDNS
does not cross subnets, so a device on that network relays adbd&rsquo;s rotating port here for
a developer sitting elsewhere. Nothing here is a measurement, and the entries expire &mdash;
a port older than a few minutes has probably already rotated.</p>
<div class="recs">
{{range .}}
<div class="rec">
<div class="rec-head">
<span class="id">{{if .DeviceName}}{{.DeviceName}}{{else}}{{.Device}}{{end}}</span>
<span class="tag v-unknown">{{ago .AgeS}}</span>
</div>
<ul class="readout">
<li><span class="k">adb connect</span><span class="lead"></span>
<span class="v">{{.Host}}:{{.Port}}</span></li>
<li><span class="k">device</span><span class="lead"></span><span class="v">{{.Device}}</span></li>
<li><span class="k">reported from</span><span class="lead"></span>
<span class="v">{{if .SourceIP}}{{.SourceIP}}{{else}}&mdash;{{end}}</span></li>
</ul>
{{with .Note}}<div class="why">{{.}}</div>{{end}}
</div>
{{end}}
</div>
{{end}}
{{else if eq .Page "devices"}}
<h2>{{if .Admin}}Devices{{else}}Your devices{{end}}</h2>
{{with .Link}}
<div class="panel">
<h3>Enrolment link</h3>
<p class="lede">Single use, valid 24 hours. Treat it like a password until it is spent.</p>
<!-- On the phone being enrolled this is the whole procedure: the scheme is registered by the
app, so following the link hands it the token directly. Copying a 200-character string
between two devices is the step that goes wrong, and it does not have to happen at all. -->
{{with $.LinkHref}}<p><a class="btn" href="{{.}}">Open in the Echolot app</a></p>{{end}}
<p class="muted">Works on the phone you are enrolling. From another device, scan this:</p>
{{with $.LinkQR}}<div class="qr">{{.}}</div>{{end}}
<p class="muted">Or copy the link into the app's enrolment field, or deliver it over adb.</p>
<p><code>{{.}}</code></p>
<p class="muted mono">adb shell am start -a android.intent.action.VIEW -d "{{.}}"</p>
</div>
{{end}}
{{if .Admin}}
<form method="post" action="/enroll-tokens">
<input type="hidden" name="csrf" value="{{.CSRF}}">
<button>Create enrolment link</button>
</form>
{{end}}
{{if .Rows}}<div class="recs">
{{range .Rows}}
<div class="rec">
<div class="rec-head">
<span class="id">{{if .Name}}{{.Name}}{{else}}{{.ID}}{{end}}</span>
{{if .LinkedToAccount}}<span class="tag v-green">{{.AccountName}}</span>
{{else}}<span class="tag v-unknown">no account</span>{{end}}
</div>
<ul class="readout">
<li><span class="k">device</span><span class="lead"></span><span class="v">{{.ID}}</span></li>
<li><span class="k">enrolled</span><span class="lead"></span>
<span class="v">{{.Enrolled.Format "2006-01-02 15:04"}}</span></li>
<li><span class="k">runs</span><span class="lead"></span><span class="v">{{.Runs}}</span></li>
</ul>
{{if $.Admin}}
<form method="post" action="/devices/{{.ID}}/revoke">
<input type="hidden" name="csrf" value="{{$.CSRF}}">
<button class="danger">Revoke</button>
</form>
{{end}}
</div>
{{end}}
</div>{{else}}
<p class="empty">{{if .Admin}}No devices yet. Create an enrolment link and open it on the phone
you want to measure from.{{else}}No devices yet. Sign in from the Echolot app on your phone to
link one to this account.{{end}}</p>
{{end}}
{{else if eq .Page "runs"}}
<h2>{{if .Admin}}Uploaded runs{{else}}Your uploaded runs{{end}}</h2>
<p class="lede">Each run is shown exactly as it arrived, at the privacy level its uploader chose.
Nothing here can un-redact one.</p>
{{if .Rows}}
<div class="trace">{{range .Rows}}<i class="{{verdictClass .Verdict}}"></i>{{end}}</div>
<p class="trace-key"><b>oldest &rarr; newest</b>
<b><em class="v-green"></em>clean</b>
<b><em class="v-yellow"></em>worth a look</b>
<b><em class="v-red"></em>faults</b>
<b><em class="v-inconclusive"></em>inconclusive</b></p>
{{end}}
{{if .Rows}}<div class="recs">
{{range .Rows}}
<div class="rec">
<div class="rec-head">
<span class="id">{{.UploadedAt.Format "2006-01-02 15:04"}}</span>
<span class="tag {{verdictClass .Verdict}}">{{verdictLabel .Verdict}}</span>
</div>
<ul class="readout">
<li><span class="k">device</span><span class="lead"></span><span class="v">{{.DeviceName}}</span></li>
<li><span class="k">findings</span><span class="lead"></span><span class="v">{{.FindingCount}}</span></li>
<li><span class="k">size</span><span class="lead"></span><span class="v">{{kb .SizeBytes}} kB</span></li>
<li><span class="k">privacy</span><span class="lead"></span><span class="v">{{.Anonymization}}</span></li>
</ul>
<div class="open"><a href="/runs/{{.DeviceID}}/{{.ID}}">Open run</a></div>
</div>
{{end}}
</div>{{else}}
<p class="empty">Nothing uploaded yet. Take a measurement in the app and upload it; it will
appear here.</p>
{{end}}
{{else if eq .Page "run"}}
<h2>Run {{.ID}}</h2>
<p class="lede">The document as stored, indented for reading. Nothing has been added or removed.</p>
<form method="post" action="/runs/{{.Device}}/{{.ID}}/delete">
<input type="hidden" name="csrf" value="{{.CSRF}}">
<button class="danger">Delete this run</button>
</form>
<pre>{{.JSON}}</pre>
{{end}}
</main></body></html>
`

Some files were not shown because too many files have changed in this diff Show More