Files
echolot/docs/findings-registry.md
T
mrambossekandClaude Opus 5 d65dbbc75a app: tell a broken device resolver apart from a broken network
Diagnosing a tablet that claimed "no internet" took twenty adb commands
to establish something the app should have said in one run: ping to
1.1.1.1 worked, the configured DNS server answered a raw UDP query in
65 bytes, and Android still could not resolve a hostname. The network was
fine; netd had wedged.

Those two failures look identical to a user and want opposite responses —
"look at your router" against "toggle your wifi" — so dns.resolver asks
the network's own servers directly and compares the answer against what
the platform returns for the same name. The query is hand-rolled over a
plain DatagramSocket on purpose: anything routed through a resolver API
would inherit the very fault being looked for.

dns.system_resolver_broken fires only on the pairing that is otherwise
unattributable: server answered, platform did not. Per network, because a
phone can have wedged wifi and working cellular at once.

Also records Android's own verdict per network — validated, captive
portal, partial connectivity — which the app reproduced with its own HTTP
probes but never stored. It is free, it is what the user sees in the
status bar, and its disagreement with our measurements is exactly what
identified the tablet. NET_CAPABILITY_PARTIAL_CONNECTIVITY is @SystemApi
so the constant is inlined with its rationale, in the manner of OsAbi.kt,
and read defensively enough to report unknown rather than false.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 09:44:49 +02:00

7.0 KiB

Echolot findings registry

Closes open item 1 of measurement-schema.md §9.

A finding code is the stable, machine-readable half of a result. The prose around it changes freely; the code is what a dashboard groups by, what a diff between two runs keys on, and what someone greps a year of archived runs for. That only works if a code means exactly one thing, forever.

This document is the contract. It is kept in step with echolot-app/core-measurement/.../FindingRegistry.kt by a test that fails when either side has a code the other does not — a registry that drifts from its documentation is worse than none, because it looks authoritative.

Rules

  1. The prefix determines the category, and the category determines which verdict light the finding rolls up into (§7.3). A nat.* code appearing under connectivity is not a naming quibble; it changes which light turns red. Two codes were renamed from nat.* to connectivity.* for exactly this reason.
  2. One code per concept. Two emitters independently produced connectivity.downstream_loss and connectivity.loss_downstream for the same claim before this registry existed. Anyone aggregating either would have silently seen half their data.
  3. Codes are declared, not typed. Emitters reference a FindingSpec, so a typo is a compile error and no two call sites can disagree about a finding's category or default severity.
  4. Severity in the registry is the default. An emitter may escalate for a specific run; it may not quietly reclassify the finding in general.
  5. Say what is ruled out, where that is the useful half. "Loss upstream" is worth far more when it also states that the return path is clean, because that halves where to look next.
  6. Renaming a code is a breaking change once runs are archived at scale. Before 1.0 it is cheap; after, it needs an alias and a deprecation window.

Registry

connectivity

code severity means rules out
connectivity.udp_unreachable high No UDP echo replies came back from the server at all.
connectivity.udp_unreachable_upstream high The server received none of the probes, so traffic is dropped on the way out. The return path: nothing arrived to be replied to.
connectivity.udp_loss medium A large fraction of round-trip probes were lost, direction unknown.
connectivity.loss_upstream medium Probes were lost on the way to the server. The return path: replies came back for everything that arrived.
connectivity.loss_downstream medium Packets were lost on the way back from the server. The outbound path: the server received what it was answering.
connectivity.downstream_blocked high Server-initiated packets never arrive, although round trips work. Basic reachability: the path forwards replies, just not unsolicited traffic.
connectivity.downstream_reorder low Downstream packets arrive in a different order than they were sent.
connectivity.captive_portal medium A captive portal is intercepting connectivity checks.
connectivity.no_internet high Android's own connectivity checks fail on this network.

mtu

code severity means rules out
mtu.reduced_downstream low The downstream path MTU is below the usual 1500 bytes.
mtu.downstream_blackhole medium Datagrams above the path MTU are dropped downstream, fragmented or not.
mtu.fragments_blocked medium IP fragments do not reach this device even when sent in order.
mtu.fragment_reorder_sensitive low Fragments are delivered in order but dropped when reordered or delayed. Fragmentation itself: in-order fragments arrive fine.

nat

code severity means rules out
nat.udp_rebinding medium A NAT remapped the UDP source port mid-flow.
nat.symmetric medium The NAT assigns a different external port per destination.

perf

code severity means rules out
perf.throughput_no_delivery high No throughput traffic arrived, although the server sent it.
perf.throughput_below_offered low Less throughput arrived than the server sent for the whole run.

dns

code severity means rules out
dns.answer_rewritten high A resolver returned an answer that differs from the authoritative record.
dns.authoritative_unreachable medium The canary zone's authoritative server could not be reached.

v6

The prefix is v6., matching the test-type registry (v6.brokenness, v6.happy_eyeballs, …). These were ipv6.* while declaring Category.IPV6; since the prefix map only knows v6, they rolled up under connectivity instead — the third occurrence of rule 1 being broken.

code severity means rules out
dns.system_resolver_broken high The network's DNS server answers, but this device cannot resolve names through it. A network fault: the server replied to a query sent from this device.
measurement.vpn_constrained info A VPN was active, so the networks underneath it could not be measured. Nothing — this run says little about the underlying network either way.
v6.no_default_route medium The device has a global IPv6 address but no IPv6 default route. Guesswork: this is read from the routing table, not inferred from silence.
v6.route_without_address medium The network advertises an IPv6 default route but the device has no global IPv6 address. A working IPv6 setup: SLAAC did not produce a usable address on this link.
v6.no_icmp_reply low IPv6 is configured but ICMPv6 echo gets no reply. Nothing on its own: IPv6 may work fine with ICMP filtered.
v6.not_offered info This network does not offer IPv6.

v6.no_icmp_reply was v6.broken until a phone reported it while loading an IPv6-only site over TCP perfectly well. The only evidence behind it is ICMPv6 echo, which is widely filtered on networks where IPv6 works — so the finding now states what was observed and names both explanations instead of choosing one. It is still worth reporting: filtered ICMPv6 breaks Path MTU Discovery. Corroborating it with a real IPv6 connection would let the two cases be separated, and is the proper fix.

v6.not_offered is info and must stay info. Most networks still do not offer IPv6 and that is not a fault; reporting it as a warning lights a yellow verdict on a healthy network, which teaches people to ignore the light — the one thing a diagnostic must never do.

Adding a finding

  1. Add a FindingSpec to FindingRegistry, and to its all list.
  2. Add the row here, under the section its prefix names.
  3. Emit it with finding(FindingRegistry.YOUR_CODE, …).

The registry test checks 1 and 2 agree, that every prefix maps to the category it claims, and that no two entries share a code.