From f7701c2d2f471f2d059decf7c752bbe02eb2dcd7 Mon Sep 17 00:00:00 2001 From: mrambossek Date: Sat, 1 Aug 2026 14:14:20 +0200 Subject: [PATCH] engine: throughput is opt-in in the run config; document the work A 5-second run at 50 Mbps moves ~30 MB. On a metered connection that is the user's money, and a measurement tool that spends it unasked is not one people keep installed. Co-Authored-By: Claude Fable 5 --- docs/build-status.md | 45 +++++++++++++++++++ .../app/echo_lot/engine/ServerMeasurement.kt | 17 +++++++ 2 files changed, 62 insertions(+) diff --git a/docs/build-status.md b/docs/build-status.md index 0d3a25f..d3dcb31 100644 --- a/docs/build-status.md +++ b/docs/build-status.md @@ -774,3 +774,48 @@ Six live tests against fmr, all green, no device involved: `LiveServerTest`, `Li `LiveGranted`, `LiveDownstream`, `LiveUpload`, `LiveCompat`, `LiveEnrollment`. Plus 74 client unit tests and the full Go suite. Everything in the last several entries is verified from the PC; the app's UI (settings, history, deep-link enrollment) and `mtu.pmtud_up` remain device-only. + +### throughput: a rate, plus the qualifier that makes it a measurement (server-v0.6.1 … v0.6.2) +A throughput test reports the *smallest* limit on the path — and the sender's own ceiling is one of +the candidates. If the server is asked for 50 Mbps and 50 Mbps arrives, the network was never the +constraint and "50 Mbps" says nothing about it. So `perf.throughput_udp` always carries +`limited_by` (duration | budget | rate | send_error) and `measures_network`, and a finding is +raised only when the path is actually implicated. The live run against fmr reports 20 Mbit/s with +`measures_network: false`, which is the correct and useful answer. + +Loss is computed against the **sender's own count**, fetched from the observations API, not against +the requested rate. A receiver alone cannot tell "the network dropped it" from "the sender never +sent it", and guessing turns a healthy server-side limit into a phantom network fault. The server +keeps one summary per action rather than per-packet records — a ten-second run at 50 Mbps is half a +million packets, and a struct each would turn a measurement into memory exhaustion. + +Sending is **paced**, on an absolute schedule. Unpaced would measure the server's NIC and the first +queue it meets, then collapse into loss that reads as a network fault; sleep-per-packet would +accumulate scheduler error and drift the rate down over a ten-second run. + +Throughput gets its own grant budget sized from the request, so every *other* action stays bounded +at 8 MiB. When the byte cap binds before the clock does, the **duration is shortened and reported** +rather than the run being truncated: promising thirty seconds and delivering twenty-one is the same +information with a surprise attached, and it keeps "the clock ended the run" as the normal case — +the only case where the rate is a clean property of the path. That behaviour came out of a test +that failed honestly (30 s at 100 Mbps needs 375 MB against a 256 MB cap). + +It is **opt-in** in the run config, default off. A 5-second run at 50 Mbps moves ~30 MB; on a +metered mobile connection that is the user's money, and a tool that spends it without being asked +is not one people keep installed. + +#### The bug the live test found +The first live run delivered 104 packets and stopped after 50 ms. The grant's rate check exempted +the first 50 ms entirely, meaning to be lenient at startup — the effect was the opposite. A sender +could dump an unbounded burst into that free window, and the instant the check switched on it +compared those bytes against 50 ms worth of allowance and refused everything until real time caught +up. **Every short test passed** (downtrain sends 50 packets, big_send seven); every sustained send +died fifty milliseconds in. + +Replaced with a token bucket (`allowance = burst + rate × elapsed`), which is smooth from t=0. +The burst is 100 ms of the allowed rate, floored at one ordinary datagram — deliberately one, since +at 8 kbps a 64 KB floor is sixty-four seconds' worth, exactly the instant dump the ceiling exists to +prevent. The pre-existing rate test caught that when I first tried the generous floor, and it was +right to. Second half of the same bug: callers treated *any* refusal as terminal, so `TryAllow` now +says why — a sender paces through a transient "too fast just now" and still stops dead on a spent +budget or an expired grant. Both halves are pinned by regression tests. diff --git a/echolot-app/core-engine/src/main/kotlin/app/echo_lot/engine/ServerMeasurement.kt b/echolot-app/core-engine/src/main/kotlin/app/echo_lot/engine/ServerMeasurement.kt index 9ddffab..41f023e 100644 --- a/echolot-app/core-engine/src/main/kotlin/app/echo_lot/engine/ServerMeasurement.kt +++ b/echolot-app/core-engine/src/main/kotlin/app/echo_lot/engine/ServerMeasurement.kt @@ -47,6 +47,14 @@ class ServerMeasurement( * it is a flag rather than an assumption. */ val downstream: Boolean = true, + /** + * Throughput moves real data — a 5-second run at 50 Mbps is about 30 MB — so it is off + * unless asked for. On a metered mobile connection that is the user's money, and a + * measurement tool that spends it without being told to is not one people keep installed. + */ + val throughput: Boolean = false, + @Suppress("unused") val throughputSeconds: Int = 5, + @Suppress("unused") val throughputKbps: Int = 50_000, ) fun run(cfg: Config): MeasurementDocument { @@ -89,6 +97,15 @@ class ServerMeasurement( tests.addAll(dsTests) allFindings.addAll(dsFindings) } + + if (cfg.throughput && profile.supports("throughput")) { + val (tpTest, tpFindings) = ThroughputMeasurement(ids).run( + cfg.credential, session.sessionId, control, ps, sessionRef = "sess-1", + durationS = cfg.throughputSeconds, kbps = cfg.throughputKbps, + ) + tests.add(tpTest) + allFindings.addAll(tpFindings) + } } control.deleteSession(cfg.credential, session.sessionId)