Give up the radio for the link, not for the search

Two failures on the tablet, one of them mine from an hour ago.

The real one first. A trainer that drops mid-ride could not get back:
its supervisor reconnects on its own (FR-1.10) and those attempts never
pass through `DeviceRegistry::connect`, so nothing suspended the device
list for them. On Android a GATT link that is discovering services while
a scan is running is killed by the platform, and the log has it exactly —
a reconnect at 13:07:18 dying with `Disconnected while discovering
services`, inside a list-scan session opened at 13:07:02.

My first fix was to suspend the list whenever the trainer was
mid-connect. That was drawn around the wrong thing. `Connecting` covers
the 15 s *search*, `Reconnecting` covers the backoff between attempts,
and against an asleep trainer those alternate for the whole of
RECONNECT_ATTEMPTS — so the list scan went off the air for minutes and
every other device starved with it. A pod that dropped could never be
seen again, which is what "the pods disconnect after 40 seconds" was.

The window that matters is narrower than either: connect, then discover
services. `scan::gatt_setup` marks it — an RAII guard taken by the
trainer, pod and heart rate paths the moment their search returns a
peripheral — and the device list yields only for that. Measured on the
tablet: 1.2 s of yielding for a heart rate connect, then straight back to
one session per 21 s.

Also: the "+ pod seen; the − pod speaks for the pair" line is logged once
per run of refusals rather than once per sighting. The device list
republishes several times a second and every pass re-reported a visible
pod — 274 identical lines in five minutes, burying the connect failures
the log was being read for.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-27 13:24:39 +02:00
co-authored by Claude Opus 5
parent 9b1749d959
commit bbd11757ee
7 changed files with 93 additions and 5 deletions
+21 -3
View File
@@ -620,6 +620,9 @@ async fn run(
// back whenever the `−` pod is reachable, so it only ever expires against a
// `−` pod that is genuinely not coming (see `PLUS_GRACE`).
let mut plus_gate = tokio::time::Instant::now() + PLUS_GRACE;
// Whether the "+ pod is redundant" refusal has already been logged for the
// current state of affairs. See the `Seen` arm.
let mut plus_declined = false;
let (attempt_tx, mut attempt_rx) = mpsc::channel::<Attempt>(4);
let mut housekeeping = tokio::time::interval(Duration::from_secs(5));
@@ -686,9 +689,20 @@ async fn run(
tokio::time::Instant::now() >= plus_gate,
)
{
tracing::debug!(
"controller: + pod seen; the − pod speaks for the pair"
);
// Once per run of refusals, not once per sighting.
// The device list republishes several times a
// second and every pass re-reports a visible pod,
// so this logged twice a second for as long as the
// + pod was in the room — 274 lines in five minutes
// on the tablet, burying the connect failures we
// were reading the log for.
if !plus_declined {
tracing::debug!(
"controller: + pod seen; the − pod speaks for the pair \
(silenced until this changes)"
);
plus_declined = true;
}
continue;
}
let slot = slot_mut(&mut minus, &mut plus, pod);
@@ -793,6 +807,10 @@ async fn run(
// grace period lets the `+` pod in.
if !minus.idle() {
plus_gate = tokio::time::Instant::now() + PLUS_GRACE;
} else {
// The − pod is no longer speaking for the pair, so the next
// refusal — if there is one — is news again.
plus_declined = false;
}
// Converge on one link. The `+` pod may have connected first —