a5c016833ddf4e045700bbd8d4bb7eba893e71fe
12
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
a5c016833d |
fix: install channel callbacks before any node runs
ThreadSanitizer reported ten data races on a plain multi-node network, all
the same one:
Read in Channel::try_push -> push_callback_() (worker thread)
Write in Channel::set_push_callback -> push_callback_ = ... (main thread)
A node's push and space callbacks are std::function members living on
channels it shares with its neighbours. register_callbacks() wrote them from
inside start(), and a network starts its nodes one at a time — so by the
time node N is being started, nodes 1..N-1 are already running and pushing
into N's input channel, reading the very std::function that start() is
assigning. Concurrent read and write of a std::function is a data race on
its vtable pointer and buffer, not a benign one.
This is the cause of the symptom
|
||
|
|
f53af260a2 |
fix: make the submit gate a single atomic
|
||
|
|
5628447ea8 |
fix: never self-move the parked output tuple
push_outputs ended with
else pending_ = std::move(result);
and the retry path calls it as push_outputs(std::move(*pending_), …), so on
that path `result` is the parked tuple itself. The assignment was a
self-move-assignment. std::tuple's is elementwise, and libstdc++'s
std::vector does not guard against self-move: _M_move_assign swaps its data
into a temporary, which is then destroyed. The vector ends up empty.
So the first park was clean — the argument there is a local temporary — and
the second erased the payload. The value was still delivered, still in
order, still counted, just empty. Downstream cannot distinguish that from a
frame on which the node genuinely found nothing, which is why it would never
surface as an error: in scene-actor-extraction it reads as "no faces in this
frame" and the run completes with a quietly wrong answer.
Scope, stated precisely because I first got it wrong: this needs a node with
*two or more* outputs. With one output the only thing that resubmits a
parked node is that output's own space callback, which by definition fires
when there is room, so the retry always succeeds and never reassigns. With
two, output A draining resubmits the node while output B is still full — the
retry skips A (already delivered, tracked in pending_done_) and fails on B,
and that is the reassignment that eats B's payload.
Every node in the scene-actor-extraction pipeline currently has exactly one
output, and the fanout is a separate class that does not use pending_, so
this is latent there rather than active. It is reachable by any multi-output
node under backpressure, which the library supports and documents.
Verified in both directions: on
|
||
|
|
6a4f45f111 |
fix: a filter or router must not drop on a full output
RouterNode and FilterNode were the last nodes on a data path still using the
throwing push() and swallowing the result:
try { out_ch_->push(val); } catch (const ChannelOverflowError&) {}
|
||
|
|
c73edffe5c |
chore: delete the unreachable duplicate of the parked-retry block
PoolNode::fire_once carried the pending_ retry block twice, verbatim. The first copy returns on every path through it — parked, drained, or not pending at all — so the second was dead code from the moment it appeared. PoolObjectNode, which is otherwise a line-for-line twin of PoolNode, has it once. No behaviour change; the deleted 25 lines were unreachable. Worth doing before the fixes queued behind it, each of which has to be applied once per copy of this function. The duplication is a symptom: PoolNode and PoolObjectNode are ~400 lines of near-identical code maintained by parallel edit, and a block getting pasted twice into one of them is exactly the failure that arrangement invites. Factoring the shared body out is a larger change and wants its own review. |
||
|
|
a8cfe7300a |
fix: a lossless fanout, a node that starts awake, and the instrumentation that found them
Three changes from one debugging session on the intermittent wedge, kept together because the instrumentation is what made the other two findable. **Fanout was never made lossless.** |
||
|
|
9c5ce5f34a |
fix: never drop a wake — a node must not sleep with one outstanding
|
||
|
|
28e06675f5 |
fix: park nodes on a full output instead of blocking the worker
push_blocking parked a scheduler worker inside the push. Nodes own a private single-thread pool, so the parked thread was the only one that could drain that node's own input — hold-and-wait, and under sustained backpressure four nodes of a five-node chain slept in nanosleep at once. channel.hpp already warned about this for sentinels; it applies just as much to data pushes. The scheduler was purely input-driven: on_input_ready() wakes a node when input arrives, with no counterpart for "my output has room". Lacking that signal, blocking the thread was the only way to handle a full output. This adds the missing half. - Channel::try_push + has_space + set_space_callback; the callback fires from both pop() and try_pop_now(). - PoolNode/PoolObjectNode keep a one-slot pending_ buffer with per-element done flags, so a retry cannot duplicate an already-accepted element. One slot suffices because queued_ admits at most one fire_once per node. - The re-check after clearing queued_ closes the lost-wakeup race where a space callback fires while the flag is still up and is swallowed. Two bugs surfaced once nodes actually parked, both fixed here: - pop_one reports an *empty* channel as ChannelClosedError, which is also the node's "upstream finished, self-stop" signal. A node woken by output space with empty inputs therefore killed itself. fire_once now releases the worker when its inputs are not ready rather than falling through. - The drained-park path resubmitted unconditionally instead of via on_input_ready(), firing nodes with nothing to read. compute_priority is now output-aware: mean output fill is deducted from mean input fill, mapped as 0.5·(1 + in - out). Input fill alone asks only "how much work is waiting for me"; a node whose outputs are already full cannot deliver, so running it just parks it again and wastes the slot while the node that would drain that channel waits behind it. The scheduler now favours whoever is furthest downstream of a bottleneck. Also adds a network-level error listener. A node's exception was discarded at the node boundary and survived only as a Closed event, which reports that a node stopped but not why — that missing detail is what made the above slow to diagnose. INode::set_network_error_callback plus StaticNetwork::set_error_handler forward it to the application. Tests: 121/121. test_backpressure_deadlock drives a five-node chain with capacity-2 channels against a slow sink and fails on the old code. The four test_pool_node overflow tests now assert parking rather than the removed drop-and-report behaviour. Known-incomplete: a rare hang remains, roughly 1 run in 20 against a 300s timeout, down from every run failing. Committed because the fix is a large strict improvement and the residual case needs its own reproduction. |
||
|
|
6595e6e925 |
fix: node outputs block instead of dropping on a full channel
Every node output used the throwing push(), so a consumer falling behind cost values rather than time. push_blocking() already existed on Channel and OutputPort — "wait for the consumer to drain instead of dropping; the producer just runs slower" — but nothing called it. A dropped frame does not degrade a downstream result, it silently changes one, and the consumer has no way to tell it happened. For any pipeline whose output is a claim about its input, that is corruption rather than degradation. Safe because sentinels are already handled out-of-band, above this path: only data blocks, so the EOF token that unwinds the network can always overtake a stalled data path. That is exactly the hold-and-wait deadlock the push_sentinel comment warns about, and the reason it is not reachable here. Measured on a downstream consumer (face pipeline, 77s clip at 5 fps, expected 385 sampled frames): before 65 frames written, 320 dropped at one node, 29s after 385 frames written, 0 dropped, 17s Faster, not slower — a dropped frame has already cost its decode, and the overflow exception cost more. Two consecutive runs now produce byte-identical output, which they did not before: what got dropped depended on timing, so the same command could yield different results. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
19f5a2b0ae |
Deliver EOF sentinels out-of-band to prevent teardown deadlock
Channel::push() drops values on overflow (the intended backpressure policy for data), and PoolNode swallows the resulting ChannelOverflowError. For a control sentinel like EOF this is fatal: a single dropped EOF under backpressure wedges every downstream pop() forever, so the pipeline never tears down. Deliver sentinels out-of-band instead. Channel::push_sentinel() stores the token in a dedicated slot that does not consume ring capacity, so it can never overflow and — crucially — never blocks the caller. That non-blocking property is essential: each KPN node has a single worker thread, so a *blocking* push would park that thread and stop it draining its own input, cascading into a hold-and-wait deadlock under backpressure. The consumer's pop()/try_pop_now() drain the ring first, then deliver the sentinel, so it always arrives after every value pushed before it. approx_size() (which node readiness checks call) counts a pending sentinel as consumable work, so a channel carrying only a sentinel still schedules its consumer's next fire — without this the token would sit undelivered and the pipeline would still deadlock at teardown. PoolNode/PoolObjectNode route values carrying an eof flag (direct .eof or nested .source.eof) through push_sentinel via a SFINAE-safe is_sentinel_value trait; all other values keep the existing lossy throwing push. The trait compiles to false for types without an eof convention, so this is a no-op for pipelines that don't use one. Verified end-to-end: scene_analyze now reaches EOF, flushes its output, and exits cleanly instead of hanging. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
6f384dc4b5 |
Added callbacks for node errors and fifo overflow
Add new doc system which should/might deploy to pages. |
||
|
|
f6bcaa15b0 |
Performance improvements, better readme and complete python bindings
🧪 Test / test (push) Failing after 28m30s
|