`docs/display-and-extension.md` §2 fixes a decision rule in advance: if the
99th percentile of a frame sits inside 16 ms, tiled computation is rewritten
as a scheduling concern for export rather than built on the interactive path.
Nothing in the tree could answer that, so the rule had nothing to act on.
This is the instrument. It renders a 60 MP synthetic source through the real
`render_detailed` at three viewport sizes and four chain lengths, fit and
zoomed to 1:1, and reports nearest-rank percentiles rather than means — a
slider drag is judged by its worst frame.
Three things it does that a simpler timer would not:
- It separates the fused pass from the neighbourhood stage. "Every operation
active" mixes one dispatch together with a chain of convolutions, and §2's
question is about the first of those. `point` is every operation that
contributes a fragment to the fused shader; `all` adds the four with
kernels, and M3 times those alone by moving only a detail parameter so
`render_detailed`'s colour reuse skips the fused dispatch. The reuse is
reported rather than assumed — the `colour` column counts fused dispatches
and must be zero for an M3 row to mean what it says.
- It times the CPU half separately. Composition runs per frame in
`DevelopSession::render`, so it is inside the budget whether or not anyone
has looked at it, and if shader assembly were the expensive half then no
tile scheduler could help.
- It builds the "every operation" chain from `EditGraph::capabilities` rather
than from a list, so declaring a new node does not quietly turn that row
into a shorter chain wearing a longer chain's label.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>