The manual described the review (7c9a4ee) without a picture, because the
demo library holds no duplicates. The new duplicates scene makes two:
it copies two New York frames into a bck folder beside their own,
restarts the app so the scan finds them, waits for the sidebar row the
sweep's dating brings (54aee50), opens the review from it, presses
Check and takes the page. It deletes the copies and restarts on the
library as it was; record.sh's snapshot restore would remove them too.
It is registered last, so no other scene sees the copies.
The picture shows both groups proved the same file, the camera-named
copy outside bck marked Stays, and "Move 2 copies to trash" ready.
Every scene was recorded again on a release build of this commit with
the automation feature. Master changed what nearly every picture shows
after they were taken: scrollbars on the develop column, the grid, the
sidebar and Settings; a "?" beside Settings in develop's top bar, with
its controls regrouped; the Film row opening its list as a popup.
The film scene pressed the list's rows through the develop column, which
no longer holds them: the list is a popup, and the automation hook
reports its contents relative to it. The scene now opens it with
film_list_open, turns the wheel down it and back so the popup and its
scrollbar are seen scrolling, and clicks Velvia at its popup position
plus the popup's origin. The caption says so. film_reach passed in the
same run: the last stock was reached by the wheel, a drag, the scrollbar
and the keys.
Looked at as contact sheets of each GIF's middle and last frames and
each PNG. launch, launch-folder, library-nesting.png and
library-collection-menu came out byte-identical after optipng and are
unchanged. develop-zoom's deepest frames are smooth on Xvfb as before;
its caption does not claim blocks. Settings shows version 0.15.0,
the build's own, until the release commit bumps it.
The film list bug gave no failure anywhere: the data was right, the
markup compiled, and the list rendered. Two guards now check that the
list can be walked to its end, both by driving input rather than by
reading markup.
tests/film_list_reaches_every_stock.rs runs in CI and needs no display.
It builds the real AppWindow on Slint's testing backend and gives it 28
stocks. It dispatches window events through the same routing a window
uses: popup, Flickables, arbitration. It then checks that the last stock
is on screen, that is, not clipped away:
- after Down past the end, and that Enter chooses it;
- after drags on the list;
- after a run of wheel events with a still pointer;
- after dragging the scrollbar thumb.
Element queries need the Slint compiler's debug tables, which build.rs
emitted only for the `automation` feature. It now emits them for every
debug build too. Release builds, the ones that ship, are unchanged. The
testing backend is a dev-dependency at the same pinned version the
automation feature already uses, so no new crate enters the lockfile.
The test is compiled out of release test runs.
film_reach in tools/manual/scenes.py is the same check on the recording
rig: a real X pointer from xdotool, the release build, and the demo
library. It makes no picture, so it adds nothing to the manual. It runs
with every recording, or alone with `record.sh LIBRARY film_reach`, and
fails the run if the last stock (Ilford HP5 Plus) is out of reach by
the wheel, a drag, the scrollbar or the keys.
Both have to add the popup's position back. The testing backend reports
anything inside a popup relative to the popup, and so does the
automation hook built on it. They take the popup's position from the
Film row and Slint's clamp into the window.
Every scene was recorded again on a build of this branch rebased onto the
keyboard work and TD-1, since nearly every scene depends on files those
changed: the develop top bar now carries a star strip and Pick/Reject, the
roll shows flags and stars, and the grid's selection bar gains Label and
Flag. `--changed` could not be trusted to find them, because the rebase
made each picture's commit newer than the sources it was recorded from.
The develop-zoom caption said the wheel goes on "until the pixels are
blocks". On Xvfb the deepest frames come out smooth even though the app
draws past 1:1 nearest-neighbour on a real display (confirmed by eye on
the desktop), and a GIF shrunk to 960 wide could not show 3-pixel blocks
anyway. The caption now says what the clip shows; the prose above it,
which describes what the app does, stays.
The traceability job now runs `tools/manual/record.sh --check`: every
picture docs/manual/README.md shows must be made by a scene in
tools/manual/scenes.py, and every picture a scene makes must be shown.
It reads the two files and nothing else, so it needs no app, display or
LFS pull.
--changed now dates a scene by the newest commit among its pictures
rather than each picture alone. A scene that also makes a picture which
re-records byte for byte (panorama-aligned beside panorama.gif) no
longer stays listed for ever. A scene all of whose pictures come out
identical (launch) stays listed until one differs, which costs one
harmless re-run.
scenes.py aimed every press at window pixels, and the develop column had
already moved under it: Compose now sits above Adjust, so the old
exposure coordinate lands on a straighten slider. Every scene now names
what it presses by its accessible label through the automation hook,
places points on the photograph relative to the canvas, and opens its
photographs by file name. Each starts from a known place and undoes what
it did, so one can be recorded alone; the few that continue another's
state name it, and running one runs that first into a scratch folder.
Each scene also declares the pictures it makes and the sources they
depend on. `record.sh --check` fails when the manual shows a picture no
scene makes, or a scene makes one it does not show; it reads two files.
`record.sh --changed` re-records the scenes whose sources, or own code,
changed since the commit that last touched their pictures. record.sh
builds with the automation feature, restores the library from
DR_LIBRARY_SNAPSHOT, starts from a fresh profile and pins inference to
the CPU; the launch screen is recorded from an empty profile of its own.
Re-recorded with the ported scenes, and looked at frame by frame. What
differs from the pictures they replace:
- develop, presets, settings, local, compose, film, wb, light: the
current develop column (Compose with Vertical and Horizontal above
Adjust, the Label button), otherwise the same moments.
- library pictures: the filter bar's colour-label chips; no collection
left over from an earlier run in the sidebar; library-selection is the
twelve alpine frames rather than eight of them and four New York ones.
- library-rating rates two frames nobody had rated, so the stars are set
and not cleared.
- develop-zoom goes on past 1:1 with the wheel and ends on the file's
pixels as hard-edged blocks.
- repair covers a real mark on the road, with a size that fits it; film
is shown on the Chinatown frame instead of the road.
- panorama tries Perspective, Spherical and Cylindrical before filling.
- launch, launch-folder and panorama-aligned came out byte-identical.
Every scene in tools/manual aimed at window pixels written in by hand, so
a panel that gained a row moved every slider under it and the recording
went on dragging where the slider used to be. The develop column has
already moved that way (Compose now sits above Adjust), and nothing said.
A build with the `automation` feature listens on the Unix socket named
by DR_AUTOMATION and answers where an element is: by its accessible
label, the name a screen reader reads, or by its markup id for the few
things that are not controls (the canvas, the crop rectangle). It uses
Slint's element queries, which need the compiler's debug tables, so the
feature also turns those on in build.rs. It only answers questions; the
input is still xdotool's real pointer. No default build has the feature,
and one that has it listens only when the variable is set.
drive.py gains click-on, drag-on, hold-on, wait-for, wait-gone, labels
and ids. The grid's cells are now named by their file, each rating star
by its value, the sidebar's + as "New collection", and the Adjust
heading's reset as "Reset all adjustments" - controls a screen reader
could not reach before either.
The gesture book is generated from GESTURE tags, so it could not describe a
gesture nobody tagged, but nothing made anyone tag one. The arrow keys, Enter,
P, X, U, Delete, F1 and F2 all worked in the grid with no line in the help
sheet, and a tag could name a key whose handler had gone.
Key handlers now compare one canonical string, Keys.chord(event) == "Ctrl+Z",
instead of reading event.text and the modifiers themselves. keys.slint folds
the key and its modifiers into that spelling, so the literal in the handler is
the whole binding and the checker reads exactly what the handler dispatches
on. Each handler carries a KEYMAP comment naming the gesture-book section its
keys belong to, and a tag's keys field names its keys between backticks.
gestures-check now fails when a handler binds a key no tag in that section
names, when a tag names a key no handler there binds, when any .slint file
other than keys.slint reads event.text, when a compared literal is not
canonical, and when keys.slint's named keys drift from the Rust list.
Spellings are normalised in one place, chord.rs: Ctrl+z, Control+Z and
LeftArrow all mean what the handler's "Ctrl+Z" and "Left" mean. Shift and Alt
count only for letters and named keys, because on the French layout every
digit needs shift and a 6 has to be a 6 however it was typed.
A Rust keymap that both dispatched and was read by the generator was the
alternative. It would have moved the handlers' decisions away from the Slint
state they depend on, and a window that forgot to install it would have had
no working keys at all.
The keys that were already bound and undocumented are now tagged.
The help sheet says which move does a thing, and the manual has a
picture of the thing being done, but nothing joined the two: a user
reading "Pinch it with two fingers" had no way from there to the GIF of
it.
A GESTURE tag takes an optional `manual:` field naming a heading of
docs/manual/README.md by its anchor. The scan checks every one against
the anchors the bundled page is rendered with and fails when the manual
has no such heading, so renaming a section cannot leave the sheet
linking to the top of the page; gestures-check carries the same failure
into CI. The anchor goes into gesture_book.rs as a new field, and into
docs/gestures.md as a "See it" link to manual/README.md#anchor. The help
sheet draws a "See it" button beside the title of each gesture that has
one, which opens the bundled manual at that section.
The field is additive: a tag without it is unchanged, and no gesture
carries one yet.
The manual existed only as docs/manual/README.md, which the forge renders
and nothing else does. An installed copy of the application, on a laptop
with no network or on a tablet, had no manual it could open.
`traces manual` renders the README to docs/manual/index.html with
pulldown-cmark (already in the tree as Slint's Markdown parser, so this
adds a dependency edge and no crate). The page is one file with an inline
stylesheet that follows the system's light or dark preference, a
contents list of every section and subsection, and the pictures by their
relative media/ paths. Each heading carries the id the forge gives it, so
README.md#rating-and-flagging and index.html#rating-and-flagging are the
same link. A picture alone in its paragraph becomes a figure whose alt
text is shown as the caption, and every picture reserves its 16:11 box
before it loads, so a jump into the middle of the page lands where it
aimed rather than a screenful above. Links to design documents, which the
installed page has no copy of, point at the forge.
The page is committed rather than rendered at build time, as the gesture
book is: it is user-facing text reviewed in the diff, and the three
packagers then only copy it. `traces manual-check` fails in CI when the
committed page is not the render of the README, and the pre-commit hook
regenerates it when the README is staged.
Nothing made a release. CI built the APK and the installer on the master
push and kept them as workflow artefacts, the Linux binary was not kept
at all, and most tags went out with no downloads until they were
attached by hand.
build-and-test now also runs on v* tags. On a tag the desktop job keeps
its release binary, and a release job that needs desktop, Android and
Windows collects the three, names them with the version and runs
tools/publish-release.sh. The script titles and describes the release
from the annotated tag's message as the server holds it, writes
SHA256SUMS, and attaches what is not already there, so a re-run after
an interrupted upload finishes the job instead of duplicating it. The
same script is how a release is made or finished by hand.
Tried on v0.14.1, whose release was made by hand with the same files:
it found the release, reported all four files attached, and changed
nothing.
docs/ had 26 developer documents flat beside the manual, and the two
audiences are very differently sized: most readers want the manual and
the gesture reference, a few want the register, the designs and the
measurements. The manual and gestures.md stay at the top; everything for
someone changing the code moves to docs/dev/, and the two documents that
name their own successors — the v0.1 milestone and the UI-refinement plan
— go to docs/dev/archive/ rather than being deleted, since both are still
cited. docs/README.md is the index, users first.
Every reference follows: code comments, Cargo manifests, the workflows,
the pre-commit hook, the bench and traceability tools (which locate the
repo root by docs/dev/requirements.md now), packaging, the Docker READMEs,
CLAUDE.md, CONTRIBUTING.md and the README. The matrix links one level
deeper and is regenerated. Links out of the moved documents into the tree
gain a level; a link checker over every Markdown file finds none broken.
Measured on a Radeon RX 7900 XT against Arch's onnxruntime-rocm 1.29
(docs/inference.md §1.3): MIGraphX fp16 runs the detectors at 2.4–3.4 ms
against 10–58 ms on the CPU provider, the inpainter at 8 ms against 514,
with a 15–135 s compile per graph the first time and under a second from
its cache after. A compiling rung on TensorRT's terms, wired the same way.
The ROCm execution provider is gone (removed in ONNX Runtime 1.23), so the
AMD ladder is MIGraphX then the CPU, with no non-compiling rung between.
MIGraphX is registered through the runtime's generic key/value entry
point rather than ort's builder: 1.29 reads the legacy options struct for
its precision flags only, and the compiled-program cache directory
(`migraphx_model_cache_dir`) only travels the generic way. The provider's
cache key omits the precision, so f32 and fp16 programs get their own
directories. The probe fingerprint now includes the provider libraries
beside the runtime and the ROCm version, since a distribution's CPU and
ROCm builds are the same file at the same path.
`status().failed` reports only the rungs above the selection, so an AMD
desktop's About line says why MIGraphX won rather than that the NVIDIA
providers are not in the build.
Two examples: `ep_probe` times each provider cold and from cache, and
`ladder` drives `init` as the app does to watch the first-run sequence.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A scene that makes a parent, nests two collections in it by drag and by
the menu, files frames into a child and opens the parent to see it count
both; stills of the tree and of the menu. The collections recording is
re-made now that the bitmap under the cursor is the photograph.
drive.py grows a multi-leg drag: a diagonal with much vertical in it is
taken by the grid's Flickable as a scroll before the DragArea can claim
it, so a drag to the sidebar goes sideways first.
The recording sampled a red brick wall and moved the sliders by three
units, which at GIF size is a click that does nothing. The scene now
drags the frame cold first and picks a white air conditioner, so the
correction is visible and the picker's being absolute - set from the
photograph, not from where the sliders were - is what the picture
shows. The text says so, and says a blown highlight is refused.
The scene clicked 40px to the right of "pick", on "reset", so the
recording showed a neutral group being reset and a click on the wall
that panned. Re-recorded with the picker fixed: the word lights, the
sample moves temperature and tint, and Before shows what it corrected.
docs/manual/README.md is a tour for a photographer opening DarkRoom for
the first time — one picture per thing, moving where movement is the
point. tools/manual/ is how the pictures are made: drive.py puppeteers the
desktop build on a private Xvfb (launch, click, drag, type, screenshot,
record), scenes.py is each picture as a script, and record.sh runs them
all over a folder and writes the results into docs/manual/media/.
The media is in LFS, with the CI pulls excluding it as they exclude the
fixtures; a screenshot changes wholesale when the interface does.
Nothing in the pictures shows a person, by design: the demo library is
seventy urban and alpine frames, chosen from the catalog's rows that face
detection found nobody in.
The traceability matrix is regenerated here after the rebase that
brought this branch up to master.
The reference desktop's only system ONNX Runtime is Arch's
onnxruntime-opt-cuda: 1.29, built without TensorRT and against cuDNN 8
on a cuDNN 9 machine. The probe rejects both providers correctly and
the app runs on the CPU provider, which is right and not what anyone
wants. runtime/ beside the models is now searched ahead of /usr/lib,
tools/fetch-desktop-runtime.sh fills it with the four libraries from
the current onnxruntime-gpu wheel (cuDNN 9, TensorRT 10), and the
About caption lists every rung that lost and why, not only the first.
Verified: the app selects TensorRT from that directory with no
environment variable set.
Sargsyan et al., ICCV 2023; MIT code and weights (models/LICENCE.md),
exported by tools/export-migan.sh at a fixed 1×4×512×512 from the
authors' checkpoint — six operator types, 28 MB, in LFS like the rest.
The package installs it beside the scene model and the APK unpacks it
with the others.
The first int8 files found no faces at all, and for two reasons the
tool now guards against. The calibration set was landscape photographs
with no faces in them, so the score head's ranges had never seen the
face regime; the set is now proxies from the library itself. And ONNX
Runtime's strided and moving-average calibration modes both degrade
these graphs measurably (a quarter of the faces at eight images, none
at ninety-six), while driving the calibrator in chunks by hand gives
ranges identical to a single pass — so the tool does that, four images
at a time, and feeds quantize_static through its range cache.
Measured against f32 over 400 proxies (docs/inference.md §10.1): the
10g form finds every face above 32 px the f32 form finds; 500m and
2.5g find 96%, and what they lose sits at a median confidence of 0.52
against the 0.50 threshold. Shipped with the number on record.
The Android unpack list gains the three int8 files; without that the
tablet never saw them. D13's runtime half records the reopening.
tools/quantise-models.sh writes the QDQ form QNN's HTP backend takes
whole: opset 17, per-channel int8 weights, uint8 activations, ranges
from running the f32 graph over photographs fed exactly as the app
feeds them. The calibration is strided, four images at a time, because
every ONNX Runtime calibrator holds each image's whole set of
activations until it folds them — a gigabyte an image on the 10g
detector, and an OOM kill with no message when folded once at the end.
Release-time, never on the device (docs/inference.md §5): it needs
real photographs and a person reading the recall measurement that
gates whether each file is offered.
The desktop names where a package may have put libonnxruntime — an
override variable, beside the executable, the package's own library
directory, the Flatpak prefix, the system library directory — and
Android points at the APK's native library directory, which is also
what Qualcomm's DSP loader must be told for the Hexagon skel. Android
starts the engine at the end of the model unpack rather than at launch,
because the probe fingerprints the model files and a first launch has
none until then.
The About panel gains an Inference row beside Graphics, re-read every
two seconds while the probe runs and engines land, and faces.model_id
carries the detector's form: an int8 detector finds a different set of
faces and is a different population (docs/inference.md §7). A
low-memory signal drops every idle session with the GPU caches.
The APK assembly bundles ONNX Runtime and the Qualcomm HTP libraries
from Maven, fetched by tools/fetch-android-runtime.sh with their
published checksums; RUNTIME_DIR=none builds the tract-only APK, which
is a slower app and not a broken one. The desktop packages carry no
runtime yet.
Two probe fixes from the first desktop run: the floor must not be
built with CPU fallback disabled, and a versioned libonnxruntime.so is
a runtime too. On the reference desktop the probe now loads ONNX
Runtime 1.30, measures 30 ms on the CPU provider, and selects TensorRT
at 1.5 ms.
Twelve real frames from the fixture set now align in 4.5 s — 4.4 s of
matching, 118 ms of bundle adjustment — where the first run took 51 s and
left the first two frames out.
The matcher computes each pair's similarity matrix once, across the
cores, with a dot product written to vectorise; both nearest-neighbour
directions read it. The frames that failed were portrait: fitted into the
landscape input they used 512 of 1024 px, and their thin overlap did not
survive at half resolution. The same weights are now exported at 768×1024
as well and the detector picks the shape by aspect. The example aligns
from embedded previews and draws the set on a cylinder; on the fixture the
sweep is 152° at a fitted 47.9 mm against the EXIF's 50, RMS 1.5 px, and
the overlaps show no ghosting.
tools/onnx-probe-on-device.sh cross-builds dr-segment's onnx_probe
without the embedded segmentation model, pushes it with a model to the
attached device and times two runs. The 768×1024 XFeat export takes
~400 ms on the reference tablet's NEON cores against ~300 ms on the
desktop, with identical output ranges — inside NFR-MRG-1's 1 s per frame.
The blend half of S15.4 waits for a chunked blend to exist.
tools/export-xfeat.sh exports the convolutional network alone at 768×1024
grayscale, on the pattern of export-seg-model.sh: thirteen standard
operator types, no dynamic axes, the keypoint decoding left to Rust.
examples/onnx_probe loads it through the ort-over-tract backend the app
ships with nothing unsupported and runs it in ~300 ms on the desktop CPU.
The weights are Apache-2.0, read from the repository's LICENSE, with no
grant on the checkpoint — recorded in models/LICENCE.md before they land,
as FR-MRG-8 asks. The probe stays: the next model will need the same
check.
2d106det for the eye contours, OCEC for open or closed, SGC for
sunglasses — all three pinned to a batch of one by the same script as
the pair, and installed by every packager so the eyes-open filter works
out of the box. The two classifiers are MIT, code and weights; the
README records their provenance, SGC's undocumented training set, and
the hashes as fetched and as shipped.
The register said two things about plugins. §7 had listed "Plugin API"
as deferred since the first draft, in a bare row; §3.10 then specified
it in 23 clauses that counted against coverage. Twenty-one of them had
no implementation of any kind, and could not have: no crate loads
anything at runtime. The coverage figure was measuring the contradiction.
Decided 2026-09-19: §7 is right. §3.10 stays as the design of record,
each of its clauses is marked "(post-v1)" on its defining line, and
NFR-SEC-6 — which exists only for plugins — goes with them, as does D16.
The traceability tool learns the marker. A deferred requirement is still
defined, so a tag naming it is not an orphan, but it leaves the
denominator and is listed in its own table rather than under "not yet
tagged". The marker must sit on the definition line; a mention of
"post-v1" in prose changes nothing, and where an ID is defined twice the
deferral on either line wins. Both are tested. Coverage moves from 72.2%
of 194 to 80.6% of 170 without a line of application code changing,
which is the honest figure: it now measures what v1 owes.
faces.md §12.3 measured what the cheapest detector costs: the small
faces in every group shot, and a dog embedded a dozen times. Which
trade is right depends on the machine doing the sweep — a desktop left
overnight and a tablet on a battery want different answers — so the
detector is now a per-device setting, Fast / Balanced / Thorough on
the settings page beside the indexing button, persisted with the rest
of the settings file.
A detector is half of a model id. Every face, marker, shard and
calibration is keyed on faces.model_id precisely so that a model change
is a new id and a re-index rather than a silent change under existing
data, and a detector change is a model change: it decides which faces
exist and where the landmarks that align them land. So each choice
names its own pipeline. 500M keeps the bare "w600k_mbf" every existing
library was written under, so an upgrade disturbs nothing; the others
are qualified. Choosing one restarts coverage from zero under the new
id, the sweep re-detects, confirmed names carry across by box overlap,
and the sync shards are keyed by the same id so a peer on another
setting neither adopts nor pollutes them. The library controller
carries the id into the sync the same way it carries the cache budget,
because the sync starts from places that have no settings in reach.
All three shape-fixed exports ship — APK, Arch, Flatpak — since a
tablet has no other way to obtain the one it was not installed with;
the APK grows by twenty megabytes for the choice.
Every release commit in the history reads `Release X.Y.Z` and none of them
carries the message this script would have written, so nobody has ever
passed it `--commit` — and the reason is in the message it wrote: a
Co-Authored-By trailer naming an assistant, which no commit in this repository
carries and none should.
The trailer goes, and so does the paragraph above it: the script's own header
already says why it exists, and a release commit is the one place a one-line
subject is the whole convention.
There was no benchmark harness of any kind -- no benches/, no criterion,
no synthetic fixture -- while §8 promised a suite run per commit that
fails the build on regression. Ten performance requirements could be
neither passed nor failed.
tools/bench builds a deterministic 50,000-row catalog over a pool of
twelve generated JPEGs, about 14 MB, reproducible from a seed, with a
stamp so it rebuilds rather than silently comparing against a different
workload. It depends on nothing GPU or UI, which is what makes the CI job
affordable.
NFR-P1 and NFR-P3 are gated and tagged. NFR-P7, NFR-P8 and R2 are
measured but deliberately untagged: the export gate is one-sided, the
memory figure is the catalog layer's share rather than the whole, and
R2's first sentence is a 60 fps scroll a catalog benchmark cannot claim.
Every recorded value in the baseline is null. Nobody has run this on the
reference desktop, and a fabricated figure would make every later
comparison a comparison against a guess.
First run on this machine: catalog opens in 70 ms against a 2 s budget,
and thumbnail throughput measures 37 img/s against a target of 100 --
reported rather than asserted here, and the first evidence that the
target may not hold.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
clippy::print_literal on the results table header, and rustfmt's first
look at code whose author could not run cargo. The four constructs the
author flagged as risky -- scoped-thread lanes, a seventeen-argument
params!, is_some_and over a closure, a refutable let-else -- all compiled
untouched.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
docs/requirements.md §8 has said since it was written that performance is
verified by "an automated benchmark suite against a synthetic 50k catalog, run
per-commit … A regression beyond stated tolerance fails the build." There was
none. No benches/, no [[bench]], no criterion, no synthetic catalog, and three
CI workflows that between them measured nothing. Ten performance requirements
could therefore be neither passed nor failed, and five of them carried a
TRACES: tag regardless.
tools/bench is the half of that promise that can be kept honestly on a runner
with no GPU and no display.
# The fixture
Rows are cheap and pixels are not, so it builds fifty thousand catalog rows
over a pool of a dozen real files, each referenced by several thousand of them.
Everything the catalog half touches is rows and is exact at full scale;
everything the pixel half touches is one file at a time and does not care how
many rows point at it. Fourteen megabytes on disk instead of two terabytes, and
neither half is flattered by the trade. It is reproducible from a seed, and a
stamp beside it — seed, row count, source size, dr-catalog's schema version —
rebuilds it rather than letting a run be compared against a baseline that
describes a different library.
# What it can now pass or fail
NFR-P1, and R2's second sentence with it: Catalog::open plus the count, first
window and timeline the grid cannot paint without. The interesting part turned
out to be the open itself — schema::backfill runs three passes over the images
table on every open, which is O(library) work on a path whose budget is stated
in absolute seconds. Tagged TRACES: NFR-P1, on a gate that fails if it breaks.
NFR-P3: thumbnail throughput on the embedded preview path, through the same
per-image work spawn_thumbnail_sweep does and in the same shape — chunks of 96,
lanes owning disjoint slices, the single thread that owns the store writing the
finished chunk. Mirrored rather than called, because that function takes a
RemoteBackend and would measure somebody's network. Tagged TRACES: NFR-P3.
# What it deliberately does not claim
NFR-P7 is the whole chain, and only the encode half of it runs without an
adapter. So the export row is a one-sided gate — over two seconds in the encode
alone violates the requirement; under it proves nothing — and there is no
TRACES: NFR-P7 anywhere. NFR-P8 is about the application at idle, and the probe
is a process holding the catalog and nothing else, so it records the catalog
layer's share and carries no budget until somebody decides what that share
should be. No tag there either. CONTRIBUTING.md asks that a requirement be
closed by a test that would fail if the behaviour were removed, and two more
plumbing tags is what this repository already has too many of.
NFR-P8 also gets the answer §4.1 demands: RSS is exclusive of device-local GPU
allocations and cannot be made otherwise, because such an allocation never
enters the process's address space. The requirement should be restated as two
figures, and docs/benchmarks.md says so.
# Two gates, and why one of them steps aside off the reference desktop
The budget is the requirement's own number and never moves. The baseline is
what the reference desktop last measured, and drifting 15% past it fails the
build even while still inside the budget — which is how performance rot
actually arrives, never over the line, always a little worse.
A budget written for twenty-four threads cannot be asserted on a two-core
container. §8 names the reference desktop, not CI, so each metric declares
whether its budget is machine-sensitive; those are asserted under --reference
and reported everywhere else. Catalog open is not one of them: two seconds
against an expected figure two orders of magnitude smaller is a threshold any
machine can be held to. This is the trap core/dr-gpu/tests/frame_budget.rs
already refuses — a red gate everybody learns to ignore.
# The baseline ships with no numbers in it
Every recorded field is null, because nobody has run it yet. Writing
plausible-looking figures would make every later comparison a comparison
against a guess, and the first real regression would be invisible. Run
`dr-bench record --reference` on the reference desktop and commit the diff;
until then the budget gate works and the report says the other one cannot.
# CI
.gitea/workflows/benchmark.yml, and its own workflow rather than a step in
build-and-test.yml: a red "Build and test" says the code is wrong, a red
"Benchmarks" says it got slower, and the second must not be reachable by
retrying a flaky compile. The cpu job runs on every push and builds -p dr-bench
alone — which is why that crate depends on no GPU and no UI crate. The gpu job
is the frame budget that already exists and already skips without an adapter,
on workflow_dispatch, because building wgpu on every commit to rediscover that
the runner has no device is not a use of anybody's minutes.
The module doc said an extractor keyed on string literals would "lose six
genuine tags to save two false ones". The six is right — `schema.rs` carries
that many `-- TRACES:` lines inside Rust literals — but the two undercounted
the false ones, which were R1 twice, FR-CAT-1 twice, FR-CAT-2, NFR-P1, one in
`gestures.rs` and two emitted by `dr-pipeline/build.rs`. The comparison it was
drawing does not need a number on that side to hold.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
NFR-OPS-1 asks for structured levelled logging to a rotating, size-capped
on-disk log in the XDG state or Android app directory, automatic redaction of
credentials and tokens, and a one-click diagnostics bundle with an explicit
preview-and-consent step. Its two tags were on `compute_coverage` and on the
traceability tool's gesture extractor.
Neither is diagnostics under any reading. One computes a ratio and the other
generates a markdown document; neither writes a log, and no rotating on-disk
log exists anywhere in the tree — logging goes to stderr and to logcat.
These were real tags, not the fixtures the extractor was just taught to
ignore, which makes them the more instructive case: the tool was correct and
the tags were wrong. NFR-OPS-1 is untagged again, and outstanding.md §9 now
says what is actually missing rather than that the requirement is covered.
`gestures.rs` keeps its FR-UI-4 tag, which is a separate claim and unaffected.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The traceability tool scans `tools/`, which is its own source, and a line was
taken for a tag whenever `TRACES:` appeared anywhere on it. Its unit-test
fixtures are therefore tags. R1 — cross-platform output within a bounded
tolerance, the requirement with no acceptance criterion at all — was reported
implemented on the strength of two string literals in
`context_looks_forward_then_backward`.
R1 was only the visible case because it had no other coverage. The same
fixtures also contributed sites to FR-CAT-1, FR-CAT-2 and NFR-P1, `gestures.rs`
contributed one to FR-UI-4 from a `push_str`, and `dr-pipeline/build.rs`
contributed FR-DEV-3a and FR-DEV-3c from the tag it *emits* into generated
code. Those four requirements keep real tags elsewhere, so nothing but noise is
lost by dropping them.
The rule is about position, not about string literals. It cannot be about
string literals: `schema.rs` writes six genuine tags inside Rust string
literals, because the SQL it embeds is commented with `--`, and an extractor
that refused those would lose more than it saved. What separates the two is
where on the line the tag is. A tag written to be read is the first word of its
comment; a tag quoted inside an expression never is. So `tag_body` asks for a
comment opener at the start of the line and `TRACES:` immediately after it.
That closes every shape but one: a multi-line literal whose lines really do
begin with `///`, which no line-oriented reader can tell from source. There is
one such fixture and its ids are now UT and IT, which `is_requirement` already
excludes from coverage — the mechanism existed and was simply never used on
the tool itself. `this_crates_own_fixtures_cannot_reach_the_register` enforces
that: any requirement id below `mod tests` in this crate fails the test and
says to use a UT- or IT- id instead. Tagging the tool's real code is still
allowed.
On `SOURCE_SUFFIXES`, which cannot reach `AndroidManifest.xml`, the Flatpak
manifest, the Dockerfile or the CI workflows: it is deliberately left alone,
and the reasoning is recorded beside it. A tag on a manifest asserts that a
comment exists next to a line nothing checks, which is the weak form
CONTRIBUTING.md warns about. The convention already in the tree — a Rust test
that `include_str!`s the file and asserts what must be in it, with the tag on
the test — is what a tag is supposed to mean.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`models/LICENCE.md` recorded, on 2026-08-21, that no YOLO model trained on
ADE20K existed in usable form — the stuff classes photography cares about,
sky and vegetation and water, had no model to come from. Re-checked
2026-08-30: Ultralytics now ships a `semantic` task with ADE20K
checkpoints, so `models/scene/` holds `yolo26s-sem-ade20k`.
This is an addition, not a replacement. A semantic model labels every
pixel but merges same-class pixels into one region, so it cannot tell
three people apart — which is exactly what clicking a subject needs, and
exactly what `segment/`'s COCO instance model already does. The scene tab
grades per category and does not care that instances are merged. Keeping
both is the point.
## The export is truncated, deliberately
Ultralytics ends the graph with `Resize -> ArgMax -> Cast` and hands back
a `[1, 640, 640]` u8 label map. The script cuts that tail and exposes the
classifier's `[1, 150, 80, 80]` f32 logits instead, for two reasons.
Cost: the Resize materialises 150 x 640 x 640 x f32, 246 MB, and ArgMax
then reduces across the channel axis, striding 409,600 elements per
comparison. On one loaded machine the full graph ran ~1160 ms against
~500 ms truncated — roughly four fifths of the time spent on work the
application discards. Those numbers were measured under contention and
are upper bounds, but the ratio is structural.
Softness: ArgMax destroys the per-class scores, and the scene tab needs
them. Softmax over the 150 channels, summed within each photographic
category, yields per-category weights summing to 1 at every pixel.
Feathering a partition of unity cannot double-grade a boundary, whereas
feathering hard labels outward from two adjacent categories paints both
grades into the overlap and haloes every horizon.
The discarded upsample was never information: the graph's true spatial
resolution is the 80x80 logit grid, and the application can resample from
that itself.
The tail is matched by op type and asserted before cutting, so an
upstream graph change fails loudly in the exporter rather than quietly
shipping a differently-shaped model.
Nothing reads these weights yet — the decode path, the category
descriptor grouping 150 classes into ~8 photographic ones, and the scene
tab are still to come. At 24 MB this model also wants the runtime-asset
treatment `models/face/` already gets on Android rather than
`include_bytes!`; embedding it would put ~35 MB of weights in the binary.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`mktemp -d` lands in `/tmp`, which on this and most current Linux
distributions is a tmpfs — memory, not disk, sized at half of RAM. The
venv this script builds installs torch into it, several gigabytes, and
the failure mode is not subtle:
error: Failed to install: torch-2.13.0-...whl
Caused by: No space left on device (os error 28)
on a machine with 102 GB free on the filesystem holding `/var/tmp`. The
quieter version of the same bug is worse: when it does fit, it evicts
whatever the user had in page cache to make room.
`${TMPDIR:-/var/tmp}` respects an explicit TMPDIR and otherwise picks the
disk-backed directory, which is what a multi-gigabyte throwaway wants.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The weights were in two places: face detection and recognition in
`models/face/`, segmentation in `core/dr-segment/models/`. Nothing was
wrong with either path, but between them there was nowhere to look to
answer "how much model does this application carry", and that number is
about to start growing.
So the crate-local copy moves up beside the other. `models/` now holds
`face/` and `segment/`, and a `du -sh` of one directory is the whole
answer.
No content changes: the .onnx and its vocabulary are byte-identical, and
`LICENCE.md` moves up a level to cover the tree rather than one crate.
The LFS pattern in `.gitattributes` is `*.onnx` and already matched both
locations, so only its comment needed the new path.
`include_bytes!` is relative to the source file and `build.rs` runs with
the crate root as its working directory, which is why the two paths climb
a different number of levels.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Three holes, all found by the gate catching itself out.
**Prose that mentions the tag was read as a tag.** `dr-ui`'s module list
carries a comment saying where the generated table comes from, and it
names `GESTURE:` in passing; the scan extracted that sentence fragment as
a gesture with no place and no way to perform it. A tag must now *open*
its comment. A line that merely mentions it is describing the mechanism,
not declaring a member of it, and position is the only thing that tells
the two apart — which also makes the string-literal guard fall out for
free rather than being a special case.
**Neither artefact was regenerated on commit.** They cite line numbers,
so they go stale on anything that moves a line — the sheet commit made
the document wrong about every gesture in `library.slint` without
touching a single one. The pre-commit hook that already keeps the matrix
in step now keeps these too, and unlike the matrix it *fails* rather than
shrugging when the scan does: a matrix that will not build leaves a stale
one in place, where a malformed gesture block means a user about to be
told the wrong thing.
**CI did not check them at all.** It does now, blocking. The matrix is
read; the gesture table is *shown to somebody using the application*, and
a stale one tells them to perform a gesture that no longer exists — from
which they will conclude the application is broken rather than the page.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Every gesture the application has was documented in the comment beside
the `TouchArea` that implements it. Excellent comments, and unreachable
by anyone not reading the source — which is the FR-UI-4 failure in a
different costume: a gesture nobody can find is a feature only its author
knows about.
Writing them out again in a hand-kept help page is the failure this
avoids. Two descriptions of one gesture drift, and it is always the prose
that drifts: the code is exercised every time somebody uses the
application and the page is exercised never. A help screen confidently
describing a double tap the grid stopped honouring last week is worse
than no help screen — and the grid did stop honouring one, in the commit
before this.
So the comment beside the implementation stays the only copy, and a
`GESTURE:` block beside it is scanned into two artefacts: `docs/gestures.md`
for a reader, and a Rust table for the application to draw a help sheet
from. Both committed, both gated, so neither can quietly stop describing
the code.
It lives in the traceability crate because it is the same operation on
the same input — walk the tree, pull structured tags out of comments,
render, fail if the committed artefact has moved. Only the vocabulary is
new. It scans `ui` and `apps` alone: a gesture needs an interface to be
performed on, and excluding `tools` is also what stops the scanner
extracting its own worked examples as broken gestures.
Fifteen gestures so far, across the library grid and the People screen.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The similarity scan picks its dot product per machine, and the NEON one
is the kernel that ships to the phone and the tablet — and the one a
desktop cargo test never executes. A wrong lane index or a mishandled
tail there is a silent wrong answer on exactly the devices nobody runs
the suite on, which is a poor place for the only untested code path.
dr-face carries no weights and touches no display, so its tests are a
plain ARM64 binary that runs under adb shell with nothing installed.
The script builds it against the SDK's newest NDK, pushes it, runs it
and cleans up. It checks for the device first, so a tablet that is not
plugged in costs a second rather than the two minutes it takes to
compile for it.
Not wired into CI, which has no device attached.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`platform` was missing from the traceability scanner's source roots, so
every TRACES tag in `dr-plat` — the secret store, volume discovery, and
now display-profile acquisition — was invisible to the matrix. The
FR-PLAT-* family is exactly what that crate exists to satisfy, so the
omission understated coverage by the requirements it was meant to
count.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Neither InsightFace export parses as shipped -- SCRFD fails at its input
node, ArcFace at the first Conv -- which is the same wall dr-segment hit
on YOLO's dynamic export. Both load cleanly with the input dims frozen,
so the pure-Rust runtime holds for the face pipeline too.
tools/fix-face-model-shapes.sh does the freezing, and exists so the
artefact is reproducible rather than a binary someone once produced. It
takes two forms because the two graphs need different ones: ArcFace's
batch is a named dim_param, SCRFD's H and W are dynamic but unnamed.
Also notes YuNet loading with no intervention, which matters for the
licence question in faces.md 2.3.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Choosing Ilford HP5 Plus showed an inverted grey frame. So did Double-X.
They are negatives, and upstream leaves `target_print` null on every
monochrome stock, so nothing was ever printed and the scan was all there was.
A colour negative at least announces itself -- the orange mask says plainly
that you are looking at a negative. A monochrome one just looks broken.
They print on Kodak 2302 now, which is a monochrome print film and is what
such a negative is actually printed onto; Double-X onto 2302 is the standard
cine chain. For the Ilford stocks it stands in for an Ilford paper, which
nobody has measured, and is at least the right kind of material.
The scan is still reachable through the Scanned/Printed toggle. It is a thing
to choose now rather than the only thing on offer.
`every_shipped_stock_bakes` did not catch this, and could not: it derives
"should this be inverted?" from the stock's kind *and whether it names a
paper*, so it looked at an inverted HP5, concluded that was right for an
unprinted negative, and passed. The assertion was self-consistent and the
situation was still wrong. The new test asserts the thing that actually
matters -- a camera negative must name a paper, that paper must be a printing
stock, and it must be the same kind of material, so a monochrome negative
cannot end up on colour paper.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Pushing was not a thing to simulate. It was measured data being thrown
away: Double-X and 2302 each ship five characteristic curves, one per
development time, and this shipped the 6.5-minute column and discarded
four. All five now ship and interpolate.
The axis is real. Double-X runs 4 to 12 minutes, and across it the average
gradient goes 0.472 to 1.034 while Dmax goes 1.19 to 2.56.
The control is in stops, because that is what a photographer means, and one
stop is a factor of about 1.41 in time. That mapping is checked rather than
assumed: against Double-X's own axis it lands within 2% of the 9-minute
column for +1, and near 12 minutes for +2, which are the times the datasheet
gives for exactly that. There is a test.
**Pushing must not recover shadow detail, and this does not.** Across the
whole measured range the speed point moves about a third of a stop while the
gradient doubles; three stops under mid-grey, density goes from 0.008 to
0.035, which is still nothing. Developing longer multiplies what was already
recorded and cannot record what never hit the film. A push built as added
exposure or global contrast brightens those shadows instead and looks
convincing until someone who shoots film sees it, so that property has a
test of its own.
Interpolated in *log* time, because development is multiplicative: 4 to 5
minutes is the same amount of push as 9 to 12, and interpolating linearly
would bunch the control at one end. Clamped at both ends, because past the
published range there is no data and extrapolating a contrast curve invents
an emulsion nobody tested. A stock measured at one process ignores the
control entirely rather than inventing a curve for it -- Portra 800's pushes
are separate *measured* profiles, which is the honest way to offer those.
Costs nothing per pixel and changes no shader. The curves are a per-stock
table, so the interpolation happens on the CPU at bake time, where choosing a
stock and moving its sliders already rebakes. The Vulkan shader is untouched.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
They were asked for and they are here, but not on the same footing as the
Kodak profiles, and the files say so in their first line.
What I had claimed, and had to withdraw: that Delta 100 is "quoted around 9"
and HP5 "around 12". Ilford publish no such figures. The word granularity does
not occur anywhere in their technical information -- grain is described as
"fine" and "finest" and nothing more. That claim was in this crate's
documentation as though it came from a datasheet; it is corrected there too.
Two further traps found while looking:
- Kodak colour negatives publish Print Grain Index, not RMS granularity.
PGI is a perceptual scale from viewer surveys -- 25 is roughly the
threshold of visibility, four units a just-noticeable difference -- and
Kodak state it cannot be compared to RMS. So a Portra number cannot be
dropped into the granularity field, and none has been.
- RMS proper is published mostly for black-and-white, reversal and motion
picture stocks. Every shipped stock therefore still carries the same
default, which means grain does not yet tell one film from another. That
is per-stock data, not code, and is now written down where somebody will
find it.
So the Ilford profiles are built rather than extracted, and each part rests on
something different:
speed published and exact -- ISO 400/27 for HP5 is a fact
contrast ISO 6:1993's normal development, average gradient 0.62
spectral borrowed from Kodak Double-X, a *measured* panchromatic
negative, shifted by the speed difference. Conventional
panchromatic sensitisation is much alike across black-and-white
films, and this is far better founded than reading pixels off a
printed curve
silver neutral, which is not an approximation: developed silver
absorbs flat, and Double-X's measurement is flat
granularity estimated, ordered by each film's known relative grain
They render as a film of that speed and contrast. They are not a measurement
of that emulsion, and the two stocks that share a speed differ only in the
estimated part.
`every_shipped_stock_bakes` is tightened to match, because a constructed
profile fails in a way a measured one does not: the curve parses, bakes, and
sits entirely off one end of its own exposure range, rendering every frame
black or blown while passing a finiteness check. It now asserts mid-grey lands
somewhere photographic and that the tone response runs the way the stock's
kind says it should.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>