Files
DarkRoom/docs/outstanding.md
T
dtourolle 4b6c110816 Count the sensor's own numbers, so a cull can see headroom the render hides
FR-CULL-3's remaining two bullets. What existed was a *display* histogram
tagged FR-DSP-7: it binds AdjustPass's Rgba8Unorm output, recovers an 8-bit
code value, and counts clipping as `r == 255`. Its own documentation says a
clipped bin means "a highlight that is actually gone rather than one the
transform might still recover", which is the opposite of what a culling
decision needs. FR-CULL-3 asks for the histogram of the sensor data, on the
explicit grounds that a rendered image "systematically lies about what is
recoverable in the raw", and a readout that measures the render cannot answer
that however it is presented.

So this is a second instrument beside the first rather than a setting on it.
Both are true; they are true about different things; the panel offers both
behind a chip row and the words travel with the numbers, because a raw
saturation figure drawn under a heading saying Highlights would be mislabelled
exactly where the difference matters.

**What is reduced over, and what it cost to decide.** ARCH §5.5 specified the
pre-demosaic CFA samples. This reduces over the demosaiced scene-linear
texture instead, and §5.5 is amended to record the choice rather than let the
specification and the code disagree in silence. The texture is camera-native —
unbalanced, unmatrixed, uncurved — and normalised by the sensor's own black
and white levels, so 1.0 is saturation by construction and the distribution
below it is the headroom question with no calibration to carry. Retaining the
CFA samples would mean keeping the packed u32 buffer Demosaicer::run currently
drops: 48 MB at 24 MP, 120 MB at 60 MP, resident per open photograph whether
or not anyone looks at the histogram, on a platform §6.2 exists because memory
is scarce on.

Three things it therefore cannot say, written into the module docs and into
§5.5 rather than left to be discovered: it counts pixels not photosites, so a
saturated site drags its interpolated neighbours up and per-channel clipping
is smeared by about a demosaic kernel; it cannot see above white, because
demosaic.wgsl clamps each photosite at 1.0 for its own good reasons (a Canon
6D reads to 16383 against a declared 15070) so "at saturation" and "a stop
past it" share a bin; and it is measured after the CFA pattern is gone, so it
can name which colour clipped in the reconstructed image but not which
photosite went first.

The axis is stops below saturation, 16 bins per stop over 256 bins — the same
bin count the display reduction uses, so the fold into drawable columns is
shared and a divergence between the two plots would have to be deliberate. A
linear axis spends half its width on the top stop, which is why nobody has
ever drawn a useful linear raw histogram. The fourth series is the brightest
channel rather than luma: these values are unbalanced, so any weighted sum of
them is a number about nothing, and the brightest channel is the one that
saturates first and so the one the headroom question is actually about.

It is a property of the file and not of the render, which has two
consequences. It is computed once per photograph and cached — nothing
downstream of the demosaic can move a count in it — so a cull does not pay the
display histogram's per-frame cost three thousand times. And it describes the
whole frame rather than the visible region, deliberately opposite to
DevelopSession::histogram: a crop changes what is on screen and changes
nothing about what the sensor recorded.

Tags are on the reduction, the type, its constructor and the presentation
arithmetic, each of which has a test that fails if the behaviour goes. The
Slint panel and the push from lib.rs keep their reasoning as prose: nothing
asserts them, and a tag would claim coverage the assertions are not making.
2026-08-29 23:36:21 +02:00

23 KiB

DarkRoom — Outstanding work

Status: Living document · first written 2026-08-29 Companion to: requirements.md §7, technical-debt.md, traceability.md

What is specified and not built, and for each cluster whether that is a decision, a dependency, or a gap nobody has looked at.

This document exists because traceability.md cannot tell those apart. It reports one number — the share of requirements carrying a TRACES tag — and a missing tag means either "nobody has built this" or "somebody built it and did not say so". Both read the same way in the summary table, which makes that figure pessimistic and uninformative at once: it understates what works while hiding which of the remainder matters. Eight requirements gained a tag on this branch because the code already satisfied them and nobody had said so. Everything below is the other kind.

It is also not a plan. requirements.md §7 records what was deferred deliberately and needs no argument; this records what is still nominally in scope, so that the distance between the register and the binary is visible rather than something a reader has to reconstruct from a percentage. Where the honest answer is "this requirement should be amended rather than met", it says so — an unbuilt requirement that nobody intends to build is worse than a deferred one, because it keeps costing attention.

Three of these are being built right now, in parallel worktrees, and are marked ⟳ in progress where they appear: burst grouping (FR-CULL-5), Flatpak packaging (FR-PLAT-LIN-3), and Android platform integration (FR-PLAT-AND-2/4/5/6). Strike those lines as they land rather than rewriting around them.


1. Plugins — 21 requirements, and a contradiction to resolve before any of them

Untagged: FR-PLG-1, -1a, -2a, -2b, -2c, -3, -3a, -4, -4a, -5, -5a, -5b, -5c, -6, -6a, -7, -8, -9, -10, -11, -12.

No plugin host exists. No crate loads anything at runtime: there is no manifest reader, no WASM or Lua engine, no registry, no signature check, no install path, no capability grant, no per-plugin failure ledger. declared/mod.rs says as much in its own documentation — the operation format is "not a plugin directory read at startup".

Two of §3.10's requirements are met, and they are the interesting two. FR-PLG-2 and FR-PLG-2d — the declarative node format — are built and tagged: core/dr-pipeline/ops/*.yaml compiled by build.rs, with the restricted expression grammar in declared/expr.rs and a parity test asserting a declared operation and a hand-written one produce identical output. code-health.md §3 calls it "a working plugin system that happens to resolve at build time", and that is exactly right. What is missing is not the format; it is everything that would let somebody who is not in this repository use it.

The contradiction. requirements.md §7 lists | Plugin API | — | among the things deferred for v1 — a bare row, where most deferrals carry a justifying note. §3.10 then spends roughly 280 lines and 23 requirement IDs specifying that same Plugin API in detail. Both statements are in the register of record, and the traceability denominator counts the second one: 21 IDs, 12% of all 179 defined requirements, worth about twelve points of coverage on their own — and nearly a third of everything the matrix reports as uncovered. A reader looking at the coverage figure has no way to know that, or that the subsystem behind it is one the same document says is not in this version.

And D16 is open. Decision D16 — plugin licensing — records that GPLv3 answers the derivative-work question differently for each of §3.10's three plugin forms, and that this "must be answered before an ecosystem exists, not after", because a term introduced later cannot be applied to plugins already written. D16 explicitly does not block FR-PLG-2; it blocks publishing a third-party format as stable.

What would resolve this: an edit to requirements.md, not code. Either §7 drops the row, or §3.10 is marked deferred with the two built requirements carved out. Until one of those happens the coverage figure is measuring a decision that has already been taken, and taking it again every time somebody reads the matrix.


2. Culling — the stated differentiator, half built

D11 names culling "the core differentiator". FR-CULL-1, -2, -3, -4 and -8 through -12 are built. Three are not.

FR-CULL-3 — Raw-truth overlays. Built, all three bullets. Focus peaking is core/dr-gpu/src/focus.rs and ui/dr-ui/src/peaking.rs; the raw histogram and the raw clipping indicators are core/dr-gpu/src/raw_histogram.rs and the second reading of the panel in histogram.slint.

Worth recording, because it is the thing this entry previously got wrong and the next reader will have to check again. The display histogram — dr-gpu/src/histogram.rs, ui/dr-ui/src/histogram.rs — is not this requirement and never was: it is tagged FR-DSP-7, it reads AdjustPass's 8-bit output, it counts clipping as r == 255, and so it describes the frame the display is about to show, after the whole develop chain. FR-CULL-3 asks for the sensor data, on the explicit grounds that a rendered image "systematically lies about what is recoverable in the raw". The two now sit in one panel behind a chip row, which is the arrangement that keeps them from being mistaken for each other: they answer different questions and both are true.

What the raw reduction cannot answer is written down rather than left to be discovered — architecture.md §5.5 records why it reduces over the demosaiced texture instead of the CFA samples §5.5 originally specified, and what that costs in what it can say.

FR-CULL-5 — Burst and near-duplicate grouping. Absent. Worth knowing before it is built: core/dr-face/src/calibrate.rs already assumes it exists — "since FR-CULL-5 already groups bursts, positives are bootstrapped from bursts" — and in fact bootstraps from confirmed labels instead. That comment is a forward reference to this requirement and will need correcting either way. core/dr-catalog/src/dedup.rs is not this: it is re-import detection under FR-CAT-11, matching a file against one already catalogued, not two photographs against each other. ⟳ in progress.

FR-CULL-6 — Compare and survey. Absent. No side-by-side view, no synchronised zoom or pan. This is the one of the four with no adjacent machinery at all, and it is also the one that most directly distinguishes culling from browsing.

FR-CULL-7 — Culling on tablet. Absent, and blocked by the three above rather than independent of them: there is no separate tablet culling surface to build until there is something to put on it.


3. FR-DEV-3g — AI denoise

Promoted into v1 by D11, and named there as the precondition for deferring AI masking — the argument being that one learned stage earns the runtime that a second could then reuse. Only classical noise reduction exists: ops/noise_reduction.rs, a bilateral filter in two arrangements, exact for luminance and separable for chroma. It is good, and it is not this.

models/ holds two face models and nothing else; core/dr-segment/models/ holds a YOLO segmentation model for subject masks. There is no denoise model, no learned demosaic, and no inference path that is not face or segmentation.

The obstacle is not the pipeline. It is that D13 — model licensing — is still open for the models that already ship, and adding a third learned stage adds a third licence to answer for. Building the runtime before that is settled means owning the same problem in one more place.


4. The render path — FR-DSP-2, FR-DSP-4, NFR-RES-2

FR-DSP-2 — Tiled computation. Unbuilt, and under challenge. architecture.md §6.2 calls for tiling "from day one" on the grounds that retrofitting it is a rewrite. It was not built, and the evidence has since moved. core/dr-gpu/tests/frame_budget.rs carries the argument in its own header: one fused dispatch over a viewport-sized target is comfortably inside the frame budget, and "if that stops being true, the recommendation to strike tiled computation from the interactive path stops being supported, and this test is what says so." technical-debt.md TD-4 reaches the same place from the other direction — a tiled convolution at clarity's radius reads nearly twice the taps that an untiled one does, so the stage that looks most like it wants a tile cache is the stage that would be hurt most by one.

What exists is the declaration and not the mechanism: DetailPass::radius is documented as the halo a tile would have to be grown by, with a test that pins it, and there is no scheduler to read it. That is deliberate plumbing, not an oversight.

So the open question here is not "when is tiling built" but "is FR-DSP-2 still a requirement". Two measurements say it costs more than it saves on the interactive path. Neither says anything about the export path or about a device under memory pressure, which is where the case for it actually lives — and that is spike S6, which has not run.

FR-DSP-4 — Progressive refinement. Unbuilt. FR-DSP-1's proxy rendering and TD-4's quarter-resolution base are adjacent and are not it: both are fixed choices about what resolution to compute at, where FR-DSP-4 asks for a first frame that is deliberately cheap and a second that replaces it. Nothing tracks a "this frame is provisional" state.

NFR-RES-2 — Images larger than GPU memory. No answer, and §4.3 knows it: the requirement text itself asks the reader to "decide explicitly" how ARCH §6.4 and NFR-RES-2 are reconciled. There is no headroom budget, no allocation-failure fallback, and no spill. Spike S6 — a tiled pipeline on a mid-range Android device with an image larger than available GPU memory — is the one that would settle both this and FR-DSP-2, and there is no evidence it has run.


5. Android beyond running, and Flatpak

The Android app is not a stub — it builds an APK, runs the whole application, unpacks bundled face models, and has been measured on a tablet (faces.md §12.1, technical-debt.md TD-1). What is missing is the platform contract around it.

FR-PLAT-AND-1 is tagged and should not be relied on. The requirement demands that library access be obtained exclusively through the Storage Access Framework. There is no SAF code: no ACTION_OPEN_DOCUMENT_TREE, no takePersistableUriPermission, no DocumentsContract. The two tags rest on a SourceRef::Document variant that nothing constructs and a volumes helper, which is the "plumbing a future feature would use" case CONTRIBUTING.md and code-health.md CH-4 both warn about. Android reaches a library through a Nextcloud account or a folder, over paths, like the desktop.

That has a consequence for the rest of the cluster: FR-PLAT-AND-2 — detecting the loss of a granted tree permission and marking images offline rather than deleting rows — cannot be built until there is a permission to lose. It is listed here as unbuilt, but it is blocked, not skipped.

FR-PLAT-AND-4 (managed background execution, foreground service for exports, stated Doze behaviour): the manifest declares one activity, no service, and neither FOREGROUND_SERVICE nor POST_NOTIFICATIONS. FR-PLAT-AND-5 (onTrimMemory with a stated eviction order): no callback is registered, though the eviction order it is supposed to drive is specified in FR-NC-6's text. FR-PLAT-AND-6 (view and share intents, FileProvider): the only intent filter is MAIN/LAUNCHER. ⟳ in progress for this group.

FR-PLAT-LIN-3 — Flatpak. packaging/ holds an Arch PKGBUILD and a .desktop entry. There is no Flatpak manifest, nothing goes through a portal, and platform/dr-plat/src/secrets.rs talks to the Secret Service directly rather than through the portal the requirement names. ⟳ in progress.

NFR-COMPAT-2 — distribution channels. Unstated, and this is the requirement that makes the others binding: §4.8 observes that the decision to publish on Play is what turns SAF from a preference into a constraint. Spike S11, the Play permissions dry-run that would settle it, has not run. Related, NFR-COMPAT-1's baseline is real but scattered — API 28/36 live in the Android Dockerfile and are checked in CI against the built ELF, which is good — while the items the requirement singles out are missing: whether shaderFloat16 and 16-bit storage are required (the one it flags as jeopardising R1), minimum RAM, minimum desktop Mesa, and a named reference device from a second GPU vendor.

NFR-OPS-2 and NFR-OPS-4. Crash reporting is a log::error! panic hook on Android and nothing at all on desktop: no local crash record, no backtrace capture, no upload path and therefore no opt-in gate to guard it. Update and first run are undefined; the concrete reason NFR-OPS-4 gives — that D2 pins rawler at a non-SemVer alpha whose camera-support fixes users will need — is unaddressed, and there is no update mechanism of any kind.


6. Accessibility and internationalisation — the hard half is done and the easy half is not

NFR-A11Y-1 — Localisation. @tr( appears zero times across 14,482 lines of Slint. That number overstates the problem, because the part that is genuinely architectural was got right: LocalizedKey keeps display strings out of core/ entirely, every operation publishes a key rather than a label, and labels::resolve is the single point where a key becomes text. What that single point does, however, is a hardcoded English match in Rust source — so changing a translation requires a recompile, which is the one thing the requirement explicitly forbids. There is no message catalogue in any format, no locale-resolution rule, and no decision recorded about RTL.

The work left is therefore smaller than it looks and entirely mechanical: a catalogue format, a load path behind resolve, and @tr( around the Slint literals. The design it needs already exists.

NFR-A11Y-2 — Accessibility. accessible-* appears five times in the whole interface, all five on one control — the parameter slider in adjust.slint — and nothing is set from the Rust side at all. Everything else in eighteen Slint files is unnamed to AT-SPI and TalkBack. The requirement's own caveat, that Slint's Android accessibility needs verifying, is spike S13, which has not run.

NFR-A11Y-3 — Colour-independent status. No compliance work found. This is cheap to satisfy while a control is being written and expensive to retrofit across forty of them, which is an argument for doing it as part of the NFR-A11Y-2 pass rather than after it.


7. Catalog and sync

FR-CAT-14 — Migration import. Reading ratings, labels, keywords and collections out of a Lightroom .lrcat or a darktable library.db. Unbuilt. The destination is not: keywords, collections, ratings and the cross-device merge rules are all built and tested, and keywords.rs already anticipates the arrival ("an import from Lightroom can bring in…"). What is missing is only the two source adapters — which is a comparatively contained piece of work for a requirement that decides whether somebody can try this software on a library they already have.

FR-NC-11 — Initial catalog build. Using WebDAV SEARCH (RFC 5323) against /remote.php/dav/, filtered by mimetype and paginated, in preference to walking folders with PROPFIND. Unbuilt: no SEARCH request is issued anywhere. The PROPFIND walk this exists to replace is fully built and well optimised — ETag pruning under FR-NC-4 turns an unchanged 50k library into one request — so the gap is narrower than it reads. It is the first build against a large remote library that pays, and that is the moment a new user meets.

FR-CAT-13 — XMP interoperability, tagged and not met. Read and write standard XMP sidecars. The single tag sits on keywords.rs, which stores keywords; no XMP is parsed or written anywhere in the tree, and dr-export's metadata module says so about its own half ("neither is read by dr-decode today"). Listed here rather than silently, because a tag makes a gap invisible and this one is load-bearing for interoperating with the editors FR-CAT-14 imports from.


8. The performance targets are unverified, not unmet

Eleven of the fifteen §4.1 targets carry no tag: NFR-P2, -P3, -P4, -P6, -P7, -P8, -P10, -P11, -P12, -P14, -P15. That is the uninteresting part of this section.

The interesting part is that §8 and §4.1 both require the same thing, in the same words, and it does not exist: an automated benchmark suite against a synthetic 50k catalog, run per commit, where "a regression beyond a stated tolerance is a build failure, not a notification." There is no benches/ directory in the workspace, no criterion dependency, and no synthetic catalog. The three CI workflows run cargo fmt --check, clippy, cargo test --workspace, a release build, an Android cross-build and a layering check. None of them measures anything, so there is no baseline to regress against and no tolerance to exceed.

What does exist is narrower and genuinely good: dr-gpu/examples/frame_budget is a real instrument, its results are committed in frame-budget.md with the machine and profile named, and TD-4's before-and-after was measured with it. But it is run by hand — frame-budget.md's own instruction is "rerun and diff this file" — and the guard version that does live in CI skips itself where there is no GPU adapter, which the workflow notes is the normal case on a runner, while asserting its CPU half only when debug_assertions is off, which a dev-profile cargo test is not. In CI it therefore asserts approximately nothing.

The claim to take from this is precise. Nothing here says the performance targets are missed. Several are plausibly met. It says that if one were broken tomorrow, nobody would find out — which is the failure mode §8 was written to prevent, and the reason it belongs in this document rather than in a backlog.


9. Two core requirements that cannot be closed as written

R1 — Cross-platform output within a bounded tolerance. §2 states that the threshold "must be fixed before spike S9", because S9 both validates R1 and calibrates what tolerance is achievable. The threshold was never fixed and S9 has not run, so R1 currently has no acceptance criterion at all — there is nothing a test could assert.

Worse, the matrix reports R1 as covered. Both of its tags are string literals inside the traceability tool's own unit tests (tools/traceability/src/lib.rs), which the tool scans along with everything else, because a fixture demonstrating tag extraction is indistinguishable from a tag. NFR-OPS-1 is covered the same way, from a tag on compute_coverage — and no rotating, size-capped on-disk log exists; logging goes to stderr and logcat. These are two of the cases CONTRIBUTING.md already warns about, now named.

R2 — Efficient display of huge RAW libraries. Its acceptance criterion contains "(figure TBD)" — the scroll velocity below which no cell may render as a placeholder — and asks for a stated prefetch margin and cache-hit rate. No figure is stated anywhere in the tree, neither quantity is measured, and TD-2 and TD-3 both describe the thumbnail path falling short of it in ways that were measured. R2 was deliberately left untagged on this branch for that reason: the machinery is substantial and the criterion is unmet and partly undefined.

Both belong with §8 above. A requirement whose threshold was never chosen and a target nothing measures fail in the same way — not by being wrong, but by being unfalsifiable.


10. Spikes

§9 defines fourteen validation spikes and says of three of them: "S1, S2 and S10 are the three that can invalidate the architecture."

Only S1 (Slint + wgpu zero-copy on Linux) and S14 (the face pipeline on a real library) have recorded results. S14's are the best evidence of any spike — a dedicated document, a measured pass over an 18,143-face library, a named device and a reproducible command — though D13's licensing half remains open.

S6, S9, S10, S11 and S13 show no evidence of having run at all. Each is referenced only from the requirement text that asks for it:

Spike Would settle Blocked on
S6 FR-DSP-2, NFR-RES-2 — tiling and images larger than GPU memory Nothing; needs a device and a large image
S9 R1's tolerance threshold, and therefore R1 Nothing; the threshold is defined by running it
S10 Whether SAF at 10k files meets NFR-P1/P3 §5 — there is no SAF code to measure
S11 NFR-COMPAT-2, and whether Play makes SAF binding Nothing
S13 NFR-A11Y-2 on Android §6 — there is almost nothing to test with

S2, S3, S4, S5, S7, S8 and S12 are also unrun, several with acknowledgements in the code that say so (dr-sync/src/upload.rs on S8, dr-sync-nextcloud/src/lib.rs on S3). S2 is one of the three architecture-invalidating spikes and needs Adreno and Mali hardware, which the manifest notes no emulator represents.

The pattern is worth stating rather than leaving to be inferred: the spikes that ran are the ones whose subject was being built anyway. The ones that did not are the ones that would have said whether something should be built — which is the opposite of the order §9 asks for.


11. D12, which governs all of the above

Decision D12 — scope versus pace is still OPEN, and says:

The calibration selected an ambitious feature set — full tablet editing, full ingest, culling as a differentiator, complete GPU masking, AI denoise, Fuji-first colour, deep sync, sidecar durability — against a stated pace of evenings and weekends, indefinitely.

Those are not compatible as stated.

Sections 1 through 10 are what that incompatibility looks like eleven versions later, and they land almost exactly where D12 predicted: tablet editing carries SAF at unproven scale, background execution limits and two GPU vendors to validate (§5), and every one of those is unbuilt or unrun. The parts that were built — the develop pipeline, sync, faces, the catalog — are the parts that did not need a decision first.

D12 is not resolved by choosing to work faster. It is resolved by moving requirements across the line into §7, which costs nothing but the admission, and which this document is intended to make easy: every cluster above is a candidate, and each says what it would take to build and what it would cost to drop. Resolving D12 sets D3 and architecture.md §10's Phase 2.