# DarkRoom — Technical debt **Status:** Living document · first written 2026-08-26 **Companion to:** [architecture.md](architecture.md) Deliberate compromises: things the code does knowing they are wrong, because the alternative was worse at the time. Each entry says what the debt is, what it cost to take on, what it would take to pay off, and how you would know it had been paid. Not a bug list. A bug is something nobody chose. Everything here was chosen, and the point of writing it down is that the reasoning outlives whoever chose it — so the next person can tell a constraint from an accident, and does not "fix" something load-bearing or preserve something that has quietly stopped being necessary. --- ## TD-1 — The Android develop view reads pixels back through the CPU ✅ PAID OFF **Breaks:** [architecture.md §12 / 6.1](architecture.md) — GPU results never round-trip through the CPU — and AC-8, on Android only. Desktop is unaffected and keeps the zero-copy path. ### What it does `DevelopSession::render` on Android runs the compute passes on the GPU as usual, then calls `AdjustPass::export_pixels` and hands the frame to Slint as a `SharedPixelBuffer`. That is exactly the GPU→CPU→GPU transfer §6.1 exists to forbid, and it is on the frame path. ### Why Zero-copy needs Slint to draw with wgpu. On Android that means wgpu's Vulkan swapchain, which hardcodes `preTransform = VK_SURFACE_TRANSFORM_IDENTITY_BIT_KHR` ([gfx-rs/wgpu#3345](https://github.com/gfx-rs/wgpu/issues/3345)) — wgpu-hal says so in a comment beside the line. On a tablet whose panel is mounted landscape, a portrait window then hands Android an unrotated buffer, every present returns `VK_SUBOPTIMAL_KHR`, and frames arrive torn. Measured on the device, same build, only the tablet rotated: | orientation | `bufferTransform` | composition | result | |---|---|---|---| | landscape | `ROT_180` | `DEVICE (2)` | clean | | portrait | `ROT_270` | `CLIENT (1)` | torn | Setting `preTransform` is not a fix available to us: the field is a *promise* that the content is already rotated, so honouring it needs the renderer to rotate what it draws, which wgpu cannot do on Skia's behalf. So the choice was never fast-develop against slow-develop. It was a develop view that costs a readback against a grid that tears in the orientation a tablet is mostly held in. ### What it costs Less than §6.1's headline numbers, because `render` fits the pass to the canvas before it runs — the readback is at viewport resolution, not sensor resolution. The 7.43 ms at 4K in §12 is the ceiling, not the bill. **It has not been measured on the device**, which is the first thing to do if the develop view feels heavy on the tablet; do not assume this is the cause without a number. ### And a second transfer, while focus peaking is on Added 2026-08-29 with FR-CULL-3. The focus-peaking overlay is a compute pass writing its own `Rgba8Unorm` texture, which on desktop reaches the compositor with no copy — but on Android there is no more a path for *that* texture than for the frame it belongs to, and an overlay that stayed on the device while the picture underneath it did not would simply never be seen. So `FocusPeakPass::read_overlay` follows the frame back through memory, and the Android frame path carries **two** full-resolution `copy_texture_to_buffer` transfers instead of one. This is recorded under TD-1 rather than as its own entry because it is not an independent choice. It exists only because TD-1 exists, it is bounded by the same thing — `render` fits the pass to the canvas, so both transfers are at viewport resolution — and TD-1's "Done when" already covers it: whichever of the three fixes above lands removes the readback for the frame and the overlay together, because both are the same missing capability. Two things worth saying plainly. The doubling is **reasoned, not measured on the device** — the same gap TD-1 admits about its own cost, and the reason neither number should be quoted as a measurement. And it is paid only while the photographer has the overlay switched on: `DevelopSession::focus_overlay` returns on its first line when peaking is off, so with it off there is no dispatch and no transfer, and the Android frame path is exactly what it was before this feature existed. ### Paying it off Any one of these removes it: - wgpu implements pre-rotation (#3345), and Android goes back on `unstable-wgpu-29`. - Slint's Skia Vulkan surface handles `preTransform` and Android uses that instead of OpenGL. - Skia over OpenGL grows a way to sample an external texture that wgpu can write. **Done when:** `ui/dr-ui/Cargo.toml` no longer scopes `renderer-femtovg-wgpu` and `unstable-wgpu-29` to non-Android, the `#[cfg(target_os = "android")]` arm of `DevelopSession::render` is gone, and the tablet is clean in portrait. ### Paid off By none of the three routes above. It came from a fourth that the list missed: patching both halves locally. "wgpu cannot do [the rotation] on Skia's behalf" was true and beside the point, because Skia can rotate its own canvas. Slint's Skia renderer already does it for rotated panels on linuxkms, just not on its wgpu surface. So two small patches, carried in `third_party/` ([README](../../third_party/README.md)), do it for us: - **wgpu-hal 29.0.4.** `vulkan::Surface::set_pre_transform` lets a caller choose the swapchain's `preTransform`, and `current_transform` reads the surface's. It is opt-in, and the default stays `IDENTITY`, so desktop is unchanged. - **i-slint-renderer-skia 1.17.1.** On Android the wgpu surface sizes its swapchain in the panel's orientation, promises the display's transform, and concatenates the matching rotation onto the canvas before Slint draws. The imported develop texture goes through the same matrix as everything else. The surface re-reads the transform every frame, because a half turn does not resize the window. The item renderer's pixel snapping now accepts right-angle rotations, or portrait would have lost it everywhere. With those in place, `unstable-wgpu-29` is common to both platforms, `shared_gpu` hands the one device to Slint on Android too, and both readbacks are gone (the frame through `export_pixels` and the focus overlay through `read_overlay`). The develop frame and its overlay reach the compositor as textures, as they do on desktop. **Verified by eye, not by instrument.** On 2026-09-25 the user checked the release-signed build on the tablet and found it clean. That is the portrait tear this entry was opened over. The `dumpsys SurfaceFlinger` readings that would complete the table above (composition and `bufferTransform` per orientation, expected `DEVICE` in all three) were not taken, because adb would not hold the device that morning. Nor was the develop frame time measured before or after. The readback's cost was never measured either, so the saving is reasoned, not measured. **The cost moved rather than vanished:** two upstream crates are pinned by path. A Slint or wgpu bump has to carry the patches forward (the README says how), and cargo only *warns* when a patch no longer matches, then silently builds the unpatched crate. Both patches should go upstream: the Skia half is the Android counterpart of a feature Slint already has, and the wgpu half is #3345. --- ## TD-2 — Thumbnails are fetched one at a time **Where:** `library::spawn_thumbnails` — the `for req in to_fetch` loop. ### What it does The interactive thumbnail batch fetches serially: one image at a time, and two HTTP round trips each (a header read, then the preview's byte range). A window of a few hundred cells is that many sequential round trips against the server. ### Why it is debt rather than a bug It is correct, and it was fast enough when a window was one screenful. It is the *ordering* that kept it survivable: since `fetch_rank`, on-screen cells are requested first, so the cells a person is looking at arrive first even though the queue as a whole is slow. Portrait makes it worse by construction — a narrow window means smaller cells, more rows, and two to three times as many cells on screen at once, all of them ahead of the ones below in a queue that never runs more than one request. ### Paying it off `spawn_thumbnail_sweep` already has the pattern: `SWEEP_LANES` disjoint lanes over a chunk, joined, with the store written on the one thread that owns it. Striping a *priority-ordered* chunk across lanes keeps `fetch_rank`'s ordering while running several requests at once. Not done yet because it multiplies concurrent requests against the user's Nextcloud during a scroll, and that is a behaviour change worth deciding on deliberately rather than inheriting from a performance fix. **Done when:** the interactive batch runs on more than one lane, priority order is preserved across the lanes, and a slow server still cannot stall the visible cells behind offscreen ones. --- ## TD-3 — The thumbnail drain applies an unbounded batch on the UI thread **Where:** `library_ui::drain_thumbnails` — the `loop` inside the timer callback. ### What it does Every message queued when the timer fires is applied in that one callback, with no ceiling. On a library whose thumbnails are already in the store, the worker delivers a whole window at once, so a single callback can do hundreds of `to_slint_image` calls back to back — each an allocation and a full RGBA copy — while the grid is mid-flick. The copy cannot move off the UI thread: `slint::SharedPixelBuffer` is not `Send`, so decoded bytes can only become an `Image` on the thread that draws. Only the *amount done per wake* is ours to choose, and right now it is "all of it". ### Cost Measured with a temporary probe, **debug build**, so treat the shape rather than the size: | class | per thumbnail | × a 280-cell window | |---|---|---| | grid, 256 px | 1.93 ms | 539 ms | | large, 512 px | 7.78 ms | 2.18 s | A release measurement was started and never completed — do not quote these as release figures. ### Paying it off A time budget per wake and a shorter interval: apply for a few milliseconds, return without stopping the timer, and finish on the next tick. A batch then lands in frame-sized slices rather than one lump between two frames. Draft written and discarded during the investigation; it is a small change. **Done when:** one wake of the drain cannot exceed a frame, and a fully-cached window still fills in well under a second. --- ## TD-4 — The local-contrast base is computed at full render resolution ✅ PAID OFF **Where:** `dr_pipeline::ops::local_contrast::LocalContrast::passes` — the `base` and `combine` passes, and the stage that dispatches them, `dr_pipeline::detail`. Breaks **FR-DSP-3** at large viewports. Measured, and the numbers are in [frame-budget.md](frame-budget.md) §M3. ### What it does Clarity's Gaussian σ is 1.2% of the frame's shorter edge, truncated at 2σ, so its kernel radius is a property of the *viewport*: 29 px at 1920 × 1200, 38 px at 2560 × 1600, **52 px at 4K**. The two separable passes therefore run 105 taps each over 8.3 M pixels at 4K, which is 1.7 billion texture reads for one control. | viewport | radius | clarity alone, p99 | |---|---:|---:| | 1920 × 1200 | 29 | 5.99 ms | | 2560 × 1600 | 38 | 12.44 ms | | 3840 × 2160 | 52 | **33.89 ms** | RTX 3050 laptop, `examples/frame_budget`, fused dispatch reused so this is the convolutions alone. Clarity is 97% of the cost of all four neighbourhood operations together at every size. For scale: the entire fused chain — every point operation active, film stock included — costs 4.5 ms at the same 4K viewport. **A single slider is seven times the rest of the pipeline.** ### Why Because the stage cannot do otherwise yet. `dr_pipeline::detail` dispatches every pass at the render size; there is no way to express "read this target and write a smaller one". The module's own documentation has said so since it was written: > The right optimisation is a base computed at reduced resolution, which needs a detail stage that > can write a smaller target than it reads; that is a change to `crate::detail`, not to this file. It was the right call to ship the correct answer slowly rather than a fast approximation nobody had checked — the halo behaviour is the hard part of this operation and it is tested. ### Not a tiling problem Worth saying because ARCH §5.3 offers a tile cache and this is the stage that looks like it wants one. It does not: a tiled convolution reads a halo per tile, so at a 52-pixel radius, 256-pixel tiles would read (256 + 104)² instead of 256² — very nearly **twice** the taps. [display-and-extension.md](display-and-extension.md) §2's decision rule was resolved on this evidence; see [frame-budget.md](frame-budget.md). ### Paying it off A detail pass that declares an output scale, so the base can be computed at a quarter resolution and sampled back up in `combine`. A quarter-resolution base is 1/16 the pixels at 1/4 the radius — about **1/64 of the work** — and is visually identical, because a base at σ = 26 px holds no content above the quarter-resolution Nyquist to lose. Texture's σ is a decade finer and must stay at full resolution; the scale therefore belongs on the `DetailPass`, not on the stage. **Done when:** clarity at 100% is inside the frame budget at 3840 × 2160, the halo tests in `tests/local_contrast.rs` still pass unchanged, and `examples/frame_budget`'s M3 table in [frame-budget.md](frame-budget.md) has been rerun and committed. ### Paid off A `DetailPass` now declares `output_scale`, and clarity's base is computed on a grid a quarter the size on each axis. Measured before and after on the same machine, same adapter, same build profile, with only the change between them — see [frame-budget.md](frame-budget.md) §"The reduced base, measured": | viewport, fit | before p50 | after p50 | | before p99 | after p99 | |---|---:|---:|---:|---:|---:| | 1920 × 1200 | 2.30 ms | 1.56 ms | 1.5× | 3.94 ms | 1.97 ms | | 2560 × 1600 | 4.41 ms | 1.95 ms | 2.3× | 10.40 ms | 2.37 ms | | 3840 × 2160 | **10.94 ms** | **3.88 ms** | **2.8×** | 25.05 ms | 4.17 ms | **Quote the p50 column.** The baseline run's p99 figures are contaminated — its `fit` rows spread 2.3× between median and 99th percentile where the after run spreads 1.1×, and [frame-budget.md](frame-budget.md)'s own independent measurement of the same baseline on the same card reports 4.67 ms p99 at 2560 × 1600 against the 10.40 ms here. The p99 improvement is real and larger than 2.8×; this run cannot say by how much. Clarity is no longer the stage that misses the budget, and no longer dominates the neighbourhood stage: at 4K it is 4.17 ms against 4.61 ms for all four neighbourhood operations together, where it was 97% of that total at every size. That comparison is within one run, so the contention does not touch it. **Two things worth recording, because neither is visible in the table.** The declared halo is now quantised to multiples of `output_scale`. The kernel truncates at 2σ and that rounding now happens on the reduced grid, so 1920 × 1200 reports 28 render pixels where it reported 29, and 2560 × 1600 reports 40 where it reported 38. At 2σ the Gaussian is already down to `e⁻²` of its peak, and the cross-form test holds the difference to 0.03 stops of peak excursion and 2% of frame reach — but it is a real change in reach, not a pure speed-up, and a tile scheduler would see it. The measurement was taken on an AMD RX 5700 XT, not the RTX 3050 the M1/M2/M3 tables above were measured on, so the *absolute* figures are not comparable with those. The before/after is, because both halves of it were measured on the same card minutes apart. --- ## TD-5 — The fused shader is reassembled from strings on every frame **Where:** `dr_pipeline::operation::compose_full`, called from `DevelopSession::render`. ### What it does `EditGraph::compose` walks the active operations and formats a WGSL source string, per frame, on the UI thread. On a full chain that is **2.8–5.2 ms** — at 1920 × 1200 it is larger than the entire fused dispatch it precedes, and on the chains that also carry a detail stage it is a third of what is left of the 16 ms budget after the GPU has taken its share. It does not vary with resolution, because it is not pixel work. ### Why Because it was free until the chain got long. Composition was written when an edit was two or three operations, and the cost is roughly linear in generated source: the colour mixer emits twelve hue bands and the tone curve emits a spline evaluator, so a full chain is a large string built from scratch sixty times a second. ### Paying it off The generated source depends only on the *structure* of the graph — which is precisely what `ComposedShader::structure_hash` already identifies, and precisely what does not change while a slider is being dragged. `AdjustPass` relies on that already: it caches compiled pipelines against that hash and does not recompile during a drag. Caching the source string against the same hash and rebuilding only the uniforms — a handful of floats per operation — takes this to approximately nothing on the path that needs it most. The care needed is in what the hash covers. It deliberately excludes parameter *magnitudes*, so a cache keyed on it is sound for the source and would be wrong for anything else in `ComposedShader`. **Done when:** the `shader` column of [frame-budget.md](frame-budget.md)'s M1 table is under a millisecond for the `point` and `all` chains, and the codegen tests still pass byte for byte. --- ## TD-6 — The quietest ink does not reach WCAG AA, and the rule does not reach 3:1 **Breaks:** [requirements.md](requirements.md) NFR-A11Y-2 — "non-canvas UI meets WCAG AA contrast". ### What it does `style.yaml` sets three inks and four surfaces. Measured as WCAG 2 contrast ratios (sRGB relative luminance, the standard formula), against the surfaces each ink is actually drawn on: | ink | on `ground` | on `surface` | on `surface-raised` | on `hover` | on `selected` | |---|---|---|---|---|---| | `ink` #EDEEF0 | 16.02 | 14.69 | 13.03 | 11.39 | 9.68 | | `ink-dim` #9EA1A6 | 7.18 | 6.58 | 5.84 | 5.10 | **4.34** | | `ink-faint` #71747A | **3.97** | **3.64** | **3.23** | **2.82** | **2.40** | | `warn-ink` #C9A05A | 7.67 | 7.03 | 6.24 | 5.45 | 4.64 | | `rule` #323438 | **1.49** | **1.37** | **1.21** | **1.06** | **1.11** | Every text size in the application is 11px, 13px, 17px or 24px, and WCAG's "large text" relief begins at 18.66px bold or 24px regular — so all four of those thresholds are the normal-text one, **4.5:1**, except the masthead. Bold entries fail it. The inverted cases pass and are worth stating so nobody re-measures them: `ground` on `active` (#FFFFFF) is 18.60, on `active-dim` 11.20, on `active-pressed` 5.95, on `selected-ring` 13.02. The near-white fills that `Button.primary`, `FilterChip.active` and the held tool-rail entry use are the *best*-contrasting text in the interface, not the worst. So the failures are exactly two, and neither is where one would guess: - **`ink-faint` reaches 4.5:1 nowhere at all.** It is the ink for `Caption`, `PanelHeading`, `Disclosure`, `Value`'s placeholder state and `FilterChip`'s count — every hint, every section name, every "3 photographs" under a title. - **`rule` reaches 3:1 nowhere.** WCAG 1.4.11 asks 3:1 of the boundary of a control the user must perceive, and `rule` is the border of every `Button`, `Field`, `Panel`, `ChoiceChip` and `IconButton`. An unfilled secondary button is a 1.4:1 outline on a 1.2:1 background. `ink-dim` on `selected` at 4.34 is a third case, marginal enough that a two-point lift fixes it. ### Why Not an oversight — the direct consequence of the palette's own argument, which `style.yaml`'s preamble makes at length and correctly. The chrome is deliberately quiet because a bright surround biases how a photograph is judged, and hue is banned outright because an accent beside the image shifts the perception of nearby colours. What is left to signal with is luminance, and the palette spends its luminance range on the *photograph*, keeping the chrome inside a narrow band above the ground. A narrow band is precisely what a contrast ratio measures. `ink-faint` exists to be skipped by the reader who did not stop to look; that is a real design intent, and "text you are meant to skip" and "text everyone can read" are in genuine tension rather than one being a mistake. ### What it costs The photographer who cannot read a caption cannot read *any* caption, on any screen — this is one token, so it fails everywhere at once. The hints under the settings switches say what a setting costs, the section names say what a panel is, and the counts say how big a filter is. None of it is decorative. ### Paying it off Two token changes, and the second is the awkward one. `ink-faint` needs roughly #8A8D93 to clear 4.5:1 against `surface-raised`, the darkest surface it is drawn on that matters — which puts it about where `ink-dim` sits today and collapses the three-ink scale to two. So the real fix is to re-derive all three inks against the surfaces rather than to nudge one: the scale wants to start higher and keep its steps, not compress. `rule` needs about #4A4D52 for 3:1 against `surface`. That is a visibly stronger line, and the preamble's "instrument rather than absence" reasoning applies to it as much as to the greys — this is a look change, not a number change, and it should be looked at rather than computed. Both are decisions about how the application appears next to a photograph, which is the one thing this palette was designed around. They want a screenshot and an opinion, not a patch. **Done when:** every `Theme` ink reaches 4.5:1 against every surface it is drawn on, `rule` reaches 3:1 against `surface` and `surface-raised`, and a test recomputes those ratios from `style.yaml` so the next palette edit cannot quietly undo it. The table above is the baseline to compare against. ### Not in scope The histogram's `plot-*` inks (2.36 for `plot-luma` on `ground`) are drawn *on* the canvas, and NFR-A11Y-2 scopes contrast to non-canvas UI. NFR-A11Y-3 covers what those need instead, and is already met — the readouts name the channel in words. --- ## TD-7 — Platform font scaling is not honoured **Breaks:** [requirements.md](requirements.md) NFR-A11Y-2 — "platform font scaling is honoured without clipping". ### What it does Nothing at all, which is the entry. Every type size is a constant in `style.yaml` — 11, 13, 17, 24 — read as `Theme.text-sm` and friends at 65 call sites, and there is no multiplier anywhere between the platform's font-size preference and those numbers. `scale_factor()` is read in `display_ui.rs` and in `lib.rs`, but only to size the canvas in physical pixels for the render; it is display DPI, which Slint already applies to logical lengths, and it is not the user's text-size setting. A photographer who sets 130% text on GNOME or Android gets an application that ignores it. ### Why Because honouring it is not a multiplier, and pretending it is would be worse than not doing it. The layout is built on constants that are not derived from the type size: `control-height` 28, `touch-target` 44, `row-height` 26, `rail-entry-height` 54, `panel-width` 360, and a dozen fixed heights written at their call sites — `ParamSlider`'s 46px, the folder picker's 220px box, the readout column's 30px. Scaling the type alone clips against every one of them, silently, because Slint elides rather than errors. `SwatchSlider` is the sharpest case: 12 hue bands × 3 channels in a 360px column, sized so that a track and a swatch and a three-character readout fit on one line. And the failure is invisible to the person shipping it. The requirement's own phrase is "without clipping", and clipping is exactly what a screenshot at 100% cannot show — the memory note on verifying Slint changes exists because these files have a history of compiling, rendering and being wrong. ### What it costs The user for whom this matters most is not the screen-reader user the rest of this branch serves — it is the one with usable but poor sight, who reads the interface and needs it larger. The application is unusable to them at any setting, and there is no partial credit: text scaling is a system-wide preference, so an app that ignores it is the one thing on the desktop that did. ### Paying it off In the order the pieces depend on each other: 1. **A scale token.** `Theme.text-sm` and the rest become `base × Theme.type-scale`, with the scale an `in-out` property Rust writes at startup from the platform. `build.rs` already emits `in-out` tokens under `live-style`, so the codegen half of this exists and is proven — that feature is the mechanism, one line from being general. 2. **A source for the number.** GNOME publishes `text-scaling-factor` over the settings portal; Android has `Configuration.fontScale` through JNI, beside the calls `lib.rs` already makes for `ACTION_VIEW`. Both want a default of 1.0 and a sane clamp — 0.8 to 2.0 — because a user who has set 300% for a phone launcher has not asked for a 72px slider readout. 3. **The constants that are not type.** Every fixed height a *label* sits inside has to follow the scale; every touch target must not shrink and need not grow. That is the work, and it is where the 46px and 220px literals get read one at a time. 4. **Evidence.** Screenshots at 1.0, 1.3 and 2.0 of the develop column, the settings page and the colour mixer — the three densest layouts — because "without clipping" is a claim about the worst case and nothing else will show it. **Done when:** the develop column, the settings page and the colour mixer render at a 2.0 scale with no elided label and no touch target under 44 logical pixels. --- ## Related, and deliberately not here The window-move rule, the grid's ordering index and the whole-library readout cache were *fixed* rather than deferred — see the commits around `6d6ef8d`. They are mentioned only so that a reader looking for "why was the grid slow" finds the answer in the code and its comments rather than assuming it is still outstanding.