Files
DarkRoom/Cargo.toml
T
dtourolle e8e96eed40 Measure the performance targets §8 has been promising, and fail on a regression
docs/requirements.md §8 has said since it was written that performance is
verified by "an automated benchmark suite against a synthetic 50k catalog, run
per-commit … A regression beyond stated tolerance fails the build." There was
none. No benches/, no [[bench]], no criterion, no synthetic catalog, and three
CI workflows that between them measured nothing. Ten performance requirements
could therefore be neither passed nor failed, and five of them carried a
TRACES: tag regardless.

tools/bench is the half of that promise that can be kept honestly on a runner
with no GPU and no display.

# The fixture

Rows are cheap and pixels are not, so it builds fifty thousand catalog rows
over a pool of a dozen real files, each referenced by several thousand of them.
Everything the catalog half touches is rows and is exact at full scale;
everything the pixel half touches is one file at a time and does not care how
many rows point at it. Fourteen megabytes on disk instead of two terabytes, and
neither half is flattered by the trade. It is reproducible from a seed, and a
stamp beside it — seed, row count, source size, dr-catalog's schema version —
rebuilds it rather than letting a run be compared against a baseline that
describes a different library.

# What it can now pass or fail

NFR-P1, and R2's second sentence with it: Catalog::open plus the count, first
window and timeline the grid cannot paint without. The interesting part turned
out to be the open itself — schema::backfill runs three passes over the images
table on every open, which is O(library) work on a path whose budget is stated
in absolute seconds. Tagged TRACES: NFR-P1, on a gate that fails if it breaks.

NFR-P3: thumbnail throughput on the embedded preview path, through the same
per-image work spawn_thumbnail_sweep does and in the same shape — chunks of 96,
lanes owning disjoint slices, the single thread that owns the store writing the
finished chunk. Mirrored rather than called, because that function takes a
RemoteBackend and would measure somebody's network. Tagged TRACES: NFR-P3.

# What it deliberately does not claim

NFR-P7 is the whole chain, and only the encode half of it runs without an
adapter. So the export row is a one-sided gate — over two seconds in the encode
alone violates the requirement; under it proves nothing — and there is no
TRACES: NFR-P7 anywhere. NFR-P8 is about the application at idle, and the probe
is a process holding the catalog and nothing else, so it records the catalog
layer's share and carries no budget until somebody decides what that share
should be. No tag there either. CONTRIBUTING.md asks that a requirement be
closed by a test that would fail if the behaviour were removed, and two more
plumbing tags is what this repository already has too many of.

NFR-P8 also gets the answer §4.1 demands: RSS is exclusive of device-local GPU
allocations and cannot be made otherwise, because such an allocation never
enters the process's address space. The requirement should be restated as two
figures, and docs/benchmarks.md says so.

# Two gates, and why one of them steps aside off the reference desktop

The budget is the requirement's own number and never moves. The baseline is
what the reference desktop last measured, and drifting 15% past it fails the
build even while still inside the budget — which is how performance rot
actually arrives, never over the line, always a little worse.

A budget written for twenty-four threads cannot be asserted on a two-core
container. §8 names the reference desktop, not CI, so each metric declares
whether its budget is machine-sensitive; those are asserted under --reference
and reported everywhere else. Catalog open is not one of them: two seconds
against an expected figure two orders of magnitude smaller is a threshold any
machine can be held to. This is the trap core/dr-gpu/tests/frame_budget.rs
already refuses — a red gate everybody learns to ignore.

# The baseline ships with no numbers in it

Every recorded field is null, because nobody has run it yet. Writing
plausible-looking figures would make every later comparison a comparison
against a guess, and the first real regression would be invisible. Run
`dr-bench record --reference` on the reference desktop and commit the diff;
until then the budget gate works and the report says the other one cannot.

# CI

.gitea/workflows/benchmark.yml, and its own workflow rather than a step in
build-and-test.yml: a red "Build and test" says the code is wrong, a red
"Benchmarks" says it got slower, and the second must not be reachable by
retrying a flaky compile. The cpu job runs on every push and builds -p dr-bench
alone — which is why that crate depends on no GPU and no UI crate. The gpu job
is the frame budget that already exists and already skips without an adapter,
on workflow_dispatch, because building wgpu on every commit to rediscover that
the runner has no device is not a use of anybody's minutes.
2026-08-30 10:40:10 +02:00

244 lines
11 KiB
TOML
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
[workspace]
resolver = "2"
members = [
"core/dr-types",
"core/dr-catalog",
"core/dr-thumbs",
"core/dr-decode",
"core/dr-export",
"core/dr-face",
"core/dr-film",
"core/dr-ingest",
"core/dr-gpu",
"core/dr-lens",
"core/dr-pipeline",
"core/dr-preset-xmp",
"core/dr-segment",
"core/dr-sync",
"core/dr-sync-folder",
"core/dr-sync-nextcloud",
"platform/dr-plat",
"ui/dr-ui",
"apps/darkroom-desktop",
"apps/darkroom-android",
"tools/bench",
"tools/traceability",
]
[workspace.package]
version = "0.9.0"
edition = "2021"
rust-version = "1.92"
license = "GPL-3.0-or-later"
repository = "https://github.com/dtourolle/DarkRoom"
[workspace.dependencies]
# Internal
dr-types = { path = "core/dr-types" }
dr-catalog = { path = "core/dr-catalog" }
dr-thumbs = { path = "core/dr-thumbs" }
dr-decode = { path = "core/dr-decode" }
dr-export = { path = "core/dr-export" }
# Stated explicitly for the same reason as `dr-segment` below: no dependant
# should drag in an ONNX runtime by accident. Members opt in with
# `features = ["inference"]`.
dr-face = { path = "core/dr-face", default-features = false }
dr-film = { path = "core/dr-film" }
dr-ingest = { path = "core/dr-ingest" }
dr-gpu = { path = "core/dr-gpu" }
dr-lens = { path = "core/dr-lens" }
dr-pipeline = { path = "core/dr-pipeline" }
dr-preset-xmp = { path = "core/dr-preset-xmp" }
# `default-features = false` belongs *here*, not on each dependant: a member
# inheriting a workspace dependency cannot turn its default features off, so
# writing it below would silently do nothing and every crate touching
# `dr-segment` would drag in tract and 11 MB of weights. Members opt in with
# `features = ["semantic", "embedded-model"]` instead.
dr-segment = { path = "core/dr-segment", default-features = false }
dr-plat = { path = "platform/dr-plat" }
dr-sync = { path = "core/dr-sync" }
dr-sync-folder = { path = "core/dr-sync-folder" }
dr-sync-nextcloud = { path = "core/dr-sync-nextcloud" }
dr-ui = { path = "ui/dr-ui" }
# GPU + UI
#
# The wgpu version is not a free choice: it is dictated by Slint. Importing a
# texture into the scene (ARCH §6.1, spike S1) requires it to come from the
# *same* `wgpu::Device` Slint renders with, and Slint will only hand out a
# device of the version it was compiled against. Slint 1.17 offers
# `unstable-wgpu-28` and `unstable-wgpu-29` and nothing older, so 29 it is —
# pinned to the same `29.0.4` floor Slint itself requires, because two
# semver-compatible-but-different wgpu crates in one tree are two *types*, and
# the device would not typecheck across them.
#
# Consequently: bumping Slint may force a wgpu bump, and wgpu cannot be bumped
# on its own. They move together or not at all.
wgpu = "29.0.4"
slint = { version = "1.17", default-features = false }
slint-build = "1.17"
# UI token codegen (S2): style.yaml -> theme.slint. serde_yaml was deprecated
# by its maintainer in 2024 and serde_yml, the first fork, has since been
# deprecated too; serde_norway is the fork still receiving releases. Its
# mappings preserve insertion order, which is what lets the generated Slint
# keep the token ordering the YAML author chose.
serde_norway = "0.9"
# Foundations
anyhow = "1"
thiserror = "2"
log = "0.4"
env_logger = "0.11"
pollster = "0.4"
# Networking — no mature Nextcloud crate exists; the connector is hand-rolled
# over reqwest (D7). reqwest_dav was evaluated and is too thin to build on.
# `rustls-no-provider` rather than `rustls`: the latter defaults to the
# aws-lc-rs crypto provider, whose aws-lc-sys crate is C and fails to
# cross-compile for Android — precisely the NDK pain D1 chose Rust to avoid.
# ring is pure Rust apart from a small asm core that does build under the NDK.
#
# `rustls-tls-webpki-roots-no-provider` rather than plain `rustls-no-provider`:
# the latter verifies against rustls-platform-verifier, which reaches the
# Android trust store over JNI and panics mid-handshake unless initialised from
# Java first — the crash D7 predicted and spike S3 exists to resolve properly.
# The panic surfaces inside tokio, which catches task panics itself, so it
# reaches the UI as a worker that stopped rather than as an error.
#
# webpki-roots is the escape hatch D7 records: a root store compiled into the
# binary, no JNI, identical on both platforms. The trade is real and belongs in
# S3's scope — user-installed and enterprise CAs are not consulted, and the
# roots go stale with the release rather than with the OS.
reqwest = { version = "0.13", default-features = false, features = ["rustls-no-provider", "webpki-roots", "stream", "json"] }
rustls = { version = "0.23", default-features = false, features = ["ring", "std", "tls12"] }
quick-xml = "0.41"
tokio = { version = "1", features = ["rt-multi-thread", "macros", "sync", "time"] }
url = "2.5"
async-trait = "0.1"
serde = { version = "1", features = ["derive"] }
serde_json = "1"
base64 = "0.23"
# Display-server clients, for FR-DSP-8's per-display profile acquisition.
#
# Neither is a new cost: winit already builds both, so the versions are the
# ones Slint's backend has resolved to and pinning anything else here would
# compile a second copy. Both are pure Rust — x11rb speaks the X11 wire
# protocol itself rather than binding libxcb, and wayland-client binds
# libwayland only under a feature that is off — which keeps the Android
# cross-compile a plain Rust dependency graph, the same criterion as the TLS
# and SQLite choices above. They are declared under a target predicate that
# excludes Android, where neither display server exists.
#
# `staging` on wayland-protocols is what carries `wp_color_manager_v1`: the
# colour-management extension is still staging upstream, which is the
# protocol-level statement of the thing FR-DSP-8 anticipates when it says
# Wayland's colour management "is not universally available".
x11rb = { version = "0.13", features = ["randr"] }
wayland-client = "0.31"
wayland-protocols = { version = "0.32", features = ["client", "staging"] }
# Platform secure storage: Secret Service on Linux, Keystore on Android
# (FR-NC-2). Credentials never touch the catalog or a plain file.
# keyring 4 restructured its features: `v1` is the default set and brings
# the zbus Secret Service backend, which is what GNOME Keyring and KWallet
# (via ksecretd) both speak.
keyring = { version = "4", features = ["v1"] }
# The Android half of the same project: a keyring-core CredentialStore backed
# by AndroidKeyStore AES-GCM over SharedPreferences (FR-PLAT-AND-1). It reads
# the JavaVM and Context from ndk-context, which android-activity populates
# before `android_main` runs, so no Kotlin shim of our own is needed.
#
# This is the keyring-core API, not the v1 `Entry` API the Linux path uses;
# the two impls are deliberately separate rather than sharing a code path.
android-native-keyring-store = "1.0.0"
keyring-core = "1"
# Decode. rawler is the pure-Rust decoder (D2); zune-jpeg decodes the
# embedded previews rawler extracts.
# Catalog. `bundled` compiles SQLite from source rather than linking the
# system library — the same cross-compilation reasoning as the TLS choice
# above: no system dependency to satisfy under the Android NDK.
#
# `backup` is not optional in practice: it is what takes a consistent snapshot
# of a live WAL database for upload. A filesystem copy of `catalog.sqlite`
# while a `-wal` exists beside it uploads a torn file.
rusqlite = { version = "0.40", features = ["bundled", "backup"] }
rawler = "0.7"
zune-jpeg = "0.4.21"
# Thumbnails are stored encoded, not as raw RGBA: a 256px RGBA buffer is
# ~256 KB against ~20 KB as JPEG, and the store syncs to Nextcloud where that
# 13× is transfer cost on every client. Pure Rust, no C dependency — the same
# criterion behind the TLS and SQLite choices above.
jpeg-encoder = "0.7"
bytemuck = { version = "1", features = ["derive"] }
# Lens correction profiles. A pure-Rust port of Lensfun rather than a binding
# to the C library, for the same cross-compilation reason as the TLS and
# SQLite choices above: liblensfun would be a third C dependency to satisfy
# under the Android NDK.
#
# The database ships *inside* the crate — 56 XML files, gzipped at build time
# and decompressed on first lookup. That matters beyond convenience: Android
# gives us no filesystem path (ARCH §6.9), so a database loaded from a
# system directory would have nowhere to live there.
#
# Licence: LGPL-3.0-or-later, which upgrades cleanly into our GPLv3 (D8).
# The upstream Lensfun *database* is CC-BY-SA and is redistributed by the
# crate; attribution belongs in the about screen.
#
# Caveat worth remembering: this is a third-party port at 0.7.0, not upstream
# Lensfun. Verified working against the bundled database (interpolation
# between calibration points, and an unknown lens returning empty rather than
# panicking), but the pipeline talks to it through its own profile types so
# swapping it out is not a pipeline change.
lensfun = "0.7"
# Neural inference for semantic segmentation (S15 arm B, D14).
#
# D13 framed this as a choice between `ort` (fast, best operator coverage, and
# a C++ dependency to cross-compile under the NDK) and a pure-Rust runtime
# (policy-compliant, unproven coverage). That framing turned out to be a false
# choice: `ort` 2.0's `alternative-backend` feature *disables the linking
# entirely* and lets a different engine supply the `OrtApi`, and `ort-tract` —
# same authors, MIT/Apache — supplies it from `tract`, which is pure Rust.
#
# So we get `ort`'s API with no C at all. `download-binaries` and `tls-native`
# are off with `default-features = false`, which is the point: nothing is
# fetched at build time and nothing is linked, so the Android cross-compile
# sees an ordinary Rust dependency graph. That is the same reasoning as rustls
# over aws-lc-rs and bundled SQLite, applied to inference — D13's largest
# tolerated exception turns out not to be needed.
#
# The trade is real and belongs on the record: tract is slower than the C++
# runtime and covers fewer operators. Both were measured rather than assumed
# before this landed — yolo26n-seg loads with **zero unsupported operators**
# and runs 640x640 in ~470 ms on the reference desktop's CPU. That is fine for
# a once-per-image precompute off the frame path (ARCH §6.1) and would not be
# fine for anything per-frame, which is a constraint on what may be built on
# top rather than on this choice.
#
# Pinned to an rc: `ort` 2.0 has been in rc for a long while and `ort-tract`
# exists only against it. Worth revisiting at 2.0 final.
ort = { version = "2.0.0-rc.13", default-features = false, features = ["alternative-backend", "ndarray", "std"] }
ort-tract = "0.4"
# Not a free choice: it is the version `ort` exposes its tensors through, so
# two semver-incompatible ndarrays would not typecheck across the boundary —
# the same coupling wgpu has with Slint above.
ndarray = "0.17"
[profile.dev]
# Dependencies optimised even in dev builds — wgpu and image decoding are
# unusably slow otherwise, and they rarely need debugging.
opt-level = 0
[profile.dev.package."*"]
opt-level = 2
[profile.release]
lto = "thin"
codegen-units = 1