d8b9b5a4bb6b2bbd667ff217d32795c77c1c55f8
3
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
e8e96eed40 |
Measure the performance targets §8 has been promising, and fail on a regression
docs/requirements.md §8 has said since it was written that performance is verified by "an automated benchmark suite against a synthetic 50k catalog, run per-commit … A regression beyond stated tolerance fails the build." There was none. No benches/, no [[bench]], no criterion, no synthetic catalog, and three CI workflows that between them measured nothing. Ten performance requirements could therefore be neither passed nor failed, and five of them carried a TRACES: tag regardless. tools/bench is the half of that promise that can be kept honestly on a runner with no GPU and no display. # The fixture Rows are cheap and pixels are not, so it builds fifty thousand catalog rows over a pool of a dozen real files, each referenced by several thousand of them. Everything the catalog half touches is rows and is exact at full scale; everything the pixel half touches is one file at a time and does not care how many rows point at it. Fourteen megabytes on disk instead of two terabytes, and neither half is flattered by the trade. It is reproducible from a seed, and a stamp beside it — seed, row count, source size, dr-catalog's schema version — rebuilds it rather than letting a run be compared against a baseline that describes a different library. # What it can now pass or fail NFR-P1, and R2's second sentence with it: Catalog::open plus the count, first window and timeline the grid cannot paint without. The interesting part turned out to be the open itself — schema::backfill runs three passes over the images table on every open, which is O(library) work on a path whose budget is stated in absolute seconds. Tagged TRACES: NFR-P1, on a gate that fails if it breaks. NFR-P3: thumbnail throughput on the embedded preview path, through the same per-image work spawn_thumbnail_sweep does and in the same shape — chunks of 96, lanes owning disjoint slices, the single thread that owns the store writing the finished chunk. Mirrored rather than called, because that function takes a RemoteBackend and would measure somebody's network. Tagged TRACES: NFR-P3. # What it deliberately does not claim NFR-P7 is the whole chain, and only the encode half of it runs without an adapter. So the export row is a one-sided gate — over two seconds in the encode alone violates the requirement; under it proves nothing — and there is no TRACES: NFR-P7 anywhere. NFR-P8 is about the application at idle, and the probe is a process holding the catalog and nothing else, so it records the catalog layer's share and carries no budget until somebody decides what that share should be. No tag there either. CONTRIBUTING.md asks that a requirement be closed by a test that would fail if the behaviour were removed, and two more plumbing tags is what this repository already has too many of. NFR-P8 also gets the answer §4.1 demands: RSS is exclusive of device-local GPU allocations and cannot be made otherwise, because such an allocation never enters the process's address space. The requirement should be restated as two figures, and docs/benchmarks.md says so. # Two gates, and why one of them steps aside off the reference desktop The budget is the requirement's own number and never moves. The baseline is what the reference desktop last measured, and drifting 15% past it fails the build even while still inside the budget — which is how performance rot actually arrives, never over the line, always a little worse. A budget written for twenty-four threads cannot be asserted on a two-core container. §8 names the reference desktop, not CI, so each metric declares whether its budget is machine-sensitive; those are asserted under --reference and reported everywhere else. Catalog open is not one of them: two seconds against an expected figure two orders of magnitude smaller is a threshold any machine can be held to. This is the trap core/dr-gpu/tests/frame_budget.rs already refuses — a red gate everybody learns to ignore. # The baseline ships with no numbers in it Every recorded field is null, because nobody has run it yet. Writing plausible-looking figures would make every later comparison a comparison against a guess, and the first real regression would be invisible. Run `dr-bench record --reference` on the reference desktop and commit the diff; until then the budget gate works and the report says the other one cannot. # CI .gitea/workflows/benchmark.yml, and its own workflow rather than a step in build-and-test.yml: a red "Build and test" says the code is wrong, a red "Benchmarks" says it got slower, and the second must not be reachable by retrying a flaky compile. The cpu job runs on every push and builds -p dr-bench alone — which is why that crate depends on no GPU and no UI crate. The gpu job is the frame budget that already exists and already skips without an adapter, on workflow_dispatch, because building wgpu on every commit to rediscover that the runner has no device is not a use of anybody's minutes. |
||
|
|
95847e3a31 |
Say which channels v1 ships through, and that the Flatpak cannot reach a library
NFR-COMPAT-2 asks for the v1 channels to be stated, and says why in its own second sentence: the channel decision and the storage design are coupled. There was nowhere that statement lived. packaging/ held a PKGBUILD and a desktop entry, which is a recipe rather than a decision, and the coupling the requirement points at was therefore invisible. docs/distribution.md states five channels and, more usefully, which two of them exist only as promises. It also records what every channel has to get right independently of format — the one identifier that appears in four places, the metainfo, Vulkan being a requirement rather than a preference while NFR-R8 is open, a secrets daemon being optional rather than required, and the LFS pointer check that stops a package shipping 130 bytes where an 11 MB model should be. §4 is the part worth reading. Preparing a Flatpak is what surfaced that FR-PLAT-LIN-3 is not satisfied and cannot be satisfied by packaging alone: a folder library is chosen by typing an absolute path into an EndpointOnly field that checks it with std::fs, and nothing in the tree calls the FileChooser portal. Inside a sandbox that path does not exist, so the launch screen refuses it. Import fails one step earlier, because a sandboxed process reads its own mount namespace and a card mounted on the host is not in it. That is written down rather than fixed with --filesystem=host, and the argument for not fixing it that way is §2: the Arch package and an AppImage both hand the application the same unrestricted process the developer runs it in, so Flatpak is the only Linux channel that tests whether a design assumed unrestricted access. Granting the permission removes the only reason to ship it. The reverse coupling on Android is recorded too. NFR-COMPAT-2 says Play distribution is what makes ARCH §6.9 binding; §6.9 is verified rather than assumed, so SAF is already unconditional and a sideloaded build would gain nothing by asking for more. Play is deferred over the GPLv3 question, which is a licence-reading exercise and blocks nothing in the storage design. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
617262b4da |
Give a contributor a way in
There was none. 14 documents, 177 numbered requirements, and all of it written for someone who has already decided to work on this — nothing that tells a newcomer which door is unlocked, what a first build needs, or why it takes so long. CONTRIBUTING.md points the first door at the operation format, because "add a develop node" is a genuinely one-file contribution and the best first experience this codebase can offer: no Rust, no shader edit, no UI change, and its tests declared in the same file. It teaches the architecture's central idea on the way through, which is why the invariant test added in the previous commit is named there rather than left to be discovered. Three things that were folklore are now written down: Git LFS is a prerequisite, the first build resolves 826 crates and is not hanging, and Slint needs pkg-config, libfontconfig1-dev and libxkbcommon-dev. The LFS one proved itself while writing this — a fresh worktree hit exactly the failure `dr-segment`'s build script is written to catch, which is the argument for saying so before it happens rather than after. rust-toolchain.toml pins 1.92.0 because `build-and-test.yml` already does and says why: a floating toolchain turns an unrelated push into a mystery failure. The two checks that gate every push are the two most sensitive to compiler version — rustfmt's output changes between releases, so a contributor on a newer stable can produce a diff nobody wrote on a line nobody touched, and `clippy -D warnings` is the same story with new lints. `rust-version = "1.92"` in the manifest stays where it is; it is a minimum, and this is the upper bound it cannot express. Also states the convention the tooling cannot enforce, from code-health.md CH-4: close a requirement with a test that would fail if the behaviour were removed. Coverage that moves slowly and means something beats coverage that moves quickly. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |