dtourolleandClaude Opus 5 e596eb0657 Give the similarity scan the machine's SIMD, and its cache
The scan is O(n²) dot products and nothing else, so its speed is the
face subsystem's speed — and it was running at 0.7 flops per cycle.

Two separate faults, both measured over the reference 18,143-face
library on twenty cores. It walked the whole embedding array once per
row, ~336 GB of traffic, where a column tile that fits in L2 is read
once per tile of rows: 4.64s → 2.81s. And the workspace builds for
baseline x86-64 — SSE2, no FMA — into which the portable loop was not
being vectorised at all: 2.81s → 0.86s, 195 GFLOP/s.

So the dot product is now chosen per machine. AVX2 + FMA where
is_x86_feature_detected! finds it; NEON unconditionally on aarch64,
since Advanced SIMD is in that baseline and every Android device the app
builds for has it — with the explicit vfmaq, because LLVM will not fuse
a multiply and an add without being told to. The portable loop stays as
the definition the others are tested against, and
the_fastest_kernel_agrees_with_the_portable_one is the only check the
NEON path gets on a machine that is not aarch64.

Faces::embeddings is one flat buffer rather than a Vec per face: the
pointer chase defeated both the prefetcher and the tiling, and it is
also the layout a GPU pass would want.

Behaviour is unchanged and that is checked rather than asserted — the
same 1,531,969 pairs from all three kernels, and on the real library the
same 2,518 groups holding the same 16,246 faces with the same confidence
distribution. A full regroup there goes from 10.0s to 5.9s; the rest is
the agglomeration, which is a sequential heap walk and is where the next
look should go.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 11:33:15 +02:00
2026-08-26 10:08:51 +02:00
2026-08-28 22:24:58 +02:00
2026-08-28 22:24:58 +02:00
2026-08-27 19:33:37 +02:00

DarkRoom

A cross-platform, non-destructive RAW photo editor for Linux and Android.

Status: early. v0.1 is a remote library viewer — see docs/milestone-v0.1.md.

Documentation

Document Contents
requirements.md What the software must do — 122 numbered requirements
architecture.md How it is built — crates, GPU pipeline, data model, sync
milestone-v0.1.md The first buildable milestone
faces.md Face detection and identity — the models, the licence problem, and what S14 measures

Building

Desktop:

cargo run -p darkroom-desktop

Android (containerised toolchain, see docker/android):

./docker/android/build.sh cargo ndk -t arm64-v8a build --release

Current state

Working: workspace, GPU context and compute pass, adaptive Slint shell, Android cross-compilation of the core crates.

Not yet working: the zero-copy display path. The build currently uploads frames through the CPU, which is exactly what ARCH §6.1 forbids — measured at 96% of frame time at 4K. Replacing it is spike S1, the project's highest priority.

cargo run -p dr-gpu --example bench --features readback

reproduces that measurement.

Licence

GPL-3.0-or-later.

S
Description
No description provided
Readme GPL-3.0
1 GiB
2026-10-07 11:27:59 +00:00
Languages
Rust 86.1%
Slint 10.3%
Python 1.1%
Shell 1%
WGSL 0.9%
Other 0.6%