dtourolleandClaude Opus 5 ebb7d3cf5c Score a suggestion against the people the user has named
The number beside a suggestion was the mean calibrated probability
between the face and the rest of its group, which measures the wrong
thing twice. It punishes coverage: a person with two hundred faces over
fifteen years is *meant* to have members a given photograph is
orthogonal to, so a correct suggestion onto a well-photographed person
scored low for being well photographed. And it never asked who else the
face might be — a face matching Anna at 0.95 and nobody else, and one
matching Anna at 0.95 and her sister at 0.93, came out identical, when
the second is the only one worth the user's attention.

dr_face::assign answers both, and multiplies them: the mean of the best
ten calibrated matches into the identity (the old mean, capped, which is
what stops coverage counting against it), times that identity's share of
the evidence against every *named* rival.

Only named people compete, and per person rather than per group. Both
halves of that had to be measured on a real 18,000-face library rather
than reasoned about. Normalising across every group made the number
useless — median suggestion 21%, four in five under half — because
clustering leaves one person spread over many groups, so a face competed
against itself; and keying rivals by group left Catherine competing with
Catherine, median 39%. Per named person: median 99.5%.

Rivals are gathered below the merge threshold, down to even odds: a
named person matching at 0.6 will never be merged into but is exactly
the competition to discount for. That would be a second similarity scan,
the expensive half of regrouping a library, so cluster_scored scans once
at the looser floor and hands the merge engine the subset at or above
the threshold — pair for pair what it would have scanned for itself,
held to that by a test.

Leave-one-out over that library's 2,702 confirmations across 54 named
people: 99.33% of faces placed on the right person against the old
mean's 99.15%, and the number shown for the right person moves from a
median of 90.4% to 99.3%. It errs low — 100% correct wherever it states
80% or more — which is the safe direction, and docs/faces.md §9.1 says
plainly that the low bands are not calibrated.

The example that measures it comes too: this is a claim about a
library's numbers, and nobody should have to take it on faith.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-29 10:44:01 +02:00
2026-08-26 10:08:51 +02:00
2026-08-28 22:24:58 +02:00
2026-08-28 22:24:58 +02:00
2026-08-27 19:33:37 +02:00

DarkRoom

A cross-platform, non-destructive RAW photo editor for Linux and Android.

Status: early. v0.1 is a remote library viewer — see docs/milestone-v0.1.md.

Documentation

Document Contents
requirements.md What the software must do — 122 numbered requirements
architecture.md How it is built — crates, GPU pipeline, data model, sync
milestone-v0.1.md The first buildable milestone
faces.md Face detection and identity — the models, the licence problem, and what S14 measures

Building

Desktop:

cargo run -p darkroom-desktop

Android (containerised toolchain, see docker/android):

./docker/android/build.sh cargo ndk -t arm64-v8a build --release

Current state

Working: workspace, GPU context and compute pass, adaptive Slint shell, Android cross-compilation of the core crates.

Not yet working: the zero-copy display path. The build currently uploads frames through the CPU, which is exactly what ARCH §6.1 forbids — measured at 96% of frame time at 4K. Replacing it is spike S1, the project's highest priority.

cargo run -p dr-gpu --example bench --features readback

reproduces that measurement.

Licence

GPL-3.0-or-later.

S
Description
No description provided
Readme GPL-3.0
1 GiB
2026-10-07 11:27:59 +00:00
Languages
Rust 86.1%
Slint 10.3%
Python 1.1%
Shell 1%
WGSL 0.9%
Other 0.6%