Require judgement anywhere, and evidence that never becomes a verdict
Rating and flagging were reachable from the grid alone, so a photograph opened in develop could not be judged without leaving it; FR-UI-5 said "rating" without qualifying the view and was built as though it had. And FR-UI-1's expanded row has said "filmstrip" since it was written while the roll stayed on demand in both classes. Both are amended to say what they meant: judgement follows the photograph, without auto-advance outside the culling mode, and the roll is open by default where there is room for it. The larger change is a rule. Per-face signals — eye state from a classifier, head pose from the five landmarks the detector already yields — are worth having for culling, and §3.9.1 excluded detecting a blink outright. The exclusion was always of judgement, not of knowing: a blink is a fact about a frame of the same kind as a clipped highlight. FR-CULL-8a specifies the two signals; FR-CULL-13 says what any signal may do (be shown, filtered, sorted, propose a burst representative) and what none may (write a rating or flag without a user action between). R7 states the same thing as a user need. Licensing was read before either was written. OCEC's eye-state weights are MIT with a clean data chain; every open gaze model is trained on Gaze360 or its peers, whose licences restrict derived models by name, so gaze is deferred in §7 and head pose stands in for it. D13 records both so they are not re-searched. Replacing a closed-eyed face from a neighbouring frame was raised and is written down as D17 rather than built: it is the multi-source schema question §7 already defers for panorama and HDR, with its non-goals — never automatic, provenance declared — fixed now. Traceability regenerated: three new IDs, none yet tagged.
This commit is contained in:
+142
-2
@@ -57,6 +57,15 @@ These are the user's stated requirements, restated as testable criteria.
|
||||
| **R4** | HW acceleration and parallelism | All per-pixel work runs on GPU compute. CPU work (decode, I/O) is parallelised across cores. The UI executor never blocks on image work (NFR-ARCH-1). |
|
||||
| **R5** | Work on downscaled proxies for display | Display pipeline operates at viewport resolution, not source resolution. The tiling clauses this criterion used to carry have been struck — see below. |
|
||||
| **R6** | Nextcloud integration | Browse, download, and upload images and edit metadata against a Nextcloud instance, offline-capable. |
|
||||
| **R7** | Judge anywhere, on evidence | Rating and flag are reachable from every view that shows a photograph, and apply to the one on screen. Everything the app computes about a frame in aid of culling — clipping, focus, burst membership, per-face state — is shown as evidence the photographer reads, and **no code path writes a rating or flag without a user action** (FR-CULL-13). |
|
||||
|
||||
**On R7, added 2026-09-19.** Stated from use rather than from the brief. Two things prompted it.
|
||||
Judgement keys had been built into the grid alone, so a photograph opened in develop — the view a
|
||||
photographer is most sure about — could not be rated without leaving it; the rule is that judging
|
||||
follows the photograph, not the view. And specifying per-face signals (FR-CULL-8a) forced the
|
||||
question of what a signal is *for*, which sharpened what this document already said in FR-CULL-5:
|
||||
the app may know a great deal about a frame and may say all of it, and it never holds the pen.
|
||||
R7 is the user-level statement; FR-CULL-13 is the testable one.
|
||||
|
||||
**On R1's tolerance.** An earlier draft required output to be *bit-identical* across platforms.
|
||||
That is not achievable and the requirement has been corrected. Floating-point compute results
|
||||
@@ -613,6 +622,14 @@ mode.
|
||||
> fires only on a desktop window dragged narrow. What portrait needs is the dock, which is the
|
||||
> aspect axis above and not a second class (D-N7).
|
||||
|
||||
> **Amended 2026-09-19.** The expanded row has said "filmstrip" since the table was written, and
|
||||
> the build had the photo roll on demand in both classes — the compact row applied everywhere. In
|
||||
> the expanded class the roll is **open by default** and closable, and it is shown for a set of
|
||||
> files named on the command line as much as for a library, because a set of photographs is a set.
|
||||
> It carries the place FR-UI-8 remembers — how many, which one, its name, and what the filter is
|
||||
> narrowing to — so that develop and the grid read as one interface with two views of the same
|
||||
> set, not two screens joined by a button.
|
||||
|
||||
**FR-UI-2 — Input modality.** The interface detects and adapts to the active input method, which
|
||||
is independent of layout class: a tablet may have a keyboard and pointer attached, and a desktop
|
||||
may have a touchscreen. Modality affects control sizing and affordances (FR-DEV-3b), not layout —
|
||||
@@ -630,6 +647,16 @@ functionality is touch-only.
|
||||
menus, and scroll-wheel adjustment on numeric controls. Keyboard shortcuts cover navigation,
|
||||
rating, and common adjustments. Neither is required for any operation to be reachable.
|
||||
|
||||
> **Amended 2026-09-19.** "Rating" above is not qualified by view and was built as though it were:
|
||||
> the judgement keys lived in the grid alone. They shall work in every view that shows a
|
||||
> photograph, applying to the **one on screen** — in develop, the open photograph, never a
|
||||
> selection left behind in the grid — with a pointer and touch equivalent in the same view, since
|
||||
> a tablet has no number row. Judging in develop does **not** advance to the next frame:
|
||||
> auto-advance belongs to FR-CULL-4's mode, where moving on is the point, and in develop the
|
||||
> photographer is working on the frame in front of them. The current rating and flag are shown
|
||||
> wherever they can be set, and on the roll's cells, so stepping along a set shows what has been
|
||||
> judged.
|
||||
|
||||
**FR-UI-6 — Shared component library.** Touch and desktop presentations are variants of shared
|
||||
components, not parallel implementations. A new operation (FR-DEV-3c) becomes usable on both
|
||||
without frontend work.
|
||||
@@ -1088,6 +1115,11 @@ mandatory.
|
||||
- **One-key reject**, plus the full rating, flag, and colour-label axes
|
||||
- **Filter to unjudged**, so a session resumes where it stopped
|
||||
- Keyboard-driven on desktop; single-thumb reachable on tablet
|
||||
- **Evidence beside the frame** (added 2026-09-19) — FR-CULL-13's signals, readable at a glance
|
||||
without leaving the mode or opening the frame
|
||||
- **The last import as a scope** (added 2026-09-19) — the set a culling session most often starts
|
||||
from, reachable in one step and counted, expressed as a term of the §5 selector language rather
|
||||
than as a special collection
|
||||
|
||||
**FR-CULL-5 — Burst and near-duplicate grouping.** Group frames by capture-time proximity and image
|
||||
similarity, allowing a burst to collapse to one representative and be judged as a unit.
|
||||
@@ -1106,6 +1138,31 @@ originals — a substantially smaller sync problem than develop parity.
|
||||
*Design note:* pinch-zoom accidentally triggering ratings is a documented defect in Lightroom
|
||||
mobile. Gesture and rating targets must not overlap.
|
||||
|
||||
**FR-CULL-13 — Evidence, never verdicts.** Added 2026-09-19; numbered after the people clauses
|
||||
because it governs them too. Everything the app computes about a frame in aid of culling is
|
||||
**evidence**: shown beside the frame, filterable and sortable through the §5 selector language,
|
||||
and — within a burst — permitted to *propose* which frame represents the group. The signals are
|
||||
the raw histogram, clipping and focus (FR-CULL-3), burst membership (FR-CULL-5), and per-face eye
|
||||
state and head pose (FR-CULL-8a). The list grows by adding to it, never by adding a second kind of
|
||||
thing.
|
||||
|
||||
What evidence may not do: change a rating, a flag, a colour label, or trash membership. There shall
|
||||
be **no code path from a signal to a judgement write** without a user action between them, and a
|
||||
proposal — a representative, "three of these have eyes closed" — is accepted by a press, never by a
|
||||
timeout or a default. This restates FR-CULL-5's ground rather than extending it: automated
|
||||
selection is distrusted because its documented failure is rejecting the only frame of a moment
|
||||
because someone blinked, and the remedy is not a better classifier, it is that the classifier does
|
||||
not hold the pen.
|
||||
|
||||
Evidence says what it measured. A focus figure names its region; a per-face count names the
|
||||
faces; a signal that could not be computed — no faces found, the original not on this device — is
|
||||
shown as absent, never as zero.
|
||||
|
||||
*Acceptance:* a test enumerates every write to the rating and flag axes and shows each reachable
|
||||
only from an input event. The evidence for a frame is visible in FR-CULL-4's mode and on its grid
|
||||
cell without opening it, and a filter on any one signal returns exactly the set whose chips show
|
||||
it.
|
||||
|
||||
### 3.9.1 People
|
||||
|
||||
Face recognition was deferred in §7 through the 2026-08-08 calibration. It is undeferred here in a
|
||||
@@ -1121,10 +1178,17 @@ frames with the bride in them" across a 4,000-image wedding is a culling operati
|
||||
the differentiator.
|
||||
|
||||
**What is deliberately not in scope**, because it is the failure FR-CULL-5 names: no automated
|
||||
*selection*. Nothing here rejects a frame, ranks a face, scores a smile, or detects a blink. The
|
||||
*selection*. Nothing here rejects a frame, ranks a face, scores a smile, ~~or detects a blink~~. The
|
||||
feature produces a **filter**, never a judgement. The user's rating axes remain the only thing that
|
||||
rejects a photograph.
|
||||
|
||||
> **Amended 2026-09-19.** "Detects a blink" is struck. The exclusion was always of *judgement*,
|
||||
> and a blink is a fact about a frame of the same kind as a clipped highlight: reporting it is
|
||||
> evidence, acting on it is the failure. FR-CULL-8a detects eye state and head pose; FR-CULL-13
|
||||
> says what may be done with them — shown, filtered, sorted, proposed — and what may not. The
|
||||
> sentence that survives is the one that matters: the user's rating axes remain the only thing
|
||||
> that rejects a photograph.
|
||||
|
||||
**FR-CULL-8 — Face detection.** The app shall detect faces in library images as a background job,
|
||||
producing per-face a bounding box, five-point landmarks, a detector confidence, and a 512-dimension
|
||||
embedding.
|
||||
@@ -1178,6 +1242,36 @@ interaction target at any point, and survives being killed and restarted with no
|
||||
beyond the in-flight image. No face is stored whose aligned crop was upsampled beyond a stated
|
||||
factor; the crop source resolution is recorded per face (`crop_px`) and is auditable.
|
||||
|
||||
**FR-CULL-8a — Per-face state.** Added 2026-09-19. For every detected face the app shall record,
|
||||
as derived data under FR-CULL-12:
|
||||
|
||||
- **Eye state** — one open-probability per eye, from a classifier over a crop around each eye
|
||||
landmark. Per face this reads as both open, one closed, or both closed; per frame, as a count.
|
||||
- **Head pose** — yaw, pitch and roll, solved from the five landmarks against a generic face
|
||||
template. No model: this is geometry the detector has already paid for. *Facing the camera* is
|
||||
the pose within a stated band, and is the proxy for eye contact this document adopts — because
|
||||
gaze estimation has no redistributable weights (D13), and because in the photographs where it
|
||||
matters, groups and children and events, a turned head is the decision and averted eyes are
|
||||
often the better frame.
|
||||
|
||||
Both are terms in the §5 selector language (FR-CULL-11) — "everyone's eyes open", "facing the
|
||||
camera" — composable with a person and with every other term. Both feed FR-CULL-13's evidence and
|
||||
may propose a burst representative (FR-CULL-5); neither may set a rating or a flag.
|
||||
|
||||
**The weights ship under the same test as every other model**: redistributable under a licence
|
||||
compatible with GPLv3 and with Flatpak, F-Droid and Play, with the training data's terms read as
|
||||
well as the weights' (D13). The candidate identified 2026-09-19 is OCEC — MIT for code and
|
||||
weights, data under ODC-By 1.0 and Apache 2.0, six variants from 112 KB to 6.4 MB, a 24×40 crop
|
||||
per eye, sub-millisecond on CPU, opset 17 with batch-norm already folded. Its published F1 of
|
||||
0.99 is on its own crops; ours are cut from a five-point landmark, so the number is measured on
|
||||
this library before it is believed.
|
||||
|
||||
*Acceptance:* on a labelled set of at least 500 faces from the reference library, eye state is
|
||||
within a stated tolerance of its published F1, reported separately for glasses, profile, and
|
||||
faces under 60 px; facing-the-camera agrees with a hand-labelled split at a stated rate. Both are
|
||||
recomputed by re-indexing, neither is written to a sidecar, and the indexing pass stays inside
|
||||
FR-CULL-8's acceptance.
|
||||
|
||||
**FR-CULL-9 — Calibrated identity.** Face similarity shall be expressed as a **calibrated
|
||||
probability that two faces are the same person**, not as a raw embedding distance. Every threshold
|
||||
in the subsystem — clustering, suggestion, auto-confirmation — shall be stated in that probability
|
||||
@@ -1834,7 +1928,7 @@ Settled by requirements calibration, 2026-08-08.
|
||||
| Ingest | Full workflow — template rename, checksum verify, dual-destination |
|
||||
| Colour defaults | Good, not obsessive — matrices plus per-body base curve |
|
||||
| Film simulation | Fujifilm explicitly targeted |
|
||||
| AI | Denoise in v1; masking deferred |
|
||||
| AI | Denoise in v1; masking deferred. Per-face eye state and head pose are in v1 **as culling evidence, not AI** (FR-CULL-8a, FR-CULL-13); gaze deferred (§7) |
|
||||
| Local adjustments | Full masking, GPU-rasterised |
|
||||
| Sync | The reason the project exists |
|
||||
| Durability | Sidecar-first |
|
||||
@@ -1901,6 +1995,15 @@ desktop window, not as a second interface.
|
||||
>
|
||||
> **The licensing half is untouched.** The InsightFace weights are still non-commercial and still
|
||||
> unusable here. That remains what S14 has to resolve first.
|
||||
>
|
||||
> **Updated 2026-09-19.** Two further models were read for FR-CULL-8a. **Eye state:** OCEC
|
||||
> (PINTO0309) is MIT for code and weights and trained on ODC-By 1.0 and Apache 2.0 data — the
|
||||
> first face-adjacent weights found with a clean chain end to end, and the smallest by two orders
|
||||
> of magnitude. **Gaze:** none. MobileGaze, L2CS-Net and their descendants carry MIT on the
|
||||
> repository and Gaze360 in the weights, and Gaze360's research licence restricts *"models trained
|
||||
> on dataset"* by name; MPIIGaze and ETH-XGaze are no better. Head pose needs no weights at all.
|
||||
> None of this moves the detector and embedder, which are still the buffalo grant and still what
|
||||
> S14 resolves first: clean eye weights behind an unshippable detector ship nothing.
|
||||
|
||||
§3.9.1 needs to run two neural networks locally. That collides with two settled positions, and
|
||||
neither collision is small enough to leave implicit.
|
||||
@@ -1952,6 +2055,41 @@ rather than thousands.
|
||||
|
||||
---
|
||||
|
||||
### D17 — cross-frame face repair · **OPEN**
|
||||
|
||||
Raised 2026-09-19: pair FR-CULL-8a's eye state with the burst, and let a closed-eyed face in the
|
||||
chosen frame be replaced by the same person's open-eyed face from a neighbouring frame. It inverts
|
||||
FR-CULL-5's fear — the tool rescues the only frame of the moment instead of rejecting it — and it
|
||||
collides with four things this document says.
|
||||
|
||||
1. **§1.3, "not a pixel editor."** Survivable by the FR-DEV-8 precedent: a repair is numbers in
|
||||
the graph and no pixels are stored. A face repair is a spot whose source is another frame,
|
||||
aligned by the similarity the face subsystem already fits between two landmark sets and blended
|
||||
by the membrane heal already shipped. It reopens `spot-removal.md`'s non-goal that the source is
|
||||
"a patch from the same photograph", which would be revised deliberately, not quietly.
|
||||
2. **§5.1, single-source `Image`.** The real one. §7 defers panorama, HDR and focus stacking on the
|
||||
grounds that the schema cannot express an image derived from several sources. This is that case
|
||||
in a milder form — the output is still frame A, but A's sidecar now names B — and it brings the
|
||||
rules the tiers already have: B's original must be present to render A (FR-NC-6c: visible and
|
||||
priced before starting), trashing B must know A depends on it (FR-CAT-15), and `Version::merge`
|
||||
has never seen a cross-reference. Reopening §5.1 for this reopens it for the three deferred
|
||||
rows at once, which is the argument for doing it once and properly rather than for this alone.
|
||||
3. **Tone.** The source patch goes through A's chain, not B's — B demosaiced and run through A's
|
||||
parameters to the head of the detail chain, then sampled. A second small pipeline at proxy
|
||||
resolution; a second full demosaic at export. The heal hides lighting drift between frames; it
|
||||
does not hide a turned head, and no blend does.
|
||||
4. **Never automatic.** `spot-removal.md`'s own rule: a false positive silently alters a
|
||||
photograph, which is the failure this application must not have. The tool may propose — same
|
||||
person, closed here, open two frames on, small alignment residual — and applies on a press.
|
||||
|
||||
And one thing the document does not say: **provenance.** The audience is RAW-literate (D11). A
|
||||
composite declares itself — the source frame in the sidecar and the history, and in export
|
||||
metadata. Nothing here specifies C2PA; that is a decision this one depends on.
|
||||
|
||||
Deferred under D12 until FR-CULL-8a exists and spot removal's disc has, in its own words, been
|
||||
finished and used. The first cut, when it comes, is "clone from a neighbouring frame" as a spot
|
||||
source; the face-aware proposal is a layer over that.
|
||||
|
||||
### D16 — plugin licensing · **OPEN**
|
||||
|
||||
D8 puts the application under GPLv3. §3.10 admits third-party plugins in three forms, and the
|
||||
@@ -1988,6 +2126,8 @@ note where deferring now constrains the design later.
|
||||
| Tethered shooting | — |
|
||||
| Panorama and HDR merge | **Keep the schema open** — these produce images derived from multiple sources, which §5.1's single-source `Image` cannot express. |
|
||||
| Focus stacking | Same provenance consideration. |
|
||||
| Cross-frame face repair ("best take") | A face from a neighbouring frame of the same burst, aligned by its landmarks and blended by FR-DEV-8's heal. Mechanically a spot whose source is another photograph; **the same multi-source schema question as the two rows above, arriving early** — D17. Deferred rather than refused, with its non-goals fixed now: never automatic, geometry not corrected, source frame declared in sidecar, history and export. |
|
||||
| Gaze and eye-contact estimation | Every open gaze model read on 2026-09-19 is trained on Gaze360, MPIIGaze or ETH-XGaze, all research-only, and Gaze360's licence restricts *models trained on it* by name — the InsightFace situation again (D13). Head pose from the five landmarks is the proxy (FR-CULL-8a). Revisit when weights with a clean data chain exist; iris offset within the eye crop is the licence-free fallback if the proxy proves too weak. |
|
||||
| Print layout | — |
|
||||
| Soft proofing | Parameterise the output colour stage by an arbitrary profile so this becomes a UI addition, not a pipeline change. FR-EXP-3's print-dimension mode already half-commits to print workflows. |
|
||||
| ~~Face recognition~~ | **Undeferred 2026-08-09**, in the narrower form specified in §3.9.1 (FR-CULL-8 … FR-CULL-12): people *grouping and search*, no automated selection. Reclassified as culling rather than AI — it is the same mechanical-grouping category as FR-CULL-5, not the taste operation AI masking is. Gated on spike S14 and decision D13. |
|
||||
|
||||
@@ -11,17 +11,17 @@ Denominators are parsed from [`requirements.md`](requirements.md) at run time, n
|
||||
|---|---|
|
||||
| Source files scanned | 354 |
|
||||
| TRACES tags found | 1495 |
|
||||
| Requirements defined | 191 |
|
||||
| Requirements defined | 194 |
|
||||
| Requirements covered | 140 |
|
||||
| **Coverage** | **73.3%** (140/191) |
|
||||
| **Coverage** | **72.2%** (140/194) |
|
||||
|
||||
### By type
|
||||
|
||||
| Type | Covered | Defined |
|
||||
|---|---|---|
|
||||
| FR | 105 | 136 |
|
||||
| FR | 105 | 138 |
|
||||
| NFR | 32 | 49 |
|
||||
| R | 3 | 6 |
|
||||
| R | 3 | 7 |
|
||||
|
||||
## Orphan tags
|
||||
|
||||
@@ -176,13 +176,15 @@ _None._
|
||||
|
||||
## Not yet tagged
|
||||
|
||||
51 of 191 requirements have no implementation tag. Expected while the codebase is young; each should gain one as it is built.
|
||||
54 of 194 requirements have no implementation tag. Expected while the codebase is young; each should gain one as it is built.
|
||||
|
||||
<details><summary>Show untagged requirements</summary>
|
||||
|
||||
- FR-CAT-14
|
||||
- FR-CULL-13
|
||||
- FR-CULL-6
|
||||
- FR-CULL-7
|
||||
- FR-CULL-8a
|
||||
- FR-DEV-19
|
||||
- FR-DEV-3g
|
||||
- FR-DSP-2
|
||||
@@ -231,5 +233,6 @@ _None._
|
||||
- R1
|
||||
- R2
|
||||
- R5
|
||||
- R7
|
||||
|
||||
</details>
|
||||
|
||||
Reference in New Issue
Block a user