The markers the previous commit stops writing are already in the
catalogs — 2 on the desktop, 429 on the tablet — and in the shards
both have exchanged. Renaming them to the faces' own id with a fresh
time is what makes the export send each image again, under an entry
newer than the empty one `held_model` would otherwise pick. Where the
old write had inserted its marker beside the right one, the wrong one
goes and the right one is refreshed for the same reason: its entry in
the shards is older than the empty one.
Images V14 left with faces and no marker are not touched. That state is
the quality pass's cue, and the fixed write marks them correctly when
it reaches them.
Checked against copies of both real catalogs: the desktop renames 2,
the tablet deletes 429, both in under 200 ms.
`record_updates` — the write behind the quality, eye and crop passes —
re-marked the image as indexed under the pipeline the pass ran as, and
left the faces it had updated under the id of the detector that found
them. On a desktop set to Thorough that put `scrfd_10g+w600k_mbf` over
faces spelled `w600k_mbf`; on the tablet, `scrfd_10g_i8+w600k_mbf` over
faces it had adopted from the desktop's thorough pass.
Every reader takes the marker and the faces to agree. `marker_under`
reads the marker as the detector having examined the image, so the
upgrade repair never revisits it. The shard store keys each face by
its pipeline id, so `export_to_shards` selects an image's faces by the
marker's id, finds none, and sends an entry that says the thorough
detector looked and found nothing — over photographs with named faces
on them. The desktop's shard index holds 54 such entries beside real
faces; the tablet's eye pass over the faces it had adopted made 430
more, and both devices have exchanged them. `held_model` takes the
newest entry for an image, which is the empty one. Nothing has been
lost yet only because the two spellings of the thorough detector rank
equal and neither side adopts the other's; a third device, or either
one after a reinstall, would adopt "nothing here" for 484 images. And
the desktop's eye pass is 4,739 images from doing the same to every
face from before V14 — which are the ones that only exist on the
desktop, and would then never reach anywhere.
The marker now takes the id the faces carry; the pass's own id is used
only when it dropped the last of them and there is no detector left to
name. A stale marker under another spelling of the same embedder is
removed in the same transaction, so one embedder has one marker.
The tablet showed a fraction of each person: 681 of the desktop's 3,851
confirmations, and none of Ian's 746, Catherine's 626 or my own 480.
Every face that existed on both devices agreed on who it was, and the
people rows were identical — the merge was fine. The missing 3,170
confirmations were on faces the tablet did not hold at all: the
desktop's 16,080 faces from the original detector, on 4,310 images,
detected before schema V14 kept the quality reading.
Those faces were in shards the tablet had already downloaded, in
August's export. `import_from_shards` looked at them on every sync pass
and declined each one, because a face without a quality reading was
"work this device cannot finish": adopting it would write the run
marker, and the marker was what stopped an image being looked at again.
That was true when it was written and has not been since the quality
repair existed — that pass lists its work by `f.quality IS NULL`, not by
the marker, exactly as the eye pass does, and faces without an eye
reading were already adopted on that reasoning.
The refusal had no exit. V14 had deleted the markers of every image
holding such faces so the quality pass would find them, and
`export_to_shards` walks the markers, so the desktop never re-exported
them either; the unmeasured August copies were the only ones there
would ever be. The tablet's answer was to queue all 17,727 images for a
re-detection of its own, a fetch of the whole library, while holding
the faces on disk.
Adopt them. The receiving device's quality pass measures them when it
reaches them, and the desktop's confirmations match onto them by box
overlap on the next catalog merge. The test that asserted the refusal
now asserts the adoption and that the image is still owed to the pass.
Eighteen declared operations, not fifteen; JPEG XL is an export format;
the grid does not filter by keyword, only the catalog's query can; and a
panorama's provenance is a sidecar beside the composite, not a history
step in it.
It said 0.9.0 against a 0.13.1 tree, listed focus peaking and burst
grouping as unbuilt when both have shipped, and opened with a page of
prose about the display path before saying what the application does.
Lead with what it is and a picture of it, say how to get it on each
platform and what state each channel is in, keep the honest account of
what is missing, and put the manual first in the documentation table.
A CLAUDE.md at the root, for anyone changing this code: the ways a redraw
and a catalog read came to cost half a second per click, what each fix
looked like, and how to measure the next one against a copy of a real
catalog.
"How many images still owe a quality reading" was a correlated EXISTS per
image over `faces`, and the face row is 8 KB of embedding and crop before
the column it looks at, so each count opened every row. Six such counts
run on every open of the Identity screen and at the end of every sweep:
160 ms on the reference library.
V19 adds three partial indexes holding only the faces still owing each
pass, keyed on the image and carrying the model id the predicate reads,
and replaces `faces_image` with `(image_id, model_id)` so "does this image
hold this embedder's faces" is answered from the index too. The planner
takes a partial index when the count is driven from `faces` and ignores it
inside the EXISTS, so `Needs::Face` carries the per-face fragment and
`repairs::count` spells the query from the faces' side; the list and the
per-image check keep the EXISTS. A test holds the two spellings to the
same answer for every repair.
Every `move_to` guaranteed its destination's parent with a `MKCOL` for
each ancestor down from the account root, and a trash folder under a
library root several levels deep meant three round trips answering
`405 Method Not Allowed` before the one `MOVE` that did anything — for
every image of a delete, on a connection built for that job.
The backend now records the collections it has confirmed exist and asks
about each once. It lives for one job, so a folder another client removes
mid-batch is the one case this misses, and the `MOVE` then reports the
`409` rather than hiding it.
`faces::people` grouped `face_person` after a LEFT JOIN over every person
and sorted the lot by name; the rail then discarded the empty, unnamed
groups a regrouping pass leaves behind — 17,000 of 19,000 rows on the
reference library. `people_in_use` filters them in the WHERE and joins
`people` to face counts aggregated first (2,000 groups), so the sort sees
only the rows that will be drawn. `count_unassigned` replaces fetching
2,400 ids to take their length. `load_people` 22 ms → 10 ms.
`confirm_all` called `faces::confirm` per face, and `split_off` called
`reject` then `confirm` per face: each opens and commits its own
transaction, so a click on a group of several hundred was several hundred
commits. `faces::confirm_all` is two statements — clear the rejections the
confirmations override, then flip the rows — and `faces::reassign` does a
split's reject-and-confirm for every face under one commit. 16 ms → 2 ms
and 22 ms → 4 ms on the largest group.
A confirm or a reject changes one row and redraws the whole grid, and the
redraw re-read every crop blob of the selected person (4 MB for the
largest) and decoded every one — 316 ms per click on the reference
library's 754-face person, to arrive at the pixels already on screen.
`load_faces` now takes the crops the previous load decoded, keyed by face,
and moves each into its new cell; the blob read is skipped when every face
is already in hand. `refresh` drains the old cells into it rather than
cloning them. The redraw is 2.6 ms.
Every click on the Identity screen's face grid — confirm, reject, split,
rename, merge — redrew the whole screen, and the redraw recomputed the
coverage line. That line lists every repair's outstanding images to count
them: six scans of the images table with a correlated EXISTS over the
8 KB face rows, an ORDER BY the job's visiting order, a Target with its
path per row, and a thumbnail-index query per image with faces. On the
reference library (24k images, 19k faces) that was ~200 ms of the
~540 ms each click cost, spent computing a figure a confirm cannot change.
`refresh` now takes what changed: `Changed::Identities` re-reads the rail
and the grid and leaves the coverage line alone; `Changed::Library` — an
open, a sweep ending or stopped, the face data deleted — re-reads it too.
For the times it does run, `repairs::counts` counts instead of building
and dropping the lists, and the thumbnail store is read once
(`ThumbStore::held`) rather than probed once per image in the audit, the
outstanding list and the proxy repair.
`identity_bench` is the measurement: the reads a click performs and the
batch writes, timed against a copy of a real catalog.
docs/manual/README.md is a tour for a photographer opening DarkRoom for
the first time — one picture per thing, moving where movement is the
point. tools/manual/ is how the pictures are made: drive.py puppeteers the
desktop build on a private Xvfb (launch, click, drag, type, screenshot,
record), scenes.py is each picture as a script, and record.sh runs them
all over a folder and writes the results into docs/manual/media/.
The media is in LFS, with the CI pulls excluding it as they exclude the
fixtures; a screenshot changes wholesale when the interface does.
Nothing in the pictures shows a person, by design: the demo library is
seventy urban and alpine frames, chosen from the catalog's rows that face
detection found nobody in.
The traceability matrix is regenerated here after the rebase that
brought this branch up to master.
Without focus the keys typed into the keywords sheet went to the grid
behind the scrim, and the Return meant for the keyword opened a
photograph. The naming sheet already takes focus on open; do the same.
An empty trash said "No images found — check the library folder and which
formats are ticked", which sends someone off to fix a library that is
fine.
An empty device destination read "Ask each time" on the settings page,
and nothing asks: an export made with the field blank is refused with
"no export folder is set". Say what will happen.
wgpu reports a device out of memory by panicking, and a twelve-frame
merge on a GPU another process is using is where that happens. The panic
unwound the worker, the sender went with it, and the page sat on "Stop"
with every control disabled and nothing to say why — the crash record on
disk was the only sign. Catch the panic and send it as a failure, and
treat a closed channel with no final event as a dead worker too.
The eyes are per layer and outlive the mode, so a photographer coming
back finds the layers they were looking at still lit. But the tint is a
way of looking at a mask, and outside Local there is no mask being looked
at: the sky stayed red through Repair and back in Photo, a mode that had
been left leaving its overlay behind — the fault ui-navigation.md D-N1
exists to prevent.
Five chips in one row declare 440px, and the develop column takes the
widest panel's request — so selecting a category mask levered the sidebar
past the window's edge, clipping the histogram, the group strip and the
subject list. The same trap ChipGrid's comment records for film formats.
A heading was drawn only on a cell that both began a month and began a
row, so at seven columns most months were never named, and the one
heading on screen — always on the window's first cell — was wrong about
every row below it. Worse, two headings drawn on the same cell overprinted
each other. Now a row carries a heading whenever its first cell's month is
not the one last announced: a month starting mid-row is named on the next
row it opens, one row late and right about everything under it.
Every shard is in WAL mode and every put opens its own connection, so
while thumbnails are being generated on several threads — which is when
the first sync pass runs — the log is never checkpointed and the main
file holds whatever the last quiet moment left in it. For a shard created
seconds earlier that is nothing: a zero-byte file with the schema still
in the log. The sync read that file and uploaded it, and every other
device merging it failed with "no such table: thumbs" on every pass.
Copy the shard through SQLite's backup API into scratch first, which
serialises against writers and carries the log, and upload that.
Confirming "/" in the folder picker set an empty root, which the launch
model read as no root at all: "Open library" stayed disabled after the
question had plainly been answered, and a folder library — whose folder
is the whole library — could never be opened without first descending
into a subfolder of it. The empty string was carrying two meanings.
Record the choice as its own fact on the account (`root_chosen`, defaulted
so existing configuration loads unchanged), treat a folder endpoint as
chosen by definition, and let the launch screen say so: a folder is shown
as a LIBRARY rather than an ACCOUNT, the second question becomes an
optional "scan only a subfolder", and the library header names the folder
instead of calling it "· whole account".
The scan reads the first 256 KB of a file for its metadata. A camera
writes its IFDs at the front, so that is the whole structure; the linear
DNG a merge writes puts its first IFD after the pixels, and rawler,
given the head alone, finds no decoder in it. The composite was
catalogued without a date and sorted to the very end of the grid, after
every dated photograph — which is where a panorama merged on the tablet
went unfound.
dr-decode's own TIFF reader now reads through a head and a tail at a
known offset; trailing_ifd says where the tail starts and
metadata_split reads the two together. The scan, when the head fails
and points beyond itself, fetches from the IFD to the end — kilobytes —
and dates the file from both. Tested against the writer's own output.
The Malvar "R at green in R row" kernel weights the two greens two
sites away along the row at -1 and the pair up and down the column at
+1/2. The shader had the two swapped, in the comment as well as the
code, so the transcription checked against itself. Both sum to zero
and reconstruct a flat patch exactly, which is all the tests fed it.
On an edge the correction at green sites is half strength and the
false colour doubles: 0.375 against 0.19 on a grey step, and a
blue/yellow zipper around every clipped highlight at 1:1. The other
three kernels and the CFA tables were right.
A grey vertical step now runs through the pass; the transposed kernel
fails it at 0.375.
The reference desktop's only system ONNX Runtime is Arch's
onnxruntime-opt-cuda: 1.29, built without TensorRT and against cuDNN 8
on a cuDNN 9 machine. The probe rejects both providers correctly and
the app runs on the CPU provider, which is right and not what anyone
wants. runtime/ beside the models is now searched ahead of /usr/lib,
tools/fetch-desktop-runtime.sh fills it with the four libraries from
the current onnxruntime-gpu wheel (cuDNN 9, TensorRT 10), and the
About caption lists every rung that lost and why, not only the first.
Verified: the app selects TensorRT from that directory with no
environment variable set.
The unpack list gained migan-512.onnx without its length following;
nothing on the desktop compiles that crate, and the first Android build
of 0.13.0 stopped there.
A fill that went wrong took a seven-minute merge to look at again. Now
DR_FILL_DUMP=dir makes the merge write what the filler was given, and
the fill example runs fill_border on that, or a crop of it, on the engine
and writes coarse, each band and the feathered result as PPMs — seconds
per attempt on TensorRT. Both examples take DARKROOM_ORT_DIR as the app
does, and --wait-engines lets a compiling rung finish before timing.
A Border choice beside the projection — crop to the picture, or fill it
— that redraws the preview filled so the invented pixels are seen before
they are confirmed (FR-MRG-1), greyed with the reason when the model is
not there. The job fills at half the composite's resolution in a
display-ish space (white balance, matrix, gamma; invertible) and samples
the result back into the linear DNG wherever no frame reached; the
sidecar's merge line says border filled and with which knobs.
Experimental because the fill is right in thin borders and wrong in deep
corners, where the model's Places2 prior puts clouds in sky and water
under grass; so its six knobs — working scale, edge erosion, coarse pass,
band width, mirror depth, seam feather — are sliders under the choice,
each committing a redraw, until the defaults are right.
dr_pano::fill owns everything the model does not — which tiles, what
context, how to blend — behind an Inpainter trait, and dr_pano::migan is
that trait over the shipped generator on the inference engine.
The known content is mirrored across the coverage edge into the hole and
a 256-px ring, the nearest 48 px folded, so the model interpolates between
real and mirrored sky rather than extrapolating into nothing. A coarse
pass at a quarter decides the structure with the whole border in a few
tiles; fine passes in 96-px bands from the edge outward texture it; the
seam is blended over a feather inside the real edge. Every knob is a
Params field, and an Observer hears each stage for whoever is looking at
why a fill went wrong.
A border fill acquires the filler once a tile, and each acquire hashed
the 28 MB model twice — 60 ms a tile, a third of the tile's run on a
throttled TensorRT. The Model keeps its hash from open.
Sargsyan et al., ICCV 2023; MIT code and weights (models/LICENCE.md),
exported by tools/export-migan.sh at a fixed 1×4×512×512 from the
authors' checkpoint — six operator types, 28 MB, in LFS like the rest.
The package installs it beside the scene model and the APK unpacks it
with the others.
The fused shader stored black with alpha 1 for a pixel whose source
coordinate left the frame, and the merge's warp averaged it in like any
other: a dark, badly interpolated fringe along every frame's edge, visible
as a seam wherever a frame ended and, later, as the edge the border fill
continued. The display keeps its opaque black; CameraLinear stores alpha 0
and the warp weights each sample by the alpha it interpolated, dropping a
sample that has none.
MI-GAN is plain convolutions, so every rung serves it and none needs a
special form; the role exists so resolve_model and the probe's fingerprint
know the model, and so the merge job can open it through the engine rather
than tract, which takes 7.4 s a tile for it.
A library's records are never all complete at once. A face found before
its quality was kept has no quality; one found before the eye models
existed has no reading; one adopted from a peer's shard has no crop; an
image the fast detector examined on a 1024 px proxy has boxes the current
detector would not have drawn; an image the scan stat'ed has no capture
date. On the reference library that is 17,762 faces under the bare
w600k_mbf id with no quality, no reading and no dense landmarks, 4,144 of
them without a crop, beside 12,217 images the fast detector examined and
found nothing in. Every one of those gaps was its own pass — V14's
measuring pass, §17.5's eye pass, the sweep's proxy repair, the sweep's
detector upgrade — with its own work list, its own count and its own idea
of done, and adding a per-face field meant adding a pass. There was no
pass at all for the case the library is actually in: boxes and landmarks
drawn by a weaker detector on a proxy, which every later per-face pass
would have read from.
dr_ui::repairs replaces them with one job over a registry. A Repair names
one thing a record can lack — the predicate that says which images still
owe it, the input its handler needs (a header, the original, or a native
render), the handler, and what to record for an image that can never be
done. The job unions the predicates into one work list, fetches each
image once at the most any claimant asks for, renders it at most once,
and runs every handler whose predicate that image still matches, checked
again before each because a detection writes every field a per-face
handler would fill. The registry today: face-proxy, face-quality,
face-eyes, face-crop, face-detection, face-upgrade, metadata — the last
there to say that this is not a face job. Adding a field is one entry.
A repair's predicate is the only definition of its work: the count the
settings page shows, the list the job fetches and the check before its
handler run are one predicate, so the job converges. That is why the
registry is cut to what the device can do rather than listing what it
skips — an entry is a count and a set of originals to fetch — and why an
eye reading that cannot be cut is not a criterion.
The catalog side is generic to match: record_updates writes whichever
fields a FaceUpdate carries and re-marks the image so the shards export
it; faces_needing and count_needing answer a predicate the caller
supplies, replacing the measuring pass's three special cases.
Two buttons on the settings page run the job and differ in one
predicate. "Index faces" converges on coverage: has anything examined
this image. "Re-index every face" converges on provenance: face-detection
claims every image with no marker under the chosen detector, in either
of its forms (FaceDetector::model_ids, so a desktop in f32 and a tablet
on the Hexagon do not re-index each other's work), and a marker saying a
weaker one looked is not that. An original over the fetch budget is left
exactly as it was under the re-index, where the sweep marks it examined:
a re-detection with nothing found would delete the faces, and "cannot
fetch" is not "no faces".
record_detections replaces an image's faces and carried only the user's
confirmations onto the new ones, by box overlap above 0.5 IoU. Everything
else on the old faces was dropped: the suggestions the last grouping pass
made, and the people the user had said a face was not. On the reference
library that is 13,011 suggestions and 77 rejections beside 3,778
confirmations — a re-detection of it would have been correct by
FR-CULL-12's letter, since suggestions are derived data, and would have
handed back a People screen of strangers.
Now every old face is read before the delete — box, vector, assignment,
rejections — and matched to the new faces one-to-one, best pair first. A
pair qualifies when the boxes overlap at all and either the overlap alone
says so (IoU above 0.5, the old rule) or the embeddings do (cosine above
SAME_FACE_COSINE, 0.45, the reference library's P≈0.95 line). The
embedding route claims the box a low-resolution pass drew badly enough
that overlap alone would not; the vector is also what breaks the tie in a
group photograph, where two neighbouring faces overlap both new boxes.
Overlap is required on both routes, because the same vector elsewhere in
the frame — a mirror, a print on the wall — is not the same face and must
not take its name. Onto the matched face go the assignment as it was,
confirmed or suggested with its probability, and every rejection.
The merge's match_faces still matches by overlap alone across devices; it
is the same question and is not changed here.
ONNX Runtime's default is already its fullest level. ort-tract maps any
level but disabled to tract's optimiser, whose slice pass divides by
zero inside the segmenter's graph — a panic across the C API and so an
abort, which is what stopped dr-ui's develop test. The app never asked
tract for that and does not start now.
On the tablet the engine compiled arcface for the NPU: the routing
compared the form a rung wants with the form on offer, and for the
embedder both are f32, so nothing said no. A rung now says which roles
it serves at all, and the Hexagon does not serve the embedder (§7 —
its vectors must compare across devices). Tested at the routing seam.
dr-segment's onnx_probe example still named ort-tract, which is what
stopped the workspace test build.
The strict flag refused the Hexagon over the ten quantise/dequantise
nodes at the graph's edges that QNN declines by policy, which cost
microseconds. A provider that hands real work to the CPU is slower than
the CPU floor and the timing already rejects it; the tablet measured
2.3 ms on the NPU against a 29.7 ms floor.
ONNX Runtime's errors open with a source path and a template signature;
the first 160 characters of a CUDA failure were all signature. The
reason now starts at the first word a person can act on.
XFeat's two exports are a Keypoints role now; the crate no longer names
tract, and the app compiles TensorRT engines for both ahead of the
first merge. The probe picks the smallest *detector* rather than the
smallest file: the tablet's first run chose the 112 KB eye classifier,
which has no int8 form, and reported the Hexagon as failed for want of
one.