Commit Graph
7 Commits
Author SHA1 Message Date
dtourolle 7fba28f7d8 Upload a snapshot of a thumbnail shard, never the live file
Every shard is in WAL mode and every put opens its own connection, so
while thumbnails are being generated on several threads — which is when
the first sync pass runs — the log is never checkpointed and the main
file holds whatever the last quiet moment left in it. For a shard created
seconds earlier that is nothing: a zero-byte file with the schema still
in the log. The sync read that file and uploaded it, and every other
device merging it failed with "no such table: thumbs" on every pass.

Copy the shard through SQLite's backup API into scratch first, which
serialises against writers and carries the log, and upload that.
2026-09-20 00:21:14 +02:00
dtourolleandClaude Opus 5 9e47133304 Run the formatter over the shard-naming change
Build and test / android-image (push) Canceled after 1m38s
Build and test / Android (aarch64) (push) Canceled after 0s
Build and test / Desktop (Linux) (push) Canceled after 1m34s
Build and test / Layer separation (push) Canceled after 0s
🐳 Android image / Build and push (push) Canceled after 1m38s
Traceability / Requirement traces (push) Successful in 1m2s
`cargo fmt --all -- --check` is a CI gate and 9a51cc8 landed two files past
it, so the build has been red on master since regardless of what came after.
Whitespace only — a wrapped `Ok(...)` and two `assert_eq!`s split across
lines. No logic is touched and the tests are unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-16 22:29:28 +02:00
dtourolleandClaude Opus 5 9a51cc88d6 Give a shard's remote name the client that wrote it
Build and test / Desktop (Linux) (push) Failing after 50s
Build and test / Layer separation (push) Successful in 24s
🐳 Android image / Build and push (push) Successful in 3s
Build and test / android-image (push) Successful in 3s
Traceability / Requirement traces (push) Failing after 59s
Build and test / Android (aarch64) (push) Failing after 9m39s
Shard ids are per store: every client fills its own numbering from 0, so
"shard 3" names different thumbnails on every device. The derived sync
published them into a flat shard-NNNN.sqlite namespace anyway, which left
two clients writing one name.

Both failures that follow were live. On upload, a client's open shard
overwrote a peer's file of the same id — content the peer still believed
was published and would never restore, because its own copy was sealed and
the name existed. On download, the loop skipped any remote id it already
held locally, which is the only safe reading of a name that says nothing
about who wrote it, so a client holding shards 0..5 never fetched the
peer's 0..5 at all. Between them, two populated clients exchanged almost
nothing: only shards numbered above the other's highest. A fresh device
worked, having no local shards to collide with, which is why this went
unnoticed — it is exactly the case the feature was written for.

The name is now shard-<client>-NNNN.sqlite. The client id is minted per
store in index.sqlite, beside the numbering it qualifies rather than in
settings: a store deleted and rebuilt restarts at shard 0 and must not
claim the remote names its predecessor wrote. Since our own ids now say
nothing about what we have taken from others, index.sqlite also keeps a
ledger of adopted remote names and the size each had when merged. A size
rather than a flag, because a peer's sealed shard never returns but its
open one grows, and re-merging the grown copy is how the thumbnails it
gained since arrive.

Flat names already on servers still parse, reporting no owner, so each
client adopts them once, and nothing is written under that form again. One
whose id and byte size match a local shard is that client's own earlier
upload by the same identity argument the upload path already makes for
sealed shards, so the rename does not cost every client a re-download of
its whole store. Older builds ignore the new names and stop receiving
shards until updated; their own uploads are still adopted, so nothing is
lost.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-16 21:15:12 +02:00
dtourolleandClaude Opus 5 03326242a1 Make the CI checks say what they mean, and format the workspace
Build and test / Desktop (Linux) (push) Successful in 1h23m26s
Build and test / Android (aarch64) (push) Failing after 2s
Build and test / Layer separation (push) Successful in 50s
Traceability / Requirement traces (push) Failing after 1m9s
The Android job's "Verify minimum API level" step has never verified the
minimum API level. It took the first `*.so` anywhere under the target
directory, which is a host proc-macro from debug/deps — an x86-64 object
built by the runner's gcc, whose .comment section cannot mention Android
and so can never contradict the expected value. It now reads the artifact
under the target triple, compares against MIN_API parsed from the
Dockerfile rather than a second copy of the number, and fails on a
mismatch. Both sides are checked non-empty first: two failed parses would
otherwise compare equal and pass, which is the same silent success in a
new costume.

The Android image installs one SDK package per layer and keeps the
output. sdkmanager is a JVM program that aborts when it cannot get memory,
and the single `> /dev/null` step reported that as a bare "exit code 134"
while a retry re-downloaded everything that had already succeeded.

tools/ci-local.sh runs all four jobs — desktop, android, layering,
traceability — against the host toolchain, which is pinned to the same
1.92.0 CI installs. Its matrix check compares regeneration against the
working tree rather than against HEAD: CI starts from a clean checkout, so
git's answer is the right one there and reports every local run stale here.

The rest is rustfmt across the workspace, and the clippy findings that
surfaced once it did: manual_contains in dr-thumbs and collections_ui, a
map iterated as pairs for its keys, an index loop over a slice, and two
runtime assertions on a constant now made at compile time.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-16 12:02:51 +02:00
dtourolleandClaude Opus 5 cd75e5a4c6 Run the library from local data when the server is unreachable
Also carries in-flight work that shared these files: the zoom structure-key
fix in the adjust pipeline, nearest-neighbour filtering past 1:1, the
timeline scrub marker correction, the 423-Locked retry in the metadata
sweep, and the thumbnail size-class migration.

# Offline mode (FR-CAT-9)

The app previously assumed the server was reachable and treated its absence
as a series of unrelated per-operation failures. A launch without a
connection produced an empty grid, even with a complete catalog on disk and
every thumbnail already in the shards.

Reachability is now inferred from traffic the app was already making, rather
than probed for. `RemoteError::indicates_offline` draws the line that makes
this possible: a dead connection is offline, a 403 or a 500 is not — the
server answered, so blanking the library over one forbidden file would be a
worse error than the one being reported. `Reachability` turns those outcomes
into a state, so a library browsing happily never issues a probe at all.

Going offline takes one failure, because the user is already experiencing it.
Coming back requires evidence — a completed scan or a fetched thumbnail —
with a capped exponential backoff behind the manual retry, so twelve sweep
lanes failing together do not schedule twelve immediate probes.

What keeps working: the catalog opens even when the scan that normally
provides it failed, so the grid fills from the last successful scan.
Thumbnails come from the shards. Rating, flagging and collecting are catalog
writes that never touched the network. What stops is opening an original that
was never stored locally, and it now says so in those words instead of
reporting "network error: connection refused" over a photograph.

Work that is pure network is refused rather than left to fail slowly: the
metadata sweep, derived sync, and sidecar writes. The sweep would otherwise
spend a timeout per image across the whole library while the progress bar
implied something was happening. Deferring sidecars is a real gap rather than
a hidden one — a rating made offline reaches its sidecar only when that image
is judged again while connected — and it is recorded as such at the call site.

# The "On this device" filter

A chip beside the rating filters, narrowing the grid to images whose original
is held locally. It composes with the rating terms rather than replacing them,
so "five-star frames I can actually edit on this train" is one filter. The
predicate is SQL, like the rating terms and for the same reason: the count in
the header has to agree with the cells drawn.

It reads `image_cache.tier_actual`, which nothing writes yet — the next
commit fills it. Until then the chip honestly reports zero.

`Tier` gains an explicit on-disk encoding. The variants are ordered by
generosity and the derived `Ord` invites reordering them, which would
silently reinterpret every cached row; the round-trip test is what holds the
two in agreement.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 20:36:50 +02:00
dtourolleandClaude Opus 5 d49b4b41de Add thumbnail size classes and grid zoom; fix the scrub ordinal
The grid now zooms, which needs thumbnails at two resolutions rather than
one, and exposed a scrub that landed in the wrong place.

**Two thumbnail size classes.** `ThumbSize::Grid` (256px, ~10 KB) and
`Large` (1024px, ~45 KB), with the class part of the store key so both
coexist. Storing everything large would take the reference library from
~200 MB to ~860 MB, and shards sync, so that is transfer cost on every
device rather than only disk. A store written before the class existed
migrates in place: its entries are all grid-sized, which is what the
column defaults to, so nothing already fetched is discarded.

`forget` now drops every size for an image. Reading a single row left the
other class's bytes on the shard's tally for good, sealing it early on
space nothing occupied.

**Grid zoom.** Ctrl+wheel and pinch resize cells between 90px and 420px in
geometric steps, so the gesture feels the same at either end where a fixed
pixel step would be imperceptible at 400px and violent at 90px. Crossing
256px switches to the large class, so a zoomed cell is sharp rather than
upscaled. Columns and window capacity already derived from cell size, so
the grid reflows for free.

**The scrub landed about half a library too high.** It counted only dated
images while the grid shows all of them — 10,733 dated against 19,841
rows — and ignored `shadowed_by`. Verified against the live catalog: the
old formula gave 10,887, the new one 10,732, the true grid position
10,732. The scrub's count and the grid's window must use identical
predicates and ordering; a test now fails if they diverge.

**Timeline gestures are continuous.** Scrub and pan were quantised to
whole buckets, so a slow drag did nothing until it crossed a boundary and
then jumped a month. Both work in fractions of the visible span now, and
pinch-to-zoom arrives for tablet, where there is no wheel to reach the
axis with.

The pinch accumulator was wrong on first writing: it took at most one step
per update, so an 8x spread — three doublings — yielded one zoom level.
`log2().trunc()` now extracts every whole doubling and carries the
remainder. The original test asserted the wrong number and defended it in
a comment, which is worth remembering: a test can entrench a bug as
readily as catch one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-09 21:58:32 +02:00
dtourolle d6ddd0703b Add dr-thumbs: a sharded, syncable thumbnail store
A thumbnail is the one derived artefact worth sending over the wire: it costs
a range fetch plus a decode to produce and is identical for every client
looking at the same file. A second device that downloads a shard gets a full
grid without fetching a byte of RAW.

Sharded at 25 MB, filled sequentially. The cap is about sync granularity, not
SQLite's limits — one growing database means every client re-downloads it
whenever a single thumbnail is added, whereas with sequential fill only the
newest shard is ever dirty and sealed shards are safe to cache forever.

Stored JPEG-encoded rather than as raw RGBA: a 256px RGBA buffer is ~256 KB
against ~20 KB encoded, and that 13x is transfer cost on every client.

Keyed on Nextcloud's oc:fileid, stable across server-side rename and move.

Assisted-by: LLM
2026-08-09 20:38:12 +02:00