2fb3eb5d2dfb5699a63aaeba63de8af2280bf595
19
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
f9e331eed2 |
Let a test build be opened up, when asked
`adb shell run-as` refuses on a release build — "package not debuggable" — and the app's private storage is then unreachable from the host. That storage is where the face shards, the thumbnail store and the catalog live, so when a device disagrees with the desktop about what it has synced there is no way to find out which of them is right. An evening was spent guessing at exactly that. `DARKROOM_DEBUGGABLE=1 ./docker/android/package.sh --install` now sets `android:debuggable` through aapt2's `--debug-mode`, and nothing else changes. Set through aapt2 rather than written into `AndroidManifest.xml` on purpose: the flag then exists only for the build that asked for it, and a release build cannot inherit it because somebody forgot to take it out again. A debuggable APK lets any process on the device read this app's files, so it belongs on a test tablet and nowhere else. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
2d95807542 |
Package the models on every platform, not just the phone
The Android bundling landed the weights under that platform's asset directory, which was the wrong home the moment a second packager wanted them. `makepkg -si` produced a desktop install with no model at all — the same "no face model is installed" the phone used to show, for the same reason: nothing put the files anywhere the app looks. So `models/face/` at the root is the one copy, and both packagers read it: assemble-apk.sh bundles it as APK assets, and the PKGBUILD installs it to /usr/share/darkroom/models. Both refuse an LFS pointer rather than shipping a 130-byte file that fails inside the graph loader on a user's machine. `face_models` now searches three places, most specific first: the account's own directory, the shared user directory, then $XDG_DATA_DIRS. So a packaged pair is found automatically and a pair the user placed by hand still outranks it — which is what keeps a deliberate choice of weights from being overridden by an upgrade. $XDG_DATA_DIRS rather than a hard-coded /usr/share: that is the variable a distribution, a prefix install or a Nix-style store already sets to say where its data went, and its documented default is exactly the two paths that would otherwise have been hard-coded. Empty on Android, which has no such directories — there the APK's copy is unpacked into the shared user directory instead, because an asset inside a package is not a path anything can read from. Verified: the APK still carries both models at assets/models/, the PKGBUILD parses and installs from the new path, 467 tests pass. Includes the pkgver 0.6.0 → 0.7.0 bump that was already sitting uncommitted in the working tree. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
eaafacc3fb |
Give the phone the model it had no way to obtain
Face indexing was compiled into the APK all along — dr-ui takes dr-face with `inference` on every target, so SCRFD, alignment, MBF, calibration and clustering were all in there. What was missing was the weights, and on Android there was no way to supply them. Route C (docs/faces.md §2.2) says the user obtains the model and the app loads it. On a desktop that is a real gesture: drop two files in ~/.local/share/darkroom/models/ and indexing starts working. On Android it is not a gesture at all. `internal_data_path` is app-private, `run-as` needs a debuggable build, and the in-app fetch route C specifies was never built — so the settings page reported "no face model is installed" on every launch with nothing behind the message. Not "off until you supply weights"; off. So the shape-fixed pair goes into LFS under the APK's assets, assemble-apk.sh copies it into the package, and `android_main` unpacks it to the shared models directory before anything asks whether a model is present. Three things that are not incidental: The models directory is now shared across accounts rather than per-account. Weights are identified by `faces.model_id`, not by who is signed in, so two accounts had no reason to hold two copies — and the unpack runs before any session exists to key a per-account path off. `face_models` still prefers a per-account directory when one is populated, so anyone mid-migration keeps the ability to pin one library to its own pair. The unpack writes under a temporary name and renames. `face_models` decides availability on `is_file()` alone, so a copy truncated by the process being killed would leave a file that passes that test and fails inside tract — reported to the user as a broken model rather than a missing one. assemble-apk.sh refuses an LFS pointer. At ~130 bytes it looks exactly like a model to `cp`, and unchecked it reaches the device and fails in the graph loader instead of telling someone to run `git lfs pull` — the same guard dr-segment's build script applies to yolo26n-seg.onnx. The licensing half is unchanged and recorded in §2.2a: the InsightFace grant is research-only, this is a private repository and a self-installed build, and these files come back out before anything is published. The weights are still not a cargo build input — dr-face has no `models/` directory and no `embedded-model` feature, and nothing in the build reads them. The APK assembly step copies two files and is the only thing in the tree that knows they exist. Verified on device: both models unpack on first launch (2524817 and 13616095 bytes) and the APK carries them at assets/models/. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
871a0eac28 |
Check the image has the commands before CI finds out it does not
🐳 Android image / Build and push (push) Successful in 1s
Build and test / android-image (push) Successful in 1s
Build and test / Desktop (Linux) (push) Successful in 19m21s
Build and test / Layer separation (push) Successful in 29s
Traceability / Requirement traces (push) Successful in 23s
Build and test / Android (aarch64) (push) Failing after 33m3s
Two failures in a row were the same shape: a command the workflow calls was not in the image. git-lfs, then file(1). Each cost a full run to learn, and the android job is expensive to be wrong in -- the step that fails is at the end, so every attempt paid twenty-eight minutes of cross-compile first to reach the line that could not work. Both were visible in ten seconds from here. `docker run <image> command -v file` is the whole diagnosis; it just never occurred to anybody to ask before pushing. So the question gets asked automatically. This reads the `run:` blocks out of the workflows, pulls the commands worth doubting -- the ones a minimal Debian plausibly lacks, not `cd` -- and checks each against the image its job declares. It does not run the workflow and is not a replacement for one. It answers exactly the question that was expensive to answer. `git lfs` is handled specially and the comment says why: it is a subcommand, so the first word of the line is `git`, which is always there. Taking first words alone would have missed the original bug -- and did, in the first version of this script. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
8f863ec56f |
Put file(1) in the Android image too
Build and test / Desktop (Linux) (push) Successful in 1h22m29s
Build and test / Layer separation (push) Successful in 35s
Traceability / Requirement traces (push) Successful in 25s
🐳 Android image / Build and push (push) Successful in 10m19s
Build and test / android-image (push) Successful in 10m19s
Build and test / Android (aarch64) (push) Failing after 33m30s
The android job now fetches the model and cross-compiles the whole app --
28 minutes of it -- and then dies on
file: not found (exit 127)
`Verify minimum API level` reads the linked API out of the .so's ELF
notes with file(1), and the image has never had it. Like git-lfs, the
absence could not show until something got that far: every previous run
panicked in dr-segment's build script long before this line, so the step
that was going to fail never ran.
Audited the rest of what the remaining steps invoke against the image
rather than find the next one the same expensive way -- zip, keytool,
base64, mktemp, shred, find, sed, awk, and aapt2/zipalign/apksigner/d8
from build-tools are all present. file was the only gap left.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
40e6334bb1 |
Sign the APK with a real key when one is configured
🐳 Android image / Build and push (push) Successful in 3s
Build and test / android-image (push) Successful in 3s
Build and test / Desktop (Linux) (push) Successful in 21m23s
Build and test / Layer separation (push) Successful in 29s
Traceability / Requirement traces (push) Failing after 28s
Build and test / Android (aarch64) (push) Failing after 33m28s
The APK has been debug-signed with a key generated on the spot, which is right for putting a build on a test device and useless for anything else: a different signature every run, so nothing can ever update in place. Four secrets now select a real signature -- ANDROID_KEYSTORE_BASE64 and its password, alias and key password. The names are JellyTau's, because that repo already signs its Android build this way against this same runner and one convention across both is one thing to remember. Absence of the secrets is not an error. A fork or a branch build has no access to them and should still produce an installable APK, so the debug path stays exactly as it was. The reverse is an error: if a keystore is supplied and cannot be read, the build fails rather than quietly falling back to a debug key, because a release that is silently debug-signed is worse than no release. Passwords reach apksigner and keytool as `env:`, never `pass:`. `pass:` puts the password in the process table for anything on the box to read. The keystore is written to a 0700 mktemp directory and never into the workspace, which is both what actions/cache saves and what the upload step globs. Also: upload-artifact drops from v4 to v3. v4 was a guess about what this Gitea supports. v3 is what JellyTau uploads its APK with on this runner today, which makes it the version known to work rather than the one that ought to. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
1d38015a7b |
Put git-lfs in the Android image, which never had it
Build and test / Desktop (Linux) (push) Successful in 1h21m48s
Build and test / Layer separation (push) Successful in 27s
🐳 Android image / Build and push (push) Successful in 10m40s
Build and test / android-image (push) Successful in 10m40s
Traceability / Requirement traces (push) Successful in 26s
Build and test / Android (aarch64) (push) Failing after 33m49s
The Android job has never fetched the model. Not because of the header
collision the desktop job hit -- that one is fixed and the desktop job
now pulls all 11 MB -- but because the image has no git-lfs at all:
git: 'lfs' is not a git command. See 'git --help'.
The fetch step dies on its first line, `git lfs install --local`, before
any of the auth handling runs. The build then panics in dr-segment's
build script with a message telling you to run `git lfs install && git
lfs pull` -- advice that could not have worked, because the client it
names was never in the image to run.
Both jobs failing their fetch step at the same time made this look like
one bug with one cause. It was two, in two different images, and the
desktop one was noisier: it had a client, so it got as far as an HTTP
error worth reading. The android one had nothing to say beyond the name
of a missing command.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
5741ec5e00 |
Lift the APK assembly out of package.sh so CI can run it too
package.sh does two things: it decides how the host reaches the image, and it assembles an APK once inside it. Only the first half is host-specific. CI already runs in that image, so the second half was about to be copied into a workflow step -- two copies of aapt2/zipalign/apksigner ordering, drifting apart at whatever rate the toolchain moves. So it moves to docker/android/assemble-apk.sh, which assumes it is inside the image and takes its paths from the environment, because the callers disagree about them: the container mounts the repo at /work, the runner checks it out wherever it likes. Every default reproduces what package.sh did, so the host path is unchanged. Two things stop being hard-coded on the way. The build-tools version and the compile SDK are resolved from what is installed rather than written out as 36.0.0 and android-36 -- the versions are Dockerfile ARGs, and a second copy is a second thing to miss when they move. --min-sdk-version now comes from that same ARG instead of a literal 28, which is the number the API-level check in CI already reads. The intermediates are removed at the end. They were harmless in a cache directory nobody looks at; beside a published artefact they are four more files for a glob to pick up by mistake. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
5a0a9719eb |
Set the version once, in one place, for every artefact
The workspace read 0.1.0 for two releases, so every desktop binary reported a version two releases stale. The APK was worse: `AndroidManifest.xml` states no version at all, so a device showed `versionName=null` and `versionCode=0` while the library inside the APK knew exactly what it was. A version edited by hand in several files is a version that is wrong in at least one of them. `tools/set-version.sh` is now the only thing that sets one. It takes the version from the latest git tag, or is told, and writes the two files that must state it before anything is built: the workspace `Cargo.toml`, from which every crate inherits, and `packaging/PKGBUILD`, which pacman reads before a build exists. It refreshes `Cargo.lock`, because members appear there by version and CI builds `--locked`. `--commit` commits the result. Android is not in that list on purpose. `package.sh` reads the version out of `Cargo.toml` and hands it to `aapt2 link`, so the APK cannot drift from the binary it contains — there is no third file to forget. `versionCode` has to be one increasing integer, which a semantic version is not, so it is packed as MAJOR*10000 + MINOR*100 + PATCH: ordered the way Android requires, and readable at a glance. A version that is not MAJOR.MINOR.PATCH is refused rather than coerced. It is a contract with whoever reads a bug report, and silently turning "0.4" into something else is worse than being asked to type it again. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
3cfa78cde2 |
Give the app a face and a name on the launcher
There was no icon anywhere, and on Android that was not a missing line in the manifest. `aapt2 link` was being handed a manifest and nothing else, so the APK carried no res/ and no resources.arsc — there was no table for an `@mipmap/...` reference to resolve against even if one had been written. Packaging now compiles the resource tree first and links the result in, which is the two steps aapt2 insists on: link reads compiled input only, never a directory. That absent table is also why the launcher caption was blank, which had looked like a second, separate bug. `android:label="DarkRoom"` was there and correct the whole time, and Settings' App info read it fine; the launcher could not, because resolving a label goes through the package's Resources and there were none to open. Nothing about the label changed here. It came back with the table under it. `android:icon` then names one drawable for both icon generations, because the `anydpi-v26` qualifier is what separates them. API 26 and up take the adaptive icon and its three layers; the third of those, monochrome, is what lets Android 13 recolour it rather than drop the app out of the themed set. Below 26 the same name lands on a density-matched PNG. `roundIcon` is deliberately absent — a launcher old enough to read it is one that would ignore the adaptive XML, and minSdk is 28. The desktop icon is one `@image-url` on the window, and the only raster asset in a UI that is otherwise entirely Path. The reasoning at the top of icons.slint does not reach it: that is about glyphs a font might not carry, and this image is never drawn by us at all. It goes to the window manager, which wants pixels and composites them unmasked, so it is pre-shaped with rounded corners rather than square the way the Android layers are. Which exposed Slint's resource default. An `@image-url` compiles down to the absolute path it had on the build machine, to be opened at runtime — already wrong for Android, where the build happens under /work inside a container and no such directory exists on the device, and wrong silently, as an image that loads empty. `EmbedFiles` puts the bytes in the binary instead. It reaches nothing else, since every glyph is a Path. Verified on a device: the APK installs and the home screen draws both the icon and "DarkRoom" under it, where before it had neither. In the link step the adaptive icon resolves at all six densities and resources.arsc lands uncompressed, which API 30 requires and the existing zipalign preserves. On the desktop by reading _NET_WM_ICON off the running window — 256x256, as handed over. Where that actually shows is narrower than it sounds, and the comment says so: Wayland ignores the property in favour of matching app_id against an installed .desktop file, which this repo does not install. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b044a8c067 |
Build the Android CI image in CI, not on a laptop
The Android job ran in gitea.tourolle.paris/dtourolle/darkroom-android:latest, a tag that had never been pushed. The image existed only as a local darkroom-android:latest on one machine, so every Android job died at docker pull with "manifest unknown" before reaching a step. The registry API confirms it: that manifest is a 404 while the other builder images answer 200. android-image.yml now builds and pushes it, following KPN's docker.yaml — host runner rather than a container, so it has the Docker daemon and the host's cached registry credentials, and a plain-git checkout because that host has no Node for actions/checkout. Where it diverges from KPN: that workflow gates on dorny/paths-filter running inside the builder image, which here would need the very image that is missing. The tag is the git tree hash of docker/android instead, which changes when and only when a file there changes. An unrelated push reuses the image, a Dockerfile edit cannot keep serving a stale latest, and a missing tag rebuilds itself without a manual step. The presence probe is curl against the registry API, not `docker manifest inspect`. The latter exits 1 on this registry even for tags that are plainly there — jellytau-builder:latest answers HTTP 200 while docker reports "manifest unknown" for it — and trusting it would have rebuilt 7 GB on every push. The HEAD request also yields Docker-Content-Digest, so latest is repointed only when the digests actually disagree, without pulling any layers. A probe that cannot authenticate falls through to building. Rebuilding when it was unnecessary costs minutes; skipping a build that was needed is the failure this commit exists to remove. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
03326242a1 |
Make the CI checks say what they mean, and format the workspace
The Android job's "Verify minimum API level" step has never verified the minimum API level. It took the first `*.so` anywhere under the target directory, which is a host proc-macro from debug/deps — an x86-64 object built by the runner's gcc, whose .comment section cannot mention Android and so can never contradict the expected value. It now reads the artifact under the target triple, compares against MIN_API parsed from the Dockerfile rather than a second copy of the number, and fails on a mismatch. Both sides are checked non-empty first: two failed parses would otherwise compare equal and pass, which is the same silent success in a new costume. The Android image installs one SDK package per layer and keeps the output. sdkmanager is a JVM program that aborts when it cannot get memory, and the single `> /dev/null` step reported that as a bare "exit code 134" while a retry re-downloaded everything that had already succeeded. tools/ci-local.sh runs all four jobs — desktop, android, layering, traceability — against the host toolchain, which is pinned to the same 1.92.0 CI installs. Its matrix check compares regeneration against the working tree rather than against HEAD: CI starts from a clean checkout, so git's answer is the right one there and reports every local run stale here. The rest is rustfmt across the workspace, and the clippy findings that surfaced once it did: manual_contains in dr-thumbs and collections_ui, a map iterated as pairs for its keys, an index loop over a slice, and two runtime assertions on a constant now made at compile time. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
4d78041d1d | Many imorovments | ||
|
|
d49b4b41de |
Add thumbnail size classes and grid zoom; fix the scrub ordinal
The grid now zooms, which needs thumbnails at two resolutions rather than one, and exposed a scrub that landed in the wrong place. **Two thumbnail size classes.** `ThumbSize::Grid` (256px, ~10 KB) and `Large` (1024px, ~45 KB), with the class part of the store key so both coexist. Storing everything large would take the reference library from ~200 MB to ~860 MB, and shards sync, so that is transfer cost on every device rather than only disk. A store written before the class existed migrates in place: its entries are all grid-sized, which is what the column defaults to, so nothing already fetched is discarded. `forget` now drops every size for an image. Reading a single row left the other class's bytes on the shard's tally for good, sealing it early on space nothing occupied. **Grid zoom.** Ctrl+wheel and pinch resize cells between 90px and 420px in geometric steps, so the gesture feels the same at either end where a fixed pixel step would be imperceptible at 400px and violent at 90px. Crossing 256px switches to the large class, so a zoomed cell is sharp rather than upscaled. Columns and window capacity already derived from cell size, so the grid reflows for free. **The scrub landed about half a library too high.** It counted only dated images while the grid shows all of them — 10,733 dated against 19,841 rows — and ignored `shadowed_by`. Verified against the live catalog: the old formula gave 10,887, the new one 10,732, the true grid position 10,732. The scrub's count and the grid's window must use identical predicates and ordering; a test now fails if they diverge. **Timeline gestures are continuous.** Scrub and pan were quantised to whole buckets, so a slow drag did nothing until it crossed a boundary and then jumped a month. Both work in fractions of the visible span now, and pinch-to-zoom arrives for tablet, where there is no wheel to reach the axis with. The pinch accumulator was wrong on first writing: it took at most one step per update, so an 8x spread — three doublings — yielded one zoom level. `log2().trunc()` now extracts every whole doubling and carries the remainder. The original test asserted the wrong number and defended it in a comment, which is worth remembering: a test can entrench a bug as readily as catch one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
9a24623e35 |
Fix the workspace build off-device
Two breaks that only appeared on a full `cargo test --workspace`. `slint::android` exists only when compiling for Android, so darkroom-android failed to compile on the host even though it is a workspace member. The entry point is now gated on the target rather than on a feature. The timeline forwarded `scrub` where the Timeline component declares `scrub-to`, which the Slint compiler rejects. Assisted-by: LLM |
||
|
|
7900184383 |
Point the Android Java build at the installed SDK jar
ANDROID_PLATFORM means "link native code for API 28" to cargo-ndk, but the android-build crate reads the same variable as "compile Java against platforms/android-28/android.jar" — a directory that does not exist in the image, because only the compile SDK is installed. Slint's Android backend builds a Java helper through that crate, so it panicked with "No Android platforms found" while android.jar sat in android-36. ANDROID_JAR is checked ahead of the platform lookup and settles it: Java compiles against the compile SDK, native code still links against MIN_API. The two are meant to differ; only the variable name is overloaded. Assisted-by: LLM |
||
|
|
5dc1279429 |
Add the Android app shell; cap cross-build parallelism
Cap both halves of the container build: CARGO_BUILD_JOBS limits how many rustc processes cargo starts, while --cpus limits what the container gets regardless of what nested build scripts spawn — cc, cmake, and ring's asm build all parallelise on their own account and do not consult cargo. Without both, a full cross-compile takes every thread on the host and makes the machine unusable for the length of a background build. Assisted-by: LLM |
||
|
|
cc1c5c892d |
Support requesting VFS hydration; fix Android TLS cross-compilation
Correcting the previous commit: I claimed VFS placeholders could not be
downloaded. That was wrong. The desktop client exposes a socket at
$XDG_RUNTIME_DIR/Nextcloud/socket speaking newline-delimited
COMMAND:argument, and MAKE_AVAILABLE_LOCALLY:<path> does fetch the file.
Verified against client 4.0.7: a 1-byte stub became a real 2.8MB file in
2.8 seconds.
Implemented as dr-sync-nextcloud::desktop_client, deliberately optional.
Android has no desktop client, no XDG_RUNTIME_DIR socket and no
placeholders, so detect() returns None there and callers fall back to the
connector. It earns its place only because it is ~30 lines with no
dependencies: where a library already lives in a VFS folder, asking the
client to fetch beats downloading a second copy over WebDAV and leaving
the client's placeholder state inconsistent.
What this does not change: hydration is whole-file, so it suits the
original tier and never browsing. Filling a grid this way downloads the
entire library. Range extraction remains the only mechanism satisfying
FR-NC-3, and ARCH §9.0 now says so precisely.
Also fixes two real Android build failures found by cross-compiling:
- reqwest's `rustls` feature defaults to aws-lc-rs, whose aws-lc-sys
crate is C and fails under the NDK — exactly the pain D1 chose Rust
to avoid. Switched to rustls-no-provider + ring, installing the
provider in the constructor so no caller can build a client that
panics on first use.
- ring itself needs CC/AR per target; cargo-ndk sets only the linker.
Added them to the container.
87 tests passing. dr-sync-nextcloud cross-compiles for aarch64-linux-android.
|
||
|
|
82a5e21ec6 |
Initial workspace: GPU context, compute pass, adaptive Slint shell
Establishes the v0.1 foundations on both platforms: - dr-types: SourceRef (never a filesystem path — Android SAF has none), Format, Availability, Validator with ETag quote normalisation - dr-gpu: wgpu device, compute pass writing a storage texture, resize - dr-ui: Slint shell with FR-UI-1 adaptive layout, computed in Rust to avoid a binding loop - docker/android: pinned toolchain, verified producing API 28 ARM binaries Measured the cost of the temporary CPU readback path (dr-gpu bench): compute is 0.06-0.28ms across sizes while readback is 0.63-7.43ms, so readback is 90-96% of frame time and scales with area. Recorded in ARCH §6.1 — this is why spike S1 is the priority. Mitigations pending S1: reuse the staging buffer, apply at most one resize per frame, and cap render resolution at 2048 on the long edge. 10 tests passing; core crates cross-compile for aarch64-linux-android. |