The model fetch has been failing on every run with
LFS: Client error: .../info/lfs/objects/0672d7a7...
which reads like a rejected credential and is not one. The object is on
the server and downloads fine; what fails is the shape of the request.
`git lfs pull` makes two calls. The first, to `/info/lfs/objects/batch`,
succeeds -- and Gitea answers it with a short-lived `Bearer` JWT scoped
to that one object, for git-lfs to use on the second. git-lfs sends that
JWT *and* the `Authorization` header this step had installed in git
config, and two `Authorization` headers is a 400 from Gitea. Hence a
client error on the object one step after the batch call it just made
successfully, which is what made this look like an auth problem rather
than a duplication.
Confirmed directly against the server: the JWT alone on that URL is a
200, the JWT plus any second `Authorization` is a 400, and a lone token
header that is merely wrong is a 401 -- so the scheme was never the
issue. `lfs: true` on the checkout fails the same way and for the same
reason, because actions/checkout persists a header of its own; the
comment here blaming a credential the endpoint would not accept was
wrong on both counts.
So the headers are stripped -- checkout's included, since nothing later
in either job talks to the remote -- and the token is handed to git-lfs
as an ordinary credential instead. It authenticates the batch call and
leaves the per-object JWT alone.
This is what fails the Android job today: the build script sees a
133-byte pointer and panics by design, which is the message it is
supposed to give and the one nobody could act on.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Two CI failures, unrelated except that both were mine.
**The test binary faulted under parallel threads.** Every GPU test opened its
own `GpuContext`, and `cargo test` runs on as many threads as there are cores —
so a full run asked the driver to bring up a dozen Vulkan devices at once and
died with SIGSEGV. Serially it passed, which made it look like flakiness rather
than a fault in the harness. One device now, behind a `OnceLock`: a
`GpuContext` is an `Arc<Device>` and an `Arc<Queue>`, so sharing it is a
refcount, and the losers of the race block until the winner is done. 416 tests
now pass in parallel, in half the time twenty-six devices took.
**`lfs: true` on the checkout made the checkout fail.** The intent was right —
the model is in LFS, a plain checkout writes a 133-byte pointer, and the build
script panics on it — but on this server `git lfs fetch` is rejected at
`/info/lfs/objects/<oid>` with a client error: the credential `actions/checkout`
installs for git is not one the LFS endpoint accepts. So a fetch problem
presented as a checkout problem and took the whole job with it.
The object is on the server; a clean clone over SSH with `git lfs install
--local` pulls all 11 MB of it. It is now its own step with an explicit token,
and `continue-on-error` so a credential problem cannot masquerade as a broken
checkout — if it fails, the build still runs and fails with the build script's
own message, which names the real problem.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Both failing jobs failed for the same reason, and it was not the reason either
of them appeared to fail. Desktop reported a Clippy failure and Android
reported the API-level check; both were `dr-segment`'s build script panicking
because `models/yolo26n-seg.onnx` was a 133-byte LFS pointer rather than the
11 MB model.
No checkout step asked for LFS objects. `.gitattributes` predicted this exactly
— "a clone without git-lfs gets a ~130-byte pointer file where the model should
be" — and the build script fails loudly by design rather than embedding a
pointer and failing at inference. That design worked; nothing was reading the
message.
It stayed hidden because the Android job built `-p dr-gpu`, which never reaches
dr-segment. Building the app crate does, which is how one fix surfaced another.
Only the two jobs that compile get `lfs: true`. `layering` runs `cargo tree`
and builds nothing, so it has no reason to fetch 11 MB.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The note above the core-crate check said the UI and app crates would join
"once the Android shell exists". They already had. Building `darkroom-android`
for aarch64 takes 4m23s and produces `libdarkroom.so`, linked for Android 28 —
which is what the step below now checks, and what a device installs.
Rewritten to say what the core check is actually for: a fast, link-free gate
that fails early and names a smaller suspect, ahead of the full build.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The step built `-p dr-gpu` and then looked for a `*.so` to read the API level
out of with `file`. dr-gpu declares no `crate-type`, so it produces an rlib —
an archive of object files no linker has touched, and about which `file` has
nothing to say. The step could therefore only ever reach its own "no aarch64
.so was produced" branch, whatever the linker did, which is the one thing it
exists to rule out.
Builds `darkroom-android` instead: the workspace's only
`crate-type = ["cdylib"]`, and the artefact that ships in the APK. The API
level verified is now the one a device will refuse to install against.
The neighbouring comment claiming only core crates cross-compile is left
alone until the build proves otherwise; it is either stale or this commit is
wrong, and the same run answers both.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The Android job ran in gitea.tourolle.paris/dtourolle/darkroom-android:latest,
a tag that had never been pushed. The image existed only as a local
darkroom-android:latest on one machine, so every Android job died at docker
pull with "manifest unknown" before reaching a step. The registry API confirms
it: that manifest is a 404 while the other builder images answer 200.
android-image.yml now builds and pushes it, following KPN's docker.yaml — host
runner rather than a container, so it has the Docker daemon and the host's
cached registry credentials, and a plain-git checkout because that host has no
Node for actions/checkout.
Where it diverges from KPN: that workflow gates on dorny/paths-filter running
inside the builder image, which here would need the very image that is missing.
The tag is the git tree hash of docker/android instead, which changes when and
only when a file there changes. An unrelated push reuses the image, a Dockerfile
edit cannot keep serving a stale latest, and a missing tag rebuilds itself
without a manual step.
The presence probe is curl against the registry API, not `docker manifest
inspect`. The latter exits 1 on this registry even for tags that are plainly
there — jellytau-builder:latest answers HTTP 200 while docker reports "manifest
unknown" for it — and trusting it would have rebuilt 7 GB on every push. The
HEAD request also yields Docker-Content-Digest, so latest is repointed only when
the digests actually disagree, without pulling any layers.
A probe that cannot authenticate falls through to building. Rebuilding when it
was unnecessary costs minutes; skipping a build that was needed is the failure
this commit exists to remove.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The Android job's "Verify minimum API level" step has never verified the
minimum API level. It took the first `*.so` anywhere under the target
directory, which is a host proc-macro from debug/deps — an x86-64 object
built by the runner's gcc, whose .comment section cannot mention Android
and so can never contradict the expected value. It now reads the artifact
under the target triple, compares against MIN_API parsed from the
Dockerfile rather than a second copy of the number, and fails on a
mismatch. Both sides are checked non-empty first: two failed parses would
otherwise compare equal and pass, which is the same silent success in a
new costume.
The Android image installs one SDK package per layer and keeps the
output. sdkmanager is a JVM program that aborts when it cannot get memory,
and the single `> /dev/null` step reported that as a bare "exit code 134"
while a retry re-downloaded everything that had already succeeded.
tools/ci-local.sh runs all four jobs — desktop, android, layering,
traceability — against the host toolchain, which is pinned to the same
1.92.0 CI installs. Its matrix check compares regeneration against the
working tree rather than against HEAD: CI starts from a clean checkout, so
git's answer is the right one there and reports every local run stale here.
The rest is rustfmt across the workspace, and the clippy findings that
surfaced once it did: manual_contains in dr-thumbs and collections_ui, a
map iterated as pairs for its keys, an index loop over a slice, and two
runtime assertions on a constant now made at compile time.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Ports JellyTau's traceability tooling to Rust, carrying across the bug it
was repaired for. That gate divided a traced count by frozen literal
denominators; the requirements file outgrew them and it reported 158%
coverage, so it could never fail its own threshold.
Two rules, both enforced by the extractor's own tests:
- denominators parsed from docs/requirements.md at run time
- coverage is |traced ∩ defined| / |defined|, never a raw traced count
The gate additionally fails hard on a misconfigured run — zero
requirements parsed or zero files scanned — rather than reporting a
plausible 0%, and on any orphan tag naming a requirement that does not
exist.
Adapted for DarkRoom: IDs are FR-CAT-1 / NFR-P13 / FR-DEV-3a shapes
rather than JellyTau's fixed three digits, and decisions (D), spikes (S),
milestone items (M) and test ids remain taggable while being excluded
from the denominator — counting them inflated it by 25.
Also adds dr-sync: the RemoteBackend trait and capability model, so the
Nextcloud connector is one implementation rather than the only shape the
engine understands. No mature Nextcloud crate exists (reqwest_dav is too
thin), so the connector will be hand-rolled over reqwest per D7.
Gitea workflows follow the same style: containerised, commented with the
reasoning, desktop and Android on every push, plus a CI check that no
core/ crate depends on the UI toolkit (ARCH §6.5a).
Coverage today: 13.3% (19/143). 50 tests passing.