The engine knew f32 and int8, and gave the Hexagon int8 for every role it served. Measured on the tablet itself (inference.md §1.5), int8 lost 5% of the detector's faces at 40-80 px, moved the landmarks 1.5 px, emptied the segmenter's scores and cost the denoiser 5-9 dB; fp16 the HTP refuses outright. `Form` gains A16W8 and A16W16, and `Rung::form` now names one per role: detectors and landmarks A16W8, the segmenter, scene model, border filler and denoiser A16W16, XFeat int8. The embedder and the eye classifiers stay on the CPU. Each loader resolves its `<stem>.<form>.onnx` sibling; the segmenter and XFeat, compiled into the binary, embed their quantised forms on Android only and pick through `choose_embedded`. The probe, the compile step and the cache fingerprint follow the form instead of assuming int8. Detectors on the new form write `scrfd_*_a16+w600k_mbf`, and `model_ids` answers for all three spellings. On the tablet (ORT 1.29 + QNN 2.42), each shipped file against f32 on the same inputs, and against the CPU's f32 time: SCRFD 500m/2.5g/10g A16W8 100% of faces in every band 4.2/5.1/9.0 ms vs 17/56/198 landmarks A16W8 0.25 px in the 192 crop 0.5 ms vs 2.8 YOLO26n-seg A16W16 98.2% found, mask IoU 0.994 12.9 ms vs 90 scene model A16W16 98.9% of cells agree 15 ms vs 151 MI-GAN A16W16 41 dB from f32 in the fill 87 ms vs 488 XFeat int8 pano alignment 0.45 px (f32's own spread 0.41) 6.5 ms vs 58 denoiser A16W16 0.00 dB at every ISO 95 ms vs 1510 a tile Face numbers are over public COCO val2017 photographs, not a library. The APK carries the siblings (BUNDLED 15 -> 19; the old int8 detectors removed), about 43 MB more. The Windows installer and its CI count skip them; the Arch and Flatpak packages list their files and never had them. The ladder example takes a role per model, which is how the per-role forms above were seen landing on the NPU from the real probe.
8.1 KiB
Model weights — licensing
Two Ultralytics checkpoints ship here, both exported by
tools/export-seg-model.sh, each with its class vocabulary written out by the
same script:
| File | Checkpoint | Trained on | Used by |
|---|---|---|---|
segment/yolo26n-seg.onnx |
yolo26n-seg.pt |
COCO, 80 thing classes | local adjustments, subject selection |
scene/yolo26s-sem-ade20k.onnx |
yolo26s-sem-ade20k.pt |
ADE20K, 150 classes | the scene tab's per-category grades |
Both come from https://huggingface.co/Ultralytics/YOLO26. The face weights in
face/ are a separate matter with a separate grant — see face/README.md.
The keypoint weights in keypoints/ and the border filler in inpaint/ are
the other two, and the easiest — see the last two sections.
Every quantised sibling — *.int8.onnx, *.a16w8.onnx, *.a16w16.onnx, the
forms the tablet's Hexagon runs (tools/quantise-models.sh) — is the same
weights rounded, and carries exactly the grant of the file it was made from.
What calibration adds is one minimum and maximum per tensor: from public COCO
val2017 photographs (CC-BY 4.0) for the image models, and for the denoiser from
the same training tiles its weights were learned from. No image is in the files.
The grant
Ultralytics releases YOLO under AGPL-3.0, and the weights carry the same
grant as the framework — the HuggingFace repository declares agpl-3.0 for the
checkpoints themselves, not merely for the training code. A commercial licence
is offered separately; DarkRoom does not use it and does not need it.
What that means for DarkRoom
DarkRoom is GPL-3.0-or-later. GPLv3 §13 explicitly permits combination with AGPL-3.0 code, so redistributing these weights inside this repository is allowed — this is not the situation the InsightFace "buffalo" weights would have created, where a non-commercial research grant is simply incompatible with the project's licence and with F-Droid, Flatpak and Play distribution (NFR-COMPAT-2, D13).
The consequence, and it is a real one: the combined work is effectively AGPL-3.0. §13's permission runs one way — the AGPL's §13 network-use condition attaches to the portion under that licence. For a local-first desktop and Android photo editor that condition has no practical bite, because there is no network service offering the combined work to remote users. It would acquire bite the moment any hosted or server-side rendering appeared, and that is the thing to remember rather than rediscover.
This was decided deliberately (D14), not arrived at by accident, and
docs/dev/segmentation.md §7 records the reasoning.
Class vocabulary — a caveat worth reading
docs/dev/segmentation.md §4 specified YOLO pretrained on ADE20K, whose 150
classes include the stuff categories that matter most in photography — sky,
vegetation, water, wall, mountain.
This was true when written and is not any more. Checked 2026-08-21, no
YOLO/ADE20K combination existed: Ultralytics shipped YOLO26-seg on COCO
only, and the one HuggingFace repository claiming otherwise
(laxmacl/yolov8-ade20k) was empty. Re-checked 2026-08-30: Ultralytics now
ships a semantic task with ADE20K checkpoints
(https://docs.ultralytics.com/tasks/semantic), and yolo26s-sem-ade20k is
what scene/ holds.
So the two vocabularies divide the work rather than compete:
segment/, COCO, 80 things. Separates instances — clicking one of three people selects that person. This is what local adjustments need, and a semantic model cannot do it: it would return one "person" region covering all three.scene/, ADE20K, 150 classes. Labels every pixel, including the stuff COCO has no word for — sky, vegetation, water, mountain, wall. This is what the scene tab's per-category grades need, and it does not care that instances are merged, because a per-category grade applies to the whole category.
Neither replaces the other. Keeping both is the deliberate choice.
The loader treats each vocabulary as model metadata rather than compiled-in knowledge, which is what made adding the second model a file plus a descriptor rather than a code change — as this document predicted it would be.
keypoints/ — XFeat, Apache-2.0
| File | Source | Trained on | Used by |
|---|---|---|---|
keypoints/xfeat-1024.onnx |
weights/xfeat.pt from https://github.com/verlab/accelerated_features |
MegaDepth + synthetic warps, by the authors | panorama alignment (FR-MRG-8), landscape frames |
keypoints/xfeat-768.onnx |
the same weights | — | the same, portrait frames |
Exported by tools/export-xfeat.sh at fixed grayscale inputs of 1024×768
and 768×1024 — the same weights twice, because tract needs a static shape
and a portrait frame in a landscape input wastes half of it. Only the
convolutional network is in each file; the keypoint decoding is Rust.
The repository and its weights are Apache-2.0, read on 2026-09-19 from the
LICENSE at its root, with no separate grant on the checkpoint and no
non-commercial clause anywhere in the tree. Apache-2.0 is GPLv3-compatible
one way — code and weights under it may be combined into a GPLv3 work — so
this is neither the InsightFace situation (D13, a use restriction that binds
every user) nor the Ultralytics one (D14, where the combined work becomes
AGPL). It is the licence position this document would have wanted for every
model in it, and it was chosen over stronger detectors partly for that reason:
SuperPoint and SuperGlue are non-commercial, R2D2 and SiLK are CC BY-NC.
The training data is the authors' concern, not a licence on the weights: XFeat trains on MegaDepth, which is itself a research dataset, but the weights are released under the repository's licence without a data-derived restriction — unlike the gaze models §7 of the requirements declined, where the dataset licence restricts models trained on it by name.
inpaint/ — MI-GAN, MIT
| File | Source | Trained on | Used by |
|---|---|---|---|
inpaint/migan-512.onnx |
migan_512_places2.pt from https://github.com/Picsart-AI-Research/MI-GAN (Sargsyan et al., ICCV 2023), fine-tuned in the darkroom-infill repository (2026-09-20, second model that evening: trained against MI-GAN's own discriminator) |
Places2 by the authors, then ~7 400 of the maintainer's own photographs with border-shaped voids | the panorama border fill (FR-MRG-4) |
The bare 512 generator at a fixed 1×4×512×512, six operator types; the
tiling, the context and the blend are Rust (dr_pano::fill). Since
2026-09-20 the shipped file is the fine-tune (docs/dev/panorama.md §14),
exported by python -m infill.export in darkroom-infill;
tools/export-migan.sh still produces the stock generator from the upstream
checkpoint, which the fine-tune starts from. The fine-tuned weights are a
derivative of the MIT weights trained on photographs the maintainer owns,
and carry the same MIT grant.
MIT, code and weights alike — LICENSE and LICENSE-WEIGHTS in the
repository, both read on 2026-09-19, both the plain MIT text with no further
grant. GPL-compatible, store-compatible, nothing to read around: the cleanest
position of any model here. The training set is Places2, a research dataset,
but the weights are released under the repository's licence without a
data-derived restriction (contrast the gaze models §7 of the requirements
declined, and the InsightFace grant of D13).
denoise/ — the mosaic denoiser, the project's own
| File | Source | Trained on | Used by |
|---|---|---|---|
denoise/mosaic-1408.onnx |
trained from scratch in the darkroom-denoise repository (2026-10-03, run m2, 60 000 steps) |
427 of the maintainer's own base-ISO Canon EOS 6D raws, with the 6D's measured noise added | the learned demosaic and denoise (FR-DEV-3g) |
A U-Net of plain 3×3 convolutions, ReLU, strided and transposed
convolutions and additive skips — no third-party architecture code or
weights — at a fixed 1×1×1408×1408 for mosaic and sigma, exported by
python -m denoise.export in darkroom-denoise. Trained only on
photographs the maintainer owns, so the weights carry no grant but the
project's own: GPL-3.0-or-later, like the code (denoise.md §10).