Files
dtourolle a9271c4850 Make Best one network, and retire Medium and the mixture
Best was a mixture of two experts and a gate, 110 GMAC a megapixel;
Medium a single network at 48 that was softer on real edges. fb-combo
(darkroom-denoise, 20 000 steps from fb-edges2, taught by the mixture
with a quarter of its crops from the edge-rich parts of the frames) is
Medium's shape and holds the mixture's edges on real photographs:
edge PSNR within 0.04-0.06 dB at ISO 1600/6400/25600, more sharpness
kept at all three, the chart's edge 0.89 photosites wide against 0.82.
It is 0.27 dB short on smooth areas at ISO 25600. It becomes Best, and
the methods are Bilinear, Fast and Best.

Saved edits keep their numbers: 2, which was Medium, is now Best, and
3, which was Best, is past the end and reads as the default, Best.
The network ships as mosaic-hq, a new name: the result cache keys a
model by name and size, and this one is byte for byte the old Medium's
size. Its tablet form (A16W16) lost 0.00 dB in simulated QDQ at every
ISO and at most 0.09 dB across the noise bracket.
2026-10-07 06:58:24 -04:00

9.0 KiB
Raw Permalink Blame History

Model weights — licensing

Two Ultralytics checkpoints ship here, both exported by tools/export-seg-model.sh, each with its class vocabulary written out by the same script:

File Checkpoint Trained on Used by
segment/yolo26n-seg.onnx yolo26n-seg.pt COCO, 80 thing classes local adjustments, subject selection
scene/yolo26s-sem-ade20k.onnx yolo26s-sem-ade20k.pt ADE20K, 150 classes the scene tab's per-category grades

Both come from https://huggingface.co/Ultralytics/YOLO26. The face weights in face/ are a separate matter with a separate grant — see face/README.md. The keypoint weights in keypoints/ and the border filler in inpaint/ are the other two, and the easiest — see the last two sections.

Every quantised sibling — *.int8.onnx, *.a16w8.onnx, *.a16w16.onnx, the forms the tablet's Hexagon runs (tools/quantise-models.sh) — is the same weights rounded, and carries exactly the grant of the file it was made from. What calibration adds is one minimum and maximum per tensor: from public COCO val2017 photographs (CC-BY 4.0) for the image models, and for the denoiser from the same training tiles its weights were learned from. No image is in the files.

The grant

Ultralytics releases YOLO under AGPL-3.0, and the weights carry the same grant as the framework — the HuggingFace repository declares agpl-3.0 for the checkpoints themselves, not merely for the training code. A commercial licence is offered separately; DarkRoom does not use it and does not need it.

What that means for DarkRoom

DarkRoom is GPL-3.0-or-later. GPLv3 §13 explicitly permits combination with AGPL-3.0 code, so redistributing these weights inside this repository is allowed — this is not the situation the InsightFace "buffalo" weights would have created, where a non-commercial research grant is simply incompatible with the project's licence and with F-Droid, Flatpak and Play distribution (NFR-COMPAT-2, D13).

The consequence, and it is a real one: the combined work is effectively AGPL-3.0. §13's permission runs one way — the AGPL's §13 network-use condition attaches to the portion under that licence. For a local-first desktop and Android photo editor that condition has no practical bite, because there is no network service offering the combined work to remote users. It would acquire bite the moment any hosted or server-side rendering appeared, and that is the thing to remember rather than rediscover.

This was decided deliberately (D14), not arrived at by accident, and docs/dev/segmentation.md §7 records the reasoning.

Class vocabulary — a caveat worth reading

docs/dev/segmentation.md §4 specified YOLO pretrained on ADE20K, whose 150 classes include the stuff categories that matter most in photography — sky, vegetation, water, wall, mountain.

This was true when written and is not any more. Checked 2026-08-21, no YOLO/ADE20K combination existed: Ultralytics shipped YOLO26-seg on COCO only, and the one HuggingFace repository claiming otherwise (laxmacl/yolov8-ade20k) was empty. Re-checked 2026-08-30: Ultralytics now ships a semantic task with ADE20K checkpoints (https://docs.ultralytics.com/tasks/semantic), and yolo26s-sem-ade20k is what scene/ holds.

So the two vocabularies divide the work rather than compete:

  • segment/, COCO, 80 things. Separates instances — clicking one of three people selects that person. This is what local adjustments need, and a semantic model cannot do it: it would return one "person" region covering all three.
  • scene/, ADE20K, 150 classes. Labels every pixel, including the stuff COCO has no word for — sky, vegetation, water, mountain, wall. This is what the scene tab's per-category grades need, and it does not care that instances are merged, because a per-category grade applies to the whole category.

Neither replaces the other. Keeping both is the deliberate choice.

The loader treats each vocabulary as model metadata rather than compiled-in knowledge, which is what made adding the second model a file plus a descriptor rather than a code change — as this document predicted it would be.

keypoints/ — XFeat, Apache-2.0

File Source Trained on Used by
keypoints/xfeat-1024.onnx weights/xfeat.pt from https://github.com/verlab/accelerated_features MegaDepth + synthetic warps, by the authors panorama alignment (FR-MRG-8), landscape frames
keypoints/xfeat-768.onnx the same weights — the same, portrait frames

Exported by tools/export-xfeat.sh at fixed grayscale inputs of 1024×768 and 768×1024 — the same weights twice, because tract needs a static shape and a portrait frame in a landscape input wastes half of it. Only the convolutional network is in each file; the keypoint decoding is Rust.

The repository and its weights are Apache-2.0, read on 2026-09-19 from the LICENSE at its root, with no separate grant on the checkpoint and no non-commercial clause anywhere in the tree. Apache-2.0 is GPLv3-compatible one way — code and weights under it may be combined into a GPLv3 work — so this is neither the InsightFace situation (D13, a use restriction that binds every user) nor the Ultralytics one (D14, where the combined work becomes AGPL). It is the licence position this document would have wanted for every model in it, and it was chosen over stronger detectors partly for that reason: SuperPoint and SuperGlue are non-commercial, R2D2 and SiLK are CC BY-NC.

The training data is the authors' concern, not a licence on the weights: XFeat trains on MegaDepth, which is itself a research dataset, but the weights are released under the repository's licence without a data-derived restriction — unlike the gaze models §7 of the requirements declined, where the dataset licence restricts models trained on it by name.

inpaint/ — MI-GAN, MIT

File Source Trained on Used by
inpaint/migan-512.onnx migan_512_places2.pt from https://github.com/Picsart-AI-Research/MI-GAN (Sargsyan et al., ICCV 2023), fine-tuned in the darkroom-infill repository (2026-09-20, second model that evening: trained against MI-GAN's own discriminator) Places2 by the authors, then ~7 400 of the maintainer's own photographs with border-shaped voids the panorama border fill (FR-MRG-4)

The bare 512 generator at a fixed 1×4×512×512, six operator types; the tiling, the context and the blend are Rust (dr_pano::fill). Since 2026-09-20 the shipped file is the fine-tune (docs/dev/panorama.md §14), exported by python -m infill.export in darkroom-infill; tools/export-migan.sh still produces the stock generator from the upstream checkpoint, which the fine-tune starts from. The fine-tuned weights are a derivative of the MIT weights trained on photographs the maintainer owns, and carry the same MIT grant.

MIT, code and weights alike — LICENSE and LICENSE-WEIGHTS in the repository, both read on 2026-09-19, both the plain MIT text with no further grant. GPL-compatible, store-compatible, nothing to read around: the cleanest position of any model here. The training set is Places2, a research dataset, but the weights are released under the repository's licence without a data-derived restriction (contrast the gaze models §7 of the requirements declined, and the InsightFace grant of D13).

denoise/ — the mosaic denoiser, the project's own

File Source Trained on Used by
denoise/mosaic-hq-1408.onnx trained in the darkroom-denoise repository (2026-10-07, run fb-combo, 20 000 steps, from fb-edges2 ← student-m), taught by the mixture of experts that was Best until 0.24 (run final) at a half share, with 10 % drawn scenes and 25 % crops from the edge-rich parts of the training frames 1,701 of the maintainer's own base-ISO raws and 6,000 synthetic scenes the repository draws itself, with the Canon EOS 6D's measured noise added the learned demosaic and denoise, Best (FR-DEV-3g)
denoise/mosaic-fast-1408.onnx distilled from final (2026-10-04, run student-s, 30 000 steps, from scratch) the same Fast
denoise/mosaic-{hq,fast}.onnx the two networks above with any height and width, by tools/export_whole.py in the same repository from the same checkpoints; identical to the 1408 files at 1408² the same the same methods, a whole frame at a time on a GPU (denoise.md §14)

U-Nets of plain 3×3 convolutions, ReLU, strided and transposed convolutions and additive skips — no third-party architecture code or weights — and, for Best, two of them blended per photosite by a small gate network of the same parts. Each at a fixed 1×1×1408×1408 for mosaic and sigma, exported by python -m denoise.export in darkroom-denoise; the .a16w16.onnx siblings are the same networks quantised for the Hexagon by tools/quantise-models.sh. Trained only on photographs the maintainer owns, so the weights carry no grant but the project's own: GPL-3.0-or-later, like the code (denoise.md §10).