docs/ had 26 developer documents flat beside the manual, and the two audiences are very differently sized: most readers want the manual and the gesture reference, a few want the register, the designs and the measurements. The manual and gestures.md stay at the top; everything for someone changing the code moves to docs/dev/, and the two documents that name their own successors — the v0.1 milestone and the UI-refinement plan — go to docs/dev/archive/ rather than being deleted, since both are still cited. docs/README.md is the index, users first. Every reference follows: code comments, Cargo manifests, the workflows, the pre-commit hook, the bench and traceability tools (which locate the repo root by docs/dev/requirements.md now), packaging, the Docker READMEs, CLAUDE.md, CONTRIBUTING.md and the README. The matrix links one level deeper and is regenerated. Links out of the moved documents into the tree gain a level; a link checker over every Markdown file finds none broken.
449 lines
19 KiB
Rust
449 lines
19 KiB
Rust
//! The synthetic 50k catalog, and the handful of real files it points at.
|
||
//!
|
||
//! `docs/dev/requirements.md` §8 asks for "an automated benchmark suite against a
|
||
//! synthetic 50k catalog". The hard part of that sentence is *50k*: a real
|
||
//! library of that size is several terabytes and cannot live in a repository,
|
||
//! in a CI cache, or on a laptop that also has to compile the thing.
|
||
//!
|
||
//! # The trick, and what it costs
|
||
//!
|
||
//! Rows are cheap and pixels are not. So this builds **fifty thousand catalog
|
||
//! rows** over a **pool of a dozen real image files**, each referenced by
|
||
//! several thousand of them. Everything the catalog half of the suite measures
|
||
//! — opening, counting, windowing, bucketing a timeline — touches only rows,
|
||
//! and is therefore exact. Everything the pixel half measures — decode,
|
||
//! downscale, orient, encode — touches one file at a time and does not care
|
||
//! how many rows point at it. The fixture is ~14 MB on disk instead of ~2 TB
|
||
//! and neither half is flattered by that.
|
||
//!
|
||
//! What it *does* cost is stated rather than hidden: the file pool is small
|
||
//! enough to sit in the OS page cache, so [`crate::thumbnails`] measures CPU
|
||
//! throughput with the read already paid for. That is the right thing to
|
||
//! measure for NFR-P3 — the target is written about the embedded preview path,
|
||
//! not about a disk — but it is not a claim about a cold library on spinning
|
||
//! rust, and the harness does not make one.
|
||
//!
|
||
//! # Reproducible from a seed
|
||
//!
|
||
//! Every value comes from [`Rng`], seeded once. Two machines running the same
|
||
//! seed build byte-comparable catalogs, which is the property that lets a
|
||
//! number measured on the reference desktop be compared with a number measured
|
||
//! anywhere else. [`Stamp`] records what a directory was built from, so a
|
||
//! fixture is reused when it matches and rebuilt when it does not — including
|
||
//! when `dr-catalog`'s schema version moves, since a catalog built by an older
|
||
//! build would otherwise be measured through a migration that a user's would
|
||
//! not run.
|
||
//!
|
||
//! # The sources are generated, not committed
|
||
//!
|
||
//! No photograph in this repository is licensed for redistribution, and a
|
||
//! dozen camera previews would be megabytes of binary in git for ever. So the
|
||
//! pool is synthesised: a coarse gradient with a fine dither on top, which is
|
||
//! the same shape `core/dr-gpu/examples/frame_budget.rs` synthesises its source
|
||
//! from and for the same reason. A flat frame lets the memory system serve
|
||
//! every sample from one cache line, which flatters a box filter by an amount
|
||
//! that has nothing to do with photographs; pure noise defeats the JPEG
|
||
//! encoder's entropy coder in the other direction and would make the encode
|
||
//! half of a thumbnail look worse than any real image ever does.
|
||
|
||
use std::path::{Path, PathBuf};
|
||
|
||
use anyhow::{Context, Result};
|
||
use dr_catalog::Catalog;
|
||
use serde::{Deserialize, Serialize};
|
||
|
||
use crate::stats::Rng;
|
||
|
||
/// How many distinct image files the pool holds.
|
||
///
|
||
/// Twelve rather than one, so that a decode measured over a batch is not one
|
||
/// file's quirks repeated — a single frame that happened to compress unusually
|
||
/// well would set the whole number — and rather than fifty thousand, so the
|
||
/// fixture stays a directory a person can look at.
|
||
pub const SOURCE_POOL: usize = 12;
|
||
|
||
/// The size of one pooled file, in pixels.
|
||
///
|
||
/// 1620×1080 is not a round number: it is what `core/dr-decode/src/preview.rs`
|
||
/// records a Canon CR2 carrying in IFD2, and the embedded preview is what
|
||
/// NFR-P3 names. A camera JPEG is 24 MP and a camera *preview* is about this,
|
||
/// so measuring the preview path against a 24 MP file would measure something
|
||
/// the sweep never does.
|
||
pub const PREVIEW: (u32, u32) = (1620, 1080);
|
||
|
||
/// Folders the rows are spread across.
|
||
///
|
||
/// Spread evenly and without regard to capture date, because what a folder
|
||
/// count decides is the cost of the folder filter's `IN (SELECT …)` and the
|
||
/// size of the `folders` table — not which image is in which.
|
||
const FOLDERS: usize = 400;
|
||
|
||
/// The earliest capture time in the fixture: 13 December 2015, UTC.
|
||
///
|
||
/// Fixed rather than relative to the clock. A library whose dates moved with
|
||
/// the calendar would make `timeline` bucket differently from one month to the
|
||
/// next, and a benchmark that measures a different query each time it runs is
|
||
/// not measuring a regression.
|
||
const EPOCH: i64 = 1_450_000_000;
|
||
|
||
/// The span capture times are drawn from: twelve years.
|
||
///
|
||
/// Long enough that the timeline query has real structure to bucket — at
|
||
/// monthly granularity that is ~144 buckets, which is the shape the scrubber
|
||
/// actually draws — and not so long that a year holds too few frames to look
|
||
/// like a library.
|
||
const SPAN: i64 = 12 * 365 * 86_400;
|
||
|
||
/// Bodies and lenses, for the columns the camera and lens filters read.
|
||
const CAMERAS: [&str; 6] = [
|
||
"Canon EOS R5",
|
||
"Nikon Z 7II",
|
||
"Sony ILCE-7RM5",
|
||
"Fujifilm X-T5",
|
||
"Panasonic DC-S5M2",
|
||
"OM SYSTEM OM-1",
|
||
];
|
||
|
||
const LENSES: [&str; 6] = [
|
||
"RF24-70mm F2.8 L IS USM",
|
||
"NIKKOR Z 50mm f/1.8 S",
|
||
"FE 85mm F1.4 GM",
|
||
"XF16-55mmF2.8 R LM WR",
|
||
"LUMIX S 20-60mm F3.5-5.6",
|
||
"M.Zuiko Digital ED 12-40mm F2.8",
|
||
];
|
||
|
||
/// What a fixture directory was built from.
|
||
///
|
||
/// Written beside the catalog and compared on every run. A mismatch rebuilds:
|
||
/// silently reusing a fixture built from a different seed, a different row
|
||
/// count or an older schema would compare two numbers that describe two
|
||
/// different workloads, which is worse than having no number at all.
|
||
#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)]
|
||
pub struct Stamp {
|
||
/// Bumped by hand whenever anything in this file changes what gets built.
|
||
/// The seed cannot carry that: the same seed through different generation
|
||
/// code produces a different library.
|
||
pub generator: u32,
|
||
pub seed: u64,
|
||
pub images: usize,
|
||
pub sources: usize,
|
||
pub folders: usize,
|
||
pub preview_width: u32,
|
||
pub preview_height: u32,
|
||
/// `dr-catalog`'s schema version at build time.
|
||
pub schema_version: i64,
|
||
}
|
||
|
||
/// Bump on any change to what [`build`] writes.
|
||
const GENERATOR: u32 = 1;
|
||
|
||
/// A built fixture on disk.
|
||
pub struct Fixture {
|
||
pub dir: PathBuf,
|
||
pub catalog: PathBuf,
|
||
pub thumbs: PathBuf,
|
||
pub sources: Vec<PathBuf>,
|
||
pub stamp: Stamp,
|
||
/// The catalog file's size, reported because it is the thing an open has
|
||
/// to read and because it is the honest denominator for "is 2 s a lot".
|
||
pub catalog_bytes: u64,
|
||
}
|
||
|
||
/// Where a fixture lives by default.
|
||
///
|
||
/// The temporary directory rather than `target/`, for two reasons. It survives
|
||
/// `cargo clean`, so a fixture is built once per machine rather than once per
|
||
/// clean; and it is not inside anything CI caches, so a 14 MB catalog is not
|
||
/// uploaded and downloaded on every push to save the two seconds it takes to
|
||
/// generate. `DR_BENCH_DIR` overrides it.
|
||
pub fn default_dir() -> PathBuf {
|
||
match std::env::var_os("DR_BENCH_DIR") {
|
||
Some(dir) => PathBuf::from(dir),
|
||
None => std::env::temp_dir().join("darkroom-bench"),
|
||
}
|
||
}
|
||
|
||
/// Build the fixture under `dir`, or confirm the one already there.
|
||
///
|
||
/// Returns whether it had to be built, so the caller can say so: a run that
|
||
/// includes fixture generation has a warm page cache for the catalog file it
|
||
/// is about to open, and a reader comparing two numbers deserves to know which
|
||
/// of them was measured that way.
|
||
pub fn build(dir: &Path, seed: u64, images: usize) -> Result<(Fixture, bool)> {
|
||
let stamp = Stamp {
|
||
generator: GENERATOR,
|
||
seed,
|
||
images,
|
||
sources: SOURCE_POOL,
|
||
folders: FOLDERS,
|
||
preview_width: PREVIEW.0,
|
||
preview_height: PREVIEW.1,
|
||
schema_version: dr_catalog::schema::SCHEMA_VERSION,
|
||
};
|
||
|
||
let catalog = dir.join("catalog.sqlite");
|
||
let stamp_path = dir.join("stamp.json");
|
||
let sources: Vec<PathBuf> = (0..SOURCE_POOL)
|
||
.map(|i| dir.join("sources").join(format!("preview-{i:02}.jpg")))
|
||
.collect();
|
||
|
||
let usable = matches_stamp(&stamp_path, &stamp)
|
||
&& catalog.is_file()
|
||
&& sources.iter().all(|p| p.is_file());
|
||
|
||
if !usable {
|
||
std::fs::create_dir_all(dir.join("sources"))
|
||
.with_context(|| format!("creating the fixture directory {}", dir.display()))?;
|
||
// The stamp goes last. A build interrupted halfway leaves no stamp, so
|
||
// the next run rebuilds rather than measuring a truncated catalog.
|
||
let _ = std::fs::remove_file(&stamp_path);
|
||
write_sources(&sources, seed)?;
|
||
write_catalog(&catalog, seed, images)?;
|
||
std::fs::write(&stamp_path, serde_json::to_vec_pretty(&stamp)?)
|
||
.with_context(|| format!("writing {}", stamp_path.display()))?;
|
||
}
|
||
|
||
let catalog_bytes = std::fs::metadata(&catalog)
|
||
.with_context(|| format!("stat {}", catalog.display()))?
|
||
.len();
|
||
|
||
Ok((
|
||
Fixture {
|
||
dir: dir.to_path_buf(),
|
||
catalog,
|
||
thumbs: dir.join("thumbs"),
|
||
sources,
|
||
stamp,
|
||
catalog_bytes,
|
||
},
|
||
!usable,
|
||
))
|
||
}
|
||
|
||
fn matches_stamp(path: &Path, want: &Stamp) -> bool {
|
||
let Ok(text) = std::fs::read_to_string(path) else {
|
||
return false;
|
||
};
|
||
matches!(serde_json::from_str::<Stamp>(&text), Ok(have) if have == *want)
|
||
}
|
||
|
||
// ---------------------------------------------------------------------------
|
||
// The file pool
|
||
// ---------------------------------------------------------------------------
|
||
|
||
/// Write the pool of JPEGs the pixel half decodes.
|
||
///
|
||
/// Encoded through [`dr_thumbs::encode_rgba`] rather than a second encoder
|
||
/// call of this crate's own. That is the quality the store already uses (82),
|
||
/// which is a little below what a camera writes its previews at, and it is one
|
||
/// fewer place for an encoder setting to drift. Stated because it is visible
|
||
/// in the result: a slightly softer source decodes marginally faster than a
|
||
/// camera's own preview would.
|
||
fn write_sources(paths: &[PathBuf], seed: u64) -> Result<()> {
|
||
let (w, h) = PREVIEW;
|
||
for (i, path) in paths.iter().enumerate() {
|
||
let rgba = plausible_frame(w, h, seed ^ (i as u64));
|
||
let jpeg = dr_thumbs::encode_rgba(w, h, &rgba)
|
||
.map_err(|e| anyhow::anyhow!("encoding the fixture source {}: {e}", path.display()))?;
|
||
std::fs::write(path, jpeg).with_context(|| format!("writing {}", path.display()))?;
|
||
}
|
||
Ok(())
|
||
}
|
||
|
||
/// RGBA with detail at every scale: a coarse gradient plus a fine dither.
|
||
///
|
||
/// See this module's header for why neither a flat frame nor pure noise would
|
||
/// do. `salt` moves the gradient and the dither together so the twelve files
|
||
/// differ from one another rather than being twelve copies with different
|
||
/// names — a JPEG encoder that saw the same image twelve times would have the
|
||
/// same cache behaviour every time, which a library does not.
|
||
pub fn plausible_frame(w: u32, h: u32, salt: u64) -> Vec<u8> {
|
||
let mut rgba = vec![0u8; (w as usize) * (h as usize) * 4];
|
||
let bias = (salt % 97) as u32;
|
||
for y in 0..h as usize {
|
||
let row = y * (w as usize) * 4;
|
||
for x in 0..w as usize {
|
||
// A cheap integer hash, so neighbouring pixels differ and the
|
||
// encoder has real high-frequency content to spend bits on.
|
||
let n = (x.wrapping_mul(2_654_435_761) ^ y.wrapping_mul(1_640_531_527)) >> 13;
|
||
let dither = (n & 0x1f) as u32;
|
||
let gx = (x * 200 / (w as usize).max(1)) as u32;
|
||
let gy = (y * 55 / (h as usize).max(1)) as u32;
|
||
let px = &mut rgba[row + x * 4..row + x * 4 + 4];
|
||
px[0] = (30 + bias + gx + dither).min(255) as u8;
|
||
px[1] = (40 + gy + dither).min(255) as u8;
|
||
px[2] = (60 + gx / 2 + gy + dither).min(255) as u8;
|
||
px[3] = 255;
|
||
}
|
||
}
|
||
rgba
|
||
}
|
||
|
||
// ---------------------------------------------------------------------------
|
||
// The catalog
|
||
// ---------------------------------------------------------------------------
|
||
|
||
/// Write a catalog holding `images` rows, plus a default version for each.
|
||
///
|
||
/// The default versions are not decoration. `Catalog::open` backfills them for
|
||
/// any image that lacks one (see `schema::backfill`), so a fixture without
|
||
/// them would charge every measured open for fifty thousand inserts once and
|
||
/// nothing thereafter — a first number that bore no relation to the second,
|
||
/// and a benchmark whose result depended on whether it had been run before.
|
||
fn write_catalog(path: &Path, seed: u64, images: usize) -> Result<()> {
|
||
for suffix in ["", "-wal", "-shm"] {
|
||
let mut p = path.as_os_str().to_os_string();
|
||
p.push(suffix);
|
||
let _ = std::fs::remove_file(PathBuf::from(p));
|
||
}
|
||
|
||
let catalog = Catalog::open(path)
|
||
.map_err(|e| anyhow::anyhow!("creating the fixture catalog at {}: {e}", path.display()))?;
|
||
let conn = catalog.connection();
|
||
let mut rng = Rng::new(seed);
|
||
|
||
conn.execute_batch("BEGIN")?;
|
||
|
||
conn.execute(
|
||
"INSERT INTO roots(id, kind, label, last_seen, scan_generation)
|
||
VALUES (1, 'local', '/library', ?1, 1)",
|
||
rusqlite::params![EPOCH],
|
||
)?;
|
||
|
||
{
|
||
let mut folder = conn.prepare(
|
||
"INSERT INTO folders(id, root_id, parent_id, path, mtime, entry_count,
|
||
scanned_generation)
|
||
VALUES (?1, 1, NULL, ?2, ?3, ?4, 1)",
|
||
)?;
|
||
for f in 0..FOLDERS {
|
||
let id = f as i64 + 1;
|
||
let folder_path = format!("/library/{:04}/{:02}", 2016 + f / 12, f % 12 + 1);
|
||
let mtime = EPOCH + f as i64 * 86_400;
|
||
let entries = (images / FOLDERS.max(1)) as i64;
|
||
folder.execute(rusqlite::params![id, folder_path, mtime, entries])?;
|
||
}
|
||
}
|
||
|
||
{
|
||
let mut image = conn.prepare(
|
||
"INSERT INTO images(id, root_id, folder_id, source_ref, format, w, h,
|
||
captured_at, captured_offset, camera, lens, iso,
|
||
aperture, shutter, availability, file_size,
|
||
file_mtime, metadata_state, added_at)
|
||
VALUES (?1, 1, ?2, ?3, 'CR3', ?4, ?5, ?6, ?7, ?8, ?9, ?10, ?11, ?12,
|
||
?13, ?14, ?15, ?16, ?17)",
|
||
)?;
|
||
let mut version = conn.prepare(
|
||
"INSERT INTO versions(id, image_id, uuid, name, is_default, rating,
|
||
label, flag)
|
||
VALUES (?1, ?1, ?2, 'Original', 1, ?3, ?4, ?5)",
|
||
)?;
|
||
|
||
// Every value is bound to a local before it reaches `params!`. Not
|
||
// style: the macro takes a reference to each argument, and an
|
||
// expression like `TABLE[rng.below(n) as usize]` inside it borrows an
|
||
// element of a temporary array while `rng` is also being borrowed
|
||
// mutably. Locals make the evaluation order and the lifetimes obvious.
|
||
const OFFSETS: [i64; 5] = [0, 60, 120, -300, 540];
|
||
const ISOS: [i64; 7] = [100, 200, 400, 800, 1600, 3200, 6400];
|
||
const APERTURES: [f64; 6] = [1.4, 1.8, 2.8, 4.0, 5.6, 8.0];
|
||
const SHUTTERS: [f64; 6] = [0.004, 0.008, 0.0167, 0.005, 0.002, 0.5];
|
||
// Mostly metadata-only, as a large library on a laptop is: some
|
||
// previewed, a few with the original present.
|
||
const AVAILABILITY: [i64; 6] = [0, 0, 0, 1, 1, 2];
|
||
|
||
for i in 0..images {
|
||
let id = i as i64 + 1;
|
||
let bucket = i % FOLDERS;
|
||
let folder = bucket as i64 + 1;
|
||
let source_ref = format!(
|
||
"/library/{:04}/{:02}/IMG_{id:05}.CR3",
|
||
2016 + bucket / 12,
|
||
bucket % 12 + 1
|
||
);
|
||
let captured = EPOCH + rng.below(SPAN as u64) as i64;
|
||
// A quarter of the library shot in portrait, which is what makes
|
||
// the thumbnail path's orientation permutation a real cost rather
|
||
// than a branch that is never taken.
|
||
let (w, h) = if i % 4 == 3 {
|
||
(4000i64, 6000i64)
|
||
} else {
|
||
(6000i64, 4000i64)
|
||
};
|
||
// Two per cent still awaiting full EXIF — a library is never
|
||
// entirely finished being read, and the grid has to render that
|
||
// state (`metadata_state` 1).
|
||
let state: i64 = if rng.below(50) == 0 { 1 } else { 2 };
|
||
let camera = CAMERAS[rng.below(CAMERAS.len() as u64) as usize];
|
||
let lens = LENSES[rng.below(LENSES.len() as u64) as usize];
|
||
// Minutes east of UTC: a library shot in a handful of places.
|
||
let offset = OFFSETS[rng.below(OFFSETS.len() as u64) as usize];
|
||
let iso = ISOS[rng.below(ISOS.len() as u64) as usize];
|
||
let aperture = APERTURES[rng.below(APERTURES.len() as u64) as usize];
|
||
let shutter = SHUTTERS[rng.below(SHUTTERS.len() as u64) as usize];
|
||
let availability = AVAILABILITY[rng.below(AVAILABILITY.len() as u64) as usize];
|
||
let file_size = 20_000_000i64 + rng.below(30_000_000) as i64;
|
||
let file_mtime = captured + 60;
|
||
let added_at = captured + 3600;
|
||
|
||
image.execute(rusqlite::params![
|
||
id,
|
||
folder,
|
||
source_ref,
|
||
w,
|
||
h,
|
||
captured,
|
||
offset,
|
||
camera,
|
||
lens,
|
||
iso,
|
||
aperture,
|
||
shutter,
|
||
availability,
|
||
file_size,
|
||
file_mtime,
|
||
state,
|
||
added_at,
|
||
])?;
|
||
|
||
// Ratings skewed the way a culled library is: most unrated, a few
|
||
// picks, fewer still at five stars.
|
||
let rating: i64 = match rng.below(100) {
|
||
0..=69 => 0,
|
||
70..=84 => 1,
|
||
85..=93 => 2,
|
||
94..=97 => 3,
|
||
98 => 4,
|
||
_ => 5,
|
||
};
|
||
let flag: i64 = match rng.below(100) {
|
||
0..=79 => 0,
|
||
80..=94 => 1,
|
||
_ => 2,
|
||
};
|
||
let label: Option<i64> = match rng.below(100) {
|
||
0..=89 => None,
|
||
n => Some((n % 5) as i64 + 1),
|
||
};
|
||
let a = rng.next_u64();
|
||
let b = rng.next_u64();
|
||
let uuid = format!("{a:016x}{b:016x}");
|
||
version.execute(rusqlite::params![id, uuid, rating, label, flag])?;
|
||
}
|
||
}
|
||
|
||
conn.execute_batch("COMMIT")?;
|
||
|
||
// Deliberately no `ANALYZE`. The application never runs one, so a fixture
|
||
// that did would be measuring a query plan no user's catalog gets — and a
|
||
// plan chosen from statistics is exactly the sort of thing that would make
|
||
// the benchmark faster than the product.
|
||
|
||
// Dropping the connection checkpoints the WAL, so the file the next open
|
||
// reads is the whole catalog rather than a stub plus a journal.
|
||
drop(catalog);
|
||
Ok(())
|
||
}
|