Benchmarks / CPU and I/O (per commit) (push) Failing after 30s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 47s
Build and test / Layer separation (push) Successful in 27s
🐳 Android image / Build and push (push) Successful in 3s
Build and test / android-image (push) Successful in 4s
🐳 Windows image / Build and push (push) Successful in 1s
Build and test / windows-image (push) Successful in 1s
Traceability / Requirement traces (push) Failing after 38s
Build and test / Android (aarch64) (push) Failing after 2m24s
Build and test / Windows (x86_64, cross) (push) Failing after 3m21s
docs/manual/README.md is a tour for a photographer opening DarkRoom for the first time — one picture per thing, moving where movement is the point. tools/manual/ is how the pictures are made: drive.py puppeteers the desktop build on a private Xvfb (launch, click, drag, type, screenshot, record), scenes.py is each picture as a script, and record.sh runs them all over a folder and writes the results into docs/manual/media/. The media is in LFS, with the CI pulls excluding it as they exclude the fixtures; a screenshot changes wholesale when the interface does. Nothing in the pictures shows a person, by design: the demo library is seventy urban and alpine frames, chosen from the catalog's rows that face detection found nobody in. The traceability matrix is regenerated here after the rebase that brought this branch up to master.
194 lines
8.4 KiB
YAML
194 lines
8.4 KiB
YAML
name: Benchmarks
|
|
|
|
# The suite docs/requirements.md §8 has been promising since it was written:
|
|
# "an automated benchmark suite against a synthetic 50k catalog, run per-commit
|
|
# … A regression beyond stated tolerance fails the build."
|
|
#
|
|
# Its own workflow rather than a step inside build-and-test.yml, and the reason
|
|
# is what a failure here means. A red `Build and test` says the code is wrong; a
|
|
# red `Benchmarks` says the code is slower than it was, which is a different
|
|
# conversation, is read by different people, and must not be reachable by
|
|
# retrying a flaky compile.
|
|
#
|
|
# # Why this is split in two
|
|
#
|
|
# §8 names "the reference desktop", not CI, and it is right to. So:
|
|
#
|
|
# cpu — runs on every push. It needs no adapter and no display, and the
|
|
# budgets it asserts (a 50k catalog opening inside two seconds) have
|
|
# two orders of magnitude of headroom, so a modest runner can be held
|
|
# to them honestly. Machine-sensitive budgets — throughput targets
|
|
# written for a 24-thread desktop — are reported here rather than
|
|
# asserted; `dr-bench` decides that per metric and says so in its
|
|
# report. Asserting them on a two-core container would produce exactly
|
|
# what core/dr-gpu/tests/frame_budget.rs refused to produce: a red gate
|
|
# everybody learns to ignore.
|
|
#
|
|
# gpu — the frame budget, which already exists and already skips itself where
|
|
# there is no adapter. Not on push: it would build wgpu and naga on
|
|
# every commit to establish, every time, that this runner has no GPU. It
|
|
# runs on demand (Actions → Run workflow) so that a runner that *does*
|
|
# have one can be pointed at it, and the numbers it produces belong in
|
|
# docs/frame-budget.md by hand, as they already are.
|
|
|
|
on:
|
|
push:
|
|
branches: [main, master, develop]
|
|
pull_request:
|
|
branches: [main, master, develop]
|
|
workflow_dispatch:
|
|
|
|
jobs:
|
|
cpu:
|
|
runs-on: linux/amd64
|
|
name: CPU and I/O (per commit)
|
|
# Node for actions/checkout and actions/cache, which the bare runner image
|
|
# cannot execute. Rust is installed below.
|
|
container:
|
|
image: catthehacker/ubuntu:act-latest
|
|
|
|
env:
|
|
# Same reasoning as the desktop job in build-and-test.yml: incremental
|
|
# state exists to make the *second* build in a working tree fast, which is
|
|
# not a thing a fresh checkout has, and it fills the runner's disk.
|
|
CARGO_INCREMENTAL: 0
|
|
# The fixture, out of the workspace so actions/cache never picks it up.
|
|
# A 14 MB synthetic catalog is two seconds to regenerate and would
|
|
# otherwise be uploaded and downloaded on every push to save them.
|
|
DR_BENCH_DIR: /tmp/darkroom-bench
|
|
|
|
steps:
|
|
- name: Checkout
|
|
uses: actions/checkout@v4
|
|
|
|
# No `git lfs pull` here, deliberately. `dr-bench` depends on the catalog,
|
|
# the decoder, the thumbnail store and the encoder, and on nothing that
|
|
# reaches `dr-segment` — so the model this repository keeps in LFS is not
|
|
# part of this job's dependency graph and fetching it would be a minute
|
|
# spent on a file nothing opens.
|
|
|
|
- name: Cache cargo
|
|
uses: actions/cache@v4
|
|
with:
|
|
path: |
|
|
~/.cargo/registry
|
|
~/.cargo/git
|
|
target
|
|
key: bench-${{ runner.os }}-${{ hashFiles('**/Cargo.lock') }}
|
|
|
|
# Pinned to the workspace rust-version, as every other job here is: a
|
|
# floating toolchain turns an unrelated push into a mystery failure, and
|
|
# for a benchmark it would turn one into a mystery *regression*.
|
|
#
|
|
# rust-analyzer is named for the reason build-and-test.yml gives: rustup
|
|
# reconciles rust-toolchain.toml on the first cargo call whatever this
|
|
# step asks for, so naming it keeps the download inside the step that says
|
|
# it is installing things.
|
|
- name: Install Rust 1.92.0
|
|
run: |
|
|
set -e
|
|
curl -fsSL https://sh.rustup.rs | sh -s -- \
|
|
-y --no-modify-path --profile minimal --default-toolchain 1.92.0 \
|
|
--component rust-analyzer
|
|
echo "$HOME/.cargo/bin" >> "$GITHUB_PATH"
|
|
|
|
# `-p dr-bench`, not `--workspace`. The whole point of that crate having
|
|
# no GPU and no UI dependency is that this job resolves the catalog, the
|
|
# decoder and the encoders and stops there — a few minutes rather than the
|
|
# release build of Slint and wgpu the desktop job pays for.
|
|
#
|
|
# Release, and it is not optional: the workspace builds its own crates at
|
|
# opt-level = 0 in dev, and every figure this produces is dominated by
|
|
# this workspace's own code. A debug run would measure rustc.
|
|
- name: Build the suite
|
|
run: cargo build --release -p dr-bench
|
|
|
|
# Exit 1 is a violated budget or a regression past tolerance; exit 2 is
|
|
# the harness failing to run at all. Both fail the job, and the report
|
|
# above the failure says which.
|
|
- name: Measure, and gate
|
|
run: cargo run --release -p dr-bench -- check
|
|
|
|
- name: Disk after
|
|
if: always()
|
|
run: df -h /workspace 2>/dev/null || df -h .
|
|
|
|
gpu:
|
|
# On demand only — see the header. A runner with a Vulkan device can be
|
|
# pointed at this; one without will skip the measurement and say so, which
|
|
# is the same posture the rest of this repository's device tests take.
|
|
if: github.event_name == 'workflow_dispatch'
|
|
runs-on: linux/amd64
|
|
name: Frame budget (on demand)
|
|
container:
|
|
image: catthehacker/ubuntu:act-latest
|
|
|
|
env:
|
|
CARGO_INCREMENTAL: 0
|
|
|
|
steps:
|
|
- name: Checkout
|
|
uses: actions/checkout@v4
|
|
|
|
# `dr-gpu` depends on `dr-segment` for the watershed's pixel passes. Its
|
|
# default features are off, so no weights are compiled in — but the fetch
|
|
# is cheap insurance and its failure is not fatal. The header of the same
|
|
# step in build-and-test.yml explains why the extraheader is stripped
|
|
# rather than reused: two Authorization headers is a 400 from Gitea, one
|
|
# step after the batch call that had just succeeded.
|
|
- name: Fetch the segmentation model
|
|
continue-on-error: true
|
|
env:
|
|
LFS_TOKEN: ${{ secrets.GITEA_TOKEN || github.token }}
|
|
run: |
|
|
set -e
|
|
git lfs install --local
|
|
git config --local --get-regexp '^http\..*extraheader$' \
|
|
| cut -d' ' -f1 | sort -u \
|
|
| while read -r key; do git config --local --unset-all "$key"; done || true
|
|
git config --local lfs.url \
|
|
"https://x-access-token:${LFS_TOKEN}@gitea.tourolle.paris/dtourolle/DarkRoom.git/info/lfs"
|
|
git lfs pull --exclude="fixtures/**,docs/manual/media/**"
|
|
|
|
- name: Cache cargo
|
|
uses: actions/cache@v4
|
|
with:
|
|
path: |
|
|
~/.cargo/registry
|
|
~/.cargo/git
|
|
target
|
|
key: bench-gpu-${{ runner.os }}-${{ hashFiles('**/Cargo.lock') }}
|
|
|
|
- name: Build dependencies
|
|
run: |
|
|
apt-get update -qq
|
|
apt-get install -y -qq pkg-config libfontconfig1-dev libxkbcommon-dev
|
|
|
|
- name: Install Rust 1.92.0
|
|
run: |
|
|
set -e
|
|
curl -fsSL https://sh.rustup.rs | sh -s -- \
|
|
-y --no-modify-path --profile minimal --default-toolchain 1.92.0 \
|
|
--component rust-analyzer
|
|
echo "$HOME/.cargo/bin" >> "$GITHUB_PATH"
|
|
|
|
# The guard, in release. Its own module documentation is explicit that a
|
|
# release run checks strictly more than a dev one: the CPU half of a frame
|
|
# is shader-string assembly, which is several times slower unoptimised, so
|
|
# it is folded into the assertion only when debug_assertions is off.
|
|
#
|
|
# With no adapter this prints "skipping: no GPU adapter" and passes. A
|
|
# test that cannot run is not evidence either way, and turning that into a
|
|
# failure would make the job useless on the runner it usually lands on.
|
|
- name: Frame budget (FR-DSP-3)
|
|
run: cargo test --release -p dr-gpu --test frame_budget -- --nocapture
|
|
|
|
# The instrument behind docs/frame-budget.md. It exits non-zero with no
|
|
# adapter, which is right for a tool a person runs deliberately and wrong
|
|
# for a job that usually has none — hence continue-on-error. Its table is
|
|
# in the log for whoever asked for this run; the committed numbers are
|
|
# still updated by hand, as that file says.
|
|
- name: Frame budget table
|
|
continue-on-error: true
|
|
run: cargo run --release -p dr-gpu --example frame_budget
|