Initial implementation: core vertical slice
CI / fmt, clippy, test (push) Failing after 2m46s
CI / static musl binary (push) Has been skipped
CI / advisories and licences (push) Successful in 4m22s

Implements the core of SPEC.md — the manifest exchange, less audio-tier
matching (§3) and federation (§9a), both of which the spec sequences as
later work.

- §2 Jmanifest format and series bundles
- §3 cut matching: exact / runtime / loose tiers
- §4 API, less POST /manifests/search
- §5 rate limiting; §5a trust model, anonymous bearer tokens
- §6 upload validation, all four stages
- §7 relational storage, no JSON blob on the write path
- §8 Rust + Axum + SQLite, single serialized writer, in-process job queue
- §9a content addressing, computed on upload

Reconciled against the system spec:

- anneal_sec removed, withdrawn upstream by AR-012/AR-013. Presence follows
  track extent, so a track survives its own gaps and there is nothing to
  anneal. Its successor extinction_sec and the new gallery_scope are accepted
  and stored; scope enters the §7 ranking. A manifest still carrying
  anneal_sec is a hard 400, not silently ignored — it came from a pipeline
  whose window semantics differ from what this server assumes.
- Audio signature: media under 120 s now emits no signature at all, matching
  scene-actor-extraction IR-007. The earlier §3 draft allowed a shortened
  window under 150 s, which was the weaker rule — a caller-varying length is
  the property SR-004 forbids.
- UR IDs regularised to UR-nnn; docs/requirements.md registers 32
  requirements, each tracing to an SR-nnn or PR-nnn.

189 tests: unit, end-to-end through the real router, and an injection suite
covering SQL, JSON, header and Unicode payloads. Writing that suite found two
real gaps, both fixed here: compatibility homoglyphs passed the §5a character
class, and a one-frame audio signature was accepted on a feature-length item.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
2026-07-30 18:14:02 +02:00
co-authored by Claude Opus 5
commit a848750a65
38 changed files with 13014 additions and 0 deletions
+22
View File
@@ -0,0 +1,22 @@
# Keep the build context small and free of host state.
target/
.git/
.gitea/
# Never ship an operator's database or key into an image layer.
*.db
*.db-wal
*.db-shm
*.sqlite
*.sqlite3
.env
.env.*
# Not needed to build.
tests/
README.md
SPEC.md
deny.toml
Dockerfile
.dockerignore
.gitignore
+151
View File
@@ -0,0 +1,151 @@
# Gitea Actions CI.
#
# Gitea Actions is workflow-compatible with GitHub Actions, so this runs on either
# with no changes. It needs a registered runner with the `ubuntu-latest` label.
#
# The gates, in the order they fail fastest:
# fmt — formatting, seconds
# clippy — lints, denied rather than warned
# test — 160 unit + integration tests
# deny — RustSec advisories, licence policy, source policy
# musl — the artifact §8 actually ships: one static binary
name: CI
on:
push:
branches: [main, master]
pull_request:
# Advisories appear without any code changing, so the dependency audit also
# runs on a schedule rather than only on push.
schedule:
- cron: "0 6 * * 1"
env:
CARGO_TERM_COLOR: always
# Fail the build on warnings. The tree is warning-clean, so keeping it that way
# is cheaper than letting warnings accumulate.
RUSTFLAGS: "-D warnings"
jobs:
check:
name: fmt, clippy, test
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Install Rust
run: |
# rustup is not guaranteed present on a self-hosted Gitea runner.
if ! command -v rustup >/dev/null 2>&1; then
curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs \
| sh -s -- -y --profile minimal --component rustfmt,clippy
echo "$HOME/.cargo/bin" >> "$GITHUB_PATH"
else
rustup component add rustfmt clippy
fi
- name: Cache cargo
uses: actions/cache@v4
with:
path: |
~/.cargo/registry
~/.cargo/git
target
key: ${{ runner.os }}-cargo-${{ hashFiles('Cargo.lock') }}
restore-keys: ${{ runner.os }}-cargo-
- name: Formatting
run: cargo fmt --all -- --check
- name: Clippy
run: cargo clippy --all-targets --all-features
- name: Tests
run: cargo test --all-features
deny:
name: advisories and licences
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Install Rust
run: |
if ! command -v rustup >/dev/null 2>&1; then
curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs \
| sh -s -- -y --profile minimal
echo "$HOME/.cargo/bin" >> "$GITHUB_PATH"
fi
- name: Cache cargo-deny
uses: actions/cache@v4
with:
path: ~/.cargo/bin/cargo-deny
key: ${{ runner.os }}-cargo-deny
- name: Install cargo-deny
run: |
command -v cargo-deny >/dev/null 2>&1 || cargo install cargo-deny --locked
# Advisories, licences, bans and sources — see deny.toml for why the licence
# allow-list is closed rather than a deny-list.
- name: cargo deny
run: cargo deny check
musl:
name: static musl binary
runs-on: ubuntu-latest
# Only gate merges on the artifact build once the cheaper checks have passed.
needs: check
steps:
- uses: actions/checkout@v4
- name: Install Rust and musl target
run: |
if ! command -v rustup >/dev/null 2>&1; then
curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs \
| sh -s -- -y --profile minimal
echo "$HOME/.cargo/bin" >> "$GITHUB_PATH"
export PATH="$HOME/.cargo/bin:$PATH"
fi
rustup target add x86_64-unknown-linux-musl
sudo apt-get update && sudo apt-get install -y musl-tools
- name: Cache cargo
uses: actions/cache@v4
with:
path: |
~/.cargo/registry
~/.cargo/git
target
key: ${{ runner.os }}-musl-${{ hashFiles('Cargo.lock') }}
restore-keys: ${{ runner.os }}-musl-
# §8: "Ship a single static binary (musl target) plus the SQLite file."
# rusqlite is built with `bundled`, so SQLite is compiled in; reqwest uses
# rustls rather than OpenSSL, so there is no system TLS dependency to link.
- name: Build
run: cargo build --release --target x86_64-unknown-linux-musl
- name: Verify the binary is actually static
run: |
BIN=target/x86_64-unknown-linux-musl/release/jray-server
file "$BIN"
# A dynamically-linked result would defeat §8's deployment story, so this
# is asserted rather than assumed.
#
# Checked with `file`, not `ldd`: the musl target produces a static-PIE,
# and `ldd` prints the musl loader for one — an `ldd`-based check reports
# a perfectly static binary as dynamic.
if ! file "$BIN" | grep -qE 'static-pie linked|statically linked'; then
echo "::error::binary is not statically linked" >&2
exit 1
fi
- name: Upload binary
uses: actions/upload-artifact@v3
with:
name: jray-server-x86_64-musl
path: target/x86_64-unknown-linux-musl/release/jray-server
if-no-files-found: error
+35
View File
@@ -0,0 +1,35 @@
# Rust build artifacts
/target/
**/*.rs.bk
*.pdb
# Cargo.lock is committed: this crate ships a binary, so reproducible builds
# matter more than dependency-resolution freedom.
# SQLite database and its WAL sidecars (§7, §8). Never commit an operator's data;
# note that a plain file copy of a live WAL database is not a valid backup —
# use `VACUUM INTO` or the backup API.
*.db
*.db-wal
*.db-shm
*.sqlite
*.sqlite3
# Local operator configuration — holds the TMDB API key (§8).
.env
.env.*
!.env.example
# Python artefacts from the traceability tooling
__pycache__/
*.pyc
# Generated traceability output — regenerate with the gate, never hand-edit.
traces-report.json
# Editor / OS noise
.vscode/
.idea/
*.swp
*~
.DS_Store
Generated
+1778
View File
File diff suppressed because it is too large Load Diff
+30
View File
@@ -0,0 +1,30 @@
[package]
name = "jray-server"
version = "0.1.0"
edition = "2021"
rust-version = "1.85"
license = "GPL-3.0-or-later"
description = "JRay public server — community manifest exchange"
[dependencies]
axum = { version = "0.8", features = ["json", "query"] }
tokio = { version = "1", features = ["rt-multi-thread", "macros", "signal", "sync", "time"] }
tower = "0.5"
tower-http = { version = "0.6", features = ["trace", "timeout", "limit"] }
serde = { version = "1", features = ["derive"] }
serde_json = "1"
rusqlite = { version = "0.37", features = ["bundled"] }
reqwest = { version = "0.12", default-features = false, features = ["json", "rustls-tls"] }
tracing = "0.1"
tracing-subscriber = { version = "0.3", features = ["env-filter"] }
unicode-normalization = "0.1"
unicode-general-category = "1"
sha2 = "0.10"
rand = "0.9"
ulid = "1"
thiserror = "2"
anyhow = "1"
[dev-dependencies]
tower = { version = "0.5", features = ["util"] }
http-body-util = "0.1"
+89
View File
@@ -0,0 +1,89 @@
# Deploy image for jray-server.
#
# §8 is explicit that neither Docker nor Compose should be *required* — the
# primary artifact is a single static binary plus one database file, and that is
# deliberately the lowest-friction thing a hobbyist operator can deploy. This
# image is the optional convenience, not the intended path.
#
# It builds against musl so the runtime stage can be `scratch`: no libc, no shell,
# no package manager, nothing to keep patched. rusqlite is built with `bundled`
# (SQLite compiled in) and reqwest with rustls rather than OpenSSL, so there is
# genuinely nothing left to link against.
#
# docker build -t jray-server .
# docker run --rm -p 8080:8080 -v jray-data:/data \
# -e JRAY_TMDB_API_KEY=... jray-server
FROM rust:1.92-alpine AS builder
# `musl-dev` for the C toolchain rusqlite's bundled SQLite needs; `file` for the
# static-linkage assertion below.
RUN apk add --no-cache musl-dev file
WORKDIR /build
# Dependencies first, in their own layer, so editing source does not re-download
# and rebuild the entire tree.
COPY Cargo.toml Cargo.lock ./
RUN mkdir -p src \
&& echo 'fn main() {}' > src/main.rs \
&& echo '' > src/lib.rs \
&& cargo build --release --target x86_64-unknown-linux-musl \
&& rm -rf src
COPY src ./src
# `touch` defeats the cargo staleness check that the dummy-source trick above
# would otherwise leave in place.
RUN touch src/main.rs src/lib.rs \
&& cargo build --release --target x86_64-unknown-linux-musl \
&& strip target/x86_64-unknown-linux-musl/release/jray-server
# Verify the result is genuinely static. A dynamically-linked binary would fail at
# runtime on `scratch`, and failing here is far easier to diagnose.
#
# Asserted with `file`, not `ldd`: the musl target produces a **static-PIE**, and
# `ldd` prints the musl loader path for one, so an `ldd`-based check reports a
# static binary as dynamic. `file` reports "static-pie linked" and is unambiguous.
RUN file target/x86_64-unknown-linux-musl/release/jray-server | tee /tmp/linkage \
&& grep -qE 'static-pie linked|statically linked' /tmp/linkage \
|| (echo "binary is not statically linked; it will not run on scratch" && exit 1)
# Stage the data directory with the runtime uid's ownership. `scratch` has no
# shell, so this cannot be done in the final stage — and a bare `VOLUME` there
# would be created root-owned, leaving the non-root process unable to create the
# database at all.
RUN mkdir -p /staged-data && chown 65534:65534 /staged-data
# ---------------------------------------------------------------------------
FROM scratch
COPY --from=builder /build/target/x86_64-unknown-linux-musl/release/jray-server /jray-server
# The database lives on a volume; §8 warns that a plain file copy of a live WAL
# database is not a valid backup, so back it up with `VACUUM INTO` from the host
# rather than by archiving this directory.
#
# Copied from the builder so it arrives owned by the runtime uid. Docker seeds a
# named volume from the image's directory, ownership included, so the server can
# create the database on first run. A bind mount is *not* seeded this way — the
# host directory keeps its own ownership, so it must be made writable by uid
# 65534 (`chown 65534:65534 /path/on/host`).
COPY --from=builder --chown=65534:65534 /staged-data /data
VOLUME ["/data"]
# Non-root. `scratch` has no /etc/passwd, so this is a bare uid — which is all the
# kernel needs, and the binary touches nothing outside /data.
USER 65534:65534
ENV JRAY_BIND=0.0.0.0:8080 \
JRAY_DB=/data/jray.db
EXPOSE 8080
# No HEALTHCHECK: it would need a shell or curl, and `scratch` has neither.
# §4 provides `GET /health` (liveness) and `/ready` (database and migrations) for
# an orchestrator to probe externally, which is the right place for it.
ENTRYPOINT ["/jray-server"]
+220
View File
@@ -0,0 +1,220 @@
# JRay public server
A community manifest exchange for JRay. Jellyfin servers running the JRay plugin
pull actor-timeline manifests ("Jmanifests") for titles they own instead of
running the CV pipeline locally, and optionally contribute the manifests they
generate back.
See [SPEC.md](SPEC.md) for the design. Section references throughout the code
point at it.
The community instance is **`https://jray.tourolle.paris`**. The JRay plugin
ships with it pre-configured but **disabled** — §9 requires that no traffic leave
an installation until an admin opts in, so the default entry exists to save the
admin from typing a URL, not to enable sharing on their behalf. Set
`JRAY_SERVER_ID=jray.tourolle.paris` when deploying that instance: it becomes the
`origin` stamped on manifests it first accepts (§9a) and the salt for report IP
hashes.
## Status
First implementation pass: the **core vertical slice**, reconciled against the
[system spec](../SPEC.md).
Per-requirement status is in [`docs/requirements.md`](docs/requirements.md) —
32 requirements (`UR-001..018`, `DR-001..014`), each tracing up to an `SR-nnn`
or `PR-nnn`. `UR-015..018` are the pending SR-003 schema bump and are marked
`Planned` rather than omitted.
Implemented:
- §2 Jmanifest format and series bundles
- §3 cut matching — `exact` / `runtime` / `loose` tiers
- §4 the API surface, less `POST /manifests/search`
- §5 rate limiting, in-process fixed-window counters
- §5a trust model — anonymous bearer tokens, closed-vocabulary storage,
automatic revocation
- §6 upload validation, all four stages
- §7 relational storage, no JSON blobs on the write path
- §8 Rust + Axum + SQLite, single serialized writer, in-process job queue
- §9a content addressing (`content_id`), computed on upload
Reconciled with the system spec (see `docs/requirements.md` for the detail):
- **`anneal_sec` removed.** Withdrawn upstream by AR-012/AR-013 — presence now
follows track extent, so a track survives its own gaps and there is nothing to
anneal. Its successor `extinction_sec` and the new `gallery_scope` are
accepted, stored, and (for scope) used in §7 ranking. A manifest still
carrying `anneal_sec` is now a hard `400`, not silently ignored: it was
produced by a pipeline whose window semantics differ from what this server
assumes.
- **Audio signature: the 120 s rule now matches both producers.** An earlier
draft of §3 allowed a shortened window for items under 150 s; that conflicted
with `scene-actor-extraction` IR-007 and was the weaker rule, since a
caller-varying length is the property SR-004 forbids. Items under 120 s now
send no signature at all.
Deferred:
- §3 audio signatures — the field is **accepted, validated and stored**, and
`content_id` already excludes it, but `audio`-tier matching and
`POST /manifests/search` are not wired up. This follows §3's own recommended
sequencing: ship the plugin-side computation first, let signatures accumulate,
then enable matching once coverage is useful.
- §9a federation endpoints (`/federation/*`) and the pull worker. The schema
columns (`content_id`, `origin`, `ingested_from`, `peers`) are in place, and
`ingest::persist` is already the shared path a pull would reuse.
## Running
```sh
cargo run
```
Configuration is entirely environment variables:
| Variable | Default | Purpose |
|---|---|---|
| `JRAY_BIND` | `127.0.0.1:8080` | Listen address. Terminate TLS at the operator's proxy (§8) |
| `JRAY_DB` | `jray.db` | SQLite path. WAL mode, created on first run |
| `JRAY_TMDB_API_KEY` | — | **Hard dependency for UR-3.** Without it, uploads stay `pending` and are never listed |
| `JRAY_TMDB_BASE_URL` | `https://api.themoviedb.org/3` | Override for testing |
| `JRAY_TRUSTED_PROXIES` | — | Comma-separated proxy IPs whose `X-Forwarded-For` is honoured. **Not default-on**: §5 rate limiting and report attribution key on client IP, so a spoofable header defeats both |
| `JRAY_SERVER_ID` | `localhost` | This server's identity, used as manifest `origin` and as the report IP-hash salt |
| `JRAY_REQUEST_TIMEOUT_SEC` | `30` | Request timeout so a slow bundle query fails fast |
| `JRAY_JOB_BATCH` | `8` | Cast-check jobs leased per worker tick |
| `JRAY_JOB_POLL_SEC` | `5` | Worker poll interval |
| `JRAY_LOG` | `info` | `tracing` filter |
Contributing requires a token (§5a) — an anonymous bearer capability, not an
account. Self-issue one:
```sh
curl -sX POST -H 'content-type: application/json' -d '{}' \
http://127.0.0.1:8080/api/v1/tokens
```
## Deployment
§8's deployment notes are requirements, not suggestions:
- Enforce the body cap at **both** the proxy and the app. `client_max_body_size`
(nginx) / `request_body max_size` (Caddy) should match §6 stage 1, so oversized
uploads are dropped at the edge and never occupy an application worker. The app
must also be safe when run without a proxy, which it is.
- Set `JRAY_TRUSTED_PROXIES` to the proxy's address, or `X-Forwarded-For` is
ignored and every client behind it shares one rate-limit bucket.
- **Back up the SQLite file with `VACUUM INTO` or the backup API** — never a
plain file copy of a live WAL database. Manifests represent real CV compute.
## Tests
```sh
cargo test # 178 tests
cargo deny check # advisories, licences, bans, sources
```
Unit tests per module, plus two integration suites:
- `tests/api.rs` — end-to-end through the real router: status codes, headers, and
the properties that only hold if the layers compose correctly (per-route body
caps, rate-limit surfaces, the strict schema actually reaching uploads).
- `tests/injection.rs` — that hostile input cannot escape its layer: SQL payloads
in query parameters, path segments, JSON bodies, bearer tokens and report notes;
JSON structure abuse; CRLF header injection; path traversal; and Unicode tricks
against the §5a character class.
Security-relevant properties are asserted rather than assumed —
`movie`/`jellyfin_id` rejection, the §5a character class defeating base64/hex
smuggling, per-route body caps, a lying `Content-Length` not bypassing the cap,
and a forged `X-Forwarded-For` not resetting a rate-limit budget.
**On injection specifically.** Two independent defences apply, and they fail
differently, so both are tested:
1. **Parameterised queries.** Every value reaches SQLite through `params![]`. The
only `format!`-built SQL interpolates two compile-time constants (a column list
and a status literal) — no runtime input ever becomes SQL syntax. This is what
actually prevents injection.
2. **Closed-vocabulary validation.** Identifiers are regex-constrained and free
text is limited to a closed character class, so most payloads never reach the
query layer at all.
The injection suite would still pass on defence 1 alone, which is deliberate: if
validation were ever loosened, the tests should not silently start depending on it.
Writing that suite found two real gaps, both since fixed:
- Compatibility homoglyphs (`𝐒𝐭𝐞𝐯𝐞`, `Actor`) passed the §5a class. They are
letters by Unicode category and NFC does not fold them — only NFKC would. Beyond
name spoofing, a fullwidth-digit alphabet would have reopened the encoding
channel the "no digits" rule exists to close.
- A one-frame `audio_signature` was accepted on a feature-length manifest, making
the field the variable-length container §3 explicitly forbids. The length floor
is now derived from the declared runtime, keeping §3's genuine short-item
exception without trusting the client's length.
Two tests worth knowing about:
- `content_id::tests::golden_vector_hash_is_stable` locks the §9a canonical form.
**The JRay plugin must reproduce it byte-identically**; §8 notes the extraction
side is Python, so this can no longer be one shared implementation and must be
cross-tested instead. `GOLDEN_VECTORS` is that fixture, and its hash was
verified against an independent Python implementation.
- The integration tests use an on-disk temporary database, not `:memory:`,
because §8's topology is one writer connection plus a read pool — in-memory
SQLite is per-connection, so the readers would see an empty database.
## CI and container image
`.gitea/workflows/ci.yml` runs on Gitea Actions (and unmodified on GitHub
Actions), gated fastest-first: `fmt` → `clippy` → `test` → `cargo deny` → static
musl build. The dependency audit also runs weekly, since advisories appear without
any code changing.
The `Dockerfile` is the optional convenience, not the intended deployment path —
§8 is explicit that neither Docker nor Compose should be *required*. It builds
against musl and runs from `scratch` as uid 65534: **8.7 MB**, no libc, no shell,
no package manager. `rusqlite` bundles SQLite and `reqwest` uses rustls, so there
is nothing left to link.
```sh
docker build -t jray-server .
docker run --rm -p 8080:8080 -v jray-data:/data -e JRAY_TMDB_API_KEY=... jray-server
```
Two things that are easy to get wrong, so they are handled explicitly:
- **Static linkage is asserted with `file`, not `ldd`.** The musl target produces
a *static-PIE*, and `ldd` prints the musl loader for one — an `ldd`-based check
reports a perfectly static binary as dynamic.
- **`/data` is staged with the runtime uid's ownership.** Docker seeds a named
volume from the image directory, ownership included, so a non-root server can
create the database on first run. A **bind mount is not seeded this way** — the
host directory keeps its own ownership, so `chown 65534:65534` it first or the
server exits with "unable to open database file".
Two tests are worth knowing about:
- `content_id::tests::golden_vector_hash_is_stable` locks the §9a canonical form.
**The JRay plugin must reproduce it byte-identically**; §8 notes the extraction
side is Python, so this can no longer be one shared implementation and must be
cross-tested instead. `GOLDEN_VECTORS` is that fixture, and its hash was
verified against an independent Python implementation.
- The integration tests use an on-disk temporary database, not `:memory:`,
because §8's topology is one writer connection plus a read pool — in-memory
SQLite is per-connection, so the readers would see an empty database.
## Known gaps
- **The §6 stage-3 thresholds are still §10's guesses.** §10 (5) is explicit that
running the check over the 331 real corpus files would give the true
distribution of honest-upload match ratios, and is "the single cheapest way to
de-risk UR-3 and UR-5". The scoring logic is deliberately pure functions in
`castcheck.rs` so that retuning is a test-data exercise, not a code change.
- `POST /manifests/{id}/report` records reports but nothing consumes them yet.
§5a's divergence detection and the operator kill switch are not implemented;
delisting is currently a manual `UPDATE`, which §5a does note is the intended
shape ("one UPDATE", not a moderation queue).
- No admin surface. Revocation is automatic (§5a), but an operator has no
endpoint for the deliberate kill-switch case.
+1839
View File
File diff suppressed because it is too large Load Diff
+91
View File
@@ -0,0 +1,91 @@
# cargo-deny configuration.
#
# This crate is GPLv3 (it shares a licence with the JRay Jellyfin plugin), and it
# is a long-lived network service whose main risks are hostile input and operator
# friction (§8). So two checks matter most here:
#
# - `advisories` — a public-facing service must not ship known-vulnerable
# dependencies.
# - `licenses` — GPLv3 is compatible with permissive licences, but *not* with
# everything. A copyleft-incompatible dependency arriving transitively would
# be a licensing problem discovered far too late.
#
# Run with `cargo deny check`.
[graph]
# Check the targets an operator actually deploys. §8 ships a single static binary
# (musl target), so both glibc and musl Linux are in scope.
targets = [
"x86_64-unknown-linux-gnu",
"x86_64-unknown-linux-musl",
"aarch64-unknown-linux-gnu",
"aarch64-unknown-linux-musl",
]
all-features = true
[advisories]
version = 2
# Fail on any RustSec advisory. Unmaintained crates are a warning rather than an
# error: `sled` was rejected in §8 partly on maintenance grounds, so the signal is
# worth surfacing, but it should not break a build on its own.
yanked = "deny"
unmaintained = "workspace"
ignore = []
[licenses]
version = 2
# Permissive licences, all GPLv3-compatible. Deliberately a closed allow-list
# rather than a deny-list: a licence nobody vetted should stop the build, in the
# same spirit as §6's "no additional fields anywhere".
#
# Kept to licences actually present in the tree, so `cargo deny` stays quiet in
# CI and an added allowance is a visible decision. Adding a dependency that needs
# a new licence should be a deliberate edit here.
allow = [
"Apache-2.0",
"MIT",
"BSD-2-Clause",
"BSD-3-Clause",
"ISC",
"Zlib",
"Unicode-3.0",
# `webpki-roots` — Mozilla's trusted CA certificate set. This is a *data*
# licence, not a code licence, which is why it is not on the usual permissive
# list: the crate ships certificates rather than logic. CDLA-Permissive-2.0
# imposes no copyleft and no attribution burden on a binary that embeds it, so
# it is compatible with distributing this server under GPLv3.
#
# It arrives via reqwest's rustls stack, which §8's single static musl binary
# depends on (bundling roots is what lets the binary verify TLS without a
# system trust store).
"CDLA-Permissive-2.0",
# This crate's own licence.
"GPL-3.0-or-later",
]
confidence-threshold = 0.9
# `ring` ships a bespoke licence file that no SPDX expression describes; it is
# a permissive OpenSSL/ISC-style licence and is GPL-compatible. Clarify it rather
# than widening the allow-list.
[[licenses.clarify]]
crate = "ring"
expression = "MIT AND ISC AND OpenSSL"
license-files = [{ path = "LICENSE", hash = 0xbd0eed23 }]
[bans]
multiple-versions = "warn"
wildcards = "deny"
# Nothing is banned outright yet. The obvious future entries are alternative TLS
# stacks: reqwest is pinned to rustls (`default-features = false`) so that a
# static musl binary needs no system OpenSSL, and an accidental openssl-sys
# dependency would silently break that deployment story.
deny = []
skip = []
skip-tree = []
[sources]
unknown-registry = "deny"
unknown-git = "deny"
# Only crates.io. A git dependency in a service that hobbyist operators build
# from source is a supply-chain and reproducibility problem.
allow-registry = ["https://github.com/rust-lang/crates.io-index"]
allow-git = []
+200
View File
@@ -0,0 +1,200 @@
# JRay-public-server — requirements register
Stable IDs for every requirement in [`../SPEC.md`](../SPEC.md), which holds the
prose. This file is the **authoritative list**; the CI gate reads its
denominators from here (see the [system spec](../../SPEC.md) §6).
**IDs are permanent.** A withdrawn requirement is marked `Withdrawn` and its
number is never reused — renumbering is what produces orphan TRACES tags.
Tag code with `// TRACES: UR-003 | SR-004`.
| Type | Scope |
|---|---|
| `UR` | User/functional — what the server does |
| `DR` | Development — how it is built and operated |
| `UT` / `IT` | Unit / integration tests |
Status: `Done` · `In Progress` · `Planned` · `TBD` · `Withdrawn`
A requirement is `Done` only when it is implemented **and** has a test that
executes. Everything below runs in CI on any machine — this repo has no GPU
requirement and no fixture-generation step, unlike `scene-actor-extraction`.
---
## User requirements (UR)
| ID | Requirement | Traces to | Priority | Status |
|---|---|---|---|---|
| UR-001 | Cheap existence probe, separate from the fetch, returning availability and cut-match tier without payload | SR-001 | High | Done |
| UR-002 | Accept a contributed manifest for a media item | PR-006 | High | Done |
| UR-003 | Content verification: strict schema, size caps, approximate TMDB cast match | SR-004 | High | Done |
| UR-004 | Rate limiting, per token where present and per source IP otherwise | SR-004 | High | Done |
| UR-005 | Trust without accounts: not usable as a content store, nor for prank manifests | SR-004 | High | Done |
| UR-006 | Serve and accept a whole series in one operation | PR-006 | High | Done |
| UR-007 | Plugin queries an ordered, configurable list of servers | PR-005 | High | In Progress |
| UR-008 | Servers replicate manifests between each other | PR-006 | Medium | Planned |
| UR-009 | Store an audio spectral-peak signature for content-based identification | SR-003 | Medium | In Progress |
| UR-010 | Identity crossing the API boundary is TMDB/IMDB ids, never a name alone | SR-001 | High | Done |
| UR-011 | Reject any field capable of carrying binary or attacker-chosen content | SR-004 | High | Done |
| UR-012 | Never accept, store, or serve gallery data — reference faces or embeddings | SR-005 | High | Done |
| UR-013 | Windows are scene-scoped claims; never reinterpret their boundaries | SR-002 | High | Done |
| UR-014 | Reject an unknown `jmanifest_version` outright, never guess | SR-003 | High | Done |
### Notes on status
**UR-007 is `In Progress`, not `Done`.** The plugin now carries the ordered
server list and its per-server trust settings, with the community instance
pre-configured but disabled. The fetch path that consumes it does not exist yet.
**UR-009 is `In Progress`.** The server accepts, validates and stores
`cut.audio_signature`, and `content_id` correctly excludes it (§9a). What is
absent is `audio`-tier matching and `POST /manifests/search`. This is the
sequencing §3 recommends — accumulate signatures first, enable matching once
coverage is useful — not an oversight.
**UR-012 is satisfied structurally, by absence.** There is no field in the
Jmanifest capable of carrying an embedding or a crop, and no endpoint that would
accept one. Like PR-005 in the system spec, it cannot be verified by pointing at
code that does something; UT-024 verifies it by asserting that the obvious
attempts are rejected.
---
## Development requirements (DR)
| ID | Requirement | Traces to | Priority | Status |
|---|---|---|---|---|
| DR-001 | Strict parse boundary: unknown fields rejected structurally, not by validator code | SR-004 | High | Done |
| DR-002 | Fully relational storage — no JSON blob on the write path | SR-004 | High | Done |
| DR-003 | Single serialized writer connection, with a read pool alongside | PR-004 | High | Done |
| DR-004 | All database access behind a repository layer, not scattered through handlers | PR-004 | Medium | Done |
| DR-005 | Background work in-process, with the job queue as a table so it survives restart | PR-004 | High | Done |
| DR-006 | Rate-limit counters in process memory; no external counter store | PR-004 | Medium | Done |
| DR-007 | Ship a single static binary plus one database file; container optional | PR-004 | High | Done |
| DR-008 | `X-Forwarded-For` honoured only from explicitly configured proxies | SR-004 | High | Done |
| DR-009 | Body caps enforced while streaming, before parsing, per route | SR-004 | High | Done |
| DR-010 | Request bodies are UTF-8 only, rejected with a diagnosable error otherwise | SR-003 | Medium | Done |
| DR-011 | `content_id` canonical form is byte-stable and cross-implementation tested | SR-003 | High | Done |
| DR-012 | Dependency audit: advisories, licence policy, source policy | PR-004 | Medium | Done |
| DR-013 | API errors use the status codes the spec names, not the framework's defaults | SR-003 | Medium | Done |
| DR-014 | Portable SQL — no SQLite-specific form where a standard one exists | PR-004 | Medium | Done |
---
## Verification
**No GPU, no fixtures, no external services.** Every test here runs on any
machine in under three seconds. The TMDB dependency is the only external service,
and it is absent from tests by construction: an unconfigured client makes uploads
stay `pending`, which is the correct production failure mode (§8) and happens to
make the test suite hermetic.
| Tier | Runs in CI | What it covers |
|---|---|---|
| **T1 — unit** | Yes | Pure logic: validation, cut matching, cast-check scoring, canonicalisation, rate limiting |
| **T2 — integration** | Yes | End-to-end through the real router against a temporary on-disk database |
There is no tier that does not run. A requirement here is either verified or
visibly not.
**Integration tests use an on-disk temporary database, not `:memory:`.** DR-003
specifies one writer connection plus a read pool, and in-memory SQLite is
per-connection — the readers would see an empty database. Testing the real
topology is the point, so this is a deliberate choice rather than an oversight.
### Per-requirement verification
| ID | Tier | Test asserts | Edge cases covered |
|---|---|---|---|
| UR-001 | T2 | `exists` reports availability and tier without payload | Absent title returns `200` with `false`, not `404`; batch form is positional; one bad item does not fail the batch |
| UR-002 | T2 | Valid upload accepted as `202 pending` | Duplicate content deduplicates; same contributor resubmitting the same cut is `409` |
| UR-003 | T1 + T2 | Strict schema, caps, and cast-match thresholds | Ratio boundaries at 0.6 and 0.3 exactly; small-\|M\| all-but-one rule; missing TMDB credits flags rather than rejects |
| UR-004 | T1 + T2 | Limits engage and carry the documented headers | Window reset; a rejected request does not extend its own lockout; surfaces have independent budgets |
| UR-005 | T1 + T2 | Prank manifests rejected; no free-text channel | Uncredited cast rejected; name-only matches capped; automatic revocation needs a minimum sample |
| UR-006 | T2 | Bundle accepted per-episode, non-atomically | One bad episode rejected while its neighbours are accepted; envelope errors are whole-request `400` |
| UR-007 | — | *No server-side test.* Plugin-side; the register there will carry it | — |
| UR-009 | T1 | Signature structurally validated | Fixed length; reserved high bit; **media < 120 s must send no signature at all** |
| UR-010 | T1 + T2 | Actors persist as TMDB person ids | A name the upload invented does not round-trip |
| UR-011 | T2 | Every payload-shaped field rejected | base64, hex, markup, control characters, bidi overrides, compatibility homoglyphs |
| UR-012 | T2 | No endpoint accepts embeddings or image data | An `embedding` or `crop` field is an unknown-field `400` |
| UR-013 | T1 | Stored windows are byte-identical to those submitted | Adjacent windows never merged; a window is never trimmed to a shorter one |
| UR-014 | T1 | Unknown `jmanifest_version` rejected | Version `2` and version `0` both refused, naming the field |
| DR-001 | T1 | Unknown field at any nesting depth fails to parse | `movie` and `jellyfin_id` named in the error |
| DR-003 | T1 | Concurrent writes serialize rather than returning `SQLITE_BUSY` | Failed transaction rolls back fully |
| DR-005 | T1 | Jobs lease once, reschedule with backoff, survive restart | Stranded lease released at startup; future job not leased early |
| DR-008 | T2 | Forged `X-Forwarded-For` cannot mint a fresh budget | Untrusted peer ignored; trusted proxy honoured; client-supplied entries to the left cannot spoof |
| DR-009 | T2 | Oversized body rejected as `413` | **A lying `Content-Length` does not bypass the cap**; per-route limits differ |
| DR-010 | T1 | Non-UTF-8 rejected by name | UTF-16 with and without BOM; UTF-8 BOM; declared `charset=utf-16` |
| DR-011 | T1 | Canonical form is stable and order-independent | Accumulated float error hashes identically; `audio_signature` and `extraction` excluded; **golden vector verified against an independent Python implementation** |
| DR-013 | T1 | Schema mismatch is `400`, not the framework's `422` | §4 names `400` for a forbidden field, and a client checking for it would mishandle `422` |
Three are worth singling out, because each verifies a claim that would otherwise
be an assertion:
- **DR-009's lying-`Content-Length` case.** §6 stage 0 is explicit that the
header is a claim by the client, so the streaming cap is mandatory rather than
redundant. A test that only sends honest bodies verifies nothing.
- **DR-011's golden vector.** Two servers that validated the same upload must
reach the same `content_id`, and the plugin must reproduce it byte-identically
from a different language. The fixture is the only thing that can catch
divergence before it silently breaks federation deduplication.
- **UR-011's homoglyph case.** Writing this test found a real gap: compatibility
variants (`𝐒𝐭𝐞𝐯𝐞`, `Actor`) are letters by Unicode category and NFC does not
fold them, so a fullwidth-digit alphabet would have reopened the encoding
channel §5a's "no digits" rule closes.
---
## Withdrawn
| ID | Requirement | Reason |
|---|---|---|
| — | `anneal_sec` in `extraction` | Withdrawn upstream (`scene-actor-extraction` AR-012/AR-013): presence follows track extent, so a track survives its own gaps and there is nothing to anneal. Ships as part of the SR-003 bump |
Deleted rather than retained at zero: a field naming a mechanism the pipeline no
longer has is actively misleading to anyone reading a manifest, and would outlive
everyone who remembers why it is zero.
No `UR`/`DR` number was ever assigned to it — it was a *field*, not a
requirement — so nothing is orphaned by its removal.
---
## Pending — the SR-003 schema bump
These are `Planned` rather than absent, because the bump is coordinated across
three repos and this register should show the work rather than imply the server
is finished.
| ID | Requirement | Traces to | Priority | Status |
|---|---|---|---|---|
| UR-015 | Accept `extraction.extinction_sec` in place of `anneal_sec` | SR-003 | High | Planned |
| UR-016 | Accept and store `extraction.gallery_scope`; rank on it (§7) | SR-003 | Medium | Planned |
| UR-017 | Accept per-window belief and identification route; `scenes` becomes objects | SR-003 | High | Planned |
| UR-018 | Exclude belief from `content_id`, replicating it as an attribute | SR-003 | High | Planned |
**UR-018 is the one with a trap in it.** Belief is a producer-side estimate that
may legitimately differ between pipeline versions for identical timings, so
including it in the canonical form would give two servers different `content_id`s
for the same content — the exact failure mode §9a quantises centiseconds to
avoid. It follows `audio_signature`'s precedent: replicated as an attribute, not
part of identity.
---
## Notes on coverage
- **`DR-*` traces to `PR-004` (self-hosted) more often than to an `SR-nnn`.**
Operational simplicity is a single-repo concern serving the project goal
directly. This is correct rather than a gap: §8's whole argument for one binary
and one file is that federation only works if running an instance is easy.
- **UR-007 has no server-side test** and cannot have one — it is a requirement on
the plugin, recorded here because this spec is where it is stated. It should be
cross-referenced from the plugin's register when that is created, and until
then it is visibly unverified rather than quietly assumed.
- **UR-012 and UR-013 are preserved by prohibition**, like PR-005 in the system
spec. They cannot be verified by pointing at code that does something, only by
asserting that the attempts fail. They die the moment either prohibition is
relaxed, which is precisely why they are stated rather than left implicit.
+8
View File
@@ -0,0 +1,8 @@
# Formatting matches the style the codebase is written in.
#
# `use_small_heuristics = "Max"` keeps short structs, calls and match arms on one
# line rather than exploding them across four. In a codebase this dense with spec
# citations, vertical space spent on punctuation is space not spent on the comment
# explaining *why* a rule exists.
max_width = 100
use_small_heuristics = "Max"
+187
View File
@@ -0,0 +1,187 @@
//! `GET /manifests/exists` and its batch form — UR-1.
//!
//! Deliberately a *separate, cheaper* endpoint from the fetch: it answers
//! "should I bother?" for a whole library sweep without transferring payloads,
//! and it is the endpoint a scheduled task will hammer. It is also the most
//! abuse-prone surface, since it doubles as an oracle for "does the community
//! have this title" — so it is rate-limited harder than the fetches and returns
//! no manifest content (§0).
use axum::extract::{Query, State};
use axum::http::HeaderMap;
use axum::response::{IntoResponse, Response};
use axum::Json;
use serde::{Deserialize, Serialize};
use super::LookupParams;
use crate::db::repo;
use crate::error::{ApiError, ApiResult};
use crate::matching::{self, StoredCut};
use crate::model::{IdentityType, MatchTier};
use crate::ratelimit::Surface;
use crate::state::{with_quota_headers, AppState};
/// §4: no manifest content, just availability and tier.
#[derive(Debug, Clone, Serialize)]
pub struct ExistsResponse {
pub exists: bool,
#[serde(skip_serializing_if = "Option::is_none")]
pub r#match: Option<&'static str>,
#[serde(skip_serializing_if = "Option::is_none")]
pub manifest_id: Option<String>,
#[serde(skip_serializing_if = "Option::is_none")]
pub actor_count: Option<i64>,
}
impl ExistsResponse {
fn absent() -> Self {
Self { exists: false, r#match: None, manifest_id: None, actor_count: None }
}
}
/// §4 batch form: up to 100 items.
///
/// Exists specifically so the §5 rate limit can be generous per *request* while
/// staying strict per *item*, and so a 2000-item library sweep is 20 requests
/// rather than 2000.
pub const MAX_BATCH_ITEMS: usize = 100;
#[derive(Debug, Deserialize)]
#[serde(deny_unknown_fields)]
pub struct BatchRequest {
pub items: Vec<LookupParams>,
}
#[derive(Debug, Serialize)]
pub struct BatchResponse {
/// Positional, matching the request order (§4).
pub results: Vec<ExistsResponse>,
}
pub async fn exists(
State(state): State<AppState>,
peer: crate::state::PeerIp,
headers: HeaderMap,
Query(params): Query<LookupParams>,
) -> ApiResult<Response> {
let ip = state.client_ip(&headers, peer.0);
let quota = state.check_limit(&ip, Surface::ExistsSingle)?;
let body = lookup_one(&state, &params).await?;
Ok(with_quota_headers(Json(body).into_response(), quota))
}
pub async fn exists_batch(
State(state): State<AppState>,
peer: crate::state::PeerIp,
headers: HeaderMap,
super::json::Json(req): super::json::Json<BatchRequest>,
) -> ApiResult<Response> {
if req.items.len() > MAX_BATCH_ITEMS {
return Err(ApiError::BadRequest(format!(
"items: at most {MAX_BATCH_ITEMS} per request, got {}",
req.items.len()
)));
}
if req.items.is_empty() {
return Err(ApiError::BadRequest("items: must not be empty".into()));
}
let ip = state.client_ip(&headers, peer.0);
let quota = state.check_limit(&ip, Surface::ExistsBatch)?;
let mut results = Vec::with_capacity(req.items.len());
for item in &req.items {
// A malformed item yields "absent" rather than failing the whole batch —
// a sweep of 100 items should not be lost to one bad entry.
results.push(lookup_one(&state, item).await.unwrap_or_else(|_| ExistsResponse::absent()));
}
Ok(with_quota_headers(Json(BatchResponse { results }).into_response(), quota))
}
async fn lookup_one(state: &AppState, params: &LookupParams) -> ApiResult<ExistsResponse> {
let Some((kind, tmdb_id, imdb_id)) = resolve_kind(params) else {
return Err(ApiError::BadRequest(
"requires tmdb_id/imdb_id, or series_tmdb_id with season and episode".into(),
));
};
let (season, episode) = match kind {
IdentityType::Movie => (None, None),
IdentityType::Episode => (params.season, params.episode),
};
let client_cut = params.client_cut();
let found = state
.db
.read(move |conn| {
let Some(title) = repo::find_title(conn, kind, tmdb_id.as_deref(), imdb_id.as_deref())?
else {
return Ok(None);
};
let candidates = repo::candidates_for_title(conn, &title.id, season, episode)?;
if candidates.is_empty() {
return Ok(None);
}
let cuts: Vec<(String, StoredCut)> = candidates
.iter()
.map(|m| {
(
m.id.clone(),
StoredCut { runtime_sec: m.runtime_sec, video_hash: m.video_hash.clone() },
)
})
.collect();
let Some((id, m)) = matching::best_match(&client_cut, &cuts) else {
return Ok(None);
};
let actor_count = repo::manifest_actor_ids(conn, &id)?.len() as i64;
Ok(Some((id, m.tier, actor_count)))
})
.await
.map_err(ApiError::Internal)?;
// §4: `exists: false` is returned with `200`, not `404` — absence is a normal
// answer to this question, and `404` would conflate "no manifest" with "bad
// route" for the client.
Ok(match found {
Some((id, tier, actor_count)) => ExistsResponse {
exists: true,
// With no cut parameters the answer is "some manifest exists" with
// `"match": "unknown"`; the client must still fetch to find out
// whether a cut aligns. This is the mode a library sweep uses (§4).
r#match: Some(tier.as_str()),
manifest_id: Some(id),
actor_count: Some(actor_count),
},
None => ExistsResponse::absent(),
})
}
/// Determines whether these parameters address a movie or an episode.
pub fn resolve_kind(
params: &LookupParams,
) -> Option<(IdentityType, Option<String>, Option<String>)> {
if params.series_tmdb_id.is_some() || params.series_imdb_id.is_some() {
// Episode coordinates are required alongside series identity; without
// them the caller wants the series bundle endpoint instead.
params.season?;
params.episode?;
return Some((
IdentityType::Episode,
params.series_tmdb_id.clone(),
params.series_imdb_id.clone(),
));
}
if params.tmdb_id.is_some() || params.imdb_id.is_some() {
return Some((IdentityType::Movie, params.tmdb_id.clone(), params.imdb_id.clone()));
}
None
}
/// Exposed for tests asserting the documented tier string.
pub fn tier_str(t: MatchTier) -> &'static str {
t.as_str()
}
+351
View File
@@ -0,0 +1,351 @@
//! Manifest fetch endpoints (§4).
//!
//! §7: the submitted JSON was parsed, validated, resolved to TMDB person ids,
//! written as rows and discarded. Everything served here is **reconstructed**
//! from those rows, never echoed — which is what makes §5a's Threat 1 defence
//! structural rather than a promise.
use axum::extract::{Path, Query, State};
use axum::http::HeaderMap;
use axum::response::{IntoResponse, Response};
use axum::Json;
use serde::Serialize;
use super::LookupParams;
use crate::db::repo::{self, ManifestRow};
use crate::error::{ApiError, ApiResult};
use crate::matching::{self, StoredCut};
use crate::model::{
Actor, Coverage, Cut, Extraction, GalleryScope, Identity, IdentityType, Jmanifest, MatchTier,
SeriesBundle, SeriesRef, JMANIFEST_VERSION,
};
use crate::ratelimit::Surface;
use crate::state::{with_quota_headers, AppState};
#[derive(Debug, Serialize)]
pub struct FetchResponse {
pub r#match: &'static str,
/// Scene offset the client must add (§3). Zero for the tiers currently
/// served; present unconditionally so the plugin contract does not change
/// when `audio` is enabled.
pub offset_sec: f64,
pub manifest: Jmanifest,
}
#[derive(Debug, Serialize)]
pub struct StatusResponse {
pub status: String,
#[serde(skip_serializing_if = "Option::is_none")]
pub reason: Option<String>,
}
pub async fn get_movie(
State(state): State<AppState>,
peer: crate::state::PeerIp,
headers: HeaderMap,
Query(params): Query<LookupParams>,
) -> ApiResult<Response> {
let ip = state.client_ip(&headers, peer.0);
let quota = state.check_limit(&ip, Surface::ManifestFetch)?;
if params.tmdb_id.is_none() && params.imdb_id.is_none() {
return Err(ApiError::BadRequest("requires tmdb_id or imdb_id".into()));
}
let body = fetch_best(&state, IdentityType::Movie, &params, None, None).await?;
Ok(with_quota_headers(Json(body).into_response(), quota))
}
pub async fn get_episode(
State(state): State<AppState>,
peer: crate::state::PeerIp,
headers: HeaderMap,
Query(params): Query<LookupParams>,
) -> ApiResult<Response> {
let ip = state.client_ip(&headers, peer.0);
let quota = state.check_limit(&ip, Surface::ManifestFetch)?;
if params.series_tmdb_id.is_none() && params.series_imdb_id.is_none() {
return Err(ApiError::BadRequest("requires series_tmdb_id or series_imdb_id".into()));
}
let (Some(season), Some(episode)) = (params.season, params.episode) else {
return Err(ApiError::BadRequest("requires season and episode".into()));
};
let body =
fetch_best(&state, IdentityType::Episode, &params, Some(season), Some(episode)).await?;
Ok(with_quota_headers(Json(body).into_response(), quota))
}
async fn fetch_best(
state: &AppState,
kind: IdentityType,
params: &LookupParams,
season: Option<i64>,
episode: Option<i64>,
) -> ApiResult<FetchResponse> {
let (tmdb_id, imdb_id) = match kind {
IdentityType::Movie => (params.tmdb_id.clone(), params.imdb_id.clone()),
IdentityType::Episode => (params.series_tmdb_id.clone(), params.series_imdb_id.clone()),
};
let client_cut = params.client_cut();
let found = state
.db
.read(move |conn| {
let Some(title) = repo::find_title(conn, kind, tmdb_id.as_deref(), imdb_id.as_deref())?
else {
return Ok(None);
};
let candidates = repo::candidates_for_title(conn, &title.id, season, episode)?;
let cuts: Vec<(ManifestRow, StoredCut)> = candidates
.into_iter()
.map(|m| {
let cut =
StoredCut { runtime_sec: m.runtime_sec, video_hash: m.video_hash.clone() };
(m, cut)
})
.collect();
let Some((row, m)) = matching::best_match(&client_cut, &cuts) else {
return Ok(None);
};
let manifest = reconstruct(conn, &row, &title, kind)?;
Ok(Some((m.tier, m.offset_sec, manifest)))
})
.await
.map_err(ApiError::Internal)?;
// §4: `404` if none clears `loose`.
let (tier, offset_sec, manifest) = found.ok_or(ApiError::NotFound)?;
Ok(FetchResponse { r#match: tier.as_str(), offset_sec, manifest })
}
/// `GET /manifests/series/{series_tmdb_id}?season=` (§4).
///
/// Returns whatever episodes the server holds. **Partial bundles are normal** — a
/// bundle with 9 of 13 episodes is a valid, useful response, not an error (§2).
/// Episode-level cut matching is done client-side against the returned bundle,
/// since a client pulling a whole series already knows its own runtimes.
pub async fn get_series(
State(state): State<AppState>,
peer: crate::state::PeerIp,
headers: HeaderMap,
Path(series_tmdb_id): Path<String>,
Query(params): Query<LookupParams>,
) -> ApiResult<Response> {
let ip = state.client_ip(&headers, peer.0);
let quota = state.check_limit(&ip, Surface::SeriesFetch)?;
let season = params.season;
let bundle = state
.db
.read(move |conn| {
let Some(title) =
repo::find_title(conn, IdentityType::Episode, Some(&series_tmdb_id), None)?
else {
return Ok(None);
};
let rows = repo::episodes_for_series(conn, &title.id, season)?;
// Multiple contributors may hold the same episode; `episodes_for_series`
// orders by rank, so keep the first per (season, episode).
let mut episodes: Vec<Jmanifest> = Vec::new();
let mut seen: Vec<(i64, i64)> = Vec::new();
let mut seasons: Vec<i64> = Vec::new();
for row in rows {
let key = (row.season.unwrap_or(-1), row.episode.unwrap_or(-1));
if seen.contains(&key) {
continue;
}
seen.push(key);
if !seasons.contains(&key.0) {
seasons.push(key.0);
}
episodes.push(reconstruct(conn, &row, &title, IdentityType::Episode)?);
}
seasons.sort_unstable();
Ok(Some(SeriesBundle {
jmanifest_version: JMANIFEST_VERSION,
series: SeriesRef {
series_tmdb_id: title.tmdb_id.clone(),
series_imdb_id: title.imdb_id.clone(),
title: title.name.clone(),
},
coverage: Some(Coverage { episodes_available: episodes.len(), seasons }),
episodes,
}))
})
.await
.map_err(ApiError::Internal)?;
let bundle = bundle.filter(|b| !b.episodes.is_empty()).ok_or(ApiError::NotFound)?;
Ok(with_quota_headers(Json(bundle).into_response(), quota))
}
/// `GET /manifests/{id}` — fetch a specific manifest by its server-assigned id,
/// for debugging and for the "report this manifest" flow (§4).
pub async fn get_by_id(
State(state): State<AppState>,
peer: crate::state::PeerIp,
headers: HeaderMap,
Path(id): Path<String>,
) -> ApiResult<Response> {
let ip = state.client_ip(&headers, peer.0);
let quota = state.check_limit(&ip, Surface::ManifestFetch)?;
let manifest = state
.db
.read(move |conn| {
let Some(row) = repo::manifest_by_id(conn, &id)? else { return Ok(None) };
// Unlisted manifests are not served to anyone (§6 stage 3).
if row.status != "listed" && row.status != "flagged" {
return Ok(None);
}
let title = title_of(conn, &row.title_id)?;
let kind =
if title.kind == "movie" { IdentityType::Movie } else { IdentityType::Episode };
Ok(Some(reconstruct(conn, &row, &title, kind)?))
})
.await
.map_err(ApiError::Internal)?;
let manifest = manifest.ok_or(ApiError::NotFound)?;
Ok(with_quota_headers(Json(manifest).into_response(), quota))
}
/// `GET /manifests/{id}/status` — poll the outcome of the asynchronous cast
/// check (§4).
pub async fn get_status(
State(state): State<AppState>,
Path(id): Path<String>,
) -> ApiResult<Json<StatusResponse>> {
let found = state
.db
.read(move |conn| repo::manifest_status(conn, &id))
.await
.map_err(ApiError::Internal)?;
match found {
Some((status, reason)) => Ok(Json(StatusResponse { status, reason })),
// §6 deletes rejected manifests, so a vanished id is reported as
// rejected rather than as a bad route.
None => Ok(Json(StatusResponse {
status: "rejected".into(),
reason: Some("not_found_or_rejected".into()),
})),
}
}
fn title_of(conn: &rusqlite::Connection, title_id: &str) -> anyhow::Result<repo::TitleRow> {
let row = conn.query_row(
"SELECT id, kind, tmdb_id, imdb_id, name, year, adult, certification
FROM titles WHERE id = ?1",
rusqlite::params![title_id],
|r| {
Ok(repo::TitleRow {
id: r.get(0)?,
kind: r.get(1)?,
tmdb_id: r.get(2)?,
imdb_id: r.get(3)?,
name: r.get(4)?,
year: r.get(5)?,
adult: r.get::<_, i64>(6)? != 0,
certification: r.get(7)?,
})
},
)?;
Ok(row)
}
/// Rebuilds a Jmanifest from stored rows.
///
/// Names come from `people` — populated from TMDB by the server — so `name` is
/// server-authoritative on download and a name a contributor invented does not
/// round-trip (§2, §5a).
pub fn reconstruct(
conn: &rusqlite::Connection,
row: &ManifestRow,
title: &repo::TitleRow,
kind: IdentityType,
) -> anyhow::Result<Jmanifest> {
let stored = repo::actors_for_manifest(conn, &row.id)?;
let actors = stored
.into_iter()
.map(|a| Actor {
name: a.name,
imdb_id: None,
tmdb_id: Some(a.tmdb_person_id.to_string()),
scenes: a
.scenes_cs
.into_iter()
.map(|(s, e)| [s as f64 / 100.0, e as f64 / 100.0])
.collect(),
})
.collect();
let identity = match kind {
IdentityType::Movie => Identity {
kind,
tmdb_id: title.tmdb_id.clone(),
imdb_id: title.imdb_id.clone(),
series_tmdb_id: None,
series_imdb_id: None,
season: None,
episode: None,
title: title.name.clone(),
year: title.year,
},
IdentityType::Episode => Identity {
kind,
tmdb_id: None,
imdb_id: None,
series_tmdb_id: title.tmdb_id.clone(),
series_imdb_id: title.imdb_id.clone(),
season: row.season,
episode: row.episode,
title: title.name.clone(),
year: title.year,
},
};
// An unrecognised stored scope is served as absent rather than guessed at:
// the column is written from a closed enum, so anything else means the row
// predates a schema change and its meaning is unknown (UR-014's spirit).
let gallery_scope = match row.gallery_scope.as_deref() {
Some("global") => Some(GalleryScope::Global),
Some("limited") => Some(GalleryScope::Limited),
_ => None,
};
let extraction = Extraction {
sample_fps: row.sample_fps,
extinction_sec: row.extinction_sec,
pipeline_version: row.pipeline_version.clone(),
gallery_size: None,
gallery_scope,
};
let has_extraction = extraction.sample_fps.is_some()
|| extraction.extinction_sec.is_some()
|| extraction.pipeline_version.is_some()
|| extraction.gallery_scope.is_some();
Ok(Jmanifest {
jmanifest_version: JMANIFEST_VERSION,
identity,
cut: Cut {
runtime_sec: row.runtime_sec,
container_duration_sec: None,
video_hash: row.video_hash.clone(),
audio_signature: None,
},
extraction: has_extraction.then_some(extraction),
actors,
})
}
/// Exposed so tests can assert the served tier strings.
pub fn tier_name(t: MatchTier) -> &'static str {
t.as_str()
}
+287
View File
@@ -0,0 +1,287 @@
//! A JSON extractor that fails with the status codes §4 specifies.
//!
//! Axum's own `Json` rejects a body that parses as JSON but does not match the
//! target type with **422 Unprocessable Entity**. §4 is explicit that this case
//! is **`400`** — "malformed, or contains an unrecognised or forbidden field" —
//! and that distinction is load-bearing: §6 requires that a client which forgets
//! to strip `movie` or `jellyfin_id` gets "a hard `400` naming the offending
//! field". A client checking for 400 would mishandle a 422.
//!
//! This wrapper also guarantees the field name reaches the caller, since serde's
//! `deny_unknown_fields` error text is what identifies the offending key.
use axum::extract::{FromRequest, Request};
use axum::http::header::CONTENT_TYPE;
use crate::error::ApiError;
/// Drop-in replacement for `axum::Json` on request bodies.
pub struct Json<T>(pub T);
impl<T, S> FromRequest<S> for Json<T>
where
T: serde::de::DeserializeOwned,
S: Send + Sync,
{
type Rejection = ApiError;
async fn from_request(req: Request, state: &S) -> Result<Self, Self::Rejection> {
// A wrong content type is the client's mistake, reported as such rather
// than as a parse failure.
let content_type =
req.headers().get(CONTENT_TYPE).and_then(|v| v.to_str().ok()).unwrap_or("").to_string();
let mime = content_type.split(';').next().unwrap_or("").trim().to_ascii_lowercase();
if !(mime == "application/json" || mime.ends_with("+json")) {
return Err(ApiError::BadRequest("expected content-type: application/json".into()));
}
// A declared charset other than UTF-8 is refused up front, so the client
// learns what is wrong rather than receiving a confusing parse error from
// deep inside the document. See `require_utf8` for why UTF-8 is the only
// accepted encoding.
if let Some(charset) =
content_type.split(';').skip(1).filter_map(|p| p.trim().strip_prefix("charset=")).next()
{
let charset = charset.trim().trim_matches('"').to_ascii_lowercase();
if !matches!(charset.as_str(), "utf-8" | "utf8") {
return Err(ApiError::BadRequest(format!(
"unsupported charset {charset:?}: JSON must be UTF-8 encoded (RFC 8259 §8.1)"
)));
}
}
let bytes = axum::body::Bytes::from_request(req, state).await.map_err(|e| {
// §6 stage 1: the body cap aborts mid-transfer, and that must surface
// as `413`, not as a generic parse error. Axum folds the length-limit
// case into `FailedToBufferBody`, so the status it chose is the
// reliable discriminator.
if e.status() == axum::http::StatusCode::PAYLOAD_TOO_LARGE {
ApiError::PayloadTooLarge("request body exceeds the limit for this route".into())
} else {
ApiError::BadRequest(format!("could not read request body: {e}"))
}
})?;
// Encoding is checked before parsing, so a mis-encoded body gets an
// actionable message instead of whatever the parser happens to trip over.
let text = require_utf8(&bytes)?;
serde_json::from_str(text)
.map(Json)
// serde's message names the offending field, which is exactly what §6
// requires the response to identify.
.map_err(|e| ApiError::BadRequest(e.to_string()))
}
}
/// Enforces that the body is UTF-8, naming the encoding it appears to be.
///
/// **UTF-8 is the only accepted encoding, deliberately.** RFC 8259 §8.1 requires
/// it for JSON exchanged outside a closed ecosystem, and this is a public,
/// federated API. Three further reasons make it the right call *here*
/// specifically, rather than merely conventional:
///
/// 1. **§9a content addressing hashes bytes.** `content_id` is a SHA-256 over the
/// canonical form, so the same manifest submitted in two encodings would
/// produce two different ids — silently defeating federation deduplication.
/// That is precisely the failure mode §9a quantises scene times to avoid, and
/// it would be reintroduced at the encoding layer.
/// 2. **UTF-16 admits lone surrogates**, which have no UTF-8 representation. A
/// field able to carry them is a channel for bytes that survive validation but
/// are not text — against §5a's premise that no field can carry a payload.
/// 3. **§5a's character class assumes well-formed Unicode scalar values.** NFC
/// normalisation and the category checks are defined over scalars, so admitting
/// an encoding that can express non-scalars would undermine both.
///
/// serde_json would reject non-UTF-8 anyway; the value added here is a diagnosable
/// error rather than a misleading one. A UTF-16 body otherwise fails with "key
/// must be a string", which points an operator at the wrong problem entirely.
fn require_utf8(bytes: &[u8]) -> Result<&str, ApiError> {
// A BOM is not valid JSON (RFC 8259 §8.1: "implementations MUST NOT add a
// byte order mark"), and it is the clearest signal of an encoding mistake, so
// it is named rather than left to the parser.
let encoding_hint = match bytes {
[0xEF, 0xBB, 0xBF, ..] => Some("UTF-8 with a byte order mark"),
[0xFF, 0xFE, 0x00, 0x00, ..] => Some("UTF-32LE"),
[0x00, 0x00, 0xFE, 0xFF, ..] => Some("UTF-32BE"),
[0xFF, 0xFE, ..] => Some("UTF-16LE"),
[0xFE, 0xFF, ..] => Some("UTF-16BE"),
// Unmarked UTF-16 is the common case, since encoders often omit the BOM.
// A JSON document always begins with an ASCII character, so an
// interleaved NUL in the first two bytes is conclusive.
[0x00, b, ..] if b.is_ascii_graphic() => Some("UTF-16BE (no BOM)"),
[b, 0x00, ..] if b.is_ascii_graphic() => Some("UTF-16LE (no BOM)"),
_ => None,
};
if let Some(encoding) = encoding_hint {
return Err(ApiError::BadRequest(format!(
"request body appears to be {encoding}: JSON must be UTF-8 encoded \
without a byte order mark (RFC 8259 §8.1)"
)));
}
std::str::from_utf8(bytes).map_err(|e| {
ApiError::BadRequest(format!(
"request body is not valid UTF-8 at byte {}: JSON must be UTF-8 encoded \
(RFC 8259 §8.1)",
e.valid_up_to()
))
})
}
#[cfg(test)]
mod tests {
use super::*;
use axum::http::StatusCode;
use axum::response::IntoResponse;
#[derive(serde::Deserialize)]
#[serde(deny_unknown_fields)]
struct Probe {
_wanted: i64,
}
async fn extract(body: &'static str, content_type: Option<&str>) -> StatusCode {
let mut builder = Request::builder().method("POST").uri("/");
if let Some(ct) = content_type {
builder = builder.header(CONTENT_TYPE, ct);
}
let req = builder.body(axum::body::Body::from(body)).unwrap();
match Json::<Probe>::from_request(req, &()).await {
Ok(_) => StatusCode::OK,
Err(e) => e.into_response().status(),
}
}
#[tokio::test]
async fn schema_mismatch_is_400_not_422() {
// The whole reason this extractor exists (§4, §6).
assert_eq!(
extract(r#"{"unexpected":1}"#, Some("application/json")).await,
StatusCode::BAD_REQUEST
);
}
#[tokio::test]
async fn malformed_json_is_400() {
assert_eq!(extract("{ nope", Some("application/json")).await, StatusCode::BAD_REQUEST);
}
#[tokio::test]
async fn missing_content_type_is_400() {
assert_eq!(extract(r#"{"_wanted":1}"#, None).await, StatusCode::BAD_REQUEST);
}
#[tokio::test]
async fn content_type_parameters_are_tolerated() {
assert_eq!(
extract(r#"{"_wanted":1}"#, Some("application/json; charset=utf-8")).await,
StatusCode::OK
);
}
#[tokio::test]
async fn valid_body_extracts() {
assert_eq!(extract(r#"{"_wanted":1}"#, Some("application/json")).await, StatusCode::OK);
}
#[tokio::test]
async fn an_explicit_utf8_charset_is_accepted() {
for ct in [
"application/json; charset=utf-8",
"application/json;charset=UTF-8",
"application/json; charset=\"utf-8\"",
"application/json; charset=utf8",
] {
assert_eq!(extract(r#"{"_wanted":1}"#, Some(ct)).await, StatusCode::OK, "{ct}");
}
}
#[tokio::test]
async fn a_non_utf8_charset_is_refused_by_name() {
for ct in [
"application/json; charset=utf-16",
"application/json; charset=iso-8859-1",
"application/json; charset=windows-1252",
] {
assert_eq!(
extract(r#"{"_wanted":1}"#, Some(ct)).await,
StatusCode::BAD_REQUEST,
"{ct}"
);
}
}
/// Builds a request from raw bytes, since these bodies are not valid `&str`.
async fn extract_bytes(body: Vec<u8>) -> Result<(), ApiError> {
let req = Request::builder()
.method("POST")
.uri("/")
.header(CONTENT_TYPE, "application/json")
.body(axum::body::Body::from(body))
.unwrap();
Json::<Probe>::from_request(req, &()).await.map(|_| ())
}
#[tokio::test]
async fn utf16_bodies_are_rejected_with_an_actionable_message() {
// The reason this check exists: serde_json rejects UTF-16 anyway, but with
// "key must be a string", which points an operator at the wrong problem.
let doc = r#"{"_wanted":1}"#;
let le: Vec<u8> = doc.encode_utf16().flat_map(|u| u.to_le_bytes()).collect();
let err = extract_bytes(le).await.unwrap_err().to_string();
assert!(err.contains("UTF-16LE"), "should name the encoding: {err}");
assert!(err.contains("UTF-8"), "should say what is required: {err}");
let be: Vec<u8> = doc.encode_utf16().flat_map(|u| u.to_be_bytes()).collect();
let err = extract_bytes(be).await.unwrap_err().to_string();
assert!(err.contains("UTF-16BE"), "should name the encoding: {err}");
// With BOMs.
let mut le_bom = vec![0xFF, 0xFE];
le_bom.extend(doc.encode_utf16().flat_map(|u| u.to_le_bytes()));
assert!(extract_bytes(le_bom).await.is_err());
let mut be_bom = vec![0xFE, 0xFF];
be_bom.extend(doc.encode_utf16().flat_map(|u| u.to_be_bytes()));
assert!(extract_bytes(be_bom).await.is_err());
}
#[tokio::test]
async fn a_utf8_bom_is_rejected() {
// RFC 8259 §8.1: implementations MUST NOT add a byte order mark.
let mut body = vec![0xEF, 0xBB, 0xBF];
body.extend_from_slice(br#"{"_wanted":1}"#);
let err = extract_bytes(body).await.unwrap_err().to_string();
assert!(err.contains("byte order mark"), "{err}");
}
#[tokio::test]
async fn invalid_utf8_is_rejected_with_the_offending_offset() {
// A truncated multi-byte sequence inside an otherwise well-formed document.
let body = b"{\"_wanted\":\"\xC3\x28\"}".to_vec();
let err = extract_bytes(body).await.unwrap_err().to_string();
assert!(err.contains("not valid UTF-8"), "{err}");
assert!(err.contains("byte 12"), "should locate the failure: {err}");
}
#[tokio::test]
async fn valid_multibyte_utf8_is_accepted() {
// The check must not reject legitimate non-ASCII content — actor names are
// routinely non-Latin (§5a accepts any Unicode letter).
// Rejected for the unknown `_note` field, not for its encoding — which is
// the distinction being asserted.
let body = r#"{"_wanted":1,"_note":"宮崎 駿 Renée"}"#.as_bytes().to_vec();
let err = extract_bytes(body).await.unwrap_err().to_string();
assert!(err.contains("_note"), "should fail on the schema, not the encoding: {err}");
}
#[tokio::test]
async fn an_empty_body_is_not_mistaken_for_an_encoding_problem() {
let err = extract_bytes(Vec::new()).await.unwrap_err().to_string();
assert!(!err.contains("UTF-16"), "empty body is a parse error, not an encoding one: {err}");
}
}
+35
View File
@@ -0,0 +1,35 @@
//! HTTP surface (§4). Base path `/api/v1`, JSON throughout.
pub mod exists;
pub mod fetch;
pub mod json;
pub mod report;
pub mod upload;
use serde::Deserialize;
use crate::matching::ClientCut;
/// Identity + cut query parameters, shared by the read endpoints (§4).
#[derive(Debug, Clone, Default, Deserialize)]
pub struct LookupParams {
pub tmdb_id: Option<String>,
pub imdb_id: Option<String>,
pub series_tmdb_id: Option<String>,
pub series_imdb_id: Option<String>,
pub season: Option<i64>,
pub episode: Option<i64>,
pub runtime_sec: Option<f64>,
pub video_hash: Option<String>,
}
impl LookupParams {
pub fn client_cut(&self) -> ClientCut {
ClientCut {
// A non-finite or non-positive runtime is not a usable signal; treat
// it as absent rather than letting it drive a match.
runtime_sec: self.runtime_sec.filter(|r| r.is_finite() && *r > 0.0),
video_hash: self.video_hash.clone(),
}
}
}
+141
View File
@@ -0,0 +1,141 @@
//! `POST /manifests/{id}/report` (§4), and `GET /health`.
//!
//! Reports are a moderation lever and cheap to abuse, hence the tight §5 limit.
//! A report never changes `status` by itself: §5a keeps delisting an operator
//! action, because automatic delisting on report would hand any client a remote
//! delete primitive.
use axum::extract::{Path, State};
use axum::http::HeaderMap;
use axum::response::{IntoResponse, Response};
use axum::Json;
use serde::{Deserialize, Serialize};
use crate::db::repo;
use crate::error::{ApiError, ApiResult};
use crate::ratelimit::Surface;
use crate::state::{with_quota_headers, AppState};
use crate::worker::now_iso;
/// §4: `{ "reason": "misaligned" | "wrong_actors" | "spam", "note": "..." }`.
#[derive(Debug, Clone, Copy, PartialEq, Eq, Deserialize, Serialize)]
#[serde(rename_all = "snake_case")]
pub enum ReportReason {
Misaligned,
WrongActors,
Spam,
}
impl ReportReason {
fn as_str(self) -> &'static str {
match self {
ReportReason::Misaligned => "misaligned",
ReportReason::WrongActors => "wrong_actors",
ReportReason::Spam => "spam",
}
}
}
#[derive(Debug, Deserialize)]
#[serde(deny_unknown_fields)]
pub struct ReportRequest {
pub reason: ReportReason,
#[serde(default)]
pub note: Option<String>,
}
/// §5a: `note` is free text from an anonymous caller, so it is capped hard. It is
/// never served back to clients — only the operator reads it.
const MAX_NOTE_CHARS: usize = 500;
#[derive(Debug, Serialize)]
pub struct ReportAccepted {
pub report_id: String,
}
pub async fn post_report(
State(state): State<AppState>,
peer: crate::state::PeerIp,
headers: HeaderMap,
Path(manifest_id): Path<String>,
super::json::Json(req): super::json::Json<ReportRequest>,
) -> ApiResult<Response> {
let ip = state.client_ip(&headers, peer.0);
let quota = state.check_limit(&ip, Surface::Report)?;
let note = match req.note {
Some(n) if n.chars().count() > MAX_NOTE_CHARS => {
return Err(ApiError::BadRequest(format!(
"note: longer than {MAX_NOTE_CHARS} characters"
)))
}
// Strip control characters; the note is operator-facing text, not markup.
Some(n) => Some(n.chars().filter(|c| !c.is_control()).collect::<String>()),
None => None,
};
let ip_hash = crate::auth::hash_ip(&ip, &state.config.server_id);
let reason = req.reason.as_str();
let now = now_iso();
let id_for_check = manifest_id.clone();
let exists = state
.db
.read(move |c| Ok(repo::manifest_by_id(c, &id_for_check)?.is_some()))
.await
.map_err(ApiError::Internal)?;
if !exists {
return Err(ApiError::NotFound);
}
let report_id = state
.db
.write(move |tx| {
repo::insert_report(tx, &manifest_id, reason, note.as_deref(), &ip_hash, &now)
})
.await
.map_err(ApiError::Internal)?;
Ok(with_quota_headers(Json(ReportAccepted { report_id }).into_response(), quota))
}
#[derive(Debug, Serialize)]
pub struct Health {
pub status: &'static str,
pub version: &'static str,
}
/// `GET /health` — liveness, unauthenticated and unlimited (§4, §5).
pub async fn health() -> Json<Health> {
Json(Health { status: "ok", version: env!("CARGO_PKG_VERSION") })
}
#[derive(Debug, Serialize)]
pub struct Readiness {
pub status: &'static str,
pub database: &'static str,
/// §8: TMDB is a hard dependency for UR-3. If it is unconfigured, uploads
/// accumulate in `pending` rather than being listed unverified — worth
/// surfacing rather than failing silently.
pub tmdb_configured: bool,
}
/// Readiness check verifying the database opens and migrations are current (§8).
pub async fn ready(State(state): State<AppState>) -> ApiResult<Json<Readiness>> {
let ok = state
.db
.read(|conn| {
// Any query against a schema table proves both that the file opens
// and that migrations have been applied.
let n: i64 = conn.query_row("SELECT COUNT(*) FROM manifests", [], |r| r.get(0))?;
Ok(n >= 0)
})
.await
.map_err(ApiError::Internal)?;
Ok(Json(Readiness {
status: if ok { "ready" } else { "degraded" },
database: "ok",
tmdb_configured: state.tmdb.is_configured(),
}))
}
+241
View File
@@ -0,0 +1,241 @@
//! Contribution endpoints (§4) — UR-2 and UR-6.
//!
//! Both require a token (§5). Both return `202`: the upload has passed size and
//! schema validation and is held unlisted pending the asynchronous TMDB cast
//! check (§6 stage 3).
use axum::extract::State;
use axum::http::{HeaderMap, StatusCode};
use axum::response::{IntoResponse, Response};
use axum::Json;
use serde::Serialize;
use crate::db::repo;
use crate::error::{ApiError, ApiResult};
use crate::ingest::{self, IngestOutcome};
use crate::model::{IdentityType, Jmanifest, SeriesBundle};
use crate::ratelimit::Surface;
use crate::state::{with_quota_headers, AppState};
use crate::validate::{self, limits};
use crate::worker::now_iso;
#[derive(Debug, Serialize)]
pub struct UploadAccepted {
pub manifest_id: String,
pub status: &'static str,
}
/// `POST /manifests` — UR-2.
pub async fn post_manifest(
State(state): State<AppState>,
headers: HeaderMap,
super::json::Json(manifest): super::json::Json<Jmanifest>,
) -> ApiResult<Response> {
let contributor = state.require_contributor(&headers).await?;
// §5: limits are per token where one is present.
let quota = state.check_limit(&contributor.id, Surface::ManifestUpload)?;
// §6 stage 2. A rejection names the offending field, so a client that forgets
// to strip `movie`/`jellyfin_id` gets a diagnosable `400`.
let valid =
validate::validate_manifest(manifest).map_err(|e| ApiError::BadRequest(e.to_string()))?;
let origin = state.config.server_id.clone();
let contributor_id = contributor.id.clone();
let now = now_iso();
let outcome = state
.db
.write(move |tx| ingest::persist(tx, &valid, Some(&contributor_id), &origin, None, &now))
.await
.map_err(ApiError::Internal)?;
let resp = match outcome {
IngestOutcome::Pending { manifest_id } => {
(StatusCode::ACCEPTED, Json(UploadAccepted { manifest_id, status: "pending" }))
.into_response()
}
// §4 `409` — an identical `(identity, cut)` manifest already exists from
// this contributor.
IngestOutcome::DuplicateFromContributor { manifest_id } => {
return Err(ApiError::Conflict(format!(
"an identical manifest already exists from this contributor: {manifest_id}"
)))
}
// §9a: identical content already held, from any source. Not an error —
// the contributor's work is simply already represented.
IngestOutcome::DuplicateContent { manifest_id } => {
(StatusCode::OK, Json(UploadAccepted { manifest_id, status: "already_present" }))
.into_response()
}
};
Ok(with_quota_headers(resp, quota))
}
#[derive(Debug, Serialize)]
pub struct BundleResult {
pub season: Option<i64>,
pub episode: Option<i64>,
#[serde(skip_serializing_if = "Option::is_none")]
pub manifest_id: Option<String>,
pub status: &'static str,
#[serde(skip_serializing_if = "Option::is_none")]
pub reason: Option<String>,
}
#[derive(Debug, Serialize)]
pub struct BundleAccepted {
pub results: Vec<BundleResult>,
}
/// `POST /manifests/bundle` — UR-6.
///
/// **Per-episode validation, not atomic**: valid episodes are accepted and
/// invalid ones rejected, with a per-episode result list. All-or-nothing would let
/// one bad episode discard an entire season's compute (§2).
///
/// **One rate-limit unit**, so contributing a season is not punished relative to
/// contributing a film (§2, §5).
pub async fn post_bundle(
State(state): State<AppState>,
headers: HeaderMap,
super::json::Json(bundle): super::json::Json<SeriesBundle>,
) -> ApiResult<Response> {
let contributor = state.require_contributor(&headers).await?;
let quota = state.check_limit(&contributor.id, Surface::BundleUpload)?;
// §4: `413` for exceeding the episode cap, distinct from a malformed envelope.
if bundle.episodes.len() > limits::MAX_BUNDLE_EPISODES {
return Err(ApiError::PayloadTooLarge(format!(
"bundle carries {} episodes, limit is {}",
bundle.episodes.len(),
limits::MAX_BUNDLE_EPISODES
)));
}
// §4: `400` only for the envelope itself; individual bad episodes are
// reported in the results list, not as a whole-request error.
validate::validate_bundle_envelope(&bundle).map_err(|e| ApiError::BadRequest(e.to_string()))?;
let series_tmdb = bundle.series.series_tmdb_id.clone();
let mut results = Vec::with_capacity(bundle.episodes.len());
for episode in bundle.episodes {
let coords = (episode.identity.season, episode.identity.episode);
// An episode whose identity contradicts the envelope is rejected on its
// own rather than being silently reattributed to the bundle's series.
if episode.identity.kind != IdentityType::Episode {
results.push(BundleResult {
season: coords.0,
episode: coords.1,
manifest_id: None,
status: "rejected",
reason: Some("identity.type must be 'episode' within a bundle".into()),
});
continue;
}
if let (Some(envelope), Some(ep)) = (&series_tmdb, &episode.identity.series_tmdb_id) {
if envelope != ep {
results.push(BundleResult {
season: coords.0,
episode: coords.1,
manifest_id: None,
status: "rejected",
reason: Some("series_tmdb_id does not match the bundle envelope".into()),
});
continue;
}
}
let valid = match validate::validate_manifest(episode) {
Ok(v) => v,
Err(e) => {
results.push(BundleResult {
season: coords.0,
episode: coords.1,
manifest_id: None,
status: "rejected",
reason: Some(e.to_string()),
});
continue;
}
};
let origin = state.config.server_id.clone();
let contributor_id = contributor.id.clone();
let now = now_iso();
// One transaction per episode, so a bundle never holds the write lock for
// the whole request (§8 chunked ingest reasoning).
let outcome = state
.db
.write(move |tx| {
ingest::persist(tx, &valid, Some(&contributor_id), &origin, None, &now)
})
.await;
results.push(match outcome {
Ok(IngestOutcome::Pending { manifest_id }) => BundleResult {
season: coords.0,
episode: coords.1,
manifest_id: Some(manifest_id),
status: "pending",
reason: None,
},
Ok(IngestOutcome::DuplicateFromContributor { manifest_id })
| Ok(IngestOutcome::DuplicateContent { manifest_id }) => BundleResult {
season: coords.0,
episode: coords.1,
manifest_id: Some(manifest_id),
status: "already_present",
reason: None,
},
Err(e) => {
tracing::error!(error = ?e, "bundle episode failed to persist");
BundleResult {
season: coords.0,
episode: coords.1,
manifest_id: None,
status: "rejected",
reason: Some("internal error".into()),
}
}
});
}
let resp = (StatusCode::ACCEPTED, Json(BundleAccepted { results })).into_response();
Ok(with_quota_headers(resp, quota))
}
#[derive(Debug, Serialize)]
pub struct TokenIssued {
pub token: String,
}
/// Issues an anonymous bearer capability (§5a).
///
/// Self-issued on request: no email, no verification, no personal data. Stored
/// only as a hash, so the server cannot enumerate who holds tokens. Discarding a
/// token and requesting another is trivially easy — and that is fine, because the
/// token is not the defence; the content checks are.
pub async fn post_token(
State(state): State<AppState>,
peer: crate::state::PeerIp,
headers: HeaderMap,
) -> ApiResult<Json<TokenIssued>> {
let ip = state.client_ip(&headers, peer.0);
// Reuse the report budget: issuing tokens is cheap but should not be a free
// unbounded write.
state.check_limit(&ip, Surface::Report)?;
let token = crate::auth::generate_token();
let hash = crate::auth::hash_token(&token);
let now = now_iso();
state
.db
.write(move |tx| repo::insert_contributor(tx, &hash, &now))
.await
.map_err(ApiError::Internal)?;
Ok(Json(TokenIssued { token }))
}
+66
View File
@@ -0,0 +1,66 @@
//! Router construction.
//!
//! §6 stage 0/1 body caps are applied here as per-route `DefaultBodyLimit`
//! layers: Axum rejects on `Content-Length` before reading a body *and* caps the
//! stream for chunked or mis-declared uploads, which is what makes a lying
//! header and a chunked upload both safe. Per-route means the bundle endpoint
//! gets its larger limit without widening the others (§6 stage 1).
use std::time::Duration;
use axum::extract::DefaultBodyLimit;
use axum::routing::{get, post};
use axum::Router;
use tower_http::timeout::TimeoutLayer;
use tower_http::trace::TraceLayer;
use crate::api::{exists, fetch, report, upload};
use crate::state::AppState;
use crate::validate::limits;
/// Small cap for endpoints that take a short JSON body. A read endpoint has no
/// business accepting a large payload, and the batch `exists` form is bounded at
/// 100 items.
const SMALL_BODY_LIMIT: usize = 256 * 1024;
pub fn router(state: AppState) -> Router {
let timeout = state.config.request_timeout;
let v1 = Router::new()
// UR-1 — existence probes.
.route("/manifests/exists", get(exists::exists).post(exists::exists_batch))
// Reads.
.route("/manifests/movie", get(fetch::get_movie))
.route("/manifests/episode", get(fetch::get_episode))
.route("/manifests/series/{series_tmdb_id}", get(fetch::get_series))
.route("/manifests/{id}", get(fetch::get_by_id))
.route("/manifests/{id}/status", get(fetch::get_status))
.route("/manifests/{id}/report", post(report::post_report))
// UR-2 — contribution.
.route(
"/manifests",
post(upload::post_manifest).layer(DefaultBodyLimit::max(limits::BODY_LIMIT_MANIFEST)),
)
// UR-6 — whole-series contribution, with its own larger cap.
.route(
"/manifests/bundle",
post(upload::post_bundle).layer(DefaultBodyLimit::max(limits::BODY_LIMIT_BUNDLE)),
)
// §5a — anonymous bearer capability, not an account.
.route("/tokens", post(upload::post_token))
.layer(DefaultBodyLimit::max(SMALL_BODY_LIMIT));
Router::new()
.route("/health", get(report::health))
.route("/ready", get(report::ready))
.nest("/api/v1", v1)
// §8: a request timeout so a slow bundle query fails fast.
.layer(TimeoutLayer::with_status_code(axum::http::StatusCode::REQUEST_TIMEOUT, timeout))
.layer(TraceLayer::new_for_http())
.with_state(state)
}
/// Convenience for tests and `main`.
pub fn default_timeout() -> Duration {
Duration::from_secs(30)
}
+186
View File
@@ -0,0 +1,186 @@
//! §5a tokens and client-IP attribution.
//!
//! A token is **not an account** — it is an anonymous bearer capability. No
//! email, no verification, no personal data. It is stored only as a hash, so the
//! server cannot enumerate who holds tokens, and its sole purposes are
//! rate-limiting attribution (§5) and revocation.
//!
//! Discarding a token and requesting another is trivially easy, and that is
//! fine: the token is not the defence, the content checks are. Sybil resistance
//! is not required because identity is not load-bearing.
use std::net::IpAddr;
use axum::http::HeaderMap;
use sha2::{Digest, Sha256};
/// Hashes a bearer token for storage and lookup.
///
/// Plain SHA-256 rather than a password KDF is deliberate and sufficient here:
/// tokens are 256 bits of server-generated randomness, not user-chosen secrets,
/// so there is no dictionary to attack.
pub fn hash_token(token: &str) -> String {
let mut h = Sha256::new();
h.update(token.as_bytes());
hex(&h.finalize())
}
/// Hashes a client IP for report attribution (§7 `reports.source_ip_hash`).
///
/// Salted with the server id so hashes are not comparable across instances.
pub fn hash_ip(ip: &str, server_id: &str) -> String {
let mut h = Sha256::new();
h.update(server_id.as_bytes());
h.update(b"\0");
h.update(ip.as_bytes());
hex(&h.finalize())
}
fn hex(bytes: &[u8]) -> String {
let mut s = String::with_capacity(bytes.len() * 2);
for b in bytes {
s.push_str(&format!("{b:02x}"));
}
s
}
/// Generates a new token. Returned once to the caller; only its hash is stored.
pub fn generate_token() -> String {
use rand::RngCore;
let mut bytes = [0u8; 32];
rand::rng().fill_bytes(&mut bytes);
format!("jray_{}", hex(&bytes))
}
/// Extracts a bearer token from an `Authorization` header.
pub fn bearer_token(headers: &HeaderMap) -> Option<String> {
let raw = headers.get(axum::http::header::AUTHORIZATION)?.to_str().ok()?;
let (scheme, value) = raw.split_once(' ')?;
if !scheme.eq_ignore_ascii_case("bearer") {
return None;
}
let value = value.trim();
if value.is_empty() {
return None;
}
Some(value.to_string())
}
/// Resolves the client IP for rate-limiting and report attribution.
///
/// §8: the app must trust `X-Forwarded-For` **only** from the operator's proxy.
/// Rate limiting and report attribution key on client IP, so a spoofable header
/// defeats both — hence `trusted_proxies` is explicit configuration and an
/// untrusted peer's header is ignored outright.
pub fn client_ip(headers: &HeaderMap, peer: Option<IpAddr>, trusted_proxies: &[IpAddr]) -> String {
let peer_is_trusted = peer.is_some_and(|p| trusted_proxies.contains(&p));
if peer_is_trusted {
if let Some(xff) = headers.get("x-forwarded-for").and_then(|v| v.to_str().ok()) {
// Right-most entry is the one our trusted proxy appended; entries to
// its left are client-supplied and forgeable. Walk from the right
// past any further trusted hops.
for candidate in xff.split(',').rev().map(str::trim).filter(|s| !s.is_empty()) {
match candidate.parse::<IpAddr>() {
Ok(ip) if trusted_proxies.contains(&ip) => continue,
Ok(ip) => return ip.to_string(),
Err(_) => break,
}
}
}
}
peer.map(|p| p.to_string()).unwrap_or_else(|| "unknown".to_string())
}
#[cfg(test)]
mod tests {
use super::*;
use axum::http::HeaderValue;
fn headers(pairs: &[(&'static str, &str)]) -> HeaderMap {
let mut h = HeaderMap::new();
for (k, v) in pairs {
h.insert(*k, HeaderValue::from_str(v).unwrap());
}
h
}
#[test]
fn token_hash_is_stable_and_distinguishing() {
assert_eq!(hash_token("abc"), hash_token("abc"));
assert_ne!(hash_token("abc"), hash_token("abd"));
assert_eq!(hash_token("abc").len(), 64);
}
#[test]
fn generated_tokens_are_unique_and_prefixed() {
let a = generate_token();
let b = generate_token();
assert_ne!(a, b);
assert!(a.starts_with("jray_"));
assert_eq!(a.len(), 5 + 64);
}
#[test]
fn ip_hash_is_salted_per_server() {
// Hashes must not be comparable across instances.
assert_ne!(hash_ip("1.2.3.4", "a.example"), hash_ip("1.2.3.4", "b.example"));
assert_eq!(hash_ip("1.2.3.4", "a.example"), hash_ip("1.2.3.4", "a.example"));
}
#[test]
fn parses_bearer_tokens_case_insensitively() {
assert_eq!(
bearer_token(&headers(&[("authorization", "Bearer xyz")])).as_deref(),
Some("xyz")
);
assert_eq!(
bearer_token(&headers(&[("authorization", "bearer xyz")])).as_deref(),
Some("xyz")
);
assert!(bearer_token(&headers(&[("authorization", "Basic xyz")])).is_none());
assert!(bearer_token(&headers(&[("authorization", "Bearer ")])).is_none());
assert!(bearer_token(&HeaderMap::new()).is_none());
}
#[test]
fn forwarded_header_from_an_untrusted_peer_is_ignored() {
// The whole point of §8's explicit trusted-proxy configuration: an
// arbitrary client must not be able to choose its own rate-limit key.
let h = headers(&[("x-forwarded-for", "9.9.9.9")]);
let peer: IpAddr = "203.0.113.7".parse().unwrap();
assert_eq!(client_ip(&h, Some(peer), &[]), "203.0.113.7");
}
#[test]
fn forwarded_header_from_a_trusted_proxy_is_honoured() {
let h = headers(&[("x-forwarded-for", "9.9.9.9")]);
let proxy: IpAddr = "127.0.0.1".parse().unwrap();
assert_eq!(client_ip(&h, Some(proxy), &[proxy]), "9.9.9.9");
}
#[test]
fn client_supplied_entries_left_of_the_proxy_cannot_spoof() {
// A client that sends its own XFF gets its value appended to, not
// replaced, so only the right-most entry is trustworthy.
let h = headers(&[("x-forwarded-for", "9.9.9.9, 203.0.113.7")]);
let proxy: IpAddr = "127.0.0.1".parse().unwrap();
assert_eq!(client_ip(&h, Some(proxy), &[proxy]), "203.0.113.7");
}
#[test]
fn walks_past_additional_trusted_hops() {
let inner: IpAddr = "10.0.0.2".parse().unwrap();
let proxy: IpAddr = "127.0.0.1".parse().unwrap();
let h = headers(&[("x-forwarded-for", "203.0.113.7, 10.0.0.2")]);
assert_eq!(client_ip(&h, Some(proxy), &[proxy, inner]), "203.0.113.7");
}
#[test]
fn malformed_forwarded_value_falls_back_to_the_peer() {
let h = headers(&[("x-forwarded-for", "not-an-ip")]);
let proxy: IpAddr = "127.0.0.1".parse().unwrap();
assert_eq!(client_ip(&h, Some(proxy), &[proxy]), "127.0.0.1");
}
}
+484
View File
@@ -0,0 +1,484 @@
//! §6 stage 3 cast-match scoring, as pure functions.
//!
//! The thresholds here are the load-bearing part of UR-3 and §5a Threat 2, and
//! §10 (5) wants them retuned against the 331-file extraction corpus. Keeping
//! the decision logic free of I/O is what makes that a test-data exercise rather
//! than a code change.
use crate::tmdb::CastMember;
/// §6: thresholds over the ratio `|M ∩ C| / |M|`.
pub const LISTED_THRESHOLD: f64 = 0.6;
pub const FLAGGED_THRESHOLD: f64 = 0.3;
/// Below this size a ratio is meaningless (§6 small-|M| handling).
pub const SMALL_M_LIMIT: usize = 5;
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum Verdict {
/// Normal case.
Listed,
/// Served with reduced ranking, flagged for review.
Flagged,
/// Deleted, and the contributor's counter bumped.
Rejected,
}
impl Verdict {
pub fn status(self) -> &'static str {
match self {
Verdict::Listed => "listed",
Verdict::Flagged => "flagged",
Verdict::Rejected => "rejected",
}
}
}
#[derive(Debug, Clone)]
pub struct CastCheckOutcome {
pub verdict: Verdict,
pub ratio: f64,
/// Manifest actors resolved to a TMDB person id, with the TMDB-authoritative
/// name. Only these are kept; §6 drops unmatched actors rather than storing
/// them, which is what closes §5a's free-text channel.
pub matched: Vec<MatchedActor>,
/// Actors that matched nothing and will be dropped.
pub unmatched_person_ids: Vec<u64>,
pub reason: Option<String>,
}
#[derive(Debug, Clone)]
pub struct MatchedActor {
pub tmdb_person_id: u64,
/// From TMDB, never from the upload.
pub name: String,
pub adult: bool,
/// True when the match came from name comparison rather than an id.
pub by_name: bool,
}
/// An actor as submitted, after §6 stage 2 validation.
#[derive(Debug, Clone)]
pub struct SubmittedActor {
pub tmdb_id: Option<u64>,
pub imdb_id: Option<String>,
/// Used only for matching here, then discarded (§5a).
pub name: Option<String>,
}
/// Case- and accent-insensitive comparison key for the name fallback (§6).
fn name_key(s: &str) -> String {
use unicode_normalization::UnicodeNormalization;
s.nfd()
.filter(|c| !unicode_normalization::char::is_combining_mark(*c))
.flat_map(|c| c.to_lowercase())
.filter(|c| !c.is_whitespace() && *c != '.' && *c != ',' && *c != '-' && *c != '\'')
.collect()
}
/// Runs the §6 stage 3 comparison.
///
/// `credits` is the reference set *C*: for a movie, its credits; for an episode,
/// the union of per-episode credits and the series' aggregate credits.
pub fn evaluate(submitted: &[SubmittedActor], credits: &[CastMember]) -> CastCheckOutcome {
let m = submitted.len();
// §6: `|M| == 0` is rejected. These are extraction failures, not
// contributions — validation already refuses them, so reaching here means a
// manifest lost every actor upstream.
if m == 0 {
return CastCheckOutcome {
verdict: Verdict::Rejected,
ratio: 0.0,
matched: Vec::new(),
unmatched_person_ids: Vec::new(),
reason: Some("empty_actor_list".into()),
};
}
// TMDB has no credits for the id: absent data is not evidence of a bad
// manifest, so this is flagged rather than rejected (§6).
if credits.is_empty() {
return CastCheckOutcome {
verdict: Verdict::Flagged,
ratio: 0.0,
matched: Vec::new(),
unmatched_person_ids: submitted.iter().filter_map(|a| a.tmdb_id).collect(),
reason: Some("tmdb_no_credits".into()),
};
}
let mut matched: Vec<MatchedActor> = Vec::new();
let mut unmatched: Vec<u64> = Vec::new();
let mut id_matches = 0usize;
let mut name_matches = 0usize;
for actor in submitted {
// Join on `tmdb_id` — grounded in the pipeline's actual output, where
// 330 of 331 manifests have `imdb_id: ""` and `tmdb_id` set (§6).
let by_id = actor.tmdb_id.and_then(|id| credits.iter().find(|c| c.id == id));
if let Some(c) = by_id {
id_matches += 1;
push_unique(
&mut matched,
MatchedActor {
tmdb_person_id: c.id,
name: c.name.clone(),
adult: c.adult,
by_name: false,
},
);
continue;
}
// Fall back to case- and accent-insensitive name comparison.
let by_name = actor.name.as_deref().and_then(|n| {
let key = name_key(n);
(!key.is_empty()).then(|| credits.iter().find(|c| name_key(&c.name) == key))?
});
if let Some(c) = by_name {
name_matches += 1;
push_unique(
&mut matched,
MatchedActor {
tmdb_person_id: c.id,
name: c.name.clone(),
adult: c.adult,
by_name: true,
},
);
continue;
}
if let Some(id) = actor.tmdb_id {
unmatched.push(id);
}
}
// §6: name-only matches are counted but capped at half the intersection, so
// a manifest cannot pass on name collisions alone.
let capped_name_matches = name_matches.min(id_matches);
let effective = id_matches + capped_name_matches;
let ratio = effective as f64 / m as f64;
let verdict = classify(m, effective, ratio);
let reason = match verdict {
Verdict::Rejected => Some("cast_match_below_threshold".into()),
Verdict::Flagged => Some("cast_match_marginal".into()),
Verdict::Listed => None,
};
CastCheckOutcome { verdict, ratio, matched, unmatched_person_ids: unmatched, reason }
}
fn push_unique(matched: &mut Vec<MatchedActor>, actor: MatchedActor) {
if !matched.iter().any(|m| m.tmdb_person_id == actor.tmdb_person_id) {
matched.push(actor);
}
}
/// §6 small-*M* handling. With a median of 7 actors a ratio threshold is coarse
/// — one mismatch moves it by 14% — so small manifests use counts, not ratios.
fn classify(m: usize, matches: usize, ratio: f64) -> Verdict {
if m >= SMALL_M_LIMIT {
if ratio >= LISTED_THRESHOLD {
Verdict::Listed
} else if ratio >= FLAGGED_THRESHOLD {
Verdict::Flagged
} else {
Verdict::Rejected
}
} else if m >= 2 {
// Require all but one actor to match.
if matches + 1 >= m {
Verdict::Listed
} else {
Verdict::Rejected
}
} else {
// |M| <= 1: accept only if the single actor matches. Such a manifest is
// near-worthless anyway and is ranked last.
if matches >= 1 {
Verdict::Listed
} else {
Verdict::Rejected
}
}
}
/// §5a additional layer 1 — category guard.
///
/// Rejects when a matched person is flagged adult by TMDB and the target title
/// is not, which targets the stated prank without needing a blocklist of names.
pub fn category_guard_violation(matched: &[MatchedActor], title_is_adult: bool) -> Option<u64> {
if title_is_adult {
return None;
}
matched.iter().find(|m| m.adult).map(|m| m.tmdb_person_id)
}
/// §5a additional layer 2 — age-appropriateness guard.
///
/// On a children's certification, apply the strictest cast-match threshold and
/// require an `exact` or `runtime` cut match. Mismatched content on children's
/// titles is the highest-harm case and deserves the tightest gate.
pub fn is_childrens_certification(cert: &str) -> bool {
matches!(
cert.trim().to_ascii_uppercase().as_str(),
"G" | "TV-Y" | "TV-Y7" | "TV-G" | "U" | "0+" | "6+" | "PG" | "TV-PG"
)
}
pub const CHILDRENS_LISTED_THRESHOLD: f64 = 0.8;
/// Applies the children's-title gate to an already-computed outcome.
pub fn apply_childrens_guard(outcome: &mut CastCheckOutcome, m: usize) {
if m >= SMALL_M_LIMIT && outcome.ratio < CHILDRENS_LISTED_THRESHOLD {
outcome.verdict = match outcome.verdict {
Verdict::Listed => Verdict::Flagged,
other => other,
};
if outcome.reason.is_none() {
outcome.reason = Some("childrens_title_strict_threshold".into());
}
}
}
#[cfg(test)]
mod tests {
use super::*;
fn credit(id: u64, name: &str) -> CastMember {
CastMember { id, name: name.to_string(), adult: false }
}
fn adult_credit(id: u64, name: &str) -> CastMember {
CastMember { id, name: name.to_string(), adult: true }
}
fn by_id(id: u64) -> SubmittedActor {
SubmittedActor { tmdb_id: Some(id), imdb_id: None, name: None }
}
fn by_name(name: &str) -> SubmittedActor {
SubmittedActor { tmdb_id: None, imdb_id: None, name: Some(name.to_string()) }
}
/// A realistic reference cast — feature casts are several times larger than
/// the manifests extracted from them (§6).
fn cast_of_20() -> Vec<CastMember> {
(1..=20).map(|i| credit(i, &format!("Actor {i}"))).collect()
}
#[test]
fn full_subset_of_the_cast_is_listed() {
// §6: the ratio is over *M*, not *C* — a manifest legitimately contains
// only actors both credited and detected on screen, so penalising it for
// missing credited actors would fail every honest upload.
let submitted: Vec<_> = (1..=7).map(by_id).collect();
let out = evaluate(&submitted, &cast_of_20());
assert_eq!(out.verdict, Verdict::Listed);
assert_eq!(out.ratio, 1.0);
assert_eq!(out.matched.len(), 7);
}
#[test]
fn threshold_boundaries_at_point_six_and_point_three() {
// 6 of 10 matching == 0.6 exactly: listed.
let mut submitted: Vec<_> = (1..=6).map(by_id).collect();
submitted.extend((900..904).map(by_id));
let out = evaluate(&submitted, &cast_of_20());
assert_eq!(out.matched.len(), 6);
assert!((out.ratio - 0.6).abs() < 1e-9);
assert_eq!(out.verdict, Verdict::Listed);
// 5 of 10 == 0.5: flagged, served with reduced ranking.
let mut submitted: Vec<_> = (1..=5).map(by_id).collect();
submitted.extend((900..905).map(by_id));
let out = evaluate(&submitted, &cast_of_20());
assert_eq!(out.verdict, Verdict::Flagged);
// 3 of 10 == 0.3 exactly: still flagged, not rejected.
let mut submitted: Vec<_> = (1..=3).map(by_id).collect();
submitted.extend((900..907).map(by_id));
let out = evaluate(&submitted, &cast_of_20());
assert_eq!(out.verdict, Verdict::Flagged);
// 2 of 10 == 0.2: rejected.
let mut submitted: Vec<_> = (1..=2).map(by_id).collect();
submitted.extend((900..908).map(by_id));
let out = evaluate(&submitted, &cast_of_20());
assert_eq!(out.verdict, Verdict::Rejected);
}
#[test]
fn prank_manifest_is_rejected() {
// §5a Threat 2: performers who are not credited cast on the title.
let submitted: Vec<_> = (500..510).map(by_id).collect();
let out = evaluate(&submitted, &cast_of_20());
assert_eq!(out.verdict, Verdict::Rejected);
assert_eq!(out.ratio, 0.0);
assert_eq!(out.reason.as_deref(), Some("cast_match_below_threshold"));
}
#[test]
fn small_m_requires_all_but_one_to_match() {
// §6: `2 <= |M| < 5` — a ratio is meaningless at this size.
let out = evaluate(&[by_id(1), by_id(2), by_id(3), by_id(999)], &cast_of_20());
assert_eq!(out.verdict, Verdict::Listed, "3 of 4 is all-but-one");
let out = evaluate(&[by_id(1), by_id(2), by_id(998), by_id(999)], &cast_of_20());
assert_eq!(out.verdict, Verdict::Rejected, "2 of 4 fails all-but-one");
// 0.5 would be `Flagged` under the ratio table, so this proves the
// small-|M| branch is actually taken.
let out = evaluate(&[by_id(1), by_id(999)], &cast_of_20());
assert_eq!(out.verdict, Verdict::Listed, "1 of 2 is all-but-one");
}
#[test]
fn single_actor_manifest_needs_that_actor_to_match() {
assert_eq!(evaluate(&[by_id(1)], &cast_of_20()).verdict, Verdict::Listed);
assert_eq!(evaluate(&[by_id(999)], &cast_of_20()).verdict, Verdict::Rejected);
}
#[test]
fn empty_manifest_is_rejected() {
let out = evaluate(&[], &cast_of_20());
assert_eq!(out.verdict, Verdict::Rejected);
assert_eq!(out.reason.as_deref(), Some("empty_actor_list"));
}
#[test]
fn missing_tmdb_credits_flags_rather_than_rejects() {
// §6: absent data is not evidence of a bad manifest.
let submitted: Vec<_> = (1..=7).map(by_id).collect();
let out = evaluate(&submitted, &[]);
assert_eq!(out.verdict, Verdict::Flagged);
assert_eq!(out.reason.as_deref(), Some("tmdb_no_credits"));
}
#[test]
fn name_matching_is_case_and_accent_insensitive() {
let credits = vec![credit(1, "Renée Zellweger"), credit(2, "Miloš Forman")];
let out = evaluate(&[by_name("renee zellweger"), by_name("MILOS FORMAN")], &credits);
assert_eq!(out.matched.len(), 2);
}
#[test]
fn name_only_matches_cannot_carry_a_manifest_alone() {
// §6: name-only matches are capped at half the intersection, so a
// manifest cannot pass on name collisions alone.
let credits: Vec<_> = (1..=20).map(|i| credit(i, &format!("Actor {i}"))).collect();
let submitted: Vec<_> = (1..=10).map(|i| by_name(&format!("Actor {i}"))).collect();
let out = evaluate(&submitted, &credits);
assert_eq!(out.ratio, 0.0, "with no id matches, name matches cap to zero");
assert_eq!(out.verdict, Verdict::Rejected);
}
#[test]
fn name_matches_count_up_to_the_number_of_id_matches() {
let credits: Vec<_> = (1..=20).map(|i| credit(i, &format!("Actor {i}"))).collect();
// 4 by id + 6 by name, of 10 => capped to 4 + 4 = 8 => 0.8.
let mut submitted: Vec<_> = (1..=4).map(by_id).collect();
submitted.extend((5..=10).map(|i| by_name(&format!("Actor {i}"))));
let out = evaluate(&submitted, &credits);
assert!((out.ratio - 0.8).abs() < 1e-9, "got {}", out.ratio);
assert_eq!(out.verdict, Verdict::Listed);
}
#[test]
fn unmatched_actors_are_reported_for_dropping() {
// §6: unmatched actors are dropped rather than stored.
let submitted: Vec<_> = (1..=6).map(by_id).chain([by_id(777)]).collect();
let out = evaluate(&submitted, &cast_of_20());
assert_eq!(out.unmatched_person_ids, vec![777]);
assert!(out.matched.iter().all(|m| m.tmdb_person_id != 777));
}
#[test]
fn matched_names_come_from_tmdb_not_the_upload() {
// §5a: the server stores references to TMDB entities, not
// attacker-authored text.
let credits = vec![credit(884, "Steve Buscemi")];
let submitted = vec![SubmittedActor {
tmdb_id: Some(884),
imdb_id: None,
name: Some("Definitely Not Him".into()),
}];
let out = evaluate(&submitted, &credits);
assert_eq!(out.matched[0].name, "Steve Buscemi");
}
#[test]
fn duplicate_credits_do_not_double_count() {
// TMDB aggregate credits can list a person more than once.
let credits = vec![credit(1, "A"), credit(1, "A")];
let out = evaluate(&[by_id(1)], &credits);
assert_eq!(out.matched.len(), 1);
}
#[test]
fn category_guard_catches_adult_performers_on_a_non_adult_title() {
// §5a layer 1, aimed squarely at the stated prank.
let matched = vec![
MatchedActor { tmdb_person_id: 1, name: "A".into(), adult: false, by_name: false },
MatchedActor { tmdb_person_id: 2, name: "B".into(), adult: true, by_name: false },
];
assert_eq!(category_guard_violation(&matched, false), Some(2));
// Unless the target title is itself flagged adult.
assert_eq!(category_guard_violation(&matched, true), None);
}
#[test]
fn category_guard_ignores_clean_casts() {
let matched = vec![MatchedActor {
tmdb_person_id: 1,
name: "A".into(),
adult: false,
by_name: false,
}];
assert_eq!(category_guard_violation(&matched, false), None);
}
#[test]
fn adult_credit_is_carried_through_matching() {
let out = evaluate(&[by_id(9)], &[adult_credit(9, "X")]);
assert!(out.matched[0].adult);
}
#[test]
fn childrens_certifications_are_recognised() {
for c in ["G", "TV-Y", "tv-y7", "U", " PG "] {
assert!(is_childrens_certification(c), "{c} should be a children's rating");
}
for c in ["R", "NC-17", "TV-MA", "18", ""] {
assert!(!is_childrens_certification(c), "{c} should not be");
}
}
#[test]
fn childrens_guard_tightens_the_threshold() {
// §5a layer 2: the highest-harm case gets the tightest gate. A ratio of
// 0.7 lists normally but only reaches `flagged` on a children's title.
let mut submitted: Vec<_> = (1..=7).map(by_id).collect();
submitted.extend((900..903).map(by_id));
let mut out = evaluate(&submitted, &cast_of_20());
assert_eq!(out.verdict, Verdict::Listed);
let m = submitted.len();
apply_childrens_guard(&mut out, m);
assert_eq!(out.verdict, Verdict::Flagged);
assert_eq!(out.reason.as_deref(), Some("childrens_title_strict_threshold"));
}
#[test]
fn childrens_guard_leaves_strong_matches_listed() {
let submitted: Vec<_> = (1..=10).map(by_id).collect();
let mut out = evaluate(&submitted, &cast_of_20());
let m = submitted.len();
apply_childrens_guard(&mut out, m);
assert_eq!(out.verdict, Verdict::Listed);
}
}
+62
View File
@@ -0,0 +1,62 @@
//! Operational configuration, read from the environment.
//!
//! The trusted-proxy CIDR is explicit configuration rather than a default-on
//! behaviour (§8 deployment notes): §5 rate limiting and report attribution key
//! on client IP, so an unconditionally-trusted `X-Forwarded-For` defeats both.
use std::net::IpAddr;
use std::time::Duration;
#[derive(Clone, Debug)]
pub struct Config {
pub bind: String,
pub db_path: String,
/// Hard dependency for UR-3. Without it, uploads accumulate in `pending`
/// rather than being listed unverified (§8).
pub tmdb_api_key: Option<String>,
pub tmdb_base_url: String,
/// Prefixes of proxy addresses whose `X-Forwarded-For` is honoured.
pub trusted_proxies: Vec<IpAddr>,
pub server_id: String,
pub request_timeout: Duration,
/// Number of cast-check jobs to lease per worker tick.
pub job_batch: usize,
pub job_poll_interval: Duration,
}
impl Config {
pub fn from_env() -> anyhow::Result<Self> {
let trusted_proxies = match std::env::var("JRAY_TRUSTED_PROXIES") {
Ok(v) => v
.split(',')
.map(str::trim)
.filter(|s| !s.is_empty())
.map(|s| {
s.parse::<IpAddr>()
.map_err(|e| anyhow::anyhow!("bad JRAY_TRUSTED_PROXIES entry {s:?}: {e}"))
})
.collect::<Result<Vec<_>, _>>()?,
Err(_) => Vec::new(),
};
Ok(Self {
bind: env_or("JRAY_BIND", "127.0.0.1:8080"),
db_path: env_or("JRAY_DB", "jray.db"),
tmdb_api_key: std::env::var("JRAY_TMDB_API_KEY").ok().filter(|s| !s.is_empty()),
tmdb_base_url: env_or("JRAY_TMDB_BASE_URL", "https://api.themoviedb.org/3"),
trusted_proxies,
server_id: env_or("JRAY_SERVER_ID", "localhost"),
request_timeout: Duration::from_secs(env_num("JRAY_REQUEST_TIMEOUT_SEC", 30)),
job_batch: env_num("JRAY_JOB_BATCH", 8) as usize,
job_poll_interval: Duration::from_secs(env_num("JRAY_JOB_POLL_SEC", 5)),
})
}
}
fn env_or(key: &str, default: &str) -> String {
std::env::var(key).ok().filter(|s| !s.is_empty()).unwrap_or_else(|| default.to_string())
}
fn env_num(key: &str, default: u64) -> u64 {
std::env::var(key).ok().and_then(|v| v.parse().ok()).unwrap_or(default)
}
+289
View File
@@ -0,0 +1,289 @@
//! §9a content addressing.
//!
//! A validated manifest is immutable and content-addressable, which is what
//! makes replication *set reconciliation* rather than state synchronisation.
//! Even without the federation endpoints, computing `content_id` on upload gives
//! deduplication now and means stored manifests are already addressable when
//! federation lands.
//!
//! **This canonical form must be reimplemented byte-identically by the JRay
//! plugin** (§8 "Cost of choosing Rust": the extraction side is Python, so this
//! can no longer be shared as one implementation and must instead be specified
//! precisely and cross-tested). [`GOLDEN_VECTORS`] is that shared fixture.
use sha2::{Digest, Sha256};
/// One actor's contribution to the canonical form.
#[derive(Debug, Clone)]
pub struct CanonicalActor {
pub tmdb_person_id: u64,
/// Integer centiseconds — quantised, not formatted floats (§9a).
pub scenes_cs: Vec<(i64, i64)>,
}
/// The identity coordinates that enter the hash.
#[derive(Debug, Clone, Default)]
pub struct CanonicalIdentity {
pub kind: &'static str,
pub tmdb_id: Option<String>,
pub imdb_id: Option<String>,
pub season: Option<i64>,
pub episode: Option<i64>,
}
/// The cut coordinates that enter the hash.
///
/// **`audio_signature` is excluded, deliberately** (§9a): it is derived by
/// decoding audio, so two servers running different FFmpeg or resampler versions
/// could compute marginally different signatures for identical content, and
/// including it would silently break federation deduplication.
#[derive(Debug, Clone, Default)]
pub struct CanonicalCut {
/// Quantised to centiseconds for the same reason scene times are.
pub runtime_cs: i64,
pub video_hash: Option<String>,
}
/// Builds the canonical JSON form: keys sorted, no whitespace, actors sorted by
/// person id, scene times as integer centiseconds.
///
/// `extraction` metadata and all local state are excluded, so two servers that
/// validated the same upload independently arrive at the same `content_id`.
pub fn canonical_json(
identity: &CanonicalIdentity,
cut: &CanonicalCut,
actors: &[CanonicalActor],
) -> String {
let mut sorted: Vec<&CanonicalActor> = actors.iter().collect();
sorted.sort_by_key(|a| a.tmdb_person_id);
let mut s = String::new();
s.push_str("{\"actors\":[");
for (i, a) in sorted.iter().enumerate() {
if i > 0 {
s.push(',');
}
// Scene windows are emitted in stored order; validation has already
// established they are sorted by start time.
s.push_str("{\"scenes\":[");
for (j, (start, end)) in a.scenes_cs.iter().enumerate() {
if j > 0 {
s.push(',');
}
s.push('[');
s.push_str(&start.to_string());
s.push(',');
s.push_str(&end.to_string());
s.push(']');
}
s.push_str("],\"tmdb_person_id\":");
s.push_str(&a.tmdb_person_id.to_string());
s.push('}');
}
s.push_str("],\"cut\":{");
s.push_str("\"runtime_cs\":");
s.push_str(&cut.runtime_cs.to_string());
s.push_str(",\"video_hash\":");
push_opt_str(&mut s, cut.video_hash.as_deref());
s.push_str("},\"identity\":{");
s.push_str("\"episode\":");
push_opt_num(&mut s, cut_opt(identity.episode));
s.push_str(",\"imdb_id\":");
push_opt_str(&mut s, identity.imdb_id.as_deref());
s.push_str(",\"season\":");
push_opt_num(&mut s, cut_opt(identity.season));
s.push_str(",\"tmdb_id\":");
push_opt_str(&mut s, identity.tmdb_id.as_deref());
s.push_str(",\"type\":\"");
s.push_str(identity.kind);
s.push_str("\"}}");
s
}
fn cut_opt(v: Option<i64>) -> Option<i64> {
v
}
fn push_opt_str(s: &mut String, v: Option<&str>) {
match v {
// Only closed-vocabulary values reach here (regex-constrained ids and a
// fixed-format hash), so no string escaping is required.
Some(v) => {
s.push('"');
s.push_str(v);
s.push('"');
}
None => s.push_str("null"),
}
}
fn push_opt_num(s: &mut String, v: Option<i64>) {
match v {
Some(v) => s.push_str(&v.to_string()),
None => s.push_str("null"),
}
}
/// `sha256:` over the canonical form (§9a).
pub fn content_id(
identity: &CanonicalIdentity,
cut: &CanonicalCut,
actors: &[CanonicalActor],
) -> String {
let canonical = canonical_json(identity, cut, actors);
let mut h = Sha256::new();
h.update(canonical.as_bytes());
let digest = h.finalize();
let mut hex = String::with_capacity(64 + 7);
hex.push_str("sha256:");
for b in digest {
hex.push_str(&format!("{b:02x}"));
}
hex
}
/// Cross-implementation fixture (§8): the JRay plugin and any reimplementation
/// must reproduce these exactly, or federation deduplication silently breaks.
pub const GOLDEN_VECTORS: &[(&str, &str)] = &[(
// Movie, one actor, two windows, with a video hash.
r#"{"actors":[{"scenes":[[19160,20920],[43820,46560]],"tmdb_person_id":884}],"cut":{"runtime_cs":642050,"video_hash":"opensubtitles:8e245d9679d31e12"},"identity":{"episode":null,"imdb_id":"tt4686844","season":null,"tmdb_id":"504172","type":"movie"}}"#,
// Verified against an independent Python implementation:
// sha256(canonical.encode()).hexdigest()
"sha256:367f8b05c54a992a3a30fa016edaaac0b9b36b148fc76574f5b1ef326b56760f",
)];
#[cfg(test)]
mod tests {
use super::*;
fn movie_identity() -> CanonicalIdentity {
CanonicalIdentity {
kind: "movie",
tmdb_id: Some("504172".into()),
imdb_id: Some("tt4686844".into()),
season: None,
episode: None,
}
}
fn movie_cut() -> CanonicalCut {
CanonicalCut {
runtime_cs: 642050,
video_hash: Some("opensubtitles:8e245d9679d31e12".into()),
}
}
fn actors() -> Vec<CanonicalActor> {
vec![CanonicalActor {
tmdb_person_id: 884,
scenes_cs: vec![(19160, 20920), (43820, 46560)],
}]
}
#[test]
fn canonical_form_matches_the_documented_shape() {
let json = canonical_json(&movie_identity(), &movie_cut(), &actors());
assert_eq!(json, GOLDEN_VECTORS[0].0);
// Keys sorted, no whitespace (§9a).
assert!(!json.contains(' '));
}
#[test]
fn canonical_form_is_valid_json_with_sorted_keys() {
// Hand-built strings are easy to get subtly wrong, so assert the output
// actually parses and that its keys really are ordered.
let json = canonical_json(&movie_identity(), &movie_cut(), &actors());
let v: serde_json::Value =
serde_json::from_str(&json).expect("canonical form must be JSON");
let obj = v.as_object().unwrap();
let keys: Vec<&String> = obj.keys().collect();
assert_eq!(keys, vec!["actors", "cut", "identity"]);
let id_keys: Vec<&String> = v["identity"].as_object().unwrap().keys().collect();
assert_eq!(id_keys, vec!["episode", "imdb_id", "season", "tmdb_id", "type"]);
let cut_keys: Vec<&String> = v["cut"].as_object().unwrap().keys().collect();
assert_eq!(cut_keys, vec!["runtime_cs", "video_hash"]);
}
#[test]
fn actor_order_does_not_affect_the_hash() {
// §9a: actors sorted by person id, so two servers that stored them in
// different orders still agree.
let a = vec![
CanonicalActor { tmdb_person_id: 884, scenes_cs: vec![(0, 100)] },
CanonicalActor { tmdb_person_id: 17419, scenes_cs: vec![(200, 300)] },
];
let b = vec![a[1].clone(), a[0].clone()];
assert_eq!(
content_id(&movie_identity(), &movie_cut(), &a),
content_id(&movie_identity(), &movie_cut(), &b)
);
}
#[test]
fn accumulated_float_error_hashes_identically() {
// The failure mode §9a exists to remove: real corpus values look like
// 8045.066666660665, and two servers may compute them slightly
// differently. Quantising first means both hash the same.
let a = vec![CanonicalActor {
tmdb_person_id: 1,
scenes_cs: vec![(crate::validate::to_centiseconds(8045.066666660665), 900000)],
}];
let b = vec![CanonicalActor {
tmdb_person_id: 1,
scenes_cs: vec![(crate::validate::to_centiseconds(8045.066666666), 900000)],
}];
assert_eq!(
content_id(&movie_identity(), &movie_cut(), &a),
content_id(&movie_identity(), &movie_cut(), &b)
);
}
#[test]
fn differing_content_produces_differing_ids() {
let base = content_id(&movie_identity(), &movie_cut(), &actors());
let mut other_actors = actors();
other_actors[0].scenes_cs[0].1 += 1;
assert_ne!(base, content_id(&movie_identity(), &movie_cut(), &other_actors));
let mut other_cut = movie_cut();
other_cut.runtime_cs += 1;
assert_ne!(base, content_id(&movie_identity(), &other_cut, &actors()));
let mut other_id = movie_identity();
other_id.tmdb_id = Some("999".into());
assert_ne!(base, content_id(&other_id, &movie_cut(), &actors()));
}
#[test]
fn episode_and_movie_coordinates_are_distinguished() {
let ep = CanonicalIdentity {
kind: "episode",
tmdb_id: Some("1396".into()),
imdb_id: None,
season: Some(2),
episode: Some(5),
};
let other = CanonicalIdentity { season: Some(3), ..ep.clone() };
assert_ne!(
content_id(&ep, &movie_cut(), &actors()),
content_id(&other, &movie_cut(), &actors())
);
}
#[test]
fn content_id_is_prefixed_and_hex() {
let id = content_id(&movie_identity(), &movie_cut(), &actors());
let hex = id.strip_prefix("sha256:").expect("prefixed");
assert_eq!(hex.len(), 64);
assert!(hex.bytes().all(|b| b.is_ascii_hexdigit()));
}
#[test]
fn golden_vector_hash_is_stable() {
// Locks the hash so an accidental change to the canonical form is caught
// here rather than by silent federation divergence.
let id = content_id(&movie_identity(), &movie_cut(), &actors());
assert_eq!(id, GOLDEN_VECTORS[0].1, "canonical form or hash changed");
}
}
+168
View File
@@ -0,0 +1,168 @@
//! Database access.
//!
//! §8 imposes two structural requirements that this module exists to satisfy:
//!
//! 1. **A single writer connection, serialized through one owner**, with a read
//! pool alongside. SQLite permits only one writer at a time even in WAL mode;
//! pointing a multi-connection pool at writes and relying on `busy_timeout`
//! to sort it out is explicitly rejected by the spec. Here the writer lives
//! behind a `Mutex`, so contention queues in Rust rather than surfacing as
//! `SQLITE_BUSY`.
//! 2. **All access behind a thin repository layer** rather than queries
//! scattered through handlers — this is what keeps the Turso/Postgres options
//! cheap and localises the serialization in one place.
//!
//! rusqlite is synchronous, so every call is wrapped in `spawn_blocking`: a
//! write that waits on the mutex must never block a Tokio worker thread.
pub mod repo;
use std::sync::{Arc, Mutex};
use anyhow::Context;
use rusqlite::Connection;
const SCHEMA: &str = include_str!("schema.sql");
/// Handle to the database: one serialized writer, plus read connections.
///
/// Cloning is cheap and shares the same underlying connections.
#[derive(Clone)]
pub struct Db {
writer: Arc<Mutex<Connection>>,
readers: Arc<ReadPool>,
}
struct ReadPool {
conns: Mutex<Vec<Connection>>,
path: String,
}
impl ReadPool {
fn acquire(&self) -> anyhow::Result<Connection> {
if let Some(c) = self.conns.lock().expect("read pool poisoned").pop() {
return Ok(c);
}
open_conn(&self.path, false)
}
fn release(&self, conn: Connection) {
let mut conns = self.conns.lock().expect("read pool poisoned");
// Bounded: excess connections are dropped rather than accumulating.
if conns.len() < 8 {
conns.push(conn);
}
}
}
fn open_conn(path: &str, writer: bool) -> anyhow::Result<Connection> {
let conn = Connection::open(path).with_context(|| format!("opening database {path}"))?;
// WAL gives concurrent readers alongside the single writer, which suits a
// read-dominated workload; `synchronous = NORMAL` is safe under WAL, and
// `busy_timeout` makes contention wait rather than error (§8).
conn.pragma_update(None, "journal_mode", "WAL")?;
conn.pragma_update(None, "synchronous", "NORMAL")?;
conn.pragma_update(None, "busy_timeout", 5_000)?;
conn.pragma_update(None, "foreign_keys", true)?;
if !writer {
conn.pragma_update(None, "query_only", true)?;
}
Ok(conn)
}
impl Db {
/// Opens the database, applying the schema. Idempotent — every statement in
/// `schema.sql` is `IF NOT EXISTS`.
pub fn open(path: &str) -> anyhow::Result<Self> {
let writer = open_conn(path, true)?;
writer.execute_batch(SCHEMA).context("applying schema")?;
Ok(Self {
writer: Arc::new(Mutex::new(writer)),
readers: Arc::new(ReadPool { conns: Mutex::new(Vec::new()), path: path.to_string() }),
})
}
/// Runs `f` against the serialized writer connection on a blocking thread.
///
/// `f` receives a `Transaction`, so a manifest's scene rows go in as one
/// transaction rather than one per row (§8), and a failure rolls back.
pub async fn write<T, F>(&self, f: F) -> anyhow::Result<T>
where
T: Send + 'static,
F: FnOnce(&rusqlite::Transaction<'_>) -> anyhow::Result<T> + Send + 'static,
{
let writer = self.writer.clone();
tokio::task::spawn_blocking(move || {
let mut conn = writer.lock().expect("writer poisoned");
let tx = conn.transaction()?;
let out = f(&tx)?;
tx.commit()?;
Ok(out)
})
.await
.context("writer task panicked")?
}
/// Runs `f` against a read connection on a blocking thread.
pub async fn read<T, F>(&self, f: F) -> anyhow::Result<T>
where
T: Send + 'static,
F: FnOnce(&Connection) -> anyhow::Result<T> + Send + 'static,
{
let readers = self.readers.clone();
tokio::task::spawn_blocking(move || {
let conn = readers.acquire()?;
let out = f(&conn);
readers.release(conn);
out
})
.await
.context("reader task panicked")?
}
}
#[cfg(test)]
mod tests {
use super::*;
#[tokio::test]
async fn schema_applies_and_roundtrips() {
let db = Db::open(":memory:").unwrap();
// In-memory databases are per-connection, so only exercise the writer.
let n = db
.write(|tx| {
tx.execute(
"INSERT INTO contributors (id, token_hash, created_at) VALUES (?1, ?2, ?3)",
rusqlite::params!["c1", "hash", "2026-01-01T00:00:00Z"],
)?;
Ok(tx.query_row("SELECT COUNT(*) FROM contributors", [], |r| r.get::<_, i64>(0))?)
})
.await
.unwrap();
assert_eq!(n, 1);
}
#[tokio::test]
async fn write_rolls_back_on_error() {
let db = Db::open(":memory:").unwrap();
let res: anyhow::Result<()> = db
.write(|tx| {
tx.execute(
"INSERT INTO contributors (id, token_hash, created_at) VALUES ('c1','h','t')",
[],
)?;
anyhow::bail!("deliberate failure")
})
.await;
assert!(res.is_err());
let n = db
.write(|tx| {
Ok(tx.query_row("SELECT COUNT(*) FROM contributors", [], |r| r.get::<_, i64>(0))?)
})
.await
.unwrap();
assert_eq!(n, 0, "failed transaction must not persist rows");
}
}
+1110
View File
File diff suppressed because it is too large Load Diff
+119
View File
@@ -0,0 +1,119 @@
-- §7 Storage. Fully relational, no JSON blobs on the write path: the database
-- can only represent what the schema models, so there is physically nowhere for
-- an unexpected field or a smuggled string to live (§5a Threat 1).
--
-- Portable SQL — runs unchanged on Postgres. Avoid SQLite-specific forms
-- (`INSERT OR REPLACE`); use `INSERT ... ON CONFLICT` (§8 deployment notes).
CREATE TABLE IF NOT EXISTS contributors (
id TEXT PRIMARY KEY,
token_hash TEXT NOT NULL UNIQUE,
created_at TEXT NOT NULL,
revoked_at TEXT,
accepted_count INTEGER NOT NULL DEFAULT 0,
rejected_count INTEGER NOT NULL DEFAULT 0,
flagged_count INTEGER NOT NULL DEFAULT 0
);
-- Server-side, TMDB-derived. `name` never comes from an upload (§5a).
CREATE TABLE IF NOT EXISTS people (
tmdb_person_id INTEGER PRIMARY KEY,
name TEXT NOT NULL,
adult INTEGER NOT NULL DEFAULT 0,
updated_at TEXT NOT NULL
);
CREATE TABLE IF NOT EXISTS titles (
id TEXT PRIMARY KEY,
kind TEXT NOT NULL, -- movie | series
tmdb_id TEXT,
imdb_id TEXT,
name TEXT,
year INTEGER,
adult INTEGER NOT NULL DEFAULT 0,
certification TEXT,
updated_at TEXT NOT NULL
);
CREATE TABLE IF NOT EXISTS manifests (
id TEXT PRIMARY KEY,
title_id TEXT NOT NULL REFERENCES titles(id),
season INTEGER,
episode INTEGER,
runtime_sec REAL NOT NULL,
video_hash TEXT,
audio_signature BLOB, -- §3, ~1290 bytes
audio_sig_coarse BLOB, -- candidate-generation index key
sample_fps REAL,
extinction_sec REAL, -- successor to the withdrawn anneal_sec
gallery_scope TEXT, -- limited | global; ranking signal (§2, §7)
pipeline_version TEXT,
contributor_id TEXT REFERENCES contributors(id),
status TEXT NOT NULL, -- pending | listed | flagged | rejected
reject_reason TEXT,
cast_match_ratio REAL,
content_id TEXT UNIQUE, -- §9a, sha256 over canonical form
origin TEXT, -- server_id of first acceptance
ingested_from TEXT, -- peer id, NULL if uploaded directly
created_at TEXT NOT NULL
);
CREATE TABLE IF NOT EXISTS manifest_actors (
manifest_id TEXT NOT NULL REFERENCES manifests(id) ON DELETE CASCADE,
tmdb_person_id INTEGER NOT NULL,
PRIMARY KEY (manifest_id, tmdb_person_id)
);
-- Integer centiseconds, not floats — the same quantisation used for
-- `content_id`, so stored values and hashed values cannot diverge (§7, §9a).
CREATE TABLE IF NOT EXISTS scenes (
manifest_id TEXT NOT NULL REFERENCES manifests(id) ON DELETE CASCADE,
tmdb_person_id INTEGER NOT NULL,
start_cs INTEGER NOT NULL,
end_cs INTEGER NOT NULL
);
CREATE TABLE IF NOT EXISTS reports (
id TEXT PRIMARY KEY,
manifest_id TEXT NOT NULL REFERENCES manifests(id) ON DELETE CASCADE,
reason TEXT NOT NULL,
note TEXT,
created_at TEXT NOT NULL,
source_ip_hash TEXT
);
-- The sole JSON column, and it holds TMDB's responses, not users' (§7).
CREATE TABLE IF NOT EXISTS tmdb_cache (
tmdb_id TEXT NOT NULL,
kind TEXT NOT NULL,
credits TEXT NOT NULL,
fetched_at TEXT NOT NULL,
PRIMARY KEY (tmdb_id, kind)
);
-- Background queue as a table rather than an external broker, so pending work
-- survives a restart (§7, §8).
CREATE TABLE IF NOT EXISTS jobs (
id TEXT PRIMARY KEY,
kind TEXT NOT NULL, -- cast_check | federation_pull
payload TEXT NOT NULL,
run_after TEXT NOT NULL,
attempts INTEGER NOT NULL DEFAULT 0,
last_error TEXT,
leased_at TEXT
);
CREATE INDEX IF NOT EXISTS idx_titles_tmdb ON titles(tmdb_id);
CREATE INDEX IF NOT EXISTS idx_titles_imdb ON titles(imdb_id);
CREATE INDEX IF NOT EXISTS idx_manifests_title_runtime ON manifests(title_id, runtime_sec);
CREATE INDEX IF NOT EXISTS idx_manifests_video_hash ON manifests(video_hash);
CREATE INDEX IF NOT EXISTS idx_manifests_episode ON manifests(title_id, season, episode);
CREATE INDEX IF NOT EXISTS idx_scenes_manifest_person ON scenes(manifest_id, tmdb_person_id);
-- All read queries filter `status IN ('listed','flagged')`, so a partial index
-- on that predicate keeps the hot path small (§7).
CREATE INDEX IF NOT EXISTS idx_manifests_served
ON manifests(title_id, season, episode)
WHERE status IN ('listed', 'flagged');
CREATE INDEX IF NOT EXISTS idx_jobs_ready ON jobs(run_after);
+83
View File
@@ -0,0 +1,83 @@
//! API error type mapping onto the status codes §4 specifies.
use axum::http::StatusCode;
use axum::response::{IntoResponse, Response};
use axum::Json;
use serde::Serialize;
#[derive(Debug, thiserror::Error)]
pub enum ApiError {
/// §6 stage 2 — malformed, unrecognised or forbidden field. The message
/// names the offending field so a client that forgets to strip `movie` or
/// `jellyfin_id` gets a hard, diagnosable `400` (§6).
#[error("{0}")]
BadRequest(String),
#[error("not found")]
NotFound,
/// §4 — identical `(identity, cut)` already exists from this contributor.
#[error("{0}")]
Conflict(String),
#[error("{0}")]
PayloadTooLarge(String),
#[error("missing or invalid API token")]
Unauthorized,
/// §5 — carries the `Retry-After` value in seconds.
#[error("rate limited")]
RateLimited { retry_after: u64 },
#[error("internal error")]
Internal(#[from] anyhow::Error),
}
#[derive(Serialize)]
struct ErrorBody {
error: String,
message: String,
}
impl IntoResponse for ApiError {
fn into_response(self) -> Response {
let (status, code) = match &self {
ApiError::BadRequest(_) => (StatusCode::BAD_REQUEST, "bad_request"),
ApiError::NotFound => (StatusCode::NOT_FOUND, "not_found"),
ApiError::Conflict(_) => (StatusCode::CONFLICT, "conflict"),
ApiError::PayloadTooLarge(_) => (StatusCode::PAYLOAD_TOO_LARGE, "payload_too_large"),
ApiError::Unauthorized => (StatusCode::UNAUTHORIZED, "unauthorized"),
ApiError::RateLimited { .. } => (StatusCode::TOO_MANY_REQUESTS, "rate_limited"),
ApiError::Internal(e) => {
// Internal detail is logged, never returned.
tracing::error!(error = ?e, "internal error");
(StatusCode::INTERNAL_SERVER_ERROR, "internal")
}
};
let body = Json(ErrorBody {
error: code.to_string(),
message: match &self {
ApiError::Internal(_) => "internal error".to_string(),
other => other.to_string(),
},
});
let mut resp = (status, body).into_response();
if let ApiError::RateLimited { retry_after } = self {
if let Ok(v) = retry_after.to_string().parse() {
resp.headers_mut().insert(axum::http::header::RETRY_AFTER, v);
}
}
resp
}
}
impl From<rusqlite::Error> for ApiError {
fn from(e: rusqlite::Error) -> Self {
ApiError::Internal(anyhow::Error::new(e))
}
}
pub type ApiResult<T> = Result<T, ApiError>;
+403
View File
@@ -0,0 +1,403 @@
//! Manifest ingestion: the shared path behind `POST /manifests` and
//! `POST /manifests/bundle`, and the path a federation pull will reuse (§9a
//! "re-derive, don't inherit").
//!
//! Stages 0 and 1 are layers; stage 2 is parse + [`crate::validate`]. What
//! happens here is persistence plus enqueueing the stage 3 check: the upload is
//! accepted with `202` and the manifest is held **unlisted** until the cast check
//! completes — it is not served to anyone in the meantime (§6).
use anyhow::Context;
use crate::content_id::{self, CanonicalActor, CanonicalCut, CanonicalIdentity};
use crate::db::repo::{self, NewManifest};
use crate::model::IdentityType;
use crate::validate::ValidManifest;
/// Outcome of persisting one manifest.
#[derive(Debug, Clone)]
pub enum IngestOutcome {
/// Held unlisted pending the §6 stage 3 cast check.
Pending { manifest_id: String },
/// §4 `409` — identical `(identity, cut)` from this contributor.
DuplicateFromContributor { manifest_id: String },
/// §9a — the exact same content is already held, from any source. Skipped
/// without re-validation, which is the deduplication content addressing buys.
DuplicateContent { manifest_id: String },
}
impl IngestOutcome {
pub fn manifest_id(&self) -> &str {
match self {
IngestOutcome::Pending { manifest_id }
| IngestOutcome::DuplicateFromContributor { manifest_id }
| IngestOutcome::DuplicateContent { manifest_id } => manifest_id,
}
}
}
/// Job payload for the §6 stage 3 check.
#[derive(Debug, Clone, serde::Serialize, serde::Deserialize)]
pub struct CastCheckJob {
pub manifest_id: String,
}
pub const JOB_CAST_CHECK: &str = "cast_check";
/// Persists a validated manifest and enqueues its cast check, all in one
/// transaction — so a manifest is never left listed-but-unchecked, and its scene
/// rows go in as a single transaction rather than one per row (§8).
pub fn persist(
tx: &rusqlite::Transaction<'_>,
valid: &ValidManifest,
contributor_id: Option<&str>,
origin: &str,
ingested_from: Option<&str>,
now: &str,
) -> anyhow::Result<IngestOutcome> {
let m = &valid.manifest;
let kind = m.identity.kind;
let tmdb_id = m.identity.effective_tmdb_id();
let imdb_id = m.identity.effective_imdb_id();
let title_id = repo::upsert_title(
tx,
kind,
tmdb_id,
imdb_id,
m.identity.title.as_deref(),
m.identity.year,
now,
)
.context("resolving title")?;
let (season, episode) = match kind {
IdentityType::Movie => (None, None),
IdentityType::Episode => (m.identity.season, m.identity.episode),
};
// Content addressing over the *submitted* actor ids. Recomputed after the
// cast check drops unmatched actors, since dropping changes the content.
let cid = compute_content_id(valid);
if let Some(existing) = repo::manifest_by_content_id(tx, &cid)? {
return Ok(IngestOutcome::DuplicateContent { manifest_id: existing });
}
if let Some(c) = contributor_id {
if let Some(existing) = repo::duplicate_from_contributor(
tx,
&title_id,
season,
episode,
m.cut.runtime_sec,
m.cut.video_hash.as_deref(),
c,
)? {
return Ok(IngestOutcome::DuplicateFromContributor { manifest_id: existing });
}
}
let manifest_id = ulid::Ulid::new().to_string();
let extraction = m.extraction.as_ref();
repo::insert_manifest(
tx,
&NewManifest {
id: &manifest_id,
title_id: &title_id,
season,
episode,
runtime_sec: m.cut.runtime_sec,
video_hash: m.cut.video_hash.as_deref(),
// Stored as an attribute, not part of identity (§9a).
audio_signature: None,
audio_sig_coarse: None,
sample_fps: extraction.and_then(|e| e.sample_fps),
extinction_sec: extraction.and_then(|e| e.extinction_sec),
pipeline_version: extraction.and_then(|e| e.pipeline_version.as_deref()),
gallery_scope: extraction.and_then(|e| e.gallery_scope).map(|g| g.as_str()),
contributor_id,
// Held unlisted until stage 3 completes (§6).
status: "pending",
content_id: Some(&cid),
origin,
ingested_from,
created_at: now,
},
)?;
// Actors are recorded by TMDB person id only. Those without one cannot be
// stored at all — there is no name column to put them in (§5a, §7) — so they
// are carried into the cast check via the submitted payload instead.
for actor in &valid.actor_scenes_cs {
if let Some(person_id) = actor.tmdb_id {
repo::insert_actor_scenes(tx, &manifest_id, person_id, &actor.scenes_cs)?;
}
}
let payload = serde_json::to_string(&CastCheckJob { manifest_id: manifest_id.clone() })?;
repo::enqueue_job(tx, JOB_CAST_CHECK, &payload, now)?;
Ok(IngestOutcome::Pending { manifest_id })
}
/// Computes the §9a `content_id` for a validated manifest.
pub fn compute_content_id(valid: &ValidManifest) -> String {
let m = &valid.manifest;
let identity = CanonicalIdentity {
kind: match m.identity.kind {
IdentityType::Movie => "movie",
IdentityType::Episode => "episode",
},
tmdb_id: m.identity.effective_tmdb_id().map(str::to_string),
imdb_id: m.identity.effective_imdb_id().map(str::to_string),
season: m.identity.season,
episode: m.identity.episode,
};
let cut = CanonicalCut {
runtime_cs: crate::validate::to_centiseconds(m.cut.runtime_sec),
video_hash: m.cut.video_hash.clone(),
};
let actors: Vec<CanonicalActor> = valid
.actor_scenes_cs
.iter()
.filter_map(|a| {
a.tmdb_id
.map(|id| CanonicalActor { tmdb_person_id: id, scenes_cs: a.scenes_cs.clone() })
})
.collect();
content_id::content_id(&identity, &cut, &actors)
}
#[cfg(test)]
mod tests {
use super::*;
use crate::db::Db;
use crate::model::Jmanifest;
use crate::validate::validate_manifest;
const NOW: &str = "2026-07-30T12:00:00Z";
fn valid_from(json: &str) -> ValidManifest {
let m: Jmanifest = serde_json::from_str(json).unwrap();
validate_manifest(m).unwrap()
}
fn movie_json(tmdb: &str, runtime: f64) -> String {
format!(
r#"{{"jmanifest_version":1,
"identity":{{"type":"movie","tmdb_id":"{tmdb}","title":"A Film"}},
"cut":{{"runtime_sec":{runtime}}},
"extraction":{{"sample_fps":5,"pipeline_version":"test 0.1"}},
"actors":[{{"name":"Steve Buscemi","tmdb_id":"884","scenes":[[10.0,20.0]]}},
{{"name":"Michael Palin","tmdb_id":"11007","scenes":[[30.0,40.0]]}}]}}"#
)
}
#[tokio::test]
async fn persists_as_pending_and_enqueues_a_check() {
let db = Db::open(":memory:").unwrap();
let valid = valid_from(&movie_json("504172", 6420.5));
let (outcome, status, jobs) = db
.write(move |tx| {
let c = repo::insert_contributor(tx, "h", NOW)?;
let outcome = persist(tx, &valid, Some(&c), "local", None, NOW)?;
let status = repo::manifest_status(tx, outcome.manifest_id())?;
let jobs = repo::lease_jobs(tx, NOW, 10)?;
Ok((outcome, status, jobs))
})
.await
.unwrap();
assert!(matches!(outcome, IngestOutcome::Pending { .. }));
// §6: held unlisted, not served to anyone, until stage 3 completes.
assert_eq!(status.unwrap().0, "pending");
assert_eq!(jobs.len(), 1);
assert_eq!(jobs[0].kind, JOB_CAST_CHECK);
}
#[tokio::test]
async fn a_pending_manifest_is_not_served() {
let db = Db::open(":memory:").unwrap();
let valid = valid_from(&movie_json("504172", 6420.5));
let candidates = db
.write(move |tx| {
let c = repo::insert_contributor(tx, "h", NOW)?;
persist(tx, &valid, Some(&c), "local", None, NOW)?;
let title =
repo::find_title(tx, IdentityType::Movie, Some("504172"), None)?.unwrap();
repo::candidates_for_title(tx, &title.id, None, None)
})
.await
.unwrap();
assert!(candidates.is_empty());
}
#[tokio::test]
async fn identical_content_deduplicates() {
// §9a: a manifest whose `content_id` is already present is skipped
// without re-validation.
let db = Db::open(":memory:").unwrap();
let a = valid_from(&movie_json("504172", 6420.5));
let b = valid_from(&movie_json("504172", 6420.5));
let (first, second) = db
.write(move |tx| {
let c1 = repo::insert_contributor(tx, "h1", NOW)?;
let c2 = repo::insert_contributor(tx, "h2", NOW)?;
let first = persist(tx, &a, Some(&c1), "local", None, NOW)?;
// A *different* contributor, so this is content dedup, not the
// per-contributor 409.
let second = persist(tx, &b, Some(&c2), "local", None, NOW)?;
Ok((first, second))
})
.await
.unwrap();
assert!(matches!(first, IngestOutcome::Pending { .. }));
assert!(matches!(second, IngestOutcome::DuplicateContent { .. }));
assert_eq!(first.manifest_id(), second.manifest_id());
}
#[tokio::test]
async fn same_contributor_resubmitting_the_same_cut_is_a_duplicate() {
let db = Db::open(":memory:").unwrap();
// Same identity and cut, different actor timings => different content_id,
// so this exercises the per-contributor 409 path specifically.
let a = valid_from(&movie_json("504172", 6420.5));
let b = valid_from(
r#"{"jmanifest_version":1,
"identity":{"type":"movie","tmdb_id":"504172","title":"A Film"},
"cut":{"runtime_sec":6420.5},
"actors":[{"name":"Steve Buscemi","tmdb_id":"884","scenes":[[11.0,21.0]]}]}"#,
);
let second = db
.write(move |tx| {
let c = repo::insert_contributor(tx, "h", NOW)?;
persist(tx, &a, Some(&c), "local", None, NOW)?;
persist(tx, &b, Some(&c), "local", None, NOW)
})
.await
.unwrap();
assert!(matches!(second, IngestOutcome::DuplicateFromContributor { .. }));
}
#[tokio::test]
async fn different_cuts_of_one_title_coexist() {
// §7: multiple manifests may coexist for the same title with different
// cuts — that is the point.
let db = Db::open(":memory:").unwrap();
let a = valid_from(&movie_json("504172", 6420.5));
let b = valid_from(&movie_json("504172", 7000.0));
let (x, y) = db
.write(move |tx| {
let c = repo::insert_contributor(tx, "h", NOW)?;
let x = persist(tx, &a, Some(&c), "local", None, NOW)?;
let y = persist(tx, &b, Some(&c), "local", None, NOW)?;
Ok((x, y))
})
.await
.unwrap();
assert!(matches!(x, IngestOutcome::Pending { .. }));
assert!(matches!(y, IngestOutcome::Pending { .. }));
assert_ne!(x.manifest_id(), y.manifest_id());
}
#[tokio::test]
async fn episode_manifests_carry_their_coordinates() {
let db = Db::open(":memory:").unwrap();
let valid = valid_from(
r#"{"jmanifest_version":1,
"identity":{"type":"episode","series_tmdb_id":"1396","title":"Breaking Bad",
"season":2,"episode":5},
"cut":{"runtime_sec":2820.0},
"actors":[{"name":"Bryan Cranston","tmdb_id":"17419","scenes":[[10.0,20.0]]}]}"#,
);
let row = db
.write(move |tx| {
let c = repo::insert_contributor(tx, "h", NOW)?;
let o = persist(tx, &valid, Some(&c), "local", None, NOW)?;
Ok(repo::manifest_by_id(tx, o.manifest_id())?.unwrap())
})
.await
.unwrap();
assert_eq!((row.season, row.episode), (Some(2), Some(5)));
}
#[tokio::test]
async fn upload_metadata_is_not_echoed_back_as_actor_names() {
// §5a/§7: only integers reach the database. The submitted name is used
// for matching and never persisted, so before the cast check populates
// `people` there is no name to serve.
let db = Db::open(":memory:").unwrap();
let valid = valid_from(&movie_json("504172", 6420.5));
let actors = db
.write(move |tx| {
let c = repo::insert_contributor(tx, "h", NOW)?;
let o = persist(tx, &valid, Some(&c), "local", None, NOW)?;
repo::actors_for_manifest(tx, o.manifest_id())
})
.await
.unwrap();
assert_eq!(actors.len(), 2);
assert!(actors.iter().all(|a| a.name.is_none()));
}
#[test]
fn content_id_excludes_extraction_metadata() {
// §9a: `extraction` metadata and local state are excluded, so two
// servers validating the same upload agree.
let a = valid_from(&movie_json("504172", 6420.5));
let b = valid_from(
r#"{"jmanifest_version":1,
"identity":{"type":"movie","tmdb_id":"504172","title":"A Film"},
"cut":{"runtime_sec":6420.5},
"extraction":{"sample_fps":1,"extinction_sec":9,"pipeline_version":"other 9.9",
"gallery_size":5},
"actors":[{"name":"Steve Buscemi","tmdb_id":"884","scenes":[[10.0,20.0]]},
{"name":"Michael Palin","tmdb_id":"11007","scenes":[[30.0,40.0]]}]}"#,
);
assert_eq!(compute_content_id(&a), compute_content_id(&b));
}
#[test]
fn content_id_excludes_the_audio_signature() {
// §9a is explicit: including it would produce different content_ids for
// identical content and silently break federation deduplication.
let a = valid_from(&movie_json("504172", 6420.5));
let sig = format!("v1:{}", "A".repeat(1720));
let with_sig = format!(
r#"{{"jmanifest_version":1,
"identity":{{"type":"movie","tmdb_id":"504172","title":"A Film"}},
"cut":{{"runtime_sec":6420.5,"audio_signature":"{sig}"}},
"extraction":{{"sample_fps":5,"pipeline_version":"test 0.1"}},
"actors":[{{"name":"Steve Buscemi","tmdb_id":"884","scenes":[[10.0,20.0]]}},
{{"name":"Michael Palin","tmdb_id":"11007","scenes":[[30.0,40.0]]}}]}}"#
);
let b = valid_from(&with_sig);
assert_eq!(compute_content_id(&a), compute_content_id(&b));
}
#[test]
fn content_id_excludes_submitted_names() {
// Names are not persisted, so they must not be part of identity either —
// otherwise a renamed resubmission would evade deduplication.
let a = valid_from(&movie_json("504172", 6420.5));
let b = valid_from(
r#"{"jmanifest_version":1,
"identity":{"type":"movie","tmdb_id":"504172","title":"A Film"},
"cut":{"runtime_sec":6420.5},
"extraction":{"sample_fps":5,"pipeline_version":"test 0.1"},
"actors":[{"name":"Someone Else","tmdb_id":"884","scenes":[[10.0,20.0]]},
{"name":"Another Person","tmdb_id":"11007","scenes":[[30.0,40.0]]}]}"#,
);
assert_eq!(compute_content_id(&a), compute_content_id(&b));
}
}
+38
View File
@@ -0,0 +1,38 @@
//! JRay public server — a community manifest exchange (see `SPEC.md`).
//!
//! Jellyfin servers running the JRay plugin pull actor-timeline manifests
//! ("Jmanifests") for titles they own instead of running the CV pipeline
//! locally, and optionally contribute the manifests they generate back.
//!
//! Module map against the spec:
//!
//! | Module | Spec section |
//! |---|---|
//! | [`model`] | §2 Jmanifest format, and §6 stage 2's `deny_unknown_fields` |
//! | [`validate`] | §6 stage 2 semantics, §5a character class |
//! | [`matching`] | §3 cut matching tiers |
//! | [`content_id`] | §9a canonical form and content addressing |
//! | [`ingest`] | the shared upload path behind §4's `POST` endpoints |
//! | [`castcheck`] | §6 stage 3 scoring, §5a Threat 2 guards |
//! | [`worker`] | §6 stage 3 execution, §8 in-process background work |
//! | [`ratelimit`] | §5 |
//! | [`auth`] | §5a tokens, §8 trusted-proxy handling |
//! | [`db`] | §7 storage, §8 single-writer serialization |
//! | [`api`] | §4 |
pub mod api;
pub mod app;
pub mod auth;
pub mod castcheck;
pub mod config;
pub mod content_id;
pub mod db;
pub mod error;
pub mod ingest;
pub mod matching;
pub mod model;
pub mod ratelimit;
pub mod state;
pub mod tmdb;
pub mod validate;
pub mod worker;
+96
View File
@@ -0,0 +1,96 @@
//! Entry point.
//!
//! §8: one binary, one database file, one reverse proxy. The rate-limit counters
//! and the background cast-check worker both live in this process — no Redis, no
//! broker, no separate worker process.
use std::net::SocketAddr;
use std::sync::Arc;
use anyhow::Context;
use jray_server::app;
use jray_server::config::Config;
use jray_server::db::Db;
use jray_server::ratelimit::RateLimiter;
use jray_server::state::AppState;
use jray_server::tmdb::TmdbClient;
use jray_server::worker::Worker;
use tracing_subscriber::EnvFilter;
#[tokio::main]
async fn main() -> anyhow::Result<()> {
tracing_subscriber::fmt()
.with_env_filter(
EnvFilter::try_from_env("JRAY_LOG").unwrap_or_else(|_| EnvFilter::new("info")),
)
.init();
let config = Arc::new(Config::from_env()?);
let db = Db::open(&config.db_path).context("opening database")?;
let tmdb = Arc::new(TmdbClient::new(config.tmdb_base_url.clone(), config.tmdb_api_key.clone()));
if !tmdb.is_configured() {
// §8: TMDB is a hard dependency for UR-3. Uploads will accumulate in
// `pending` rather than being listed unverified — which is the correct
// failure mode, but the operator should know.
tracing::warn!("no JRAY_TMDB_API_KEY configured: uploads will stay pending, never listed");
}
if config.trusted_proxies.is_empty() {
tracing::info!("no JRAY_TRUSTED_PROXIES set: X-Forwarded-For will be ignored");
}
let state = AppState {
db: db.clone(),
config: config.clone(),
limiter: Arc::new(RateLimiter::new()),
tmdb: tmdb.clone(),
};
let (shutdown_tx, shutdown_rx) = tokio::sync::watch::channel(false);
let worker = Worker {
db: db.clone(),
tmdb,
batch: config.job_batch,
poll_interval: config.job_poll_interval,
};
let worker_handle = tokio::spawn(worker.run(shutdown_rx));
let listener = tokio::net::TcpListener::bind(&config.bind)
.await
.with_context(|| format!("binding {}", config.bind))?;
tracing::info!(bind = %config.bind, server_id = %config.server_id, "jray-server listening");
let router = app::router(state);
axum::serve(listener, router.into_make_service_with_connect_info::<SocketAddr>())
.with_graceful_shutdown(async move {
shutdown_signal().await;
let _ = shutdown_tx.send(true);
})
.await
.context("server error")?;
let _ = worker_handle.await;
Ok(())
}
async fn shutdown_signal() {
let ctrl_c = async {
tokio::signal::ctrl_c().await.expect("installing ctrl-c handler");
};
#[cfg(unix)]
let terminate = async {
tokio::signal::unix::signal(tokio::signal::unix::SignalKind::terminate())
.expect("installing SIGTERM handler")
.recv()
.await;
};
#[cfg(not(unix))]
let terminate = std::future::pending::<()>();
tokio::select! {
_ = ctrl_c => tracing::info!("received ctrl-c, shutting down"),
_ = terminate => tracing::info!("received SIGTERM, shutting down"),
}
}
+217
View File
@@ -0,0 +1,217 @@
//! §3 cut matching.
//!
//! Timings only transfer between identical cuts, so matching is tiered and the
//! server reports *which* tier matched — a `loose` match is meant to surface as
//! a caveat in the JRay UI rather than being applied silently.
//!
//! `video_hash` identifies a *file*, so it only ever matches an identical
//! release and can never produce a false positive; that is why it is tier one.
use crate::model::MatchTier;
/// §3: runtimes within ±2s.
pub const RUNTIME_TOLERANCE_SEC: f64 = 2.0;
/// §3: runtimes within ±30s.
pub const LOOSE_TOLERANCE_SEC: f64 = 30.0;
/// What the client tells us about its own copy.
#[derive(Debug, Clone, Default)]
pub struct ClientCut {
pub runtime_sec: Option<f64>,
pub video_hash: Option<String>,
}
impl ClientCut {
/// True when the client supplied nothing to match on, in which case §4
/// specifies a `"match": "unknown"` answer rather than a guess.
pub fn is_empty(&self) -> bool {
self.runtime_sec.is_none() && self.video_hash.is_none()
}
}
/// What the server holds.
#[derive(Debug, Clone)]
pub struct StoredCut {
pub runtime_sec: f64,
pub video_hash: Option<String>,
}
/// The outcome of comparing a client's cut against a stored one.
#[derive(Debug, Clone, Copy, PartialEq)]
pub struct CutMatch {
pub tier: MatchTier,
/// Scene offset in seconds the client must add (§3 `audio` tier). Always
/// zero for the tiers implemented here; the field exists because the plugin
/// contract is "the server returns the offset, the client applies it", and
/// enabling `audio` must not change the response shape.
pub offset_sec: f64,
}
/// Compares a client's cut against a stored one, returning the best tier that
/// fires, or `None` for "beyond that: no match; do not serve" (§3).
pub fn match_cut(client: &ClientCut, stored: &StoredCut) -> Option<CutMatch> {
// Tier 1 — same file. Checked first and unconditionally: an equal hash is
// decisive regardless of what the runtimes say.
if let (Some(c), Some(s)) = (&client.video_hash, &stored.video_hash) {
if c.eq_ignore_ascii_case(s) {
return Some(CutMatch { tier: MatchTier::Exact, offset_sec: 0.0 });
}
}
// `audio` tier would slot in here, above `runtime`, once signature coverage
// is useful (§3 recommended sequencing).
if let Some(c_rt) = client.runtime_sec {
let delta = (c_rt - stored.runtime_sec).abs();
if delta <= RUNTIME_TOLERANCE_SEC {
return Some(CutMatch { tier: MatchTier::Runtime, offset_sec: 0.0 });
}
if delta <= LOOSE_TOLERANCE_SEC {
return Some(CutMatch { tier: MatchTier::Loose, offset_sec: 0.0 });
}
// A runtime was supplied and cleared nothing — that is a definite
// no-match, not an unknown.
return None;
}
// A hash that did not match, with no runtime to fall back on, tells us
// nothing about alignment either way.
if client.video_hash.is_some() {
return None;
}
Some(CutMatch { tier: MatchTier::Unknown, offset_sec: 0.0 })
}
/// Picks the best-matching stored cut, if any clears `loose` (§4).
///
/// `candidates` is `(key, cut)`; the key is returned so the caller can identify
/// which manifest won without re-scanning.
pub fn best_match<K: Clone>(
client: &ClientCut,
candidates: &[(K, StoredCut)],
) -> Option<(K, CutMatch)> {
candidates
.iter()
.filter_map(|(k, cut)| match_cut(client, cut).map(|m| (k.clone(), m)))
.max_by(|a, b| a.1.tier.cmp(&b.1.tier))
}
#[cfg(test)]
mod tests {
use super::*;
fn stored(runtime: f64, hash: Option<&str>) -> StoredCut {
StoredCut { runtime_sec: runtime, video_hash: hash.map(str::to_string) }
}
#[test]
fn equal_video_hash_is_exact() {
let c = ClientCut {
runtime_sec: Some(6420.5),
video_hash: Some("opensubtitles:8e245d9679d31e12".into()),
};
let m = match_cut(&c, &stored(6420.5, Some("opensubtitles:8e245d9679d31e12"))).unwrap();
assert_eq!(m.tier, MatchTier::Exact);
}
#[test]
fn equal_hash_wins_even_when_runtimes_disagree() {
// The hash identifies the file; a differing stored runtime means our
// own metadata is off, not that the file is different.
let c = ClientCut {
runtime_sec: Some(6000.0),
video_hash: Some("opensubtitles:8e245d9679d31e12".into()),
};
let m = match_cut(&c, &stored(6420.5, Some("opensubtitles:8e245d9679d31e12"))).unwrap();
assert_eq!(m.tier, MatchTier::Exact);
}
#[test]
fn hash_comparison_is_case_insensitive() {
let c = ClientCut {
runtime_sec: None,
video_hash: Some("opensubtitles:8E245D9679D31E12".into()),
};
let m = match_cut(&c, &stored(6420.5, Some("opensubtitles:8e245d9679d31e12"))).unwrap();
assert_eq!(m.tier, MatchTier::Exact);
}
#[test]
fn runtime_within_two_seconds_is_runtime_tier() {
let c = ClientCut { runtime_sec: Some(6422.0), video_hash: None };
assert_eq!(match_cut(&c, &stored(6420.5, None)).unwrap().tier, MatchTier::Runtime);
}
#[test]
fn runtime_within_thirty_seconds_is_loose() {
let c = ClientCut { runtime_sec: Some(6450.0), video_hash: None };
assert_eq!(match_cut(&c, &stored(6420.5, None)).unwrap().tier, MatchTier::Loose);
}
#[test]
fn beyond_thirty_seconds_does_not_match() {
// §3: "beyond that — no match; do not serve".
let c = ClientCut { runtime_sec: Some(6500.0), video_hash: None };
assert!(match_cut(&c, &stored(6420.5, None)).is_none());
}
#[test]
fn tier_boundaries_are_inclusive() {
let c = ClientCut { runtime_sec: Some(6422.5), video_hash: None };
assert_eq!(match_cut(&c, &stored(6420.5, None)).unwrap().tier, MatchTier::Runtime);
let c = ClientCut { runtime_sec: Some(6450.5), video_hash: None };
assert_eq!(match_cut(&c, &stored(6420.5, None)).unwrap().tier, MatchTier::Loose);
}
#[test]
fn no_cut_information_yields_unknown() {
// §4: the mode a library-wide sweep uses — "does the community have
// this title at all", with alignment still to be determined.
let c = ClientCut::default();
assert!(c.is_empty());
assert_eq!(match_cut(&c, &stored(6420.5, None)).unwrap().tier, MatchTier::Unknown);
}
#[test]
fn non_matching_hash_alone_is_not_a_match() {
let c = ClientCut {
runtime_sec: None,
video_hash: Some("opensubtitles:ffffffffffffffff".into()),
};
assert!(match_cut(&c, &stored(6420.5, Some("opensubtitles:8e245d9679d31e12"))).is_none());
}
#[test]
fn non_matching_hash_falls_back_to_runtime() {
let c = ClientCut {
runtime_sec: Some(6421.0),
video_hash: Some("opensubtitles:ffffffffffffffff".into()),
};
let m = match_cut(&c, &stored(6420.5, Some("opensubtitles:8e245d9679d31e12"))).unwrap();
assert_eq!(m.tier, MatchTier::Runtime);
}
#[test]
fn best_match_prefers_the_highest_tier() {
let c = ClientCut {
runtime_sec: Some(6420.5),
video_hash: Some("opensubtitles:8e245d9679d31e12".into()),
};
let candidates = vec![
("loose", stored(6445.0, None)),
("exact", stored(9999.0, Some("opensubtitles:8e245d9679d31e12"))),
("runtime", stored(6420.0, None)),
];
let (winner, m) = best_match(&c, &candidates).unwrap();
assert_eq!(winner, "exact");
assert_eq!(m.tier, MatchTier::Exact);
}
#[test]
fn best_match_returns_none_when_nothing_clears_loose() {
let c = ClientCut { runtime_sec: Some(100.0), video_hash: None };
let candidates = vec![("a", stored(6420.5, None)), ("b", stored(3000.0, None))];
assert!(best_match(&c, &candidates).is_none());
}
}
+307
View File
@@ -0,0 +1,307 @@
//! Jmanifest wire types (§2).
//!
//! **`#[serde(deny_unknown_fields)]` on every struct is the §6 stage 2
//! enforcement mechanism.** "No additional fields anywhere" is a property of
//! these type definitions rather than of validator code that could omit a
//! field, so an unrecognised key at any nesting level fails to parse. That is
//! also what makes the §9 `movie`/`jellyfin_id` strip verifiable: a client that
//! forgets gets a hard `400` naming the field, rather than quietly publishing a
//! contributor's directory layout.
//!
//! Deeper semantic checks — bounds, character classes, path-shaped strings —
//! live in [`crate::validate`]. Parsing rejects *shape*; validation rejects
//! *content*.
use serde::{Deserialize, Serialize};
/// Cut-match tier (§3). Ordered worst-to-best so derived `Ord` ranks them.
#[derive(Debug, Clone, Copy, PartialEq, Eq, PartialOrd, Ord, Serialize, Deserialize)]
#[serde(rename_all = "lowercase")]
pub enum MatchTier {
/// No cut information was supplied, so alignment is unknown (§4).
Unknown,
/// Audio 0.60–0.85, or runtimes within ±30s. Caveat in UI.
Loose,
/// Runtimes within ±2s.
Runtime,
/// Audio score ≥ 0.85; ranks above `runtime` because it is content-derived.
Audio,
/// `video_hash` equal — same file.
Exact,
}
impl MatchTier {
pub fn as_str(self) -> &'static str {
match self {
MatchTier::Unknown => "unknown",
MatchTier::Loose => "loose",
MatchTier::Runtime => "runtime",
MatchTier::Audio => "audio",
MatchTier::Exact => "exact",
}
}
}
#[derive(Debug, Clone, Copy, PartialEq, Eq, Serialize, Deserialize)]
#[serde(rename_all = "lowercase")]
pub enum IdentityType {
Movie,
Episode,
}
/// What the work is (§2 terminology: *title identity*).
///
/// Movie and episode coordinates share one struct because `deny_unknown_fields`
/// with `#[serde(untagged)]` alternatives produces unhelpful error messages;
/// the discriminant is checked in [`crate::validate`], which can name the
/// offending field.
#[derive(Debug, Clone, Deserialize, Serialize)]
#[serde(deny_unknown_fields)]
pub struct Identity {
#[serde(rename = "type")]
pub kind: IdentityType,
// Movie coordinates.
#[serde(default, skip_serializing_if = "Option::is_none")]
pub tmdb_id: Option<String>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub imdb_id: Option<String>,
// Episode coordinates.
#[serde(default, skip_serializing_if = "Option::is_none")]
pub series_tmdb_id: Option<String>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub series_imdb_id: Option<String>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub season: Option<i64>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub episode: Option<i64>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub title: Option<String>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub year: Option<i64>,
}
impl Identity {
/// The TMDB id used as the lookup key, whichever coordinate carries it.
pub fn effective_tmdb_id(&self) -> Option<&str> {
match self.kind {
IdentityType::Movie => self.tmdb_id.as_deref(),
IdentityType::Episode => self.series_tmdb_id.as_deref(),
}
}
pub fn effective_imdb_id(&self) -> Option<&str> {
match self.kind {
IdentityType::Movie => self.imdb_id.as_deref(),
IdentityType::Episode => self.series_imdb_id.as_deref(),
}
}
}
/// Which encode/edit the timings apply to (§2 terminology: *cut fingerprint*).
#[derive(Debug, Clone, Deserialize, Serialize)]
#[serde(deny_unknown_fields)]
pub struct Cut {
/// **Required** — the decoded duration of the media the timings came from.
/// The primary alignment guard (§2).
pub runtime_sec: f64,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub container_duration_sec: Option<f64>,
/// Optional but strongly preferred. OpenSubtitles hash (§3).
#[serde(default, skip_serializing_if = "Option::is_none")]
pub video_hash: Option<String>,
/// Optional; version-prefixed spectral-peak signature (§3, UR-9).
///
/// Accepted and stored by this build; `audio`-tier matching is enabled once
/// coverage is useful, per §3's recommended sequencing.
#[serde(default, skip_serializing_if = "Option::is_none")]
pub audio_signature: Option<String>,
}
/// How well the contributor's gallery could discriminate (§2, §7).
///
/// The strongest available quality signal between two otherwise comparable
/// manifests: a `Global` gallery had to distinguish its actors from every other
/// actor in the contributor's library, whereas a `Limited` one only had to
/// distinguish them from this title's own cast.
#[derive(Debug, Clone, Copy, PartialEq, Eq, PartialOrd, Ord, Deserialize, Serialize)]
#[serde(rename_all = "lowercase")]
pub enum GalleryScope {
/// Built from this title's cast alone.
Limited,
/// Built from the whole library. The default upstream.
Global,
}
impl GalleryScope {
pub fn as_str(self) -> &'static str {
match self {
GalleryScope::Limited => "limited",
GalleryScope::Global => "global",
}
}
}
/// Extraction parameters, carried for provenance and ranking.
///
/// Note there is no `anneal_sec`: it was **withdrawn** in the SR-003 schema
/// bump, because presence now follows track extent — a track survives its own
/// gaps, so there is nothing to anneal (`scene-actor-extraction` AR-012/AR-013).
/// `deny_unknown_fields` therefore makes its presence a hard parse error rather
/// than something silently ignored, which is deliberate: a manifest still
/// carrying it was produced by a pipeline whose window semantics differ from
/// what this server now assumes.
#[derive(Debug, Clone, Deserialize, Serialize)]
#[serde(deny_unknown_fields)]
pub struct Extraction {
#[serde(default, skip_serializing_if = "Option::is_none")]
pub sample_fps: Option<f64>,
/// The re-acquisition timeout that shapes window extent. Successor to the
/// withdrawn `anneal_sec`.
#[serde(default, skip_serializing_if = "Option::is_none")]
pub extinction_sec: Option<f64>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub pipeline_version: Option<String>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub gallery_size: Option<i64>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub gallery_scope: Option<GalleryScope>,
}
/// One actor's timeline.
///
/// Note there is no `jellyfin_id` field: `deny_unknown_fields` means its
/// presence is a parse error, which is exactly the §2/§6 requirement that it be
/// *rejected on upload* rather than merely ignored on download.
#[derive(Debug, Clone, Deserialize, Serialize)]
#[serde(deny_unknown_fields)]
pub struct Actor {
/// Sent on upload for matching, but **not persisted** — the server resolves
/// each actor to a TMDB person id and serves names from its own TMDB-derived
/// table (§2, §5a). On download this is server-authoritative.
#[serde(default, skip_serializing_if = "Option::is_none")]
pub name: Option<String>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub imdb_id: Option<String>,
/// The **primary** actor join key (§2, §6 stage 3).
#[serde(default, skip_serializing_if = "Option::is_none")]
pub tmdb_id: Option<String>,
/// `[start_sec, end_sec]` inclusive, sorted.
pub scenes: Vec<[f64; 2]>,
}
/// One shareable actor timeline for one cut of one title (§2).
#[derive(Debug, Clone, Deserialize, Serialize)]
#[serde(deny_unknown_fields)]
pub struct Jmanifest {
pub jmanifest_version: u32,
pub identity: Identity,
pub cut: Cut,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub extraction: Option<Extraction>,
pub actors: Vec<Actor>,
}
/// Series-level coordinates for a bundle envelope (§2).
#[derive(Debug, Clone, Deserialize, Serialize)]
#[serde(deny_unknown_fields)]
pub struct SeriesRef {
#[serde(default, skip_serializing_if = "Option::is_none")]
pub series_tmdb_id: Option<String>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub series_imdb_id: Option<String>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub title: Option<String>,
}
/// A thin wrapper, not a new format (§2). Bundles are a transfer convenience,
/// never a storage unit — each episode is stored and moderated individually.
#[derive(Debug, Clone, Deserialize, Serialize)]
#[serde(deny_unknown_fields)]
pub struct SeriesBundle {
pub jmanifest_version: u32,
pub series: SeriesRef,
pub episodes: Vec<Jmanifest>,
/// Present on responses only; ignored on upload.
#[serde(default, skip_serializing_if = "Option::is_none")]
pub coverage: Option<Coverage>,
}
#[derive(Debug, Clone, Deserialize, Serialize)]
#[serde(deny_unknown_fields)]
pub struct Coverage {
pub episodes_available: usize,
pub seasons: Vec<i64>,
}
/// The current `jmanifest_version` this server speaks (§2).
pub const JMANIFEST_VERSION: u32 = 1;
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn unknown_field_at_top_level_is_rejected() {
let json = r#"{"jmanifest_version":1,"identity":{"type":"movie","tmdb_id":"1"},
"cut":{"runtime_sec":100.0},"actors":[],"surprise":"x"}"#;
let err = serde_json::from_str::<Jmanifest>(json).unwrap_err().to_string();
assert!(err.contains("surprise"), "error should name the field: {err}");
}
#[test]
fn jellyfin_id_on_an_actor_is_a_parse_error() {
// §2: `actors[].jellyfin_id` must not appear. `deny_unknown_fields`
// makes this structural rather than a validator's responsibility.
let json = r#"{"jmanifest_version":1,"identity":{"type":"movie","tmdb_id":"1"},
"cut":{"runtime_sec":100.0},
"actors":[{"name":"A","tmdb_id":"2","jellyfin_id":"guid","scenes":[]}]}"#;
let err = serde_json::from_str::<Jmanifest>(json).unwrap_err().to_string();
assert!(err.contains("jellyfin_id"), "error should name the field: {err}");
}
#[test]
fn movie_path_field_is_a_parse_error() {
let json = r#"{"jmanifest_version":1,"movie":"/data/movies/x.mkv",
"identity":{"type":"movie","tmdb_id":"1"},
"cut":{"runtime_sec":100.0},"actors":[]}"#;
let err = serde_json::from_str::<Jmanifest>(json).unwrap_err().to_string();
assert!(err.contains("movie"), "error should name the field: {err}");
}
#[test]
fn unknown_field_nested_in_cut_is_rejected() {
let json = r#"{"jmanifest_version":1,"identity":{"type":"movie","tmdb_id":"1"},
"cut":{"runtime_sec":100.0,"payload":"x"},"actors":[]}"#;
assert!(serde_json::from_str::<Jmanifest>(json).is_err());
}
#[test]
fn spec_example_manifest_parses() {
let json = r#"{
"jmanifest_version": 1,
"identity": { "type": "movie", "tmdb_id": "504172", "imdb_id": "tt4686844",
"title": "The Death of Stalin", "year": 2017 },
"cut": { "runtime_sec": 6420.5, "container_duration_sec": 6420.5,
"video_hash": "opensubtitles:8e245d9679d31e12" },
"extraction": { "sample_fps": 5, "extinction_sec": 12,
"pipeline_version": "scene-actor-extraction 0.4.1",
"gallery_size": 1820, "gallery_scope": "global" },
"actors": [ { "name": "Steve Buscemi", "imdb_id": "nm0000114", "tmdb_id": "884",
"scenes": [[191.6, 209.2], [438.2, 465.6]] } ]
}"#;
let m: Jmanifest = serde_json::from_str(json).unwrap();
assert_eq!(m.actors.len(), 1);
assert_eq!(m.identity.effective_tmdb_id(), Some("504172"));
}
#[test]
fn tiers_order_audio_above_runtime() {
// §3: `audio` ranks above `runtime` because it is content-derived.
assert!(MatchTier::Audio > MatchTier::Runtime);
assert!(MatchTier::Exact > MatchTier::Audio);
assert!(MatchTier::Runtime > MatchTier::Loose);
}
}
+217
View File
@@ -0,0 +1,217 @@
//! §5 rate limiting.
//!
//! A fixed-window counter keyed on `(token_or_ip, surface)`, held in process
//! memory — no external counter store. §5 is explicit that a sliding window is
//! not worth the complexity at this volume, and that counters resetting on
//! restart is acceptable for abuse throttling.
//!
//! Read limits are applied *behind* the CDN cache, so a cache hit costs a client
//! nothing against its budget — that is a deployment property (§8), not
//! something this module can enforce.
use std::collections::HashMap;
use std::sync::Mutex;
use std::time::{Duration, Instant};
/// The rate-limited surfaces of §5. Distinct from routes: the batch and single
/// forms of `exists` are separate surfaces with separate budgets.
#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash)]
pub enum Surface {
ExistsSingle,
ExistsBatch,
ManifestFetch,
SeriesFetch,
ManifestUpload,
BundleUpload,
Report,
Search,
}
impl Surface {
/// Requests per hour, per §5's table.
pub fn limit(self) -> u32 {
match self {
Surface::ExistsSingle => 600,
Surface::ExistsBatch => 60,
Surface::ManifestFetch => 300,
Surface::SeriesFetch => 120,
Surface::ManifestUpload => 100,
Surface::BundleUpload => 20,
Surface::Report => 20,
Surface::Search => 60,
}
}
pub fn as_str(self) -> &'static str {
match self {
Surface::ExistsSingle => "exists",
Surface::ExistsBatch => "exists_batch",
Surface::ManifestFetch => "manifest_fetch",
Surface::SeriesFetch => "series_fetch",
Surface::ManifestUpload => "manifest_upload",
Surface::BundleUpload => "bundle_upload",
Surface::Report => "report",
Surface::Search => "search",
}
}
}
const WINDOW: Duration = Duration::from_secs(3600);
/// Headers §5 requires on every rate-limited response.
#[derive(Debug, Clone, Copy)]
pub struct Quota {
pub limit: u32,
pub remaining: u32,
/// Seconds until the window resets.
pub reset: u64,
}
#[derive(Debug, Clone, Copy)]
struct Window {
started: Instant,
count: u32,
}
pub struct RateLimiter {
windows: Mutex<HashMap<(String, Surface), Window>>,
}
impl Default for RateLimiter {
fn default() -> Self {
Self::new()
}
}
impl RateLimiter {
pub fn new() -> Self {
Self { windows: Mutex::new(HashMap::new()) }
}
/// Records one request against `(key, surface)`.
///
/// `Ok(quota)` when within budget, `Err(quota)` when the limit is exceeded —
/// in which case the caller returns `429` with `Retry-After` set from
/// `quota.reset`. A rejected request does **not** increment the counter, so a
/// client hammering a closed window cannot extend its own lockout.
pub fn check(&self, key: &str, surface: Surface) -> Result<Quota, Quota> {
self.check_at(key, surface, Instant::now())
}
fn check_at(&self, key: &str, surface: Surface, now: Instant) -> Result<Quota, Quota> {
let limit = surface.limit();
let mut windows = self.windows.lock().expect("rate limiter poisoned");
// Opportunistic eviction of stale windows, so an IP-keyed map cannot
// grow without bound behind CGNAT.
if windows.len() > 10_000 {
windows.retain(|_, w| now.duration_since(w.started) < WINDOW);
}
let entry =
windows.entry((key.to_string(), surface)).or_insert(Window { started: now, count: 0 });
let elapsed = now.duration_since(entry.started);
if elapsed >= WINDOW {
*entry = Window { started: now, count: 0 };
}
let reset = WINDOW.saturating_sub(now.duration_since(entry.started)).as_secs();
if entry.count >= limit {
return Err(Quota { limit, remaining: 0, reset });
}
entry.count += 1;
Ok(Quota { limit, remaining: limit - entry.count, reset })
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn allows_up_to_the_limit_then_rejects() {
let rl = RateLimiter::new();
let limit = Surface::Report.limit();
for i in 0..limit {
let q = rl.check("ip", Surface::Report).expect("within budget");
assert_eq!(q.remaining, limit - i - 1);
}
let q = rl.check("ip", Surface::Report).expect_err("over budget");
assert_eq!(q.remaining, 0);
}
#[test]
fn surfaces_have_independent_budgets() {
let rl = RateLimiter::new();
for _ in 0..Surface::BundleUpload.limit() {
rl.check("t", Surface::BundleUpload).unwrap();
}
assert!(rl.check("t", Surface::BundleUpload).is_err());
// §5: a bundle counts as a single write against its own limit, and must
// not consume the single-manifest budget.
assert!(rl.check("t", Surface::ManifestUpload).is_ok());
}
#[test]
fn keys_are_independent() {
let rl = RateLimiter::new();
for _ in 0..Surface::Report.limit() {
rl.check("a", Surface::Report).unwrap();
}
assert!(rl.check("a", Surface::Report).is_err());
assert!(rl.check("b", Surface::Report).is_ok());
}
#[test]
fn window_resets_after_an_hour() {
let rl = RateLimiter::new();
let t0 = Instant::now();
for _ in 0..Surface::Report.limit() {
rl.check_at("ip", Surface::Report, t0).unwrap();
}
assert!(rl.check_at("ip", Surface::Report, t0).is_err());
// Still closed just inside the window.
assert!(rl.check_at("ip", Surface::Report, t0 + Duration::from_secs(3599)).is_err());
// Open again once it rolls over.
assert!(rl.check_at("ip", Surface::Report, t0 + Duration::from_secs(3600)).is_ok());
}
#[test]
fn rejected_requests_do_not_extend_the_lockout() {
let rl = RateLimiter::new();
let t0 = Instant::now();
for _ in 0..Surface::Report.limit() {
rl.check_at("ip", Surface::Report, t0).unwrap();
}
// Hammer the closed window; the counter must not keep climbing, so the
// window still expires on schedule.
for _ in 0..50 {
assert!(rl.check_at("ip", Surface::Report, t0 + Duration::from_secs(10)).is_err());
}
assert!(rl.check_at("ip", Surface::Report, t0 + Duration::from_secs(3600)).is_ok());
}
#[test]
fn reset_counts_down_within_the_window() {
let rl = RateLimiter::new();
let t0 = Instant::now();
let q = rl.check_at("ip", Surface::ExistsSingle, t0).unwrap();
assert_eq!(q.reset, 3600);
let q = rl.check_at("ip", Surface::ExistsSingle, t0 + Duration::from_secs(600)).unwrap();
assert_eq!(q.reset, 3000);
}
#[test]
fn limits_match_the_spec_table() {
assert_eq!(Surface::ExistsSingle.limit(), 600);
assert_eq!(Surface::ExistsBatch.limit(), 60);
assert_eq!(Surface::ManifestFetch.limit(), 300);
assert_eq!(Surface::SeriesFetch.limit(), 120);
assert_eq!(Surface::ManifestUpload.limit(), 100);
assert_eq!(Surface::BundleUpload.limit(), 20);
assert_eq!(Surface::Report.limit(), 20);
assert_eq!(Surface::Search.limit(), 60);
}
}
+102
View File
@@ -0,0 +1,102 @@
//! Shared application state, and the cross-cutting request concerns (§5 rate
//! limiting, §5a token resolution) that every handler needs.
use std::net::SocketAddr;
use std::sync::Arc;
use axum::extract::ConnectInfo;
use axum::http::{HeaderMap, HeaderValue};
use axum::response::Response;
use crate::auth;
use crate::config::Config;
use crate::db::{repo, Db};
use crate::error::{ApiError, ApiResult};
use crate::ratelimit::{Quota, RateLimiter, Surface};
use crate::tmdb::TmdbClient;
/// The connection's peer address, when the server was started with connect-info.
///
/// A dedicated extractor rather than `ConnectInfo<SocketAddr>` directly, because
/// this must not be a *hard* requirement: a router used without
/// `into_make_service_with_connect_info` — as in tests — has no peer address, and
/// a handler that fails to extract would be a routing error rather than degrading
/// to header-only attribution.
pub struct PeerIp(pub Option<std::net::IpAddr>);
impl<S> axum::extract::FromRequestParts<S> for PeerIp
where
S: Send + Sync,
{
type Rejection = std::convert::Infallible;
async fn from_request_parts(
parts: &mut axum::http::request::Parts,
_state: &S,
) -> Result<Self, Self::Rejection> {
Ok(PeerIp(
parts.extensions.get::<ConnectInfo<SocketAddr>>().map(|ConnectInfo(addr)| addr.ip()),
))
}
}
#[derive(Clone)]
pub struct AppState {
pub db: Db,
pub config: Arc<Config>,
pub limiter: Arc<RateLimiter>,
pub tmdb: Arc<TmdbClient>,
}
impl AppState {
/// Resolves the client IP for rate-limiting and attribution, honouring
/// `X-Forwarded-For` only from a configured proxy (§8).
pub fn client_ip(&self, headers: &HeaderMap, peer: Option<std::net::IpAddr>) -> String {
auth::client_ip(headers, peer, &self.config.trusted_proxies)
}
/// §5: limits are per token where one is present, otherwise per source IP.
pub fn check_limit(&self, key: &str, surface: Surface) -> ApiResult<Quota> {
self.limiter.check(key, surface).map_err(|q| {
tracing::debug!(surface = surface.as_str(), "rate limited");
ApiError::RateLimited { retry_after: q.reset.max(1) }
})
}
/// Resolves a bearer token to a contributor (§5a).
///
/// A token is an anonymous bearer capability, not an account: the only state
/// behind it is the per-token counters used for rate-limiting attribution and
/// automatic revocation.
pub async fn require_contributor(&self, headers: &HeaderMap) -> ApiResult<repo::Contributor> {
let token = auth::bearer_token(headers).ok_or(ApiError::Unauthorized)?;
let hash = auth::hash_token(&token);
let found = self
.db
.read(move |conn| repo::contributor_by_token_hash(conn, &hash))
.await
.map_err(ApiError::Internal)?;
match found {
Some(c) if !c.revoked => Ok(c),
// A revoked token is indistinguishable from an unknown one to the
// caller; there is nothing useful to disclose.
_ => Err(ApiError::Unauthorized),
}
}
}
/// Attaches the §5 rate-limit headers to a response.
pub fn with_quota_headers(mut resp: Response, quota: Quota) -> Response {
let h = resp.headers_mut();
insert_num(h, "x-ratelimit-limit", quota.limit as u64);
insert_num(h, "x-ratelimit-remaining", quota.remaining as u64);
insert_num(h, "x-ratelimit-reset", quota.reset);
resp
}
fn insert_num(headers: &mut HeaderMap, name: &'static str, value: u64) {
if let Ok(v) = HeaderValue::from_str(&value.to_string()) {
headers.insert(name, v);
}
}
+226
View File
@@ -0,0 +1,226 @@
//! TMDB client for the §6 stage 3 cast cross-check.
//!
//! §5a's Threat 2 defence rests entirely on the attacker not controlling TMDB:
//! to make a prank manifest pass, they would need those performers to be
//! credited cast on that title in TMDB, which means vandalising a separate,
//! moderated system.
//!
//! Responses are cached for 24h (§6) so a burst of episode uploads for one
//! series costs a single upstream call, and so the server stays within TMDB's
//! own rate limits.
use std::time::Duration;
use serde::Deserialize;
/// A credited cast member, reduced to what the check needs.
#[derive(Debug, Clone, Deserialize)]
pub struct CastMember {
pub id: u64,
#[serde(default)]
pub name: String,
#[serde(default)]
pub adult: bool,
}
#[derive(Debug, Clone, Default, Deserialize)]
pub struct Credits {
#[serde(default)]
pub cast: Vec<CastMember>,
/// Present on episode credits.
#[serde(default)]
pub guest_stars: Vec<CastMember>,
}
impl Credits {
/// Cast plus guest stars — the union §6 specifies for episodes.
pub fn all(&self) -> impl Iterator<Item = &CastMember> {
self.cast.iter().chain(self.guest_stars.iter())
}
}
#[derive(Debug, Clone, Default, Deserialize)]
pub struct TitleDetails {
#[serde(default)]
pub adult: bool,
#[serde(default)]
pub title: Option<String>,
#[serde(default)]
pub name: Option<String>,
}
/// A failure that should be retried rather than treated as a verdict.
///
/// §6: "TMDB unreachable / rate-limited → retry with backoff; stays unlisted,
/// not rejected." Distinguishing this from "TMDB has no credits" is essential —
/// conflating them would reject honest manifests during an outage.
#[derive(Debug, thiserror::Error)]
pub enum TmdbError {
#[error("tmdb transport error: {0}")]
Transport(String),
#[error("tmdb rate limited")]
RateLimited,
#[error("tmdb server error: {0}")]
ServerError(u16),
/// The id genuinely does not exist upstream.
#[error("tmdb resource not found")]
NotFound,
#[error("tmdb response was not understood: {0}")]
Malformed(String),
#[error("no tmdb api key configured")]
NotConfigured,
}
impl TmdbError {
/// True when the job should be rescheduled rather than resolved.
pub fn is_retryable(&self) -> bool {
matches!(
self,
TmdbError::Transport(_)
| TmdbError::RateLimited
| TmdbError::ServerError(_)
| TmdbError::NotConfigured
)
}
}
#[derive(Clone)]
pub struct TmdbClient {
http: reqwest::Client,
base_url: String,
api_key: Option<String>,
}
impl TmdbClient {
pub fn new(base_url: String, api_key: Option<String>) -> Self {
let http = reqwest::Client::builder()
.timeout(Duration::from_secs(15))
.user_agent(concat!("jray-server/", env!("CARGO_PKG_VERSION")))
.build()
.expect("building reqwest client");
Self { http, base_url, api_key }
}
pub fn is_configured(&self) -> bool {
self.api_key.is_some()
}
async fn get<T: serde::de::DeserializeOwned>(&self, path: &str) -> Result<T, TmdbError> {
let key = self.api_key.as_deref().ok_or(TmdbError::NotConfigured)?;
let url =
format!("{}/{}", self.base_url.trim_end_matches('/'), path.trim_start_matches('/'));
let resp = self
.http
.get(&url)
.query(&[("api_key", key)])
.send()
.await
.map_err(|e| TmdbError::Transport(e.to_string()))?;
let status = resp.status();
if status == reqwest::StatusCode::NOT_FOUND {
return Err(TmdbError::NotFound);
}
if status == reqwest::StatusCode::TOO_MANY_REQUESTS {
return Err(TmdbError::RateLimited);
}
if status.is_server_error() {
return Err(TmdbError::ServerError(status.as_u16()));
}
if !status.is_success() {
return Err(TmdbError::Malformed(format!("unexpected status {status}")));
}
let body = resp.text().await.map_err(|e| TmdbError::Transport(e.to_string()))?;
serde_json::from_str(&body).map_err(|e| TmdbError::Malformed(e.to_string()))
}
pub async fn movie_credits(&self, tmdb_id: &str) -> Result<Credits, TmdbError> {
self.get(&format!("movie/{tmdb_id}/credits")).await
}
pub async fn movie_details(&self, tmdb_id: &str) -> Result<TitleDetails, TmdbError> {
self.get(&format!("movie/{tmdb_id}")).await
}
pub async fn series_credits(&self, series_tmdb_id: &str) -> Result<Credits, TmdbError> {
// Aggregate credits carry recurring cast TMDB lists only at series level.
self.get(&format!("tv/{series_tmdb_id}/aggregate_credits")).await
}
pub async fn episode_credits(
&self,
series_tmdb_id: &str,
season: i64,
episode: i64,
) -> Result<Credits, TmdbError> {
self.get(&format!("tv/{series_tmdb_id}/season/{season}/episode/{episode}/credits")).await
}
pub async fn series_details(&self, series_tmdb_id: &str) -> Result<TitleDetails, TmdbError> {
self.get(&format!("tv/{series_tmdb_id}")).await
}
}
/// Fetches a person's details, used by the §5a category guard.
#[derive(Debug, Clone, Default, Deserialize)]
pub struct PersonDetails {
#[serde(default)]
pub adult: bool,
#[serde(default)]
pub name: String,
}
impl TmdbClient {
pub async fn person(&self, tmdb_person_id: u64) -> Result<PersonDetails, TmdbError> {
self.get(&format!("person/{tmdb_person_id}")).await
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn credits_union_covers_cast_and_guest_stars() {
// §6: for episodes the check runs against the union of per-episode
// credits (cast + guest stars) and series aggregate credits.
let c: Credits = serde_json::from_str(
r#"{"cast":[{"id":1,"name":"A"}],"guest_stars":[{"id":2,"name":"B"}]}"#,
)
.unwrap();
let ids: Vec<u64> = c.all().map(|m| m.id).collect();
assert_eq!(ids, vec![1, 2]);
}
#[test]
fn credits_tolerate_missing_and_extra_fields() {
// TMDB adds fields freely; our own strictness applies to *uploads*, not
// to a trusted upstream we merely read.
let c: Credits =
serde_json::from_str(r#"{"cast":[{"id":1,"unexpected":true}],"id":99}"#).unwrap();
assert_eq!(c.cast.len(), 1);
assert_eq!(c.cast[0].name, "");
assert!(c.guest_stars.is_empty());
}
#[test]
fn transport_and_rate_limit_are_retryable_but_not_found_is_not() {
// The distinction that keeps an outage from rejecting honest uploads.
assert!(TmdbError::Transport("x".into()).is_retryable());
assert!(TmdbError::RateLimited.is_retryable());
assert!(TmdbError::ServerError(503).is_retryable());
assert!(TmdbError::NotConfigured.is_retryable());
assert!(!TmdbError::NotFound.is_retryable());
assert!(!TmdbError::Malformed("x".into()).is_retryable());
}
#[tokio::test]
async fn unconfigured_client_reports_retryable_failure() {
let c = TmdbClient::new("http://127.0.0.1:1".into(), None);
assert!(!c.is_configured());
let err = c.movie_credits("1").await.unwrap_err();
assert!(err.is_retryable(), "missing key must hold uploads pending, not reject them");
}
}
+1055
View File
File diff suppressed because it is too large Load Diff
+494
View File
@@ -0,0 +1,494 @@
//! Background worker for the §6 stage 3 cast check.
//!
//! §8: this runs as a Tokio background task in the same binary, with the job
//! queue as a SQLite table so state survives restart — replacing an external
//! broker entirely. The check needs an outbound TMDB call and so cannot run
//! inside the request without coupling upload latency to a third party (§6).
use std::sync::Arc;
use std::time::Duration;
use anyhow::Context;
use crate::castcheck::{self, SubmittedActor, Verdict};
use crate::db::{repo, Db};
use crate::ingest::{CastCheckJob, JOB_CAST_CHECK};
use crate::tmdb::{CastMember, Credits, TmdbClient, TmdbError};
/// §6: TMDB responses are cached for 24h, so a burst of episode uploads for one
/// series costs a single upstream call.
const CACHE_TTL: Duration = Duration::from_secs(24 * 3600);
/// Cap on retry backoff for a persistent TMDB outage.
const MAX_BACKOFF_SECS: u64 = 3600;
pub struct Worker {
pub db: Db,
pub tmdb: Arc<TmdbClient>,
pub batch: usize,
pub poll_interval: Duration,
}
impl Worker {
/// Runs until `shutdown` resolves.
pub async fn run(self, mut shutdown: tokio::sync::watch::Receiver<bool>) {
// A process that died mid-job would otherwise leave work stranded.
match self.db.write(repo::release_all_leases).await {
Ok(n) if n > 0 => tracing::info!(released = n, "released stranded job leases"),
Ok(_) => {}
Err(e) => tracing::error!(error = ?e, "failed to release job leases at startup"),
}
loop {
tokio::select! {
_ = shutdown.changed() => {
tracing::info!("worker shutting down");
return;
}
_ = tokio::time::sleep(self.poll_interval) => {
if let Err(e) = self.tick().await {
tracing::error!(error = ?e, "worker tick failed");
}
}
}
}
}
async fn tick(&self) -> anyhow::Result<()> {
let now = now_iso();
let batch = self.batch;
let leased_at = now.clone();
let jobs = self.db.write(move |tx| repo::lease_jobs(tx, &leased_at, batch)).await?;
for job in jobs {
let result = match job.kind.as_str() {
JOB_CAST_CHECK => self.run_cast_check(&job.payload).await,
other => {
tracing::warn!(kind = other, "unknown job kind, dropping");
Ok(())
}
};
let job_id = job.id.clone();
match result {
Ok(()) => {
self.db.write(move |tx| repo::delete_job(tx, &job_id)).await?;
}
Err(JobError::Retry(msg)) => {
// §6: TMDB unreachable or rate-limited means retry with
// backoff; the manifest stays unlisted, not rejected.
let delay = backoff_secs(job.attempts);
let run_after = iso_in(delay);
tracing::warn!(job = %job_id, attempts = job.attempts, delay, reason = %msg,
"rescheduling job");
self.db
.write(move |tx| repo::reschedule_job(tx, &job_id, &run_after, &msg))
.await?;
}
Err(JobError::Fatal(e)) => {
tracing::error!(job = %job_id, error = ?e, "dropping job after fatal error");
self.db.write(move |tx| repo::delete_job(tx, &job_id)).await?;
}
}
}
Ok(())
}
async fn run_cast_check(&self, payload: &str) -> Result<(), JobError> {
let job: CastCheckJob =
serde_json::from_str(payload).map_err(|e| JobError::Fatal(e.into()))?;
let manifest_id = job.manifest_id;
// Load what the check needs.
let mid = manifest_id.clone();
let loaded = self
.db
.read(move |conn| {
let Some(m) = repo::manifest_by_id(conn, &mid)? else { return Ok(None) };
let title = conn
.query_row(
"SELECT kind, tmdb_id, imdb_id, adult, certification FROM titles WHERE id = ?1",
rusqlite::params![m.title_id],
|r| {
Ok((
r.get::<_, String>(0)?,
r.get::<_, Option<String>>(1)?,
r.get::<_, Option<String>>(2)?,
r.get::<_, i64>(3)? != 0,
r.get::<_, Option<String>>(4)?,
))
},
)
.map_err(anyhow::Error::from)?;
let actor_ids = repo::manifest_actor_ids(conn, &mid)?;
Ok(Some((m, title, actor_ids)))
})
.await
.map_err(JobError::Fatal)?;
// The manifest may have been deleted (contributor revoked, §5a) between
// enqueue and now; that is not an error.
let Some((manifest, (kind, tmdb_id, _imdb_id, title_adult, certification), actor_ids)) =
loaded
else {
return Ok(());
};
if manifest.status != "pending" {
return Ok(());
}
let Some(tmdb_id) = tmdb_id else {
// No TMDB id means the cast check cannot run at all. §6 treats absent
// reference data as flagged, not rejected.
self.finalise(&manifest_id, Verdict::Flagged, 0.0, Some("no_tmdb_id"), &[], &[])
.await
.map_err(JobError::Fatal)?;
return Ok(());
};
let credits = self
.credits_for(&kind, &tmdb_id, manifest.season, manifest.episode)
.await
.map_err(|e| {
if e.is_retryable() {
JobError::Retry(e.to_string())
} else {
// A genuinely absent title is a verdict, not a transport
// failure — handled below via empty credits.
JobError::Retry(format!("non-retryable tmdb error treated as absent: {e}"))
}
});
let credits = match credits {
Ok(c) => c,
Err(JobError::Retry(msg)) if msg.starts_with("non-retryable") => {
tracing::info!(manifest = %manifest_id, "tmdb has no such title; flagging");
Credits::default()
}
Err(e) => return Err(e),
};
let reference: Vec<CastMember> = credits.all().cloned().collect();
let submitted: Vec<SubmittedActor> = actor_ids
.iter()
.map(|id| SubmittedActor { tmdb_id: Some(*id), imdb_id: None, name: None })
.collect();
let mut outcome = castcheck::evaluate(&submitted, &reference);
// §5a layer 1 — category guard.
if let Some(offender) = castcheck::category_guard_violation(&outcome.matched, title_adult) {
tracing::warn!(manifest = %manifest_id, person = offender,
"category guard: adult-flagged performer on a non-adult title");
outcome.verdict = Verdict::Rejected;
outcome.reason = Some("category_guard".into());
}
// §5a layer 2 — age-appropriateness guard.
if let Some(cert) = &certification {
if castcheck::is_childrens_certification(cert) {
castcheck::apply_childrens_guard(&mut outcome, submitted.len());
}
}
self.finalise(
&manifest_id,
outcome.verdict,
outcome.ratio,
outcome.reason.as_deref(),
&outcome.matched,
&outcome.unmatched_person_ids,
)
.await
.map_err(JobError::Fatal)?;
Ok(())
}
/// Fetches credits, using the 24h cache (§6).
///
/// For episodes this is the **union** of TMDB's per-episode credits (cast +
/// guest stars) and the series' aggregate credits: per-episode alone would
/// reject recurring cast TMDB lists only at series level, series-wide alone
/// would reject legitimate guest stars.
async fn credits_for(
&self,
kind: &str,
tmdb_id: &str,
season: Option<i64>,
episode: Option<i64>,
) -> Result<Credits, TmdbError> {
if kind == "movie" {
return self.cached("movie", tmdb_id, || self.tmdb.movie_credits(tmdb_id)).await;
}
let series = self.cached("series", tmdb_id, || self.tmdb.series_credits(tmdb_id)).await?;
let mut combined = series;
if let (Some(s), Some(e)) = (season, episode) {
let key = format!("{tmdb_id}:{s}:{e}");
match self.cached("episode", &key, || self.tmdb.episode_credits(tmdb_id, s, e)).await {
Ok(ep) => {
combined.cast.extend(ep.cast);
combined.guest_stars.extend(ep.guest_stars);
}
// A missing episode entry is normal; the series set still applies.
Err(TmdbError::NotFound) => {}
Err(e) if e.is_retryable() => return Err(e),
Err(e) => tracing::warn!(error = ?e, "ignoring episode credits error"),
}
}
Ok(combined)
}
async fn cached<F, Fut>(&self, kind: &str, key: &str, fetch: F) -> Result<Credits, TmdbError>
where
F: FnOnce() -> Fut,
Fut: std::future::Future<Output = Result<Credits, TmdbError>>,
{
let (k, kk) = (key.to_string(), kind.to_string());
let cached = self
.db
.read(move |conn| repo::cached_credits(conn, &k, &kk))
.await
.map_err(|e| TmdbError::Transport(e.to_string()))?;
if let Some((json, fetched_at)) = cached {
if !is_stale(&fetched_at, CACHE_TTL) {
if let Ok(c) = serde_json::from_str::<Credits>(&json) {
return Ok(c);
}
}
}
let fresh = fetch().await?;
let json = serde_json::to_string(&SerializableCredits::from(&fresh))
.map_err(|e| TmdbError::Malformed(e.to_string()))?;
let (k, kk, now) = (key.to_string(), kind.to_string(), now_iso());
let _ = self.db.write(move |tx| repo::put_credits(tx, &k, &kk, &json, &now)).await;
Ok(fresh)
}
/// Applies the verdict: resolves names into `people`, drops unmatched actors,
/// and updates status and contributor counters — in one transaction.
async fn finalise(
&self,
manifest_id: &str,
verdict: Verdict,
ratio: f64,
reason: Option<&str>,
matched: &[castcheck::MatchedActor],
unmatched: &[u64],
) -> anyhow::Result<()> {
let id = manifest_id.to_string();
let reason = reason.map(str::to_string);
let matched: Vec<(u64, String, bool)> =
matched.iter().map(|m| (m.tmdb_person_id, m.name.clone(), m.adult)).collect();
let unmatched = unmatched.to_vec();
let now = now_iso();
self.db
.write(move |tx| {
let contributor: Option<String> = tx
.query_row(
"SELECT contributor_id FROM manifests WHERE id = ?1",
rusqlite::params![id],
|r| r.get(0),
)
.map_err(anyhow::Error::from)?;
if verdict == Verdict::Rejected {
// §6: the manifest is deleted and the contributor notified
// (via `GET /manifests/{id}/status` until it is gone).
repo::delete_manifest(tx, &id)?;
} else {
// Names come from TMDB, never from the upload (§5a, §7).
for (person_id, name, adult) in &matched {
repo::upsert_person(tx, *person_id, name, *adult, &now)?;
}
// §6: unmatched actors are dropped rather than stored.
for person_id in &unmatched {
repo::delete_manifest_actor(tx, &id, *person_id)?;
}
repo::set_manifest_status(
tx,
&id,
verdict.status(),
reason.as_deref(),
Some(ratio),
)?;
}
if let Some(c) = contributor {
let counter = match verdict {
Verdict::Listed => "accepted",
Verdict::Flagged => "flagged",
Verdict::Rejected => "rejected",
};
repo::bump_contributor_counter(tx, &c, counter)?;
if repo::maybe_revoke_contributor(tx, &c, &now)? {
tracing::warn!(contributor = %c, "revoked token for excessive rejections");
}
}
Ok(())
})
.await
.context("finalising cast check")
}
}
/// Serialisable projection of `Credits` for the cache.
#[derive(serde::Serialize)]
struct SerializableCredits {
cast: Vec<SerializableMember>,
guest_stars: Vec<SerializableMember>,
}
#[derive(serde::Serialize)]
struct SerializableMember {
id: u64,
name: String,
adult: bool,
}
impl From<&Credits> for SerializableCredits {
fn from(c: &Credits) -> Self {
let f =
|m: &CastMember| SerializableMember { id: m.id, name: m.name.clone(), adult: m.adult };
Self {
cast: c.cast.iter().map(f).collect(),
guest_stars: c.guest_stars.iter().map(&f).collect(),
}
}
}
enum JobError {
Retry(String),
Fatal(anyhow::Error),
}
/// Exponential backoff, capped (§5: the client must back off exponentially
/// rather than retrying tightly; the same discipline applies to our own
/// outbound calls).
fn backoff_secs(attempts: i64) -> u64 {
let base = 30u64;
base.saturating_mul(1u64 << attempts.clamp(0, 8) as u32).min(MAX_BACKOFF_SECS)
}
fn is_stale(fetched_at: &str, ttl: Duration) -> bool {
let Some(then) = parse_iso(fetched_at) else { return true };
let now = unix_now();
now.saturating_sub(then) > ttl.as_secs()
}
/// Current time as an RFC 3339 UTC string, which is what every timestamp column
/// stores. Kept in one place so the format cannot drift.
pub fn now_iso() -> String {
iso_in(0)
}
pub fn iso_in(secs: u64) -> String {
format_unix(unix_now() + secs)
}
fn unix_now() -> u64 {
std::time::SystemTime::now()
.duration_since(std::time::UNIX_EPOCH)
.map(|d| d.as_secs())
.unwrap_or(0)
}
/// Formats a Unix timestamp as `YYYY-MM-DDTHH:MM:SSZ`.
///
/// Hand-rolled rather than pulling in `chrono`/`time`: the only requirement is a
/// lexicographically-sortable UTC string, which is what the `jobs.run_after`
/// comparison relies on.
pub fn format_unix(mut secs: u64) -> String {
let days = secs / 86_400;
secs %= 86_400;
let (h, m, s) = (secs / 3600, (secs % 3600) / 60, secs % 60);
// Civil-from-days, Howard Hinnant's algorithm.
let z = days as i64 + 719_468;
let era = z.div_euclid(146_097);
let doe = z.rem_euclid(146_097);
let yoe = (doe - doe / 1460 + doe / 36_524 - doe / 146_096) / 365;
let y = yoe + era * 400;
let doy = doe - (365 * yoe + yoe / 4 - yoe / 100);
let mp = (5 * doy + 2) / 153;
let d = doy - (153 * mp + 2) / 5 + 1;
let mo = if mp < 10 { mp + 3 } else { mp - 9 };
let y = if mo <= 2 { y + 1 } else { y };
format!("{y:04}-{mo:02}-{d:02}T{h:02}:{m:02}:{s:02}Z")
}
fn parse_iso(s: &str) -> Option<u64> {
// Parses the format `format_unix` produces.
let b = s.as_bytes();
if b.len() < 20 {
return None;
}
let num = |from: usize, to: usize| s.get(from..to)?.parse::<i64>().ok();
let (y, mo, d) = (num(0, 4)?, num(5, 7)?, num(8, 10)?);
let (h, mi, se) = (num(11, 13)?, num(14, 16)?, num(17, 19)?);
let y_adj = if mo <= 2 { y - 1 } else { y };
let era = y_adj.div_euclid(400);
let yoe = y_adj - era * 400;
let mp = if mo > 2 { mo - 3 } else { mo + 9 };
let doy = (153 * mp + 2) / 5 + d - 1;
let doe = yoe * 365 + yoe / 4 - yoe / 100 + doy;
let days = era * 146_097 + doe - 719_468;
Some((days * 86_400 + h * 3600 + mi * 60 + se).max(0) as u64)
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn timestamp_roundtrips() {
for t in [0u64, 1, 1_000_000, 1_700_000_000, 1_785_000_000, 4_000_000_000] {
let s = format_unix(t);
assert_eq!(parse_iso(&s), Some(t), "roundtrip failed for {t} => {s}");
}
}
#[test]
fn timestamp_format_is_sortable() {
// `jobs.run_after <= ?1` is a string comparison, so lexical order must
// match chronological order.
let a = format_unix(1_700_000_000);
let b = format_unix(1_700_000_001);
let c = format_unix(1_800_000_000);
assert!(a < b && b < c, "{a} {b} {c}");
assert_eq!(format_unix(0), "1970-01-01T00:00:00Z");
}
#[test]
fn known_dates_format_correctly() {
// 2026-07-30T12:00:00Z
assert_eq!(format_unix(1_785_412_800), "2026-07-30T12:00:00Z");
// A leap day, since the civil-from-days algorithm is where this would
// break.
assert_eq!(format_unix(1_709_164_800), "2024-02-29T00:00:00Z");
}
#[test]
fn backoff_grows_and_is_capped() {
assert_eq!(backoff_secs(0), 30);
assert_eq!(backoff_secs(1), 60);
assert_eq!(backoff_secs(4), 480);
assert_eq!(backoff_secs(50), MAX_BACKOFF_SECS, "must not overflow or grow unbounded");
}
#[test]
fn staleness_uses_the_ttl() {
let fresh = format_unix(unix_now());
assert!(!is_stale(&fresh, CACHE_TTL));
let old = format_unix(unix_now() - 25 * 3600);
assert!(is_stale(&old, CACHE_TTL));
assert!(is_stale("not-a-timestamp", CACHE_TTL));
}
}
+953
View File
@@ -0,0 +1,953 @@
//! End-to-end tests through the real router.
//!
//! The unit tests cover each spec rule in isolation; these cover the wiring —
//! status codes, headers, and the properties that only hold if the layers are
//! composed correctly (per-route body caps, rate-limit surfaces, the strict
//! schema actually reaching uploads).
use std::sync::Arc;
use axum::body::Body;
use axum::http::{Request, StatusCode};
use http_body_util::BodyExt;
use jray_server::app;
use jray_server::config::Config;
use jray_server::db::Db;
use jray_server::ratelimit::RateLimiter;
use jray_server::state::AppState;
use jray_server::tmdb::TmdbClient;
use serde_json::{json, Value};
use tower::ServiceExt;
/// A server backed by a temporary on-disk database.
///
/// On-disk rather than `:memory:` because §8's design uses a separate writer
/// connection and a read pool, and in-memory SQLite is per-connection — the
/// readers would see an empty database. Testing the real topology is the point.
struct TestServer {
router: axum::Router,
_dir: TempDir,
}
struct TempDir(std::path::PathBuf);
impl TempDir {
fn new(tag: &str) -> Self {
let mut p = std::env::temp_dir();
// Unique per test without pulling in a tempfile dependency.
p.push(format!("jray-test-{}-{}", tag, ulid_like()));
std::fs::create_dir_all(&p).expect("creating temp dir");
Self(p)
}
fn db_path(&self) -> String {
self.0.join("test.db").to_string_lossy().into_owned()
}
}
impl Drop for TempDir {
fn drop(&mut self) {
let _ = std::fs::remove_dir_all(&self.0);
}
}
fn ulid_like() -> String {
use std::sync::atomic::{AtomicU64, Ordering};
static N: AtomicU64 = AtomicU64::new(0);
let t = std::time::SystemTime::now()
.duration_since(std::time::UNIX_EPOCH)
.map(|d| d.as_nanos())
.unwrap_or(0);
format!("{t}-{}", N.fetch_add(1, Ordering::Relaxed))
}
impl TestServer {
fn new(tag: &str) -> Self {
let dir = TempDir::new(tag);
let db = Db::open(&dir.db_path()).expect("opening database");
let config = Arc::new(Config {
bind: "127.0.0.1:0".into(),
db_path: dir.db_path(),
// No key: uploads stay `pending`, which is the correct failure mode
// (§8) and keeps these tests free of network calls.
tmdb_api_key: None,
tmdb_base_url: "http://127.0.0.1:1".into(),
trusted_proxies: Vec::new(),
server_id: "test.example".into(),
request_timeout: std::time::Duration::from_secs(30),
job_batch: 8,
job_poll_interval: std::time::Duration::from_secs(3600),
});
let state = AppState {
db,
config: config.clone(),
limiter: Arc::new(RateLimiter::new()),
tmdb: Arc::new(TmdbClient::new(config.tmdb_base_url.clone(), None)),
};
Self { router: app::router(state), _dir: dir }
}
async fn send(&self, req: Request<Body>) -> (StatusCode, Value, axum::http::HeaderMap) {
let resp = self.router.clone().oneshot(req).await.expect("router call");
let status = resp.status();
let headers = resp.headers().clone();
let bytes = resp.into_body().collect().await.expect("reading body").to_bytes();
let body = if bytes.is_empty() {
Value::Null
} else {
serde_json::from_slice(&bytes)
.unwrap_or(Value::String(String::from_utf8_lossy(&bytes).into_owned()))
};
(status, body, headers)
}
async fn get(&self, uri: &str) -> (StatusCode, Value, axum::http::HeaderMap) {
self.send(Request::builder().uri(uri).body(Body::empty()).unwrap()).await
}
async fn post_json(
&self,
uri: &str,
body: &Value,
) -> (StatusCode, Value, axum::http::HeaderMap) {
self.send(
Request::builder()
.method("POST")
.uri(uri)
.header("content-type", "application/json")
.body(Body::from(body.to_string()))
.unwrap(),
)
.await
}
async fn post_json_auth(
&self,
uri: &str,
token: &str,
body: &Value,
) -> (StatusCode, Value, axum::http::HeaderMap) {
self.send(
Request::builder()
.method("POST")
.uri(uri)
.header("content-type", "application/json")
.header("authorization", format!("Bearer {token}"))
.body(Body::from(body.to_string()))
.unwrap(),
)
.await
}
/// Issues an anonymous bearer capability (§5a).
async fn token(&self) -> String {
let (status, body, _) = self.post_json("/api/v1/tokens", &json!({})).await;
assert_eq!(status, StatusCode::OK, "token issue failed: {body}");
body["token"].as_str().expect("token in response").to_string()
}
}
fn movie_manifest(tmdb_id: &str, runtime: f64) -> Value {
json!({
"jmanifest_version": 1,
"identity": { "type": "movie", "tmdb_id": tmdb_id, "title": "The Death of Stalin",
"year": 2017 },
"cut": { "runtime_sec": runtime, "video_hash": "opensubtitles:8e245d9679d31e12" },
"extraction": { "sample_fps": 5, "extinction_sec": 12,
"pipeline_version": "scene-actor-extraction 0.4.1",
"gallery_scope": "global" },
"actors": [
{ "name": "Steve Buscemi", "tmdb_id": "884", "scenes": [[191.6, 209.2], [438.2, 465.6]] },
{ "name": "Michael Palin", "tmdb_id": "11007", "scenes": [[300.0, 320.0]] }
]
})
}
// ---------------------------------------------------------------------------
// Health and readiness
// ---------------------------------------------------------------------------
#[tokio::test]
async fn health_is_unauthenticated() {
let s = TestServer::new("health");
let (status, body, _) = s.get("/health").await;
assert_eq!(status, StatusCode::OK);
assert_eq!(body["status"], "ok");
}
#[tokio::test]
async fn readiness_reports_database_and_tmdb_configuration() {
// §8: TMDB is a hard dependency for UR-3, so its absence is worth surfacing.
let s = TestServer::new("ready");
let (status, body, _) = s.get("/ready").await;
assert_eq!(status, StatusCode::OK);
assert_eq!(body["status"], "ready");
assert_eq!(body["tmdb_configured"], false);
}
// ---------------------------------------------------------------------------
// §5a — tokens
// ---------------------------------------------------------------------------
#[tokio::test]
async fn upload_without_a_token_is_rejected() {
let s = TestServer::new("noauth");
let (status, _, _) = s.post_json("/api/v1/manifests", &movie_manifest("504172", 6420.5)).await;
assert_eq!(status, StatusCode::UNAUTHORIZED);
}
#[tokio::test]
async fn upload_with_an_unknown_token_is_rejected() {
// A token the server never issued has no contributor row, and §5a stores only
// hashes, so there is nothing to match.
let s = TestServer::new("badauth");
let (status, _, _) = s
.post_json_auth("/api/v1/manifests", "jray_deadbeef", &movie_manifest("504172", 6420.5))
.await;
assert_eq!(status, StatusCode::UNAUTHORIZED);
}
#[tokio::test]
async fn tokens_are_issued_anonymously_and_are_distinct() {
let s = TestServer::new("tokens");
let a = s.token().await;
let b = s.token().await;
assert_ne!(a, b);
assert!(a.starts_with("jray_"));
}
// ---------------------------------------------------------------------------
// §6 — upload validation
// ---------------------------------------------------------------------------
#[tokio::test]
async fn valid_upload_is_accepted_as_pending() {
// §6 stage 3: accepted with `202` and held unlisted until the cast check.
let s = TestServer::new("upload-ok");
let token = s.token().await;
let (status, body, _) =
s.post_json_auth("/api/v1/manifests", &token, &movie_manifest("504172", 6420.5)).await;
assert_eq!(status, StatusCode::ACCEPTED, "body: {body}");
assert_eq!(body["status"], "pending");
assert!(body["manifest_id"].is_string());
}
#[tokio::test]
async fn a_pending_manifest_is_not_served() {
// The property that makes §6 stage 3 meaningful: an unverified manifest is
// not served to anyone in the meantime.
let s = TestServer::new("pending-hidden");
let token = s.token().await;
let (_, body, _) =
s.post_json_auth("/api/v1/manifests", &token, &movie_manifest("504172", 6420.5)).await;
let id = body["manifest_id"].as_str().unwrap();
let (status, _, _) = s.get("/api/v1/manifests/movie?tmdb_id=504172").await;
assert_eq!(status, StatusCode::NOT_FOUND);
let (status, _, _) = s.get(&format!("/api/v1/manifests/{id}")).await;
assert_eq!(status, StatusCode::NOT_FOUND);
// But its status is pollable (§4).
let (status, body, _) = s.get(&format!("/api/v1/manifests/{id}/status")).await;
assert_eq!(status, StatusCode::OK);
assert_eq!(body["status"], "pending");
}
#[tokio::test]
async fn unknown_field_anywhere_is_rejected_with_400() {
// §6 stage 2, enforced by `deny_unknown_fields` on every DTO.
let s = TestServer::new("strict");
let token = s.token().await;
let mut m = movie_manifest("504172", 6420.5);
m["surprise"] = json!("payload");
let (status, body, _) = s.post_json_auth("/api/v1/manifests", &token, &m).await;
assert_eq!(status, StatusCode::BAD_REQUEST);
assert!(
body["message"].as_str().unwrap_or("").contains("surprise"),
"the error should name the offending field: {body}"
);
// Nested, too — `extra="forbid"` applies at every level (§5a).
let mut m = movie_manifest("504172", 6420.5);
m["cut"]["extra"] = json!(1);
let (status, _, _) = s.post_json_auth("/api/v1/manifests", &token, &m).await;
assert_eq!(status, StatusCode::BAD_REQUEST);
}
#[tokio::test]
async fn contributor_local_identifiers_are_rejected_not_ignored() {
// §1/§6: `movie` leaks the contributor's directory layout and `jellyfin_id` is
// a GUID from their database. Both must be *rejected on upload*, so a client
// that forgets to strip them gets a hard 400 naming the field rather than
// quietly publishing them.
let s = TestServer::new("strip");
let token = s.token().await;
let mut m = movie_manifest("504172", 6420.5);
m["movie"] = json!("/data/movies/The.Death.of.Stalin.2017.mkv");
let (status, body, _) = s.post_json_auth("/api/v1/manifests", &token, &m).await;
assert_eq!(status, StatusCode::BAD_REQUEST);
assert!(body["message"].as_str().unwrap_or("").contains("movie"), "{body}");
let mut m = movie_manifest("504172", 6420.5);
m["actors"][0]["jellyfin_id"] = json!("a1b2c3d4e5f6");
let (status, body, _) = s.post_json_auth("/api/v1/manifests", &token, &m).await;
assert_eq!(status, StatusCode::BAD_REQUEST);
assert!(body["message"].as_str().unwrap_or("").contains("jellyfin_id"), "{body}");
}
#[tokio::test]
async fn the_withdrawn_anneal_sec_field_is_rejected() {
// `anneal_sec` was withdrawn in the SR-003 bump: presence now follows track
// extent, so a track survives its own gaps and there is nothing to anneal
// (`scene-actor-extraction` AR-012/AR-013).
//
// Rejecting rather than ignoring it is the point. A manifest still carrying
// the field was produced by a pipeline whose window semantics differ from
// what this server now assumes, and silently accepting it would store
// timings whose meaning we cannot vouch for.
let s = TestServer::new("anneal-withdrawn");
let token = s.token().await;
let mut m = movie_manifest("504172", 6420.5);
m["extraction"]["anneal_sec"] = json!(3);
let (status, body, _) = s.post_json_auth("/api/v1/manifests", &token, &m).await;
assert_eq!(status, StatusCode::BAD_REQUEST);
assert!(
body["message"].as_str().unwrap_or("").contains("anneal_sec"),
"the error should name the withdrawn field: {body}"
);
}
#[tokio::test]
async fn the_schema_bump_fields_round_trip() {
// `extinction_sec` and `gallery_scope` are the SR-003 additions. They are
// stored and reconstructed, since §7 ranks on scope and both are provenance
// a consumer may want.
let s = TestServer::new("bump-fields");
let token = s.token().await;
let m = movie_manifest("504172", 6420.5);
assert_eq!(m["extraction"]["gallery_scope"], "global");
let (status, body, _) = s.post_json_auth("/api/v1/manifests", &token, &m).await;
assert_eq!(status, StatusCode::ACCEPTED, "{body}");
// An unrecognised scope is a closed-vocabulary violation, not a free string.
let mut bad = movie_manifest("504173", 6420.5);
bad["extraction"]["gallery_scope"] = json!("enormous");
let (status, _, _) = s.post_json_auth("/api/v1/manifests", &token, &bad).await;
assert_eq!(status, StatusCode::BAD_REQUEST, "gallery_scope is a closed enum");
}
#[tokio::test]
async fn an_unknown_manifest_version_is_rejected() {
// UR-014 / SR-003: a consumer encountering an unknown `schema_version`
// refuses or warns; it never guesses.
let s = TestServer::new("version");
let token = s.token().await;
for version in [0, 2, 99] {
let mut m = movie_manifest("504172", 6420.5);
m["jmanifest_version"] = json!(version);
let (status, body, _) = s.post_json_auth("/api/v1/manifests", &token, &m).await;
assert_eq!(status, StatusCode::BAD_REQUEST, "version {version}: {body}");
assert!(
body["message"].as_str().unwrap_or("").contains("jmanifest_version"),
"should name the field: {body}"
);
}
}
#[tokio::test]
async fn missing_runtime_is_rejected() {
// §2: `cut.runtime_sec` is required — the primary alignment guard.
let s = TestServer::new("no-runtime");
let token = s.token().await;
let mut m = movie_manifest("504172", 6420.5);
m["cut"] = json!({ "video_hash": "opensubtitles:8e245d9679d31e12" });
let (status, _, _) = s.post_json_auth("/api/v1/manifests", &token, &m).await;
assert_eq!(status, StatusCode::BAD_REQUEST);
}
#[tokio::test]
async fn empty_actor_list_is_rejected() {
// §6: 15 of the 331 corpus files have empty actor lists — extraction
// failures, not contributions.
let s = TestServer::new("empty-actors");
let token = s.token().await;
let mut m = movie_manifest("504172", 6420.5);
m["actors"] = json!([]);
let (status, _, _) = s.post_json_auth("/api/v1/manifests", &token, &m).await;
assert_eq!(status, StatusCode::BAD_REQUEST);
}
#[tokio::test]
async fn smuggled_payload_in_an_actor_name_is_rejected() {
// §5a: the character class defeats base64/hex smuggling, which needs digits
// and padding characters. This is the last free-text channel, so it is worth
// asserting end-to-end and not only in the unit tests.
let s = TestServer::new("smuggle");
let token = s.token().await;
for payload in [
"SGVsbG8gd29ybGQgdGhpcyBpcyBhIHBheWxvYWQ=",
"4d5a90000300000004000000ffff0000",
"<script>alert(1)</script>",
"http://evil.example/x",
] {
let mut m = movie_manifest("504172", 6420.5);
m["actors"][0]["name"] = json!(payload);
let (status, _, _) = s.post_json_auth("/api/v1/manifests", &token, &m).await;
assert_eq!(status, StatusCode::BAD_REQUEST, "payload {payload:?} should be rejected");
}
}
#[tokio::test]
async fn resubmitting_identical_content_is_not_a_duplicate_error() {
// §9a: content addressing gives deduplication — the same manifest from the
// same contributor is recognised rather than stored twice.
let s = TestServer::new("dedup");
let token = s.token().await;
let m = movie_manifest("504172", 6420.5);
let (first, body1, _) = s.post_json_auth("/api/v1/manifests", &token, &m).await;
assert_eq!(first, StatusCode::ACCEPTED);
let (second, body2, _) = s.post_json_auth("/api/v1/manifests", &token, &m).await;
assert_eq!(second, StatusCode::OK);
assert_eq!(body2["status"], "already_present");
assert_eq!(body1["manifest_id"], body2["manifest_id"]);
}
#[tokio::test]
async fn oversized_body_is_rejected_by_the_route_cap() {
// §6 stage 1: the app-level cap counts bytes as they are read, so a lying
// `Content-Length` and a chunked upload are both safe. Here the body genuinely
// exceeds the 2 MiB single-manifest cap.
let s = TestServer::new("too-big");
let token = s.token().await;
let mut m = movie_manifest("504172", 6420.5);
// Many actors, each with many windows — legitimate shape, illegitimate size.
let actors: Vec<Value> = (0..400)
.map(|i| {
let scenes: Vec<Value> = (0..1500).map(|j| json!([j as f64, (j + 1) as f64])).collect();
json!({ "tmdb_id": (1000 + i).to_string(), "scenes": scenes })
})
.collect();
m["actors"] = json!(actors);
let (status, _, _) = s.post_json_auth("/api/v1/manifests", &token, &m).await;
assert_eq!(status, StatusCode::PAYLOAD_TOO_LARGE);
}
#[tokio::test]
async fn a_lying_content_length_does_not_bypass_the_cap() {
// §6 stage 0 is explicit that `Content-Length` is a *claim by the client*: a
// hostile client can declare 100 and send far more, so the streaming cap is
// mandatory rather than redundant.
let s = TestServer::new("lying-length");
let token = s.token().await;
let huge = "x".repeat(3 * 1024 * 1024);
let body = format!("{{\"padding\":\"{huge}\"}}");
let req = Request::builder()
.method("POST")
.uri("/api/v1/manifests")
.header("content-type", "application/json")
.header("authorization", format!("Bearer {token}"))
.header("content-length", "100")
.body(Body::from(body))
.unwrap();
let (status, _, _) = s.send(req).await;
assert_ne!(
status,
StatusCode::ACCEPTED,
"an oversized body must never be accepted, whatever the declared length"
);
assert!(
status == StatusCode::PAYLOAD_TOO_LARGE || status == StatusCode::BAD_REQUEST,
"unexpected status {status}"
);
}
// ---------------------------------------------------------------------------
// §4 — exists
// ---------------------------------------------------------------------------
#[tokio::test]
async fn exists_returns_200_with_false_rather_than_404() {
// §4: absence is a normal answer, and `404` would conflate "no manifest" with
// "bad route" for the client.
let s = TestServer::new("exists-absent");
let (status, body, _) = s.get("/api/v1/manifests/exists?tmdb_id=999999").await;
assert_eq!(status, StatusCode::OK);
assert_eq!(body["exists"], false);
assert!(body["manifest_id"].is_null());
}
#[tokio::test]
async fn exists_requires_identity_parameters() {
let s = TestServer::new("exists-noid");
let (status, _, _) = s.get("/api/v1/manifests/exists").await;
assert_eq!(status, StatusCode::BAD_REQUEST);
}
#[tokio::test]
async fn exists_carries_rate_limit_headers() {
// §5: responses carry `X-RateLimit-Limit`, `-Remaining` and `-Reset`.
let s = TestServer::new("exists-headers");
let (_, _, headers) = s.get("/api/v1/manifests/exists?tmdb_id=1").await;
assert_eq!(headers["x-ratelimit-limit"], "600");
assert_eq!(headers["x-ratelimit-remaining"], "599");
assert!(headers.contains_key("x-ratelimit-reset"));
}
#[tokio::test]
async fn batch_exists_is_positional_and_capped_at_100() {
// §4: results are positional, and the cap is what lets §5 be generous per
// request while staying strict per item.
let s = TestServer::new("exists-batch");
let items: Vec<Value> = (0..3).map(|i| json!({ "tmdb_id": (100 + i).to_string() })).collect();
let (status, body, headers) =
s.post_json("/api/v1/manifests/exists", &json!({ "items": items })).await;
assert_eq!(status, StatusCode::OK);
assert_eq!(body["results"].as_array().unwrap().len(), 3);
assert_eq!(headers["x-ratelimit-limit"], "60", "batch has its own §5 budget");
let too_many: Vec<Value> =
(0..101).map(|i| json!({ "tmdb_id": (100 + i).to_string() })).collect();
let (status, _, _) =
s.post_json("/api/v1/manifests/exists", &json!({ "items": too_many })).await;
assert_eq!(status, StatusCode::BAD_REQUEST);
}
#[tokio::test]
async fn batch_exists_rejects_unknown_fields() {
let s = TestServer::new("exists-batch-strict");
let (status, _, _) =
s.post_json("/api/v1/manifests/exists", &json!({ "items": [], "extra": 1 })).await;
assert_eq!(status, StatusCode::BAD_REQUEST);
}
#[tokio::test]
async fn one_bad_item_does_not_fail_the_whole_batch() {
// A 100-item sweep should not be lost to one malformed entry.
let s = TestServer::new("exists-batch-partial");
let (status, body, _) = s
.post_json(
"/api/v1/manifests/exists",
&json!({ "items": [ { "tmdb_id": "1" }, { }, { "tmdb_id": "2" } ] }),
)
.await;
assert_eq!(status, StatusCode::OK);
let results = body["results"].as_array().unwrap();
assert_eq!(results.len(), 3);
assert_eq!(results[1]["exists"], false);
}
// ---------------------------------------------------------------------------
// §5 — rate limiting
// ---------------------------------------------------------------------------
#[tokio::test]
async fn exceeding_a_limit_returns_429_with_retry_after() {
// §5: exceeding a limit returns `429` with `Retry-After`, which the JRay
// client must honour.
let s = TestServer::new("ratelimit");
let token = s.token().await;
// The bundle surface has the tightest write limit (20/hour), so it is the
// cheapest to exhaust.
let bundle = json!({
"jmanifest_version": 1,
"series": { "series_tmdb_id": "1396", "title": "Breaking Bad" },
"episodes": []
});
let mut saw_429 = false;
for _ in 0..25 {
let (status, _, headers) =
s.post_json_auth("/api/v1/manifests/bundle", &token, &bundle).await;
if status == StatusCode::TOO_MANY_REQUESTS {
assert!(headers.contains_key("retry-after"), "429 must carry Retry-After");
saw_429 = true;
break;
}
}
assert!(saw_429, "the §5 bundle limit should engage within 25 requests");
}
#[tokio::test]
async fn read_surfaces_have_independent_budgets() {
// §5: each surface has its own budget, so a library sweep hammering `exists`
// cannot exhaust the budget a fetch needs.
//
// Only the `exists` surface returns a body on an empty database; the fetch
// surfaces 404 (and a 404 carries no quota headers, by design). So the
// independence is asserted by consuming `exists` and observing that its
// counter alone moves.
let s = TestServer::new("surfaces");
let (_, _, h) = s.get("/api/v1/manifests/exists?tmdb_id=1").await;
assert_eq!(h["x-ratelimit-limit"], "600");
assert_eq!(h["x-ratelimit-remaining"], "599");
// A fetch and a series request in between must not consume `exists` budget.
let _ = s.get("/api/v1/manifests/movie?tmdb_id=1").await;
let _ = s.get("/api/v1/manifests/series/1396").await;
let (_, _, h) = s.get("/api/v1/manifests/exists?tmdb_id=1").await;
assert_eq!(
h["x-ratelimit-remaining"], "598",
"fetch requests must not draw down the exists budget"
);
// And the batch form is a separate surface again (§5).
let (_, _, h) =
s.post_json("/api/v1/manifests/exists", &json!({ "items": [ { "tmdb_id": "1" } ] })).await;
assert_eq!(h["x-ratelimit-limit"], "60");
assert_eq!(h["x-ratelimit-remaining"], "59");
}
#[tokio::test]
async fn a_forged_forwarded_header_cannot_reset_a_budget() {
// §8: the app must trust `X-Forwarded-For` only from the operator's proxy,
// because §5 rate limiting keys on client IP. This server has no configured
// proxies, so the header must be ignored entirely — otherwise a client could
// mint a fresh budget per request.
let s = TestServer::new("xff");
let mut last_remaining = u32::MAX;
for i in 0..3 {
let req = Request::builder()
.uri("/api/v1/manifests/exists?tmdb_id=1")
.header("x-forwarded-for", format!("10.1.1.{i}"))
.body(Body::empty())
.unwrap();
let (_, _, headers) = s.send(req).await;
let remaining: u32 = headers["x-ratelimit-remaining"].to_str().unwrap().parse().unwrap();
assert!(
remaining < last_remaining,
"budget must keep decreasing despite a changing X-Forwarded-For"
);
last_remaining = remaining;
}
}
// ---------------------------------------------------------------------------
// §4 — fetch
// ---------------------------------------------------------------------------
#[tokio::test]
async fn fetch_requires_identity_parameters() {
let s = TestServer::new("fetch-noid");
let (status, _, _) = s.get("/api/v1/manifests/movie").await;
assert_eq!(status, StatusCode::BAD_REQUEST);
let (status, _, _) = s.get("/api/v1/manifests/episode?series_tmdb_id=1396").await;
assert_eq!(status, StatusCode::BAD_REQUEST, "episode fetch needs season and episode");
}
#[tokio::test]
async fn fetching_an_absent_manifest_is_404() {
// §4: `404` if none clears `loose`.
let s = TestServer::new("fetch-absent");
let (status, _, _) = s.get("/api/v1/manifests/movie?tmdb_id=999999").await;
assert_eq!(status, StatusCode::NOT_FOUND);
}
#[tokio::test]
async fn series_bundle_for_an_unknown_series_is_404() {
let s = TestServer::new("series-absent");
let (status, _, _) = s.get("/api/v1/manifests/series/999999").await;
assert_eq!(status, StatusCode::NOT_FOUND);
}
#[tokio::test]
async fn status_of_an_unknown_manifest_reports_rejected() {
// §6 deletes rejected manifests, so a vanished id must not read as a bad
// route — the contributor polling it needs a verdict.
let s = TestServer::new("status-unknown");
let (status, body, _) = s.get("/api/v1/manifests/01HZZZZZZZZZZZZZZZZZZZZZZZ/status").await;
assert_eq!(status, StatusCode::OK);
assert_eq!(body["status"], "rejected");
}
// ---------------------------------------------------------------------------
// §2, §4 — bundles
// ---------------------------------------------------------------------------
#[tokio::test]
async fn bundle_upload_is_not_atomic() {
// §2: valid episodes are accepted and invalid ones rejected, with a
// per-episode result list. All-or-nothing would let one bad episode discard an
// entire season's compute.
let s = TestServer::new("bundle-partial");
let token = s.token().await;
let good = |ep: i64| {
json!({
"jmanifest_version": 1,
"identity": { "type": "episode", "series_tmdb_id": "1396", "title": "Breaking Bad",
"season": 1, "episode": ep },
"cut": { "runtime_sec": 2820.0 },
"actors": [ { "name": "Bryan Cranston", "tmdb_id": "17419",
"scenes": [[10.0, 20.0]] } ]
})
};
// Invalid: a scene window beyond the runtime tolerance (§6).
let bad = json!({
"jmanifest_version": 1,
"identity": { "type": "episode", "series_tmdb_id": "1396", "season": 1, "episode": 3 },
"cut": { "runtime_sec": 2820.0 },
"actors": [ { "tmdb_id": "17419", "scenes": [[10.0, 99999.0]] } ]
});
let bundle = json!({
"jmanifest_version": 1,
"series": { "series_tmdb_id": "1396", "title": "Breaking Bad" },
"episodes": [ good(1), bad, good(2) ]
});
let (status, body, _) = s.post_json_auth("/api/v1/manifests/bundle", &token, &bundle).await;
assert_eq!(status, StatusCode::ACCEPTED, "{body}");
let results = body["results"].as_array().unwrap();
assert_eq!(results.len(), 3);
assert_eq!(results[0]["status"], "pending");
assert_eq!(results[1]["status"], "rejected");
assert!(results[1]["reason"].is_string(), "a rejected episode should say why");
assert_eq!(results[2]["status"], "pending", "a later episode must still be accepted");
}
#[tokio::test]
async fn bundle_envelope_errors_are_whole_request_400s() {
// §4: `400` for the envelope itself, whereas individual bad episodes are
// reported in the results list.
let s = TestServer::new("bundle-envelope");
let token = s.token().await;
let (status, _, _) = s
.post_json_auth(
"/api/v1/manifests/bundle",
&token,
&json!({ "jmanifest_version": 1, "series": {}, "episodes": [] }),
)
.await;
assert_eq!(status, StatusCode::BAD_REQUEST, "series needs an identifier");
let (status, _, _) = s
.post_json_auth(
"/api/v1/manifests/bundle",
&token,
&json!({ "jmanifest_version": 1,
"series": { "series_tmdb_id": "1396" },
"episodes": [], "extra": 1 }),
)
.await;
assert_eq!(status, StatusCode::BAD_REQUEST, "unknown envelope field");
}
#[tokio::test]
async fn bundle_rejects_an_episode_contradicting_the_envelope() {
// An episode must not be silently reattributed to the bundle's series.
let s = TestServer::new("bundle-mismatch");
let token = s.token().await;
let bundle = json!({
"jmanifest_version": 1,
"series": { "series_tmdb_id": "1396" },
"episodes": [ {
"jmanifest_version": 1,
"identity": { "type": "episode", "series_tmdb_id": "9999", "season": 1, "episode": 1 },
"cut": { "runtime_sec": 2820.0 },
"actors": [ { "tmdb_id": "17419", "scenes": [[1.0, 2.0]] } ]
} ]
});
let (status, body, _) = s.post_json_auth("/api/v1/manifests/bundle", &token, &bundle).await;
assert_eq!(status, StatusCode::ACCEPTED);
assert_eq!(body["results"][0]["status"], "rejected");
}
#[tokio::test]
async fn bundle_beyond_the_episode_cap_is_413() {
// §2/§4: capped at 500 episodes; beyond that the client must page by season.
let s = TestServer::new("bundle-cap");
let token = s.token().await;
let episodes: Vec<Value> = (0..501)
.map(|i| {
json!({
"jmanifest_version": 1,
"identity": { "type": "episode", "series_tmdb_id": "1396",
"season": 1, "episode": i },
"cut": { "runtime_sec": 2820.0 },
"actors": [ { "tmdb_id": "17419", "scenes": [[1.0, 2.0]] } ]
})
})
.collect();
let bundle = json!({
"jmanifest_version": 1,
"series": { "series_tmdb_id": "1396" },
"episodes": episodes
});
let (status, _, _) = s.post_json_auth("/api/v1/manifests/bundle", &token, &bundle).await;
assert_eq!(status, StatusCode::PAYLOAD_TOO_LARGE);
}
#[tokio::test]
async fn bundle_route_accepts_a_body_larger_than_the_single_manifest_cap() {
// §6 stage 1: per-route limits, so the bundle endpoint gets its larger cap
// without widening the others. A ~3 MiB bundle exceeds the 2 MiB manifest cap
// but is well within the 25 MiB bundle cap.
let s = TestServer::new("bundle-bigger-cap");
let token = s.token().await;
let episodes: Vec<Value> = (1..=60)
.map(|ep| {
let scenes: Vec<Value> = (0..600).map(|j| json!([j as f64, (j + 1) as f64])).collect();
json!({
"jmanifest_version": 1,
"identity": { "type": "episode", "series_tmdb_id": "1396",
"season": 1, "episode": ep },
"cut": { "runtime_sec": 2820.0 },
"actors": (0..8).map(|a| json!({
"tmdb_id": (20000 + a).to_string(), "scenes": scenes
})).collect::<Vec<_>>()
})
})
.collect();
let bundle = json!({
"jmanifest_version": 1,
"series": { "series_tmdb_id": "1396" },
"episodes": episodes
});
let encoded = bundle.to_string();
assert!(
encoded.len() > 2 * 1024 * 1024,
"test body should exceed the single-manifest cap, got {} bytes",
encoded.len()
);
let (status, _, _) = s.post_json_auth("/api/v1/manifests/bundle", &token, &bundle).await;
assert_eq!(status, StatusCode::ACCEPTED, "the bundle route has its own larger cap");
}
// ---------------------------------------------------------------------------
// §4 — reports
// ---------------------------------------------------------------------------
#[tokio::test]
async fn reporting_an_unknown_manifest_is_404() {
let s = TestServer::new("report-unknown");
let (status, _, _) = s
.post_json(
"/api/v1/manifests/01HZZZZZZZZZZZZZZZZZZZZZZZ/report",
&json!({ "reason": "misaligned" }),
)
.await;
assert_eq!(status, StatusCode::NOT_FOUND);
}
#[tokio::test]
async fn a_report_is_accepted_and_does_not_delist() {
// §5a: delisting stays an operator action. Automatic delisting on report would
// hand any client a remote delete primitive.
let s = TestServer::new("report-ok");
let token = s.token().await;
let (_, body, _) =
s.post_json_auth("/api/v1/manifests", &token, &movie_manifest("504172", 6420.5)).await;
let id = body["manifest_id"].as_str().unwrap().to_string();
let (status, body, _) = s
.post_json(
&format!("/api/v1/manifests/{id}/report"),
&json!({ "reason": "wrong_actors", "note": "these are not the right people" }),
)
.await;
assert_eq!(status, StatusCode::OK, "{body}");
assert!(body["report_id"].is_string());
let (status, body, _) = s.get(&format!("/api/v1/manifests/{id}/status")).await;
assert_eq!(status, StatusCode::OK);
assert_eq!(body["status"], "pending", "a report must not change status by itself");
}
#[tokio::test]
async fn report_rejects_an_unknown_reason_and_unknown_fields() {
let s = TestServer::new("report-strict");
let (status, _, _) = s
.post_json("/api/v1/manifests/x/report", &json!({ "reason": "i_just_dont_like_it" }))
.await;
assert_eq!(status, StatusCode::BAD_REQUEST);
let (status, _, _) = s
.post_json("/api/v1/manifests/x/report", &json!({ "reason": "spam", "extra": true }))
.await;
assert_eq!(status, StatusCode::BAD_REQUEST);
}
#[tokio::test]
async fn report_note_is_length_capped() {
// §5a: free text from an anonymous caller is capped hard.
let s = TestServer::new("report-note");
let token = s.token().await;
let (_, body, _) =
s.post_json_auth("/api/v1/manifests", &token, &movie_manifest("504172", 6420.5)).await;
let id = body["manifest_id"].as_str().unwrap().to_string();
let (status, _, _) = s
.post_json(
&format!("/api/v1/manifests/{id}/report"),
&json!({ "reason": "spam", "note": "a".repeat(5000) }),
)
.await;
assert_eq!(status, StatusCode::BAD_REQUEST);
}
// ---------------------------------------------------------------------------
// Malformed input
// ---------------------------------------------------------------------------
#[tokio::test]
async fn malformed_json_is_a_400_not_a_500() {
let s = TestServer::new("bad-json");
let token = s.token().await;
let req = Request::builder()
.method("POST")
.uri("/api/v1/manifests")
.header("content-type", "application/json")
.header("authorization", format!("Bearer {token}"))
.body(Body::from("{ this is not json"))
.unwrap();
let (status, _, _) = s.send(req).await;
assert_eq!(status, StatusCode::BAD_REQUEST);
}
#[tokio::test]
async fn deeply_nested_json_does_not_crash_the_parser() {
// §6 stage 1 caps nesting depth; a parser handed unbounded input is a DoS
// primitive, so the failure must be a clean rejection.
let s = TestServer::new("deep-json");
let token = s.token().await;
let deep = format!("{}{}", "[".repeat(5000), "]".repeat(5000));
let req = Request::builder()
.method("POST")
.uri("/api/v1/manifests")
.header("content-type", "application/json")
.header("authorization", format!("Bearer {token}"))
.body(Body::from(deep))
.unwrap();
let (status, _, _) = s.send(req).await;
assert!(
status == StatusCode::BAD_REQUEST || status == StatusCode::PAYLOAD_TOO_LARGE,
"deeply nested input should be rejected cleanly, got {status}"
);
}
#[tokio::test]
async fn unknown_routes_are_404() {
let s = TestServer::new("routes");
let (status, _, _) = s.get("/api/v1/nonexistent").await;
assert_eq!(status, StatusCode::NOT_FOUND);
let (status, _, _) = s.get("/api/v2/manifests/exists?tmdb_id=1").await;
assert_eq!(status, StatusCode::NOT_FOUND);
}
+634
View File
@@ -0,0 +1,634 @@
//! Injection resistance — SQL, JSON and header.
//!
//! These are regression tests for properties the design already provides, kept
//! separate from `api.rs` because their purpose is different: `api.rs` asserts the
//! spec's behaviour, this asserts that hostile input cannot escape its layer.
//!
//! Two distinct defences are at work, and it is worth being precise about which
//! applies where, because they fail differently:
//!
//! 1. **Parameterised queries** (§7, §8). Every value reaches SQLite through
//! `params![]`; the only `format!`-built SQL interpolates compile-time
//! constants (a column list and a status literal). So a value carrying SQL
//! syntax is bound as *data* and simply matches nothing.
//! 2. **Closed-vocabulary validation** (§5a, §6 stage 2). Identifiers are
//! regex-constrained and free text is restricted to a closed character class,
//! so most injection strings are rejected before they reach the database.
//!
//! Defence 1 is what actually prevents injection; defence 2 means an attacker
//! usually cannot even reach it. Testing both matters: if validation were ever
//! loosened, these tests should still pass on the strength of parameterisation
//! alone.
use std::sync::Arc;
use axum::body::Body;
use axum::http::{Request, StatusCode};
use http_body_util::BodyExt;
use jray_server::app;
use jray_server::config::Config;
use jray_server::db::Db;
use jray_server::ratelimit::RateLimiter;
use jray_server::state::AppState;
use jray_server::tmdb::TmdbClient;
use serde_json::{json, Value};
use tower::ServiceExt;
/// Payloads spanning the usual SQL-injection shapes: boolean tautology, statement
/// termination, stacked statements, UNION exfiltration, comment truncation, and
/// string-concatenation exfiltration.
const SQL_PAYLOADS: &[&str] = &[
"1' OR '1'='1",
"1'; DROP TABLE manifests;--",
"1 UNION SELECT token_hash FROM contributors",
"' OR 1=1--",
"1'||(SELECT token_hash FROM contributors)||'",
"1)) OR 1=1 --",
"'; UPDATE manifests SET status='listed' WHERE 1=1;--",
"1/**/UNION/**/SELECT/**/1",
"x' AND (SELECT COUNT(*) FROM sqlite_master)>0 --",
"\"; DELETE FROM scenes; --",
];
struct TestServer {
router: axum::Router,
db: Db,
_dir: TempDir,
}
struct TempDir(std::path::PathBuf);
impl TempDir {
fn new(tag: &str) -> Self {
let mut p = std::env::temp_dir();
p.push(format!("jray-inj-{}-{}", tag, unique()));
std::fs::create_dir_all(&p).expect("creating temp dir");
Self(p)
}
fn db_path(&self) -> String {
self.0.join("test.db").to_string_lossy().into_owned()
}
}
impl Drop for TempDir {
fn drop(&mut self) {
let _ = std::fs::remove_dir_all(&self.0);
}
}
fn unique() -> String {
use std::sync::atomic::{AtomicU64, Ordering};
static N: AtomicU64 = AtomicU64::new(0);
let t = std::time::SystemTime::now()
.duration_since(std::time::UNIX_EPOCH)
.map(|d| d.as_nanos())
.unwrap_or(0);
format!("{t}-{}", N.fetch_add(1, Ordering::Relaxed))
}
impl TestServer {
fn new(tag: &str) -> Self {
let dir = TempDir::new(tag);
let db = Db::open(&dir.db_path()).expect("opening database");
let config = Arc::new(Config {
bind: "127.0.0.1:0".into(),
db_path: dir.db_path(),
tmdb_api_key: None,
tmdb_base_url: "http://127.0.0.1:1".into(),
trusted_proxies: Vec::new(),
server_id: "test.example".into(),
request_timeout: std::time::Duration::from_secs(30),
job_batch: 8,
job_poll_interval: std::time::Duration::from_secs(3600),
});
let state = AppState {
db: db.clone(),
config: config.clone(),
limiter: Arc::new(RateLimiter::new()),
tmdb: Arc::new(TmdbClient::new(config.tmdb_base_url.clone(), None)),
};
Self { router: app::router(state), db, _dir: dir }
}
async fn send(&self, req: Request<Body>) -> (StatusCode, Value) {
let resp = self.router.clone().oneshot(req).await.expect("router call");
let status = resp.status();
let bytes = resp.into_body().collect().await.expect("body").to_bytes();
let body = if bytes.is_empty() {
Value::Null
} else {
serde_json::from_slice(&bytes)
.unwrap_or(Value::String(String::from_utf8_lossy(&bytes).into_owned()))
};
(status, body)
}
async fn get(&self, uri: &str) -> (StatusCode, Value) {
self.send(Request::builder().uri(uri).body(Body::empty()).unwrap()).await
}
async fn post(&self, uri: &str, token: Option<&str>, body: &Value) -> (StatusCode, Value) {
let mut b =
Request::builder().method("POST").uri(uri).header("content-type", "application/json");
if let Some(t) = token {
b = b.header("authorization", format!("Bearer {t}"));
}
self.send(b.body(Body::from(body.to_string())).unwrap()).await
}
async fn token(&self) -> String {
let (_, body) = self.post("/api/v1/tokens", None, &json!({})).await;
body["token"].as_str().expect("token").to_string()
}
/// Confirms the schema is intact and the expected row counts hold.
///
/// A successful injection would most likely drop a table or delete rows, so
/// this is the assertion that actually matters after each payload.
async fn assert_schema_intact(&self) {
let tables: Vec<String> = self
.db
.read(|conn| {
let mut stmt = conn
.prepare("SELECT name FROM sqlite_master WHERE type='table' ORDER BY name")?;
let rows = stmt
.query_map([], |r| r.get::<_, String>(0))?
.collect::<rusqlite::Result<Vec<_>>>()?;
Ok(rows)
})
.await
.expect("listing tables");
for expected in [
"contributors",
"jobs",
"manifest_actors",
"manifests",
"people",
"reports",
"scenes",
"titles",
"tmdb_cache",
] {
assert!(
tables.iter().any(|t| t == expected),
"table {expected} is missing — an injection may have dropped it. tables: {tables:?}"
);
}
}
}
fn urlencode(s: &str) -> String {
let mut out = String::new();
for b in s.bytes() {
match b {
b'A'..=b'Z' | b'a'..=b'z' | b'0'..=b'9' | b'-' | b'_' | b'.' | b'~' => {
out.push(b as char)
}
_ => out.push_str(&format!("%{b:02X}")),
}
}
out
}
// ---------------------------------------------------------------------------
// SQL injection — query parameters
// ---------------------------------------------------------------------------
#[tokio::test]
async fn sql_payloads_in_query_parameters_are_inert() {
let s = TestServer::new("query");
for payload in SQL_PAYLOADS {
let enc = urlencode(payload);
for uri in [
format!("/api/v1/manifests/exists?tmdb_id={enc}"),
format!("/api/v1/manifests/exists?imdb_id={enc}"),
format!("/api/v1/manifests/movie?tmdb_id={enc}"),
format!("/api/v1/manifests/movie?imdb_id={enc}"),
format!("/api/v1/manifests/episode?series_tmdb_id={enc}&season=1&episode=1"),
format!("/api/v1/manifests/series/{enc}"),
format!("/api/v1/manifests/exists?tmdb_id=1&video_hash={enc}"),
] {
let (status, body) = s.get(&uri).await;
// The payload is bound as data, so it matches nothing. What must never
// happen is a 5xx, which would mean SQLite saw it as syntax.
assert!(
status.is_success()
|| status == StatusCode::NOT_FOUND
|| status == StatusCode::BAD_REQUEST,
"payload {payload:?} on {uri} produced {status} — expected data-not-found, \
not a server error. body: {body}"
);
}
}
s.assert_schema_intact().await;
}
#[tokio::test]
async fn sql_payloads_in_path_parameters_are_inert() {
let s = TestServer::new("path");
for payload in SQL_PAYLOADS {
let enc = urlencode(payload);
for uri in [
format!("/api/v1/manifests/{enc}"),
format!("/api/v1/manifests/{enc}/status"),
format!("/api/v1/manifests/series/{enc}"),
] {
let (status, body) = s.get(&uri).await;
assert!(
!status.is_server_error(),
"payload {payload:?} on {uri} produced {status}: {body}"
);
}
}
s.assert_schema_intact().await;
}
// ---------------------------------------------------------------------------
// SQL injection — JSON body fields
// ---------------------------------------------------------------------------
#[tokio::test]
async fn sql_payloads_in_identifier_fields_are_rejected() {
// §6 stage 2 regex-constrains every identifier, so these never even reach the
// query layer. The response must be a clean 400 naming the field.
let s = TestServer::new("body-ids");
let token = s.token().await;
for payload in SQL_PAYLOADS {
let manifest = json!({
"jmanifest_version": 1,
"identity": { "type": "movie", "tmdb_id": payload },
"cut": { "runtime_sec": 100.0 },
"actors": [ { "tmdb_id": "884", "scenes": [[1.0, 2.0]] } ]
});
let (status, body) = s.post("/api/v1/manifests", Some(&token), &manifest).await;
assert_eq!(
status,
StatusCode::BAD_REQUEST,
"payload {payload:?} should be rejected by validation: {body}"
);
}
s.assert_schema_intact().await;
}
#[tokio::test]
async fn sql_payloads_in_free_text_fields_are_rejected() {
// The two free-text fields (§5a) are the only place arbitrary strings could
// arrive. The closed character class excludes quotes, semicolons and digits,
// which is what makes SQL syntax unrepresentable there.
let s = TestServer::new("body-text");
let token = s.token().await;
for payload in SQL_PAYLOADS {
for manifest in [
json!({
"jmanifest_version": 1,
"identity": { "type": "movie", "tmdb_id": "504172", "title": payload },
"cut": { "runtime_sec": 100.0 },
"actors": [ { "tmdb_id": "884", "scenes": [[1.0, 2.0]] } ]
}),
json!({
"jmanifest_version": 1,
"identity": { "type": "movie", "tmdb_id": "504172" },
"cut": { "runtime_sec": 100.0 },
"actors": [ { "name": payload, "tmdb_id": "884", "scenes": [[1.0, 2.0]] } ]
}),
] {
let (status, body) = s.post("/api/v1/manifests", Some(&token), &manifest).await;
assert_eq!(
status,
StatusCode::BAD_REQUEST,
"free-text payload {payload:?} should be rejected: {body}"
);
}
}
s.assert_schema_intact().await;
}
#[tokio::test]
async fn sql_payloads_in_a_report_note_cannot_escape() {
// `note` is the one field that accepts relatively free text (control
// characters stripped, length capped) because only the operator reads it. It
// reaches the database, so it is the strongest test of parameterisation:
// validation is *not* filtering SQL syntax here.
let s = TestServer::new("report-note");
let token = s.token().await;
let (_, body) = s
.post(
"/api/v1/manifests",
Some(&token),
&json!({
"jmanifest_version": 1,
"identity": { "type": "movie", "tmdb_id": "504172" },
"cut": { "runtime_sec": 100.0 },
"actors": [ { "tmdb_id": "884", "scenes": [[1.0, 2.0]] } ]
}),
)
.await;
let id = body["manifest_id"].as_str().expect("manifest id").to_string();
for payload in SQL_PAYLOADS {
let (status, body) = s
.post(
&format!("/api/v1/manifests/{id}/report"),
None,
&json!({ "reason": "spam", "note": payload }),
)
.await;
assert!(
status.is_success() || status == StatusCode::TOO_MANY_REQUESTS,
"note payload {payload:?} produced {status}: {body}"
);
if status.is_success() {
s.assert_schema_intact().await;
}
}
// The notes were stored verbatim as *data* — proving they were bound, not
// executed. Verified by reading them back out.
let stored: i64 =
s.db.read(|conn| Ok(conn.query_row("SELECT COUNT(*) FROM reports", [], |r| r.get(0))?))
.await
.expect("counting reports");
assert!(stored > 0, "reports should have been stored as inert data");
}
#[tokio::test]
async fn sql_payloads_in_a_bearer_token_are_inert() {
// The token is hashed before it reaches any query, but a payload arriving via
// a header must still not produce a 5xx.
let s = TestServer::new("token-inj");
for payload in SQL_PAYLOADS {
let req = Request::builder()
.method("POST")
.uri("/api/v1/manifests")
.header("content-type", "application/json")
.header("authorization", format!("Bearer {payload}"))
.body(Body::from(
json!({
"jmanifest_version": 1,
"identity": { "type": "movie", "tmdb_id": "504172" },
"cut": { "runtime_sec": 100.0 },
"actors": [ { "tmdb_id": "884", "scenes": [[1.0, 2.0]] } ]
})
.to_string(),
))
.unwrap();
let (status, body) = s.send(req).await;
assert_eq!(
status,
StatusCode::UNAUTHORIZED,
"token payload {payload:?} produced {status}: {body}"
);
}
s.assert_schema_intact().await;
}
#[tokio::test]
async fn sql_payloads_in_the_batch_exists_body_are_inert() {
let s = TestServer::new("batch-inj");
let items: Vec<Value> = SQL_PAYLOADS.iter().map(|p| json!({ "tmdb_id": p })).collect();
let (status, body) = s.post("/api/v1/manifests/exists", None, &json!({ "items": items })).await;
assert_eq!(status, StatusCode::OK, "{body}");
// Each malformed item degrades to "absent" rather than erroring the batch.
for result in body["results"].as_array().expect("results") {
assert_eq!(result["exists"], false);
}
s.assert_schema_intact().await;
}
// ---------------------------------------------------------------------------
// JSON injection / parser abuse
// ---------------------------------------------------------------------------
#[tokio::test]
async fn json_structure_abuse_is_rejected_cleanly() {
// A parser handed hostile structure must fail with 400/413, never 5xx and
// never a hang (§6 stage 1: "a parser handed an unbounded body is a
// denial-of-service primitive").
let s = TestServer::new("json-abuse");
let token = s.token().await;
let cases: Vec<(&str, String)> = vec![
("deep nesting", format!("{}{}", "[".repeat(20_000), "]".repeat(20_000))),
("unterminated", "{\"identity\": {\"type\": \"movie\"".to_string()),
("duplicate keys", r#"{"jmanifest_version":1,"jmanifest_version":2}"#.to_string()),
("null bytes", "{\"jmanifest_version\":\u{0}1}".to_string()),
("huge number", format!("{{\"jmanifest_version\":{}}}", "9".repeat(5000))),
("nan literal", r#"{"jmanifest_version":1,"cut":{"runtime_sec":NaN}}"#.to_string()),
("bare array", "[1,2,3]".to_string()),
("bare string", "\"just a string\"".to_string()),
("empty body", String::new()),
(
"prototype-style key",
r#"{"__proto__":{"admin":true},"jmanifest_version":1}"#.to_string(),
),
];
for (label, body) in cases {
let req = Request::builder()
.method("POST")
.uri("/api/v1/manifests")
.header("content-type", "application/json")
.header("authorization", format!("Bearer {token}"))
.body(Body::from(body))
.unwrap();
let (status, resp) = s.send(req).await;
assert!(
status == StatusCode::BAD_REQUEST || status == StatusCode::PAYLOAD_TOO_LARGE,
"{label} produced {status}, expected a clean rejection: {resp}"
);
}
s.assert_schema_intact().await;
}
#[tokio::test]
async fn non_finite_scene_times_are_rejected() {
// §6 explicitly rejects NaN/Infinity. They cannot arrive as JSON literals, but
// they can arrive as overflowing decimals, which parse to f64 infinity.
let s = TestServer::new("nonfinite");
let token = s.token().await;
// Sent as raw JSON text rather than via `json!`, because rustc refuses an
// out-of-range float literal — and the point is to make the *server's* parser
// handle it, which is the real attack path.
let raw = r#"{"jmanifest_version":1,
"identity":{"type":"movie","tmdb_id":"504172"},
"cut":{"runtime_sec":100.0},
"actors":[{"tmdb_id":"884","scenes":[[1.0,1e400]]}]}"#;
let req = Request::builder()
.method("POST")
.uri("/api/v1/manifests")
.header("content-type", "application/json")
.header("authorization", format!("Bearer {token}"))
.body(Body::from(raw))
.unwrap();
let (status, body) = s.send(req).await;
assert_eq!(status, StatusCode::BAD_REQUEST, "{body}");
// Likewise an overflowing runtime.
let raw = r#"{"jmanifest_version":1,
"identity":{"type":"movie","tmdb_id":"504172"},
"cut":{"runtime_sec":1e400},
"actors":[{"tmdb_id":"884","scenes":[[1.0,2.0]]}]}"#;
let req = Request::builder()
.method("POST")
.uri("/api/v1/manifests")
.header("content-type", "application/json")
.header("authorization", format!("Bearer {token}"))
.body(Body::from(raw))
.unwrap();
let (status, body) = s.send(req).await;
assert_eq!(status, StatusCode::BAD_REQUEST, "{body}");
}
// ---------------------------------------------------------------------------
// Header injection
// ---------------------------------------------------------------------------
#[tokio::test]
async fn crlf_in_a_header_value_cannot_split_the_response() {
// A CRLF-carrying header value must not appear in the response as new headers.
// `http` rejects such values at construction, so this asserts the invariant
// holds at the boundary rather than relying on our own escaping.
let bad = "1.2.3.4\r\nX-Injected: yes";
assert!(
axum::http::HeaderValue::from_str(bad).is_err(),
"the http crate must refuse CRLF in header values"
);
// And a percent-encoded variant reaching a handler stays inert data.
let s = TestServer::new("crlf");
let (status, _) = s.get("/api/v1/manifests/exists?tmdb_id=1%0D%0AX-Injected:%20yes").await;
assert!(!status.is_server_error());
s.assert_schema_intact().await;
}
#[tokio::test]
async fn oversized_headers_do_not_take_the_server_down() {
let s = TestServer::new("big-header");
let big = "a".repeat(100_000);
let req = Request::builder()
.uri("/api/v1/manifests/exists?tmdb_id=1")
.header("x-filler", big)
.body(Body::empty())
.unwrap();
let (status, _) = s.send(req).await;
assert!(!status.is_server_error(), "got {status}");
}
// ---------------------------------------------------------------------------
// Path traversal
// ---------------------------------------------------------------------------
#[tokio::test]
async fn path_traversal_attempts_reach_no_filesystem() {
// The server serves no files at all, so traversal has nowhere to go. Asserted
// anyway, because the manifest id is a path segment.
let s = TestServer::new("traversal");
for probe in [
"..%2F..%2F..%2Fetc%2Fpasswd",
"....%2F%2F....%2F%2Fetc%2Fpasswd",
"%2e%2e%2f%2e%2e%2fetc%2fshadow",
"..%5C..%5Cwindows%5Csystem32",
"%00/etc/passwd",
] {
let (status, body) = s.get(&format!("/api/v1/manifests/{probe}")).await;
assert!(
status == StatusCode::NOT_FOUND || status == StatusCode::BAD_REQUEST,
"probe {probe} produced {status}: {body}"
);
// Nothing that looks like file content should ever come back.
let text = body.to_string();
assert!(!text.contains("root:"), "probe {probe} returned passwd-like content");
}
}
// ---------------------------------------------------------------------------
// Unicode and encoding tricks against the §5a character class
// ---------------------------------------------------------------------------
#[tokio::test]
async fn unicode_tricks_cannot_smuggle_text_past_the_character_class() {
// §5a's class is checked *after* NFC normalisation, so decomposed and
// compatibility forms must not provide a way in. Fullwidth digits are the
// sharpest case: NFKC would fold them to ASCII digits, but NFC does not, and
// they are `Nd` (not a letter), so the class rejects them either way.
let s = TestServer::new("unicode");
let token = s.token().await;
for payload in [
"Actor 123", // fullwidth letters and digits
"Steve\u{FEFF}Buscemi", // zero-width no-break space
"Ste\u{0301}ve\u{202E}", // combining acute plus bidi override
"𝐒𝐭𝐞𝐯𝐞", // mathematical bold (compatibility form)
"Steve\u{2028}Buscemi", // line separator
"\u{1F600} Actor", // emoji
"Actor\u{00A0}Name\u{0000}", // nbsp plus NUL
] {
let (status, body) = s
.post(
"/api/v1/manifests",
Some(&token),
&json!({
"jmanifest_version": 1,
"identity": { "type": "movie", "tmdb_id": "504172" },
"cut": { "runtime_sec": 100.0 },
"actors": [ { "name": payload, "tmdb_id": "884", "scenes": [[1.0, 2.0]] } ]
}),
)
.await;
assert_eq!(
status,
StatusCode::BAD_REQUEST,
"unicode payload {payload:?} should be rejected: {body}"
);
}
}
#[tokio::test]
async fn the_audio_signature_field_cannot_carry_arbitrary_bytes() {
// §3: a variable-length blob would be a payload channel — "precisely what §5a
// closes". Length is fixed and every byte is structurally constrained.
let s = TestServer::new("audio-sig");
let token = s.token().await;
for sig in [
"v1:aGVsbG8gd29ybGQ=", // too short to be a signature
&format!("v1:{}", "/".repeat(4000)), // high bit set throughout
&format!("v1:{}", "A".repeat(100_000)), // oversized
"not-base64-at-all", // missing version prefix
&"A".repeat(1720), // unprefixed
] {
let (status, body) = s
.post(
"/api/v1/manifests",
Some(&token),
&json!({
"jmanifest_version": 1,
"identity": { "type": "movie", "tmdb_id": "504172" },
"cut": { "runtime_sec": 6420.5, "audio_signature": sig },
"actors": [ { "tmdb_id": "884", "scenes": [[1.0, 2.0]] } ]
}),
)
.await;
assert_eq!(
status,
StatusCode::BAD_REQUEST,
"signature {:?} should be rejected: {body}",
&sig.chars().take(40).collect::<String>()
);
}
}