Files
JRay-public-server/SPEC.md
dtourolle 6c0c80f20b
CI / fmt, clippy, test (pull_request) Successful in 1m48s
CI / advisories and licences (pull_request) Successful in 42s
CI / static musl binary (pull_request) Successful in 1m46s
CI / debian package (pull_request) Failing after 3h10m28s
feat(deb): package the server, and ask the questions that fail silently
DR-015, DR-016. §8 already shipped a static binary and an optional
container; this adds the third form, and it packages the SAME binary the
musl job proved static rather than building its own. Two builds of the
same commit could diverge, and the whole point of that assertion is that
the artifact an operator installs is the one that was checked.

Built with dpkg-deb from an explicit staging tree rather than cargo-deb.
debconf's `config` script and `templates` live in the control archive
next to the maintainer scripts, and controlling that archive directly
beats discovering what a wrapper will copy into it. dpkg-dev is on every
Debian builder, so this adds no build dependency.

Why debconf at all: two settings fail SILENTLY when unset. Without
JRAY_TMDB_API_KEY every upload stays `pending` and is never listed;
without JRAY_TRUSTED_PROXIES the X-Forwarded-For header is ignored, so
every client shares one rate-limit bucket and every abuse report points
at the proxy. Both leave a server that works and is quietly doing the
wrong thing — the worst thing to leave to a README nobody reads.

Three properties, each a way packaging usually goes wrong:

The generated config is NOT a dpkg conffile. It is written from the
debconf answers, so shipping it as one would make dpkg prompt on every
upgrade about changes the package itself had made.

Hand edits survive. postinst rewrites only the keys debconf manages;
comments, ordering and any other setting are left alone.

A blank key on reconfigure keeps the existing one. Otherwise pressing
Enter through a dpkg-reconfigure would unpublish every future upload.

The seeding guard is worth its comment, because the obvious version is
wrong twice over. `config` seeds unanswered questions from the env file
so a reconfigure shows what is actually in force. Seeding
unconditionally overwrites a preseed — debconf-set-selections marks what
it sets as seen — so every unattended install would quietly reconfigure
itself back to whatever was on disk. Guarding on an empty value does not
work either: server-id and bind carry template Defaults, so db_get
returns "localhost" for a question nobody answered. The test is the
`seen` flag, which is the actual question being asked.

Purge keeps the database, knowingly departing from the expectation that
purge removes everything. Manifests are the output of real CV compute on
media the operator may no longer have, and §8 says federation is
explicitly not a backup. Destroying that during an `apt purge` is not a
trade worth making for tidiness; postrm names the path instead.

The nginx example is documentation, not installed configuration. The
proxy usually runs on a different host from the server, so a file
dropped into this machine's nginx would be in the wrong place — and §8
leaves the edge to the operator deliberately.

Verified by running it, not by reading it: a full lifecycle in a
bookworm container — build, preseeded install, mode-600 env file, key
absent from debconf's database afterwards, `systemd-analyze verify` on
the unit, the installed binary answering /health and /ready, reconfigure
preserving both the key and an unmanaged setting, and purge leaving the
database. It failed on the seeding bug above the first time, which is
why that guard exists. CI runs the same checks against every build.

TRACES: DR-015, DR-016 | PR-004
2026-09-06 09:58:47 +02:00

98 KiB
Raw Permalink Blame History

JRay Public Server — specification

A community manifest exchange for JRay. Jellyfin servers running the JRay plugin pull actor-timeline manifests ("Jmanifests") for titles they own instead of running the CV pipeline locally, and optionally contribute the manifests they generate back.

Status: core implemented. See docs/requirements.md for per-requirement status and README.md for what is deferred.

This is a software spec: its job is to implement the system spec, which owns everything spanning more than one repo. Requirements here trace up to an SR-nnn; the prose below is the detail.


0. Requirements

IDs are UR-nnn, zero-padded and permanent — a withdrawn requirement keeps its number, because renumbering is what produces orphan TRACES tags (system spec §6). The authoritative list with status lives in docs/requirements.md; this table is the prose anchor.

# Requirement Traces to Where addressed
UR-001 Query whether a JRay manifest for a given media item exists on the server SR-001 §4 GET /manifests/exists
UR-002 Route to post a JRay manifest for a media item PR-006 §4 POST /manifests
UR-003 Content verification: no additional JSON fields, file size limit, approximate cast match against TMDB SR-004 §6
UR-004 Rate limiting on queries SR-004 §5
UR-005 Trust without account management: server must not be usable as a content store, nor for prank/vandalism manifests SR-004 §5a
UR-006 Serve and accept a whole series in one operation, not episode-by-episode PR-006 §2 series bundles, §4 GET /manifests/series, POST /manifests/bundle
UR-007 JRay plugin must query a configurable list of servers PR-005 §9
UR-008 Servers must be able to sync/replicate manifests between each other PR-006 §9a
UR-009 Store an audio spectral-peak signature from the media centre, so a file of unknown providence can be identified and synchronised SR-003 §3 audio signature
UR-010 Identity crossing the API boundary is TMDB/IMDB ids, never a name alone SR-001 §2, §6 stage 3, §7
UR-011 Reject any manifest field capable of carrying binary or attacker-chosen content SR-004 §5a Threat 1, §6 stage 2
UR-012 Never accept, store, or serve gallery data — reference faces or embeddings SR-005 §5a, and the absence of any such field in §2
UR-013 Windows are scene-scoped claims; the server must not reinterpret their boundaries SR-002 §2 field notes, §6
UR-014 Reject a manifest whose schema_version / jmanifest_version is unknown, never guess SR-003 §2, §6 stage 2

Two notes on UR-001. An existence check is deliberately a separate, cheaper endpoint from the fetch in §4 — it answers "should I bother?" for a whole library sweep without transferring payloads, and it is the endpoint a scheduled task will hammer. It is also the most abuse-prone surface, since it doubles as an oracle for "does the community have this title" — so it is rate-limited harder than the fetches and returns no manifest content.

UR-003's three checks are different in kind and are enforced at different stages: field strictness and size are cheap and synchronous (reject at the door), whereas the TMDB cast match needs an outbound API call and so runs asynchronously after a 202. See §6.

UR-010 to UR-014 were added when this spec was reconciled against the system spec. They are not new work — each states a property the design already had, which had been left implicit because no system requirement existed to trace it to. UR-012 and UR-013 are the two worth stating explicitly: the server's refusal to carry gallery data (SR-005) and its refusal to reinterpret window boundaries (SR-002) are both invariants preserved by not doing something, and an unstated prohibition is the kind that erodes.


1. Why this needs more than the current truth file

The extraction pipeline (result_sink_node) writes:

{
  "schema_version": 1,
  "movie": "/data/movies/Movie.mkv",
  "sample_fps": 1,
  "anneal_sec": 2,
  "actors": [
    { "name": "...", "imdb_id": "...", "tmdb_id": "...", "jellyfin_id": "...",
      "scenes": [[12.0, 45.0]] }
  ]
}

Three properties of this format block sharing as-is:

  1. No portable title identity. movie is an absolute path on the machine that ran extraction. Nothing in the file says "this is The Death of Stalin (2017)". The receiving server cannot tell what it just downloaded.
  2. Installation-local identifiers. movie leaks the contributor's directory layout and jellyfin_id is a GUID from the contributor's database — meaningless and mildly identifying elsewhere. Both must be stripped on upload, not merely ignored on download.
  3. Timings are cut-specific. scenes are absolute seconds. A theatrical cut, an extended cut, a PAL speed-up, and a release with 40s of distributor logos all produce different timelines for the same TMDB id. Keying purely on TMDB id would silently serve misaligned overlays.

The Jmanifest format below is the truth file plus a portable identity block and a cut fingerprint; the actor timeline payload is unchanged.

Terminology

  • Jmanifest — one shareable actor timeline for one cut of one title.
  • Title identity — what the work is (TMDB/IMDB id + episode coordinates).
  • Cut fingerprint — which encode/edit the timings apply to (runtime, plus optional stronger signals).

2. Jmanifest format

{
  "jmanifest_version": 2,
  "identity": {
    "type": "movie",
    "tmdb_id": "504172",
    "imdb_id": "tt4686844",
    "title": "The Death of Stalin",
    "year": 2017
  },
  "cut": {
    "runtime_sec": 6420.5,
    "container_duration_sec": 6420.5,
    "audio_signature": "v1:v7fA3k…"
  },
  "extraction": {
    "sample_fps": 5,
    "extinction_sec": 12,
    "gallery_size": 1820,
    "gallery_scope": "global",
    "pipeline_version": "scene-actor-extraction 0.4.1"
  },
  "actors": [
    {
      "name": "Steve Buscemi",
      "imdb_id": "nm0000114",
      "tmdb_id": "884",
      "scenes": [
        { "start": 191.6, "end": 209.2, "belief": 0.98, "route": "live" },
        { "start": 438.2, "end": 465.6, "belief": 0.81, "route": "deferred" }
      ]
    }
  ]
}

For an episode, identity is:

{
  "type": "episode",
  "series_tmdb_id": "1396",
  "series_imdb_id": "tt0903747",
  "title": "Breaking Bad",
  "season": 2,
  "episode": 5
}

Field notes:

  • jmanifest_version — currently 2. Separate from the plugin's schema_version; this versions the exchange envelope, and the two remain independent by design. They coincide at 2 only because the SR-003 bump touched both. An unknown version is rejected outright (UR-014), never guessed at.
  • identity.tmdb_id / imdb_id — at least one required. These are the lookup keys.
  • cut.runtime_sec — required, the decoded duration of the media the timings came from. This is the primary alignment guard.
  • cut.video_hash — withdrawn. A file hash rather than a cut fingerprint; see §3 "Why there is no file-level signal". Rejected as an unknown field like any other (§6 stage 2), so a client still sending it gets a 400 naming it rather than having it silently dropped.
  • cut.audio_signature — optional; a spectral-peak signature from the media centre, version-prefixed (v1:). Enables content-based matching and offset recovery for files of unknown providence. See §3.
  • extraction.extinction_sec — the re-acquisition timeout that shapes window extent. Replaces anneal_sec; see the schema-bump note below.
  • extraction.gallery_scope — global or limited. The strongest available quality signal when ranking competing manifests for one cut (§7), since a gallery built from the whole library competes against every actor in it, whereas a per-title gallery does not.
  • actors[].jellyfin_id — must not appear. The server rejects uploads containing it (see §6).
  • movie (absolute path) — must not appear. Rejected likewise.
  • actors[].scenes — objects, not float pairs. start/end in seconds, inclusive, sorted. belief is the accumulated posterior that justified the claim, in [0, 1]; route is live, deferred or pooled (extraction AR-017). Both are optional and both are excluded from content_id (§9a). A window is a claim about scene membership, not a recognition event (UR-013, system spec SR-002): an actor who turns away or is off-camera during a reverse shot is still present. The server therefore never reinterprets, merges, splits or trims windows — it stores and serves what it was given, quantised (§9a) but not reshaped. Two windows mean a genuine departure and return.
  • actors[].tmdb_id — the primary actor join key. In practice the extraction pipeline populates this and leaves imdb_id empty (see §6 stage 3), so a manifest without actor TMDB ids will match poorly.
  • actors[].name — sent on upload for matching, but not persisted: the server resolves each actor to a TMDB person id and serves names from its own TMDB-derived table (§5a, §7). On download, name is present and server-authoritative. Contributors should not expect a name they invented to round-trip.

Schema bump — SR-003, shipped at version 2

The truth file and the Jmanifest are consumed by components that ship independently, so breaking changes are batched into one schema_version bump coordinated across all three repos (system spec SR-003). One bump has shipped, moving jmanifest_version to 2 in lockstep with the truth file's schema_version:

Change Effect here
Remove anneal_sec Withdrawn upstream: presence now follows track extent, so a track survives its own gaps and there is nothing to anneal. Deleted rather than kept as a vestigial 0 — a field naming a mechanism the pipeline no longer has is actively misleading
Add extinction_sec Its successor: the parameter that actually shapes window extent
Add gallery_scope New ranking signal (§7)
Per-window belief scenes becomes a list of objects — interval plus posterior and identification route — rather than a list of float pairs
Add audio signature Already specified here (§3, UR-009)

Per-window belief does not weaken §5a. The added fields are a bounded float and a small enumerated string, so an accepted manifest still contains only numbers and closed-vocabulary values. No free-form channel is opened, and SR-004 is preserved.

Two consequences for content_id (§9a), both of which must land with the bump rather than after it:

  • The canonical form currently hashes [start_cs, end_cs] pairs. Once windows carry belief, the canonical form must decide whether belief is part of identity. It should not be: two servers that validated the same upload must agree, and belief is a producer-side estimate that may legitimately differ between pipeline versions for identical timings. Belief is replicated as an attribute, exactly as audio_signature is (§9a).
  • Quantisation is unchanged: integer centiseconds, for the reasons in §9a.

Flag day, not dual-accept. This server accepts jmanifest_version: 2 and rejects everything else outright (UR-014), including version 1. All three components are pre-release, and a v1 read path would be the one nobody exercises — so it is the one that would rot while being carried through every later change to the reader. The consequence is that a pipeline still emitting v1 is incompatible until it is updated, which is stated plainly rather than papered over with a compatibility shim nobody tests.

Series bundles

Series are the primary unit of exchange, not episodes. A user asks for "Breaking Bad", not for 62 individual files, and per-episode round trips would mean 62 requests against the rate limit for one obvious intent.

A bundle is a thin wrapper, not a new format:

{
  "jmanifest_version": 2,
  "series": {
    "series_tmdb_id": "1396",
    "series_imdb_id": "tt0903747",
    "title": "Breaking Bad"
  },
  "episodes": [ { "...a full Jmanifest, identity.type == episode..." } ]
}

Sizing. Measured against the 316 non-empty manifests in the extraction corpus, a minimal (actors-only) manifest is ~4.8 KiB median and ~9.7 KiB at p95. So a 24-episode season is ~112 KiB median / ~228 KiB p95, and even a 62-episode series is well under 1 MiB. Whole-series transfer is therefore the sensible default rather than something to paginate defensively — the response is smaller than a single poster image.

Bundles are capped at 500 episodes and 25 MiB; beyond that the client must page by season.

Partial bundles are normal. The server returns whatever episodes it holds. A bundle with 9 of 13 episodes is a valid, useful response, not an error. Each episode carries its own cut block, so the client matches each one independently — one mismatched episode does not invalidate the rest. The bundle response includes coverage metadata so the client can report it:

{
  "coverage": { "episodes_available": 9, "seasons": [1, 2] }
}

Bundles are a transfer convenience, not a storage unit. Each episode manifest is stored, validated, versioned, reported and delisted individually (§7). There is no "series manifest" row — a bundle is assembled per request. This matters for moderation: one bad episode is delisted on its own without disturbing the other 61.

Series bundle upload

Contributing a whole series is the natural counterpart, and it is the more important half — a worker that has just processed a season should not make 24 separate POSTs, each triggering its own TMDB round trip.

POST /manifests/bundle takes the same envelope. Semantics:

  • Per-episode validation. Each episode runs the full §6 pipeline independently. The bundle is not atomic: valid episodes are accepted and invalid ones rejected, with a per-episode result list. All-or-nothing would let one bad episode discard an entire season's compute.
  • Shared TMDB fetch. All episodes of a series resolve against one cached credits fetch (§6 stage 3), so a 24-episode bundle costs one upstream call rather than 24. This is the main reason bundle upload exists.
  • One rate-limit unit. A bundle counts as a single write against the §5 limit, with a separate per-episode cap, so contributing a season is not punished relative to contributing a film.

Response is 202 with per-episode outcomes:

{
  "results": [
    { "season": 1, "episode": 1, "manifest_id": "01HZ...", "status": "pending" },
    { "season": 1, "episode": 2, "status": "rejected", "reason": "cast_match_below_threshold" }
  ]
}

3. Cut matching

Timings only transfer between identical cuts. Matching is tiered, and the server reports which tier matched so the client can decide whether to trust it.

Tier Signal Confidence
runtime runtimes within ±2s Very likely the same cut
loose runtimes within ±30s Probably same cut, different trims
— beyond that No match; do not serve

The client sends its own runtime when requesting; the server does the matching and returns the best available tier. A loose match should surface as a caveat in the JRay UI rather than being applied silently.

Why there is no file-level signal

Every tier is a claim about a cut, never about a copy. No field in the Jmanifest distinguishes two files of the same cut, and none may be added.

An earlier draft made cut.video_hash — the OpenSubtitles hash of first+last 64 KiB plus file size — the top exact tier, on the reasoning that an equal hash identifies the same file and so can never produce a false positive. It has been withdrawn, for two independent reasons:

  • It was the one field that individuated a copy rather than a work. A cut fingerprint is shared by everyone who holds that edit, however they came by it, and is therefore a statement about the film. A file hash is a statement about one person's particular encode: anyone holding a given release can compute its hash and ask GET /manifests/exists whether the community has a manifest for exactly that file. That made a read endpoint into a release-level oracle, and made a contributor's uploads a published inventory of their own files. Nothing else in the design has that property, and PR-005 is the reason it should not.
  • It bought no accuracy the cut-level tiers lack. Timings transfer between cuts. Two files of the same cut yield the same timings whether or not their bytes agree, so exact never told a client anything audio does not — it only told the server something it had no need to know.

The consequence is accepted rather than mitigated: the server cannot tell a client holding the very file a manifest was extracted from apart from one holding a different encode of the same cut. That is the intended property.

audio (below) is the top tier in its place, and is the better signal on the merits: it confirms the audio actually matches, survives re-encoding, and recovers a trim offset, none of which a byte-level hash can do.

Audio signature — UR-009

Everything above depends on knowing what the file is. When providence is unknown — no TMDB id, no usable metadata, a renamed or badly-tagged file — none of those tiers can fire. And when a release is trimmed differently (distributor logos, PAL speed-up, an extra recap), the runtime tiers correctly decline to match, but the underlying timings would have been reusable if only the offset were known.

A content-derived audio signature solves both. It is stored on every manifest as cut.audio_signature.

Why audio and not video. Audio survives what breaks video hashing: re-encoding, resolution changes, bitrate changes, colour-space conversion, letterboxing. Two releases of the same cut nearly always share an audio track that is perceptually identical even when every video byte differs.

Construction

Sampled from the centre of the media, which avoids the two regions that differ most between releases — logos and cold opens at the head, credits at the tail.

  1. Decode a 120 s window centred on the midpoint (runtime/2 - 60s to runtime/2 + 60s).
  2. Downmix to mono, resample to 11025 Hz.
  3. STFT with a 4096-sample frame, 1024-sample hop (~93 ms/frame, ~1290 frames), Hann window.
  4. Per frame, take the log-magnitude spectrum over 300–3000 Hz — the band carrying dialogue and score, and the most codec-robust.
  5. Divide that band into 32 logarithmically spaced bins and record the index of the peak bin plus a coarse 2-bit energy class.
  6. Pack each frame into one byte; the signature is the resulting ~1290-byte array, base64-encoded.

The result is ~1.7 KB per manifest — negligible against a ~4.8 KiB manifest.

Normative v1 parameters

The six steps above are not sufficient to reproduce a byte stream. Each choice below was underspecified and is now pinned; two independent implementations that differ on any one of them produce signatures that never match, which silently defeats the entire mechanism.

Parameter v1 value
Hann window Periodic (not symmetric)
Band value Mean of linear magnitudes in the band — not sum, not max, and taken before the log
Peak tie-break Lowest band index wins
Energy class log10(frame band-energy / upper-median frame energy), quantised at −0.6 / −0.2 / +0.2
Byte layout `(band << 2)
Base64 Standard alphabet, with padding
Frame count Whole frames only. Over 1 323 000 samples this yields 1288 frames, not "~1290"

The energy class is normalised against the upper-median frame energy rather than an absolute level, which is what makes it invariant to gain and to trim differences between releases. The thresholds straddle the median rather than sitting on it, so a frame near the centre of the distribution does not flip class under small perturbations.

Conformance fixture

A golden fixture is the authoritative tiebreak, because prose cannot pin floating-point behaviour: scene-actor-extraction/tests/fixtures/audio/jray_audio_v1_golden.json.

It carries the expected signature, the decoded-window PCM checksum, the full 32-entry band→FFT-bin table, and the parameter contract — an implementation can be written from that file alone. The PCM checksum is asserted separately from the signature so a codec-level divergence is distinguishable from a DSP one.

Implementations should compute the FFT themselves (radix-2, double precision) rather than depending on a library whose version could change the numerics. Verified decision margins on the fixture are 1.3% between the two strongest bands and 3.6e-3 in log10 to an energy-class edge — many orders above double-precision noise, so any two correct implementations agree.

Measured on that fixture: the peak-band sequence survives a stereo/44.1 kHz round trip and AAC 128 kbit/s re-encoding exactly (score 1.00), which is the codec-robustness this design claims.

This is deliberately a peak-bin signature rather than a full spectrum: peaks survive lossy re-encoding, loudness normalisation and channel-layout differences, whereas absolute magnitudes do not. It follows the same principle as Chromaprint/AcoustID (compact per-frame spectral features, matched by sliding alignment) but is self-contained: no external service is queried, so no lookup leaks which titles an instance holds (§9 privacy).

Matching and offset recovery

Two signatures are compared by sliding one against the other and taking the best score:

for offset in -600 .. +600 frames:        # ±56 s
    score(offset) = fraction of overlapping frames whose peak bin matches
best = argmax score
Result Interpretation
score ≥ 0.85, offset ≈ 0 Same cut, aligned. Timings apply directly
score ≥ 0.85, offset ≠ 0 Same cut, shifted. Timings apply with offset added
0.60 ≤ score < 0.85 Possibly same cut, degraded audio. Flag as loose
score < 0.60 Different content. No match

The second row is the valuable one, and the reason to do this at all: a release with 40 s of extra logos previously failed the ±2 s runtime tier outright. Now it matches, and the client shifts every scene window by the recovered offset. The server returns the offset; the client applies it — manifests are never rewritten, so one stored manifest serves every trim of the same cut.

Offset search is capped at ±56 s, which covers realistic trim differences. Speed-differing releases (PAL 4% speed-up) are not handled by a constant offset and are correctly rejected by the score threshold; a scale-and-offset search is possible later but is out of scope.

Revised tier table

Tier Signal Confidence
audio audio score ≥ 0.85 Same cut; offset returned, may be non-zero
runtime runtimes within ±2s Very likely the same cut
loose audio 0.60–0.85, or runtimes within ±30s Caveat in UI

audio ranks above runtime because it is content-derived: it confirms the audio actually matches, where equal runtimes are only circumstantial. With video_hash withdrawn it is also the top tier — there is nothing above it, and nothing above it that could be added without reintroducing a file-level signal.

With no TMDB id at all, a client can search by signature alone:

POST /manifests/search
{ "audio_signature": "base64…", "runtime_sec": 6420.5 }

The server returns candidate matches with scores, offsets and title identity — letting JRay identify an unidentified file and align to it in one step.

This endpoint is a scaling problem, not a correctness one. A naive implementation compares against every stored signature. Mitigations:

  • Prefilter by runtime (±90 s) before scoring, which eliminates almost everything.
  • Index a coarse hash of the signature (e.g. the peak-bin sequence of every 16th frame) for candidate generation, with full sliding comparison only on candidates.
  • Rate-limit hard (§5): this is the most expensive read endpoint and the most attractive to abuse.

Because it is expensive, POST /manifests/search is optional for a server to implement; GET /federation/capabilities advertises support.

Validation and abuse

The signature is attacker-supplied, so §6 applies:

  • Fixed length (1290 frames ± a small tolerance for seek and encoder differences at the window edges), base64, rejected otherwise. The length is not caller-varying: items too short for the window emit no signature at all (see "Media shorter than the window" above), so there is no legitimate short signature to accommodate. A variable-length blob would be a payload channel — precisely what §5a closes.
  • Each byte is structurally constrained (5-bit bin index + 2-bit energy class), so arbitrary bytes are invalid. This keeps §5a's "no free-form storage" property intact: the field cannot carry meaningful smuggled data.
  • Signatures are never used as a trust signal for cast validity — they establish which cut a manifest describes, nothing more.

Implementation cost — flagged honestly

This is the most expensive addition in the spec, and it is worth being clear where the work lands:

  • Extraction pipeline (C++) — optional, and best deferred. It already links libavformat/libavcodec/libavutil, but ffmpeg_decoder.hpp is video-only, so audio would need libswresample plus an FFT. Since the plugin covers the whole library (below), this is redundant work.
  • JRay plugin (C#) — the primary implementation site, and less costly than it first appears. See below.
  • Server (Rust) — comparison only, no audio decoding. rustfft plus a sliding comparison; the cheapest of the three.

Recommended sequencing: make audio_signature optional. Manifests without one continue to work exactly as today via the existing tiers. Ship the plugin-side computation first, let signatures accumulate, then enable audio-tier matching and the search endpoint once coverage is useful. Nothing above needs to land at once.

Computing the signature in the JRay plugin

The plugin is the right place for this, and it is the only place that covers the whole use case. The extraction pipeline only ever sees files it processes; the plugin sees every item in the library, including the ones with no truth data and unknown providence — which is exactly the population UR-009 targets. A signature must also be computable at query time (to identify a local file), not only at contribution time.

Jellyfin already ships FFmpeg, and the plugin can reach it. Verified against Jellyfin.Controller 10.11.5, which the plugin already references: MediaBrowser.Controller.MediaEncoding.IMediaEncoder is injectable and exposes

Member Use
EncoderPath Absolute path to the server's ffmpeg binary
ProbePath Path to ffprobe
EncoderVersion Version gating
SupportsEncoder(...) Capability check

So there is no new dependency and nothing to bundle — the plugin takes IMediaEncoder through DI (registered in ServiceRegistrator) and invokes the binary Jellyfin is already using for transcoding.

FFmpeg does the hard part. Decode, downmix, resample and format conversion are all a single invocation; the plugin never touches a codec:

ffmpeg -nostdin -v error \
  -ss <runtime/2 - 60> -t 120 \
  -i <media path> \
  -vn -ac 1 -ar 11025 -f f32le -

That streams 120 s of mono 32-bit float PCM at 11025 Hz to stdout — ~5.3 MB, read incrementally rather than buffered whole. -ss before -i makes the seek fast, which matters when sweeping a library.

What remains in C# is only the DSP, and it is modest:

  1. Hann window, 4096-sample frames, 1024 hop (~1290 frames).
  2. Real FFT per frame.
  3. Log-magnitude, 300–3000 Hz band, 32 log-spaced bins, take peak bin + 2-bit energy class.
  4. Pack one byte per frame, base64.

A radix-2 real FFT over 4096 samples is on the order of a hundred lines and has no external dependency. Avoid pulling in a DSP package: this is a fixed, well-specified transform, and vendoring a small implementation keeps the plugin's dependency surface at zero, which matters for a GPLv3 Jellyfin plugin.

Cost is dominated by the FFmpeg seek and decode, not the FFT: roughly a second or two per item, entirely I/O-bound.

Where it runs in the plugin:

  • On demand, for POST /Plugins/JRay/Items/{itemId}/Identify (§9).
  • As a scheduled task that backfills signatures for library items, so a sweep is not blocked on computing them inline. Signatures are cached against the item (keyed on item id + file mtime + size, so a replaced file recomputes).
  • Before contributing a manifest, so uploads carry cut.audio_signature.

Degradation, not failure. If IMediaEncoder is unavailable, the binary is missing, the item has no audio stream, or the file is shorter than the window, the plugin logs and proceeds without a signature. Every existing tier keeps working; UR-009 is an enhancement and must never be able to break a fetch.

Media shorter than the window — 120 s

Items under 120 s emit no signature at all, and no sync offset is applied to them. The window is runtime/2 ± 60 s, so below 120 s it underflows: there is no shortened window to compute, because the construction has no definition there. Such items fall back to the runtime tier, which is adequate — a sub-two-minute item is rarely the ambiguous-providence case UR-009 exists to solve.

The signature is therefore fixed-length by construction, not merely bounded. That is what keeps it inside SR-004: a caller cannot choose the length, so the field cannot be used as a variable-size container (§5a, and "Validation and abuse" below).

Reconciled with scene-actor-extraction IR-007. An earlier draft of this section said items under 150 s got a centred, shortened window with the frame count recorded. That described a mechanism neither producer implements, and it was the weaker rule: a caller-varying length is exactly the property SR-004 forbids. The 120 s cutoff is now identical in both producers and in this server's validator, which is what IR-007 requires — a rule that differs between producers yields signatures that never match.

The extraction pipeline (C++) is then optional for UR-009. It may compute signatures for files it processes — libswresample plus an FFT, as noted above — but since the plugin computes them for the whole library and attaches them on contribution, the pipeline need not implement this at all. That removes the libswresample work from the critical path.


4. API

Base path /api/v1. JSON throughout.

GET /manifests/exists — UR-001

Cheap existence probe. Answers whether a manifest is available for a given title and at what cut-match tier, without transferring the payload.

Query parameters are the same identity + cut parameters as the fetch endpoints: tmdb_id / imdb_id (or series_tmdb_id + season + episode), plus optional runtime_sec.

{ "exists": true, "match": "runtime", "manifest_id": "01HZ...", "actor_count": 34 }

exists: false is returned with 200, not 404 — absence is a normal answer to this question, and using 404 would conflate "no manifest" with "bad route" for the client.

If runtime_sec is omitted, the response reports whether any manifest exists for the title with "match": "unknown"; the client must still fetch to find out whether a cut actually aligns. This is the mode a library-wide sweep uses.

Batch form

A client sweeping a library should not issue one request per item. The batch form takes up to 100 items:

POST /manifests/exists
{ "items": [ { "tmdb_id": "504172", "runtime_sec": 6420.5 }, ... ] }

returning results positionally. This exists specifically so the rate limit in §5 can be generous per request while staying strict per item, and so a 2000-item library sweep is 20 requests rather than 2000. It is a POST only because the payload does not fit a query string; it is a read and requires no token.

GET /manifests/movie?tmdb_id=&imdb_id=&runtime_sec=

Returns the best-matching Jmanifest, or 404 if none clears loose.

{ "match": "runtime", "manifest": { "...": "..." } }

GET /manifests/series/{series_tmdb_id}?season=

Returns a series bundle (§2). season optional; omitted means all seasons. Episode-level cut matching is done client-side against the returned bundle, since a client pulling a whole series already knows its own runtimes.

GET /manifests/episode?series_tmdb_id=&season=&episode=&runtime_sec=

Single-episode equivalent of the movie endpoint.

POST /manifests

Contribute a manifest. Body is a Jmanifest. Requires an API token (§5).

  • 202 Accepted — passed size and schema validation; held unlisted pending the TMDB cast check (§6). Returns { "manifest_id": "...", "status": "pending" }
  • 400 — malformed, or contains an unrecognised or forbidden field (§6)
  • 409 — an identical (identity, cut) manifest already exists from this contributor
  • 413 — body exceeds the size limits (§6)
  • 429 — rate limited (§5)

POST /manifests/bundle — UR-006

Contribute a whole series in one request. Body is a series bundle (§2). Per-episode validation, non-atomic, one rate-limit unit, shared TMDB fetch — see "Series bundle upload" in §2.

  • 202 Accepted — returns per-episode outcomes
  • 400 — the bundle envelope itself is malformed (individual bad episodes are reported in the results list, not as a whole-request error)
  • 413 — exceeds 500 episodes or 25 MiB

GET /manifests/{id}/status

Poll the outcome of the asynchronous cast check for an upload: { "status": "pending" | "listed" | "flagged" | "rejected", "reason": "..." }.

GET /manifests/{id}

Fetch a specific manifest by its server-assigned id (for debugging and for the "report this manifest" flow).

POST /manifests/{id}/report

Flag a manifest as wrong (misaligned, wrong actors). Body: { "reason": "misaligned" | "wrong_actors" | "spam", "note": "..." }.

GET /health

Liveness. Unauthenticated.


5. Authentication and abuse

Reads are anonymous and cacheable. Writes require a token — an anonymous bearer capability, not an account. No email, no verification, no personal data; see §5a for why identity is deliberately not load-bearing.

Rate limiting — UR-004

Limits are per token where one is present, otherwise per source IP. Anonymous reads are keyed on IP, which is imperfect behind CGNAT; the limits below are therefore set well above what a single real server needs.

Surface Limit Rationale
GET /manifests/exists 600 / hour Sweeps should use the batch form
POST /manifests/exists (batch) 60 / hour, ≤100 items each 6000 items/hour — a large library sweeps in one pass
Manifest fetches (/movie, /episode) 300 / hour A client only fetches what exists said was there
GET /manifests/series/{id} 120 / hour Bundles are ~100–250 KiB; this is the preferred path for TV and should not be scarcer than per-episode fetching
POST /manifests 100 / hour per token Nobody uploads faster than the CV pipeline runs
POST /manifests/bundle 20 / hour per token, ≤500 episodes each One unit per bundle, so contributing a season is not penalised versus a film
POST /manifests/{id}/report 20 / hour per IP Reports are a moderation lever; cheap to abuse
POST /manifests/search (audio) 60 / hour Most expensive read endpoint (§3); sliding comparison over candidates
GET /federation/peers 60 / hour per IP Public directory, read by humans; no reason for volume
GET /federation/changes 120 / hour per peer Hourly polling is the default; this allows generous catch-up
GET /federation/manifests/{content_id} 5000 / hour per peer Bootstrap pulls are bulk by nature; capped so one peer cannot saturate egress
POST /federation/have 120 / hour per peer, ≤1000 ids each Diffing a catalogue should be a handful of requests
GET /health unlimited Liveness

Responses carry X-RateLimit-Limit, X-RateLimit-Remaining and X-RateLimit-Reset; exceeding a limit returns 429 with Retry-After. The JRay client must honour Retry-After and back off exponentially rather than retrying tightly — a scheduled library sweep that ignores this will get an instance's IP throttled.

Implemented as a fixed-window counter keyed on (token_or_ip, surface), held in process memory (§8) — no external counter store. A sliding window is not worth the complexity at this volume. Counters reset on restart, which is acceptable for abuse throttling. Read limits are applied behind the CDN cache, so a cache hit costs a client nothing against its budget.


5a. Trust model

Design goal: no accounts, no identity, no moderation queue that scales with users — and no way to use the server as a content host.

The key property that makes this tractable: a Jmanifest is not free-form content. It is a closed-vocabulary document — a title identity, a runtime, and a list of actors with timings. Everything in it is checkable against an external ground truth (TMDB) that the attacker does not control. So trust can attach to content, not to contributors. This is why the server needs no accounts: a manifest listing pornstars for a children's film fails the check regardless of who uploaded it, and a valid manifest is valid regardless of who uploaded it.

Threat 1 — using the server as a content store

The concern is the server being used to host illegal material (the worst case being CSAM) or arbitrary payloads, making the operator liable.

The structural defence is that there is nowhere to put it. After §6 stage 2, an accepted document contains only:

Field Constraint
identity.tmdb_id / imdb_id Regex-constrained to digits / tt\d{7,8}
identity.season/episode/year Bounded integers
cut.* Numbers, plus the fixed-length audio signature (§3)
extraction.* Numbers and a version string from an allow-list
actors[].tmdb_id / imdb_id Regex-constrained
actors[].scenes Pairs of floats
actors[].name, identity.title The only free-form strings

No binary. No images. No URLs. No base64 fields. No extension points — because extra="forbid" applies at every nesting level, an attacker cannot add one.

That reduces the entire content-hosting surface to two short text fields, which are then constrained further:

  • Length caps. name ≤ 200 chars, title ≤ 300. With ≤ 500 actors that is a hard ceiling of ~100 KB of attacker-controlled text per manifest, but see the next two rules, which cut it far below that.
  • Character class. Names must match a permissive-but-closed pattern: Unicode letters, marks, spaces, and . ' - , only. No digits, no / + =, no control characters, no zero-width or bidi-control codepoints, NFC normalised. This alone defeats base64/hex smuggling, which needs digits and padding characters.
  • Cross-check against a known vocabulary. Every actor name must correspond to a real TMDB person (§6 stage 3). A name that matches no TMDB person is not stored at all. An attacker therefore cannot write arbitrary strings — only strings that already exist in TMDB's person index.

Combined, the last rule is decisive: the server does not store attacker-authored text, it stores references to TMDB entities. The strongest form of this — and what I recommend for v1 — is to go one step further and not persist the submitted name string at all:

Store tmdb_person_id plus the timings. Resolve display names from the server's own TMDB-derived person table at serve time. The uploaded name field is used only for matching during validation, then discarded.

At that point the free-text channel is closed completely. The only attacker-controlled values that reach the database are integers. There is no CSAM risk and no payload-smuggling risk because there is no field capable of carrying either.

This also resolves your point about not storing JSON files: with names normalised to person ids, the natural representation is relational rather than a blob. See §7.

Threat 2 — prank and vandalism manifests

Semantically valid but wrong: casting pornstars in a children's film, or mislabelling a film's cast as a joke. Every structural check passes; only ground truth catches it.

Defence is the TMDB cast cross-check in §6 stage 3. Its effectiveness rests on the attacker not controlling TMDB: to make a pornstar manifest pass, they would need those performers to be credited cast on that title in TMDB, which means vandalising TMDB itself — a separate, moderated system with its own edit history. That is a meaningfully high bar for a prank.

Additional layers, in order of cost:

  1. Category guard. Reject any manifest where a matched TMDB person's known-for department or credits are dominated by titles TMDB flags as adult (adult: true), unless the target title is itself flagged adult. This directly targets the stated prank without needing a blocklist of names.
  2. Age-appropriateness guard. If the target title's TMDB certification is a children's rating, apply the strictest cast-match threshold and require an audio or runtime cut match. Mismatched content on children's titles is the highest-harm case and deserves the tightest gate.
  3. Divergence detection. When two manifests exist for the same (title, cut) from different sources and their actor sets disagree beyond a threshold, flag both and serve the one with the better cast-match ratio. Honest extractions of the same cut converge; a prank diverges from them.

What replaces accounts

Contribution requires a token, but a token is not an account — it is an anonymous bearer capability:

  • Self-issued on request, no email, no verification, no personal data.
  • Stored only as a hash. The server cannot enumerate who holds tokens.
  • Its sole purposes are rate-limiting attribution (§5) and revocation.
  • Discarding a token and requesting another is trivially easy — and that is fine, because the token is not the defence. The content checks are. A new token gains an attacker nothing, since every upload faces the same ground-truth validation.

This is the crucial difference from an account system: the token exists to throttle volume, not to establish identity. Sybil resistance is not required because identity is not load-bearing.

Consequently the only reputational state is per-token counters (accepted, rejected, flagged), used for one purpose: a token whose rejection rate exceeds a threshold over a minimum sample is revoked automatically, and its pending/flagged manifests are dropped. No human is in the loop for the common case.

Residual risk and the operator's lever

Two things remain that automation cannot fully close:

  1. A manifest that is plausible but wrong — correct cast, deliberately misaligned timings — degrades the overlay but carries no legal or safety risk. Reports plus divergence detection handle it.
  2. A novel abuse pattern nobody anticipated.

For both, the operator needs a kill switch, not a moderation queue: status transitions (§7) are a single column, so delisting a manifest, every manifest from a token, or every manifest for a title is one UPDATE. Delisting is instant and reversible; deletion is a separate, logged action.

Legal posture. Because the server stores only integers and references to TMDB entities, it holds no user-generated content in the sense that intermediary-liability regimes contemplate. It is a materially better position than "we store user-submitted JSON and moderate it".

Stated in full, against EU and international copyright law, in docs/legal-posture.md — which an operator should publish alongside a contact address for notices. That document is downstream of this spec, not alongside it: every claim in it is a consequence of a design property recorded here, so a change that weakens SR-004 or SR-005 silently invalidates it. Its §7 is the list of changes that would.


5b. The licence contributed manifests carry — UR-019

Contributed manifests are CC0 1.0 Universal. The full text is LICENSE-DATA; the code is separately GPL-3.0-or-later, and the two must not be conflated — the code licence says nothing about the data, and the data is the part that replicates between instances.

This follows the practice of every comparable service (MusicBrainz core data and AcousticBrainz are CC0; AcoustID is CC BY-SA), and it closes a gap rather than adding a feature: until it was stated, contributed manifests were in no declared condition at all, and §9a replication had no grant flowing through it.

Why CC0 and not a share-alike licence. A share-alike licence works by asserting a right in the data and then conditioning its use. The position throughout — docs/legal-posture.md §3 — is that presence timings are facts rather than protectable expression. Asserting copyright in them in order to license them would contradict that argument in the same repository, and that contradiction is worth more to an opponent than a share-alike licence is worth to the project. CC0 asserts nothing, which is the position actually taken.

Two further consequences, both load-bearing:

  • CC0 waives the sui generis database right by name, not merely copyright. That closes the EU-specific residual exposure from the contributor's side, in the one jurisdiction where such a right exists to be waived.
  • Federation needs no per-peer negotiation. §9a has independent operators replicating each other's catalogues wholesale; without a grant reaching every peer, each hop is unlicensed. CC0 makes each one a non-event, and means no instance can become a chokepoint by withholding permission to mirror.

The grant is taken at token issuance, and this is not incidental. There are no accounts, so there is no sign-up to attach terms to; and a manifest arrives over POST /manifests with no channel to negotiate over. Acquiring the contribute capability is therefore the only moment at which a grant can be made, so POST /tokens returns the licence and its terms alongside the token. A licence the server publishes but never delivers is one no contributor agreed to.

Scope, stated precisely and repeated in the terms themselves: the grant covers the manifest — timings, identifiers, audio signature. It does not, and cannot, license the underlying work, which is not the contributor's to license and which this server does not hold. That sentence is the whole of §§1–4 of the legal posture restated, and it belongs in the terms in exactly that form.


Client-side hardening

Independent of the server, because a compromised or hostile server must not be able to attack its clients:

  • The JRay overlay renders actor names as text nodes only, never as HTML.
  • The plugin validates downloaded manifests against the same schema it would apply to an upload — a client must not trust a manifest merely because the server served it.
  • Downloaded manifests are stored via the existing IManagedTruthStore and never written into the media library filesystem.

6. Upload validation — UR-003

Validation runs in four stages, ordered cheapest-first so that abusive uploads are rejected before they cost anything.

Stage 0 — reject on headers, before the body is read

The cheapest rejection is the one that happens before any payload is accepted.

In one line: Content-Length is the fast path; the streaming byte counter is the enforcement. The header is a claim by the client, so it rejects honest oversized uploads early and cheaply, but it cannot be the only check — a lying header, a chunked upload, or a compressed body all pass it. Implement both; in Axum they are the same one-line layer (§8).

Three mechanisms, in order of how early they fire:

1. Expect: 100-continue (earliest — body genuinely never sent). A client may send headers with Expect: 100-continue and wait before transmitting the body. The server responds 100 Continue or, if Content-Length already exceeds the cap, 413 — and the body is never transmitted at all. This is the true "reject before upload".

Clients under this project's control — the JRay plugin contributing manifests (§9) and federation peers pulling (§9a) — should use Expect: 100-continue for uploads, because it turns a rejected 25 MiB bundle into a two-header exchange. The server must handle it correctly, but must never depend on it: arbitrary clients will not send it.

2. Content-Length check (normal case). When present, the declared length is available in the request headers before the body. If it exceeds the cap for that route, respond 413 immediately and do not read the body.

This is a filter, not a guarantee: a hostile client can declare Content-Length: 100 and then send gigabytes. The streaming cap below is therefore mandatory, not redundant.

3. Chunked requests have no declared size. HTTP/1.1 Transfer-Encoding: chunked omits Content-Length entirely, so there is nothing to check up front. These must be capped while streaming.

What "before anything is uploaded" can and cannot mean. Even on an immediate 413, a client has typically already put some body bytes on the wire — they may sit in kernel or proxy buffers before the response lands. The achievable guarantee is that the server never reads, buffers, or parses an oversized body, and closes the connection promptly. It is not that zero bytes cross the network. Do not size defences on the assumption that a 413 prevents transmission.

Stage 1 — size limits while streaming (before parsing)

Enforced at the reverse proxy and again in the app, on the raw body, before JSON parsing. A parser handed an unbounded body is a denial-of-service primitive, so this must not be deferred to the schema layer.

The app-level cap counts bytes as they are read and aborts mid-transfer once exceeded, rather than reading to completion and then measuring. This is what makes a lying Content-Length and a chunked upload both safe.

Limit Value
Request body (movie or episode manifest) 2 MiB
Request body (series bundle, §2) 25 MiB
Request body after gzip decompression 8 MiB, with a max compression ratio of 20:1
JSON nesting depth 12
actors[] entries 500
scenes[] entries per actor 2000
Total scene windows across all actors 20000

For scale: the sample feature film in the extraction repo has ~30 actors and a few hundred windows. These caps are roughly an order of magnitude above anything legitimate. Oversized bodies are rejected with 413.

The decompression-ratio cap matters because a gzip bomb passes a 2 MiB body-size check trivially. Decompression must also be streamed with a running output cap — decompressing fully and then checking the size defeats the point.

Implementation. Axum's DefaultBodyLimit (§8) implements the streaming cap and honours Content-Length for early rejection, applied per-route so the bundle endpoint gets its larger limit without widening the others. Set the matching client_max_body_size (nginx) / request_body max_size (Caddy) at the proxy so oversized uploads are dropped at the edge and never occupy an application worker.

Stage 2 — strict schema (synchronous, rejects with 400)

No additional fields anywhere. Every object in the document is validated in strict mode — serde #[serde(deny_unknown_fields)] on every DTO (§8) — so an unrecognised key at any nesting level is an error, not something silently ignored. This is the default posture, not a special case for the two fields below, and it is enforced by the type definitions rather than by validator code that could omit a field.

Rejected outright:

  • any unrecognised field, at any level of the document
  • movie present, or any string anywhere that looks like an absolute filesystem path (/…, C:\…, \\…) or a file:// URI
  • actors[].jellyfin_id present and non-empty
  • missing cut.runtime_sec
  • neither tmdb_id nor imdb_id in identity
  • imdb_id not matching ^tt\d{7,8}$, actor imdb_id not matching ^nm\d{7,8}$, tmdb_id not matching ^\d{1,9}$
  • scene windows with end < start, negative times, non-finite values (NaN/Infinity), or times beyond runtime_sec + 5s tolerance
  • actor names longer than 200 characters, containing control characters, or failing Unicode normalisation to NFC
  • duplicate actors within one manifest (same imdb_id)

Rejecting unknown fields is what makes the movie / jellyfin_id strip in §9 verifiable: a client that forgets to strip them gets a hard 400 naming the offending field, rather than quietly publishing a contributor's directory layout.

Stage 3 — TMDB cast cross-check (asynchronous, after 202)

This needs an outbound TMDB call and so cannot run inside the request without coupling upload latency to a third party. The upload is accepted with 202 and the manifest is held unlisted until the check completes; it is not served to anyone in the meantime.

The check: fetch the TMDB credits for identity.tmdb_id, take the set of credited cast TMDB person ids, and compare against the actors in the manifest.

Join on tmdb_id, not imdb_id. This is grounded in the actual pipeline output, not assumed. Of the 331 manifests in the scene-actor-extraction repo, exactly one has IMDB ids populated; the other 330 have imdb_id: "" with tmdb_id set. The Jellyfin-gallery path (make_jellyfin_gallery.py) — which is the path most users will take, since it needs no TMDB key — yields TMDB ids only: 2391 of 2392 gallery entries have a tmdb_id, and none have an imdb_id.

A design keyed on IMDB ids would therefore fall back to name matching for essentially every real upload, which is exactly the weak path. tmdb_id is the join key; imdb_id is an optional secondary signal when present.

Let M = actors in the manifest, C = credited cast from TMDB.

Condition Outcome
|M ∩ C| / |M| ≥ 0.6 Listed. Normal case
0.3 ≤ ratio < 0.6 Listed, flagged for review; served with reduced ranking
ratio < 0.3 Rejected. Manifest is deleted and the contributor notified
TMDB has no credits for the id Listed, flagged — absent data is not evidence of a bad manifest
TMDB unreachable / rate-limited Retry with backoff; stays unlisted, not rejected

The match is deliberately approximate and directional. It asks "are these plausibly this film's cast?", not "is this cast list complete":

  • Ratio is over M, not C. A manifest legitimately contains only actors who were both credited and detected on screen, so it is normally a strict subset of the cast — penalising it for missing credited actors would fail every honest upload. The corpus bears this out: median 7 actors per manifest, against feature casts several times larger.
  • Uncredited appearances, cameos, and actors TMDB lists only under a differently-spelled name are exactly why the threshold is 0.6 and not 1.0.
  • Matching is on tmdb_id (see above), with imdb_id as a secondary signal when present, falling back to case- and accent-insensitive name comparison. Name-only matches are counted but capped at half the intersection, so a manifest cannot pass on name collisions alone.

Small-M handling. With a median of 7 actors, a ratio threshold is coarse — one mismatch moves it by 14%. So:

  • |M| ≥ 5: apply the ratio table above.
  • 2 ≤ |M| < 5: require all but one actor to match. A ratio is meaningless at this size.
  • |M| ≤ 1: accept only if the single actor matches; such a manifest is near-worthless anyway and is ranked last.
  • |M| == 0: reject. 15 of the 331 corpus files have empty actor lists — these are extraction failures, not contributions, and must not be uploaded. The client should refuse to submit them.

Every matched actor is resolved to a TMDB person id, and unmatched actors are dropped rather than stored. This is what closes the free-text channel described in §5a: a manifest is persisted as a set of TMDB person references, so a name that corresponds to no TMDB person never reaches the database.

For episodes the check runs against the union of TMDB's per-episode credits (cast + guest stars) and the series' aggregate credits.

Using the union rather than either alone matches what the extraction client already does: run_from_jellyfin.py defaults to --episode-cast tmdb, taking TMDB per-episode credits and falling back to series-wide when the episode has no usable credits. Checking against per-episode credits alone would reject recurring cast that TMDB lists only at series level; checking against series-wide alone would reject legitimate guest stars. The union admits both, and since the ratio is over the manifest's actors (not TMDB's cast), widening the reference set costs nothing in strictness against pranks — a pornstar is in neither set.

TMDB responses are cached (24h) so that a burst of episode uploads for one series costs a single upstream call, and so the server stays within TMDB's own rate limits.

Sanity-checked and warned, not rejected

  • total on-screen coverage implausibly high (>95% of runtime) or near zero
  • an actor whose windows sum to under a second
  • sample_fps below 1, which yields low-quality timings

Because the contributed manifest is stripped of jellyfin_id, the downloading server resolves actors locally via tmdb_id (primarily) against its own People ProviderIds — which is exactly the fallback path JRay already implements.


7. Storage

SQLite in WAL mode (§8), fully relational — no JSON blobs on the write path. The schema below is portable SQL and runs unchanged on Postgres should an instance ever outgrow SQLite.

Storing the uploaded document as a JSON payload would undermine §5a: a blob is an opaque container, so whatever the schema validator missed gets persisted verbatim and served back out. Decomposing into columns means the database can only represent what the schema models — there is physically nowhere for an unexpected field or a smuggled string to live. Normalisation is a security control here, not just tidiness.

Concretely, the submitted JSON is parsed, validated, resolved to TMDB person ids, written as rows, and discarded. The document served to clients is reconstructed from those rows, never echoed.

contributors (id, token_hash, created_at, revoked_at,
              accepted_count, rejected_count, flagged_count)

people       (tmdb_person_id PK,      -- server-side, TMDB-derived
              name,                   -- from TMDB, never from an upload
              adult bool, updated_at)

titles       (id PK, kind,            -- movie | series
              tmdb_id, imdb_id, name, year,
              adult bool, certification, updated_at)

manifests    (id PK, title_id FK, season, episode,
              runtime_sec,                 -- no video_hash: withdrawn, §3
              audio_signature blob NULL,   -- §3, ~1290 bytes
              audio_sig_coarse blob NULL,  -- candidate-generation index key
              sample_fps, extinction_sec, pipeline_version,
              gallery_size, gallery_scope,  -- ranking signal, §2
              contributor_id FK,
              status,                 -- pending | listed | flagged | rejected
              cast_match_ratio real, created_at)

manifest_actors (manifest_id FK, tmdb_person_id FK,
                 PRIMARY KEY (manifest_id, tmdb_person_id))

scenes       (manifest_id FK, tmdb_person_id FK,
              start_cs integer, end_cs integer)   -- centiseconds, see §9a

reports      (id, manifest_id FK, reason, note, created_at, source_ip_hash)
tmdb_cache   (tmdb_id, kind, credits json, fetched_at)

jobs         (id, kind,               -- cast_check | federation_pull
              payload, run_after, attempts, last_error)

jobs is the background queue (§8) — a table rather than an external broker, so pending work survives a restart.

Note what is not in this schema: no actor-name column on any upload-derived table. Display names come from people.name, populated from TMDB by the server. tmdb_cache is the sole JSON column and holds TMDB's responses, not users'.

Scene times are stored as integer centiseconds (§9a), not floats — the same quantisation used for content_id, so stored values and hashed values cannot diverge.

Indexes on titles(tmdb_id), manifests(title_id, runtime_sec), manifests(title_id, season, episode), and scenes(manifest_id, tmdb_person_id). All read queries filter status IN ('listed','flagged'), so a partial index on that predicate keeps the hot path small.

scenes is the only table with real row volume — roughly (actors × windows) per manifest, capped by §6 at 20000 rows. At corpus-realistic sizes (median 7 actors) it is a few hundred rows per manifest, so even tens of thousands of manifests stay comfortably small.

Multiple manifests may coexist for the same title with different cuts — that is the point. Multiple manifests for the same cut from different contributors are allowed too; serve the one with the best (cast_match_ratio, reports, gallery_scope, sample_fps) ranking.

gallery_scope enters the ranking because it is the strongest available quality signal between two otherwise comparable manifests (§2): a manifest extracted against a global gallery had to distinguish its actors from every other actor in the contributor's library, whereas a limited one only had to distinguish them from that title's own cast. The former surviving the cast check is stronger evidence than the latter doing so. It ranks below cast_match_ratio and reports, which are evidence about this manifest rather than about the conditions that produced it.


8. Stack

Traffic is low and read-dominated; a reverse-proxy or CDN cache in front keeps the app tier trivial.

Two capabilities beyond request handling and storage are load-bearing rather than optional:

  • Rate-limit counters (§5).
  • Asynchronous background work — the TMDB cast check (§6 stage 3) and its retries, plus federation pulls (§9a).

Both are satisfied in-process by the recommended stack below; neither requires a separate service.

The server needs a TMDB API key as operational configuration. It is a hard dependency for UR-003: if TMDB is unreachable, uploads accumulate in pending rather than being listed unverified.

Rust + Axum + SQLite, behind an operator-provided reverse proxy.

Layer Choice Version at time of writing
Framework Axum + Tower/tower-http axum 0.8
Runtime Tokio 1.x
Database SQLite (WAL mode) 3.4x
DB access sqlx (compile-time checked SQL) or rusqlite sqlx 0.9 / rusqlite 0.40
Serialization serde / serde_json 1.x
HTTP client reqwest (TMDB, federation pulls) 0.12
Observability tracing + OpenTelemetry exporter —
Edge / TLS Operator-provided (Caddy, nginx, Traefik) —
Packaging Single static binary + one DB file —

Why this fits. The workload is read-dominated, low-volume, and cache-frontable; the largest response is a ~228 KiB series bundle. Raw throughput is not the constraint — Postgres/SQLite queries and TMDB calls are. What does matter here is operational simplicity for hobbyist operators (§9a expects independent people to run instances) and strictness at the validation boundary (§6, §5a). Rust serves both: a single static binary plus one file is the lowest-friction thing an operator can deploy, and a strict type system at the parse boundary is exactly the posture §5a asks for.

Axum over Actix Web. Actix leads on raw throughput by ~10–15% under heavy load, which is irrelevant at this volume. Axum's Tower middleware composition maps directly onto what the spec needs — rate limiting (§5), body-size limits (§6 stage 1), tracing, timeouts — as composable layers rather than bespoke code. It is the mainstream default for new services and the easier codebase for occasional contributors.

Serde deny_unknown_fields is the §6 stage 2 enforcement mechanism. This is the strongest argument for Rust here. #[serde(deny_unknown_fields)] on every DTO gives the "no additional fields anywhere" requirement structurally, checked at compile time against the type definitions, with no possibility of a field being silently accepted because a validator forgot it. Combined with newtypes for TmdbId, ContentId, and centisecond timestamps, malformed input fails to parse rather than being caught later — invalid states become unrepresentable rather than merely rejected.

Axum's DefaultBodyLimit enforces the §6 stage 0/1 caps in the framework: it rejects on Content-Length before reading a body, and caps the stream for chunked or mis-declared uploads — satisfying "reject before parsing" without trusting the client's declared size.

SQLite: the write-concurrency question

SQLite is the right call, but it has one hard constraint that must be designed around rather than discovered: only one writer at a time, even in WAL mode. Concurrent write transactions return SQLITE_BUSY.

Reads are unaffected — WAL gives concurrent readers alongside the single writer, which suits a read-dominated workload well. The risk is concentrated in this spec's three bulk-write paths:

Path Write shape Risk
Single manifest upload ~106 scene rows median Negligible
Series bundle upload (§2) 24 episodes × ~106 rows ≈ 2.5k rows Moderate — one long transaction
Federation bulk ingest (§9a) Thousands of manifests This is the real one

Federation bootstrap is explicitly a bulk-write workload, and it runs concurrently with live uploads. Mitigations, which are requirements rather than suggestions:

  • WAL mode, plus busy_timeout (5s) so contention waits rather than errors, and synchronous = NORMAL (safe under WAL).
  • A single writer connection, serialized through one task/actor, with a read pool alongside. Do not point a multi-connection pool at writes and rely on busy_timeout to sort it out — serialize deliberately.
  • Chunked ingest transactions. Federation ingest commits per manifest, not per batch, so a bootstrap never holds the write lock for long. Combined with the §9a MaxIngestPerHour cap, live uploads are not starved.
  • Batch inserts within a transaction for a manifest's scene rows — one transaction per manifest, not per row.

With those, a single modest VPS handles this comfortably. SQLite does tens of thousands of writes/sec on modern hardware; the constraint is lock duration, not throughput.

When to reconsider. If an instance ever runs multiple writer processes, or federation bootstrap contention becomes visible in practice, Postgres is the escape hatch. Keep the SQL portable and use sqlx (which supports both) so the migration is a configuration change rather than a rewrite. Do not design around a hypothetical Postgres future at the cost of SQLite's simplicity now.

Alternatives considered

The single-writer constraint above is the one real weakness, so it is worth being explicit about why SQLite still wins.

Option Verdict
SQLite (rusqlite / sqlx) Chosen. Ubiquitous, unmatched track record, trivially portable, one file
Turso (SQLite rewritten in Rust, MVCC) Strong future candidate; pre-1.0 today
libSQL (C fork of SQLite) Viable, but its own maintainers now direct effort at Turso
Postgres The escape hatch, not the default — a service to operate, against §9a's goal
redb / fjall / sled Wrong data model — see below
DuckDB Analytical (OLAP); this is a transactional point-lookup workload
SurrealDB Far larger surface area than needed; not an embedded-first story

Key-value stores are the wrong shape, not merely a different one. redb is mature (v4.1, actively developed) and gives MVCC with concurrent readers plus a single writer — but it is a key-value B-tree with no SQL, no secondary indexes and no joins. §7 is a genuinely relational schema: foreign keys between manifests, scenes, manifest_actors and people, partial indexes on status, and queries that join and filter across them. On a KV store all of that becomes hand-maintained index keys and application-side joins — more code in exactly the layer where §5a demands correctness. sled is additionally out on maintenance grounds (last release October 2024).

Turso deserves a serious look, just not yet. It is a clean-room Rust rewrite of SQLite whose BEGIN CONCURRENT / MVCC mode (PRAGMA journal_mode = 'mvcc') directly removes the single-writer limitation described above — the precise weakness in this design. It is SQLite-compatible, so the schema and most queries carry over, and it is developed with deterministic simulation testing. But as of this writing it is pre-1.0; the maintainers state it powers production systems while being explicit that they have not yet reached their "SQLite-level reliability" bar. It also shifts work onto the application: MVCC transactions that touch overlapping data return a conflict error and must be retried, so the caller owns retry logic.

For a small, low-write, read-dominated service where the mitigations above already keep lock duration short, adopting a pre-1.0 database to solve a problem this workload does not yet have is the wrong trade. The recommendation is therefore:

Build on SQLite, keep the SQL standard and the data-access layer behind a thin trait. Re-evaluate Turso when it reaches 1.0 or if federation ingest contention shows up in real operation. Because Turso is SQLite-compatible, that migration is far cheaper than the Postgres one — which is itself an argument for not over-engineering now.

A note on portability. "Portable" here means two distinct things, and SQLite is best at both: the file is portable (a single database file an operator can copy, back up, or hand to someone bootstrapping a mirror), and the SQL is portable (standard enough to move to Postgres or Turso later). Any KV store sacrifices the second entirely.

What Rust changes elsewhere in the spec

  • No Redis. Rate-limit counters (§5) live in process memory (governor or a Tower layer) or in SQLite. A single-process server does not need an external counter store, and dropping Redis removes a whole moving part. Note the tradeoff: in-memory counters reset on restart, which is acceptable for abuse throttling and avoids a dependency for a hobbyist deployment.
  • No separate worker process or broker. The async TMDB cast check (§6 stage 3) and federation pulls (§9a) run as Tokio background tasks in the same binary, with the job queue as a SQLite table so state survives restart. This replaces arq/Dramatiq/Celery entirely.
  • The TMDB cache (§7 tmdb_cache) stays a table; SQLite's JSON functions cover the jsonb usage, which is only caching TMDB responses.

Net effect: one binary, one database file, one reverse proxy. That is a materially better deployment story for federation than "app + worker + Postgres + Redis", and federation only works if running an instance is easy.

Cost of choosing Rust

Stated honestly, since the alternative was Python:

  • The extraction side is Python, so validation and canonicalisation logic (notably the §9a content_id canonical form) can no longer be shared as one implementation. It must be specified precisely enough to reimplement, and cross-tested — a golden-vector test fixture shared by both sides.
  • Fewer casual contributors than a FastAPI codebase would attract.
  • Slower initial development.

These are real, and worth accepting for a long-lived service whose main risks are hostile input and operator friction — both of which Rust directly addresses.

Deployment notes

  • Ship a single static binary (musl target) plus the SQLite file. Optional container image, but neither Docker nor Compose should be required.
  • Ship a Debian package as well (DR-015). It is a third distribution form alongside the raw binary and the container, not a replacement for either: it packages the same musl binary CI has already proved static, so what an operator installs is byte-identical to the artifact that was verified. Published to the Gitea Debian registry, so apt install jray-server and ordinary upgrades work, and attached to the release for operators who would rather not add a third-party apt source.
  • Put all database access behind a thin repository trait rather than scattering queries through handlers. This is what keeps the Turso/Postgres options above cheap, and it localises the single-writer serialization described earlier in one place instead of every call site.
  • Avoid SQLite-specific SQL where a standard form exists (notably INSERT … ON CONFLICT, which is portable, versus INSERT OR REPLACE, which is not).
  • Terminate TLS at the operator's proxy; the app speaks plain HTTP on loopback and must trust X-Forwarded-For only from that proxy — §5 rate limiting and report attribution key on client IP, so a spoofable header defeats both. Make the trusted-proxy CIDR explicit configuration, not a default-on behaviour.
  • Enforce the §6 stage 1 body cap at both the proxy and DefaultBodyLimit; defence in depth, and the app must be safe when run without a proxy.
  • Set a statement timeout and request timeout (tower_http::timeout) so a slow bundle query fails fast.
  • Health checks: GET /health for liveness, plus a readiness check verifying the database opens and migrations are current.
  • Back up the SQLite file with VACUUM INTO or the backup API (never a plain file copy of a live WAL database). Manifests represent real CV compute; federation (§9a) gives partial resilience but is not a backup.

First-run configuration — DR-016

The settings an operator must get right are not discoverable from the binary, and two of them fail silently when unset: without JRAY_TMDB_API_KEY every upload stays pending and is never listed, and without JRAY_TRUSTED_PROXIES the X-Forwarded-For header is ignored, so every client shares one rate-limit bucket and every abuse report points at the proxy. Both produce a working server that is quietly doing the wrong thing — the worst kind of default to leave to a README.

So the package asks, at install time, via debconf: public hostname, listen address, trusted proxies, TMDB key, contact, and whether to publish the peer directory. It is re-runnable with dpkg-reconfigure jray-server, and preseedable for unattended installs.

Three properties this has to hold, each of which is a way packaging usually goes wrong:

  • The generated config is not a dpkg conffile. It is written from the debconf answers, so shipping it as a conffile would make dpkg prompt on every upgrade about changes the package itself had made.
  • Hand edits survive. Only the keys debconf manages are rewritten; comments, ordering and any other setting are left alone, so editing the file directly and running dpkg-reconfigure later do not fight.
  • A blank API key on reconfigure keeps the existing one. Otherwise pressing Enter through a reconfigure would silently unpublish every future upload.

The database is not removed on purge, which knowingly departs from the usual expectation. Manifests are the output of real CV compute on media the operator may no longer have, and federation is explicitly not a backup; destroying that during an apt purge is not a trade worth making for tidiness. postrm says where the file is and leaves the decision to the operator.

An example nginx site ships in /usr/share/doc/jray-server/examples/ rather than being installed into any nginx configuration directory. The proxy commonly runs on a different host from the server, so a file dropped into this machine's nginx would be in the wrong place — and §8 leaves the edge to the operator deliberately.


9. JRay plugin integration

Configuration

  • Enable manifest sharing (default off — this is a network egress feature and must be opt-in)
  • Servers — an ordered list, not a single URL. See below.
  • Contribute manifests (separate opt-in from downloading; off by default)
  • Minimum accepted match tier (audio / runtime / loose)
  • Compute audio signatures (default off) — enables audio-tier matching and unknown-providence search (§3). Uses the FFmpeg binary Jellyfin already ships, via IMediaEncoder.EncoderPath; no extra dependency.

Multiple servers

The plugin queries a user-configured ordered list of servers rather than one. Each entry is:

Field Purpose
Url Base URL
Name Display label
Token Optional; required only to contribute
Enabled Toggle without deleting
AllowContribute Per-server, independent of fetching
TrustLevel Full / FetchOnly — see below

A default entry for the community instance ships pre-configured but disabled, so no traffic leaves an installation until the admin opts in.

Resolution order. For a fetch, servers are tried in list order and the first acceptable result wins — acceptable meaning it clears the configured match tier. Order is the user's trust ranking, made explicit. Rationale for first-match over best-match: querying every server for every item multiplies egress, leaks the library to more parties, and the ordering already encodes which source the admin prefers. A Best match across servers toggle is a reasonable later addition, off by default.

For a series bundle, first-match applies per episode, not per bundle: fetch the bundle from server 1, then query server 2 only for the episodes still missing. A series is commonly split across sources, and this is where multi-server earns its keep.

Failure isolation. A server that is unreachable, slow, or returning errors is skipped after a short timeout (5s connect, 30s read) and marked temporarily failed with exponential backoff. One dead server must never stall a library sweep. Failures are surfaced per-server in the config page.

Contribution is never fanned out. A manifest is contributed only to servers with AllowContribute set, and each is an explicit choice. The plugin must not broadcast uploads to every configured server — that would multiply the privacy exposure described below without the user intending it.

Trusting third-party servers

This is the part that does not come for free. Everything in §5a is a property of a correctly operated server. Pointing the plugin at an arbitrary URL inherits none of it: a hostile server can serve malformed manifests, wrong casts, or oversized payloads.

The plugin therefore treats every server as untrusted, including the default one, and re-applies client-side what the server applies on upload:

  • Validate on receipt. Downloaded manifests are validated against the same strict schema used for uploads (§6 stage 2) — unknown fields rejected, sizes capped, scene windows bounds-checked against the item's real runtime. A manifest is never trusted merely because a server served it.
  • Response size caps enforced during streaming, so an unbounded body is aborted rather than buffered. Bundle cap 25 MiB, single manifest 2 MiB.
  • HTTPS required for non-loopback servers; certificate validation must not be disabled. A plaintext community server would let any network intermediary rewrite actor overlays.
  • Names rendered as text, never markup (§5a client-side hardening). This is the single most important control, because it holds even if every other check is bypassed.
  • TrustLevel: FetchOnly — the default for user-added servers — accepts manifests but never contributes to them and never sends library inventory beyond the single item being queried.

The honest framing for the config page: adding a third-party server means trusting its operator not to serve you deliberately wrong actor data. The structural protections above bound the damage to bad overlay content; they cannot make wrong data right.

Endpoints

Mirroring the existing Truth/Tasks controllers:

  • POST /Plugins/JRay/Items/{itemId}/Fetch — resolve the item across the configured servers in order and, on a match at or above the configured tier, store the result via the existing IManagedTruthStore. Admin key. When the response carries a non-zero offset (§3 audio tier), the plugin must add it to every scene window before storing — the stored truth file is always in the local file's own timebase, so the overlay and the jray?t= query need no offset awareness at read time.
  • POST /Plugins/JRay/Series/{seriesId}/Fetch — bundle fetch for a whole series, with per-episode gap-filling across servers as described above.
  • GET /Plugins/JRay/Servers/Status — per-server reachability and last-error, for the config page.
  • POST /Plugins/JRay/Items/{itemId}/Identify — compute the item's audio signature and search configured servers by content (§3), for items whose providence is unknown. Returns candidate titles with scores and offsets; storing a result is a separate confirmation step, never automatic.
  • A scheduled task that walks items with no truth data and attempts a fetch, reusing the same backlog logic as Tasks/Pending, using the batch exists endpoint (§4) so a sweep is a handful of requests per server.

Contribution runs the reverse: on a PUT .../Truth from a local worker, if contribution is enabled, strip movie/jellyfin_id, attach identity from the item's ProviderIds and its measured runtime, and POST to each contribute-enabled server. For a series, batch into a bundle upload rather than per-episode posts.

Uploads should set Expect: 100-continue (§6 stage 0) so a server that is going to reject the request on size or auth does so before the body is transmitted. This matters most for series bundles, where a rejected upload would otherwise push tens of MiB pointlessly.

Privacy

Contribution reveals to the server operator that some instance holds a given title. Fetching reveals the same thing. That is inherent, but it means:

  • opt-in, off by default, clearly described in the config page
  • no library-wide inventory ever sent in one request — the batch exists endpoint is capped at 100 items and a sweep is paced
  • each configured server multiplies this exposure, which the config page must say plainly; first-match resolution limits it, since later servers are only queried for what earlier ones lacked

9a. Federation and replication — UR-008

Servers can replicate manifests from each other, so a new instance can bootstrap from an existing one and independent communities need not each re-run the CV pipeline on the same films.

What makes this easy, and what makes it hard

Easy: a validated manifest is immutable and content-addressable. Its content is a fixed set of (TMDB person id, time windows) for a fixed (title, cut). Nothing about it changes after acceptance. Replication is therefore set reconciliation, not state synchronisation — there are no concurrent edits, no last-write-wins, no vector clocks, no merge conflicts. Two servers holding the same manifest hold byte-identical content.

Hard: the mutable state is exactly the part that must not replicate blindly. status, reports and cast_match_ratio encode a local operator's judgement and legal position. A server that pulls another's delisted flags as authoritative has outsourced its moderation; a server that pulls another's listed flags has outsourced its liability. §5a's guarantees are per-operator, and federation must not silently transfer them.

The design follows directly: replicate content, re-derive judgement.

Content addressing

Every manifest gets a content_id — a SHA-256 over its canonical form:

sha256(canonical_json({
  identity, cut, actors: [{tmdb_person_id, scenes}] sorted by person id
}))

Canonicalisation: keys sorted, no whitespace, and scene times quantised to whole centiseconds — round(t * 100) stored as an integer, not a rounded float. extraction metadata and all local state are excluded, so two servers that validated the same upload independently arrive at the same content_id.

audio_signature is excluded from content_id, deliberately. It is derived by decoding audio, so two servers running different FFmpeg or resampler versions could compute marginally different signatures for the same manifest — including it would produce different content_ids for identical content and silently break federation deduplication. The signature is replicated as an attribute of the manifest, not as part of its identity. A peer that already holds a manifest but lacks its signature may adopt the incoming one.

cut is therefore a single key — runtime_cs. video_hash was in the canonical form until it was withdrawn (§3), and it was removed from the form rather than retained as a vestigial null, on the same reasoning §2 applied to anneal_sec: a key naming a signal the format no longer has is actively misleading. Every content_id changed at that point, including for manifests that never carried a hash. Peers holding pre-withdrawal ids must re-derive them; there is no migration, because a content address is not a value that can be migrated — it is recomputed or it is wrong.

Quantising to integers rather than formatting floats is deliberate. Pipeline timings are derived by accumulating 1/fps, not measured, so they carry accumulated float error — real corpus values look like 8045.066666660665. Measured over 28972 scene values from the extraction corpus, 3-decimal rounding has a maximum error of 3.3e-4 and produces no boundary cases, so it is currently safe. But "currently safe" is luck: any value landing near a .0005 boundary would hash differently on two servers that computed it slightly differently, silently defeating deduplication.

Integer centiseconds remove the failure mode rather than dodging it — 10 ms is far below the precision any overlay can use (the pipeline samples at 1–10 fps), so nothing is lost. The client must canonicalise identically, and the canonicalisation routine should be shared code between server and client rather than reimplemented.

This gives deduplication for free: a pulled manifest whose content_id is already present is skipped without re-validation. It also makes "have you got this?" a cheap hash comparison rather than a content diff.

Replication protocol

Deliberately a pull-based feed, not push. Pull means a server chooses what it ingests and when; push would let any peer inject work into your validation queue, which is the same abuse surface as anonymous upload but with higher volume.

GET /federation/changes?since={cursor}&limit=1000

A monotonic, append-only change feed of locally-listed manifests.

{
  "cursor": "01HZ...",
  "server_id": "jray.example.org",
  "changes": [
    {
      "content_id": "sha256:9f2a…",
      "op": "add",
      "identity": { "type": "movie", "tmdb_id": "504172" },
      "cut": { "runtime_sec": 6420.5 },
      "actor_count": 17,
      "cast_match_ratio": 0.82,
      "origin": "jray.example.org",
      "seq": "01HZ..."
    }
  ]
}

Entries are metadata only — enough to decide whether to fetch, without transferring payloads. op is add or retract (see below). The cursor is opaque and monotonic; a peer resumes from its last cursor, making the feed resumable and idempotent.

GET /federation/manifests/{content_id}

Fetch full content by hash. The puller must verify that the returned content hashes to the requested content_id and reject it otherwise — this is what makes an intermediary or a misbehaving peer unable to substitute content.

POST /federation/have

Batch existence check by content_id (up to 1000), so a peer can diff its set against yours in one request before fetching anything.

Ingestion: re-derive, don't inherit

A pulled manifest is not trusted because a peer listed it. It enters the local pipeline as if freshly uploaded:

  1. Full §6 stage 1 and 2 validation — size caps, strict schema, bounds checks. A peer is not exempt from the checks that keep §5a's Threat 1 closed.
  2. Local TMDB cast cross-check (§6 stage 3), against this server's own TMDB cache and its own thresholds. The peer's cast_match_ratio is advisory only — useful for prioritising ingestion order, never a substitute.
  3. Local status assigned by this server's rules.

The peer's moderation decisions are recorded as signals, not verdicts:

Peer state Local effect
Peer lists it Eligible for ingestion; still fully re-validated
Peer retracts it (op: retract) Local copy flagged for review, not auto-delisted
Peer never had it No signal

The asymmetry is deliberate: a retraction is a warning worth acting on, while a listing is merely a nomination. Auto-delisting on a peer's retraction would hand any peer a remote delete primitive over your catalogue.

Exception — the abuse channel. One class of retraction should auto-delist: content withdrawn for legal reasons. A retract entry may carry reason: "abuse", and a peer explicitly configured as TrustAbuseRetractions: true will delist immediately and log it. This is opt-in per peer, and the intended configuration between operators who know each other. It exists because the alternative — a takedown propagating at the speed of manual review — is the wrong failure mode for that one case.

Origin and loop prevention

Each change carries origin, the server_id that first accepted the manifest, preserved across hops. A server ignores changes whose origin is itself, which prevents the trivial A→B→A loop. Because content is addressed by hash and ingestion is idempotent, longer cycles are harmless: the second arrival is a no-op deduplication.

origin is provenance, not authority — it does not confer trust, it just enables an operator to say "stop ingesting anything originating from X".

Peer configuration

Symmetrical with the plugin's server list (§9), and for the same reason — federation is trust-by-configuration, not trust-by-protocol:

Field Purpose
Url Peer base URL
Enabled Toggle without deleting
PullInterval Poll cadence, default hourly
TrustAbuseRetractions Auto-delist on legal retractions (default false)
IngestFilter Optional: only titles matching a filter (e.g. exclude adult-flagged)
MaxIngestPerHour Rate cap, so a peer cannot flood the validation queue
Advertise Whether to list this peering in the public directory (default false)

There is no automatic peering. Ever. A peer relationship is created only by an operator explicitly adding a URL. Nothing a remote server says, and no data returned from any endpoint, can cause a peering to be established, re-enabled, or widened. Automatic peering would let the network's trust properties be set by whoever joins, which is precisely what §5a avoids.

Federation is off by default. A server with no peers configured behaves exactly as specified in §1–§9.

Peer directory — publishing, not discovering

A server may publish the peers it has chosen, so an operator evaluating the network can see who is connected to whom. This is a human-facing directory, not a discovery mechanism.

The distinction is the whole point:

Peer directory (allowed) Auto-discovery (prohibited)
What it does Publishes a list a human can read Acts on a list a machine received
Who decides The operator, by hand The protocol
Failure mode Someone reads a stale list A hostile server injects itself into your trust set, transitively

GET /federation/peers

Returns peers this server has chosen to advertise:

{
  "server_id": "jray.example.org",
  "contact": "admin@example.org",
  "peers": [
    { "url": "https://jray.other.org", "name": "Other Community", "since": "2026-03-01" }
  ]
}

Rules that keep this a directory and not a discovery channel:

  • Advertising is per-peer opt-in on both sides. A peering appears here only if the local operator set Advertise: true and the remote operator consented to being listed. Peering with someone must not publish their existence against their wishes — for a small operator, being listed is an invitation to traffic and scrutiny they may not want.
  • The response is never ingested. The server does not parse it, store it, or act on it. It is rendered in the admin UI for a human, with entries as inert text and an explicit "add this peer" button that performs the same manual add as typing a URL. No one-click "add all".
  • Not transitive. A peer's peers are not fetched recursively. There is no crawl, so there is no network-wide topology to poison.
  • contact is for humans arranging a peering out of band, which is the intended workflow: operators talk, then each adds the other by hand.

Publishing the directory is itself optional (PublishPeerDirectory, default off). A server that would rather not disclose its topology simply doesn't.

Duplicate manifests across origins

Two servers may independently hold manifests for the same (title, cut) from different contributors — different content_id, same identity. This is already handled: §7 permits multiple manifests per cut and ranks by (cast_match_ratio, reports, gallery_scope, sample_fps) (§7). Federation just makes it more common. No deduplication beyond exact content_id match is attempted, because choosing between two plausible extractions is a ranking problem, not a merge problem.

Storage additions

peers          (id, url, name, enabled, pull_interval,
                trust_abuse_retractions, advertise, peered_since,
                last_cursor, last_pull_at, last_error)

-- manifests gains:
--   content_id  text unique   -- sha256 over canonical form
--   origin      text          -- server_id of first acceptance
--   ingested_from text NULL   -- peer id, NULL if uploaded directly

content_id carries a unique index and is the deduplication key on ingest.

What is deliberately not specified

  • No automatic peering. A published peer directory (above) is readable by humans; it is never acted on by software. No crawling, no transitive peering, no "trusted because a peer trusts them".
  • No consensus. There is no global agreement on what the catalogue contains. Each server's catalogue is its own; federation only makes it cheaper to fill.
  • No global identity. No shared contributor identity across servers. Tokens stay local, consistent with §5a — a contributor's standing on one server means nothing on another, and needs to mean nothing.
  • No deletion propagation beyond the opt-in abuse channel above.

10. Open questions

  1. Is runtime tier good enough as the default? ±2s will match most same-cut releases but will also match a different encode with identical runtime and a different logo trim. Leaning yes-with-caveat-in-UI.
  2. Should manifests be signed by the contributor? Adds provenance but also key management for hobbyist operators. Probably not for v1.
  3. Should the gallery be shareable too? Embeddings are far larger than manifests and are derived from copyrighted headshots; out of scope here, but worth a separate look — it would remove the biggest setup cost for new users.
  4. Federation is now specified in §9a. Open sub-questions:
    • Is re-running the TMDB cast check on every ingested manifest affordable? Bulk-ingesting a large peer catalogue means a TMDB lookup per distinct title. The 24h credits cache and the fact that lookups are per title (not per manifest) should make it fine, but a bootstrap of tens of thousands of titles needs a throttled backfill mode rather than the normal upload path.
    • Should a fresh server be allowed to trust a peer's cast_match_ratio during initial bootstrap only? It would make standing up a mirror far cheaper, at the cost of the guarantee in §9a. Leaning no.
    • Float canonicalisation stability — resolved. Checked against 28972 scene values from the corpus: 3dp rounding is safe today (max error 3.3e-4, no boundary cases) but fragile, since pipeline timings accumulate float error from summing 1/fps. §9a now quantises to integer centiseconds, which removes the failure mode rather than relying on luck.
  5. Is 0.6 the right cast-match threshold? Still a guess, but now a testable one: the extraction repo has 331 real output files. Running the §6 stage-3 check over them against TMDB would yield the true distribution of honest-upload match ratios and let the threshold be set at, say, the 1st percentile rather than by intuition. This is the single cheapest way to de-risk UR-003 and UR-005 and should happen before launch. Note that the corpus is heavily TV-weighted, so movie and episode thresholds may need to differ.
  6. Discarding the uploaded name string (§5a) depends on TMDB person resolution being reliable. If too many legitimate actors fail to resolve, manifests lose actors silently. The corpus run in (5) measures this too. If resolution proves lossy, the fallback is to store names but restricted to the closed character class — weaker, but still not a usable payload channel.
  7. Anonymous existence checks are a title-availability oracle. Rate limits blunt this but do not remove it. Requiring a token for exists would close it at the cost of making read-only use non-anonymous. Left open.
  8. Adult-content classification for the §5a category guard relies on TMDB's adult flag, whose coverage for performers is less consistent than for titles. The guard may need a supplementary signal.