Files
dtourolleandClaude Opus 5 413785f8be docs: PR-005 has software rows now, and the rollup should say so
The matrix section still claimed PR-005 "has no software row at all" and could
be verified only by prohibition. jRay's register has carried four rows against
it since the schema-v2 landing: JR-038 (Done), JR-034 and JR-039 (both High,
both T1, both still Planned), and JR-040 (T4). jRay is the component that
actually performs egress, so that is where the goal became verifiable rather
than merely preserved.

Three of the four are untagged in the matrix. That is unbuilt work, not a
broken chain, and saying so here keeps the rollup from reading as an
inconsistency. The structural guarantees still hold PR-005 from the other
side; nothing about SR-004 or GR-005 changes.

jRay/SPEC.md already stated this in the past tense. The vendored copy under
jRay/scripts/vendor/jray-project is a nested checkout of this repo, so it
follows on the next vendor bump rather than needing its own edit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

TRACES: JR-034, JR-039 | PR-005
2026-07-31 17:11:54 +02:00

485 lines
24 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# JRay — system specification
Status: **draft**.
This is the **system** spec: it owns the requirements that span more than one
component, and the contracts between them. Each component repo carries a
*software* spec whose job is to implement this one.
| Component | Software spec | Role |
|---|---|---|
| `scene-actor-extraction` | [`docs/SPEC.md`](scene-actor-extraction/docs/SPEC.md) | Produces presence data from media |
| `jRay` (Jellyfin plugin) | [`SPEC.md`](jRay/SPEC.md) | Consumes and displays it; owns the truth-file format |
| `JRay-public-server` | [`SPEC.md`](JRay-public-server/SPEC.md) | Exchanges presence data between instances |
A requirement belongs here if changing it requires changing more than one repo.
Everything else belongs in a component spec.
---
## 1. What the system does
**Project goal: a self-hosted alternative to Amazon X-Ray.** While watching, the
viewer can see who is on screen — for their own library, on their own hardware,
with nobody else learning what they own.
Everything below traces to that. The project requirements are the top of the
traceability chain (§6):
| ID | Project requirement | Why it is not negotiable |
|---|---|---|
| **PR-001** | Show which actors are on screen at a given moment, in the player | The product. Without it there is nothing. |
| **PR-002** | Answer at **scene** granularity, as X-Ray does | The thing being replaced works this way, and it is the more useful question (SR-002) |
| **PR-003** | Fully automatic — no per-title manual work | Manual labelling does not scale past a handful of titles; a library is thousands |
| **PR-004** | Run on the user's own hardware, against their own library | Self-hosted is the point; a cloud dependency recreates what is being replaced |
| **PR-005** | Leak nothing about what the user owns | A service that must be told which titles you hold is not an alternative to a service that already knows |
| **PR-006** | Compute once per *cut*, not once per user | Extraction is expensive; identical cuts yield identical timings, so repeating the work per-user is pure waste |
Three capabilities follow:
1. **Extract** — derive presence data from a media file (`scene-actor-extraction`).
Serves PR-001, PR-002, PR-003, PR-004.
2. **Serve** — surface it in the player at the right moment (`jRay`). Serves PR-001.
3. **Share** — let one user's extraction spare every other user the same compute
(`JRay-public-server`). Serves PR-006, constrained by PR-005.
---
## 2. Data flow
```
media file ──► extraction ──► truth file (sidecar or pushed) ──► plugin ──► player overlay
▲ │
│ ▼
gallery Jmanifest ──► public server ──► other instances
(Jellyfin + TMDB)
```
Two independent axes, deliberately:
- **Presence data** flows outward and is shareable — it is a set of timings
against public identifiers.
- **Gallery data** (actor reference faces) is built locally from the user's own
Jellyfin instance plus TMDB, and does **not** flow outward. See SR-005.
---
## 3. System requirements
### SR-001 — Identity is expressed in public identifiers
Every actor reference crossing a component boundary is a `tmdb_id` / `imdb_id`
(and, locally, a Jellyfin Person GUID). Never a name alone, never a local index.
Rationale: names are ambiguous and unstable; local indices are meaningless
outside the process that made them. Public identifiers also make the public
server's abuse model tractable — it stores references to entities that already
exist in TMDB, not attacker-authored text (`JRay-public-server/SPEC.md` §5a).
### SR-002 — Presence is scene-scoped, not instantaneous
**The question the system answers is "is this actor in this scene?", not "is this
actor's face visible in this frame?"**
This is the defining semantic of the whole system, so it is stated here rather
than in any component spec.
It is set by the ground truth we emulate. Amazon X-Ray ships one cast list per
scene span — there is no per-frame or per-second annotation anywhere in its data
(`scene-actor-extraction/docs/methodology.md`). If X-Ray credits five actors to a
30-second scene, all five are present for all 30 seconds, including the seconds
where only one is on camera.
Consequences that bind every component:
- **An actor who turns away, is occluded, or is off-camera while the shot cuts to
whoever they are speaking to is still present.** Reporting them absent is not a
stricter answer, it is an answer to a different question.
- A window is a claim about **scene membership**, not a recognition event. A
consumer must never interpret window boundaries as "the face was detected here".
- Gaps shorter than the re-acquisition timeout are absorbed *into* a window and
claimed as presence. Two separate windows therefore mean a genuine departure and
return, not merely a break in detection.
- Windows may be numerous; consumers must not assume a handful of long ones.
**Why this is the differentiator.** A plugin that decodes a single frame on pause
and runs face detection on it answers the instantaneous question. Against
scene-level truth it produces a false negative every time the camera is on someone
else — which, in dialogue, is most of the time. It also cannot know who *just*
left or is about to speak. Precomputed scene-scoped presence is a different and
strictly more useful product, and the extra cost is paid once per title rather
than on every pause.
**A window ends at the last sighting, never after it.** This is the rule that
prevents over-claiming, and the previous design broke it: `extinction_sec` kept
an actor *active* for 57 s past their last detection, so presence ran on into
whatever followed — most visibly the closing credits, inheriting the final
scene's cast (the Downton Abbey recall collapse, `lvface-deep-dive.md`).
Ending each window at the last frame the face was actually seen removes that
class of error outright. Gaps *between* two sightings are absorbed (an actor who
turns away is still in the scene); time *after* the final sighting never is.
Cuts and scene boundaries are **association hints, not presence events**. Both
say the same thing to the tracker — spatial continuity is broken, so associate on
embedding similarity rather than position — and neither closes a window. An actor
who genuinely continues across a boundary is kept; one who does not reappear
times out and closes at their last sighting.
### SR-003 — Schema changes are coordinated across all three repos
The truth file and the Jmanifest are consumed by components that ship
independently. Therefore:
- `schema_version` is a **system-level** version, incremented once per breaking
change, and every repo's spec references the same number.
- Breaking changes are **batched**. Two pending breaks ship as one bump.
- A consumer encountering an unknown `schema_version` refuses or warns; it never
guesses.
**Currently pending — these ship together as one bump:**
| Change | Origin | Effect |
|---|---|---|
| Remove `anneal_sec` | SR-002 / A6 | Field describes a mechanism that no longer exists |
| Add `extinction_sec` | A6 | Its successor: the parameter that shapes window extent |
| Add `gallery_scope` | C1 | `global` / `limited` — the strongest quality signal between competing manifests |
| Per-window belief | A6 / A9 | Windows carry the posterior that justified them, plus identification route |
| Add audio signature | C2 | Self-identifying truth files |
Removing `anneal_sec` rather than retaining it as a vestigial `0` is deliberate:
a field naming a mechanism the pipeline no longer has is actively misleading to
anyone reading a manifest, and would outlive everyone who remembers why it is
zero.
**Per-window belief changes the shape of `scenes`**, which is currently a list of
float pairs. It becomes a list of objects carrying the interval plus its
posterior and route. Note this **does not** weaken SR-004: the added fields are a
bounded float and a small enumerated string, so an accepted manifest still
contains only numbers and closed-vocabulary values. No free-form channel is
opened.
Owner of the bump: to be assigned (see §4).
### SR-004 — The public server never stores binary content
The server's abuse defence is structural, not procedural: **there is nowhere to
put a payload**. An accepted manifest contains only bounded numbers, constrained
identifiers, and — at most — names that must resolve to real TMDB persons
(`JRay-public-server/SPEC.md` §5a, Threat 1). No binary, no images, no URLs, no
base64, no extension points.
**This is load-bearing and must not be eroded.** It is what makes it safe for a
volunteer to operate an instance: there is no CSAM risk and no payload-smuggling
risk because no field can carry either. Any proposal to ship images, embeddings,
or opaque blobs through the manifest server is a proposal to delete that
property, and must be treated as such rather than as an incremental feature.
**The audio signature is the one encoded field, and it is not a precedent.** It
is base64, so the flat reading of "no base64" needs qualifying — otherwise
someone later cites SR-004 to argue the signature out, or cites the signature to
argue images in. Neither is right. What makes it acceptable is that it is
**fully constrained**, not that it is small:
- fixed length (~1290 bytes) derived from a fixed 120 s window — not
caller-chosen;
- fixed alphabet, validated on parse;
- every byte is a peak-bin index plus a 2-bit energy class, so the value space is
enumerable and a payload cannot be smuggled in it;
- it is verifiable — a server can recompute it from the same input and compare.
The test for any future encoded field is that list, not its size. A field whose
length or content is attacker-chosen fails it regardless of how small it is.
### SR-005 — Gallery data never leaves the instance
Actor reference faces are built from the user's own Jellyfin instance and TMDB
profile images. **They are never shared, exported, or transmitted — not between
instances, not to a server, and not by a user who wants to.** There is no export
path, deliberately.
This is the strict form on purpose. A user-initiated export was considered and
rejected: it moves a legal judgement onto the user while the capability, and
therefore the exposure, belongs to the software.
Three reasons, each independently sufficient:
- **Legal risk.** TMDB profile images carry their own licensing; face crops
extracted from a film are derivative works of a copyrighted title; and faces
are personal data and likenesses. Recognising one's own library locally is a
materially different act from distributing crops or their derivatives to third
parties. Not shipping them avoids the question rather than answering it.
- **SR-004.** Images cannot cross the public server without destroying its
structural abuse defence.
- **PR-005.** A mature gallery is a direct fingerprint of a library. One that has
learned an actor's profile views across forty of *your* films encodes which
forty films those were.
That third reason **strengthened** with D4. When the gallery held only downloaded
headshots it was largely public data; now it accumulates harvested embeddings
derived from the user's own media, and later human-confirmed associations. The
artifact became more valuable and more personal at the same time.
Two consequences, both accepted:
- **Gallery quality is per-instance.** Two users with the same film may get
different results. Accepted — the alternative is the exposure above.
- **Human associations do not compound across users.** Everyone labels their own
unknowns. Accepted for the same reason.
---
## 4. Human-in-the-loop actor association
*Proposed; not yet designed in any component spec.*
The pipeline necessarily produces **unidentified tracks** — a face that is
genuinely someone, sustained across many frames, that the gallery cannot name.
Today they are discarded. They are the system's best signal about its own blind
spots, and a user watching the film usually knows exactly who the person is.
**The capability:** surface unidentified tracks in the plugin, let the user
associate one with an actor, and feed that association back into the local
gallery so subsequent extractions recognise them.
This is valuable precisely where automatic extraction is weakest: character
actors, ageing across a career, heavy makeup, and cast members TMDB has no
usable headshot for.
### What each component contributes
| Component | Contribution |
|---|---|
| `scene-actor-extraction` | Retain unidentified tracks, **clustered into one entity per unknown person** (A7.4): embeddings, metadata, and **context crops** — a wider region than the 112×112 aligned face, so a human can actually recognise the person and the shot they are in |
| `jRay` | Review UI: show the cluster's crops, let the user pick an actor from the title's cast (or search TMDB), record the association |
| `scene-actor-extraction` | Ingest confirmed associations into the local gallery as additional reference embeddings for that `tmdb_id` |
**The unit of review is a person, not a track.** Because unknown tracks are
clustered before they reach a human, the question is "who is this person, who
appears in these twelve places?" rather than twelve disconnected questions about
twelve faces. One answer resolves the whole cluster. This is what makes the
feature tractable to use rather than tedious — and the clustering is already
required for recognition (A7.4), so the review UI inherits it for free.
The 112×112 aligned crop is optimised for ArcFace, not for humans — tightly
cropped, geometrically normalised, often unrecognisable out of context. A
separate, larger crop is required for the review UI. Storage is bounded by
retaining a handful of representative frames per track, not every frame.
### Where the gallery lives — settled: local only
The appeal of moving gallery ownership to a server is obvious: one user's
association would help everyone, and the gallery would improve monotonically
across the community rather than per-instance.
**It is nonetheless ruled out.** SR-005 is strict, on legal grounds, and that
disposes of both the shared-service and the user-initiated-export variants.
Two further reasons it would not have worked even setting rights aside:
- **SR-004 blocks the manifest server specifically.** Shipping crops destroys its
structural abuse defence, and embeddings are not a loophole: a 512-float array
is a 2 KB opaque binary field, exactly the extension point the design
deliberately lacks. Unit-norm constraint does not meaningfully reduce its
payload capacity.
- **A gallery-sharing service would be a different service** — accepting binary
content, therefore needing moderation, provenance, and an operator willing to
carry that liability. Bolting it onto the manifest server would forfeit the
property that makes the manifest server safe to run.
**So: associations stay on the instance that made them.** No new trust surface,
no rights exposure, no sharing benefit. This is not a compromise position — the
recognition benefit is almost entirely captured locally anyway, since the
associations that matter to a user are for the titles that user owns.
---
## 5. Open system questions
1. **Who drives the SR-003 bump?** Three repos must move together. **The
transition half is settled: flag day** — the plugin accepts `schema_version: 2`
and rejects everything else, on every path, with no dual-accept period
(`JR-003`). All three components are pre-release, and a v1 read path would be
the one nobody exercises, so it is the one that would rot while being carried
through every subsequent change to the reader.
The consequence is accepted and must be stated to users rather than
discovered: **existing v1 sidecar files stop being read** and stay dark until
the library is re-extracted. The plugin logs this per item, naming the file
and the version found, so a stale item is not mistaken for an un-extracted
one.
Still open: **who sequences the three merges**, and whether the server should
accept `jmanifest_version: 1` for a period after the plugin stops — its
exposure is different, since it holds other people's contributions rather
than one user's files.
2. **Should unidentified presence be published?** An unowned track is *someone*
on screen. Emitting it as anonymous presence ("a face is here, unnamed") would
let the plugin show "unidentified person" rather than nothing — and would give
§4's review UI its work queue directly from the truth file. It also enlarges
the format and invites consumers to display something unhelpful. Undecided.
3. **Episode-level gallery scope.** A series' recurring cast is stable across
episodes; re-deriving the gallery per episode is wasteful, and associations
made while watching episode 1 should apply to episode 12. Not yet specified
anywhere.
---
## 6. Traceability
Every requirement must be traceable **up** to the project goal and **down** to the
code that implements it. A requirement nothing traces to is either unnecessary or
unimplemented, and both are worth knowing.
### The chain
```
PR-nnn project requirement (this document, §1) — why the system exists
└─ SR-nnn system requirement (this document, §3) — what spans components
└─ component requirement (component registers) — what one repo does
├─ TRACES tag (source) — where it lives
└─ TRACES trailer (commit message) — when and why it changed
```
The commit trailer is the last link and the only one that carries *history*:
`git log --grep=AR-012` reconstructs everything that ever happened to a
requirement, which no other artifact provides. Same token and syntax as the code
tag, so one grep pattern serves both.
### ID scheme
Zero-padded three digits throughout. **IDs are permanent**: a withdrawn
requirement is marked `Withdrawn` and its number never reused, because
renumbering is exactly what produces orphan tags.
| Prefix | Scope | Register |
|---|---|---|
| `PR-nnn` | Project | This document §1 |
| `SR-nnn` | System | This document §3 |
| `AR-nnn` | Extraction — algorithm | [`scene-actor-extraction/docs/requirements.md`](scene-actor-extraction/docs/requirements.md) |
| `DP-nnn` | Extraction — deployment | ″ |
| `IR-nnn` | Extraction — integration | ″ |
| `GR-nnn` | Extraction — gallery | ″ |
| `VR-nnn` | Extraction — validation | ″ |
| `JR-nnn` | Plugin — all of it | [`jRay/docs/requirements.md`](jRay/docs/requirements.md) |
| `UR-nnn` | Public server — all of it | `JRay-public-server/docs/requirements.md` |
| `UT-nnn` / `IT-nnn` | Tests | Per-component register |
Each component keeps a `requirements.md` register — the authoritative ID list,
with prose in its `SPEC.md`. **The CI gate reads its denominators from the
register at run time.**
**`JR` is flat where extraction is themed.** Extraction splits by concern because
its five areas have genuinely different verification strategies and audiences;
the plugin is one deployable with one audience, so a flat namespace with
thematic *section headings* carries the same grouping without a prefix that can
go stale. The server's `UR` is flat for the same reason.
**Two register gaps remain, recorded rather than omitted:**
- `JRay-public-server/docs/requirements.md` **does not exist**, though its
`SPEC.md` §0 cites it as authoritative and its `UR-nnn` IDs are already
regularised. Its matrix rows stay empty until it does.
- `UT-nnn`/`IT-nnn` are per-component, so `UT-001` will exist in more than one
repo. Fine while the gate runs per repo; ambiguous the moment a rollup spans
them.
### Code annotation
**Format follows the house convention already in use in
[`JellyTau`](../JellyTau/docs/traces-quick-ref.md)** — adopt its tooling rather
than inventing a parallel scheme:
```
// TRACES: UR-001, UR-002 | DR-003
```
Pipe separates requirement *types*; comma separates IDs within a type.
```cpp
/// TRACES: AR-006, AR-009 | SR-002
struct TrackRegistry { … };
```
```python
def build_gallery(...):
"""Build a gallery from Jellyfin plus TMDB fallback.
TRACES: GR-001 | SR-005
"""
```
Rules:
- The tag names **what the code satisfies**, not where it lives. One requirement
may be tagged in several places; one place may satisfy several requirements.
- Tag the unit that *decides*, not every helper it calls. A tag on every function
is noise and rots faster than it helps.
- **Tests carry TRACES too** (`UT-nnn`, `IT-nnn`), which is what demonstrates a
requirement is *verified* rather than merely implemented.
- **A deliberate exception to an invariant is tagged** —
`EXCEPTION: AR-009 raw cosine used here because <reason>`. Per
[`CLAUDE.md`](CLAUDE.md), a bare cosine with no recorded exception is a defect.
### Verification
**Requirement: a CI gate that checks the chain.** Hand-maintained traceability
rots within weeks; only the automated check makes it worth having. The matrix is
**generated, never hand-written**.
It must report:
- **Coverage** — requirements with at least one TRACES tag, over requirements
defined.
- **Orphan tags** — a TRACES naming an ID that no spec defines. This is what
renumbering produces, and what a typo produces.
- **Untraced requirements** — a software requirement citing no `SR-nnn`, or an
`SR-nnn` citing no `PR-nnn`. Catches scope creep: work serving no stated goal.
- **Unserved goals** — a `PR-nnn` or `SR-nnn` nothing traces up to, which is how
a goal quietly stops being pursued.
**CI has no GPU** — an Intel N100 board. A requirement verifiable only on GPU
hardware must be reported as *tagged but unexecuted*, never counted as covered.
A gate that treats "has a test that never runs" as passing is the same failure
mode as the 158% coverage bug below: it reports success it cannot substantiate.
Component registers carry the per-requirement verification tier.
**Two rules inherited from JellyTau's gate, both learned the hard way** (see
`JellyTau/docs/specs/traceability-gate-repair.md`):
1. **Denominators are computed from the spec at run time — never hardcoded.**
JellyTau's gate divided by frozen literals while the requirement count grew to
211, so it reported *158% coverage* and the threshold could never trip. A gate
that cannot fail is worse than no gate, because it is trusted.
2. **Coverage above 100% is a hard failure**, not a pass. It means the
computation is broken, and it is the signal that would have caught (1)
immediately.
### Matrix
**Generated, not hand-written** — `docs/traceability.md` in each component,
produced by the extract-traces tool and timestamped. The per-requirement parent
links live in each register's `Traces to` column; the tool rolls them up.
Porting from JellyTau: `scripts/extract-traces.ts` and
`scripts/generate-traceability-matrix.sh` need a C++/Python scanner added
alongside the existing TypeScript/Rust/Svelte one, and the workflow adapting.
The gate logic transfers unchanged.
Two facts the rollup should surface, noted now because they are easy to
misread as gaps:
- **`DP-*` and `VR-*` trace to no system requirement.** Deployment and parameter
studies are single-repo concerns serving `PR-004` and `PR-002` directly. This
is correct.
- **`PR-005` (leak nothing) will look thinly covered, and that is the true
picture.** It had no software row anywhere; jRay is the component that
actually performs egress, so four rows now carry it — `JR-038` (exchange
off by default, Done), `JR-034` (contribution strips `movie` and
`jellyfin_id`, posts only to contribute-enabled servers) and `JR-039`
(batch `exists` capped at 100, sweeps paced), both High and T1, and
`JR-040` (the config page states the per-server exposure, T4). Three are
untagged because they are unbuilt, not because the chain is broken. The
structural guarantees — `SR-004` and `GR-005` — still hold the goal from
the other side, and it still dies the moment either prohibition is relaxed.