Files
jray-project/README.md
T
dtourolleandClaude Opus 5 17106f3370 Adopt the config-driven extractor; project overview README
Replaces the copy taken earlier with the newer version from
scene-actor-extraction's traceability-tooling branch, which had moved on: it
takes per-repo settings from a traceability.toml rather than the CLI flags
added here, validates them, and names languages ("rust") rather than making
each repo spell out extensions. That is the better design, so the flags go and
this becomes the single source.

Two fixes on top:

- Config discovery searched from the working directory only, so --root pointed
  at another tree found no traceability.toml and failed with
  "requirement_types is empty" while a perfectly good config sat in the
  directory named. That breaks both intended callers: CI passing --root, and a
  wrapper running the vendored copy. Discovery now starts from --root.
- The test suite had not been migrated with the Config refactor and failed on
  the branch as well as here. All 53 now pass: entry points take a Config,
  ci_executable moved to the Register which owns tier policy, fixtures write a
  real traceability.toml so config discovery is exercised rather than bypassed,
  and the live-register tests take LIVE_REGISTER from the environment since the
  project home holds no component register of its own.

The README becomes a project overview rather than a table of contents: what the
problem is, why a paused-frame answer is the wrong question, why gallery data
never leaves the instance, and why the manifest server can hold no binary.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 18:38:59 +02:00

191 lines
7.9 KiB
Markdown

# JRay
**A self-hosted alternative to Amazon X-Ray.** While watching, the viewer can
see who is on screen — for their own library, on their own hardware, with nobody
else learning what they own.
This is the project home. It owns the [system specification](SPEC.md), the
requirements that span more than one component, and the tooling they share. The
code lives in three repositories, linked below.
---
## The problem
You are watching a film. Someone appears and you know you have seen them
before — but pausing to search breaks the film, and the answer is rarely worth
the interruption. Amazon X-Ray solves this well, and only for Amazon's catalogue.
Doing the same for a personal library is harder than it looks:
- **Face recognition on a paused frame answers the wrong question.** In dialogue
the camera is usually on whoever is *not* speaking, so a per-frame answer
reports the other actor absent. X-Ray credits a whole scene's cast for the
scene's duration, and that is the more useful question (SR-002).
- **The compute is real.** Detecting, embedding and tracking every face in a
feature takes GPU time. Doing it once per viewer, for the same film, is waste.
- **The obvious fix leaks.** A service that identifies your library must be told
what your library contains, which recreates the thing being replaced (PR-005).
JRay's answer: extract presence data locally, share the *timings* rather than
the media or the faces, and make the shared artefact structurally incapable of
carrying anything else.
---
## How it fits together
```
media file ──► extraction ──► truth file (sidecar or pushed) ──► plugin ──► player overlay
▲ │
│ ▼
gallery Jmanifest ──► public server ──► other instances
(Jellyfin + TMDB)
```
| Repository | Role | Language |
|---|---|---|
| [`scene-actor-extraction`](https://gitea.tourolle.paris/dtourolle/scene-actor-extraction) | Derives presence data from a media file | C++ / Python |
| [`jRay`](https://gitea.tourolle.paris/dtourolle/jRay) | Jellyfin plugin: surfaces it in the player, owns the truth-file format | C# |
| [`JRay-public-server`](https://gitea.tourolle.paris/dtourolle/JRay-public-server) | Exchanges presence data between instances | Rust |
They ship independently — the plugin to Jellyfin's catalogue, the server as a
static binary, the pipeline to a GPU host — which is why they are separate
repositories rather than one monorepo.
**Two data axes, deliberately separate.** *Presence data* flows outward and is
shareable: it is timings against public TMDB identifiers. *Gallery data* — the
actor reference faces — is built locally from your own Jellyfin instance plus
TMDB, and never leaves the machine (SR-005). Not by export, not by opt-in, not
at all: the capability is what creates the exposure, so it does not exist.
**The manifest server holds no binary content.** An accepted manifest contains
bounded numbers, regex-constrained identifiers, and references to TMDB persons.
No images, no embeddings, no free-form strings, no extension points (SR-004).
That is what makes it safe for a volunteer to operate an instance, and it is a
property preserved by prohibition — any proposal to ship blobs through it is a
proposal to delete it.
---
## Setting up
Clone this repository, then the components beside it:
```sh
git clone git@gitea.tourolle.paris:dtourolle/jray-project.git
cd jray-project
git clone git@gitea.tourolle.paris:dtourolle/scene-actor-extraction.git
git clone git@gitea.tourolle.paris:dtourolle/jRay.git
git clone git@gitea.tourolle.paris:dtourolle/JRay-public-server.git
```
Each component has its own README with build instructions. They are
`.gitignore`d here, so they sit beside the system spec without this repository
trying to track them.
### Why the components are not submodules
A submodule pins a commit. With feature branches and worktrees in flight across
the components, every component commit would leave this repository's pointer
stale and its `git status` dirty until someone committed a bump — churn that
buys nothing, since the components are developed together in one directory.
The dependency runs the other way: **each component pulls *this* repository in**
as a submodule, for the system spec and the shared tooling, both of which change
rarely. That is the asymmetry submodules suit.
---
## Where to start reading
| Doc | Owns |
|---|---|
| [`SPEC.md`](SPEC.md) | **System requirements** — `PR-nnn` project goals, `SR-nnn` cross-component contracts |
| [`CLAUDE.md`](CLAUDE.md) | Working notes, and the invariants that must not be violated silently |
| Each repo's `SPEC.md` | That component's software requirements |
| Each repo's `docs/requirements.md` | Its stable requirement IDs, status, and verification plan |
Read the system spec first. Every component requirement traces up to an
`SR-nnn`, and every `SR-nnn` to a `PR-nnn`, so the chain explains *why* a given
piece of code exists — and makes it visible when something exists for no stated
reason.
---
## Requirement traceability
```
PR-nnn project requirement (SPEC.md §1) — why the system exists
└─ SR-nnn system requirement (SPEC.md §3) — what spans components
└─ component requirement (each repo's docs/requirements.md)
└─ TRACES tag (source)
```
Tag the code that *satisfies* a requirement — the unit that decides, not every
helper it calls:
```rust
/// TRACES: UR-003, UR-011 | SR-004
pub fn validate_manifest(m: Jmanifest) -> VResult<ValidManifest> { … }
```
A pipe separates requirement *types*; a comma separates IDs within a type.
Tests carry tags too (`UT-nnn`, `IT-nnn`), which is what shows a requirement is
*verified* rather than merely implemented. A deliberate departure from an
invariant is tagged `EXCEPTION:` with its reason — an untagged one is a defect.
### Running the gate
The tooling lives in [`scripts/traceability/`](scripts/traceability/) here and
is vendored into each component as a submodule, so there is **one
implementation**. Each component declares its own taxonomy in a
`traceability.toml` at its root:
```toml
requirement_types = ["UR", "DR"]
languages = ["rust"]
source_roots = ["src", "tests"]
```
```sh
scripts/traceability-gate.sh # from any component
```
It reports coverage, orphan tags (an ID no register defines), untraced
requirements, and requirements verifiable only on hardware CI lacks.
**Two rules inherited from JellyTau, both learned the hard way:**
- **Denominators are read from the register at run time, never hardcoded.** A
gate that divided by a frozen literal reported *158% coverage* for months
while the requirement count grew, so its threshold could never trip. A gate
that cannot fail is worse than no gate, because it is trusted.
- **Coverage above 100% is a hard failure**, not a pass. It means the
computation is broken, and it is the signal that catches the above at once.
The same reasoning is why a requirement whose only evidence is a test that never
runs is reported as *tagged but unexecuted*, never counted as covered.
---
## Status
| Component | State |
|---|---|
| `scene-actor-extraction` | Pipeline redesign in progress — presence follows track extent (AR-012), replacing per-frame recognition |
| `jRay` | Truth-file serving and overlay working; manifest-sharing configuration added, fetch path outstanding |
| `JRay-public-server` | Core implemented: 23/32 requirements traced, 189 tests. Audio-tier matching and federation deferred by design |
**One schema bump is pending across all three repos** (SR-003). It removes
`anneal_sec`, adds `extinction_sec` and `gallery_scope`, gives each window its
belief and identification route, and adds the audio signature. Breaking changes
are batched, so these ship together rather than piecemeal.
---
## Licence
GPLv3, matching the Jellyfin plugin it serves.