Initial implementation: core vertical slice
CI / fmt, clippy, test (push) Failing after 2m46s
CI / static musl binary (push) Has been skipped
CI / advisories and licences (push) Successful in 4m22s

Implements the core of SPEC.md — the manifest exchange, less audio-tier
matching (§3) and federation (§9a), both of which the spec sequences as
later work.

- §2 Jmanifest format and series bundles
- §3 cut matching: exact / runtime / loose tiers
- §4 API, less POST /manifests/search
- §5 rate limiting; §5a trust model, anonymous bearer tokens
- §6 upload validation, all four stages
- §7 relational storage, no JSON blob on the write path
- §8 Rust + Axum + SQLite, single serialized writer, in-process job queue
- §9a content addressing, computed on upload

Reconciled against the system spec:

- anneal_sec removed, withdrawn upstream by AR-012/AR-013. Presence follows
  track extent, so a track survives its own gaps and there is nothing to
  anneal. Its successor extinction_sec and the new gallery_scope are accepted
  and stored; scope enters the §7 ranking. A manifest still carrying
  anneal_sec is a hard 400, not silently ignored — it came from a pipeline
  whose window semantics differ from what this server assumes.
- Audio signature: media under 120 s now emits no signature at all, matching
  scene-actor-extraction IR-007. The earlier §3 draft allowed a shortened
  window under 150 s, which was the weaker rule — a caller-varying length is
  the property SR-004 forbids.
- UR IDs regularised to UR-nnn; docs/requirements.md registers 32
  requirements, each tracing to an SR-nnn or PR-nnn.

189 tests: unit, end-to-end through the real router, and an injection suite
covering SQL, JSON, header and Unicode payloads. Writing that suite found two
real gaps, both fixed here: compatibility homoglyphs passed the §5a character
class, and a one-frame audio signature was accepted on a feature-length item.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
2026-07-30 18:14:02 +02:00
co-authored by Claude Opus 5
commit a848750a65
38 changed files with 13014 additions and 0 deletions
+187
View File
@@ -0,0 +1,187 @@
//! `GET /manifests/exists` and its batch form — UR-1.
//!
//! Deliberately a *separate, cheaper* endpoint from the fetch: it answers
//! "should I bother?" for a whole library sweep without transferring payloads,
//! and it is the endpoint a scheduled task will hammer. It is also the most
//! abuse-prone surface, since it doubles as an oracle for "does the community
//! have this title" — so it is rate-limited harder than the fetches and returns
//! no manifest content (§0).
use axum::extract::{Query, State};
use axum::http::HeaderMap;
use axum::response::{IntoResponse, Response};
use axum::Json;
use serde::{Deserialize, Serialize};
use super::LookupParams;
use crate::db::repo;
use crate::error::{ApiError, ApiResult};
use crate::matching::{self, StoredCut};
use crate::model::{IdentityType, MatchTier};
use crate::ratelimit::Surface;
use crate::state::{with_quota_headers, AppState};
/// §4: no manifest content, just availability and tier.
#[derive(Debug, Clone, Serialize)]
pub struct ExistsResponse {
pub exists: bool,
#[serde(skip_serializing_if = "Option::is_none")]
pub r#match: Option<&'static str>,
#[serde(skip_serializing_if = "Option::is_none")]
pub manifest_id: Option<String>,
#[serde(skip_serializing_if = "Option::is_none")]
pub actor_count: Option<i64>,
}
impl ExistsResponse {
fn absent() -> Self {
Self { exists: false, r#match: None, manifest_id: None, actor_count: None }
}
}
/// §4 batch form: up to 100 items.
///
/// Exists specifically so the §5 rate limit can be generous per *request* while
/// staying strict per *item*, and so a 2000-item library sweep is 20 requests
/// rather than 2000.
pub const MAX_BATCH_ITEMS: usize = 100;
#[derive(Debug, Deserialize)]
#[serde(deny_unknown_fields)]
pub struct BatchRequest {
pub items: Vec<LookupParams>,
}
#[derive(Debug, Serialize)]
pub struct BatchResponse {
/// Positional, matching the request order (§4).
pub results: Vec<ExistsResponse>,
}
pub async fn exists(
State(state): State<AppState>,
peer: crate::state::PeerIp,
headers: HeaderMap,
Query(params): Query<LookupParams>,
) -> ApiResult<Response> {
let ip = state.client_ip(&headers, peer.0);
let quota = state.check_limit(&ip, Surface::ExistsSingle)?;
let body = lookup_one(&state, &params).await?;
Ok(with_quota_headers(Json(body).into_response(), quota))
}
pub async fn exists_batch(
State(state): State<AppState>,
peer: crate::state::PeerIp,
headers: HeaderMap,
super::json::Json(req): super::json::Json<BatchRequest>,
) -> ApiResult<Response> {
if req.items.len() > MAX_BATCH_ITEMS {
return Err(ApiError::BadRequest(format!(
"items: at most {MAX_BATCH_ITEMS} per request, got {}",
req.items.len()
)));
}
if req.items.is_empty() {
return Err(ApiError::BadRequest("items: must not be empty".into()));
}
let ip = state.client_ip(&headers, peer.0);
let quota = state.check_limit(&ip, Surface::ExistsBatch)?;
let mut results = Vec::with_capacity(req.items.len());
for item in &req.items {
// A malformed item yields "absent" rather than failing the whole batch —
// a sweep of 100 items should not be lost to one bad entry.
results.push(lookup_one(&state, item).await.unwrap_or_else(|_| ExistsResponse::absent()));
}
Ok(with_quota_headers(Json(BatchResponse { results }).into_response(), quota))
}
async fn lookup_one(state: &AppState, params: &LookupParams) -> ApiResult<ExistsResponse> {
let Some((kind, tmdb_id, imdb_id)) = resolve_kind(params) else {
return Err(ApiError::BadRequest(
"requires tmdb_id/imdb_id, or series_tmdb_id with season and episode".into(),
));
};
let (season, episode) = match kind {
IdentityType::Movie => (None, None),
IdentityType::Episode => (params.season, params.episode),
};
let client_cut = params.client_cut();
let found = state
.db
.read(move |conn| {
let Some(title) = repo::find_title(conn, kind, tmdb_id.as_deref(), imdb_id.as_deref())?
else {
return Ok(None);
};
let candidates = repo::candidates_for_title(conn, &title.id, season, episode)?;
if candidates.is_empty() {
return Ok(None);
}
let cuts: Vec<(String, StoredCut)> = candidates
.iter()
.map(|m| {
(
m.id.clone(),
StoredCut { runtime_sec: m.runtime_sec, video_hash: m.video_hash.clone() },
)
})
.collect();
let Some((id, m)) = matching::best_match(&client_cut, &cuts) else {
return Ok(None);
};
let actor_count = repo::manifest_actor_ids(conn, &id)?.len() as i64;
Ok(Some((id, m.tier, actor_count)))
})
.await
.map_err(ApiError::Internal)?;
// §4: `exists: false` is returned with `200`, not `404` — absence is a normal
// answer to this question, and `404` would conflate "no manifest" with "bad
// route" for the client.
Ok(match found {
Some((id, tier, actor_count)) => ExistsResponse {
exists: true,
// With no cut parameters the answer is "some manifest exists" with
// `"match": "unknown"`; the client must still fetch to find out
// whether a cut aligns. This is the mode a library sweep uses (§4).
r#match: Some(tier.as_str()),
manifest_id: Some(id),
actor_count: Some(actor_count),
},
None => ExistsResponse::absent(),
})
}
/// Determines whether these parameters address a movie or an episode.
pub fn resolve_kind(
params: &LookupParams,
) -> Option<(IdentityType, Option<String>, Option<String>)> {
if params.series_tmdb_id.is_some() || params.series_imdb_id.is_some() {
// Episode coordinates are required alongside series identity; without
// them the caller wants the series bundle endpoint instead.
params.season?;
params.episode?;
return Some((
IdentityType::Episode,
params.series_tmdb_id.clone(),
params.series_imdb_id.clone(),
));
}
if params.tmdb_id.is_some() || params.imdb_id.is_some() {
return Some((IdentityType::Movie, params.tmdb_id.clone(), params.imdb_id.clone()));
}
None
}
/// Exposed for tests asserting the documented tier string.
pub fn tier_str(t: MatchTier) -> &'static str {
t.as_str()
}
+351
View File
@@ -0,0 +1,351 @@
//! Manifest fetch endpoints (§4).
//!
//! §7: the submitted JSON was parsed, validated, resolved to TMDB person ids,
//! written as rows and discarded. Everything served here is **reconstructed**
//! from those rows, never echoed — which is what makes §5a's Threat 1 defence
//! structural rather than a promise.
use axum::extract::{Path, Query, State};
use axum::http::HeaderMap;
use axum::response::{IntoResponse, Response};
use axum::Json;
use serde::Serialize;
use super::LookupParams;
use crate::db::repo::{self, ManifestRow};
use crate::error::{ApiError, ApiResult};
use crate::matching::{self, StoredCut};
use crate::model::{
Actor, Coverage, Cut, Extraction, GalleryScope, Identity, IdentityType, Jmanifest, MatchTier,
SeriesBundle, SeriesRef, JMANIFEST_VERSION,
};
use crate::ratelimit::Surface;
use crate::state::{with_quota_headers, AppState};
#[derive(Debug, Serialize)]
pub struct FetchResponse {
pub r#match: &'static str,
/// Scene offset the client must add (§3). Zero for the tiers currently
/// served; present unconditionally so the plugin contract does not change
/// when `audio` is enabled.
pub offset_sec: f64,
pub manifest: Jmanifest,
}
#[derive(Debug, Serialize)]
pub struct StatusResponse {
pub status: String,
#[serde(skip_serializing_if = "Option::is_none")]
pub reason: Option<String>,
}
pub async fn get_movie(
State(state): State<AppState>,
peer: crate::state::PeerIp,
headers: HeaderMap,
Query(params): Query<LookupParams>,
) -> ApiResult<Response> {
let ip = state.client_ip(&headers, peer.0);
let quota = state.check_limit(&ip, Surface::ManifestFetch)?;
if params.tmdb_id.is_none() && params.imdb_id.is_none() {
return Err(ApiError::BadRequest("requires tmdb_id or imdb_id".into()));
}
let body = fetch_best(&state, IdentityType::Movie, &params, None, None).await?;
Ok(with_quota_headers(Json(body).into_response(), quota))
}
pub async fn get_episode(
State(state): State<AppState>,
peer: crate::state::PeerIp,
headers: HeaderMap,
Query(params): Query<LookupParams>,
) -> ApiResult<Response> {
let ip = state.client_ip(&headers, peer.0);
let quota = state.check_limit(&ip, Surface::ManifestFetch)?;
if params.series_tmdb_id.is_none() && params.series_imdb_id.is_none() {
return Err(ApiError::BadRequest("requires series_tmdb_id or series_imdb_id".into()));
}
let (Some(season), Some(episode)) = (params.season, params.episode) else {
return Err(ApiError::BadRequest("requires season and episode".into()));
};
let body =
fetch_best(&state, IdentityType::Episode, &params, Some(season), Some(episode)).await?;
Ok(with_quota_headers(Json(body).into_response(), quota))
}
async fn fetch_best(
state: &AppState,
kind: IdentityType,
params: &LookupParams,
season: Option<i64>,
episode: Option<i64>,
) -> ApiResult<FetchResponse> {
let (tmdb_id, imdb_id) = match kind {
IdentityType::Movie => (params.tmdb_id.clone(), params.imdb_id.clone()),
IdentityType::Episode => (params.series_tmdb_id.clone(), params.series_imdb_id.clone()),
};
let client_cut = params.client_cut();
let found = state
.db
.read(move |conn| {
let Some(title) = repo::find_title(conn, kind, tmdb_id.as_deref(), imdb_id.as_deref())?
else {
return Ok(None);
};
let candidates = repo::candidates_for_title(conn, &title.id, season, episode)?;
let cuts: Vec<(ManifestRow, StoredCut)> = candidates
.into_iter()
.map(|m| {
let cut =
StoredCut { runtime_sec: m.runtime_sec, video_hash: m.video_hash.clone() };
(m, cut)
})
.collect();
let Some((row, m)) = matching::best_match(&client_cut, &cuts) else {
return Ok(None);
};
let manifest = reconstruct(conn, &row, &title, kind)?;
Ok(Some((m.tier, m.offset_sec, manifest)))
})
.await
.map_err(ApiError::Internal)?;
// §4: `404` if none clears `loose`.
let (tier, offset_sec, manifest) = found.ok_or(ApiError::NotFound)?;
Ok(FetchResponse { r#match: tier.as_str(), offset_sec, manifest })
}
/// `GET /manifests/series/{series_tmdb_id}?season=` (§4).
///
/// Returns whatever episodes the server holds. **Partial bundles are normal** — a
/// bundle with 9 of 13 episodes is a valid, useful response, not an error (§2).
/// Episode-level cut matching is done client-side against the returned bundle,
/// since a client pulling a whole series already knows its own runtimes.
pub async fn get_series(
State(state): State<AppState>,
peer: crate::state::PeerIp,
headers: HeaderMap,
Path(series_tmdb_id): Path<String>,
Query(params): Query<LookupParams>,
) -> ApiResult<Response> {
let ip = state.client_ip(&headers, peer.0);
let quota = state.check_limit(&ip, Surface::SeriesFetch)?;
let season = params.season;
let bundle = state
.db
.read(move |conn| {
let Some(title) =
repo::find_title(conn, IdentityType::Episode, Some(&series_tmdb_id), None)?
else {
return Ok(None);
};
let rows = repo::episodes_for_series(conn, &title.id, season)?;
// Multiple contributors may hold the same episode; `episodes_for_series`
// orders by rank, so keep the first per (season, episode).
let mut episodes: Vec<Jmanifest> = Vec::new();
let mut seen: Vec<(i64, i64)> = Vec::new();
let mut seasons: Vec<i64> = Vec::new();
for row in rows {
let key = (row.season.unwrap_or(-1), row.episode.unwrap_or(-1));
if seen.contains(&key) {
continue;
}
seen.push(key);
if !seasons.contains(&key.0) {
seasons.push(key.0);
}
episodes.push(reconstruct(conn, &row, &title, IdentityType::Episode)?);
}
seasons.sort_unstable();
Ok(Some(SeriesBundle {
jmanifest_version: JMANIFEST_VERSION,
series: SeriesRef {
series_tmdb_id: title.tmdb_id.clone(),
series_imdb_id: title.imdb_id.clone(),
title: title.name.clone(),
},
coverage: Some(Coverage { episodes_available: episodes.len(), seasons }),
episodes,
}))
})
.await
.map_err(ApiError::Internal)?;
let bundle = bundle.filter(|b| !b.episodes.is_empty()).ok_or(ApiError::NotFound)?;
Ok(with_quota_headers(Json(bundle).into_response(), quota))
}
/// `GET /manifests/{id}` — fetch a specific manifest by its server-assigned id,
/// for debugging and for the "report this manifest" flow (§4).
pub async fn get_by_id(
State(state): State<AppState>,
peer: crate::state::PeerIp,
headers: HeaderMap,
Path(id): Path<String>,
) -> ApiResult<Response> {
let ip = state.client_ip(&headers, peer.0);
let quota = state.check_limit(&ip, Surface::ManifestFetch)?;
let manifest = state
.db
.read(move |conn| {
let Some(row) = repo::manifest_by_id(conn, &id)? else { return Ok(None) };
// Unlisted manifests are not served to anyone (§6 stage 3).
if row.status != "listed" && row.status != "flagged" {
return Ok(None);
}
let title = title_of(conn, &row.title_id)?;
let kind =
if title.kind == "movie" { IdentityType::Movie } else { IdentityType::Episode };
Ok(Some(reconstruct(conn, &row, &title, kind)?))
})
.await
.map_err(ApiError::Internal)?;
let manifest = manifest.ok_or(ApiError::NotFound)?;
Ok(with_quota_headers(Json(manifest).into_response(), quota))
}
/// `GET /manifests/{id}/status` — poll the outcome of the asynchronous cast
/// check (§4).
pub async fn get_status(
State(state): State<AppState>,
Path(id): Path<String>,
) -> ApiResult<Json<StatusResponse>> {
let found = state
.db
.read(move |conn| repo::manifest_status(conn, &id))
.await
.map_err(ApiError::Internal)?;
match found {
Some((status, reason)) => Ok(Json(StatusResponse { status, reason })),
// §6 deletes rejected manifests, so a vanished id is reported as
// rejected rather than as a bad route.
None => Ok(Json(StatusResponse {
status: "rejected".into(),
reason: Some("not_found_or_rejected".into()),
})),
}
}
fn title_of(conn: &rusqlite::Connection, title_id: &str) -> anyhow::Result<repo::TitleRow> {
let row = conn.query_row(
"SELECT id, kind, tmdb_id, imdb_id, name, year, adult, certification
FROM titles WHERE id = ?1",
rusqlite::params![title_id],
|r| {
Ok(repo::TitleRow {
id: r.get(0)?,
kind: r.get(1)?,
tmdb_id: r.get(2)?,
imdb_id: r.get(3)?,
name: r.get(4)?,
year: r.get(5)?,
adult: r.get::<_, i64>(6)? != 0,
certification: r.get(7)?,
})
},
)?;
Ok(row)
}
/// Rebuilds a Jmanifest from stored rows.
///
/// Names come from `people` — populated from TMDB by the server — so `name` is
/// server-authoritative on download and a name a contributor invented does not
/// round-trip (§2, §5a).
pub fn reconstruct(
conn: &rusqlite::Connection,
row: &ManifestRow,
title: &repo::TitleRow,
kind: IdentityType,
) -> anyhow::Result<Jmanifest> {
let stored = repo::actors_for_manifest(conn, &row.id)?;
let actors = stored
.into_iter()
.map(|a| Actor {
name: a.name,
imdb_id: None,
tmdb_id: Some(a.tmdb_person_id.to_string()),
scenes: a
.scenes_cs
.into_iter()
.map(|(s, e)| [s as f64 / 100.0, e as f64 / 100.0])
.collect(),
})
.collect();
let identity = match kind {
IdentityType::Movie => Identity {
kind,
tmdb_id: title.tmdb_id.clone(),
imdb_id: title.imdb_id.clone(),
series_tmdb_id: None,
series_imdb_id: None,
season: None,
episode: None,
title: title.name.clone(),
year: title.year,
},
IdentityType::Episode => Identity {
kind,
tmdb_id: None,
imdb_id: None,
series_tmdb_id: title.tmdb_id.clone(),
series_imdb_id: title.imdb_id.clone(),
season: row.season,
episode: row.episode,
title: title.name.clone(),
year: title.year,
},
};
// An unrecognised stored scope is served as absent rather than guessed at:
// the column is written from a closed enum, so anything else means the row
// predates a schema change and its meaning is unknown (UR-014's spirit).
let gallery_scope = match row.gallery_scope.as_deref() {
Some("global") => Some(GalleryScope::Global),
Some("limited") => Some(GalleryScope::Limited),
_ => None,
};
let extraction = Extraction {
sample_fps: row.sample_fps,
extinction_sec: row.extinction_sec,
pipeline_version: row.pipeline_version.clone(),
gallery_size: None,
gallery_scope,
};
let has_extraction = extraction.sample_fps.is_some()
|| extraction.extinction_sec.is_some()
|| extraction.pipeline_version.is_some()
|| extraction.gallery_scope.is_some();
Ok(Jmanifest {
jmanifest_version: JMANIFEST_VERSION,
identity,
cut: Cut {
runtime_sec: row.runtime_sec,
container_duration_sec: None,
video_hash: row.video_hash.clone(),
audio_signature: None,
},
extraction: has_extraction.then_some(extraction),
actors,
})
}
/// Exposed so tests can assert the served tier strings.
pub fn tier_name(t: MatchTier) -> &'static str {
t.as_str()
}
+287
View File
@@ -0,0 +1,287 @@
//! A JSON extractor that fails with the status codes §4 specifies.
//!
//! Axum's own `Json` rejects a body that parses as JSON but does not match the
//! target type with **422 Unprocessable Entity**. §4 is explicit that this case
//! is **`400`** — "malformed, or contains an unrecognised or forbidden field" —
//! and that distinction is load-bearing: §6 requires that a client which forgets
//! to strip `movie` or `jellyfin_id` gets "a hard `400` naming the offending
//! field". A client checking for 400 would mishandle a 422.
//!
//! This wrapper also guarantees the field name reaches the caller, since serde's
//! `deny_unknown_fields` error text is what identifies the offending key.
use axum::extract::{FromRequest, Request};
use axum::http::header::CONTENT_TYPE;
use crate::error::ApiError;
/// Drop-in replacement for `axum::Json` on request bodies.
pub struct Json<T>(pub T);
impl<T, S> FromRequest<S> for Json<T>
where
T: serde::de::DeserializeOwned,
S: Send + Sync,
{
type Rejection = ApiError;
async fn from_request(req: Request, state: &S) -> Result<Self, Self::Rejection> {
// A wrong content type is the client's mistake, reported as such rather
// than as a parse failure.
let content_type =
req.headers().get(CONTENT_TYPE).and_then(|v| v.to_str().ok()).unwrap_or("").to_string();
let mime = content_type.split(';').next().unwrap_or("").trim().to_ascii_lowercase();
if !(mime == "application/json" || mime.ends_with("+json")) {
return Err(ApiError::BadRequest("expected content-type: application/json".into()));
}
// A declared charset other than UTF-8 is refused up front, so the client
// learns what is wrong rather than receiving a confusing parse error from
// deep inside the document. See `require_utf8` for why UTF-8 is the only
// accepted encoding.
if let Some(charset) =
content_type.split(';').skip(1).filter_map(|p| p.trim().strip_prefix("charset=")).next()
{
let charset = charset.trim().trim_matches('"').to_ascii_lowercase();
if !matches!(charset.as_str(), "utf-8" | "utf8") {
return Err(ApiError::BadRequest(format!(
"unsupported charset {charset:?}: JSON must be UTF-8 encoded (RFC 8259 §8.1)"
)));
}
}
let bytes = axum::body::Bytes::from_request(req, state).await.map_err(|e| {
// §6 stage 1: the body cap aborts mid-transfer, and that must surface
// as `413`, not as a generic parse error. Axum folds the length-limit
// case into `FailedToBufferBody`, so the status it chose is the
// reliable discriminator.
if e.status() == axum::http::StatusCode::PAYLOAD_TOO_LARGE {
ApiError::PayloadTooLarge("request body exceeds the limit for this route".into())
} else {
ApiError::BadRequest(format!("could not read request body: {e}"))
}
})?;
// Encoding is checked before parsing, so a mis-encoded body gets an
// actionable message instead of whatever the parser happens to trip over.
let text = require_utf8(&bytes)?;
serde_json::from_str(text)
.map(Json)
// serde's message names the offending field, which is exactly what §6
// requires the response to identify.
.map_err(|e| ApiError::BadRequest(e.to_string()))
}
}
/// Enforces that the body is UTF-8, naming the encoding it appears to be.
///
/// **UTF-8 is the only accepted encoding, deliberately.** RFC 8259 §8.1 requires
/// it for JSON exchanged outside a closed ecosystem, and this is a public,
/// federated API. Three further reasons make it the right call *here*
/// specifically, rather than merely conventional:
///
/// 1. **§9a content addressing hashes bytes.** `content_id` is a SHA-256 over the
/// canonical form, so the same manifest submitted in two encodings would
/// produce two different ids — silently defeating federation deduplication.
/// That is precisely the failure mode §9a quantises scene times to avoid, and
/// it would be reintroduced at the encoding layer.
/// 2. **UTF-16 admits lone surrogates**, which have no UTF-8 representation. A
/// field able to carry them is a channel for bytes that survive validation but
/// are not text — against §5a's premise that no field can carry a payload.
/// 3. **§5a's character class assumes well-formed Unicode scalar values.** NFC
/// normalisation and the category checks are defined over scalars, so admitting
/// an encoding that can express non-scalars would undermine both.
///
/// serde_json would reject non-UTF-8 anyway; the value added here is a diagnosable
/// error rather than a misleading one. A UTF-16 body otherwise fails with "key
/// must be a string", which points an operator at the wrong problem entirely.
fn require_utf8(bytes: &[u8]) -> Result<&str, ApiError> {
// A BOM is not valid JSON (RFC 8259 §8.1: "implementations MUST NOT add a
// byte order mark"), and it is the clearest signal of an encoding mistake, so
// it is named rather than left to the parser.
let encoding_hint = match bytes {
[0xEF, 0xBB, 0xBF, ..] => Some("UTF-8 with a byte order mark"),
[0xFF, 0xFE, 0x00, 0x00, ..] => Some("UTF-32LE"),
[0x00, 0x00, 0xFE, 0xFF, ..] => Some("UTF-32BE"),
[0xFF, 0xFE, ..] => Some("UTF-16LE"),
[0xFE, 0xFF, ..] => Some("UTF-16BE"),
// Unmarked UTF-16 is the common case, since encoders often omit the BOM.
// A JSON document always begins with an ASCII character, so an
// interleaved NUL in the first two bytes is conclusive.
[0x00, b, ..] if b.is_ascii_graphic() => Some("UTF-16BE (no BOM)"),
[b, 0x00, ..] if b.is_ascii_graphic() => Some("UTF-16LE (no BOM)"),
_ => None,
};
if let Some(encoding) = encoding_hint {
return Err(ApiError::BadRequest(format!(
"request body appears to be {encoding}: JSON must be UTF-8 encoded \
without a byte order mark (RFC 8259 §8.1)"
)));
}
std::str::from_utf8(bytes).map_err(|e| {
ApiError::BadRequest(format!(
"request body is not valid UTF-8 at byte {}: JSON must be UTF-8 encoded \
(RFC 8259 §8.1)",
e.valid_up_to()
))
})
}
#[cfg(test)]
mod tests {
use super::*;
use axum::http::StatusCode;
use axum::response::IntoResponse;
#[derive(serde::Deserialize)]
#[serde(deny_unknown_fields)]
struct Probe {
_wanted: i64,
}
async fn extract(body: &'static str, content_type: Option<&str>) -> StatusCode {
let mut builder = Request::builder().method("POST").uri("/");
if let Some(ct) = content_type {
builder = builder.header(CONTENT_TYPE, ct);
}
let req = builder.body(axum::body::Body::from(body)).unwrap();
match Json::<Probe>::from_request(req, &()).await {
Ok(_) => StatusCode::OK,
Err(e) => e.into_response().status(),
}
}
#[tokio::test]
async fn schema_mismatch_is_400_not_422() {
// The whole reason this extractor exists (§4, §6).
assert_eq!(
extract(r#"{"unexpected":1}"#, Some("application/json")).await,
StatusCode::BAD_REQUEST
);
}
#[tokio::test]
async fn malformed_json_is_400() {
assert_eq!(extract("{ nope", Some("application/json")).await, StatusCode::BAD_REQUEST);
}
#[tokio::test]
async fn missing_content_type_is_400() {
assert_eq!(extract(r#"{"_wanted":1}"#, None).await, StatusCode::BAD_REQUEST);
}
#[tokio::test]
async fn content_type_parameters_are_tolerated() {
assert_eq!(
extract(r#"{"_wanted":1}"#, Some("application/json; charset=utf-8")).await,
StatusCode::OK
);
}
#[tokio::test]
async fn valid_body_extracts() {
assert_eq!(extract(r#"{"_wanted":1}"#, Some("application/json")).await, StatusCode::OK);
}
#[tokio::test]
async fn an_explicit_utf8_charset_is_accepted() {
for ct in [
"application/json; charset=utf-8",
"application/json;charset=UTF-8",
"application/json; charset=\"utf-8\"",
"application/json; charset=utf8",
] {
assert_eq!(extract(r#"{"_wanted":1}"#, Some(ct)).await, StatusCode::OK, "{ct}");
}
}
#[tokio::test]
async fn a_non_utf8_charset_is_refused_by_name() {
for ct in [
"application/json; charset=utf-16",
"application/json; charset=iso-8859-1",
"application/json; charset=windows-1252",
] {
assert_eq!(
extract(r#"{"_wanted":1}"#, Some(ct)).await,
StatusCode::BAD_REQUEST,
"{ct}"
);
}
}
/// Builds a request from raw bytes, since these bodies are not valid `&str`.
async fn extract_bytes(body: Vec<u8>) -> Result<(), ApiError> {
let req = Request::builder()
.method("POST")
.uri("/")
.header(CONTENT_TYPE, "application/json")
.body(axum::body::Body::from(body))
.unwrap();
Json::<Probe>::from_request(req, &()).await.map(|_| ())
}
#[tokio::test]
async fn utf16_bodies_are_rejected_with_an_actionable_message() {
// The reason this check exists: serde_json rejects UTF-16 anyway, but with
// "key must be a string", which points an operator at the wrong problem.
let doc = r#"{"_wanted":1}"#;
let le: Vec<u8> = doc.encode_utf16().flat_map(|u| u.to_le_bytes()).collect();
let err = extract_bytes(le).await.unwrap_err().to_string();
assert!(err.contains("UTF-16LE"), "should name the encoding: {err}");
assert!(err.contains("UTF-8"), "should say what is required: {err}");
let be: Vec<u8> = doc.encode_utf16().flat_map(|u| u.to_be_bytes()).collect();
let err = extract_bytes(be).await.unwrap_err().to_string();
assert!(err.contains("UTF-16BE"), "should name the encoding: {err}");
// With BOMs.
let mut le_bom = vec![0xFF, 0xFE];
le_bom.extend(doc.encode_utf16().flat_map(|u| u.to_le_bytes()));
assert!(extract_bytes(le_bom).await.is_err());
let mut be_bom = vec![0xFE, 0xFF];
be_bom.extend(doc.encode_utf16().flat_map(|u| u.to_be_bytes()));
assert!(extract_bytes(be_bom).await.is_err());
}
#[tokio::test]
async fn a_utf8_bom_is_rejected() {
// RFC 8259 §8.1: implementations MUST NOT add a byte order mark.
let mut body = vec![0xEF, 0xBB, 0xBF];
body.extend_from_slice(br#"{"_wanted":1}"#);
let err = extract_bytes(body).await.unwrap_err().to_string();
assert!(err.contains("byte order mark"), "{err}");
}
#[tokio::test]
async fn invalid_utf8_is_rejected_with_the_offending_offset() {
// A truncated multi-byte sequence inside an otherwise well-formed document.
let body = b"{\"_wanted\":\"\xC3\x28\"}".to_vec();
let err = extract_bytes(body).await.unwrap_err().to_string();
assert!(err.contains("not valid UTF-8"), "{err}");
assert!(err.contains("byte 12"), "should locate the failure: {err}");
}
#[tokio::test]
async fn valid_multibyte_utf8_is_accepted() {
// The check must not reject legitimate non-ASCII content — actor names are
// routinely non-Latin (§5a accepts any Unicode letter).
// Rejected for the unknown `_note` field, not for its encoding — which is
// the distinction being asserted.
let body = r#"{"_wanted":1,"_note":"宮崎 駿 Renée"}"#.as_bytes().to_vec();
let err = extract_bytes(body).await.unwrap_err().to_string();
assert!(err.contains("_note"), "should fail on the schema, not the encoding: {err}");
}
#[tokio::test]
async fn an_empty_body_is_not_mistaken_for_an_encoding_problem() {
let err = extract_bytes(Vec::new()).await.unwrap_err().to_string();
assert!(!err.contains("UTF-16"), "empty body is a parse error, not an encoding one: {err}");
}
}
+35
View File
@@ -0,0 +1,35 @@
//! HTTP surface (§4). Base path `/api/v1`, JSON throughout.
pub mod exists;
pub mod fetch;
pub mod json;
pub mod report;
pub mod upload;
use serde::Deserialize;
use crate::matching::ClientCut;
/// Identity + cut query parameters, shared by the read endpoints (§4).
#[derive(Debug, Clone, Default, Deserialize)]
pub struct LookupParams {
pub tmdb_id: Option<String>,
pub imdb_id: Option<String>,
pub series_tmdb_id: Option<String>,
pub series_imdb_id: Option<String>,
pub season: Option<i64>,
pub episode: Option<i64>,
pub runtime_sec: Option<f64>,
pub video_hash: Option<String>,
}
impl LookupParams {
pub fn client_cut(&self) -> ClientCut {
ClientCut {
// A non-finite or non-positive runtime is not a usable signal; treat
// it as absent rather than letting it drive a match.
runtime_sec: self.runtime_sec.filter(|r| r.is_finite() && *r > 0.0),
video_hash: self.video_hash.clone(),
}
}
}
+141
View File
@@ -0,0 +1,141 @@
//! `POST /manifests/{id}/report` (§4), and `GET /health`.
//!
//! Reports are a moderation lever and cheap to abuse, hence the tight §5 limit.
//! A report never changes `status` by itself: §5a keeps delisting an operator
//! action, because automatic delisting on report would hand any client a remote
//! delete primitive.
use axum::extract::{Path, State};
use axum::http::HeaderMap;
use axum::response::{IntoResponse, Response};
use axum::Json;
use serde::{Deserialize, Serialize};
use crate::db::repo;
use crate::error::{ApiError, ApiResult};
use crate::ratelimit::Surface;
use crate::state::{with_quota_headers, AppState};
use crate::worker::now_iso;
/// §4: `{ "reason": "misaligned" | "wrong_actors" | "spam", "note": "..." }`.
#[derive(Debug, Clone, Copy, PartialEq, Eq, Deserialize, Serialize)]
#[serde(rename_all = "snake_case")]
pub enum ReportReason {
Misaligned,
WrongActors,
Spam,
}
impl ReportReason {
fn as_str(self) -> &'static str {
match self {
ReportReason::Misaligned => "misaligned",
ReportReason::WrongActors => "wrong_actors",
ReportReason::Spam => "spam",
}
}
}
#[derive(Debug, Deserialize)]
#[serde(deny_unknown_fields)]
pub struct ReportRequest {
pub reason: ReportReason,
#[serde(default)]
pub note: Option<String>,
}
/// §5a: `note` is free text from an anonymous caller, so it is capped hard. It is
/// never served back to clients — only the operator reads it.
const MAX_NOTE_CHARS: usize = 500;
#[derive(Debug, Serialize)]
pub struct ReportAccepted {
pub report_id: String,
}
pub async fn post_report(
State(state): State<AppState>,
peer: crate::state::PeerIp,
headers: HeaderMap,
Path(manifest_id): Path<String>,
super::json::Json(req): super::json::Json<ReportRequest>,
) -> ApiResult<Response> {
let ip = state.client_ip(&headers, peer.0);
let quota = state.check_limit(&ip, Surface::Report)?;
let note = match req.note {
Some(n) if n.chars().count() > MAX_NOTE_CHARS => {
return Err(ApiError::BadRequest(format!(
"note: longer than {MAX_NOTE_CHARS} characters"
)))
}
// Strip control characters; the note is operator-facing text, not markup.
Some(n) => Some(n.chars().filter(|c| !c.is_control()).collect::<String>()),
None => None,
};
let ip_hash = crate::auth::hash_ip(&ip, &state.config.server_id);
let reason = req.reason.as_str();
let now = now_iso();
let id_for_check = manifest_id.clone();
let exists = state
.db
.read(move |c| Ok(repo::manifest_by_id(c, &id_for_check)?.is_some()))
.await
.map_err(ApiError::Internal)?;
if !exists {
return Err(ApiError::NotFound);
}
let report_id = state
.db
.write(move |tx| {
repo::insert_report(tx, &manifest_id, reason, note.as_deref(), &ip_hash, &now)
})
.await
.map_err(ApiError::Internal)?;
Ok(with_quota_headers(Json(ReportAccepted { report_id }).into_response(), quota))
}
#[derive(Debug, Serialize)]
pub struct Health {
pub status: &'static str,
pub version: &'static str,
}
/// `GET /health` — liveness, unauthenticated and unlimited (§4, §5).
pub async fn health() -> Json<Health> {
Json(Health { status: "ok", version: env!("CARGO_PKG_VERSION") })
}
#[derive(Debug, Serialize)]
pub struct Readiness {
pub status: &'static str,
pub database: &'static str,
/// §8: TMDB is a hard dependency for UR-3. If it is unconfigured, uploads
/// accumulate in `pending` rather than being listed unverified — worth
/// surfacing rather than failing silently.
pub tmdb_configured: bool,
}
/// Readiness check verifying the database opens and migrations are current (§8).
pub async fn ready(State(state): State<AppState>) -> ApiResult<Json<Readiness>> {
let ok = state
.db
.read(|conn| {
// Any query against a schema table proves both that the file opens
// and that migrations have been applied.
let n: i64 = conn.query_row("SELECT COUNT(*) FROM manifests", [], |r| r.get(0))?;
Ok(n >= 0)
})
.await
.map_err(ApiError::Internal)?;
Ok(Json(Readiness {
status: if ok { "ready" } else { "degraded" },
database: "ok",
tmdb_configured: state.tmdb.is_configured(),
}))
}
+241
View File
@@ -0,0 +1,241 @@
//! Contribution endpoints (§4) — UR-2 and UR-6.
//!
//! Both require a token (§5). Both return `202`: the upload has passed size and
//! schema validation and is held unlisted pending the asynchronous TMDB cast
//! check (§6 stage 3).
use axum::extract::State;
use axum::http::{HeaderMap, StatusCode};
use axum::response::{IntoResponse, Response};
use axum::Json;
use serde::Serialize;
use crate::db::repo;
use crate::error::{ApiError, ApiResult};
use crate::ingest::{self, IngestOutcome};
use crate::model::{IdentityType, Jmanifest, SeriesBundle};
use crate::ratelimit::Surface;
use crate::state::{with_quota_headers, AppState};
use crate::validate::{self, limits};
use crate::worker::now_iso;
#[derive(Debug, Serialize)]
pub struct UploadAccepted {
pub manifest_id: String,
pub status: &'static str,
}
/// `POST /manifests` — UR-2.
pub async fn post_manifest(
State(state): State<AppState>,
headers: HeaderMap,
super::json::Json(manifest): super::json::Json<Jmanifest>,
) -> ApiResult<Response> {
let contributor = state.require_contributor(&headers).await?;
// §5: limits are per token where one is present.
let quota = state.check_limit(&contributor.id, Surface::ManifestUpload)?;
// §6 stage 2. A rejection names the offending field, so a client that forgets
// to strip `movie`/`jellyfin_id` gets a diagnosable `400`.
let valid =
validate::validate_manifest(manifest).map_err(|e| ApiError::BadRequest(e.to_string()))?;
let origin = state.config.server_id.clone();
let contributor_id = contributor.id.clone();
let now = now_iso();
let outcome = state
.db
.write(move |tx| ingest::persist(tx, &valid, Some(&contributor_id), &origin, None, &now))
.await
.map_err(ApiError::Internal)?;
let resp = match outcome {
IngestOutcome::Pending { manifest_id } => {
(StatusCode::ACCEPTED, Json(UploadAccepted { manifest_id, status: "pending" }))
.into_response()
}
// §4 `409` — an identical `(identity, cut)` manifest already exists from
// this contributor.
IngestOutcome::DuplicateFromContributor { manifest_id } => {
return Err(ApiError::Conflict(format!(
"an identical manifest already exists from this contributor: {manifest_id}"
)))
}
// §9a: identical content already held, from any source. Not an error —
// the contributor's work is simply already represented.
IngestOutcome::DuplicateContent { manifest_id } => {
(StatusCode::OK, Json(UploadAccepted { manifest_id, status: "already_present" }))
.into_response()
}
};
Ok(with_quota_headers(resp, quota))
}
#[derive(Debug, Serialize)]
pub struct BundleResult {
pub season: Option<i64>,
pub episode: Option<i64>,
#[serde(skip_serializing_if = "Option::is_none")]
pub manifest_id: Option<String>,
pub status: &'static str,
#[serde(skip_serializing_if = "Option::is_none")]
pub reason: Option<String>,
}
#[derive(Debug, Serialize)]
pub struct BundleAccepted {
pub results: Vec<BundleResult>,
}
/// `POST /manifests/bundle` — UR-6.
///
/// **Per-episode validation, not atomic**: valid episodes are accepted and
/// invalid ones rejected, with a per-episode result list. All-or-nothing would let
/// one bad episode discard an entire season's compute (§2).
///
/// **One rate-limit unit**, so contributing a season is not punished relative to
/// contributing a film (§2, §5).
pub async fn post_bundle(
State(state): State<AppState>,
headers: HeaderMap,
super::json::Json(bundle): super::json::Json<SeriesBundle>,
) -> ApiResult<Response> {
let contributor = state.require_contributor(&headers).await?;
let quota = state.check_limit(&contributor.id, Surface::BundleUpload)?;
// §4: `413` for exceeding the episode cap, distinct from a malformed envelope.
if bundle.episodes.len() > limits::MAX_BUNDLE_EPISODES {
return Err(ApiError::PayloadTooLarge(format!(
"bundle carries {} episodes, limit is {}",
bundle.episodes.len(),
limits::MAX_BUNDLE_EPISODES
)));
}
// §4: `400` only for the envelope itself; individual bad episodes are
// reported in the results list, not as a whole-request error.
validate::validate_bundle_envelope(&bundle).map_err(|e| ApiError::BadRequest(e.to_string()))?;
let series_tmdb = bundle.series.series_tmdb_id.clone();
let mut results = Vec::with_capacity(bundle.episodes.len());
for episode in bundle.episodes {
let coords = (episode.identity.season, episode.identity.episode);
// An episode whose identity contradicts the envelope is rejected on its
// own rather than being silently reattributed to the bundle's series.
if episode.identity.kind != IdentityType::Episode {
results.push(BundleResult {
season: coords.0,
episode: coords.1,
manifest_id: None,
status: "rejected",
reason: Some("identity.type must be 'episode' within a bundle".into()),
});
continue;
}
if let (Some(envelope), Some(ep)) = (&series_tmdb, &episode.identity.series_tmdb_id) {
if envelope != ep {
results.push(BundleResult {
season: coords.0,
episode: coords.1,
manifest_id: None,
status: "rejected",
reason: Some("series_tmdb_id does not match the bundle envelope".into()),
});
continue;
}
}
let valid = match validate::validate_manifest(episode) {
Ok(v) => v,
Err(e) => {
results.push(BundleResult {
season: coords.0,
episode: coords.1,
manifest_id: None,
status: "rejected",
reason: Some(e.to_string()),
});
continue;
}
};
let origin = state.config.server_id.clone();
let contributor_id = contributor.id.clone();
let now = now_iso();
// One transaction per episode, so a bundle never holds the write lock for
// the whole request (§8 chunked ingest reasoning).
let outcome = state
.db
.write(move |tx| {
ingest::persist(tx, &valid, Some(&contributor_id), &origin, None, &now)
})
.await;
results.push(match outcome {
Ok(IngestOutcome::Pending { manifest_id }) => BundleResult {
season: coords.0,
episode: coords.1,
manifest_id: Some(manifest_id),
status: "pending",
reason: None,
},
Ok(IngestOutcome::DuplicateFromContributor { manifest_id })
| Ok(IngestOutcome::DuplicateContent { manifest_id }) => BundleResult {
season: coords.0,
episode: coords.1,
manifest_id: Some(manifest_id),
status: "already_present",
reason: None,
},
Err(e) => {
tracing::error!(error = ?e, "bundle episode failed to persist");
BundleResult {
season: coords.0,
episode: coords.1,
manifest_id: None,
status: "rejected",
reason: Some("internal error".into()),
}
}
});
}
let resp = (StatusCode::ACCEPTED, Json(BundleAccepted { results })).into_response();
Ok(with_quota_headers(resp, quota))
}
#[derive(Debug, Serialize)]
pub struct TokenIssued {
pub token: String,
}
/// Issues an anonymous bearer capability (§5a).
///
/// Self-issued on request: no email, no verification, no personal data. Stored
/// only as a hash, so the server cannot enumerate who holds tokens. Discarding a
/// token and requesting another is trivially easy — and that is fine, because the
/// token is not the defence; the content checks are.
pub async fn post_token(
State(state): State<AppState>,
peer: crate::state::PeerIp,
headers: HeaderMap,
) -> ApiResult<Json<TokenIssued>> {
let ip = state.client_ip(&headers, peer.0);
// Reuse the report budget: issuing tokens is cheap but should not be a free
// unbounded write.
state.check_limit(&ip, Surface::Report)?;
let token = crate::auth::generate_token();
let hash = crate::auth::hash_token(&token);
let now = now_iso();
state
.db
.write(move |tx| repo::insert_contributor(tx, &hash, &now))
.await
.map_err(ApiError::Internal)?;
Ok(Json(TokenIssued { token }))
}
+66
View File
@@ -0,0 +1,66 @@
//! Router construction.
//!
//! §6 stage 0/1 body caps are applied here as per-route `DefaultBodyLimit`
//! layers: Axum rejects on `Content-Length` before reading a body *and* caps the
//! stream for chunked or mis-declared uploads, which is what makes a lying
//! header and a chunked upload both safe. Per-route means the bundle endpoint
//! gets its larger limit without widening the others (§6 stage 1).
use std::time::Duration;
use axum::extract::DefaultBodyLimit;
use axum::routing::{get, post};
use axum::Router;
use tower_http::timeout::TimeoutLayer;
use tower_http::trace::TraceLayer;
use crate::api::{exists, fetch, report, upload};
use crate::state::AppState;
use crate::validate::limits;
/// Small cap for endpoints that take a short JSON body. A read endpoint has no
/// business accepting a large payload, and the batch `exists` form is bounded at
/// 100 items.
const SMALL_BODY_LIMIT: usize = 256 * 1024;
pub fn router(state: AppState) -> Router {
let timeout = state.config.request_timeout;
let v1 = Router::new()
// UR-1 — existence probes.
.route("/manifests/exists", get(exists::exists).post(exists::exists_batch))
// Reads.
.route("/manifests/movie", get(fetch::get_movie))
.route("/manifests/episode", get(fetch::get_episode))
.route("/manifests/series/{series_tmdb_id}", get(fetch::get_series))
.route("/manifests/{id}", get(fetch::get_by_id))
.route("/manifests/{id}/status", get(fetch::get_status))
.route("/manifests/{id}/report", post(report::post_report))
// UR-2 — contribution.
.route(
"/manifests",
post(upload::post_manifest).layer(DefaultBodyLimit::max(limits::BODY_LIMIT_MANIFEST)),
)
// UR-6 — whole-series contribution, with its own larger cap.
.route(
"/manifests/bundle",
post(upload::post_bundle).layer(DefaultBodyLimit::max(limits::BODY_LIMIT_BUNDLE)),
)
// §5a — anonymous bearer capability, not an account.
.route("/tokens", post(upload::post_token))
.layer(DefaultBodyLimit::max(SMALL_BODY_LIMIT));
Router::new()
.route("/health", get(report::health))
.route("/ready", get(report::ready))
.nest("/api/v1", v1)
// §8: a request timeout so a slow bundle query fails fast.
.layer(TimeoutLayer::with_status_code(axum::http::StatusCode::REQUEST_TIMEOUT, timeout))
.layer(TraceLayer::new_for_http())
.with_state(state)
}
/// Convenience for tests and `main`.
pub fn default_timeout() -> Duration {
Duration::from_secs(30)
}
+186
View File
@@ -0,0 +1,186 @@
//! §5a tokens and client-IP attribution.
//!
//! A token is **not an account** — it is an anonymous bearer capability. No
//! email, no verification, no personal data. It is stored only as a hash, so the
//! server cannot enumerate who holds tokens, and its sole purposes are
//! rate-limiting attribution (§5) and revocation.
//!
//! Discarding a token and requesting another is trivially easy, and that is
//! fine: the token is not the defence, the content checks are. Sybil resistance
//! is not required because identity is not load-bearing.
use std::net::IpAddr;
use axum::http::HeaderMap;
use sha2::{Digest, Sha256};
/// Hashes a bearer token for storage and lookup.
///
/// Plain SHA-256 rather than a password KDF is deliberate and sufficient here:
/// tokens are 256 bits of server-generated randomness, not user-chosen secrets,
/// so there is no dictionary to attack.
pub fn hash_token(token: &str) -> String {
let mut h = Sha256::new();
h.update(token.as_bytes());
hex(&h.finalize())
}
/// Hashes a client IP for report attribution (§7 `reports.source_ip_hash`).
///
/// Salted with the server id so hashes are not comparable across instances.
pub fn hash_ip(ip: &str, server_id: &str) -> String {
let mut h = Sha256::new();
h.update(server_id.as_bytes());
h.update(b"\0");
h.update(ip.as_bytes());
hex(&h.finalize())
}
fn hex(bytes: &[u8]) -> String {
let mut s = String::with_capacity(bytes.len() * 2);
for b in bytes {
s.push_str(&format!("{b:02x}"));
}
s
}
/// Generates a new token. Returned once to the caller; only its hash is stored.
pub fn generate_token() -> String {
use rand::RngCore;
let mut bytes = [0u8; 32];
rand::rng().fill_bytes(&mut bytes);
format!("jray_{}", hex(&bytes))
}
/// Extracts a bearer token from an `Authorization` header.
pub fn bearer_token(headers: &HeaderMap) -> Option<String> {
let raw = headers.get(axum::http::header::AUTHORIZATION)?.to_str().ok()?;
let (scheme, value) = raw.split_once(' ')?;
if !scheme.eq_ignore_ascii_case("bearer") {
return None;
}
let value = value.trim();
if value.is_empty() {
return None;
}
Some(value.to_string())
}
/// Resolves the client IP for rate-limiting and report attribution.
///
/// §8: the app must trust `X-Forwarded-For` **only** from the operator's proxy.
/// Rate limiting and report attribution key on client IP, so a spoofable header
/// defeats both — hence `trusted_proxies` is explicit configuration and an
/// untrusted peer's header is ignored outright.
pub fn client_ip(headers: &HeaderMap, peer: Option<IpAddr>, trusted_proxies: &[IpAddr]) -> String {
let peer_is_trusted = peer.is_some_and(|p| trusted_proxies.contains(&p));
if peer_is_trusted {
if let Some(xff) = headers.get("x-forwarded-for").and_then(|v| v.to_str().ok()) {
// Right-most entry is the one our trusted proxy appended; entries to
// its left are client-supplied and forgeable. Walk from the right
// past any further trusted hops.
for candidate in xff.split(',').rev().map(str::trim).filter(|s| !s.is_empty()) {
match candidate.parse::<IpAddr>() {
Ok(ip) if trusted_proxies.contains(&ip) => continue,
Ok(ip) => return ip.to_string(),
Err(_) => break,
}
}
}
}
peer.map(|p| p.to_string()).unwrap_or_else(|| "unknown".to_string())
}
#[cfg(test)]
mod tests {
use super::*;
use axum::http::HeaderValue;
fn headers(pairs: &[(&'static str, &str)]) -> HeaderMap {
let mut h = HeaderMap::new();
for (k, v) in pairs {
h.insert(*k, HeaderValue::from_str(v).unwrap());
}
h
}
#[test]
fn token_hash_is_stable_and_distinguishing() {
assert_eq!(hash_token("abc"), hash_token("abc"));
assert_ne!(hash_token("abc"), hash_token("abd"));
assert_eq!(hash_token("abc").len(), 64);
}
#[test]
fn generated_tokens_are_unique_and_prefixed() {
let a = generate_token();
let b = generate_token();
assert_ne!(a, b);
assert!(a.starts_with("jray_"));
assert_eq!(a.len(), 5 + 64);
}
#[test]
fn ip_hash_is_salted_per_server() {
// Hashes must not be comparable across instances.
assert_ne!(hash_ip("1.2.3.4", "a.example"), hash_ip("1.2.3.4", "b.example"));
assert_eq!(hash_ip("1.2.3.4", "a.example"), hash_ip("1.2.3.4", "a.example"));
}
#[test]
fn parses_bearer_tokens_case_insensitively() {
assert_eq!(
bearer_token(&headers(&[("authorization", "Bearer xyz")])).as_deref(),
Some("xyz")
);
assert_eq!(
bearer_token(&headers(&[("authorization", "bearer xyz")])).as_deref(),
Some("xyz")
);
assert!(bearer_token(&headers(&[("authorization", "Basic xyz")])).is_none());
assert!(bearer_token(&headers(&[("authorization", "Bearer ")])).is_none());
assert!(bearer_token(&HeaderMap::new()).is_none());
}
#[test]
fn forwarded_header_from_an_untrusted_peer_is_ignored() {
// The whole point of §8's explicit trusted-proxy configuration: an
// arbitrary client must not be able to choose its own rate-limit key.
let h = headers(&[("x-forwarded-for", "9.9.9.9")]);
let peer: IpAddr = "203.0.113.7".parse().unwrap();
assert_eq!(client_ip(&h, Some(peer), &[]), "203.0.113.7");
}
#[test]
fn forwarded_header_from_a_trusted_proxy_is_honoured() {
let h = headers(&[("x-forwarded-for", "9.9.9.9")]);
let proxy: IpAddr = "127.0.0.1".parse().unwrap();
assert_eq!(client_ip(&h, Some(proxy), &[proxy]), "9.9.9.9");
}
#[test]
fn client_supplied_entries_left_of_the_proxy_cannot_spoof() {
// A client that sends its own XFF gets its value appended to, not
// replaced, so only the right-most entry is trustworthy.
let h = headers(&[("x-forwarded-for", "9.9.9.9, 203.0.113.7")]);
let proxy: IpAddr = "127.0.0.1".parse().unwrap();
assert_eq!(client_ip(&h, Some(proxy), &[proxy]), "203.0.113.7");
}
#[test]
fn walks_past_additional_trusted_hops() {
let inner: IpAddr = "10.0.0.2".parse().unwrap();
let proxy: IpAddr = "127.0.0.1".parse().unwrap();
let h = headers(&[("x-forwarded-for", "203.0.113.7, 10.0.0.2")]);
assert_eq!(client_ip(&h, Some(proxy), &[proxy, inner]), "203.0.113.7");
}
#[test]
fn malformed_forwarded_value_falls_back_to_the_peer() {
let h = headers(&[("x-forwarded-for", "not-an-ip")]);
let proxy: IpAddr = "127.0.0.1".parse().unwrap();
assert_eq!(client_ip(&h, Some(proxy), &[proxy]), "127.0.0.1");
}
}
+484
View File
@@ -0,0 +1,484 @@
//! §6 stage 3 cast-match scoring, as pure functions.
//!
//! The thresholds here are the load-bearing part of UR-3 and §5a Threat 2, and
//! §10 (5) wants them retuned against the 331-file extraction corpus. Keeping
//! the decision logic free of I/O is what makes that a test-data exercise rather
//! than a code change.
use crate::tmdb::CastMember;
/// §6: thresholds over the ratio `|M ∩ C| / |M|`.
pub const LISTED_THRESHOLD: f64 = 0.6;
pub const FLAGGED_THRESHOLD: f64 = 0.3;
/// Below this size a ratio is meaningless (§6 small-|M| handling).
pub const SMALL_M_LIMIT: usize = 5;
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum Verdict {
/// Normal case.
Listed,
/// Served with reduced ranking, flagged for review.
Flagged,
/// Deleted, and the contributor's counter bumped.
Rejected,
}
impl Verdict {
pub fn status(self) -> &'static str {
match self {
Verdict::Listed => "listed",
Verdict::Flagged => "flagged",
Verdict::Rejected => "rejected",
}
}
}
#[derive(Debug, Clone)]
pub struct CastCheckOutcome {
pub verdict: Verdict,
pub ratio: f64,
/// Manifest actors resolved to a TMDB person id, with the TMDB-authoritative
/// name. Only these are kept; §6 drops unmatched actors rather than storing
/// them, which is what closes §5a's free-text channel.
pub matched: Vec<MatchedActor>,
/// Actors that matched nothing and will be dropped.
pub unmatched_person_ids: Vec<u64>,
pub reason: Option<String>,
}
#[derive(Debug, Clone)]
pub struct MatchedActor {
pub tmdb_person_id: u64,
/// From TMDB, never from the upload.
pub name: String,
pub adult: bool,
/// True when the match came from name comparison rather than an id.
pub by_name: bool,
}
/// An actor as submitted, after §6 stage 2 validation.
#[derive(Debug, Clone)]
pub struct SubmittedActor {
pub tmdb_id: Option<u64>,
pub imdb_id: Option<String>,
/// Used only for matching here, then discarded (§5a).
pub name: Option<String>,
}
/// Case- and accent-insensitive comparison key for the name fallback (§6).
fn name_key(s: &str) -> String {
use unicode_normalization::UnicodeNormalization;
s.nfd()
.filter(|c| !unicode_normalization::char::is_combining_mark(*c))
.flat_map(|c| c.to_lowercase())
.filter(|c| !c.is_whitespace() && *c != '.' && *c != ',' && *c != '-' && *c != '\'')
.collect()
}
/// Runs the §6 stage 3 comparison.
///
/// `credits` is the reference set *C*: for a movie, its credits; for an episode,
/// the union of per-episode credits and the series' aggregate credits.
pub fn evaluate(submitted: &[SubmittedActor], credits: &[CastMember]) -> CastCheckOutcome {
let m = submitted.len();
// §6: `|M| == 0` is rejected. These are extraction failures, not
// contributions — validation already refuses them, so reaching here means a
// manifest lost every actor upstream.
if m == 0 {
return CastCheckOutcome {
verdict: Verdict::Rejected,
ratio: 0.0,
matched: Vec::new(),
unmatched_person_ids: Vec::new(),
reason: Some("empty_actor_list".into()),
};
}
// TMDB has no credits for the id: absent data is not evidence of a bad
// manifest, so this is flagged rather than rejected (§6).
if credits.is_empty() {
return CastCheckOutcome {
verdict: Verdict::Flagged,
ratio: 0.0,
matched: Vec::new(),
unmatched_person_ids: submitted.iter().filter_map(|a| a.tmdb_id).collect(),
reason: Some("tmdb_no_credits".into()),
};
}
let mut matched: Vec<MatchedActor> = Vec::new();
let mut unmatched: Vec<u64> = Vec::new();
let mut id_matches = 0usize;
let mut name_matches = 0usize;
for actor in submitted {
// Join on `tmdb_id` — grounded in the pipeline's actual output, where
// 330 of 331 manifests have `imdb_id: ""` and `tmdb_id` set (§6).
let by_id = actor.tmdb_id.and_then(|id| credits.iter().find(|c| c.id == id));
if let Some(c) = by_id {
id_matches += 1;
push_unique(
&mut matched,
MatchedActor {
tmdb_person_id: c.id,
name: c.name.clone(),
adult: c.adult,
by_name: false,
},
);
continue;
}
// Fall back to case- and accent-insensitive name comparison.
let by_name = actor.name.as_deref().and_then(|n| {
let key = name_key(n);
(!key.is_empty()).then(|| credits.iter().find(|c| name_key(&c.name) == key))?
});
if let Some(c) = by_name {
name_matches += 1;
push_unique(
&mut matched,
MatchedActor {
tmdb_person_id: c.id,
name: c.name.clone(),
adult: c.adult,
by_name: true,
},
);
continue;
}
if let Some(id) = actor.tmdb_id {
unmatched.push(id);
}
}
// §6: name-only matches are counted but capped at half the intersection, so
// a manifest cannot pass on name collisions alone.
let capped_name_matches = name_matches.min(id_matches);
let effective = id_matches + capped_name_matches;
let ratio = effective as f64 / m as f64;
let verdict = classify(m, effective, ratio);
let reason = match verdict {
Verdict::Rejected => Some("cast_match_below_threshold".into()),
Verdict::Flagged => Some("cast_match_marginal".into()),
Verdict::Listed => None,
};
CastCheckOutcome { verdict, ratio, matched, unmatched_person_ids: unmatched, reason }
}
fn push_unique(matched: &mut Vec<MatchedActor>, actor: MatchedActor) {
if !matched.iter().any(|m| m.tmdb_person_id == actor.tmdb_person_id) {
matched.push(actor);
}
}
/// §6 small-*M* handling. With a median of 7 actors a ratio threshold is coarse
/// — one mismatch moves it by 14% — so small manifests use counts, not ratios.
fn classify(m: usize, matches: usize, ratio: f64) -> Verdict {
if m >= SMALL_M_LIMIT {
if ratio >= LISTED_THRESHOLD {
Verdict::Listed
} else if ratio >= FLAGGED_THRESHOLD {
Verdict::Flagged
} else {
Verdict::Rejected
}
} else if m >= 2 {
// Require all but one actor to match.
if matches + 1 >= m {
Verdict::Listed
} else {
Verdict::Rejected
}
} else {
// |M| <= 1: accept only if the single actor matches. Such a manifest is
// near-worthless anyway and is ranked last.
if matches >= 1 {
Verdict::Listed
} else {
Verdict::Rejected
}
}
}
/// §5a additional layer 1 — category guard.
///
/// Rejects when a matched person is flagged adult by TMDB and the target title
/// is not, which targets the stated prank without needing a blocklist of names.
pub fn category_guard_violation(matched: &[MatchedActor], title_is_adult: bool) -> Option<u64> {
if title_is_adult {
return None;
}
matched.iter().find(|m| m.adult).map(|m| m.tmdb_person_id)
}
/// §5a additional layer 2 — age-appropriateness guard.
///
/// On a children's certification, apply the strictest cast-match threshold and
/// require an `exact` or `runtime` cut match. Mismatched content on children's
/// titles is the highest-harm case and deserves the tightest gate.
pub fn is_childrens_certification(cert: &str) -> bool {
matches!(
cert.trim().to_ascii_uppercase().as_str(),
"G" | "TV-Y" | "TV-Y7" | "TV-G" | "U" | "0+" | "6+" | "PG" | "TV-PG"
)
}
pub const CHILDRENS_LISTED_THRESHOLD: f64 = 0.8;
/// Applies the children's-title gate to an already-computed outcome.
pub fn apply_childrens_guard(outcome: &mut CastCheckOutcome, m: usize) {
if m >= SMALL_M_LIMIT && outcome.ratio < CHILDRENS_LISTED_THRESHOLD {
outcome.verdict = match outcome.verdict {
Verdict::Listed => Verdict::Flagged,
other => other,
};
if outcome.reason.is_none() {
outcome.reason = Some("childrens_title_strict_threshold".into());
}
}
}
#[cfg(test)]
mod tests {
use super::*;
fn credit(id: u64, name: &str) -> CastMember {
CastMember { id, name: name.to_string(), adult: false }
}
fn adult_credit(id: u64, name: &str) -> CastMember {
CastMember { id, name: name.to_string(), adult: true }
}
fn by_id(id: u64) -> SubmittedActor {
SubmittedActor { tmdb_id: Some(id), imdb_id: None, name: None }
}
fn by_name(name: &str) -> SubmittedActor {
SubmittedActor { tmdb_id: None, imdb_id: None, name: Some(name.to_string()) }
}
/// A realistic reference cast — feature casts are several times larger than
/// the manifests extracted from them (§6).
fn cast_of_20() -> Vec<CastMember> {
(1..=20).map(|i| credit(i, &format!("Actor {i}"))).collect()
}
#[test]
fn full_subset_of_the_cast_is_listed() {
// §6: the ratio is over *M*, not *C* — a manifest legitimately contains
// only actors both credited and detected on screen, so penalising it for
// missing credited actors would fail every honest upload.
let submitted: Vec<_> = (1..=7).map(by_id).collect();
let out = evaluate(&submitted, &cast_of_20());
assert_eq!(out.verdict, Verdict::Listed);
assert_eq!(out.ratio, 1.0);
assert_eq!(out.matched.len(), 7);
}
#[test]
fn threshold_boundaries_at_point_six_and_point_three() {
// 6 of 10 matching == 0.6 exactly: listed.
let mut submitted: Vec<_> = (1..=6).map(by_id).collect();
submitted.extend((900..904).map(by_id));
let out = evaluate(&submitted, &cast_of_20());
assert_eq!(out.matched.len(), 6);
assert!((out.ratio - 0.6).abs() < 1e-9);
assert_eq!(out.verdict, Verdict::Listed);
// 5 of 10 == 0.5: flagged, served with reduced ranking.
let mut submitted: Vec<_> = (1..=5).map(by_id).collect();
submitted.extend((900..905).map(by_id));
let out = evaluate(&submitted, &cast_of_20());
assert_eq!(out.verdict, Verdict::Flagged);
// 3 of 10 == 0.3 exactly: still flagged, not rejected.
let mut submitted: Vec<_> = (1..=3).map(by_id).collect();
submitted.extend((900..907).map(by_id));
let out = evaluate(&submitted, &cast_of_20());
assert_eq!(out.verdict, Verdict::Flagged);
// 2 of 10 == 0.2: rejected.
let mut submitted: Vec<_> = (1..=2).map(by_id).collect();
submitted.extend((900..908).map(by_id));
let out = evaluate(&submitted, &cast_of_20());
assert_eq!(out.verdict, Verdict::Rejected);
}
#[test]
fn prank_manifest_is_rejected() {
// §5a Threat 2: performers who are not credited cast on the title.
let submitted: Vec<_> = (500..510).map(by_id).collect();
let out = evaluate(&submitted, &cast_of_20());
assert_eq!(out.verdict, Verdict::Rejected);
assert_eq!(out.ratio, 0.0);
assert_eq!(out.reason.as_deref(), Some("cast_match_below_threshold"));
}
#[test]
fn small_m_requires_all_but_one_to_match() {
// §6: `2 <= |M| < 5` — a ratio is meaningless at this size.
let out = evaluate(&[by_id(1), by_id(2), by_id(3), by_id(999)], &cast_of_20());
assert_eq!(out.verdict, Verdict::Listed, "3 of 4 is all-but-one");
let out = evaluate(&[by_id(1), by_id(2), by_id(998), by_id(999)], &cast_of_20());
assert_eq!(out.verdict, Verdict::Rejected, "2 of 4 fails all-but-one");
// 0.5 would be `Flagged` under the ratio table, so this proves the
// small-|M| branch is actually taken.
let out = evaluate(&[by_id(1), by_id(999)], &cast_of_20());
assert_eq!(out.verdict, Verdict::Listed, "1 of 2 is all-but-one");
}
#[test]
fn single_actor_manifest_needs_that_actor_to_match() {
assert_eq!(evaluate(&[by_id(1)], &cast_of_20()).verdict, Verdict::Listed);
assert_eq!(evaluate(&[by_id(999)], &cast_of_20()).verdict, Verdict::Rejected);
}
#[test]
fn empty_manifest_is_rejected() {
let out = evaluate(&[], &cast_of_20());
assert_eq!(out.verdict, Verdict::Rejected);
assert_eq!(out.reason.as_deref(), Some("empty_actor_list"));
}
#[test]
fn missing_tmdb_credits_flags_rather_than_rejects() {
// §6: absent data is not evidence of a bad manifest.
let submitted: Vec<_> = (1..=7).map(by_id).collect();
let out = evaluate(&submitted, &[]);
assert_eq!(out.verdict, Verdict::Flagged);
assert_eq!(out.reason.as_deref(), Some("tmdb_no_credits"));
}
#[test]
fn name_matching_is_case_and_accent_insensitive() {
let credits = vec![credit(1, "Renée Zellweger"), credit(2, "Miloš Forman")];
let out = evaluate(&[by_name("renee zellweger"), by_name("MILOS FORMAN")], &credits);
assert_eq!(out.matched.len(), 2);
}
#[test]
fn name_only_matches_cannot_carry_a_manifest_alone() {
// §6: name-only matches are capped at half the intersection, so a
// manifest cannot pass on name collisions alone.
let credits: Vec<_> = (1..=20).map(|i| credit(i, &format!("Actor {i}"))).collect();
let submitted: Vec<_> = (1..=10).map(|i| by_name(&format!("Actor {i}"))).collect();
let out = evaluate(&submitted, &credits);
assert_eq!(out.ratio, 0.0, "with no id matches, name matches cap to zero");
assert_eq!(out.verdict, Verdict::Rejected);
}
#[test]
fn name_matches_count_up_to_the_number_of_id_matches() {
let credits: Vec<_> = (1..=20).map(|i| credit(i, &format!("Actor {i}"))).collect();
// 4 by id + 6 by name, of 10 => capped to 4 + 4 = 8 => 0.8.
let mut submitted: Vec<_> = (1..=4).map(by_id).collect();
submitted.extend((5..=10).map(|i| by_name(&format!("Actor {i}"))));
let out = evaluate(&submitted, &credits);
assert!((out.ratio - 0.8).abs() < 1e-9, "got {}", out.ratio);
assert_eq!(out.verdict, Verdict::Listed);
}
#[test]
fn unmatched_actors_are_reported_for_dropping() {
// §6: unmatched actors are dropped rather than stored.
let submitted: Vec<_> = (1..=6).map(by_id).chain([by_id(777)]).collect();
let out = evaluate(&submitted, &cast_of_20());
assert_eq!(out.unmatched_person_ids, vec![777]);
assert!(out.matched.iter().all(|m| m.tmdb_person_id != 777));
}
#[test]
fn matched_names_come_from_tmdb_not_the_upload() {
// §5a: the server stores references to TMDB entities, not
// attacker-authored text.
let credits = vec![credit(884, "Steve Buscemi")];
let submitted = vec![SubmittedActor {
tmdb_id: Some(884),
imdb_id: None,
name: Some("Definitely Not Him".into()),
}];
let out = evaluate(&submitted, &credits);
assert_eq!(out.matched[0].name, "Steve Buscemi");
}
#[test]
fn duplicate_credits_do_not_double_count() {
// TMDB aggregate credits can list a person more than once.
let credits = vec![credit(1, "A"), credit(1, "A")];
let out = evaluate(&[by_id(1)], &credits);
assert_eq!(out.matched.len(), 1);
}
#[test]
fn category_guard_catches_adult_performers_on_a_non_adult_title() {
// §5a layer 1, aimed squarely at the stated prank.
let matched = vec![
MatchedActor { tmdb_person_id: 1, name: "A".into(), adult: false, by_name: false },
MatchedActor { tmdb_person_id: 2, name: "B".into(), adult: true, by_name: false },
];
assert_eq!(category_guard_violation(&matched, false), Some(2));
// Unless the target title is itself flagged adult.
assert_eq!(category_guard_violation(&matched, true), None);
}
#[test]
fn category_guard_ignores_clean_casts() {
let matched = vec![MatchedActor {
tmdb_person_id: 1,
name: "A".into(),
adult: false,
by_name: false,
}];
assert_eq!(category_guard_violation(&matched, false), None);
}
#[test]
fn adult_credit_is_carried_through_matching() {
let out = evaluate(&[by_id(9)], &[adult_credit(9, "X")]);
assert!(out.matched[0].adult);
}
#[test]
fn childrens_certifications_are_recognised() {
for c in ["G", "TV-Y", "tv-y7", "U", " PG "] {
assert!(is_childrens_certification(c), "{c} should be a children's rating");
}
for c in ["R", "NC-17", "TV-MA", "18", ""] {
assert!(!is_childrens_certification(c), "{c} should not be");
}
}
#[test]
fn childrens_guard_tightens_the_threshold() {
// §5a layer 2: the highest-harm case gets the tightest gate. A ratio of
// 0.7 lists normally but only reaches `flagged` on a children's title.
let mut submitted: Vec<_> = (1..=7).map(by_id).collect();
submitted.extend((900..903).map(by_id));
let mut out = evaluate(&submitted, &cast_of_20());
assert_eq!(out.verdict, Verdict::Listed);
let m = submitted.len();
apply_childrens_guard(&mut out, m);
assert_eq!(out.verdict, Verdict::Flagged);
assert_eq!(out.reason.as_deref(), Some("childrens_title_strict_threshold"));
}
#[test]
fn childrens_guard_leaves_strong_matches_listed() {
let submitted: Vec<_> = (1..=10).map(by_id).collect();
let mut out = evaluate(&submitted, &cast_of_20());
let m = submitted.len();
apply_childrens_guard(&mut out, m);
assert_eq!(out.verdict, Verdict::Listed);
}
}
+62
View File
@@ -0,0 +1,62 @@
//! Operational configuration, read from the environment.
//!
//! The trusted-proxy CIDR is explicit configuration rather than a default-on
//! behaviour (§8 deployment notes): §5 rate limiting and report attribution key
//! on client IP, so an unconditionally-trusted `X-Forwarded-For` defeats both.
use std::net::IpAddr;
use std::time::Duration;
#[derive(Clone, Debug)]
pub struct Config {
pub bind: String,
pub db_path: String,
/// Hard dependency for UR-3. Without it, uploads accumulate in `pending`
/// rather than being listed unverified (§8).
pub tmdb_api_key: Option<String>,
pub tmdb_base_url: String,
/// Prefixes of proxy addresses whose `X-Forwarded-For` is honoured.
pub trusted_proxies: Vec<IpAddr>,
pub server_id: String,
pub request_timeout: Duration,
/// Number of cast-check jobs to lease per worker tick.
pub job_batch: usize,
pub job_poll_interval: Duration,
}
impl Config {
pub fn from_env() -> anyhow::Result<Self> {
let trusted_proxies = match std::env::var("JRAY_TRUSTED_PROXIES") {
Ok(v) => v
.split(',')
.map(str::trim)
.filter(|s| !s.is_empty())
.map(|s| {
s.parse::<IpAddr>()
.map_err(|e| anyhow::anyhow!("bad JRAY_TRUSTED_PROXIES entry {s:?}: {e}"))
})
.collect::<Result<Vec<_>, _>>()?,
Err(_) => Vec::new(),
};
Ok(Self {
bind: env_or("JRAY_BIND", "127.0.0.1:8080"),
db_path: env_or("JRAY_DB", "jray.db"),
tmdb_api_key: std::env::var("JRAY_TMDB_API_KEY").ok().filter(|s| !s.is_empty()),
tmdb_base_url: env_or("JRAY_TMDB_BASE_URL", "https://api.themoviedb.org/3"),
trusted_proxies,
server_id: env_or("JRAY_SERVER_ID", "localhost"),
request_timeout: Duration::from_secs(env_num("JRAY_REQUEST_TIMEOUT_SEC", 30)),
job_batch: env_num("JRAY_JOB_BATCH", 8) as usize,
job_poll_interval: Duration::from_secs(env_num("JRAY_JOB_POLL_SEC", 5)),
})
}
}
fn env_or(key: &str, default: &str) -> String {
std::env::var(key).ok().filter(|s| !s.is_empty()).unwrap_or_else(|| default.to_string())
}
fn env_num(key: &str, default: u64) -> u64 {
std::env::var(key).ok().and_then(|v| v.parse().ok()).unwrap_or(default)
}
+289
View File
@@ -0,0 +1,289 @@
//! §9a content addressing.
//!
//! A validated manifest is immutable and content-addressable, which is what
//! makes replication *set reconciliation* rather than state synchronisation.
//! Even without the federation endpoints, computing `content_id` on upload gives
//! deduplication now and means stored manifests are already addressable when
//! federation lands.
//!
//! **This canonical form must be reimplemented byte-identically by the JRay
//! plugin** (§8 "Cost of choosing Rust": the extraction side is Python, so this
//! can no longer be shared as one implementation and must instead be specified
//! precisely and cross-tested). [`GOLDEN_VECTORS`] is that shared fixture.
use sha2::{Digest, Sha256};
/// One actor's contribution to the canonical form.
#[derive(Debug, Clone)]
pub struct CanonicalActor {
pub tmdb_person_id: u64,
/// Integer centiseconds — quantised, not formatted floats (§9a).
pub scenes_cs: Vec<(i64, i64)>,
}
/// The identity coordinates that enter the hash.
#[derive(Debug, Clone, Default)]
pub struct CanonicalIdentity {
pub kind: &'static str,
pub tmdb_id: Option<String>,
pub imdb_id: Option<String>,
pub season: Option<i64>,
pub episode: Option<i64>,
}
/// The cut coordinates that enter the hash.
///
/// **`audio_signature` is excluded, deliberately** (§9a): it is derived by
/// decoding audio, so two servers running different FFmpeg or resampler versions
/// could compute marginally different signatures for identical content, and
/// including it would silently break federation deduplication.
#[derive(Debug, Clone, Default)]
pub struct CanonicalCut {
/// Quantised to centiseconds for the same reason scene times are.
pub runtime_cs: i64,
pub video_hash: Option<String>,
}
/// Builds the canonical JSON form: keys sorted, no whitespace, actors sorted by
/// person id, scene times as integer centiseconds.
///
/// `extraction` metadata and all local state are excluded, so two servers that
/// validated the same upload independently arrive at the same `content_id`.
pub fn canonical_json(
identity: &CanonicalIdentity,
cut: &CanonicalCut,
actors: &[CanonicalActor],
) -> String {
let mut sorted: Vec<&CanonicalActor> = actors.iter().collect();
sorted.sort_by_key(|a| a.tmdb_person_id);
let mut s = String::new();
s.push_str("{\"actors\":[");
for (i, a) in sorted.iter().enumerate() {
if i > 0 {
s.push(',');
}
// Scene windows are emitted in stored order; validation has already
// established they are sorted by start time.
s.push_str("{\"scenes\":[");
for (j, (start, end)) in a.scenes_cs.iter().enumerate() {
if j > 0 {
s.push(',');
}
s.push('[');
s.push_str(&start.to_string());
s.push(',');
s.push_str(&end.to_string());
s.push(']');
}
s.push_str("],\"tmdb_person_id\":");
s.push_str(&a.tmdb_person_id.to_string());
s.push('}');
}
s.push_str("],\"cut\":{");
s.push_str("\"runtime_cs\":");
s.push_str(&cut.runtime_cs.to_string());
s.push_str(",\"video_hash\":");
push_opt_str(&mut s, cut.video_hash.as_deref());
s.push_str("},\"identity\":{");
s.push_str("\"episode\":");
push_opt_num(&mut s, cut_opt(identity.episode));
s.push_str(",\"imdb_id\":");
push_opt_str(&mut s, identity.imdb_id.as_deref());
s.push_str(",\"season\":");
push_opt_num(&mut s, cut_opt(identity.season));
s.push_str(",\"tmdb_id\":");
push_opt_str(&mut s, identity.tmdb_id.as_deref());
s.push_str(",\"type\":\"");
s.push_str(identity.kind);
s.push_str("\"}}");
s
}
fn cut_opt(v: Option<i64>) -> Option<i64> {
v
}
fn push_opt_str(s: &mut String, v: Option<&str>) {
match v {
// Only closed-vocabulary values reach here (regex-constrained ids and a
// fixed-format hash), so no string escaping is required.
Some(v) => {
s.push('"');
s.push_str(v);
s.push('"');
}
None => s.push_str("null"),
}
}
fn push_opt_num(s: &mut String, v: Option<i64>) {
match v {
Some(v) => s.push_str(&v.to_string()),
None => s.push_str("null"),
}
}
/// `sha256:` over the canonical form (§9a).
pub fn content_id(
identity: &CanonicalIdentity,
cut: &CanonicalCut,
actors: &[CanonicalActor],
) -> String {
let canonical = canonical_json(identity, cut, actors);
let mut h = Sha256::new();
h.update(canonical.as_bytes());
let digest = h.finalize();
let mut hex = String::with_capacity(64 + 7);
hex.push_str("sha256:");
for b in digest {
hex.push_str(&format!("{b:02x}"));
}
hex
}
/// Cross-implementation fixture (§8): the JRay plugin and any reimplementation
/// must reproduce these exactly, or federation deduplication silently breaks.
pub const GOLDEN_VECTORS: &[(&str, &str)] = &[(
// Movie, one actor, two windows, with a video hash.
r#"{"actors":[{"scenes":[[19160,20920],[43820,46560]],"tmdb_person_id":884}],"cut":{"runtime_cs":642050,"video_hash":"opensubtitles:8e245d9679d31e12"},"identity":{"episode":null,"imdb_id":"tt4686844","season":null,"tmdb_id":"504172","type":"movie"}}"#,
// Verified against an independent Python implementation:
// sha256(canonical.encode()).hexdigest()
"sha256:367f8b05c54a992a3a30fa016edaaac0b9b36b148fc76574f5b1ef326b56760f",
)];
#[cfg(test)]
mod tests {
use super::*;
fn movie_identity() -> CanonicalIdentity {
CanonicalIdentity {
kind: "movie",
tmdb_id: Some("504172".into()),
imdb_id: Some("tt4686844".into()),
season: None,
episode: None,
}
}
fn movie_cut() -> CanonicalCut {
CanonicalCut {
runtime_cs: 642050,
video_hash: Some("opensubtitles:8e245d9679d31e12".into()),
}
}
fn actors() -> Vec<CanonicalActor> {
vec![CanonicalActor {
tmdb_person_id: 884,
scenes_cs: vec![(19160, 20920), (43820, 46560)],
}]
}
#[test]
fn canonical_form_matches_the_documented_shape() {
let json = canonical_json(&movie_identity(), &movie_cut(), &actors());
assert_eq!(json, GOLDEN_VECTORS[0].0);
// Keys sorted, no whitespace (§9a).
assert!(!json.contains(' '));
}
#[test]
fn canonical_form_is_valid_json_with_sorted_keys() {
// Hand-built strings are easy to get subtly wrong, so assert the output
// actually parses and that its keys really are ordered.
let json = canonical_json(&movie_identity(), &movie_cut(), &actors());
let v: serde_json::Value =
serde_json::from_str(&json).expect("canonical form must be JSON");
let obj = v.as_object().unwrap();
let keys: Vec<&String> = obj.keys().collect();
assert_eq!(keys, vec!["actors", "cut", "identity"]);
let id_keys: Vec<&String> = v["identity"].as_object().unwrap().keys().collect();
assert_eq!(id_keys, vec!["episode", "imdb_id", "season", "tmdb_id", "type"]);
let cut_keys: Vec<&String> = v["cut"].as_object().unwrap().keys().collect();
assert_eq!(cut_keys, vec!["runtime_cs", "video_hash"]);
}
#[test]
fn actor_order_does_not_affect_the_hash() {
// §9a: actors sorted by person id, so two servers that stored them in
// different orders still agree.
let a = vec![
CanonicalActor { tmdb_person_id: 884, scenes_cs: vec![(0, 100)] },
CanonicalActor { tmdb_person_id: 17419, scenes_cs: vec![(200, 300)] },
];
let b = vec![a[1].clone(), a[0].clone()];
assert_eq!(
content_id(&movie_identity(), &movie_cut(), &a),
content_id(&movie_identity(), &movie_cut(), &b)
);
}
#[test]
fn accumulated_float_error_hashes_identically() {
// The failure mode §9a exists to remove: real corpus values look like
// 8045.066666660665, and two servers may compute them slightly
// differently. Quantising first means both hash the same.
let a = vec![CanonicalActor {
tmdb_person_id: 1,
scenes_cs: vec![(crate::validate::to_centiseconds(8045.066666660665), 900000)],
}];
let b = vec![CanonicalActor {
tmdb_person_id: 1,
scenes_cs: vec![(crate::validate::to_centiseconds(8045.066666666), 900000)],
}];
assert_eq!(
content_id(&movie_identity(), &movie_cut(), &a),
content_id(&movie_identity(), &movie_cut(), &b)
);
}
#[test]
fn differing_content_produces_differing_ids() {
let base = content_id(&movie_identity(), &movie_cut(), &actors());
let mut other_actors = actors();
other_actors[0].scenes_cs[0].1 += 1;
assert_ne!(base, content_id(&movie_identity(), &movie_cut(), &other_actors));
let mut other_cut = movie_cut();
other_cut.runtime_cs += 1;
assert_ne!(base, content_id(&movie_identity(), &other_cut, &actors()));
let mut other_id = movie_identity();
other_id.tmdb_id = Some("999".into());
assert_ne!(base, content_id(&other_id, &movie_cut(), &actors()));
}
#[test]
fn episode_and_movie_coordinates_are_distinguished() {
let ep = CanonicalIdentity {
kind: "episode",
tmdb_id: Some("1396".into()),
imdb_id: None,
season: Some(2),
episode: Some(5),
};
let other = CanonicalIdentity { season: Some(3), ..ep.clone() };
assert_ne!(
content_id(&ep, &movie_cut(), &actors()),
content_id(&other, &movie_cut(), &actors())
);
}
#[test]
fn content_id_is_prefixed_and_hex() {
let id = content_id(&movie_identity(), &movie_cut(), &actors());
let hex = id.strip_prefix("sha256:").expect("prefixed");
assert_eq!(hex.len(), 64);
assert!(hex.bytes().all(|b| b.is_ascii_hexdigit()));
}
#[test]
fn golden_vector_hash_is_stable() {
// Locks the hash so an accidental change to the canonical form is caught
// here rather than by silent federation divergence.
let id = content_id(&movie_identity(), &movie_cut(), &actors());
assert_eq!(id, GOLDEN_VECTORS[0].1, "canonical form or hash changed");
}
}
+168
View File
@@ -0,0 +1,168 @@
//! Database access.
//!
//! §8 imposes two structural requirements that this module exists to satisfy:
//!
//! 1. **A single writer connection, serialized through one owner**, with a read
//! pool alongside. SQLite permits only one writer at a time even in WAL mode;
//! pointing a multi-connection pool at writes and relying on `busy_timeout`
//! to sort it out is explicitly rejected by the spec. Here the writer lives
//! behind a `Mutex`, so contention queues in Rust rather than surfacing as
//! `SQLITE_BUSY`.
//! 2. **All access behind a thin repository layer** rather than queries
//! scattered through handlers — this is what keeps the Turso/Postgres options
//! cheap and localises the serialization in one place.
//!
//! rusqlite is synchronous, so every call is wrapped in `spawn_blocking`: a
//! write that waits on the mutex must never block a Tokio worker thread.
pub mod repo;
use std::sync::{Arc, Mutex};
use anyhow::Context;
use rusqlite::Connection;
const SCHEMA: &str = include_str!("schema.sql");
/// Handle to the database: one serialized writer, plus read connections.
///
/// Cloning is cheap and shares the same underlying connections.
#[derive(Clone)]
pub struct Db {
writer: Arc<Mutex<Connection>>,
readers: Arc<ReadPool>,
}
struct ReadPool {
conns: Mutex<Vec<Connection>>,
path: String,
}
impl ReadPool {
fn acquire(&self) -> anyhow::Result<Connection> {
if let Some(c) = self.conns.lock().expect("read pool poisoned").pop() {
return Ok(c);
}
open_conn(&self.path, false)
}
fn release(&self, conn: Connection) {
let mut conns = self.conns.lock().expect("read pool poisoned");
// Bounded: excess connections are dropped rather than accumulating.
if conns.len() < 8 {
conns.push(conn);
}
}
}
fn open_conn(path: &str, writer: bool) -> anyhow::Result<Connection> {
let conn = Connection::open(path).with_context(|| format!("opening database {path}"))?;
// WAL gives concurrent readers alongside the single writer, which suits a
// read-dominated workload; `synchronous = NORMAL` is safe under WAL, and
// `busy_timeout` makes contention wait rather than error (§8).
conn.pragma_update(None, "journal_mode", "WAL")?;
conn.pragma_update(None, "synchronous", "NORMAL")?;
conn.pragma_update(None, "busy_timeout", 5_000)?;
conn.pragma_update(None, "foreign_keys", true)?;
if !writer {
conn.pragma_update(None, "query_only", true)?;
}
Ok(conn)
}
impl Db {
/// Opens the database, applying the schema. Idempotent — every statement in
/// `schema.sql` is `IF NOT EXISTS`.
pub fn open(path: &str) -> anyhow::Result<Self> {
let writer = open_conn(path, true)?;
writer.execute_batch(SCHEMA).context("applying schema")?;
Ok(Self {
writer: Arc::new(Mutex::new(writer)),
readers: Arc::new(ReadPool { conns: Mutex::new(Vec::new()), path: path.to_string() }),
})
}
/// Runs `f` against the serialized writer connection on a blocking thread.
///
/// `f` receives a `Transaction`, so a manifest's scene rows go in as one
/// transaction rather than one per row (§8), and a failure rolls back.
pub async fn write<T, F>(&self, f: F) -> anyhow::Result<T>
where
T: Send + 'static,
F: FnOnce(&rusqlite::Transaction<'_>) -> anyhow::Result<T> + Send + 'static,
{
let writer = self.writer.clone();
tokio::task::spawn_blocking(move || {
let mut conn = writer.lock().expect("writer poisoned");
let tx = conn.transaction()?;
let out = f(&tx)?;
tx.commit()?;
Ok(out)
})
.await
.context("writer task panicked")?
}
/// Runs `f` against a read connection on a blocking thread.
pub async fn read<T, F>(&self, f: F) -> anyhow::Result<T>
where
T: Send + 'static,
F: FnOnce(&Connection) -> anyhow::Result<T> + Send + 'static,
{
let readers = self.readers.clone();
tokio::task::spawn_blocking(move || {
let conn = readers.acquire()?;
let out = f(&conn);
readers.release(conn);
out
})
.await
.context("reader task panicked")?
}
}
#[cfg(test)]
mod tests {
use super::*;
#[tokio::test]
async fn schema_applies_and_roundtrips() {
let db = Db::open(":memory:").unwrap();
// In-memory databases are per-connection, so only exercise the writer.
let n = db
.write(|tx| {
tx.execute(
"INSERT INTO contributors (id, token_hash, created_at) VALUES (?1, ?2, ?3)",
rusqlite::params!["c1", "hash", "2026-01-01T00:00:00Z"],
)?;
Ok(tx.query_row("SELECT COUNT(*) FROM contributors", [], |r| r.get::<_, i64>(0))?)
})
.await
.unwrap();
assert_eq!(n, 1);
}
#[tokio::test]
async fn write_rolls_back_on_error() {
let db = Db::open(":memory:").unwrap();
let res: anyhow::Result<()> = db
.write(|tx| {
tx.execute(
"INSERT INTO contributors (id, token_hash, created_at) VALUES ('c1','h','t')",
[],
)?;
anyhow::bail!("deliberate failure")
})
.await;
assert!(res.is_err());
let n = db
.write(|tx| {
Ok(tx.query_row("SELECT COUNT(*) FROM contributors", [], |r| r.get::<_, i64>(0))?)
})
.await
.unwrap();
assert_eq!(n, 0, "failed transaction must not persist rows");
}
}
+1110
View File
File diff suppressed because it is too large Load Diff
+119
View File
@@ -0,0 +1,119 @@
-- §7 Storage. Fully relational, no JSON blobs on the write path: the database
-- can only represent what the schema models, so there is physically nowhere for
-- an unexpected field or a smuggled string to live (§5a Threat 1).
--
-- Portable SQL — runs unchanged on Postgres. Avoid SQLite-specific forms
-- (`INSERT OR REPLACE`); use `INSERT ... ON CONFLICT` (§8 deployment notes).
CREATE TABLE IF NOT EXISTS contributors (
id TEXT PRIMARY KEY,
token_hash TEXT NOT NULL UNIQUE,
created_at TEXT NOT NULL,
revoked_at TEXT,
accepted_count INTEGER NOT NULL DEFAULT 0,
rejected_count INTEGER NOT NULL DEFAULT 0,
flagged_count INTEGER NOT NULL DEFAULT 0
);
-- Server-side, TMDB-derived. `name` never comes from an upload (§5a).
CREATE TABLE IF NOT EXISTS people (
tmdb_person_id INTEGER PRIMARY KEY,
name TEXT NOT NULL,
adult INTEGER NOT NULL DEFAULT 0,
updated_at TEXT NOT NULL
);
CREATE TABLE IF NOT EXISTS titles (
id TEXT PRIMARY KEY,
kind TEXT NOT NULL, -- movie | series
tmdb_id TEXT,
imdb_id TEXT,
name TEXT,
year INTEGER,
adult INTEGER NOT NULL DEFAULT 0,
certification TEXT,
updated_at TEXT NOT NULL
);
CREATE TABLE IF NOT EXISTS manifests (
id TEXT PRIMARY KEY,
title_id TEXT NOT NULL REFERENCES titles(id),
season INTEGER,
episode INTEGER,
runtime_sec REAL NOT NULL,
video_hash TEXT,
audio_signature BLOB, -- §3, ~1290 bytes
audio_sig_coarse BLOB, -- candidate-generation index key
sample_fps REAL,
extinction_sec REAL, -- successor to the withdrawn anneal_sec
gallery_scope TEXT, -- limited | global; ranking signal (§2, §7)
pipeline_version TEXT,
contributor_id TEXT REFERENCES contributors(id),
status TEXT NOT NULL, -- pending | listed | flagged | rejected
reject_reason TEXT,
cast_match_ratio REAL,
content_id TEXT UNIQUE, -- §9a, sha256 over canonical form
origin TEXT, -- server_id of first acceptance
ingested_from TEXT, -- peer id, NULL if uploaded directly
created_at TEXT NOT NULL
);
CREATE TABLE IF NOT EXISTS manifest_actors (
manifest_id TEXT NOT NULL REFERENCES manifests(id) ON DELETE CASCADE,
tmdb_person_id INTEGER NOT NULL,
PRIMARY KEY (manifest_id, tmdb_person_id)
);
-- Integer centiseconds, not floats — the same quantisation used for
-- `content_id`, so stored values and hashed values cannot diverge (§7, §9a).
CREATE TABLE IF NOT EXISTS scenes (
manifest_id TEXT NOT NULL REFERENCES manifests(id) ON DELETE CASCADE,
tmdb_person_id INTEGER NOT NULL,
start_cs INTEGER NOT NULL,
end_cs INTEGER NOT NULL
);
CREATE TABLE IF NOT EXISTS reports (
id TEXT PRIMARY KEY,
manifest_id TEXT NOT NULL REFERENCES manifests(id) ON DELETE CASCADE,
reason TEXT NOT NULL,
note TEXT,
created_at TEXT NOT NULL,
source_ip_hash TEXT
);
-- The sole JSON column, and it holds TMDB's responses, not users' (§7).
CREATE TABLE IF NOT EXISTS tmdb_cache (
tmdb_id TEXT NOT NULL,
kind TEXT NOT NULL,
credits TEXT NOT NULL,
fetched_at TEXT NOT NULL,
PRIMARY KEY (tmdb_id, kind)
);
-- Background queue as a table rather than an external broker, so pending work
-- survives a restart (§7, §8).
CREATE TABLE IF NOT EXISTS jobs (
id TEXT PRIMARY KEY,
kind TEXT NOT NULL, -- cast_check | federation_pull
payload TEXT NOT NULL,
run_after TEXT NOT NULL,
attempts INTEGER NOT NULL DEFAULT 0,
last_error TEXT,
leased_at TEXT
);
CREATE INDEX IF NOT EXISTS idx_titles_tmdb ON titles(tmdb_id);
CREATE INDEX IF NOT EXISTS idx_titles_imdb ON titles(imdb_id);
CREATE INDEX IF NOT EXISTS idx_manifests_title_runtime ON manifests(title_id, runtime_sec);
CREATE INDEX IF NOT EXISTS idx_manifests_video_hash ON manifests(video_hash);
CREATE INDEX IF NOT EXISTS idx_manifests_episode ON manifests(title_id, season, episode);
CREATE INDEX IF NOT EXISTS idx_scenes_manifest_person ON scenes(manifest_id, tmdb_person_id);
-- All read queries filter `status IN ('listed','flagged')`, so a partial index
-- on that predicate keeps the hot path small (§7).
CREATE INDEX IF NOT EXISTS idx_manifests_served
ON manifests(title_id, season, episode)
WHERE status IN ('listed', 'flagged');
CREATE INDEX IF NOT EXISTS idx_jobs_ready ON jobs(run_after);
+83
View File
@@ -0,0 +1,83 @@
//! API error type mapping onto the status codes §4 specifies.
use axum::http::StatusCode;
use axum::response::{IntoResponse, Response};
use axum::Json;
use serde::Serialize;
#[derive(Debug, thiserror::Error)]
pub enum ApiError {
/// §6 stage 2 — malformed, unrecognised or forbidden field. The message
/// names the offending field so a client that forgets to strip `movie` or
/// `jellyfin_id` gets a hard, diagnosable `400` (§6).
#[error("{0}")]
BadRequest(String),
#[error("not found")]
NotFound,
/// §4 — identical `(identity, cut)` already exists from this contributor.
#[error("{0}")]
Conflict(String),
#[error("{0}")]
PayloadTooLarge(String),
#[error("missing or invalid API token")]
Unauthorized,
/// §5 — carries the `Retry-After` value in seconds.
#[error("rate limited")]
RateLimited { retry_after: u64 },
#[error("internal error")]
Internal(#[from] anyhow::Error),
}
#[derive(Serialize)]
struct ErrorBody {
error: String,
message: String,
}
impl IntoResponse for ApiError {
fn into_response(self) -> Response {
let (status, code) = match &self {
ApiError::BadRequest(_) => (StatusCode::BAD_REQUEST, "bad_request"),
ApiError::NotFound => (StatusCode::NOT_FOUND, "not_found"),
ApiError::Conflict(_) => (StatusCode::CONFLICT, "conflict"),
ApiError::PayloadTooLarge(_) => (StatusCode::PAYLOAD_TOO_LARGE, "payload_too_large"),
ApiError::Unauthorized => (StatusCode::UNAUTHORIZED, "unauthorized"),
ApiError::RateLimited { .. } => (StatusCode::TOO_MANY_REQUESTS, "rate_limited"),
ApiError::Internal(e) => {
// Internal detail is logged, never returned.
tracing::error!(error = ?e, "internal error");
(StatusCode::INTERNAL_SERVER_ERROR, "internal")
}
};
let body = Json(ErrorBody {
error: code.to_string(),
message: match &self {
ApiError::Internal(_) => "internal error".to_string(),
other => other.to_string(),
},
});
let mut resp = (status, body).into_response();
if let ApiError::RateLimited { retry_after } = self {
if let Ok(v) = retry_after.to_string().parse() {
resp.headers_mut().insert(axum::http::header::RETRY_AFTER, v);
}
}
resp
}
}
impl From<rusqlite::Error> for ApiError {
fn from(e: rusqlite::Error) -> Self {
ApiError::Internal(anyhow::Error::new(e))
}
}
pub type ApiResult<T> = Result<T, ApiError>;
+403
View File
@@ -0,0 +1,403 @@
//! Manifest ingestion: the shared path behind `POST /manifests` and
//! `POST /manifests/bundle`, and the path a federation pull will reuse (§9a
//! "re-derive, don't inherit").
//!
//! Stages 0 and 1 are layers; stage 2 is parse + [`crate::validate`]. What
//! happens here is persistence plus enqueueing the stage 3 check: the upload is
//! accepted with `202` and the manifest is held **unlisted** until the cast check
//! completes — it is not served to anyone in the meantime (§6).
use anyhow::Context;
use crate::content_id::{self, CanonicalActor, CanonicalCut, CanonicalIdentity};
use crate::db::repo::{self, NewManifest};
use crate::model::IdentityType;
use crate::validate::ValidManifest;
/// Outcome of persisting one manifest.
#[derive(Debug, Clone)]
pub enum IngestOutcome {
/// Held unlisted pending the §6 stage 3 cast check.
Pending { manifest_id: String },
/// §4 `409` — identical `(identity, cut)` from this contributor.
DuplicateFromContributor { manifest_id: String },
/// §9a — the exact same content is already held, from any source. Skipped
/// without re-validation, which is the deduplication content addressing buys.
DuplicateContent { manifest_id: String },
}
impl IngestOutcome {
pub fn manifest_id(&self) -> &str {
match self {
IngestOutcome::Pending { manifest_id }
| IngestOutcome::DuplicateFromContributor { manifest_id }
| IngestOutcome::DuplicateContent { manifest_id } => manifest_id,
}
}
}
/// Job payload for the §6 stage 3 check.
#[derive(Debug, Clone, serde::Serialize, serde::Deserialize)]
pub struct CastCheckJob {
pub manifest_id: String,
}
pub const JOB_CAST_CHECK: &str = "cast_check";
/// Persists a validated manifest and enqueues its cast check, all in one
/// transaction — so a manifest is never left listed-but-unchecked, and its scene
/// rows go in as a single transaction rather than one per row (§8).
pub fn persist(
tx: &rusqlite::Transaction<'_>,
valid: &ValidManifest,
contributor_id: Option<&str>,
origin: &str,
ingested_from: Option<&str>,
now: &str,
) -> anyhow::Result<IngestOutcome> {
let m = &valid.manifest;
let kind = m.identity.kind;
let tmdb_id = m.identity.effective_tmdb_id();
let imdb_id = m.identity.effective_imdb_id();
let title_id = repo::upsert_title(
tx,
kind,
tmdb_id,
imdb_id,
m.identity.title.as_deref(),
m.identity.year,
now,
)
.context("resolving title")?;
let (season, episode) = match kind {
IdentityType::Movie => (None, None),
IdentityType::Episode => (m.identity.season, m.identity.episode),
};
// Content addressing over the *submitted* actor ids. Recomputed after the
// cast check drops unmatched actors, since dropping changes the content.
let cid = compute_content_id(valid);
if let Some(existing) = repo::manifest_by_content_id(tx, &cid)? {
return Ok(IngestOutcome::DuplicateContent { manifest_id: existing });
}
if let Some(c) = contributor_id {
if let Some(existing) = repo::duplicate_from_contributor(
tx,
&title_id,
season,
episode,
m.cut.runtime_sec,
m.cut.video_hash.as_deref(),
c,
)? {
return Ok(IngestOutcome::DuplicateFromContributor { manifest_id: existing });
}
}
let manifest_id = ulid::Ulid::new().to_string();
let extraction = m.extraction.as_ref();
repo::insert_manifest(
tx,
&NewManifest {
id: &manifest_id,
title_id: &title_id,
season,
episode,
runtime_sec: m.cut.runtime_sec,
video_hash: m.cut.video_hash.as_deref(),
// Stored as an attribute, not part of identity (§9a).
audio_signature: None,
audio_sig_coarse: None,
sample_fps: extraction.and_then(|e| e.sample_fps),
extinction_sec: extraction.and_then(|e| e.extinction_sec),
pipeline_version: extraction.and_then(|e| e.pipeline_version.as_deref()),
gallery_scope: extraction.and_then(|e| e.gallery_scope).map(|g| g.as_str()),
contributor_id,
// Held unlisted until stage 3 completes (§6).
status: "pending",
content_id: Some(&cid),
origin,
ingested_from,
created_at: now,
},
)?;
// Actors are recorded by TMDB person id only. Those without one cannot be
// stored at all — there is no name column to put them in (§5a, §7) — so they
// are carried into the cast check via the submitted payload instead.
for actor in &valid.actor_scenes_cs {
if let Some(person_id) = actor.tmdb_id {
repo::insert_actor_scenes(tx, &manifest_id, person_id, &actor.scenes_cs)?;
}
}
let payload = serde_json::to_string(&CastCheckJob { manifest_id: manifest_id.clone() })?;
repo::enqueue_job(tx, JOB_CAST_CHECK, &payload, now)?;
Ok(IngestOutcome::Pending { manifest_id })
}
/// Computes the §9a `content_id` for a validated manifest.
pub fn compute_content_id(valid: &ValidManifest) -> String {
let m = &valid.manifest;
let identity = CanonicalIdentity {
kind: match m.identity.kind {
IdentityType::Movie => "movie",
IdentityType::Episode => "episode",
},
tmdb_id: m.identity.effective_tmdb_id().map(str::to_string),
imdb_id: m.identity.effective_imdb_id().map(str::to_string),
season: m.identity.season,
episode: m.identity.episode,
};
let cut = CanonicalCut {
runtime_cs: crate::validate::to_centiseconds(m.cut.runtime_sec),
video_hash: m.cut.video_hash.clone(),
};
let actors: Vec<CanonicalActor> = valid
.actor_scenes_cs
.iter()
.filter_map(|a| {
a.tmdb_id
.map(|id| CanonicalActor { tmdb_person_id: id, scenes_cs: a.scenes_cs.clone() })
})
.collect();
content_id::content_id(&identity, &cut, &actors)
}
#[cfg(test)]
mod tests {
use super::*;
use crate::db::Db;
use crate::model::Jmanifest;
use crate::validate::validate_manifest;
const NOW: &str = "2026-07-30T12:00:00Z";
fn valid_from(json: &str) -> ValidManifest {
let m: Jmanifest = serde_json::from_str(json).unwrap();
validate_manifest(m).unwrap()
}
fn movie_json(tmdb: &str, runtime: f64) -> String {
format!(
r#"{{"jmanifest_version":1,
"identity":{{"type":"movie","tmdb_id":"{tmdb}","title":"A Film"}},
"cut":{{"runtime_sec":{runtime}}},
"extraction":{{"sample_fps":5,"pipeline_version":"test 0.1"}},
"actors":[{{"name":"Steve Buscemi","tmdb_id":"884","scenes":[[10.0,20.0]]}},
{{"name":"Michael Palin","tmdb_id":"11007","scenes":[[30.0,40.0]]}}]}}"#
)
}
#[tokio::test]
async fn persists_as_pending_and_enqueues_a_check() {
let db = Db::open(":memory:").unwrap();
let valid = valid_from(&movie_json("504172", 6420.5));
let (outcome, status, jobs) = db
.write(move |tx| {
let c = repo::insert_contributor(tx, "h", NOW)?;
let outcome = persist(tx, &valid, Some(&c), "local", None, NOW)?;
let status = repo::manifest_status(tx, outcome.manifest_id())?;
let jobs = repo::lease_jobs(tx, NOW, 10)?;
Ok((outcome, status, jobs))
})
.await
.unwrap();
assert!(matches!(outcome, IngestOutcome::Pending { .. }));
// §6: held unlisted, not served to anyone, until stage 3 completes.
assert_eq!(status.unwrap().0, "pending");
assert_eq!(jobs.len(), 1);
assert_eq!(jobs[0].kind, JOB_CAST_CHECK);
}
#[tokio::test]
async fn a_pending_manifest_is_not_served() {
let db = Db::open(":memory:").unwrap();
let valid = valid_from(&movie_json("504172", 6420.5));
let candidates = db
.write(move |tx| {
let c = repo::insert_contributor(tx, "h", NOW)?;
persist(tx, &valid, Some(&c), "local", None, NOW)?;
let title =
repo::find_title(tx, IdentityType::Movie, Some("504172"), None)?.unwrap();
repo::candidates_for_title(tx, &title.id, None, None)
})
.await
.unwrap();
assert!(candidates.is_empty());
}
#[tokio::test]
async fn identical_content_deduplicates() {
// §9a: a manifest whose `content_id` is already present is skipped
// without re-validation.
let db = Db::open(":memory:").unwrap();
let a = valid_from(&movie_json("504172", 6420.5));
let b = valid_from(&movie_json("504172", 6420.5));
let (first, second) = db
.write(move |tx| {
let c1 = repo::insert_contributor(tx, "h1", NOW)?;
let c2 = repo::insert_contributor(tx, "h2", NOW)?;
let first = persist(tx, &a, Some(&c1), "local", None, NOW)?;
// A *different* contributor, so this is content dedup, not the
// per-contributor 409.
let second = persist(tx, &b, Some(&c2), "local", None, NOW)?;
Ok((first, second))
})
.await
.unwrap();
assert!(matches!(first, IngestOutcome::Pending { .. }));
assert!(matches!(second, IngestOutcome::DuplicateContent { .. }));
assert_eq!(first.manifest_id(), second.manifest_id());
}
#[tokio::test]
async fn same_contributor_resubmitting_the_same_cut_is_a_duplicate() {
let db = Db::open(":memory:").unwrap();
// Same identity and cut, different actor timings => different content_id,
// so this exercises the per-contributor 409 path specifically.
let a = valid_from(&movie_json("504172", 6420.5));
let b = valid_from(
r#"{"jmanifest_version":1,
"identity":{"type":"movie","tmdb_id":"504172","title":"A Film"},
"cut":{"runtime_sec":6420.5},
"actors":[{"name":"Steve Buscemi","tmdb_id":"884","scenes":[[11.0,21.0]]}]}"#,
);
let second = db
.write(move |tx| {
let c = repo::insert_contributor(tx, "h", NOW)?;
persist(tx, &a, Some(&c), "local", None, NOW)?;
persist(tx, &b, Some(&c), "local", None, NOW)
})
.await
.unwrap();
assert!(matches!(second, IngestOutcome::DuplicateFromContributor { .. }));
}
#[tokio::test]
async fn different_cuts_of_one_title_coexist() {
// §7: multiple manifests may coexist for the same title with different
// cuts — that is the point.
let db = Db::open(":memory:").unwrap();
let a = valid_from(&movie_json("504172", 6420.5));
let b = valid_from(&movie_json("504172", 7000.0));
let (x, y) = db
.write(move |tx| {
let c = repo::insert_contributor(tx, "h", NOW)?;
let x = persist(tx, &a, Some(&c), "local", None, NOW)?;
let y = persist(tx, &b, Some(&c), "local", None, NOW)?;
Ok((x, y))
})
.await
.unwrap();
assert!(matches!(x, IngestOutcome::Pending { .. }));
assert!(matches!(y, IngestOutcome::Pending { .. }));
assert_ne!(x.manifest_id(), y.manifest_id());
}
#[tokio::test]
async fn episode_manifests_carry_their_coordinates() {
let db = Db::open(":memory:").unwrap();
let valid = valid_from(
r#"{"jmanifest_version":1,
"identity":{"type":"episode","series_tmdb_id":"1396","title":"Breaking Bad",
"season":2,"episode":5},
"cut":{"runtime_sec":2820.0},
"actors":[{"name":"Bryan Cranston","tmdb_id":"17419","scenes":[[10.0,20.0]]}]}"#,
);
let row = db
.write(move |tx| {
let c = repo::insert_contributor(tx, "h", NOW)?;
let o = persist(tx, &valid, Some(&c), "local", None, NOW)?;
Ok(repo::manifest_by_id(tx, o.manifest_id())?.unwrap())
})
.await
.unwrap();
assert_eq!((row.season, row.episode), (Some(2), Some(5)));
}
#[tokio::test]
async fn upload_metadata_is_not_echoed_back_as_actor_names() {
// §5a/§7: only integers reach the database. The submitted name is used
// for matching and never persisted, so before the cast check populates
// `people` there is no name to serve.
let db = Db::open(":memory:").unwrap();
let valid = valid_from(&movie_json("504172", 6420.5));
let actors = db
.write(move |tx| {
let c = repo::insert_contributor(tx, "h", NOW)?;
let o = persist(tx, &valid, Some(&c), "local", None, NOW)?;
repo::actors_for_manifest(tx, o.manifest_id())
})
.await
.unwrap();
assert_eq!(actors.len(), 2);
assert!(actors.iter().all(|a| a.name.is_none()));
}
#[test]
fn content_id_excludes_extraction_metadata() {
// §9a: `extraction` metadata and local state are excluded, so two
// servers validating the same upload agree.
let a = valid_from(&movie_json("504172", 6420.5));
let b = valid_from(
r#"{"jmanifest_version":1,
"identity":{"type":"movie","tmdb_id":"504172","title":"A Film"},
"cut":{"runtime_sec":6420.5},
"extraction":{"sample_fps":1,"extinction_sec":9,"pipeline_version":"other 9.9",
"gallery_size":5},
"actors":[{"name":"Steve Buscemi","tmdb_id":"884","scenes":[[10.0,20.0]]},
{"name":"Michael Palin","tmdb_id":"11007","scenes":[[30.0,40.0]]}]}"#,
);
assert_eq!(compute_content_id(&a), compute_content_id(&b));
}
#[test]
fn content_id_excludes_the_audio_signature() {
// §9a is explicit: including it would produce different content_ids for
// identical content and silently break federation deduplication.
let a = valid_from(&movie_json("504172", 6420.5));
let sig = format!("v1:{}", "A".repeat(1720));
let with_sig = format!(
r#"{{"jmanifest_version":1,
"identity":{{"type":"movie","tmdb_id":"504172","title":"A Film"}},
"cut":{{"runtime_sec":6420.5,"audio_signature":"{sig}"}},
"extraction":{{"sample_fps":5,"pipeline_version":"test 0.1"}},
"actors":[{{"name":"Steve Buscemi","tmdb_id":"884","scenes":[[10.0,20.0]]}},
{{"name":"Michael Palin","tmdb_id":"11007","scenes":[[30.0,40.0]]}}]}}"#
);
let b = valid_from(&with_sig);
assert_eq!(compute_content_id(&a), compute_content_id(&b));
}
#[test]
fn content_id_excludes_submitted_names() {
// Names are not persisted, so they must not be part of identity either —
// otherwise a renamed resubmission would evade deduplication.
let a = valid_from(&movie_json("504172", 6420.5));
let b = valid_from(
r#"{"jmanifest_version":1,
"identity":{"type":"movie","tmdb_id":"504172","title":"A Film"},
"cut":{"runtime_sec":6420.5},
"extraction":{"sample_fps":5,"pipeline_version":"test 0.1"},
"actors":[{"name":"Someone Else","tmdb_id":"884","scenes":[[10.0,20.0]]},
{"name":"Another Person","tmdb_id":"11007","scenes":[[30.0,40.0]]}]}"#,
);
assert_eq!(compute_content_id(&a), compute_content_id(&b));
}
}
+38
View File
@@ -0,0 +1,38 @@
//! JRay public server — a community manifest exchange (see `SPEC.md`).
//!
//! Jellyfin servers running the JRay plugin pull actor-timeline manifests
//! ("Jmanifests") for titles they own instead of running the CV pipeline
//! locally, and optionally contribute the manifests they generate back.
//!
//! Module map against the spec:
//!
//! | Module | Spec section |
//! |---|---|
//! | [`model`] | §2 Jmanifest format, and §6 stage 2's `deny_unknown_fields` |
//! | [`validate`] | §6 stage 2 semantics, §5a character class |
//! | [`matching`] | §3 cut matching tiers |
//! | [`content_id`] | §9a canonical form and content addressing |
//! | [`ingest`] | the shared upload path behind §4's `POST` endpoints |
//! | [`castcheck`] | §6 stage 3 scoring, §5a Threat 2 guards |
//! | [`worker`] | §6 stage 3 execution, §8 in-process background work |
//! | [`ratelimit`] | §5 |
//! | [`auth`] | §5a tokens, §8 trusted-proxy handling |
//! | [`db`] | §7 storage, §8 single-writer serialization |
//! | [`api`] | §4 |
pub mod api;
pub mod app;
pub mod auth;
pub mod castcheck;
pub mod config;
pub mod content_id;
pub mod db;
pub mod error;
pub mod ingest;
pub mod matching;
pub mod model;
pub mod ratelimit;
pub mod state;
pub mod tmdb;
pub mod validate;
pub mod worker;
+96
View File
@@ -0,0 +1,96 @@
//! Entry point.
//!
//! §8: one binary, one database file, one reverse proxy. The rate-limit counters
//! and the background cast-check worker both live in this process — no Redis, no
//! broker, no separate worker process.
use std::net::SocketAddr;
use std::sync::Arc;
use anyhow::Context;
use jray_server::app;
use jray_server::config::Config;
use jray_server::db::Db;
use jray_server::ratelimit::RateLimiter;
use jray_server::state::AppState;
use jray_server::tmdb::TmdbClient;
use jray_server::worker::Worker;
use tracing_subscriber::EnvFilter;
#[tokio::main]
async fn main() -> anyhow::Result<()> {
tracing_subscriber::fmt()
.with_env_filter(
EnvFilter::try_from_env("JRAY_LOG").unwrap_or_else(|_| EnvFilter::new("info")),
)
.init();
let config = Arc::new(Config::from_env()?);
let db = Db::open(&config.db_path).context("opening database")?;
let tmdb = Arc::new(TmdbClient::new(config.tmdb_base_url.clone(), config.tmdb_api_key.clone()));
if !tmdb.is_configured() {
// §8: TMDB is a hard dependency for UR-3. Uploads will accumulate in
// `pending` rather than being listed unverified — which is the correct
// failure mode, but the operator should know.
tracing::warn!("no JRAY_TMDB_API_KEY configured: uploads will stay pending, never listed");
}
if config.trusted_proxies.is_empty() {
tracing::info!("no JRAY_TRUSTED_PROXIES set: X-Forwarded-For will be ignored");
}
let state = AppState {
db: db.clone(),
config: config.clone(),
limiter: Arc::new(RateLimiter::new()),
tmdb: tmdb.clone(),
};
let (shutdown_tx, shutdown_rx) = tokio::sync::watch::channel(false);
let worker = Worker {
db: db.clone(),
tmdb,
batch: config.job_batch,
poll_interval: config.job_poll_interval,
};
let worker_handle = tokio::spawn(worker.run(shutdown_rx));
let listener = tokio::net::TcpListener::bind(&config.bind)
.await
.with_context(|| format!("binding {}", config.bind))?;
tracing::info!(bind = %config.bind, server_id = %config.server_id, "jray-server listening");
let router = app::router(state);
axum::serve(listener, router.into_make_service_with_connect_info::<SocketAddr>())
.with_graceful_shutdown(async move {
shutdown_signal().await;
let _ = shutdown_tx.send(true);
})
.await
.context("server error")?;
let _ = worker_handle.await;
Ok(())
}
async fn shutdown_signal() {
let ctrl_c = async {
tokio::signal::ctrl_c().await.expect("installing ctrl-c handler");
};
#[cfg(unix)]
let terminate = async {
tokio::signal::unix::signal(tokio::signal::unix::SignalKind::terminate())
.expect("installing SIGTERM handler")
.recv()
.await;
};
#[cfg(not(unix))]
let terminate = std::future::pending::<()>();
tokio::select! {
_ = ctrl_c => tracing::info!("received ctrl-c, shutting down"),
_ = terminate => tracing::info!("received SIGTERM, shutting down"),
}
}
+217
View File
@@ -0,0 +1,217 @@
//! §3 cut matching.
//!
//! Timings only transfer between identical cuts, so matching is tiered and the
//! server reports *which* tier matched — a `loose` match is meant to surface as
//! a caveat in the JRay UI rather than being applied silently.
//!
//! `video_hash` identifies a *file*, so it only ever matches an identical
//! release and can never produce a false positive; that is why it is tier one.
use crate::model::MatchTier;
/// §3: runtimes within ±2s.
pub const RUNTIME_TOLERANCE_SEC: f64 = 2.0;
/// §3: runtimes within ±30s.
pub const LOOSE_TOLERANCE_SEC: f64 = 30.0;
/// What the client tells us about its own copy.
#[derive(Debug, Clone, Default)]
pub struct ClientCut {
pub runtime_sec: Option<f64>,
pub video_hash: Option<String>,
}
impl ClientCut {
/// True when the client supplied nothing to match on, in which case §4
/// specifies a `"match": "unknown"` answer rather than a guess.
pub fn is_empty(&self) -> bool {
self.runtime_sec.is_none() && self.video_hash.is_none()
}
}
/// What the server holds.
#[derive(Debug, Clone)]
pub struct StoredCut {
pub runtime_sec: f64,
pub video_hash: Option<String>,
}
/// The outcome of comparing a client's cut against a stored one.
#[derive(Debug, Clone, Copy, PartialEq)]
pub struct CutMatch {
pub tier: MatchTier,
/// Scene offset in seconds the client must add (§3 `audio` tier). Always
/// zero for the tiers implemented here; the field exists because the plugin
/// contract is "the server returns the offset, the client applies it", and
/// enabling `audio` must not change the response shape.
pub offset_sec: f64,
}
/// Compares a client's cut against a stored one, returning the best tier that
/// fires, or `None` for "beyond that: no match; do not serve" (§3).
pub fn match_cut(client: &ClientCut, stored: &StoredCut) -> Option<CutMatch> {
// Tier 1 — same file. Checked first and unconditionally: an equal hash is
// decisive regardless of what the runtimes say.
if let (Some(c), Some(s)) = (&client.video_hash, &stored.video_hash) {
if c.eq_ignore_ascii_case(s) {
return Some(CutMatch { tier: MatchTier::Exact, offset_sec: 0.0 });
}
}
// `audio` tier would slot in here, above `runtime`, once signature coverage
// is useful (§3 recommended sequencing).
if let Some(c_rt) = client.runtime_sec {
let delta = (c_rt - stored.runtime_sec).abs();
if delta <= RUNTIME_TOLERANCE_SEC {
return Some(CutMatch { tier: MatchTier::Runtime, offset_sec: 0.0 });
}
if delta <= LOOSE_TOLERANCE_SEC {
return Some(CutMatch { tier: MatchTier::Loose, offset_sec: 0.0 });
}
// A runtime was supplied and cleared nothing — that is a definite
// no-match, not an unknown.
return None;
}
// A hash that did not match, with no runtime to fall back on, tells us
// nothing about alignment either way.
if client.video_hash.is_some() {
return None;
}
Some(CutMatch { tier: MatchTier::Unknown, offset_sec: 0.0 })
}
/// Picks the best-matching stored cut, if any clears `loose` (§4).
///
/// `candidates` is `(key, cut)`; the key is returned so the caller can identify
/// which manifest won without re-scanning.
pub fn best_match<K: Clone>(
client: &ClientCut,
candidates: &[(K, StoredCut)],
) -> Option<(K, CutMatch)> {
candidates
.iter()
.filter_map(|(k, cut)| match_cut(client, cut).map(|m| (k.clone(), m)))
.max_by(|a, b| a.1.tier.cmp(&b.1.tier))
}
#[cfg(test)]
mod tests {
use super::*;
fn stored(runtime: f64, hash: Option<&str>) -> StoredCut {
StoredCut { runtime_sec: runtime, video_hash: hash.map(str::to_string) }
}
#[test]
fn equal_video_hash_is_exact() {
let c = ClientCut {
runtime_sec: Some(6420.5),
video_hash: Some("opensubtitles:8e245d9679d31e12".into()),
};
let m = match_cut(&c, &stored(6420.5, Some("opensubtitles:8e245d9679d31e12"))).unwrap();
assert_eq!(m.tier, MatchTier::Exact);
}
#[test]
fn equal_hash_wins_even_when_runtimes_disagree() {
// The hash identifies the file; a differing stored runtime means our
// own metadata is off, not that the file is different.
let c = ClientCut {
runtime_sec: Some(6000.0),
video_hash: Some("opensubtitles:8e245d9679d31e12".into()),
};
let m = match_cut(&c, &stored(6420.5, Some("opensubtitles:8e245d9679d31e12"))).unwrap();
assert_eq!(m.tier, MatchTier::Exact);
}
#[test]
fn hash_comparison_is_case_insensitive() {
let c = ClientCut {
runtime_sec: None,
video_hash: Some("opensubtitles:8E245D9679D31E12".into()),
};
let m = match_cut(&c, &stored(6420.5, Some("opensubtitles:8e245d9679d31e12"))).unwrap();
assert_eq!(m.tier, MatchTier::Exact);
}
#[test]
fn runtime_within_two_seconds_is_runtime_tier() {
let c = ClientCut { runtime_sec: Some(6422.0), video_hash: None };
assert_eq!(match_cut(&c, &stored(6420.5, None)).unwrap().tier, MatchTier::Runtime);
}
#[test]
fn runtime_within_thirty_seconds_is_loose() {
let c = ClientCut { runtime_sec: Some(6450.0), video_hash: None };
assert_eq!(match_cut(&c, &stored(6420.5, None)).unwrap().tier, MatchTier::Loose);
}
#[test]
fn beyond_thirty_seconds_does_not_match() {
// §3: "beyond that — no match; do not serve".
let c = ClientCut { runtime_sec: Some(6500.0), video_hash: None };
assert!(match_cut(&c, &stored(6420.5, None)).is_none());
}
#[test]
fn tier_boundaries_are_inclusive() {
let c = ClientCut { runtime_sec: Some(6422.5), video_hash: None };
assert_eq!(match_cut(&c, &stored(6420.5, None)).unwrap().tier, MatchTier::Runtime);
let c = ClientCut { runtime_sec: Some(6450.5), video_hash: None };
assert_eq!(match_cut(&c, &stored(6420.5, None)).unwrap().tier, MatchTier::Loose);
}
#[test]
fn no_cut_information_yields_unknown() {
// §4: the mode a library-wide sweep uses — "does the community have
// this title at all", with alignment still to be determined.
let c = ClientCut::default();
assert!(c.is_empty());
assert_eq!(match_cut(&c, &stored(6420.5, None)).unwrap().tier, MatchTier::Unknown);
}
#[test]
fn non_matching_hash_alone_is_not_a_match() {
let c = ClientCut {
runtime_sec: None,
video_hash: Some("opensubtitles:ffffffffffffffff".into()),
};
assert!(match_cut(&c, &stored(6420.5, Some("opensubtitles:8e245d9679d31e12"))).is_none());
}
#[test]
fn non_matching_hash_falls_back_to_runtime() {
let c = ClientCut {
runtime_sec: Some(6421.0),
video_hash: Some("opensubtitles:ffffffffffffffff".into()),
};
let m = match_cut(&c, &stored(6420.5, Some("opensubtitles:8e245d9679d31e12"))).unwrap();
assert_eq!(m.tier, MatchTier::Runtime);
}
#[test]
fn best_match_prefers_the_highest_tier() {
let c = ClientCut {
runtime_sec: Some(6420.5),
video_hash: Some("opensubtitles:8e245d9679d31e12".into()),
};
let candidates = vec![
("loose", stored(6445.0, None)),
("exact", stored(9999.0, Some("opensubtitles:8e245d9679d31e12"))),
("runtime", stored(6420.0, None)),
];
let (winner, m) = best_match(&c, &candidates).unwrap();
assert_eq!(winner, "exact");
assert_eq!(m.tier, MatchTier::Exact);
}
#[test]
fn best_match_returns_none_when_nothing_clears_loose() {
let c = ClientCut { runtime_sec: Some(100.0), video_hash: None };
let candidates = vec![("a", stored(6420.5, None)), ("b", stored(3000.0, None))];
assert!(best_match(&c, &candidates).is_none());
}
}
+307
View File
@@ -0,0 +1,307 @@
//! Jmanifest wire types (§2).
//!
//! **`#[serde(deny_unknown_fields)]` on every struct is the §6 stage 2
//! enforcement mechanism.** "No additional fields anywhere" is a property of
//! these type definitions rather than of validator code that could omit a
//! field, so an unrecognised key at any nesting level fails to parse. That is
//! also what makes the §9 `movie`/`jellyfin_id` strip verifiable: a client that
//! forgets gets a hard `400` naming the field, rather than quietly publishing a
//! contributor's directory layout.
//!
//! Deeper semantic checks — bounds, character classes, path-shaped strings —
//! live in [`crate::validate`]. Parsing rejects *shape*; validation rejects
//! *content*.
use serde::{Deserialize, Serialize};
/// Cut-match tier (§3). Ordered worst-to-best so derived `Ord` ranks them.
#[derive(Debug, Clone, Copy, PartialEq, Eq, PartialOrd, Ord, Serialize, Deserialize)]
#[serde(rename_all = "lowercase")]
pub enum MatchTier {
/// No cut information was supplied, so alignment is unknown (§4).
Unknown,
/// Audio 0.60–0.85, or runtimes within ±30s. Caveat in UI.
Loose,
/// Runtimes within ±2s.
Runtime,
/// Audio score ≥ 0.85; ranks above `runtime` because it is content-derived.
Audio,
/// `video_hash` equal — same file.
Exact,
}
impl MatchTier {
pub fn as_str(self) -> &'static str {
match self {
MatchTier::Unknown => "unknown",
MatchTier::Loose => "loose",
MatchTier::Runtime => "runtime",
MatchTier::Audio => "audio",
MatchTier::Exact => "exact",
}
}
}
#[derive(Debug, Clone, Copy, PartialEq, Eq, Serialize, Deserialize)]
#[serde(rename_all = "lowercase")]
pub enum IdentityType {
Movie,
Episode,
}
/// What the work is (§2 terminology: *title identity*).
///
/// Movie and episode coordinates share one struct because `deny_unknown_fields`
/// with `#[serde(untagged)]` alternatives produces unhelpful error messages;
/// the discriminant is checked in [`crate::validate`], which can name the
/// offending field.
#[derive(Debug, Clone, Deserialize, Serialize)]
#[serde(deny_unknown_fields)]
pub struct Identity {
#[serde(rename = "type")]
pub kind: IdentityType,
// Movie coordinates.
#[serde(default, skip_serializing_if = "Option::is_none")]
pub tmdb_id: Option<String>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub imdb_id: Option<String>,
// Episode coordinates.
#[serde(default, skip_serializing_if = "Option::is_none")]
pub series_tmdb_id: Option<String>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub series_imdb_id: Option<String>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub season: Option<i64>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub episode: Option<i64>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub title: Option<String>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub year: Option<i64>,
}
impl Identity {
/// The TMDB id used as the lookup key, whichever coordinate carries it.
pub fn effective_tmdb_id(&self) -> Option<&str> {
match self.kind {
IdentityType::Movie => self.tmdb_id.as_deref(),
IdentityType::Episode => self.series_tmdb_id.as_deref(),
}
}
pub fn effective_imdb_id(&self) -> Option<&str> {
match self.kind {
IdentityType::Movie => self.imdb_id.as_deref(),
IdentityType::Episode => self.series_imdb_id.as_deref(),
}
}
}
/// Which encode/edit the timings apply to (§2 terminology: *cut fingerprint*).
#[derive(Debug, Clone, Deserialize, Serialize)]
#[serde(deny_unknown_fields)]
pub struct Cut {
/// **Required** — the decoded duration of the media the timings came from.
/// The primary alignment guard (§2).
pub runtime_sec: f64,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub container_duration_sec: Option<f64>,
/// Optional but strongly preferred. OpenSubtitles hash (§3).
#[serde(default, skip_serializing_if = "Option::is_none")]
pub video_hash: Option<String>,
/// Optional; version-prefixed spectral-peak signature (§3, UR-9).
///
/// Accepted and stored by this build; `audio`-tier matching is enabled once
/// coverage is useful, per §3's recommended sequencing.
#[serde(default, skip_serializing_if = "Option::is_none")]
pub audio_signature: Option<String>,
}
/// How well the contributor's gallery could discriminate (§2, §7).
///
/// The strongest available quality signal between two otherwise comparable
/// manifests: a `Global` gallery had to distinguish its actors from every other
/// actor in the contributor's library, whereas a `Limited` one only had to
/// distinguish them from this title's own cast.
#[derive(Debug, Clone, Copy, PartialEq, Eq, PartialOrd, Ord, Deserialize, Serialize)]
#[serde(rename_all = "lowercase")]
pub enum GalleryScope {
/// Built from this title's cast alone.
Limited,
/// Built from the whole library. The default upstream.
Global,
}
impl GalleryScope {
pub fn as_str(self) -> &'static str {
match self {
GalleryScope::Limited => "limited",
GalleryScope::Global => "global",
}
}
}
/// Extraction parameters, carried for provenance and ranking.
///
/// Note there is no `anneal_sec`: it was **withdrawn** in the SR-003 schema
/// bump, because presence now follows track extent — a track survives its own
/// gaps, so there is nothing to anneal (`scene-actor-extraction` AR-012/AR-013).
/// `deny_unknown_fields` therefore makes its presence a hard parse error rather
/// than something silently ignored, which is deliberate: a manifest still
/// carrying it was produced by a pipeline whose window semantics differ from
/// what this server now assumes.
#[derive(Debug, Clone, Deserialize, Serialize)]
#[serde(deny_unknown_fields)]
pub struct Extraction {
#[serde(default, skip_serializing_if = "Option::is_none")]
pub sample_fps: Option<f64>,
/// The re-acquisition timeout that shapes window extent. Successor to the
/// withdrawn `anneal_sec`.
#[serde(default, skip_serializing_if = "Option::is_none")]
pub extinction_sec: Option<f64>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub pipeline_version: Option<String>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub gallery_size: Option<i64>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub gallery_scope: Option<GalleryScope>,
}
/// One actor's timeline.
///
/// Note there is no `jellyfin_id` field: `deny_unknown_fields` means its
/// presence is a parse error, which is exactly the §2/§6 requirement that it be
/// *rejected on upload* rather than merely ignored on download.
#[derive(Debug, Clone, Deserialize, Serialize)]
#[serde(deny_unknown_fields)]
pub struct Actor {
/// Sent on upload for matching, but **not persisted** — the server resolves
/// each actor to a TMDB person id and serves names from its own TMDB-derived
/// table (§2, §5a). On download this is server-authoritative.
#[serde(default, skip_serializing_if = "Option::is_none")]
pub name: Option<String>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub imdb_id: Option<String>,
/// The **primary** actor join key (§2, §6 stage 3).
#[serde(default, skip_serializing_if = "Option::is_none")]
pub tmdb_id: Option<String>,
/// `[start_sec, end_sec]` inclusive, sorted.
pub scenes: Vec<[f64; 2]>,
}
/// One shareable actor timeline for one cut of one title (§2).
#[derive(Debug, Clone, Deserialize, Serialize)]
#[serde(deny_unknown_fields)]
pub struct Jmanifest {
pub jmanifest_version: u32,
pub identity: Identity,
pub cut: Cut,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub extraction: Option<Extraction>,
pub actors: Vec<Actor>,
}
/// Series-level coordinates for a bundle envelope (§2).
#[derive(Debug, Clone, Deserialize, Serialize)]
#[serde(deny_unknown_fields)]
pub struct SeriesRef {
#[serde(default, skip_serializing_if = "Option::is_none")]
pub series_tmdb_id: Option<String>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub series_imdb_id: Option<String>,
#[serde(default, skip_serializing_if = "Option::is_none")]
pub title: Option<String>,
}
/// A thin wrapper, not a new format (§2). Bundles are a transfer convenience,
/// never a storage unit — each episode is stored and moderated individually.
#[derive(Debug, Clone, Deserialize, Serialize)]
#[serde(deny_unknown_fields)]
pub struct SeriesBundle {
pub jmanifest_version: u32,
pub series: SeriesRef,
pub episodes: Vec<Jmanifest>,
/// Present on responses only; ignored on upload.
#[serde(default, skip_serializing_if = "Option::is_none")]
pub coverage: Option<Coverage>,
}
#[derive(Debug, Clone, Deserialize, Serialize)]
#[serde(deny_unknown_fields)]
pub struct Coverage {
pub episodes_available: usize,
pub seasons: Vec<i64>,
}
/// The current `jmanifest_version` this server speaks (§2).
pub const JMANIFEST_VERSION: u32 = 1;
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn unknown_field_at_top_level_is_rejected() {
let json = r#"{"jmanifest_version":1,"identity":{"type":"movie","tmdb_id":"1"},
"cut":{"runtime_sec":100.0},"actors":[],"surprise":"x"}"#;
let err = serde_json::from_str::<Jmanifest>(json).unwrap_err().to_string();
assert!(err.contains("surprise"), "error should name the field: {err}");
}
#[test]
fn jellyfin_id_on_an_actor_is_a_parse_error() {
// §2: `actors[].jellyfin_id` must not appear. `deny_unknown_fields`
// makes this structural rather than a validator's responsibility.
let json = r#"{"jmanifest_version":1,"identity":{"type":"movie","tmdb_id":"1"},
"cut":{"runtime_sec":100.0},
"actors":[{"name":"A","tmdb_id":"2","jellyfin_id":"guid","scenes":[]}]}"#;
let err = serde_json::from_str::<Jmanifest>(json).unwrap_err().to_string();
assert!(err.contains("jellyfin_id"), "error should name the field: {err}");
}
#[test]
fn movie_path_field_is_a_parse_error() {
let json = r#"{"jmanifest_version":1,"movie":"/data/movies/x.mkv",
"identity":{"type":"movie","tmdb_id":"1"},
"cut":{"runtime_sec":100.0},"actors":[]}"#;
let err = serde_json::from_str::<Jmanifest>(json).unwrap_err().to_string();
assert!(err.contains("movie"), "error should name the field: {err}");
}
#[test]
fn unknown_field_nested_in_cut_is_rejected() {
let json = r#"{"jmanifest_version":1,"identity":{"type":"movie","tmdb_id":"1"},
"cut":{"runtime_sec":100.0,"payload":"x"},"actors":[]}"#;
assert!(serde_json::from_str::<Jmanifest>(json).is_err());
}
#[test]
fn spec_example_manifest_parses() {
let json = r#"{
"jmanifest_version": 1,
"identity": { "type": "movie", "tmdb_id": "504172", "imdb_id": "tt4686844",
"title": "The Death of Stalin", "year": 2017 },
"cut": { "runtime_sec": 6420.5, "container_duration_sec": 6420.5,
"video_hash": "opensubtitles:8e245d9679d31e12" },
"extraction": { "sample_fps": 5, "extinction_sec": 12,
"pipeline_version": "scene-actor-extraction 0.4.1",
"gallery_size": 1820, "gallery_scope": "global" },
"actors": [ { "name": "Steve Buscemi", "imdb_id": "nm0000114", "tmdb_id": "884",
"scenes": [[191.6, 209.2], [438.2, 465.6]] } ]
}"#;
let m: Jmanifest = serde_json::from_str(json).unwrap();
assert_eq!(m.actors.len(), 1);
assert_eq!(m.identity.effective_tmdb_id(), Some("504172"));
}
#[test]
fn tiers_order_audio_above_runtime() {
// §3: `audio` ranks above `runtime` because it is content-derived.
assert!(MatchTier::Audio > MatchTier::Runtime);
assert!(MatchTier::Exact > MatchTier::Audio);
assert!(MatchTier::Runtime > MatchTier::Loose);
}
}
+217
View File
@@ -0,0 +1,217 @@
//! §5 rate limiting.
//!
//! A fixed-window counter keyed on `(token_or_ip, surface)`, held in process
//! memory — no external counter store. §5 is explicit that a sliding window is
//! not worth the complexity at this volume, and that counters resetting on
//! restart is acceptable for abuse throttling.
//!
//! Read limits are applied *behind* the CDN cache, so a cache hit costs a client
//! nothing against its budget — that is a deployment property (§8), not
//! something this module can enforce.
use std::collections::HashMap;
use std::sync::Mutex;
use std::time::{Duration, Instant};
/// The rate-limited surfaces of §5. Distinct from routes: the batch and single
/// forms of `exists` are separate surfaces with separate budgets.
#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash)]
pub enum Surface {
ExistsSingle,
ExistsBatch,
ManifestFetch,
SeriesFetch,
ManifestUpload,
BundleUpload,
Report,
Search,
}
impl Surface {
/// Requests per hour, per §5's table.
pub fn limit(self) -> u32 {
match self {
Surface::ExistsSingle => 600,
Surface::ExistsBatch => 60,
Surface::ManifestFetch => 300,
Surface::SeriesFetch => 120,
Surface::ManifestUpload => 100,
Surface::BundleUpload => 20,
Surface::Report => 20,
Surface::Search => 60,
}
}
pub fn as_str(self) -> &'static str {
match self {
Surface::ExistsSingle => "exists",
Surface::ExistsBatch => "exists_batch",
Surface::ManifestFetch => "manifest_fetch",
Surface::SeriesFetch => "series_fetch",
Surface::ManifestUpload => "manifest_upload",
Surface::BundleUpload => "bundle_upload",
Surface::Report => "report",
Surface::Search => "search",
}
}
}
const WINDOW: Duration = Duration::from_secs(3600);
/// Headers §5 requires on every rate-limited response.
#[derive(Debug, Clone, Copy)]
pub struct Quota {
pub limit: u32,
pub remaining: u32,
/// Seconds until the window resets.
pub reset: u64,
}
#[derive(Debug, Clone, Copy)]
struct Window {
started: Instant,
count: u32,
}
pub struct RateLimiter {
windows: Mutex<HashMap<(String, Surface), Window>>,
}
impl Default for RateLimiter {
fn default() -> Self {
Self::new()
}
}
impl RateLimiter {
pub fn new() -> Self {
Self { windows: Mutex::new(HashMap::new()) }
}
/// Records one request against `(key, surface)`.
///
/// `Ok(quota)` when within budget, `Err(quota)` when the limit is exceeded —
/// in which case the caller returns `429` with `Retry-After` set from
/// `quota.reset`. A rejected request does **not** increment the counter, so a
/// client hammering a closed window cannot extend its own lockout.
pub fn check(&self, key: &str, surface: Surface) -> Result<Quota, Quota> {
self.check_at(key, surface, Instant::now())
}
fn check_at(&self, key: &str, surface: Surface, now: Instant) -> Result<Quota, Quota> {
let limit = surface.limit();
let mut windows = self.windows.lock().expect("rate limiter poisoned");
// Opportunistic eviction of stale windows, so an IP-keyed map cannot
// grow without bound behind CGNAT.
if windows.len() > 10_000 {
windows.retain(|_, w| now.duration_since(w.started) < WINDOW);
}
let entry =
windows.entry((key.to_string(), surface)).or_insert(Window { started: now, count: 0 });
let elapsed = now.duration_since(entry.started);
if elapsed >= WINDOW {
*entry = Window { started: now, count: 0 };
}
let reset = WINDOW.saturating_sub(now.duration_since(entry.started)).as_secs();
if entry.count >= limit {
return Err(Quota { limit, remaining: 0, reset });
}
entry.count += 1;
Ok(Quota { limit, remaining: limit - entry.count, reset })
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn allows_up_to_the_limit_then_rejects() {
let rl = RateLimiter::new();
let limit = Surface::Report.limit();
for i in 0..limit {
let q = rl.check("ip", Surface::Report).expect("within budget");
assert_eq!(q.remaining, limit - i - 1);
}
let q = rl.check("ip", Surface::Report).expect_err("over budget");
assert_eq!(q.remaining, 0);
}
#[test]
fn surfaces_have_independent_budgets() {
let rl = RateLimiter::new();
for _ in 0..Surface::BundleUpload.limit() {
rl.check("t", Surface::BundleUpload).unwrap();
}
assert!(rl.check("t", Surface::BundleUpload).is_err());
// §5: a bundle counts as a single write against its own limit, and must
// not consume the single-manifest budget.
assert!(rl.check("t", Surface::ManifestUpload).is_ok());
}
#[test]
fn keys_are_independent() {
let rl = RateLimiter::new();
for _ in 0..Surface::Report.limit() {
rl.check("a", Surface::Report).unwrap();
}
assert!(rl.check("a", Surface::Report).is_err());
assert!(rl.check("b", Surface::Report).is_ok());
}
#[test]
fn window_resets_after_an_hour() {
let rl = RateLimiter::new();
let t0 = Instant::now();
for _ in 0..Surface::Report.limit() {
rl.check_at("ip", Surface::Report, t0).unwrap();
}
assert!(rl.check_at("ip", Surface::Report, t0).is_err());
// Still closed just inside the window.
assert!(rl.check_at("ip", Surface::Report, t0 + Duration::from_secs(3599)).is_err());
// Open again once it rolls over.
assert!(rl.check_at("ip", Surface::Report, t0 + Duration::from_secs(3600)).is_ok());
}
#[test]
fn rejected_requests_do_not_extend_the_lockout() {
let rl = RateLimiter::new();
let t0 = Instant::now();
for _ in 0..Surface::Report.limit() {
rl.check_at("ip", Surface::Report, t0).unwrap();
}
// Hammer the closed window; the counter must not keep climbing, so the
// window still expires on schedule.
for _ in 0..50 {
assert!(rl.check_at("ip", Surface::Report, t0 + Duration::from_secs(10)).is_err());
}
assert!(rl.check_at("ip", Surface::Report, t0 + Duration::from_secs(3600)).is_ok());
}
#[test]
fn reset_counts_down_within_the_window() {
let rl = RateLimiter::new();
let t0 = Instant::now();
let q = rl.check_at("ip", Surface::ExistsSingle, t0).unwrap();
assert_eq!(q.reset, 3600);
let q = rl.check_at("ip", Surface::ExistsSingle, t0 + Duration::from_secs(600)).unwrap();
assert_eq!(q.reset, 3000);
}
#[test]
fn limits_match_the_spec_table() {
assert_eq!(Surface::ExistsSingle.limit(), 600);
assert_eq!(Surface::ExistsBatch.limit(), 60);
assert_eq!(Surface::ManifestFetch.limit(), 300);
assert_eq!(Surface::SeriesFetch.limit(), 120);
assert_eq!(Surface::ManifestUpload.limit(), 100);
assert_eq!(Surface::BundleUpload.limit(), 20);
assert_eq!(Surface::Report.limit(), 20);
assert_eq!(Surface::Search.limit(), 60);
}
}
+102
View File
@@ -0,0 +1,102 @@
//! Shared application state, and the cross-cutting request concerns (§5 rate
//! limiting, §5a token resolution) that every handler needs.
use std::net::SocketAddr;
use std::sync::Arc;
use axum::extract::ConnectInfo;
use axum::http::{HeaderMap, HeaderValue};
use axum::response::Response;
use crate::auth;
use crate::config::Config;
use crate::db::{repo, Db};
use crate::error::{ApiError, ApiResult};
use crate::ratelimit::{Quota, RateLimiter, Surface};
use crate::tmdb::TmdbClient;
/// The connection's peer address, when the server was started with connect-info.
///
/// A dedicated extractor rather than `ConnectInfo<SocketAddr>` directly, because
/// this must not be a *hard* requirement: a router used without
/// `into_make_service_with_connect_info` — as in tests — has no peer address, and
/// a handler that fails to extract would be a routing error rather than degrading
/// to header-only attribution.
pub struct PeerIp(pub Option<std::net::IpAddr>);
impl<S> axum::extract::FromRequestParts<S> for PeerIp
where
S: Send + Sync,
{
type Rejection = std::convert::Infallible;
async fn from_request_parts(
parts: &mut axum::http::request::Parts,
_state: &S,
) -> Result<Self, Self::Rejection> {
Ok(PeerIp(
parts.extensions.get::<ConnectInfo<SocketAddr>>().map(|ConnectInfo(addr)| addr.ip()),
))
}
}
#[derive(Clone)]
pub struct AppState {
pub db: Db,
pub config: Arc<Config>,
pub limiter: Arc<RateLimiter>,
pub tmdb: Arc<TmdbClient>,
}
impl AppState {
/// Resolves the client IP for rate-limiting and attribution, honouring
/// `X-Forwarded-For` only from a configured proxy (§8).
pub fn client_ip(&self, headers: &HeaderMap, peer: Option<std::net::IpAddr>) -> String {
auth::client_ip(headers, peer, &self.config.trusted_proxies)
}
/// §5: limits are per token where one is present, otherwise per source IP.
pub fn check_limit(&self, key: &str, surface: Surface) -> ApiResult<Quota> {
self.limiter.check(key, surface).map_err(|q| {
tracing::debug!(surface = surface.as_str(), "rate limited");
ApiError::RateLimited { retry_after: q.reset.max(1) }
})
}
/// Resolves a bearer token to a contributor (§5a).
///
/// A token is an anonymous bearer capability, not an account: the only state
/// behind it is the per-token counters used for rate-limiting attribution and
/// automatic revocation.
pub async fn require_contributor(&self, headers: &HeaderMap) -> ApiResult<repo::Contributor> {
let token = auth::bearer_token(headers).ok_or(ApiError::Unauthorized)?;
let hash = auth::hash_token(&token);
let found = self
.db
.read(move |conn| repo::contributor_by_token_hash(conn, &hash))
.await
.map_err(ApiError::Internal)?;
match found {
Some(c) if !c.revoked => Ok(c),
// A revoked token is indistinguishable from an unknown one to the
// caller; there is nothing useful to disclose.
_ => Err(ApiError::Unauthorized),
}
}
}
/// Attaches the §5 rate-limit headers to a response.
pub fn with_quota_headers(mut resp: Response, quota: Quota) -> Response {
let h = resp.headers_mut();
insert_num(h, "x-ratelimit-limit", quota.limit as u64);
insert_num(h, "x-ratelimit-remaining", quota.remaining as u64);
insert_num(h, "x-ratelimit-reset", quota.reset);
resp
}
fn insert_num(headers: &mut HeaderMap, name: &'static str, value: u64) {
if let Ok(v) = HeaderValue::from_str(&value.to_string()) {
headers.insert(name, v);
}
}
+226
View File
@@ -0,0 +1,226 @@
//! TMDB client for the §6 stage 3 cast cross-check.
//!
//! §5a's Threat 2 defence rests entirely on the attacker not controlling TMDB:
//! to make a prank manifest pass, they would need those performers to be
//! credited cast on that title in TMDB, which means vandalising a separate,
//! moderated system.
//!
//! Responses are cached for 24h (§6) so a burst of episode uploads for one
//! series costs a single upstream call, and so the server stays within TMDB's
//! own rate limits.
use std::time::Duration;
use serde::Deserialize;
/// A credited cast member, reduced to what the check needs.
#[derive(Debug, Clone, Deserialize)]
pub struct CastMember {
pub id: u64,
#[serde(default)]
pub name: String,
#[serde(default)]
pub adult: bool,
}
#[derive(Debug, Clone, Default, Deserialize)]
pub struct Credits {
#[serde(default)]
pub cast: Vec<CastMember>,
/// Present on episode credits.
#[serde(default)]
pub guest_stars: Vec<CastMember>,
}
impl Credits {
/// Cast plus guest stars — the union §6 specifies for episodes.
pub fn all(&self) -> impl Iterator<Item = &CastMember> {
self.cast.iter().chain(self.guest_stars.iter())
}
}
#[derive(Debug, Clone, Default, Deserialize)]
pub struct TitleDetails {
#[serde(default)]
pub adult: bool,
#[serde(default)]
pub title: Option<String>,
#[serde(default)]
pub name: Option<String>,
}
/// A failure that should be retried rather than treated as a verdict.
///
/// §6: "TMDB unreachable / rate-limited → retry with backoff; stays unlisted,
/// not rejected." Distinguishing this from "TMDB has no credits" is essential —
/// conflating them would reject honest manifests during an outage.
#[derive(Debug, thiserror::Error)]
pub enum TmdbError {
#[error("tmdb transport error: {0}")]
Transport(String),
#[error("tmdb rate limited")]
RateLimited,
#[error("tmdb server error: {0}")]
ServerError(u16),
/// The id genuinely does not exist upstream.
#[error("tmdb resource not found")]
NotFound,
#[error("tmdb response was not understood: {0}")]
Malformed(String),
#[error("no tmdb api key configured")]
NotConfigured,
}
impl TmdbError {
/// True when the job should be rescheduled rather than resolved.
pub fn is_retryable(&self) -> bool {
matches!(
self,
TmdbError::Transport(_)
| TmdbError::RateLimited
| TmdbError::ServerError(_)
| TmdbError::NotConfigured
)
}
}
#[derive(Clone)]
pub struct TmdbClient {
http: reqwest::Client,
base_url: String,
api_key: Option<String>,
}
impl TmdbClient {
pub fn new(base_url: String, api_key: Option<String>) -> Self {
let http = reqwest::Client::builder()
.timeout(Duration::from_secs(15))
.user_agent(concat!("jray-server/", env!("CARGO_PKG_VERSION")))
.build()
.expect("building reqwest client");
Self { http, base_url, api_key }
}
pub fn is_configured(&self) -> bool {
self.api_key.is_some()
}
async fn get<T: serde::de::DeserializeOwned>(&self, path: &str) -> Result<T, TmdbError> {
let key = self.api_key.as_deref().ok_or(TmdbError::NotConfigured)?;
let url =
format!("{}/{}", self.base_url.trim_end_matches('/'), path.trim_start_matches('/'));
let resp = self
.http
.get(&url)
.query(&[("api_key", key)])
.send()
.await
.map_err(|e| TmdbError::Transport(e.to_string()))?;
let status = resp.status();
if status == reqwest::StatusCode::NOT_FOUND {
return Err(TmdbError::NotFound);
}
if status == reqwest::StatusCode::TOO_MANY_REQUESTS {
return Err(TmdbError::RateLimited);
}
if status.is_server_error() {
return Err(TmdbError::ServerError(status.as_u16()));
}
if !status.is_success() {
return Err(TmdbError::Malformed(format!("unexpected status {status}")));
}
let body = resp.text().await.map_err(|e| TmdbError::Transport(e.to_string()))?;
serde_json::from_str(&body).map_err(|e| TmdbError::Malformed(e.to_string()))
}
pub async fn movie_credits(&self, tmdb_id: &str) -> Result<Credits, TmdbError> {
self.get(&format!("movie/{tmdb_id}/credits")).await
}
pub async fn movie_details(&self, tmdb_id: &str) -> Result<TitleDetails, TmdbError> {
self.get(&format!("movie/{tmdb_id}")).await
}
pub async fn series_credits(&self, series_tmdb_id: &str) -> Result<Credits, TmdbError> {
// Aggregate credits carry recurring cast TMDB lists only at series level.
self.get(&format!("tv/{series_tmdb_id}/aggregate_credits")).await
}
pub async fn episode_credits(
&self,
series_tmdb_id: &str,
season: i64,
episode: i64,
) -> Result<Credits, TmdbError> {
self.get(&format!("tv/{series_tmdb_id}/season/{season}/episode/{episode}/credits")).await
}
pub async fn series_details(&self, series_tmdb_id: &str) -> Result<TitleDetails, TmdbError> {
self.get(&format!("tv/{series_tmdb_id}")).await
}
}
/// Fetches a person's details, used by the §5a category guard.
#[derive(Debug, Clone, Default, Deserialize)]
pub struct PersonDetails {
#[serde(default)]
pub adult: bool,
#[serde(default)]
pub name: String,
}
impl TmdbClient {
pub async fn person(&self, tmdb_person_id: u64) -> Result<PersonDetails, TmdbError> {
self.get(&format!("person/{tmdb_person_id}")).await
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn credits_union_covers_cast_and_guest_stars() {
// §6: for episodes the check runs against the union of per-episode
// credits (cast + guest stars) and series aggregate credits.
let c: Credits = serde_json::from_str(
r#"{"cast":[{"id":1,"name":"A"}],"guest_stars":[{"id":2,"name":"B"}]}"#,
)
.unwrap();
let ids: Vec<u64> = c.all().map(|m| m.id).collect();
assert_eq!(ids, vec![1, 2]);
}
#[test]
fn credits_tolerate_missing_and_extra_fields() {
// TMDB adds fields freely; our own strictness applies to *uploads*, not
// to a trusted upstream we merely read.
let c: Credits =
serde_json::from_str(r#"{"cast":[{"id":1,"unexpected":true}],"id":99}"#).unwrap();
assert_eq!(c.cast.len(), 1);
assert_eq!(c.cast[0].name, "");
assert!(c.guest_stars.is_empty());
}
#[test]
fn transport_and_rate_limit_are_retryable_but_not_found_is_not() {
// The distinction that keeps an outage from rejecting honest uploads.
assert!(TmdbError::Transport("x".into()).is_retryable());
assert!(TmdbError::RateLimited.is_retryable());
assert!(TmdbError::ServerError(503).is_retryable());
assert!(TmdbError::NotConfigured.is_retryable());
assert!(!TmdbError::NotFound.is_retryable());
assert!(!TmdbError::Malformed("x".into()).is_retryable());
}
#[tokio::test]
async fn unconfigured_client_reports_retryable_failure() {
let c = TmdbClient::new("http://127.0.0.1:1".into(), None);
assert!(!c.is_configured());
let err = c.movie_credits("1").await.unwrap_err();
assert!(err.is_retryable(), "missing key must hold uploads pending, not reject them");
}
}
+1055
View File
File diff suppressed because it is too large Load Diff
+494
View File
@@ -0,0 +1,494 @@
//! Background worker for the §6 stage 3 cast check.
//!
//! §8: this runs as a Tokio background task in the same binary, with the job
//! queue as a SQLite table so state survives restart — replacing an external
//! broker entirely. The check needs an outbound TMDB call and so cannot run
//! inside the request without coupling upload latency to a third party (§6).
use std::sync::Arc;
use std::time::Duration;
use anyhow::Context;
use crate::castcheck::{self, SubmittedActor, Verdict};
use crate::db::{repo, Db};
use crate::ingest::{CastCheckJob, JOB_CAST_CHECK};
use crate::tmdb::{CastMember, Credits, TmdbClient, TmdbError};
/// §6: TMDB responses are cached for 24h, so a burst of episode uploads for one
/// series costs a single upstream call.
const CACHE_TTL: Duration = Duration::from_secs(24 * 3600);
/// Cap on retry backoff for a persistent TMDB outage.
const MAX_BACKOFF_SECS: u64 = 3600;
pub struct Worker {
pub db: Db,
pub tmdb: Arc<TmdbClient>,
pub batch: usize,
pub poll_interval: Duration,
}
impl Worker {
/// Runs until `shutdown` resolves.
pub async fn run(self, mut shutdown: tokio::sync::watch::Receiver<bool>) {
// A process that died mid-job would otherwise leave work stranded.
match self.db.write(repo::release_all_leases).await {
Ok(n) if n > 0 => tracing::info!(released = n, "released stranded job leases"),
Ok(_) => {}
Err(e) => tracing::error!(error = ?e, "failed to release job leases at startup"),
}
loop {
tokio::select! {
_ = shutdown.changed() => {
tracing::info!("worker shutting down");
return;
}
_ = tokio::time::sleep(self.poll_interval) => {
if let Err(e) = self.tick().await {
tracing::error!(error = ?e, "worker tick failed");
}
}
}
}
}
async fn tick(&self) -> anyhow::Result<()> {
let now = now_iso();
let batch = self.batch;
let leased_at = now.clone();
let jobs = self.db.write(move |tx| repo::lease_jobs(tx, &leased_at, batch)).await?;
for job in jobs {
let result = match job.kind.as_str() {
JOB_CAST_CHECK => self.run_cast_check(&job.payload).await,
other => {
tracing::warn!(kind = other, "unknown job kind, dropping");
Ok(())
}
};
let job_id = job.id.clone();
match result {
Ok(()) => {
self.db.write(move |tx| repo::delete_job(tx, &job_id)).await?;
}
Err(JobError::Retry(msg)) => {
// §6: TMDB unreachable or rate-limited means retry with
// backoff; the manifest stays unlisted, not rejected.
let delay = backoff_secs(job.attempts);
let run_after = iso_in(delay);
tracing::warn!(job = %job_id, attempts = job.attempts, delay, reason = %msg,
"rescheduling job");
self.db
.write(move |tx| repo::reschedule_job(tx, &job_id, &run_after, &msg))
.await?;
}
Err(JobError::Fatal(e)) => {
tracing::error!(job = %job_id, error = ?e, "dropping job after fatal error");
self.db.write(move |tx| repo::delete_job(tx, &job_id)).await?;
}
}
}
Ok(())
}
async fn run_cast_check(&self, payload: &str) -> Result<(), JobError> {
let job: CastCheckJob =
serde_json::from_str(payload).map_err(|e| JobError::Fatal(e.into()))?;
let manifest_id = job.manifest_id;
// Load what the check needs.
let mid = manifest_id.clone();
let loaded = self
.db
.read(move |conn| {
let Some(m) = repo::manifest_by_id(conn, &mid)? else { return Ok(None) };
let title = conn
.query_row(
"SELECT kind, tmdb_id, imdb_id, adult, certification FROM titles WHERE id = ?1",
rusqlite::params![m.title_id],
|r| {
Ok((
r.get::<_, String>(0)?,
r.get::<_, Option<String>>(1)?,
r.get::<_, Option<String>>(2)?,
r.get::<_, i64>(3)? != 0,
r.get::<_, Option<String>>(4)?,
))
},
)
.map_err(anyhow::Error::from)?;
let actor_ids = repo::manifest_actor_ids(conn, &mid)?;
Ok(Some((m, title, actor_ids)))
})
.await
.map_err(JobError::Fatal)?;
// The manifest may have been deleted (contributor revoked, §5a) between
// enqueue and now; that is not an error.
let Some((manifest, (kind, tmdb_id, _imdb_id, title_adult, certification), actor_ids)) =
loaded
else {
return Ok(());
};
if manifest.status != "pending" {
return Ok(());
}
let Some(tmdb_id) = tmdb_id else {
// No TMDB id means the cast check cannot run at all. §6 treats absent
// reference data as flagged, not rejected.
self.finalise(&manifest_id, Verdict::Flagged, 0.0, Some("no_tmdb_id"), &[], &[])
.await
.map_err(JobError::Fatal)?;
return Ok(());
};
let credits = self
.credits_for(&kind, &tmdb_id, manifest.season, manifest.episode)
.await
.map_err(|e| {
if e.is_retryable() {
JobError::Retry(e.to_string())
} else {
// A genuinely absent title is a verdict, not a transport
// failure — handled below via empty credits.
JobError::Retry(format!("non-retryable tmdb error treated as absent: {e}"))
}
});
let credits = match credits {
Ok(c) => c,
Err(JobError::Retry(msg)) if msg.starts_with("non-retryable") => {
tracing::info!(manifest = %manifest_id, "tmdb has no such title; flagging");
Credits::default()
}
Err(e) => return Err(e),
};
let reference: Vec<CastMember> = credits.all().cloned().collect();
let submitted: Vec<SubmittedActor> = actor_ids
.iter()
.map(|id| SubmittedActor { tmdb_id: Some(*id), imdb_id: None, name: None })
.collect();
let mut outcome = castcheck::evaluate(&submitted, &reference);
// §5a layer 1 — category guard.
if let Some(offender) = castcheck::category_guard_violation(&outcome.matched, title_adult) {
tracing::warn!(manifest = %manifest_id, person = offender,
"category guard: adult-flagged performer on a non-adult title");
outcome.verdict = Verdict::Rejected;
outcome.reason = Some("category_guard".into());
}
// §5a layer 2 — age-appropriateness guard.
if let Some(cert) = &certification {
if castcheck::is_childrens_certification(cert) {
castcheck::apply_childrens_guard(&mut outcome, submitted.len());
}
}
self.finalise(
&manifest_id,
outcome.verdict,
outcome.ratio,
outcome.reason.as_deref(),
&outcome.matched,
&outcome.unmatched_person_ids,
)
.await
.map_err(JobError::Fatal)?;
Ok(())
}
/// Fetches credits, using the 24h cache (§6).
///
/// For episodes this is the **union** of TMDB's per-episode credits (cast +
/// guest stars) and the series' aggregate credits: per-episode alone would
/// reject recurring cast TMDB lists only at series level, series-wide alone
/// would reject legitimate guest stars.
async fn credits_for(
&self,
kind: &str,
tmdb_id: &str,
season: Option<i64>,
episode: Option<i64>,
) -> Result<Credits, TmdbError> {
if kind == "movie" {
return self.cached("movie", tmdb_id, || self.tmdb.movie_credits(tmdb_id)).await;
}
let series = self.cached("series", tmdb_id, || self.tmdb.series_credits(tmdb_id)).await?;
let mut combined = series;
if let (Some(s), Some(e)) = (season, episode) {
let key = format!("{tmdb_id}:{s}:{e}");
match self.cached("episode", &key, || self.tmdb.episode_credits(tmdb_id, s, e)).await {
Ok(ep) => {
combined.cast.extend(ep.cast);
combined.guest_stars.extend(ep.guest_stars);
}
// A missing episode entry is normal; the series set still applies.
Err(TmdbError::NotFound) => {}
Err(e) if e.is_retryable() => return Err(e),
Err(e) => tracing::warn!(error = ?e, "ignoring episode credits error"),
}
}
Ok(combined)
}
async fn cached<F, Fut>(&self, kind: &str, key: &str, fetch: F) -> Result<Credits, TmdbError>
where
F: FnOnce() -> Fut,
Fut: std::future::Future<Output = Result<Credits, TmdbError>>,
{
let (k, kk) = (key.to_string(), kind.to_string());
let cached = self
.db
.read(move |conn| repo::cached_credits(conn, &k, &kk))
.await
.map_err(|e| TmdbError::Transport(e.to_string()))?;
if let Some((json, fetched_at)) = cached {
if !is_stale(&fetched_at, CACHE_TTL) {
if let Ok(c) = serde_json::from_str::<Credits>(&json) {
return Ok(c);
}
}
}
let fresh = fetch().await?;
let json = serde_json::to_string(&SerializableCredits::from(&fresh))
.map_err(|e| TmdbError::Malformed(e.to_string()))?;
let (k, kk, now) = (key.to_string(), kind.to_string(), now_iso());
let _ = self.db.write(move |tx| repo::put_credits(tx, &k, &kk, &json, &now)).await;
Ok(fresh)
}
/// Applies the verdict: resolves names into `people`, drops unmatched actors,
/// and updates status and contributor counters — in one transaction.
async fn finalise(
&self,
manifest_id: &str,
verdict: Verdict,
ratio: f64,
reason: Option<&str>,
matched: &[castcheck::MatchedActor],
unmatched: &[u64],
) -> anyhow::Result<()> {
let id = manifest_id.to_string();
let reason = reason.map(str::to_string);
let matched: Vec<(u64, String, bool)> =
matched.iter().map(|m| (m.tmdb_person_id, m.name.clone(), m.adult)).collect();
let unmatched = unmatched.to_vec();
let now = now_iso();
self.db
.write(move |tx| {
let contributor: Option<String> = tx
.query_row(
"SELECT contributor_id FROM manifests WHERE id = ?1",
rusqlite::params![id],
|r| r.get(0),
)
.map_err(anyhow::Error::from)?;
if verdict == Verdict::Rejected {
// §6: the manifest is deleted and the contributor notified
// (via `GET /manifests/{id}/status` until it is gone).
repo::delete_manifest(tx, &id)?;
} else {
// Names come from TMDB, never from the upload (§5a, §7).
for (person_id, name, adult) in &matched {
repo::upsert_person(tx, *person_id, name, *adult, &now)?;
}
// §6: unmatched actors are dropped rather than stored.
for person_id in &unmatched {
repo::delete_manifest_actor(tx, &id, *person_id)?;
}
repo::set_manifest_status(
tx,
&id,
verdict.status(),
reason.as_deref(),
Some(ratio),
)?;
}
if let Some(c) = contributor {
let counter = match verdict {
Verdict::Listed => "accepted",
Verdict::Flagged => "flagged",
Verdict::Rejected => "rejected",
};
repo::bump_contributor_counter(tx, &c, counter)?;
if repo::maybe_revoke_contributor(tx, &c, &now)? {
tracing::warn!(contributor = %c, "revoked token for excessive rejections");
}
}
Ok(())
})
.await
.context("finalising cast check")
}
}
/// Serialisable projection of `Credits` for the cache.
#[derive(serde::Serialize)]
struct SerializableCredits {
cast: Vec<SerializableMember>,
guest_stars: Vec<SerializableMember>,
}
#[derive(serde::Serialize)]
struct SerializableMember {
id: u64,
name: String,
adult: bool,
}
impl From<&Credits> for SerializableCredits {
fn from(c: &Credits) -> Self {
let f =
|m: &CastMember| SerializableMember { id: m.id, name: m.name.clone(), adult: m.adult };
Self {
cast: c.cast.iter().map(f).collect(),
guest_stars: c.guest_stars.iter().map(&f).collect(),
}
}
}
enum JobError {
Retry(String),
Fatal(anyhow::Error),
}
/// Exponential backoff, capped (§5: the client must back off exponentially
/// rather than retrying tightly; the same discipline applies to our own
/// outbound calls).
fn backoff_secs(attempts: i64) -> u64 {
let base = 30u64;
base.saturating_mul(1u64 << attempts.clamp(0, 8) as u32).min(MAX_BACKOFF_SECS)
}
fn is_stale(fetched_at: &str, ttl: Duration) -> bool {
let Some(then) = parse_iso(fetched_at) else { return true };
let now = unix_now();
now.saturating_sub(then) > ttl.as_secs()
}
/// Current time as an RFC 3339 UTC string, which is what every timestamp column
/// stores. Kept in one place so the format cannot drift.
pub fn now_iso() -> String {
iso_in(0)
}
pub fn iso_in(secs: u64) -> String {
format_unix(unix_now() + secs)
}
fn unix_now() -> u64 {
std::time::SystemTime::now()
.duration_since(std::time::UNIX_EPOCH)
.map(|d| d.as_secs())
.unwrap_or(0)
}
/// Formats a Unix timestamp as `YYYY-MM-DDTHH:MM:SSZ`.
///
/// Hand-rolled rather than pulling in `chrono`/`time`: the only requirement is a
/// lexicographically-sortable UTC string, which is what the `jobs.run_after`
/// comparison relies on.
pub fn format_unix(mut secs: u64) -> String {
let days = secs / 86_400;
secs %= 86_400;
let (h, m, s) = (secs / 3600, (secs % 3600) / 60, secs % 60);
// Civil-from-days, Howard Hinnant's algorithm.
let z = days as i64 + 719_468;
let era = z.div_euclid(146_097);
let doe = z.rem_euclid(146_097);
let yoe = (doe - doe / 1460 + doe / 36_524 - doe / 146_096) / 365;
let y = yoe + era * 400;
let doy = doe - (365 * yoe + yoe / 4 - yoe / 100);
let mp = (5 * doy + 2) / 153;
let d = doy - (153 * mp + 2) / 5 + 1;
let mo = if mp < 10 { mp + 3 } else { mp - 9 };
let y = if mo <= 2 { y + 1 } else { y };
format!("{y:04}-{mo:02}-{d:02}T{h:02}:{m:02}:{s:02}Z")
}
fn parse_iso(s: &str) -> Option<u64> {
// Parses the format `format_unix` produces.
let b = s.as_bytes();
if b.len() < 20 {
return None;
}
let num = |from: usize, to: usize| s.get(from..to)?.parse::<i64>().ok();
let (y, mo, d) = (num(0, 4)?, num(5, 7)?, num(8, 10)?);
let (h, mi, se) = (num(11, 13)?, num(14, 16)?, num(17, 19)?);
let y_adj = if mo <= 2 { y - 1 } else { y };
let era = y_adj.div_euclid(400);
let yoe = y_adj - era * 400;
let mp = if mo > 2 { mo - 3 } else { mo + 9 };
let doy = (153 * mp + 2) / 5 + d - 1;
let doe = yoe * 365 + yoe / 4 - yoe / 100 + doy;
let days = era * 146_097 + doe - 719_468;
Some((days * 86_400 + h * 3600 + mi * 60 + se).max(0) as u64)
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn timestamp_roundtrips() {
for t in [0u64, 1, 1_000_000, 1_700_000_000, 1_785_000_000, 4_000_000_000] {
let s = format_unix(t);
assert_eq!(parse_iso(&s), Some(t), "roundtrip failed for {t} => {s}");
}
}
#[test]
fn timestamp_format_is_sortable() {
// `jobs.run_after <= ?1` is a string comparison, so lexical order must
// match chronological order.
let a = format_unix(1_700_000_000);
let b = format_unix(1_700_000_001);
let c = format_unix(1_800_000_000);
assert!(a < b && b < c, "{a} {b} {c}");
assert_eq!(format_unix(0), "1970-01-01T00:00:00Z");
}
#[test]
fn known_dates_format_correctly() {
// 2026-07-30T12:00:00Z
assert_eq!(format_unix(1_785_412_800), "2026-07-30T12:00:00Z");
// A leap day, since the civil-from-days algorithm is where this would
// break.
assert_eq!(format_unix(1_709_164_800), "2024-02-29T00:00:00Z");
}
#[test]
fn backoff_grows_and_is_capped() {
assert_eq!(backoff_secs(0), 30);
assert_eq!(backoff_secs(1), 60);
assert_eq!(backoff_secs(4), 480);
assert_eq!(backoff_secs(50), MAX_BACKOFF_SECS, "must not overflow or grow unbounded");
}
#[test]
fn staleness_uses_the_ttl() {
let fresh = format_unix(unix_now());
assert!(!is_stale(&fresh, CACHE_TTL));
let old = format_unix(unix_now() - 25 * 3600);
assert!(is_stale(&old, CACHE_TTL));
assert!(is_stale("not-a-timestamp", CACHE_TTL));
}
}