fix(tmdb): stop a broken provider tmdb_id from suppressing enrichment (#1239)

* fix(tmdb): stop a broken provider tmdb_id from suppressing enrichment

Providers ship dead and stale tmdb_id values, and enrich() trusted them
unconditionally:

    parseProviderTmdbId(query.tmdbId) ?? await resolveIdBySearch(...)

A garbage-but-integer id short-circuited the title search entirely. The
details fetch then 404'd, the outer catch swallowed it, and the item was
left permanently unenriched — no plot, no cast, no artwork — for a title
the search would have matched. Failed detail fetches cache nothing, so
the wasted request repeated on every re-open. The stale-but-valid case
was worse: it never threw, nothing sanity-checked the resolved title, and
we confidently rendered another film's metadata.

enrich() now treats the provider id as a hint. If it fails to resolve, or
resolves to something whose title matches none of the search variants we
would have queried, the confidence-gated title search gets its turn — and
proven-bad ids are negative-cached (7d, language-independent row) so the
404 is not repeated forever.

Deliberately NOT a hard rejection on title mismatch: TMDB returns titles
in the REQUEST language, so a Russian provider title legitimately fails
the name check against an en-US payload. A mismatch only lets the search
compete; when the search finds nothing confident, the provider payload is
kept. The change can therefore only add enrichment, never remove it.

Extracts the search resolution and the bad-id cache into
TmdbIdResolverService — tmdb-enrichment.service.ts was at 290 lines
against the 300-line target, and the resolver is independently testable.

Tests: new tmdb-enrichment.service.spec.ts covers the happy path issuing
exactly one details call and no search, 404 fallback, stale-id override,
the keep-the-payload safety property, bad-id skip, and the no-match case;
matcher spec covers detailsMatchProviderTitle and the namespaced cache key.

Refs docs/architecture/tmdb-roadmap.md A1.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(tmdb): only blame a provider id when TMDB confirms it does not exist

Review found the bad-id negative cache too eager in two ways, both of
which could deny enrichment to items whose id was fine.

1. Any failure recorded the verdict. A 401, 429, 5xx or an offline blip
   would mark a perfectly valid id as dead for seven days, so after the
   service recovered — or the user fixed their API key — titles that the
   search cannot resolve confidently stayed unenriched until the marker
   expired. TmdbApiService now throws a typed TmdbApiError carrying the
   status, and only a confirmed 404 is recorded.

2. Title mismatches were recorded too. That id EXISTS; it is merely wrong
   for this item. The row is keyed by id alone and shared across
   playlists, so a stale mapping on one item disabled the direct lookup
   for every other item that legitimately used the same id. Mismatches
   are no longer cached at all — the search verdict is cached anyway, so
   the repeat cost is a single details fetch.

Documents the row kind in the cache contract, which listed only two of
the (now six) lookup_key shapes.

Tests: 404 records, 429 does not, network error does not, mismatch does
not.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(tmdb): keep provider details when the competing search fails

detailsForProviderId only runs the search to see whether it can beat a
title-mismatched provider payload. A throw from that best-effort search
(offline, rate limit, 5xx) propagated to enrich()'s outer catch and threw
away details we already had — the searched-details fetch right below it
was already tolerant. Fail to the details in hand instead.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(tmdb): decide a suspect provider id on evidence, not on the title

The title check alone was both too weak and too dangerous.

Too weak: normalizeTitle strips trailing years, so "Blade Runner 2049"
carrying the 1982 film's id matched and the wrong film was rendered —
exactly the stale-id case this was meant to catch.

Too dangerous: an ALL-CAPS leading token reads as a language tag, so
"IT - Chapter Two" normalizes to "chapter two". The correct payload
failed the name check, and a year-less search for "chapter two" would
confidently return the 1979 film and overwrite it. Master trusted the
provider id here and got it right.

assessProviderId weighs both signals: title or year agrees means use the
details; both years known and incompatible means the search may take
over; a title-only mismatch is inconclusive and keeps the details. The
search branch now always has a year, so its own gate corroborates
whatever it returns instead of matching on name alone.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(tmdb): do not search after a transient provider-id failure

enrich() reads a null from detailsForProviderId as "the id is unusable,
try the search". A 401/429/5xx/offline failure gave it that null, so an
outage turned into a second request that would fail too — and if it did
come back, a title match replaced a provider id that was probably fine.
Only a 404 falls through to the search now; everything else rethrows and
leaves the id retryable.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(tmdb): add the release note for the provider-id fix

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
4grayandClaude Opus 4.8 authored and GitHub committed 2026-07-26 00:34:17 +02:00
1 parent b4ec68c1fa
commit d76b2d2a57
11 files changed
+869 -100

No files matched your search

+43 -3
View File
@@ -83,8 +83,29 @@ in the settings section validates the API key against `/configuration`.
Wrong metadata is worse than no metadata, so id resolution is conservative:
1. If the provider returns a usable `tmdb_id` (Xtream VOD info often does),
it is trusted fully and no search runs. Series have no show-level
`tmdb_id`, so they always go through search.
its details are fetched directly and normally used as-is. Series have no
show-level `tmdb_id`, so they always go through search. The id is a
strong hint rather than gospel — panels ship dead and stale ones — so
the payload it returns is weighed against the provider item
(`assessProviderId`):
- a matching title **or** a compatible release year (±1; for series,
any earlier premiere) → **corroborated**, use it;
- both years known and incompatible → **contradicted**, the stale-id
signature ("Blade Runner 2049" carrying the 1982 film's id): the
title search may take over, and does so only if it finds a confident
match of its own;
- title differs with no year to arbitrate → **inconclusive**, keep the
details. TMDB returns titles in the request language and
normalization strips stylized prefixes ("IT - Chapter Two" →
"chapter two"), so a name mismatch alone says more about our inputs
than about the id.
A 404 is the one hard verdict: the id is recorded as dead
(`badProviderId:<id>` row), skipped next time, and the title search
takes over. Transient failures (auth, rate limit, 5xx, offline) neither
disable the id nor trigger a search — that request would hit the same
outage, and a title match that did come back would be weaker evidence
than the id already in hand.
2. Otherwise `/search/movie` (or `/search/tv`) runs with the normalized
title. Normalization strips bracketed tags, quality markers (`4K`,
`1080p`, `MULTI`, …), leading language prefixes (`EN - `), diacritics,
@@ -247,14 +268,18 @@ Filmography has two scopes:
## Cache
Single table with two row kinds discriminated by `lookup_key` prefix:
Single table with several row kinds discriminated by the `lookup_key`
prefix:
```
tmdb_metadata (
media_type 'movie' | 'tv' | 'person',
lookup_key 'id:<tmdbId>|v2' -- details payload row
'id:<tmdbId>|season:<n>' -- season payload row
'title:<normalized>|year:<y>|v2' -- search resolution row
'person:<personId>' -- person payload row
'trending:week' -- trending list row
'badProviderId:<tmdbId>' -- id confirmed 404 by TMDB
language TEXT, -- TMDB language code
tmdb_id INTEGER, -- NULL on a search row = negative cache
payload TEXT, -- raw JSON details, NULL for search rows
@@ -343,3 +368,18 @@ dashboard rail, artwork upgrade for M3U VOD, persistent PWA cache
heroes additionally show the tracked "S{n}·E{n}" badge from the playback
position (no TMDB involved); the watch-progress bar is limited to
movie/series heroes.
### `badProviderId:` rows
Providers ship `tmdb_id` values that do not exist. A failed details fetch
caches nothing, so without a marker the same 404 is re-issued on every
detail open, forever. These rows record that verdict: `tmdb_id` NULL,
language `any` (a dead id is dead in every language), read with the
negative-match TTL.
Only a **confirmed 404** is recorded. Transient failures (401, 429, 5xx,
offline) leave no marker — they say nothing about the id. Neither does a
title mismatch: that id exists and may be correct for a *different* item,
and since the row is keyed by id alone and shared across playlists,
recording per-item mismatches here would deny the direct lookup to every
other item that legitimately uses the same id.