fix(tmdb): a new series no longer matches its older, better-known namesake (#1648)

Metadata is looked up in the app's own language, so an unrelated older foreign
series can come back under exactly the same localized name as a recent
local-language one. `pickConfidentMatch` admitted the older row through the
series "premiered earlier" tolerance — portals report the running season's
year while TMDB reports the premiere — and then let `pickMostPopular` decide
across every admitted candidate, discarding the year evidence that had just
admitted them. The better-known show won on votes, and the newer series
rendered its poster, cast, genres and rating.

Rank admitted candidates by year evidence first (`yearEvidenceTier`: the
provider's exact year, then a year off by one, then the series tolerance) and
let popularity break ties only inside the strongest tier any candidate
reached. The tolerance stays — three of eight real lookups from one install
depend on it — but it is a last resort, not an equal. Measured over 400
Cyrillic series titles sampled from a real catalog, 20 normalized keys had a
same-titled older series and 16 of those were the more popular row.

The mirror case is accepted knowingly: a long-running show whose stated
season year happens to BE another same-titled show's premiere year now
resolves to the newer show. Only the older show's season air dates could
separate the two and a search response does not carry them, while that shape
needs three coincidences at once against one that needs none.

Search cache keys move to `|v4` with a matching startup cleanup, because a
positive row naming the wrong show stays fresh for 30 days.

Merged with `Build on windows x64` red: the checked-in Windows Embedded MPV
runtime pin points at an upstream release whose retention expired, so that
job fails repository-wide on a cold cache. Unrelated to this change; tracked
separately.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
4grayandClaude Opus 5 authored and GitHub committed 2026-09-21 08:18:56 +02:00
1 parent 42b41ebabc
commit b02d79805b
9 files changed
+267 -65

No files matched your search

+10
View File
@@ -0,0 +1,10 @@
---
type: fix
area: tmdb
---
A new series no longer picks up the poster, cast and plot of an older show
that happens to share its title. Metadata is looked up in your own language,
so an unrelated foreign series can be listed under the very same name — and
the better-known one used to win. The release year the provider states now
decides first.
+2 -2
View File
@@ -1630,9 +1630,9 @@ stream_id`); it drops `series_id`/`movie_id`, so the builder pins the
- Actor page "All portals" scope (Electron only): batched `DB_MATCH_TITLES` worker op (trigram FTS over all imported Xtream playlists, `apps/electron-backend/src/app/database/operations/title-match.operations.ts`); `normalizeTitle` is shared renderer/worker via `libs/shared/interfaces/src/lib/title-normalization.util.ts`
- All `DB_MATCH_TITLES` consumers (Trending rail, "Because you watched" recommendations rail, cross-portal Similar rail, actor "All portals" scope) resolve the worker's flat result list through the shared `groupTitleMatchesByKey()` + `pickTitleMatch()` in `libs/services/src/lib/catalog-title-match.service.ts`. The grouping keeps EVERY row per `type:exactNormalizedTitle` on purpose — the year that separates same-titled rows belongs to the lookup, which the grouping cannot see, so collapsing first made a catalog holding both "Dune 1984" and "Dune 2021" drop whichever copy the user actually owns. `pickTitleMatch` then ranks year-compatible rows by evidence (exact year → untagged → any compatible) across all title aliases at once; only the recommendations rail passes an alias (TMDB `original_title`, via `candidateLookup()`). Multi-source VOD discovery deliberately stays off these helpers: there every copy is a distinct selectable source, not one best answer
- Opt-in via `Settings > Metadata (TMDB)` (sends titles to TMDB); the section also has a "check key" button and a cache panel (row count + payload size, with a clear button). Distributed builds ship without a shared key; users supply their own. `DEFAULT_TMDB_API_KEY` in `libs/services/src/lib/tmdb/tmdb-config.ts` is empty by default; `tools/tmdb/inject-tmdb-key.mjs` still supports optional CI injection from `TMDB_API_KEY`, but is a no-op when it is unset. A user key takes precedence over any injected default; without either key, enrichment stays inactive even when enabled. Requiring a personal key is project policy, not a categorical TMDB terms restriction; see the canonical "Settings and API Key" section in `docs/architecture/tmdb-metadata-enrichment.md`.
- Match confidence: a provider `tmdb_id` is a strong hint, not gospel — its payload is weighed against the item (`assessProviderId`: title or year agrees → use it; both years known and incompatible → the search may take over; title-only mismatch → keep it, since TMDB localizes titles). A 404 marks the id dead (`badProviderId:<id>` row); transient failures never do. Without a usable id: normalized-title + year (±1) search with a strict gate — no confident match means no enrichment
- Match confidence: a provider `tmdb_id` is a strong hint, not gospel — its payload is weighed against the item (`assessProviderId`: title or year agrees → use it; both years known and incompatible → the search may take over; title-only mismatch → keep it, since TMDB localizes titles). A 404 marks the id dead (`badProviderId:<id>` row); transient failures never do. Without a usable id: normalized-title + year (±1) search with a strict gate — no confident match means no enrichment. Several admitted results are ranked by year evidence FIRST (`yearEvidenceTier` in `tmdb-matcher.ts`: provider's exact year → off by one → the series "premiered earlier" tolerance), and only tie-broken by `vote_count`/`popularity` inside the strongest tier reached. That tolerance covers portals reporting the running season's year, but it is a last resort: ranked as an equal it handed every new series its older, better-known namesake, because TMDB returns titles in the REQUEST language (an unrelated 2018 foreign series comes back under the same `ru-RU` name as a 2026 local-language series and outvotes it — 20 of 400 sampled Cyrillic series titles had such a collision, 16 of them won by the older row). The mirror case is accepted knowingly: only the older show's season air dates could separate it, and a search response does not carry them
- Detail views render provider data immediately; enrichment patches the selection asynchronously (staleness-guarded)
- Cached in SQLite `tmdb_metadata` (Electron, via DB worker ops `DB_GET/SET_TMDB_METADATA`, plus `DB_GET_TMDB_CACHE_STATS` / `DB_CLEAR_TMDB_METADATA` behind the settings cache panel) or in-memory (PWA); localized via the app language setting. Search-match lookup keys are versioned (`|v3`), and connection startup removes the rows of every retired generation once, each under its own app-state marker (`migration:tmdb-search-lookup-v2-cache-cleanup:v1` for unversioned rows, `migration:tmdb-search-lookup-v3-cache-cleanup:v1` for `|v2` rows). The TMDB search query is `cleanTitleForSearch` (provider spelling, tags/brackets/season/year stripped), NOT the folded `normalizeTitle` key used by `pickConfidentMatch`; variant deduplication uses the lowercased query (`searchQueryIdentity`) and every attempted variant is cached under its own key, since "Феик"/"Фейк" fold to one key but are different searches and two items can share an original title while walking different variant lists: NFD folding rewrites Cyrillic "й"→"и" and "ё"→"е" and splits Arabic hamza forms, and TMDB answers such a query with nothing ("Фейк (10 серий)" never matched). Never send the folded key over the wire.
- Cached in SQLite `tmdb_metadata` (Electron, via DB worker ops `DB_GET/SET_TMDB_METADATA`, plus `DB_GET_TMDB_CACHE_STATS` / `DB_CLEAR_TMDB_METADATA` behind the settings cache panel) or in-memory (PWA); localized via the app language setting. Search-match lookup keys are versioned (`|v4`), and connection startup removes the rows of every retired generation once, each under its own app-state marker (`migration:tmdb-search-lookup-v2-cache-cleanup:v1` for unversioned rows, `…-v3-…` for `|v2` rows, `…-v4-…` for the `|v3` rows resolved before year evidence was tiered). The TMDB search query is `cleanTitleForSearch` (provider spelling, tags/brackets/season/year stripped), NOT the folded `normalizeTitle` key used by `pickConfidentMatch`; variant deduplication uses the lowercased query (`searchQueryIdentity`) and every attempted variant is cached under its own key, since "Феик"/"Фейк" fold to one key but are different searches and two items can share an original title while walking different variant lists: NFD folding rewrites Cyrillic "й"→"и" and "ё"→"е" and splits Arabic hamza forms, and TMDB answers such a query with nothing ("Фейк (10 серий)" never matched). Never send the folded key over the wire.
- Service layer: `libs/services/src/lib/tmdb/`; store glue: `libs/portal/xtream/data-access/src/lib/stores/xtream-tmdb-enrichment.ts` and `libs/portal/stalker/data-access/src/lib/stores/stalker-tmdb-enrichment.ts` (hooked in `withStalkerSelection().setSelectedItem`)
- TMDB attribution (logo + disclaimer) is required and shown in the settings TMDB section and About
- See `docs/architecture/tmdb-metadata-enrichment.md`
+30 -8
View File
@@ -141,7 +141,27 @@ Wrong metadata is worse than no metadata, so id resolution is conservative:
provider year — portals report the current season's year while TMDB's
`first_air_date` is the premiere. Without a year, the exact-title match
must be unambiguous (single hit).
4. No confident match → the provider data stays untouched, and the negative
4. Several admitted results are ranked by **year evidence first**
(`yearEvidenceTier`): the provider's exact year beats a year off by one,
which beats the series "premiered earlier" tolerance. Popularity
(`vote_count`, then `popularity`) only breaks ties inside the strongest
tier any candidate reached.
The tolerance is a last resort, not an equal. Ranked alongside the
exact-year tier it handed every new series its older, better-known
namesake: because TMDB returns titles in the request language, an
unrelated 2018 foreign series came back under the same `ru-RU` name as a
2026 local-language series and outvoted it, rendering its poster, cast
and genres. Over 400 Cyrillic series titles sampled from a real catalog,
20 normalized keys had a same-titled older series and 16 of those were
the more popular row.
The mirror case is accepted knowingly: a long-running show whose stated
season year happens to BE another same-titled show's premiere year now
resolves to the newer show. Separating those two needs the older show's
season air dates, which a search response does not carry — and the shape
needs three coincidences at once, against one that needs none.
5. No confident match → the provider data stays untouched, and the negative
verdict is cached (shorter TTL) so browsing back doesn't re-search.
The year filter is applied client-side rather than via TMDB's strict
@@ -475,7 +495,7 @@ tmdb_metadata (
media_type 'movie' | 'tv' | 'person',
lookup_key 'id:<tmdbId>|v2' -- details payload row
'id:<tmdbId>|season:<n>' -- season payload row
'title:<query lowercased>|year:<y>|v3' -- search resolution row
'title:<query lowercased>|year:<y>|v4' -- search resolution row
'person:<personId>' -- person payload row
'trending:week' -- trending list row
'badProviderId:<tmdbId>' -- id confirmed 404 by TMDB
@@ -490,19 +510,21 @@ tmdb_metadata (
TTLs (enforced at read time in `TmdbCacheService.isFresh`): details and
positive matches 30 days, negative matches 7 days.
Search keys carry a `|v3` and details keys a `|v2` version suffix
Search keys carry a `|v4` and details keys a `|v2` version suffix
(`buildSearchLookupKey` / `buildDetailsLookupKey` in `tmdb-matcher.ts`): for
search rows so normalization or query changes cannot reuse stale positive or
negative resolutions, for details rows because payloads now include videos
via `append_to_response` and pre-videos cache rows had to be invalidated.
Search v2 → v3 retired the rows written while the folded comparison key was
also the wire query (see "Search query vs. comparison key" below). Database
also the wire query (see "Search query vs. comparison key" below); v3 → v4
retired the rows resolved before year evidence was tiered, where a positive
row naming the wrong show would otherwise stay fresh for 30 days. Database
startup deletes the rows of every retired search-key generation once, each
under its own `app_state` marker so a skipped release still runs the cleanups
it missed (`migration:tmdb-search-lookup-v2-cache-cleanup:v1` for the
unversioned rows, `migration:tmdb-search-lookup-v3-cache-cleanup:v1` for the
`|v2` rows; `LEGACY_TMDB_SEARCH_CACHE_CLEANUPS` in `connection.ts`); details
and person cache rows are unaffected.
unversioned rows, `…-v3-…` for the `|v2` rows, `…-v4-…` for the `|v3` rows;
`LEGACY_TMDB_SEARCH_CACHE_CLEANUPS` in `connection.ts`); details and person
cache rows are unaffected.
### Search query vs. comparison key
@@ -513,7 +535,7 @@ both sides fold the same way. `query` is `cleanTitleForSearch`
stripping, but the letters left as the provider wrote them. The search's
identity is the query with only its case removed (`searchQueryIdentity`):
variants are deduplicated by it and every attempted variant is cached under
its own key (`title:<identity>|year:<y>|v3`, in the language that variant
its own key (`title:<identity>|year:<y>|v4`, in the language that variant
was searched in), never by the folded key and never only under the first
variant — "Феик" and "Фейк" fold to one key but are different searches with
different answers, so a verdict for one must not be read back for the other;
@@ -87,14 +87,14 @@ describe('TmdbIdResolverService.resolveBySearch', () => {
expect(searchTv).toHaveBeenCalledWith('Фейк', null, 'ru-RU', 'key');
expect(cacheSet).toHaveBeenCalledWith({
mediaType: 'tv',
lookupKey: 'title:фейк|year:2026|v3',
lookupKey: 'title:фейк|year:2026|v4',
language: 'ru-RU',
tmdbId: 317869,
payload: null,
});
});
it('caches the miss under the folded v3 key', async () => {
it('caches the miss under the current key, not the folded one', async () => {
searchTv.mockResolvedValue([]);
const service = createService();
@@ -112,7 +112,7 @@ describe('TmdbIdResolverService.resolveBySearch', () => {
);
expect(cacheSet).toHaveBeenCalledWith(
expect.objectContaining({
lookupKey: 'title:молодой шерлок|year:2026|v3',
lookupKey: 'title:молодой шерлок|year:2026|v4',
tmdbId: null,
})
);
@@ -133,7 +133,7 @@ describe('TmdbIdResolverService.resolveBySearch', () => {
expect(id).toBe(317869);
expect(cacheGet).toHaveBeenCalledWith(
'tv',
'title:фейк|year:2026|v3',
'title:фейк|year:2026|v4',
'ru-RU'
);
expect(searchTv).not.toHaveBeenCalled();
@@ -157,8 +157,8 @@ describe('TmdbIdResolverService.resolveBySearch', () => {
expect(
cacheSet.mock.calls.map(([row]) => [row.lookupKey, row.tmdbId])
).toEqual([
['title:феик|year:2026|v3', null],
['title:фейк|year:2026|v3', 317869],
['title:феик|year:2026|v4', null],
['title:фейк|year:2026|v4', 317869],
]);
});
@@ -167,7 +167,7 @@ describe('TmdbIdResolverService.resolveBySearch', () => {
// "феик". Item B shares the original title but displays "Фейк":
// the cached miss must not suppress B's own second variant.
cacheGet.mockImplementation(async (_type: string, key: string) =>
key === 'title:феик|year:2026|v3' ? { tmdbId: null } : null
key === 'title:феик|year:2026|v4' ? { tmdbId: null } : null
);
const service = createService();
const cacheService = (service as unknown as { cache: TmdbCacheService })
@@ -186,8 +186,8 @@ describe('TmdbIdResolverService.resolveBySearch', () => {
expect(searchTv).toHaveBeenCalledTimes(1);
expect(searchTv).toHaveBeenCalledWith('Фейк', null, 'ru-RU', 'key');
expect(cacheGet.mock.calls.map(([, key]) => key)).toEqual([
'title:феик|year:2026|v3',
'title:фейк|year:2026|v3',
'title:феик|year:2026|v4',
'title:фейк|year:2026|v4',
]);
});
@@ -80,10 +80,10 @@ describe('extractYear', () => {
describe('lookup keys', () => {
it('builds stable search and details keys', () => {
expect(buildSearchLookupKey('The Matrix', 1999)).toBe(
'title:the matrix|year:1999|v3'
'title:the matrix|year:1999|v4'
);
expect(buildSearchLookupKey('The Matrix', null)).toBe(
'title:the matrix|year:|v3'
'title:the matrix|year:|v4'
);
});
@@ -92,7 +92,7 @@ describe('lookup keys', () => {
// searches with different answers; a verdict cached for one must
// never be read back for the other.
expect(buildSearchLookupKey('Фейк', 2026)).toBe(
'title:фейк|year:2026|v3'
'title:фейк|year:2026|v4'
);
expect(buildSearchLookupKey('Феик', 2026)).not.toBe(
buildSearchLookupKey('Фейк', 2026)
@@ -329,6 +329,84 @@ describe('pickConfidentMatch', () => {
).toBe(theBoys);
});
it('prefers an exact-year series over an older, more popular namesake', () => {
// Regression: TMDB returns titles in the REQUEST language, so an
// unrelated older foreign series (2018, 26 votes) came back under the
// same localized name as the catalog's own 2026 series (4 votes). The
// older row is admitted only by the running-season tolerance and used
// to win the popularity tiebreak. Titles here are stand-ins.
const translatedNamesake: TmdbSearchResult = {
id: 1,
name: 'Nightfall',
original_name: 'Yoru no Tobari',
first_air_date: '2018-12-14',
vote_count: 26,
popularity: 3.4,
};
const localSeries: TmdbSearchResult = {
id: 2,
name: 'Nightfall',
original_name: 'Nightfall',
first_air_date: '2026-07-24',
vote_count: 4,
popularity: 2.4,
};
expect(
pickConfidentMatch(
[localSeries, translatedNamesake],
{ title: 'Nightfall (12 episodes)', year: 2026 },
'tv'
)
).toBe(localSeries);
});
it('still falls back to an older series when nothing matches the year', () => {
const theBoys: TmdbSearchResult = {
id: 76479,
name: 'The Boys',
first_air_date: '2019-07-26',
vote_count: 12000,
};
const unrelatedOlder: TmdbSearchResult = {
id: 999,
name: 'The Boys',
first_air_date: '2010-01-01',
vote_count: 5,
};
expect(
pickConfidentMatch(
[unrelatedOlder, theBoys],
{ title: 'The Boys s05', year: 2026 },
'tv'
)
).toBe(theBoys);
});
it('prefers the exact release year over one off by one', () => {
const exact: TmdbSearchResult = {
id: 1,
title: 'The Matrix',
release_date: '1999-03-31',
vote_count: 10,
};
const adjacent: TmdbSearchResult = {
id: 2,
title: 'The Matrix',
release_date: '1998-03-31',
vote_count: 9000,
};
expect(
pickConfidentMatch(
[adjacent, exact],
{ title: 'The Matrix', year: 1999 },
'movie'
)
).toBe(exact);
});
it('still rejects movies with a year that differs by more than one', () => {
expect(
pickConfidentMatch(
+76 -16
View File
@@ -121,7 +121,10 @@ export function buildSearchLookupKey(
// under v2 for a title with "й"/"ё" is that bug, not a missing title, and
// must not block the retry for its 7-day TTL. Rows are keyed by the
// query since then.
return `title:${searchQueryIdentity(query)}|year:${year ?? ''}|v3`;
// v4: year evidence is tiered (see `yearEvidenceTier`), so every v3 row
// resolved by popularity across tiers may name the wrong show — and a
// positive row stays fresh for 30 days.
return `title:${searchQueryIdentity(query)}|year:${year ?? ''}|v4`;
}
export function buildDetailsLookupKey(tmdbId: number): string {
@@ -256,6 +259,52 @@ function resultYear(
);
}
/**
* How strongly one candidate's own year backs the year the provider stated.
* Lower is stronger; `null` means the candidate is not admissible at all.
*
* The series tier is what makes long-running shows work: a portal reports
* the CURRENT season's year ("The Boys s05" → 2026) while TMDB's
* `first_air_date` is the 2019 premiere. It is a last resort, though, not an
* equal — ranked alongside the exact-year tier with popularity deciding, it
* hands every NEW series its older, better-known namesake. TMDB returns
* titles in the REQUEST language, so in a non-English catalog those
* collisions are routine rather than exotic: a 2026 local-language drama
* (4 votes) lost to an unrelated 2018 foreign show TMDB lists under the same
* localized name (26 votes), and rendered its poster, cast and genres. Over
* 400 Cyrillic series titles sampled from a real catalog, 20 normalized keys
* had a same-titled older series and 16 of those were the more popular row.
*
* The mirror case survives on purpose: a long-running show whose stated
* season year happens to BE another same-titled show's premiere year now
* resolves to the newer show. Only the older show's season air dates could
* separate the two, and a search response does not carry them — while the
* shape needs three coincidences at once, against one that needs none.
*/
const YEAR_TIER_EXACT = 0;
const YEAR_TIER_ADJACENT = 1;
const YEAR_TIER_EARLIER_SERIES = 2;
function yearEvidenceTier(
year: number | null,
wantedYear: number,
mediaType: TmdbMediaType
): number | null {
if (year === null) {
return null;
}
if (year === wantedYear) {
return YEAR_TIER_EXACT;
}
if (Math.abs(year - wantedYear) === 1) {
return YEAR_TIER_ADJACENT;
}
return mediaType === 'tv' && year < wantedYear
? YEAR_TIER_EARLIER_SERIES
: null;
}
/**
* Pick the search result that confidently matches the queried title/year.
* Returns `null` when confidence is insufficient — enrichment must never
@@ -283,25 +332,36 @@ export function pickConfidentMatch(
const wantedYear = query.year;
if (wantedYear !== null) {
const yearMatches = exactTitleMatches.filter((result) => {
const year = resultYear(result, mediaType);
if (year === null) {
return false;
}
if (Math.abs(year - wantedYear) <= 1) {
return true;
}
// Series: providers often report the CURRENT season's year
// ("The Boys s05" → 2026) while TMDB's first_air_date is the
// show's premiere (2019) — accept shows that started earlier.
return mediaType === 'tv' && year < wantedYear;
});
// Popularity only breaks ties INSIDE the strongest tier any
// candidate reached — see `yearEvidenceTier`.
const ranked = exactTitleMatches
.map((result) => ({
result,
tier: yearEvidenceTier(
resultYear(result, mediaType),
wantedYear,
mediaType
),
}))
.filter(
(
candidate
): candidate is {
result: TmdbSearchResult;
tier: number;
} => candidate.tier !== null
);
if (yearMatches.length === 0) {
if (ranked.length === 0) {
return null;
}
return pickMostPopular(yearMatches);
const bestTier = Math.min(...ranked.map((candidate) => candidate.tier));
return pickMostPopular(
ranked
.filter((candidate) => candidate.tier === bestTier)
.map((candidate) => candidate.result)
);
}
// Without a year the title must be unambiguous
@@ -159,20 +159,24 @@ describe('TMDB search lookup cache cleanup', () => {
cleanupLegacyTmdbSearchCache(sqlite);
expect(transaction).toHaveBeenCalledTimes(2);
expect(deleteRun).toHaveBeenCalledTimes(2);
expect(transaction).toHaveBeenCalledTimes(3);
expect(deleteRun).toHaveBeenCalledTimes(3);
expect(markerRun.mock.calls.map(([key]) => key)).toEqual([
'migration:tmdb-search-lookup-v2-cache-cleanup:v1',
'migration:tmdb-search-lookup-v3-cache-cleanup:v1',
'migration:tmdb-search-lookup-v4-cache-cleanup:v1',
]);
const [unversionedDelete, v2Delete] = deleteStatements(sqlite);
const [unversionedDelete, v2Delete, v3Delete] =
deleteStatements(sqlite);
expect(unversionedDelete).toContain(
"lookup_key LIKE 'title:%|year:%' AND lookup_key NOT LIKE 'title:%|year:%|v%'"
);
// Only v2 search rows: the v3 rows the resolver writes now, and the
// `id:`/`person:`/`badProviderId:` rows, must survive.
// Each generation deletes only its own rows: the v4 rows the resolver
// writes now, and the `id:`/`person:`/`badProviderId:` rows, survive.
expect(v2Delete).toContain("WHERE lookup_key LIKE 'title:%|year:%|v2'");
expect(v2Delete).not.toContain('v3');
expect(v3Delete).toContain("WHERE lookup_key LIKE 'title:%|year:%|v3'");
expect(v3Delete).not.toContain('v4');
});
it('runs only the generations that have not completed yet', () => {
@@ -193,13 +197,14 @@ describe('TMDB search lookup cache cleanup', () => {
cleanupLegacyTmdbSearchCache(sqlite);
expect(transaction).toHaveBeenCalledTimes(1);
expect(markerRun).toHaveBeenCalledTimes(1);
expect(markerRun).toHaveBeenCalledWith(
'migration:tmdb-search-lookup-v3-cache-cleanup:v1'
);
expect(transaction).toHaveBeenCalledTimes(2);
expect(markerRun.mock.calls.map(([key]) => key)).toEqual([
'migration:tmdb-search-lookup-v3-cache-cleanup:v1',
'migration:tmdb-search-lookup-v4-cache-cleanup:v1',
]);
expect(deleteStatements(sqlite)).toEqual([
expect.stringContaining("lookup_key LIKE 'title:%|year:%|v2'"),
expect.stringContaining("lookup_key LIKE 'title:%|year:%|v3'"),
]);
});
@@ -212,7 +217,7 @@ describe('TMDB search lookup cache cleanup', () => {
expect(transaction).not.toHaveBeenCalled();
// One marker read per retired generation, nothing else
expect(prepare).toHaveBeenCalledTimes(2);
expect(prepare).toHaveBeenCalledTimes(3);
});
});
@@ -45,6 +45,8 @@ const TMDB_SEARCH_LOOKUP_V2_CACHE_CLEANUP_MIGRATION_KEY =
'migration:tmdb-search-lookup-v2-cache-cleanup:v1';
const TMDB_SEARCH_LOOKUP_V3_CACHE_CLEANUP_MIGRATION_KEY =
'migration:tmdb-search-lookup-v3-cache-cleanup:v1';
const TMDB_SEARCH_LOOKUP_V4_CACHE_CLEANUP_MIGRATION_KEY =
'migration:tmdb-search-lookup-v4-cache-cleanup:v1';
const EPG_PROGRAM_SOURCE_URL_BACKFILL_BATCH_SIZE = 50_000;
function readTraceFlag(name: string): boolean {
@@ -977,6 +979,10 @@ function widenTmdbMetadataMediaTypeCheck(sqliteDb: Database.Database): void {
* v2 every title with a Cyrillic "й"/"ё" was searched folded ("феик" for
* "Фейк", "елки" for "Ёлки"), got no answer, and was cached as missing for
* 7 days.
* - v3 → v4: year evidence became tiered. Under v3 a series admitted only by
* the "premiered earlier" tolerance competed with an exact-year match on
* popularity alone, so a new series resolved to its older, better-known
* namesake — and that positive row stays fresh for 30 days.
*/
const LEGACY_TMDB_SEARCH_CACHE_CLEANUPS: ReadonlyArray<{
migrationKey: string;
@@ -991,6 +997,10 @@ const LEGACY_TMDB_SEARCH_CACHE_CLEANUPS: ReadonlyArray<{
migrationKey: TMDB_SEARCH_LOOKUP_V3_CACHE_CLEANUP_MIGRATION_KEY,
rowPredicate: `lookup_key LIKE 'title:%|year:%|v2'`,
},
{
migrationKey: TMDB_SEARCH_LOOKUP_V4_CACHE_CLEANUP_MIGRATION_KEY,
rowPredicate: `lookup_key LIKE 'title:%|year:%|v3'`,
},
];
/**
@@ -28,11 +28,13 @@ it('drops only retired search rows across skipped, previous, pre-person, fresh a
console.log = () => undefined;
const V2_MARKER = 'migration:tmdb-search-lookup-v2-cache-cleanup:v1';
const V3_MARKER = 'migration:tmdb-search-lookup-v3-cache-cleanup:v1';
const V4_MARKER = 'migration:tmdb-search-lookup-v4-cache-cleanup:v1';
const ROWS = [
['tv', 'title:феик|year:2026', 'ru-RU', null],
['tv', 'title:феик|year:2026|v2', 'ru-RU', null],
['tv', 'title:the boys|year:2019|v2', 'en-US', 76479],
['tv', 'title:фейк|year:2026|v3', 'ru-RU', 317869],
['tv', 'title:nightfall|year:2026|v4', 'ru-RU', 424242],
['tv', 'id:317869|v2', 'ru-RU', 317869],
['tv', 'id:317869|season:1', 'ru-RU', 317869],
['person', 'person:287', 'en-US', 287],
@@ -61,11 +63,11 @@ it('drops only retired search rows across skipped, previous, pre-person, fresh a
hooks.runMigrations(skipped);
const skippedAfter = snapshot(skipped);
// Next startup: a row written meanwhile under the current key survives
skipped.prepare("INSERT INTO tmdb_metadata (media_type, lookup_key, language, tmdb_id) VALUES ('tv', 'title:гудовы|year:2026|v3', 'ru-RU', 318894)").run();
skipped.prepare("INSERT INTO tmdb_metadata (media_type, lookup_key, language, tmdb_id) VALUES ('tv', 'title:гудовы|year:2026|v4', 'ru-RU', 318894)").run();
hooks.runMigrations(skipped);
const repeated = snapshot(skipped);
// Previous release: the unversioned cleanup already ran; only v2 rows go
const previous = openDb([V2_MARKER]);
// Previous release: the earlier cleanups already ran; only v3 rows go
const previous = openDb([V2_MARKER, V3_MARKER]);
hooks.runMigrations(previous);
const previousAfter = snapshot(previous);
hooks.runMigrations(previous);
@@ -104,12 +106,13 @@ it('drops only retired search rows across skipped, previous, pre-person, fresh a
const V2_MARKER = 'migration:tmdb-search-lookup-v2-cache-cleanup:v1';
const V3_MARKER = 'migration:tmdb-search-lookup-v3-cache-cleanup:v1';
const V4_MARKER = 'migration:tmdb-search-lookup-v4-cache-cleanup:v1';
const survivors = [
'badProviderId:999',
'id:317869|season:1',
'id:317869|v2',
'person:287',
'title:фейк|year:2026|v3',
'title:nightfall|year:2026|v4',
'trending:week',
];
const detailsPayloads = ['{"id":317869}', '{"id":317869}'];
@@ -118,40 +121,54 @@ it('drops only retired search rows across skipped, previous, pre-person, fresh a
skippedAfter: {
keys: survivors,
payloads: detailsPayloads,
markers: [V2_MARKER, V3_MARKER],
markers: [V2_MARKER, V3_MARKER, V4_MARKER],
},
repeated: {
keys: [...survivors, 'title:гудовы|year:2026|v3'].sort(),
keys: [...survivors, 'title:гудовы|year:2026|v4'].sort(),
payloads: detailsPayloads,
markers: [V2_MARKER, V3_MARKER],
markers: [V2_MARKER, V3_MARKER, V4_MARKER],
},
previousAfter: {
// The unversioned row is that generation's business, already done
keys: [...survivors, 'title:феик|year:2026'].sort(),
// The older generations' rows are their business, already done
keys: [
...survivors,
'title:феик|year:2026',
'title:феик|year:2026|v2',
'title:the boys|year:2019|v2',
].sort(),
payloads: detailsPayloads,
markers: [V2_MARKER, V3_MARKER],
markers: [V2_MARKER, V3_MARKER, V4_MARKER],
},
previousRepeated: {
keys: [...survivors, 'title:феик|year:2026'].sort(),
keys: [
...survivors,
'title:феик|year:2026',
'title:феик|year:2026|v2',
'title:the boys|year:2019|v2',
].sort(),
payloads: detailsPayloads,
markers: [V2_MARKER, V3_MARKER],
markers: [V2_MARKER, V3_MARKER, V4_MARKER],
},
prePersonAfter: {
keys: [],
payloads: [],
markers: [V2_MARKER, V3_MARKER],
markers: [V2_MARKER, V3_MARKER, V4_MARKER],
check: true,
},
prePersonRepeated: {
keys: [],
payloads: [],
markers: [V2_MARKER, V3_MARKER],
markers: [V2_MARKER, V3_MARKER, V4_MARKER],
},
freshAfter: {
keys: [],
payloads: [],
markers: [V2_MARKER, V3_MARKER, V4_MARKER],
},
freshAfter: { keys: [], payloads: [], markers: [V2_MARKER, V3_MARKER] },
freshRepeated: {
keys: [],
payloads: [],
markers: [V2_MARKER, V3_MARKER],
markers: [V2_MARKER, V3_MARKER, V4_MARKER],
},
});
});