diff --git a/.changes/tmdb-new-series-year-match.md b/.changes/tmdb-new-series-year-match.md new file mode 100644 index 000000000..89c8bb995 --- /dev/null +++ b/.changes/tmdb-new-series-year-match.md @@ -0,0 +1,10 @@ +--- +type: fix +area: tmdb +--- + +A new series no longer picks up the poster, cast and plot of an older show +that happens to share its title. Metadata is looked up in your own language, +so an unrelated foreign series can be listed under the very same name — and +the better-known one used to win. The release year the provider states now +decides first. diff --git a/CLAUDE.md b/CLAUDE.md index c2022dc4e..01db26830 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -1630,9 +1630,9 @@ stream_id`); it drops `series_id`/`movie_id`, so the builder pins the - Actor page "All portals" scope (Electron only): batched `DB_MATCH_TITLES` worker op (trigram FTS over all imported Xtream playlists, `apps/electron-backend/src/app/database/operations/title-match.operations.ts`); `normalizeTitle` is shared renderer/worker via `libs/shared/interfaces/src/lib/title-normalization.util.ts` - All `DB_MATCH_TITLES` consumers (Trending rail, "Because you watched" recommendations rail, cross-portal Similar rail, actor "All portals" scope) resolve the worker's flat result list through the shared `groupTitleMatchesByKey()` + `pickTitleMatch()` in `libs/services/src/lib/catalog-title-match.service.ts`. The grouping keeps EVERY row per `type:exactNormalizedTitle` on purpose — the year that separates same-titled rows belongs to the lookup, which the grouping cannot see, so collapsing first made a catalog holding both "Dune 1984" and "Dune 2021" drop whichever copy the user actually owns. `pickTitleMatch` then ranks year-compatible rows by evidence (exact year → untagged → any compatible) across all title aliases at once; only the recommendations rail passes an alias (TMDB `original_title`, via `candidateLookup()`). Multi-source VOD discovery deliberately stays off these helpers: there every copy is a distinct selectable source, not one best answer - Opt-in via `Settings > Metadata (TMDB)` (sends titles to TMDB); the section also has a "check key" button and a cache panel (row count + payload size, with a clear button). Distributed builds ship without a shared key; users supply their own. `DEFAULT_TMDB_API_KEY` in `libs/services/src/lib/tmdb/tmdb-config.ts` is empty by default; `tools/tmdb/inject-tmdb-key.mjs` still supports optional CI injection from `TMDB_API_KEY`, but is a no-op when it is unset. A user key takes precedence over any injected default; without either key, enrichment stays inactive even when enabled. Requiring a personal key is project policy, not a categorical TMDB terms restriction; see the canonical "Settings and API Key" section in `docs/architecture/tmdb-metadata-enrichment.md`. -- Match confidence: a provider `tmdb_id` is a strong hint, not gospel — its payload is weighed against the item (`assessProviderId`: title or year agrees → use it; both years known and incompatible → the search may take over; title-only mismatch → keep it, since TMDB localizes titles). A 404 marks the id dead (`badProviderId:` row); transient failures never do. Without a usable id: normalized-title + year (±1) search with a strict gate — no confident match means no enrichment +- Match confidence: a provider `tmdb_id` is a strong hint, not gospel — its payload is weighed against the item (`assessProviderId`: title or year agrees → use it; both years known and incompatible → the search may take over; title-only mismatch → keep it, since TMDB localizes titles). A 404 marks the id dead (`badProviderId:` row); transient failures never do. Without a usable id: normalized-title + year (±1) search with a strict gate — no confident match means no enrichment. Several admitted results are ranked by year evidence FIRST (`yearEvidenceTier` in `tmdb-matcher.ts`: provider's exact year → off by one → the series "premiered earlier" tolerance), and only tie-broken by `vote_count`/`popularity` inside the strongest tier reached. That tolerance covers portals reporting the running season's year, but it is a last resort: ranked as an equal it handed every new series its older, better-known namesake, because TMDB returns titles in the REQUEST language (an unrelated 2018 foreign series comes back under the same `ru-RU` name as a 2026 local-language series and outvotes it — 20 of 400 sampled Cyrillic series titles had such a collision, 16 of them won by the older row). The mirror case is accepted knowingly: only the older show's season air dates could separate it, and a search response does not carry them - Detail views render provider data immediately; enrichment patches the selection asynchronously (staleness-guarded) -- Cached in SQLite `tmdb_metadata` (Electron, via DB worker ops `DB_GET/SET_TMDB_METADATA`, plus `DB_GET_TMDB_CACHE_STATS` / `DB_CLEAR_TMDB_METADATA` behind the settings cache panel) or in-memory (PWA); localized via the app language setting. Search-match lookup keys are versioned (`|v3`), and connection startup removes the rows of every retired generation once, each under its own app-state marker (`migration:tmdb-search-lookup-v2-cache-cleanup:v1` for unversioned rows, `migration:tmdb-search-lookup-v3-cache-cleanup:v1` for `|v2` rows). The TMDB search query is `cleanTitleForSearch` (provider spelling, tags/brackets/season/year stripped), NOT the folded `normalizeTitle` key used by `pickConfidentMatch`; variant deduplication uses the lowercased query (`searchQueryIdentity`) and every attempted variant is cached under its own key, since "Феик"/"Фейк" fold to one key but are different searches and two items can share an original title while walking different variant lists: NFD folding rewrites Cyrillic "й"→"и" and "ё"→"е" and splits Arabic hamza forms, and TMDB answers such a query with nothing ("Фейк (10 серий)" never matched). Never send the folded key over the wire. +- Cached in SQLite `tmdb_metadata` (Electron, via DB worker ops `DB_GET/SET_TMDB_METADATA`, plus `DB_GET_TMDB_CACHE_STATS` / `DB_CLEAR_TMDB_METADATA` behind the settings cache panel) or in-memory (PWA); localized via the app language setting. Search-match lookup keys are versioned (`|v4`), and connection startup removes the rows of every retired generation once, each under its own app-state marker (`migration:tmdb-search-lookup-v2-cache-cleanup:v1` for unversioned rows, `…-v3-…` for `|v2` rows, `…-v4-…` for the `|v3` rows resolved before year evidence was tiered). The TMDB search query is `cleanTitleForSearch` (provider spelling, tags/brackets/season/year stripped), NOT the folded `normalizeTitle` key used by `pickConfidentMatch`; variant deduplication uses the lowercased query (`searchQueryIdentity`) and every attempted variant is cached under its own key, since "Феик"/"Фейк" fold to one key but are different searches and two items can share an original title while walking different variant lists: NFD folding rewrites Cyrillic "й"→"и" and "ё"→"е" and splits Arabic hamza forms, and TMDB answers such a query with nothing ("Фейк (10 серий)" never matched). Never send the folded key over the wire. - Service layer: `libs/services/src/lib/tmdb/`; store glue: `libs/portal/xtream/data-access/src/lib/stores/xtream-tmdb-enrichment.ts` and `libs/portal/stalker/data-access/src/lib/stores/stalker-tmdb-enrichment.ts` (hooked in `withStalkerSelection().setSelectedItem`) - TMDB attribution (logo + disclaimer) is required and shown in the settings TMDB section and About - See `docs/architecture/tmdb-metadata-enrichment.md` diff --git a/docs/architecture/tmdb-metadata-enrichment.md b/docs/architecture/tmdb-metadata-enrichment.md index 6fb9a544a..e4772b125 100644 --- a/docs/architecture/tmdb-metadata-enrichment.md +++ b/docs/architecture/tmdb-metadata-enrichment.md @@ -141,7 +141,27 @@ Wrong metadata is worse than no metadata, so id resolution is conservative: provider year — portals report the current season's year while TMDB's `first_air_date` is the premiere. Without a year, the exact-title match must be unambiguous (single hit). -4. No confident match → the provider data stays untouched, and the negative +4. Several admitted results are ranked by **year evidence first** + (`yearEvidenceTier`): the provider's exact year beats a year off by one, + which beats the series "premiered earlier" tolerance. Popularity + (`vote_count`, then `popularity`) only breaks ties inside the strongest + tier any candidate reached. + + The tolerance is a last resort, not an equal. Ranked alongside the + exact-year tier it handed every new series its older, better-known + namesake: because TMDB returns titles in the request language, an + unrelated 2018 foreign series came back under the same `ru-RU` name as a + 2026 local-language series and outvoted it, rendering its poster, cast + and genres. Over 400 Cyrillic series titles sampled from a real catalog, + 20 normalized keys had a same-titled older series and 16 of those were + the more popular row. + + The mirror case is accepted knowingly: a long-running show whose stated + season year happens to BE another same-titled show's premiere year now + resolves to the newer show. Separating those two needs the older show's + season air dates, which a search response does not carry — and the shape + needs three coincidences at once, against one that needs none. +5. No confident match → the provider data stays untouched, and the negative verdict is cached (shorter TTL) so browsing back doesn't re-search. The year filter is applied client-side rather than via TMDB's strict @@ -475,7 +495,7 @@ tmdb_metadata ( media_type 'movie' | 'tv' | 'person', lookup_key 'id:|v2' -- details payload row 'id:|season:' -- season payload row - 'title:|year:|v3' -- search resolution row + 'title:|year:|v4' -- search resolution row 'person:' -- person payload row 'trending:week' -- trending list row 'badProviderId:' -- id confirmed 404 by TMDB @@ -490,19 +510,21 @@ tmdb_metadata ( TTLs (enforced at read time in `TmdbCacheService.isFresh`): details and positive matches 30 days, negative matches 7 days. -Search keys carry a `|v3` and details keys a `|v2` version suffix +Search keys carry a `|v4` and details keys a `|v2` version suffix (`buildSearchLookupKey` / `buildDetailsLookupKey` in `tmdb-matcher.ts`): for search rows so normalization or query changes cannot reuse stale positive or negative resolutions, for details rows because payloads now include videos via `append_to_response` and pre-videos cache rows had to be invalidated. Search v2 → v3 retired the rows written while the folded comparison key was -also the wire query (see "Search query vs. comparison key" below). Database +also the wire query (see "Search query vs. comparison key" below); v3 → v4 +retired the rows resolved before year evidence was tiered, where a positive +row naming the wrong show would otherwise stay fresh for 30 days. Database startup deletes the rows of every retired search-key generation once, each under its own `app_state` marker so a skipped release still runs the cleanups it missed (`migration:tmdb-search-lookup-v2-cache-cleanup:v1` for the -unversioned rows, `migration:tmdb-search-lookup-v3-cache-cleanup:v1` for the -`|v2` rows; `LEGACY_TMDB_SEARCH_CACHE_CLEANUPS` in `connection.ts`); details -and person cache rows are unaffected. +unversioned rows, `…-v3-…` for the `|v2` rows, `…-v4-…` for the `|v3` rows; +`LEGACY_TMDB_SEARCH_CACHE_CLEANUPS` in `connection.ts`); details and person +cache rows are unaffected. ### Search query vs. comparison key @@ -513,7 +535,7 @@ both sides fold the same way. `query` is `cleanTitleForSearch` stripping, but the letters left as the provider wrote them. The search's identity is the query with only its case removed (`searchQueryIdentity`): variants are deduplicated by it and every attempted variant is cached under -its own key (`title:|year:|v3`, in the language that variant +its own key (`title:|year:|v4`, in the language that variant was searched in), never by the folded key and never only under the first variant — "Феик" and "Фейк" fold to one key but are different searches with different answers, so a verdict for one must not be read back for the other; diff --git a/libs/services/src/lib/tmdb/tmdb-id-resolver.service.spec.ts b/libs/services/src/lib/tmdb/tmdb-id-resolver.service.spec.ts index aa9549eef..6b8908004 100644 --- a/libs/services/src/lib/tmdb/tmdb-id-resolver.service.spec.ts +++ b/libs/services/src/lib/tmdb/tmdb-id-resolver.service.spec.ts @@ -87,14 +87,14 @@ describe('TmdbIdResolverService.resolveBySearch', () => { expect(searchTv).toHaveBeenCalledWith('Фейк', null, 'ru-RU', 'key'); expect(cacheSet).toHaveBeenCalledWith({ mediaType: 'tv', - lookupKey: 'title:фейк|year:2026|v3', + lookupKey: 'title:фейк|year:2026|v4', language: 'ru-RU', tmdbId: 317869, payload: null, }); }); - it('caches the miss under the folded v3 key', async () => { + it('caches the miss under the current key, not the folded one', async () => { searchTv.mockResolvedValue([]); const service = createService(); @@ -112,7 +112,7 @@ describe('TmdbIdResolverService.resolveBySearch', () => { ); expect(cacheSet).toHaveBeenCalledWith( expect.objectContaining({ - lookupKey: 'title:молодой шерлок|year:2026|v3', + lookupKey: 'title:молодой шерлок|year:2026|v4', tmdbId: null, }) ); @@ -133,7 +133,7 @@ describe('TmdbIdResolverService.resolveBySearch', () => { expect(id).toBe(317869); expect(cacheGet).toHaveBeenCalledWith( 'tv', - 'title:фейк|year:2026|v3', + 'title:фейк|year:2026|v4', 'ru-RU' ); expect(searchTv).not.toHaveBeenCalled(); @@ -157,8 +157,8 @@ describe('TmdbIdResolverService.resolveBySearch', () => { expect( cacheSet.mock.calls.map(([row]) => [row.lookupKey, row.tmdbId]) ).toEqual([ - ['title:феик|year:2026|v3', null], - ['title:фейк|year:2026|v3', 317869], + ['title:феик|year:2026|v4', null], + ['title:фейк|year:2026|v4', 317869], ]); }); @@ -167,7 +167,7 @@ describe('TmdbIdResolverService.resolveBySearch', () => { // "феик". Item B shares the original title but displays "Фейк": // the cached miss must not suppress B's own second variant. cacheGet.mockImplementation(async (_type: string, key: string) => - key === 'title:феик|year:2026|v3' ? { tmdbId: null } : null + key === 'title:феик|year:2026|v4' ? { tmdbId: null } : null ); const service = createService(); const cacheService = (service as unknown as { cache: TmdbCacheService }) @@ -186,8 +186,8 @@ describe('TmdbIdResolverService.resolveBySearch', () => { expect(searchTv).toHaveBeenCalledTimes(1); expect(searchTv).toHaveBeenCalledWith('Фейк', null, 'ru-RU', 'key'); expect(cacheGet.mock.calls.map(([, key]) => key)).toEqual([ - 'title:феик|year:2026|v3', - 'title:фейк|year:2026|v3', + 'title:феик|year:2026|v4', + 'title:фейк|year:2026|v4', ]); }); diff --git a/libs/services/src/lib/tmdb/tmdb-matcher.spec.ts b/libs/services/src/lib/tmdb/tmdb-matcher.spec.ts index 4bc8860a1..d70d232bd 100644 --- a/libs/services/src/lib/tmdb/tmdb-matcher.spec.ts +++ b/libs/services/src/lib/tmdb/tmdb-matcher.spec.ts @@ -80,10 +80,10 @@ describe('extractYear', () => { describe('lookup keys', () => { it('builds stable search and details keys', () => { expect(buildSearchLookupKey('The Matrix', 1999)).toBe( - 'title:the matrix|year:1999|v3' + 'title:the matrix|year:1999|v4' ); expect(buildSearchLookupKey('The Matrix', null)).toBe( - 'title:the matrix|year:|v3' + 'title:the matrix|year:|v4' ); }); @@ -92,7 +92,7 @@ describe('lookup keys', () => { // searches with different answers; a verdict cached for one must // never be read back for the other. expect(buildSearchLookupKey('Фейк', 2026)).toBe( - 'title:фейк|year:2026|v3' + 'title:фейк|year:2026|v4' ); expect(buildSearchLookupKey('Феик', 2026)).not.toBe( buildSearchLookupKey('Фейк', 2026) @@ -329,6 +329,84 @@ describe('pickConfidentMatch', () => { ).toBe(theBoys); }); + it('prefers an exact-year series over an older, more popular namesake', () => { + // Regression: TMDB returns titles in the REQUEST language, so an + // unrelated older foreign series (2018, 26 votes) came back under the + // same localized name as the catalog's own 2026 series (4 votes). The + // older row is admitted only by the running-season tolerance and used + // to win the popularity tiebreak. Titles here are stand-ins. + const translatedNamesake: TmdbSearchResult = { + id: 1, + name: 'Nightfall', + original_name: 'Yoru no Tobari', + first_air_date: '2018-12-14', + vote_count: 26, + popularity: 3.4, + }; + const localSeries: TmdbSearchResult = { + id: 2, + name: 'Nightfall', + original_name: 'Nightfall', + first_air_date: '2026-07-24', + vote_count: 4, + popularity: 2.4, + }; + + expect( + pickConfidentMatch( + [localSeries, translatedNamesake], + { title: 'Nightfall (12 episodes)', year: 2026 }, + 'tv' + ) + ).toBe(localSeries); + }); + + it('still falls back to an older series when nothing matches the year', () => { + const theBoys: TmdbSearchResult = { + id: 76479, + name: 'The Boys', + first_air_date: '2019-07-26', + vote_count: 12000, + }; + const unrelatedOlder: TmdbSearchResult = { + id: 999, + name: 'The Boys', + first_air_date: '2010-01-01', + vote_count: 5, + }; + + expect( + pickConfidentMatch( + [unrelatedOlder, theBoys], + { title: 'The Boys s05', year: 2026 }, + 'tv' + ) + ).toBe(theBoys); + }); + + it('prefers the exact release year over one off by one', () => { + const exact: TmdbSearchResult = { + id: 1, + title: 'The Matrix', + release_date: '1999-03-31', + vote_count: 10, + }; + const adjacent: TmdbSearchResult = { + id: 2, + title: 'The Matrix', + release_date: '1998-03-31', + vote_count: 9000, + }; + + expect( + pickConfidentMatch( + [adjacent, exact], + { title: 'The Matrix', year: 1999 }, + 'movie' + ) + ).toBe(exact); + }); + it('still rejects movies with a year that differs by more than one', () => { expect( pickConfidentMatch( diff --git a/libs/services/src/lib/tmdb/tmdb-matcher.ts b/libs/services/src/lib/tmdb/tmdb-matcher.ts index 98977beb0..e21fb0cee 100644 --- a/libs/services/src/lib/tmdb/tmdb-matcher.ts +++ b/libs/services/src/lib/tmdb/tmdb-matcher.ts @@ -121,7 +121,10 @@ export function buildSearchLookupKey( // under v2 for a title with "й"/"ё" is that bug, not a missing title, and // must not block the retry for its 7-day TTL. Rows are keyed by the // query since then. - return `title:${searchQueryIdentity(query)}|year:${year ?? ''}|v3`; + // v4: year evidence is tiered (see `yearEvidenceTier`), so every v3 row + // resolved by popularity across tiers may name the wrong show — and a + // positive row stays fresh for 30 days. + return `title:${searchQueryIdentity(query)}|year:${year ?? ''}|v4`; } export function buildDetailsLookupKey(tmdbId: number): string { @@ -256,6 +259,52 @@ function resultYear( ); } +/** + * How strongly one candidate's own year backs the year the provider stated. + * Lower is stronger; `null` means the candidate is not admissible at all. + * + * The series tier is what makes long-running shows work: a portal reports + * the CURRENT season's year ("The Boys s05" → 2026) while TMDB's + * `first_air_date` is the 2019 premiere. It is a last resort, though, not an + * equal — ranked alongside the exact-year tier with popularity deciding, it + * hands every NEW series its older, better-known namesake. TMDB returns + * titles in the REQUEST language, so in a non-English catalog those + * collisions are routine rather than exotic: a 2026 local-language drama + * (4 votes) lost to an unrelated 2018 foreign show TMDB lists under the same + * localized name (26 votes), and rendered its poster, cast and genres. Over + * 400 Cyrillic series titles sampled from a real catalog, 20 normalized keys + * had a same-titled older series and 16 of those were the more popular row. + * + * The mirror case survives on purpose: a long-running show whose stated + * season year happens to BE another same-titled show's premiere year now + * resolves to the newer show. Only the older show's season air dates could + * separate the two, and a search response does not carry them — while the + * shape needs three coincidences at once, against one that needs none. + */ +const YEAR_TIER_EXACT = 0; +const YEAR_TIER_ADJACENT = 1; +const YEAR_TIER_EARLIER_SERIES = 2; + +function yearEvidenceTier( + year: number | null, + wantedYear: number, + mediaType: TmdbMediaType +): number | null { + if (year === null) { + return null; + } + if (year === wantedYear) { + return YEAR_TIER_EXACT; + } + if (Math.abs(year - wantedYear) === 1) { + return YEAR_TIER_ADJACENT; + } + + return mediaType === 'tv' && year < wantedYear + ? YEAR_TIER_EARLIER_SERIES + : null; +} + /** * Pick the search result that confidently matches the queried title/year. * Returns `null` when confidence is insufficient — enrichment must never @@ -283,25 +332,36 @@ export function pickConfidentMatch( const wantedYear = query.year; if (wantedYear !== null) { - const yearMatches = exactTitleMatches.filter((result) => { - const year = resultYear(result, mediaType); - if (year === null) { - return false; - } - if (Math.abs(year - wantedYear) <= 1) { - return true; - } - // Series: providers often report the CURRENT season's year - // ("The Boys s05" → 2026) while TMDB's first_air_date is the - // show's premiere (2019) — accept shows that started earlier. - return mediaType === 'tv' && year < wantedYear; - }); + // Popularity only breaks ties INSIDE the strongest tier any + // candidate reached — see `yearEvidenceTier`. + const ranked = exactTitleMatches + .map((result) => ({ + result, + tier: yearEvidenceTier( + resultYear(result, mediaType), + wantedYear, + mediaType + ), + })) + .filter( + ( + candidate + ): candidate is { + result: TmdbSearchResult; + tier: number; + } => candidate.tier !== null + ); - if (yearMatches.length === 0) { + if (ranked.length === 0) { return null; } - return pickMostPopular(yearMatches); + const bestTier = Math.min(...ranked.map((candidate) => candidate.tier)); + return pickMostPopular( + ranked + .filter((candidate) => candidate.tier === bestTier) + .map((candidate) => candidate.result) + ); } // Without a year the title must be unambiguous diff --git a/libs/shared/database/src/lib/connection-migrations.spec.ts b/libs/shared/database/src/lib/connection-migrations.spec.ts index 5a6c115bb..c7f2de17c 100644 --- a/libs/shared/database/src/lib/connection-migrations.spec.ts +++ b/libs/shared/database/src/lib/connection-migrations.spec.ts @@ -159,20 +159,24 @@ describe('TMDB search lookup cache cleanup', () => { cleanupLegacyTmdbSearchCache(sqlite); - expect(transaction).toHaveBeenCalledTimes(2); - expect(deleteRun).toHaveBeenCalledTimes(2); + expect(transaction).toHaveBeenCalledTimes(3); + expect(deleteRun).toHaveBeenCalledTimes(3); expect(markerRun.mock.calls.map(([key]) => key)).toEqual([ 'migration:tmdb-search-lookup-v2-cache-cleanup:v1', 'migration:tmdb-search-lookup-v3-cache-cleanup:v1', + 'migration:tmdb-search-lookup-v4-cache-cleanup:v1', ]); - const [unversionedDelete, v2Delete] = deleteStatements(sqlite); + const [unversionedDelete, v2Delete, v3Delete] = + deleteStatements(sqlite); expect(unversionedDelete).toContain( "lookup_key LIKE 'title:%|year:%' AND lookup_key NOT LIKE 'title:%|year:%|v%'" ); - // Only v2 search rows: the v3 rows the resolver writes now, and the - // `id:`/`person:`/`badProviderId:` rows, must survive. + // Each generation deletes only its own rows: the v4 rows the resolver + // writes now, and the `id:`/`person:`/`badProviderId:` rows, survive. expect(v2Delete).toContain("WHERE lookup_key LIKE 'title:%|year:%|v2'"); expect(v2Delete).not.toContain('v3'); + expect(v3Delete).toContain("WHERE lookup_key LIKE 'title:%|year:%|v3'"); + expect(v3Delete).not.toContain('v4'); }); it('runs only the generations that have not completed yet', () => { @@ -193,13 +197,14 @@ describe('TMDB search lookup cache cleanup', () => { cleanupLegacyTmdbSearchCache(sqlite); - expect(transaction).toHaveBeenCalledTimes(1); - expect(markerRun).toHaveBeenCalledTimes(1); - expect(markerRun).toHaveBeenCalledWith( - 'migration:tmdb-search-lookup-v3-cache-cleanup:v1' - ); + expect(transaction).toHaveBeenCalledTimes(2); + expect(markerRun.mock.calls.map(([key]) => key)).toEqual([ + 'migration:tmdb-search-lookup-v3-cache-cleanup:v1', + 'migration:tmdb-search-lookup-v4-cache-cleanup:v1', + ]); expect(deleteStatements(sqlite)).toEqual([ expect.stringContaining("lookup_key LIKE 'title:%|year:%|v2'"), + expect.stringContaining("lookup_key LIKE 'title:%|year:%|v3'"), ]); }); @@ -212,7 +217,7 @@ describe('TMDB search lookup cache cleanup', () => { expect(transaction).not.toHaveBeenCalled(); // One marker read per retired generation, nothing else - expect(prepare).toHaveBeenCalledTimes(2); + expect(prepare).toHaveBeenCalledTimes(3); }); }); diff --git a/libs/shared/database/src/lib/connection.ts b/libs/shared/database/src/lib/connection.ts index ff040e9ea..543f92296 100644 --- a/libs/shared/database/src/lib/connection.ts +++ b/libs/shared/database/src/lib/connection.ts @@ -45,6 +45,8 @@ const TMDB_SEARCH_LOOKUP_V2_CACHE_CLEANUP_MIGRATION_KEY = 'migration:tmdb-search-lookup-v2-cache-cleanup:v1'; const TMDB_SEARCH_LOOKUP_V3_CACHE_CLEANUP_MIGRATION_KEY = 'migration:tmdb-search-lookup-v3-cache-cleanup:v1'; +const TMDB_SEARCH_LOOKUP_V4_CACHE_CLEANUP_MIGRATION_KEY = + 'migration:tmdb-search-lookup-v4-cache-cleanup:v1'; const EPG_PROGRAM_SOURCE_URL_BACKFILL_BATCH_SIZE = 50_000; function readTraceFlag(name: string): boolean { @@ -977,6 +979,10 @@ function widenTmdbMetadataMediaTypeCheck(sqliteDb: Database.Database): void { * v2 every title with a Cyrillic "й"/"ё" was searched folded ("феик" for * "Фейк", "елки" for "Ёлки"), got no answer, and was cached as missing for * 7 days. + * - v3 → v4: year evidence became tiered. Under v3 a series admitted only by + * the "premiered earlier" tolerance competed with an exact-year match on + * popularity alone, so a new series resolved to its older, better-known + * namesake — and that positive row stays fresh for 30 days. */ const LEGACY_TMDB_SEARCH_CACHE_CLEANUPS: ReadonlyArray<{ migrationKey: string; @@ -991,6 +997,10 @@ const LEGACY_TMDB_SEARCH_CACHE_CLEANUPS: ReadonlyArray<{ migrationKey: TMDB_SEARCH_LOOKUP_V3_CACHE_CLEANUP_MIGRATION_KEY, rowPredicate: `lookup_key LIKE 'title:%|year:%|v2'`, }, + { + migrationKey: TMDB_SEARCH_LOOKUP_V4_CACHE_CLEANUP_MIGRATION_KEY, + rowPredicate: `lookup_key LIKE 'title:%|year:%|v3'`, + }, ]; /** diff --git a/libs/shared/database/src/lib/tmdb-search-cache-cleanup.spec.ts b/libs/shared/database/src/lib/tmdb-search-cache-cleanup.spec.ts index 137f453a5..753984808 100644 --- a/libs/shared/database/src/lib/tmdb-search-cache-cleanup.spec.ts +++ b/libs/shared/database/src/lib/tmdb-search-cache-cleanup.spec.ts @@ -28,11 +28,13 @@ it('drops only retired search rows across skipped, previous, pre-person, fresh a console.log = () => undefined; const V2_MARKER = 'migration:tmdb-search-lookup-v2-cache-cleanup:v1'; const V3_MARKER = 'migration:tmdb-search-lookup-v3-cache-cleanup:v1'; + const V4_MARKER = 'migration:tmdb-search-lookup-v4-cache-cleanup:v1'; const ROWS = [ ['tv', 'title:феик|year:2026', 'ru-RU', null], ['tv', 'title:феик|year:2026|v2', 'ru-RU', null], ['tv', 'title:the boys|year:2019|v2', 'en-US', 76479], ['tv', 'title:фейк|year:2026|v3', 'ru-RU', 317869], + ['tv', 'title:nightfall|year:2026|v4', 'ru-RU', 424242], ['tv', 'id:317869|v2', 'ru-RU', 317869], ['tv', 'id:317869|season:1', 'ru-RU', 317869], ['person', 'person:287', 'en-US', 287], @@ -61,11 +63,11 @@ it('drops only retired search rows across skipped, previous, pre-person, fresh a hooks.runMigrations(skipped); const skippedAfter = snapshot(skipped); // Next startup: a row written meanwhile under the current key survives - skipped.prepare("INSERT INTO tmdb_metadata (media_type, lookup_key, language, tmdb_id) VALUES ('tv', 'title:гудовы|year:2026|v3', 'ru-RU', 318894)").run(); + skipped.prepare("INSERT INTO tmdb_metadata (media_type, lookup_key, language, tmdb_id) VALUES ('tv', 'title:гудовы|year:2026|v4', 'ru-RU', 318894)").run(); hooks.runMigrations(skipped); const repeated = snapshot(skipped); - // Previous release: the unversioned cleanup already ran; only v2 rows go - const previous = openDb([V2_MARKER]); + // Previous release: the earlier cleanups already ran; only v3 rows go + const previous = openDb([V2_MARKER, V3_MARKER]); hooks.runMigrations(previous); const previousAfter = snapshot(previous); hooks.runMigrations(previous); @@ -104,12 +106,13 @@ it('drops only retired search rows across skipped, previous, pre-person, fresh a const V2_MARKER = 'migration:tmdb-search-lookup-v2-cache-cleanup:v1'; const V3_MARKER = 'migration:tmdb-search-lookup-v3-cache-cleanup:v1'; + const V4_MARKER = 'migration:tmdb-search-lookup-v4-cache-cleanup:v1'; const survivors = [ 'badProviderId:999', 'id:317869|season:1', 'id:317869|v2', 'person:287', - 'title:фейк|year:2026|v3', + 'title:nightfall|year:2026|v4', 'trending:week', ]; const detailsPayloads = ['{"id":317869}', '{"id":317869}']; @@ -118,40 +121,54 @@ it('drops only retired search rows across skipped, previous, pre-person, fresh a skippedAfter: { keys: survivors, payloads: detailsPayloads, - markers: [V2_MARKER, V3_MARKER], + markers: [V2_MARKER, V3_MARKER, V4_MARKER], }, repeated: { - keys: [...survivors, 'title:гудовы|year:2026|v3'].sort(), + keys: [...survivors, 'title:гудовы|year:2026|v4'].sort(), payloads: detailsPayloads, - markers: [V2_MARKER, V3_MARKER], + markers: [V2_MARKER, V3_MARKER, V4_MARKER], }, previousAfter: { - // The unversioned row is that generation's business, already done - keys: [...survivors, 'title:феик|year:2026'].sort(), + // The older generations' rows are their business, already done + keys: [ + ...survivors, + 'title:феик|year:2026', + 'title:феик|year:2026|v2', + 'title:the boys|year:2019|v2', + ].sort(), payloads: detailsPayloads, - markers: [V2_MARKER, V3_MARKER], + markers: [V2_MARKER, V3_MARKER, V4_MARKER], }, previousRepeated: { - keys: [...survivors, 'title:феик|year:2026'].sort(), + keys: [ + ...survivors, + 'title:феик|year:2026', + 'title:феик|year:2026|v2', + 'title:the boys|year:2019|v2', + ].sort(), payloads: detailsPayloads, - markers: [V2_MARKER, V3_MARKER], + markers: [V2_MARKER, V3_MARKER, V4_MARKER], }, prePersonAfter: { keys: [], payloads: [], - markers: [V2_MARKER, V3_MARKER], + markers: [V2_MARKER, V3_MARKER, V4_MARKER], check: true, }, prePersonRepeated: { keys: [], payloads: [], - markers: [V2_MARKER, V3_MARKER], + markers: [V2_MARKER, V3_MARKER, V4_MARKER], + }, + freshAfter: { + keys: [], + payloads: [], + markers: [V2_MARKER, V3_MARKER, V4_MARKER], }, - freshAfter: { keys: [], payloads: [], markers: [V2_MARKER, V3_MARKER] }, freshRepeated: { keys: [], payloads: [], - markers: [V2_MARKER, V3_MARKER], + markers: [V2_MARKER, V3_MARKER, V4_MARKER], }, }); });