fix(portals): fold non-ASCII case in the scan, and read years as tags

I was wrong about SQLite twice over, and both errors cost matches.

GLOB character classes are NOT ASCII-only. `patternCompare` reads them as
UTF-8 code points, so `'Он' GLOB '*[Оо][Нн]*'` is true — only `LOWER()` is
ASCII-only. The scan tier now folds the case in JavaScript, where Unicode
case mapping is real, and hands SQLite one class per character. A short
Cyrillic or Greek title stored in a different case is found instead of
being silently absent from the Sources chip. The builder returns `null`,
leaving the substring tests as the whole answer, for a token holding a
GLOB metacharacter (GLOB has no escape character) or a case mapping that
changes length. The FTS tier is untouched and still cannot fold — that
needs a stored normalized-title column.

The movie's own year came from `extractYear`, which reads a year from
anywhere in the title. That is right where a year is a search hint, wrong
where it is an identity: `2001: A Space Odyssey` was treated as a 2001
film, so every genuine 1968 copy failed the year gate and the movie had
no alternatives at all — and its pin key moved the moment enrichment
supplied the real year. `releaseTagYear` accepts only bracketed and
trailing forms; the repo's own TRAILING_YEAR_PATTERN already documented
this exact hazard.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
4grayandClaude Opus 5 committed 2026-07-29 19:48:25 +02:00
1 parent bbdc64b9dc
commit ddb299e237
9 files changed
+305 -33

No files matched your search

+1 -1
View File
@@ -897,7 +897,7 @@ engine` (restart required) or
- Finds the same movie in the user's other imported playlists and adds a "Sources N" chip to the Xtream VOD action row (only when ≥1 alternative exists), plus a `.source-caption` line reporting where playback is coming from. The chip opens a 460px anchored CDK-overlay popover (`libs/ui/components/src/lib/vod-sources/`; not `MatMenu`, which caps its width at 280px), reused unchanged in the inline player's now-playing bar and on the playback-error screen. Both chips are handed the same `matchKind` and `vodAutoFailover` and both write the setting back. The chip counts alternative **streams**; the caption ("also found in N other playlists") counts distinct **playlists** via `alternativePlaylistCount`, because the popover groups one portal's copies under that portal.
- Scope v1 is **Xtream ↔ Xtream, movies only, Electron only**. Stalker never reaches the `content` table and M3U is a JSON blob whose search forces `content_type:'live'`; both are additive later since `VodSourceCandidate.portalType` already carries all three. In the PWA every entry point is gated off by a bridge `typeof` check and the chip renders nothing.
- **Metadata provenance is the core contract.** Every field is `{value, provenance}` where `api`/`probe` are facts (plain tag), `parsed` is a title-regex guess (tag prefixed `~`, warn colour), and absent renders **no tag at all** plus a `check` chip. `factualOnly()` in `vod-source-metadata.util.ts` is the only accessor allowed for ranking/failover, so guesses are structurally unable to influence a decision. `VodSourceProbeStatus` separates `fail` (contacted and refused) from `unknown` (timed out / blocked / no capability) — an unchecked source is never shown as offline. Quality is derived from pixel **width** because letterboxing crops height.
- Discovery (`DB_FIND_TITLE_SOURCES`, trigram FTS over `content_title_fts`) is lazy and returns only what the `content` table can prove; titles whose tokens are all shorter than three characters ("Up", "It") fall back to a scan, since the trigram tokenizer cannot index them at all. A source that is never read looks exactly like one that does not exist, so: the current playlist is excluded **in SQL** and duplicates collapse there too (`GROUP BY cat.playlist_id, c.xtream_id` before the limit — one playlist's dozens of identically ranked category rows would otherwise crowd out every alternative), and the scan matches the token as a whole word (`' ' || LOWER(title) || ' ' GLOB '*[^a-z0-9]it[^a-z0-9]*'`) ordered by title length **with no row limit** — FTS keeps its 60-row window because it ranks by relevance, while a scan cannot rank, and the GLOB reads every row regardless so a limit would only truncate the answer. The year gate covers BOTH match tiers: `normalizeTitleKeys` strips bracketed segments, so "Dune (1984)" normalizes identically to "Dune" and would otherwise be an *exact* match for the 2021 film; a bracketed year is read out of the raw title and a stated disagreement rejects the row. One row inside the excluded playlist is kept when the caller names it (`keepContentId`), because a pin can point at another copy in the playlist being viewed — the host reads the pin before discovery for exactly this. Resolution is deferred to click/pin/check because `content` stores no `container_extension` and `constructVodUrl` returns `''` without one — each alternative costs a live `get_vod_info` against the foreign playlist's credentials.
- Discovery (`DB_FIND_TITLE_SOURCES`, trigram FTS over `content_title_fts`) is lazy and returns only what the `content` table can prove; titles whose tokens are all shorter than three characters ("Up", "It") fall back to a scan, since the trigram tokenizer cannot index them at all. A source that is never read looks exactly like one that does not exist, so: the current playlist is excluded **in SQL** and duplicates collapse there too (`GROUP BY cat.playlist_id, c.xtream_id` before the limit — one playlist's dozens of identically ranked category rows would otherwise crowd out every alternative), and the scan matches an ASCII token as a whole word (`' ' || LOWER(title) || ' ' GLOB '*[^a-z0-9]it[^a-z0-9]*'`) ordered by title length **with no row limit** — FTS keeps its 60-row window because it ranks by relevance, while a scan cannot rank, and the GLOB reads every row regardless so a limit would only truncate the answer. The year gate covers BOTH match tiers: `normalizeTitleKeys` strips bracketed segments, so "Dune (1984)" normalizes identically to "Dune" and would otherwise be an *exact* match for the 2021 film; a bracketed year is read out of the raw title and a stated disagreement rejects the row. A non-ASCII token cannot be folded by `LOWER()` (ASCII-only) but CAN be by a GLOB character class (UTF-8 code points), so `caseInsensitiveGlobPattern` folds the case in JS and emits one `[lowerUpper]` class per character — returning `null`, leaving the two substring tests alone, for a GLOB metacharacter or a length-changing case map (`ß`→`SS`). The movie's own year comes from `releaseTagYear` (bracketed or trailing only), never `extractYear`: a year inside the NAME ("2001: A Space Odyssey") would fail every genuine 1968 copy at the year gate and move the pin key once enrichment lands. One row inside the excluded playlist is kept when the caller names it (`keepContentId`), because a pin can point at another copy in the playlist being viewed — the host reads the pin before discovery for exactly this. Resolution is deferred to click/pin/check because `content` stores no `container_extension` and `constructVodUrl` returns `''` without one — each alternative costs a live `get_vod_info` against the foreign playlist's credentials.
- Switching = one `inlinePlayback.set({...next, startTime})`, never null-then-set, so the player and engine survive and re-seek. The carried position is read *before* the 15s persistence throttle, and `VodDetailsPlaybackService` uses a one-shot `resumeSettled` latch so a resuming engine's `timeupdate` at ~0 cannot overwrite the resume point. `handleInlineTimeUpdate` returns that verdict and the route feeds multi-source the requested `startTime` until the engine reaches it — one latch for both, or a switch during the initial seek would restart the film. Before anything plays there is no live position at all, so the controller is seeded from the persisted one (`seedResumeSeconds`, one-way: a live value always wins). Portal failures in the multi-source path log through the redacting `createLogger`/`redactSensitiveData` — an Xtream error message carries the stream URL, and that URL is built out of the username and password.
- Pins are keyed portal-agnostically (`tmdb:{id}` else `title:{base}:{year}` else the yearless `title:{base}:`, `vod_source_pins` table); enrichment supplies the id and the year late, so a pin may sit under any poorer form — three key sets (`pinKeysFor`): `lookup` passes every alias most-trusted-first, `write` holds only keys naming exactly one film, and `loaded` records where the pin on screen was found — the yearless alias is readable but never written or deleted on spec, since it is shared by every remake, with the single exception of the row this session actually read. A write stores the decision under **every** key in `write` (`setVodSourcePin(db, pin, retireKeys, aliasKeys)`: one upsert per key plus the leftover retirement, in a single transaction), because a movie's identity grows — recorded only under the enriched `tmdb:` key, a pin is invisible to the next reopen, which starts out with just a title and a year, and stays invisible for good if enrichment is off or never answers. A pin is not decoration: the primary Play action starts from the pinned source (except when that button reads Stop — an active external session wins, or the control would launch a second player), and it outranks everything else in failover ranking. The row changes only after the write lands, so a refused pin is never shown as saved. Starting a pinned source loads THAT source's own playback position — progress is keyed by (playlist, stream), so the row the page loaded belongs to the route's copy. The primary button says nothing at all until that row is in, and "is it in" is answered by comparing the loaded pin **id** rather than mere presence, or re-pinning would leave the button wearing the previous copy's timecode. An external player launched for an alternative carries the OTHER playlist's ids, so `VodDetailsPlaybackBindings.activeSource` feeds one `ownsContent()` predicate used by BOTH the session matcher and the playback-position bridge — if they disagree, the page shows Stop for a session whose progress it throws away and a later switch rewinds hours. Two identity keys: `vodMultiSourceMovieKey` (title, year, tmdbId) makes TMDB enrichment re-trigger discovery and rebuild the pin keys, while `vodMultiSourceSessionKey` (`playlistId:contentId`) decides whether that rerun is a refresh or a new session — a refresh keeps the active source, its resolved facts, the tried set, the live position and any switch in flight; only a different film resets them.
- Claims in the present tense (the "Playing from" caption and the source row's `Playing` badge) are gated on `VodDetailsRouteComponent.playbackLive`, never on `isActive` — discovery marks a source active before anything plays and it stays active after the player closes. Inline that means a `timeupdate` has arrived (`inlinePlayback()` is only the request to play); external it means the session is past `launching`. A merely selected row reads `Current`.
@@ -66,16 +66,20 @@ describe('title-sources.operations', () => {
});
it('can still find a short non-ASCII title', async () => {
// SQLite's LOWER() and GLOB classes are ASCII-only, so folding
// "Он" to "он" never happened and the film stayed invisible in
// the Sources chip entirely.
// SQLite's LOWER() is ASCII-only, so folding "Он" to "он" never
// happened and the film stayed invisible in the Sources chip.
const scan = createDbMock([]);
await findTitleSources(scan.db, { title: 'Он' });
const scanQuery = compiledQuery(scan.all);
// GLOB character classes are NOT ASCII-only — they compare UTF-8
// code points — so the case is folded in JavaScript and handed
// over one class per character. Without this a title stored in
// any case other than the two below is simply never found.
expect(scanQuery.params).toContain('*[оО][нН]*');
expect(scanQuery.sql).toContain('instr');
// Both the folded and the as-typed form, since neither alone
// matches a title stored in the other case.
// Both the folded and the as-typed form, since the class is built
// from the raw token and cannot cover a diacritic difference.
expect(scanQuery.params).toContain('он');
expect(scanQuery.params).toContain('Он');
});
@@ -5,6 +5,7 @@ import {
type VodSourceCandidateRow,
} from '@iptvnator/shared/interfaces';
import type { AppDatabase } from '../database.types';
import { caseInsensitiveGlobPattern } from './title-token-glob';
/**
* VOD multi-source discovery: find the SAME movie in the user's other
@@ -121,18 +122,22 @@ function excludePlaylistClause(
/**
* One token's presence test.
*
* SQLite's `LOWER()` and GLOB character classes are ASCII-only, so a Cyrillic
* or Greek token cannot be folded or word-bounded in stock SQLite — matching
* "он" against a stored "Он" would simply fail, and the film stayed invisible
* in the Sources chip.
* SQLite's `LOWER()` is ASCII-only — `LOWER('Он')` is `'Он'` — so a Cyrillic or
* Greek token cannot be folded by SQL, and matching "он" against a stored "Он"
* simply failed: the film stayed invisible in the Sources chip.
*
* GLOB character classes are NOT ASCII-only, though; they compare UTF-8 code
* points. So the non-ASCII branch folds the case in JavaScript, where Unicode
* case mapping is real, and hands SQLite one `[lowerUpper]` class per
* character. The two substring tests stay alongside it, because the class is
* built from the RAW token and cannot match a title whose diacritics differ
* ("Ça" vs a stored "Ca"), which the folded-token test does cover.
*
* ASCII tokens keep the word-boundary GLOB, which is what stops "it" matching
* "Titanic". A non-ASCII token falls back to a substring test tried against
* both the folded and the as-typed form: that covers a title stored in the
* same case as the request or in lower case, which is the realistic pair, and
* it deliberately does NOT claim full Unicode case folding. It is looser than
* the ASCII branch, and the normalized confirmation afterwards is what makes
* that safe — a looser filter here costs transfer, never a wrong match.
* "Titanic". The non-ASCII branch has no word boundary — `[^a-z0-9]` would
* treat any Cyrillic letter as a separator — so it is looser, and the
* normalized confirmation afterwards is what makes that safe: a looser filter
* here costs transfer, never a wrong match.
*/
function tokenPredicate(token: string, rawToken: string): SQL {
// Decided from the RAW token, not the normalized one. Normalization folds
@@ -147,7 +152,11 @@ function tokenPredicate(token: string, rawToken: string): SQL {
return sql`' ' || LOWER(c.title) || ' ' GLOB ${`*[^a-z0-9]${token}[^a-z0-9]*`}`;
}
return sql`(instr(LOWER(c.title), ${token}) > 0 OR instr(c.title, ${rawToken}) > 0)`;
const substring = sql`(instr(LOWER(c.title), ${token}) > 0 OR instr(c.title, ${rawToken}) > 0)`;
const anyCase = caseInsensitiveGlobPattern(rawToken);
// `null` means the token holds a GLOB metacharacter or a case mapping that
// changes length, and the substring tests alone are the honest answer.
return anyCase ? sql`(c.title GLOB ${anyCase} OR ${substring})` : substring;
}
function scanCandidateQuery(
@@ -0,0 +1,106 @@
import { execFileSync } from 'node:child_process';
import { caseInsensitiveGlobPattern } from './title-token-glob';
/**
* Case folding for the tokens SQLite cannot fold.
*
* The patterns are only worth anything if SQLite agrees with them, so the
* behavioural cases below are asserted against a real `sqlite3` rather than
* against a restatement of the builder's own logic.
*/
/** `true` when the running SQLite says `value GLOB pattern`. */
function sqliteGlob(value: string, pattern: string): boolean {
const out = execFileSync(
'sqlite3',
[':memory:', `SELECT ${quote(value)} GLOB ${quote(pattern)};`],
{ encoding: 'utf8' }
);
return out.trim() === '1';
}
function quote(literal: string): string {
return `'${literal.replace(/'/g, "''")}'`;
}
/**
* The suite asserts against the sqlite3 CLI, which is present on macOS and on
* the CI images but is not a declared dependency of this workspace.
*/
function hasSqlite(): boolean {
try {
execFileSync('sqlite3', ['-version'], { stdio: 'ignore' });
return true;
} catch {
return false;
}
}
const describeWithSqlite = hasSqlite() ? describe : describe.skip;
describe('caseInsensitiveGlobPattern', () => {
it('builds a two-case class for each cased character', () => {
expect(caseInsensitiveGlobPattern('Он')).toBe('*[оО][нН]*');
});
it('leaves uncased characters alone', () => {
expect(caseInsensitiveGlobPattern('7')).toBe('*7*');
});
it('refuses a token holding a GLOB metacharacter', () => {
// SQLite GLOB has no escape character, so an unescaped `*` here would
// silently become a wildcard and match every row in the table.
expect(caseInsensitiveGlobPattern('о*')).toBeNull();
expect(caseInsensitiveGlobPattern('[о')).toBeNull();
expect(caseInsensitiveGlobPattern('о-н')).toBeNull();
});
it('refuses a case mapping that changes length', () => {
// 'ß'.toUpperCase() is 'SS' — there is no single-character class that
// means the same thing, and guessing one would be a wrong pattern
// rather than an absent one.
expect(caseInsensitiveGlobPattern('ß')).toBeNull();
});
it('refuses an empty token', () => {
expect(caseInsensitiveGlobPattern('')).toBeNull();
});
});
describeWithSqlite('caseInsensitiveGlobPattern against SQLite', () => {
it('matches a Cyrillic title in any case', () => {
const pattern = caseInsensitiveGlobPattern('Он') as string;
// The bug: LOWER() is ASCII-only, so a request for "Он" never found a
// stored "ОН" and the film was missing from the Sources chip.
expect(sqliteGlob('Он', pattern)).toBe(true);
expect(sqliteGlob('ОН', pattern)).toBe(true);
expect(sqliteGlob('он (2017)', pattern)).toBe(true);
});
it('is exactly what SQLite cannot do on its own', () => {
// Guards the premise the whole helper rests on: if a future SQLite
// folded Cyrillic in LOWER(), this indirection would be dead weight.
const out = execFileSync(
'sqlite3',
[':memory:', `SELECT LOWER('Он') = 'он';`],
{ encoding: 'utf8' }
);
expect(out.trim()).toBe('0');
});
it('still refuses a title that does not hold the token', () => {
const pattern = caseInsensitiveGlobPattern('Он') as string;
expect(sqliteGlob('Дюна', pattern)).toBe(false);
});
it('matches Greek and accented Latin in any case', () => {
expect(
sqliteGlob('ΟΙ', caseInsensitiveGlobPattern('οι') as string)
).toBe(true);
expect(
sqliteGlob('ÇA', caseInsensitiveGlobPattern('Ça') as string)
).toBe(true);
});
});
@@ -0,0 +1,57 @@
/**
* Case-insensitive GLOB patterns for tokens SQLite cannot fold itself.
*
* `LOWER()` in stock SQLite is ASCII-only — `LOWER('Он')` is `'Он'` — so a
* Cyrillic or Greek token can never be folded in SQL. GLOB character classes,
* however, are NOT ASCII-only: `patternCompare` reads them as UTF-8 code
* points, so `'Он' GLOB '*[Оо][Нн]*'` is true. Folding the case in JavaScript,
* where Unicode case mapping is real, and handing SQLite one class per
* character gets the match SQL cannot compute on its own.
*/
/**
* GLOB's own metacharacters. SQLite has no escape character in GLOB patterns
* (only the `[*]` class form), so a token containing one of these is left to
* the caller's substring fallback rather than escaped by hand.
*/
const GLOB_METACHARACTERS = /[*?[\]^-]/;
/**
* A `*…*` containment pattern matching `token` in any case, or `null` when the
* token cannot be expressed safely.
*
* `null` is returned rather than a best-effort pattern for two reasons: a
* metacharacter would silently turn into a wildcard, and a character whose case
* mapping changes length (`ß` uppercases to `SS`) has no single-character class
* that means the same thing. Both are rare enough that the caller's plain
* substring test is a better answer than a subtly wrong pattern.
*/
export function caseInsensitiveGlobPattern(token: string): string | null {
if (token === '' || GLOB_METACHARACTERS.test(token)) {
return null;
}
let pattern = '*';
for (const character of token) {
const lower = character.toLowerCase();
const upper = character.toUpperCase();
if (lower === upper) {
pattern += character;
continue;
}
if (
[...lower].length !== 1 ||
[...upper].length !== 1 ||
GLOB_METACHARACTERS.test(lower) ||
GLOB_METACHARACTERS.test(upper)
) {
return null;
}
pattern += `[${lower}${upper}]`;
}
return `${pattern}*`;
}
+31 -15
View File
@@ -285,6 +285,16 @@ year, ranked *above* every fuzzy match. Discovery reads a bracketed year out of
the raw title first, and when both sides state a year and they disagree, it is
not the same movie. An unknown year still never blocks.
Which is why the movie's OWN year is read with `releaseTagYear`, not
`extractYear`. The latter takes a year from anywhere in the title, which is the
right answer where a year is only a search hint that scoring will confirm — and
the wrong one here, where the year is part of an identity. "2001: A Space
Odyssey" is not a 2001 film: calling it one makes every genuine 1968 copy fail
the gate above, so the movie has no alternatives at all, and it moves the pin
key the moment enrichment supplies the real year. Only bracketed and trailing
forms count. The trailing form stays ambiguous on purpose ("Blade Runner 2049"
is a title, not a tag) — that is what the two-tier exact/base match is for.
## Why resolution is lazy
The `content` table stores no `container_extension`, and
@@ -513,23 +523,29 @@ table first and does nothing when the runtime rejects it, leaving the working
index in place — and does NOT record itself as done, so a later app version
shipping a newer SQLite upgrades then.
**Case folding for non-ASCII is still not possible.** `LOWER()`, GLOB classes
and the trigram tokenizer all fold ASCII only, so a Cyrillic title stored with
different capitalisation in two playlists ("ОН" vs "Он") cannot be matched by
any predicate available in stock SQLite. Closing that needs a stored
normalized-title column, which is deliberately out of scope here.
**Case folding for non-ASCII, on the FTS tier, is still not possible.**
`LOWER()` and the trigram tokenizer both fold ASCII only, so a Cyrillic title
stored with different capitalisation in two playlists ("ОН" vs "Он") cannot be
folded by the indexed path. Closing that needs a stored normalized-title
column, which is deliberately out of scope here.
The **scan** tier does fold it, because GLOB character classes are *not*
ASCII-only — `patternCompare` reads them as UTF-8 code points, and
`'Он' GLOB '*[Оо][Нн]*'` is true. `caseInsensitiveGlobPattern` therefore folds
the case in JavaScript, where Unicode case mapping is real, and hands SQLite one
`[lowerUpper]` class per character. It returns `null` — leaving the substring
tests as the whole answer — for a token holding a GLOB metacharacter (SQLite
GLOB has no escape character, so an unescaped `*` would become a wildcard) or a
case mapping that changes length (`ß` uppercases to `SS`).
SQLite's `LOWER()` and GLOB character classes are ASCII-only, so a short
non-ASCII title could not be folded or word-bounded and simply never matched —
the film stayed absent from the chip. The ASCII/Unicode branch is decided from the RAW token, not the normalized
one: normalization folds diacritics, so "Ça" arrives as "ca" and looks like
plain ASCII while the stored title still reads "Ça". ASCII tokens keep the
word-boundary GLOB
(what stops "it" matching "Titanic"); a non-ASCII token falls back to a
substring test against both the folded and the as-typed form. That covers a
title stored in the same case as the request or in lower case, and deliberately
does not claim full Unicode case folding. It is looser, and the normalized
The ASCII/Unicode branch is decided from the RAW token, not the normalized one:
normalization folds diacritics, so "Ça" arrives as "ca" and looks like plain
ASCII while the stored title still reads "Ça". ASCII tokens keep the
word-boundary GLOB (what stops "it" matching "Titanic"). The non-ASCII branch
keeps both substring tests alongside the folded class, because the class is
built from the raw token and cannot match across a diacritic difference, which
the folded-token test does. It has no word boundary — `[^a-z0-9]` would treat
every Cyrillic letter as a separator — so it is looser, and the normalized
confirmation afterwards is what makes looser safe.
## Claims about the present
@@ -66,6 +66,47 @@ describe('resolveVodMultiSourceMovie', () => {
expect(resolved?.tmdbId).toBe(693134);
});
it('ignores a year that belongs to the movie’s name', () => {
const resolved = resolveVodMultiSourceMovie({
...BASE,
catalogItem: { title: '2001: A Space Odyssey' },
});
// Reading any year anywhere calls this a 2001 film. Every genuine copy
// states 1968 and fails the year gate, so the movie has no alternative
// sources at all — and its pin key moves as soon as enrichment
// supplies the real year.
expect(resolved?.year).toBeNull();
});
it('reads a release tag in either shape', () => {
expect(
resolveVodMultiSourceMovie({
...BASE,
catalogItem: { title: 'Dune (2021)' },
})?.year
).toBe(2021);
expect(
resolveVodMultiSourceMovie({
...BASE,
catalogItem: { title: 'The Matrix 1999' },
})?.year
).toBe(1999);
});
it('prefers the stated release date over anything in the title', () => {
const resolved = resolveVodMultiSourceMovie({
...BASE,
catalogItem: { title: 'The Matrix 1999' },
vodInfo: {
name: 'The Matrix 1999',
releasedate: '1999-03-31',
} as never,
});
expect(resolved?.year).toBe(1999);
});
it('falls back to the playlist id when the name is empty', () => {
const resolved = resolveVodMultiSourceMovie({
...BASE,
@@ -1,5 +1,6 @@
import {
extractYear,
releaseTagYear,
type XtreamVodInfo,
type XtreamVodStream,
} from '@iptvnator/shared/interfaces';
@@ -58,7 +59,12 @@ export function resolveVodMultiSourceMovie(input: {
playlistName: input.playlistName?.trim() || playlistId,
contentId: vodId,
title,
year: extractYear(vodInfo?.releasedate, title),
// The title is read for a release TAG only. Here the year is part of
// the movie's identity — it gates discovery and keys the pin — so a
// year that is part of the NAME ("2001: A Space Odyssey") would reject
// every genuine 1968 copy and move the pin key the moment enrichment
// supplies the real one.
year: extractYear(vodInfo?.releasedate) ?? releaseTagYear(title),
tmdbId: vodInfo?.tmdb_id,
};
}
@@ -293,3 +293,36 @@ export function extractYear(
const fromTitle = rawTitle?.match(YEAR_PATTERN)?.[0];
return fromTitle ? Number(fromTitle) : null;
}
/** Bracketed release tag: "Dune (2021)", "Dune [2021]". */
const BRACKETED_YEAR_PATTERN = /[([{]\s*(19\d{2}|20\d{2})\s*[)\]}]/;
/**
* The year a provider title states as a release TAG, never one that belongs to
* the film's name.
*
* `extractYear` reads any year anywhere in the title, which is the right answer
* where a year is only a search hint that scoring will confirm. It is the wrong
* answer wherever the year becomes part of an IDENTITY: "2001: A Space Odyssey"
* is not a 2001 film, and calling it one makes every genuine 1968 copy fail the
* year gate — the movie then has no alternative sources at all — while its pin
* key moves the moment enrichment supplies the real year.
*
* Only bracketed and trailing forms count. The trailing form stays ambiguous on
* purpose ("Blade Runner 2049" is a title, not a tag); that is what the
* `exact`/`base` two-tier match is for, and nothing here can settle it.
*/
export function releaseTagYear(
rawTitle: string | null | undefined
): number | null {
if (!rawTitle) {
return null;
}
// Before `normalizeTitleKeys`, which strips bracketed segments wholesale
// and would take the tag with them.
const bracketed = rawTitle.match(BRACKETED_YEAR_PATTERN);
return bracketed
? Number(bracketed[1])
: normalizeTitleKeys(rawTitle).trailingYear;
}