mirror of
https://github.com/4gray/iptvnator.git
synced 2026-10-11 02:46:16 -08:00
docs(tmdb): replace catalog titles in matching examples with stand-ins (#1658)
* docs(tmdb): replace catalog titles in matching examples with stand-ins The Cyrillic and Arabic examples that document title folding were taken straight from a user's own portal catalog while the folding bugs were being diagnosed, and they spread from there into comments, fixtures, assertions and the architecture doc. A public repository is not the place for someone's viewing inventory, and the TMDB ids pinned alongside them identify the exact shows. Swap in stand-ins that reproduce the property under test rather than the title: a word whose "й" folds away, one whose "ё" does, a two-word title carrying a season marker, an Arabic phrase whose initial hamza decomposes into a separate word. Every replacement was run through the real `normalizeTitle` and `cleanTitleForSearch` before being written, so the folded keys, the wire queries and the cache lookup keys are the same shape as before — "Лейка" and "Леика" still meet on one key while staying two different searches, and "AR| أمثلة تجريبية" still keys as a hamza split into a space. Real TMDB ids become synthetic ones. One channel fixture that read as a film title becomes a plain channel name. Deliberately left alone: the leading-tag rule's "Akira | 1988" / "Момо | Momo" examples in `title-normalization.util.ts`. Those are not an inventory — they are the evidence for a regex decision measured over 1.27M catalog titles, and inventing replacements would document a measurement that never happened. No release note: comments, fixtures and docs only, with no behaviour change. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(tmdb): label the folding stand-ins as illustrative, not as history Review caught that the substitution left synthetic titles inside sentences that assert observed fact: the architecture doc claimed a particular folded query "returns zero results while `Лейка` returns the show" and named the stand-ins as the titles that were searched, missed and negatively cached, and several comments read the same way. The titles never existed, so the reading is wrong in the one direction that matters — a future reader debugging a regression would take them for production evidence. Keep the claim that is actually established, which is a class-level one: under the old single-form design every Cyrillic title carrying "й"/"ё" and every Arabic title carrying a hamza form was searched folded, missed, and cached as missing for the negative TTL. Present the strings themselves as illustrative stand-ins chosen to fold the same way. The verified invariant — what NFD does to those letters, and what the two tiers therefore produce — is unchanged and still assertable, because it is a property of the text rather than of any title. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(tmdb): finish the sweep — a missed cache key and a last observed-fact claim Two spots the first pass left behind: - `tmdb-search-cache-cleanup.spec.ts` kept a catalog-derived lookup key verbatim. Only the TMDB id beside it had been swapped, because the sweep matched the title in its display casing and the key stores it lowercased — so one item of the inventory stayed in the repository, twice. - `SearchTitleVariant`'s doc comment still asserted that one specific folded string finds nothing on TMDB. State the mechanism instead: folding rewrites the letters the search matches on, so the folded form finds nothing there. `content-search.util.sqlite.spec.ts` gets a channel fixture that no longer reads as a film title. `search-text-fold.util.spec.ts` is deliberately left alone. Its "Ёлки" is a common noun sitting in a Unicode corpus beside "Amélie", "İnşaat" and "Ά", not an inventory entry — and its rows are precomposed/decomposed PAIRS (U+0401 against U+0415 U+0308). A textual substitution rewrites only the composed half and silently turns the pair into two different words; doing it failed that spec, which is what a fixture encoding an invisible property is for. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(search): finish the fold fixtures by rewriting both halves of the pair The last two occurrences lived in the Unicode fold corpus, which an earlier attempt left alone after a textual substitution failed the spec. The reason it failed is the point: one row is a precomposed/decomposed PAIR — U+0401 against U+0415 U+0308 — asserting that both spellings fold to one string. Replacing the visible text rewrites only the composed half and silently turns the pair into two different words, which is not a thing a reader or a regex can see. Rewrite both halves by code point instead, keeping the decomposed half decomposed: U+0415 U+0308 + the new stem. Verified by decoding every quoted string in the file afterwards, and the spec passes (534/534). A repository-wide sweep now finds no occurrence of the replaced titles in either normalization form. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
1 parent
11358caa7a
commit
eb61d37dc3
11 files changed
+114
-107
No files matched your search
@@ -976,9 +976,9 @@ function widenTmdbMetadataMediaTypeCheck(sqliteDb: Database.Database): void {
|
||||
* - unversioned → v2: title normalization learned to strip appended
|
||||
* language/quality tags.
|
||||
* - v2 → v3: the search query stopped being the folded comparison key. Under
|
||||
* v2 every title with a Cyrillic "й"/"ё" was searched folded ("феик" for
|
||||
* "Фейк", "елки" for "Ёлки"), got no answer, and was cached as missing for
|
||||
* 7 days.
|
||||
* v2 every title with a Cyrillic "й"/"ё" was searched folded — the fold
|
||||
* spells them "и" and "е", as in the illustrative "леика" for "Лейка" —
|
||||
* got no answer, and was cached as missing for 7 days.
|
||||
* - v3 → v4: year evidence became tiered. Under v3 a series admitted only by
|
||||
* the "premiered earlier" tolerance competed with an exact-year match on
|
||||
* popularity alone, so a new series resolved to its older, better-known
|
||||
|
||||
@@ -30,13 +30,13 @@ it('drops only retired search rows across skipped, previous, pre-person, fresh a
|
||||
const V3_MARKER = 'migration:tmdb-search-lookup-v3-cache-cleanup:v1';
|
||||
const V4_MARKER = 'migration:tmdb-search-lookup-v4-cache-cleanup:v1';
|
||||
const ROWS = [
|
||||
['tv', 'title:феик|year:2026', 'ru-RU', null],
|
||||
['tv', 'title:феик|year:2026|v2', 'ru-RU', null],
|
||||
['tv', 'title:леика|year:2026', 'ru-RU', null],
|
||||
['tv', 'title:леика|year:2026|v2', 'ru-RU', null],
|
||||
['tv', 'title:the boys|year:2019|v2', 'en-US', 76479],
|
||||
['tv', 'title:фейк|year:2026|v3', 'ru-RU', 317869],
|
||||
['tv', 'title:лейка|year:2026|v3', 'ru-RU', 101101],
|
||||
['tv', 'title:nightfall|year:2026|v4', 'ru-RU', 424242],
|
||||
['tv', 'id:317869|v2', 'ru-RU', 317869],
|
||||
['tv', 'id:317869|season:1', 'ru-RU', 317869],
|
||||
['tv', 'id:101101|v2', 'ru-RU', 101101],
|
||||
['tv', 'id:101101|season:1', 'ru-RU', 101101],
|
||||
['person', 'person:287', 'en-US', 287],
|
||||
['movie', 'badProviderId:999', 'any', null],
|
||||
['movie', 'trending:week', 'en-US', null],
|
||||
@@ -63,7 +63,7 @@ it('drops only retired search rows across skipped, previous, pre-person, fresh a
|
||||
hooks.runMigrations(skipped);
|
||||
const skippedAfter = snapshot(skipped);
|
||||
// Next startup: a row written meanwhile under the current key survives
|
||||
skipped.prepare("INSERT INTO tmdb_metadata (media_type, lookup_key, language, tmdb_id) VALUES ('tv', 'title:гудовы|year:2026|v4', 'ru-RU', 318894)").run();
|
||||
skipped.prepare("INSERT INTO tmdb_metadata (media_type, lookup_key, language, tmdb_id) VALUES ('tv', 'title:сосенка|year:2026|v4', 'ru-RU', 101103)").run();
|
||||
hooks.runMigrations(skipped);
|
||||
const repeated = snapshot(skipped);
|
||||
// Previous release: the earlier cleanups already ran; only v3 rows go
|
||||
@@ -78,7 +78,7 @@ it('drops only retired search rows across skipped, previous, pre-person, fresh a
|
||||
const prePerson = new Database(':memory:');
|
||||
hooks.createTables(prePerson);
|
||||
prePerson.exec("DROP TABLE tmdb_metadata; CREATE TABLE tmdb_metadata (id INTEGER PRIMARY KEY AUTOINCREMENT, media_type TEXT NOT NULL CHECK (media_type IN ('movie', 'tv')), lookup_key TEXT NOT NULL, language TEXT NOT NULL, tmdb_id INTEGER, payload TEXT, fetched_at TEXT DEFAULT (datetime('now'))); CREATE UNIQUE INDEX tmdb_metadata_lookup_unique ON tmdb_metadata(media_type, lookup_key, language)");
|
||||
prePerson.prepare("INSERT INTO tmdb_metadata (media_type, lookup_key, language, tmdb_id) VALUES ('tv', 'title:феик|year:2026', 'ru-RU', NULL)").run();
|
||||
prePerson.prepare("INSERT INTO tmdb_metadata (media_type, lookup_key, language, tmdb_id) VALUES ('tv', 'title:леика|year:2026', 'ru-RU', NULL)").run();
|
||||
hooks.runMigrations(prePerson);
|
||||
const prePersonAfter = { ...snapshot(prePerson), check: prePerson.prepare("SELECT sql FROM sqlite_master WHERE name = 'tmdb_metadata'").get().sql.includes("'person'") };
|
||||
hooks.runMigrations(prePerson);
|
||||
@@ -109,13 +109,13 @@ it('drops only retired search rows across skipped, previous, pre-person, fresh a
|
||||
const V4_MARKER = 'migration:tmdb-search-lookup-v4-cache-cleanup:v1';
|
||||
const survivors = [
|
||||
'badProviderId:999',
|
||||
'id:317869|season:1',
|
||||
'id:317869|v2',
|
||||
'id:101101|season:1',
|
||||
'id:101101|v2',
|
||||
'person:287',
|
||||
'title:nightfall|year:2026|v4',
|
||||
'trending:week',
|
||||
];
|
||||
const detailsPayloads = ['{"id":317869}', '{"id":317869}'];
|
||||
const detailsPayloads = ['{"id":101101}', '{"id":101101}'];
|
||||
|
||||
expect(JSON.parse(result)).toEqual({
|
||||
skippedAfter: {
|
||||
@@ -124,7 +124,7 @@ it('drops only retired search rows across skipped, previous, pre-person, fresh a
|
||||
markers: [V2_MARKER, V3_MARKER, V4_MARKER],
|
||||
},
|
||||
repeated: {
|
||||
keys: [...survivors, 'title:гудовы|year:2026|v4'].sort(),
|
||||
keys: [...survivors, 'title:сосенка|year:2026|v4'].sort(),
|
||||
payloads: detailsPayloads,
|
||||
markers: [V2_MARKER, V3_MARKER, V4_MARKER],
|
||||
},
|
||||
@@ -132,8 +132,8 @@ it('drops only retired search rows across skipped, previous, pre-person, fresh a
|
||||
// The older generations' rows are their business, already done
|
||||
keys: [
|
||||
...survivors,
|
||||
'title:феик|year:2026',
|
||||
'title:феик|year:2026|v2',
|
||||
'title:леика|year:2026',
|
||||
'title:леика|year:2026|v2',
|
||||
'title:the boys|year:2019|v2',
|
||||
].sort(),
|
||||
payloads: detailsPayloads,
|
||||
@@ -142,8 +142,8 @@ it('drops only retired search rows across skipped, previous, pre-person, fresh a
|
||||
previousRepeated: {
|
||||
keys: [
|
||||
...survivors,
|
||||
'title:феик|year:2026',
|
||||
'title:феик|year:2026|v2',
|
||||
'title:леика|year:2026',
|
||||
'title:леика|year:2026|v2',
|
||||
'title:the boys|year:2019|v2',
|
||||
].sort(),
|
||||
payloads: detailsPayloads,
|
||||
|
||||
Reference in new issue
Block a user