mirror of
https://github.com/4gray/iptvnator.git
synced 2026-10-11 11:06:16 -08:00
fix(xtream): keep normalization uppercase-only, measured on real catalogs
The previous commit widened the pipe branch of normalizeTitleKeys to any case and to Cyrillic, on the theory that nothing but a tag precedes a pipe. Checked against 1.27M real catalog titles that theory is wrong in two ways at once: "Akira | 1988" and "Coco | 2017" put the film's name before the pipe and the year after it, and Russian catalogs write "Момо | Momo" — localized title, then original. The widening corrupted 349 keys and rescued none, so it is reverted. What survives is what the data supports: the pipe-lookalike set (0 changed keys, and correct for panels that use them) and dropping the required space after a pipe (35 changed keys, genuine welded tags like "EN|Dark Shadows" and "|FR|VO|Le dernier empereur"). A leading-tag guard that refused to strip when no letter remained is also dropped: it fixes "AKA | 2023" but breaks "IT - 65", so telling those apart needs a tag vocabulary and belongs in its own change. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
1 parent
2ddd38facc
commit
29a81baec1
5 files changed
+66
-52
No files matched your search
@@ -5,6 +5,6 @@ area: xtream
|
||||
|
||||
The movie sources popover now recognizes more language tags: prefixes with
|
||||
Unicode pipes, brackets or dashes ("EN │ …", "[EN] …", "EN - …"), Cyrillic
|
||||
tags ("РУС | …") and MULTI. Those tags are also stripped when matching, so
|
||||
copies carrying them are found as the same movie, and the language filter
|
||||
now also reads the language off category names ("EN | Netflix").
|
||||
tags ("РУС | …") and MULTI. It also reads the language off category names
|
||||
("EN | Netflix"), and a tag welded to a title ("EN|Movie") no longer hides
|
||||
that copy from the other playlists' copies of the same movie.
|
||||
@@ -1209,7 +1209,7 @@ engine` (restart required) or
|
||||
|
||||
**VOD Multi-Source** (alternative sources for a movie):
|
||||
|
||||
- Finds the same movie in the user's other imported playlists and adds a "Sources N" chip to the Xtream VOD action row (only when ≥1 alternative exists), plus a `.source-caption` line reporting where playback is coming from. The chip opens a 660px anchored CDK-overlay popover (`libs/ui/components/src/lib/vod-sources/`; not `MatMenu`, which caps its width at 280px), reused unchanged in the inline player's now-playing bar and on the playback-error screen. It opens ABOVE the chip (right edges aligned, pressed state on the chip while open), height-capped by the overlay's flexible bounding box so only the source list scrolls, and flips below when less than the overlay `minHeight` remains above; filter chips (All / Available / HD+ / language select) compose with the host search, "Available" auto-runs check-all when no verdicts exist, and expanded copy rows show a parsed language chip + raw stream title with diff-only tags ("same as above" for the parent's copy). A row's language is `vodSourceLanguage` (`libs/shared/interfaces/src/lib/vod-source-language.util.ts`): the title's own prefix (pipe incl. Unicode lookalikes, bracketed, or ALL-CAPS spaced-dash form; Latin/Cyrillic 2–4 letters + `MULTI`; only the legacy pipe form is permissive — bracket/dash matches must also pass `isKnownLanguageTag`, since those positions carry quality/rip tags like `[HD]`) wins, else the language the stream's visible categories unambiguously carry ("EN | Netflix" — discovery aggregates `group_concat(cat.name, char(31))` per (playlist, stream), prefixed categories must agree, and category prefixes must pass `isKnownLanguageTag`, since `new`/`top`/`hot` are real ISO 639-3 codes but everyday category words; the route's own row reads the one category the route arrived through, overlaid late by the host's same-key `refreshRouteFacts` since cold/direct routes load categories after discovery). Both forms are parsed guesses: browse filter and chips only, never ranking/failover/dub-warning inputs. Recognition alone is not enough — `normalizeTitleKeys` must STRIP the same tag or the copy is never discovered, so its leading-tag rule shares `PROVIDER_PIPE_CLASS` and, on the pipe branch only, takes the same Latin+Cyrillic any-case alphabet with no required trailing space; dash/colon stay uppercase-Latin + spaced, since loosening them amputates "ОНО: Часть 2"/"It: Chapter Two". Checks run through a 4-slot queue and settled verdicts are cached 10 min per movie+source (`VodSourceProbeCacheService`). Both chips are handed the same `matchKind` and `vodAutoFailover` and both write the setting back. The details-page chip badge counts TOTAL **copies** across all playlists (the in-player chip still counts alternatives); the caption ("also found in N other playlists") counts distinct **playlists** via `alternativePlaylistCount`, because the popover groups one portal's copies under that portal. The action row's Favorites and Download buttons are icon-only 64px squares: filled red heart when favorited, and a download idle icon → progress ring (real percent, indeterminate spin, paused-resume) → green done-checkmark whose click reveals the file (state read from the download manager; the labeled "Play from source" secondary is gone — provider playback for a downloaded movie goes through the Sources popover).
|
||||
- Finds the same movie in the user's other imported playlists and adds a "Sources N" chip to the Xtream VOD action row (only when ≥1 alternative exists), plus a `.source-caption` line reporting where playback is coming from. The chip opens a 660px anchored CDK-overlay popover (`libs/ui/components/src/lib/vod-sources/`; not `MatMenu`, which caps its width at 280px), reused unchanged in the inline player's now-playing bar and on the playback-error screen. It opens ABOVE the chip (right edges aligned, pressed state on the chip while open), height-capped by the overlay's flexible bounding box so only the source list scrolls, and flips below when less than the overlay `minHeight` remains above; filter chips (All / Available / HD+ / language select) compose with the host search, "Available" auto-runs check-all when no verdicts exist, and expanded copy rows show a parsed language chip + raw stream title with diff-only tags ("same as above" for the parent's copy). A row's language is `vodSourceLanguage` (`libs/shared/interfaces/src/lib/vod-source-language.util.ts`): the title's own prefix (pipe incl. Unicode lookalikes, bracketed, or ALL-CAPS spaced-dash form; Latin/Cyrillic 2–4 letters + `MULTI`; only the legacy pipe form is permissive — bracket/dash matches must also pass `isKnownLanguageTag`, since those positions carry quality/rip tags like `[HD]`) wins, else the language the stream's visible categories unambiguously carry ("EN | Netflix" — discovery aggregates `group_concat(cat.name, char(31))` per (playlist, stream), prefixed categories must agree, and category prefixes must pass `isKnownLanguageTag`, since `new`/`top`/`hot` are real ISO 639-3 codes but everyday category words; the route's own row reads the one category the route arrived through, overlaid late by the host's same-key `refreshRouteFacts` since cold/direct routes load categories after discovery). Both forms are parsed guesses: browse filter and chips only, never ranking/failover/dub-warning inputs. Recognition alone is not enough — `normalizeTitleKeys` must STRIP the same tag or the copy is never discovered, so its leading-tag rule shares `PROVIDER_PIPE_CLASS` and drops the required space after a pipe. It goes no further on purpose: a wrong guess costs a filter option, a wrong strip corrupts identity, and on 1.27M real titles a case-insensitive/Cyrillic pipe rule corrupts 349 keys ("Akira | 1988", "Момо | Momo" — the name sits in the tag position) while `–`/`—` on the dash branch amputates 14 subtitled titles. Verify such widenings against the real catalog before shipping them. Checks run through a 4-slot queue and settled verdicts are cached 10 min per movie+source (`VodSourceProbeCacheService`). Both chips are handed the same `matchKind` and `vodAutoFailover` and both write the setting back. The details-page chip badge counts TOTAL **copies** across all playlists (the in-player chip still counts alternatives); the caption ("also found in N other playlists") counts distinct **playlists** via `alternativePlaylistCount`, because the popover groups one portal's copies under that portal. The action row's Favorites and Download buttons are icon-only 64px squares: filled red heart when favorited, and a download idle icon → progress ring (real percent, indeterminate spin, paused-resume) → green done-checkmark whose click reveals the file (state read from the download manager; the labeled "Play from source" secondary is gone — provider playback for a downloaded movie goes through the Sources popover).
|
||||
- Scope v1 is **Xtream ↔ Xtream, movies only, Electron only**. Stalker never reaches the `content` table and M3U is a JSON blob whose search forces `content_type:'live'`; both are additive later since `VodSourceCandidate.portalType` already carries all three. In the PWA every entry point is gated off by a bridge `typeof` check and the chip renders nothing.
|
||||
- **Metadata provenance is the core contract.** Every field is `{value, provenance}` where `api`/`probe` are facts (plain tag), `parsed` is a title-regex guess (tag prefixed `~`, warn colour), and absent renders **no tag at all** plus a `check` chip. `factualOnly()` in `vod-source-metadata.util.ts` is the only accessor allowed for ranking/failover, so guesses are structurally unable to influence a decision. `VodSourceProbeStatus` separates `fail` (contacted and refused) from `unknown` (timed out / blocked / no capability) — an unchecked source is never shown as offline. Quality is derived from pixel **width** because letterboxing crops height — but a known height vetoes the answer on every tier, since cropping only removes lines: a taller frame is a different shape (1440×1080 anamorphic or 1600×900 are not 720p, 960×540 is not 576p) and gets no tag rather than a wrong one carrying `api` provenance. The route's OWN row is never resolved, so it takes its facts from the `get_vod_info` the page already loaded (`providerVodMetadataOf`, shared with the resolver) and picks them up via `refreshRouteFacts()` even when they arrive without changing the movie identity — otherwise `audioDiffersFactually` has nothing on one side and the dub warning cannot fire on a route-to-alternative switch.
|
||||
- Discovery (`DB_FIND_TITLE_SOURCES`, trigram FTS over `content_title_fts`) is lazy and returns only what the `content` table can prove; titles whose tokens are all shorter than three characters ("Up", "It") fall back to a scan, since the trigram tokenizer cannot index them at all. A source that is never read looks exactly like one that does not exist, so: the current playlist is excluded **in SQL** and duplicates collapse there too (`GROUP BY cat.playlist_id, c.xtream_id` before the limit — one playlist's dozens of identically ranked category rows would otherwise crowd out every alternative), and the scan matches an ASCII token as a whole word (`' ' || LOWER(title) || ' ' GLOB '*[^a-z0-9]it[^a-z0-9]*'`) ordered by title length **with no row limit** — FTS keeps its 60-row window because it ranks by relevance, while a scan cannot rank, and the GLOB reads every row regardless so a limit would only truncate the answer. The year gate covers BOTH match tiers: `normalizeTitleKeys` strips bracketed segments, so "Dune (1984)" normalizes identically to "Dune" and would otherwise be an _exact_ match for the 2021 film; a bracketed year is read out of the raw title and a stated disagreement rejects the row — but the two tiers read different forms: the base tier accepts bracketed or trailing (it just stripped a trailing year, the only thing separating "Dune 1984" from "Dune 2021"), while the exact tier reads bracketed ONLY, since reaching it means both titles are the same string and a trailing number is then part of the NAME ("Blade Runner 2049" against a metadata year of 2017 would otherwise vanish once enrichment lands). A non-ASCII token cannot be folded by `LOWER()` (ASCII-only) but CAN be by a GLOB character class (UTF-8 code points), so `caseInsensitiveGlobPattern` folds the case in JS and emits one `[lowerUpper]` class per character — returning `null`, leaving the two substring tests alone, for a GLOB metacharacter or a length-changing case map (`ß`→`SS`). The movie's own year comes from `releaseTagYear` (bracketed or trailing only), never `extractYear`: a year inside the NAME ("2001: A Space Odyssey") would fail every genuine 1968 copy at the year gate and move the pin key once enrichment lands. One row inside the excluded playlist is kept when the caller names it (`keepContentId`), because a pin can point at another copy in the playlist being viewed — the host reads the pin before discovery for exactly this. Resolution is deferred to click/pin/check because `content` stores no `container_extension` and `constructVodUrl` returns `''` without one — each alternative costs a live `get_vod_info` against the foreign playlist's credentials.
|
||||
|
||||
@@ -199,12 +199,22 @@ get the row excluded by the very filter meant to find it.
|
||||
Recognizing a prefix is only half the job: the same tag also has to be
|
||||
STRIPPED by `normalizeTitleKeys`, or the tagged copy and the bare one never
|
||||
match and the row is never discovered at all. Its leading-tag rule therefore
|
||||
shares this file's pipe set (`PROVIDER_PIPE_CLASS`) and, on the pipe branch
|
||||
only, takes the same Latin+Cyrillic any-case alphabet and needs no space
|
||||
after the separator — so "РУС | Дюна", "ru| Dune" and "EN|Dune" reach the key
|
||||
"dune". Dash and colon keep their uppercase-Latin, space-required form: those
|
||||
are ordinary title punctuation, and loosening them would amputate
|
||||
"ОНО: Часть 2" the way a case-insensitive rule amputates "It: Chapter Two".
|
||||
shares this file's pipe set (`PROVIDER_PIPE_CLASS`) and, on the pipe branch,
|
||||
needs no space after the separator — so "EN │ Fallout", "EN|Fallout" and
|
||||
"|FR|VO|Le dernier empereur" reach the same key as the bare title.
|
||||
|
||||
It stops there, and the asymmetry with the reader above is the point. A wrong
|
||||
GUESS costs a junk option in a filter; a wrong STRIP corrupts a film's
|
||||
identity everywhere the key is used. Measured against 1.27M real catalog
|
||||
titles: making the pipe branch case-insensitive or Cyrillic corrupts 349 keys
|
||||
and rescues none, because "Akira | 1988" and "Момо | Momo" put the film's own
|
||||
name in the tag position; widening the dash branch to `–`/`—` amputates 14
|
||||
subtitled titles ("1918 – A Batalha de Kruty"); and Cyrillic before a dash
|
||||
does not occur at all. So normalization keeps its uppercase-Latin,
|
||||
space-required form everywhere except the pipe separator itself, and the
|
||||
reader is free to be permissive because the gate in front of its riskier
|
||||
forms — and the fact that a row only appears once it HAS matched — keeps a
|
||||
bad guess cosmetic.
|
||||
|
||||
The category path exists because many panels tag the CATEGORY ("EN | Netflix",
|
||||
"DE | Apple TV") and leave stream titles bare. Discovery aggregates every
|
||||
|
||||
@@ -113,24 +113,38 @@ describe('provider tag stripping', () => {
|
||||
);
|
||||
});
|
||||
|
||||
it('strips Cyrillic and lowercase tags before a pipe', () => {
|
||||
// A pipe is a strong tag signal, so its branch takes the wider
|
||||
// alphabet and needs no trailing space: without this the tagged copy
|
||||
// and the bare one never match, and multi-source cannot offer the
|
||||
// film at all.
|
||||
expect(normalizeTitle('РУС | Дюна')).toBe('дюна');
|
||||
expect(normalizeTitle('укр│ Дюна')).toBe('дюна');
|
||||
expect(normalizeTitle('ru| Dune')).toBe('dune');
|
||||
expect(normalizeTitle('EN|Dune')).toBe('dune');
|
||||
expect(normalizeTitle('|РУС| Дюна')).toBe('дюна');
|
||||
it('reads pipe lookalikes as the pipe they look like', () => {
|
||||
// `│`, `¦` and `|` are indistinguishable from `|` in a catalog, so
|
||||
// the same tag must not survive in one playlist and vanish in
|
||||
// another — the two copies would never match as the same film.
|
||||
expect(normalizeTitle('EN │ Fallout')).toBe('fallout');
|
||||
expect(normalizeTitle('DE ¦ Fallout')).toBe('fallout');
|
||||
expect(normalizeTitle('MULTI|Fallout')).toBe('fallout');
|
||||
});
|
||||
|
||||
it('keeps that widening away from dash and colon', () => {
|
||||
// Where the separator is ordinary punctuation, an uppercase-Latin
|
||||
// restriction is the only thing standing between a tag and a real
|
||||
// title — "ОНО: Часть 2" is the Cyrillic "It: Chapter Two".
|
||||
expect(normalizeTitle('ОНО: Часть 2')).toBe('оно часть 2');
|
||||
expect(normalizeTitle('ОНО - Часть 2')).toBe('оно часть 2');
|
||||
it('strips a pipe tag welded to the title', () => {
|
||||
// "|FR|VO|Le dernier empereur" — the wrapped tag goes first, then
|
||||
// "VO|" with no space after it.
|
||||
expect(normalizeTitle('|FR|VO|Le dernier empereur')).toBe(
|
||||
'le dernier empereur'
|
||||
);
|
||||
expect(normalizeTitle('EN|Fallout')).toBe('fallout');
|
||||
});
|
||||
|
||||
it('keeps a name that only looks like a tag before a pipe', () => {
|
||||
// Pins the UPPERCASE rule on the pipe branch. Relaxing it there is
|
||||
// tempting — nothing but a tag precedes a pipe, surely — but measured
|
||||
// against 1.27M real catalog titles a case-insensitive (or Cyrillic)
|
||||
// pipe rule corrupted 349 keys and rescued none: "name | year" and
|
||||
// the Russian "localized | original" convention both put the film's
|
||||
// own name in the tag position.
|
||||
// The base tier drops the trailing year, so the name is what must
|
||||
// survive; the exact tier shows the whole string it came from.
|
||||
expect(normalizeTitle('Akira | 1988')).toBe('akira');
|
||||
expect(normalizeTitleKeys('Akira | 1988').exact).toBe('akira 1988');
|
||||
expect(normalizeTitle('Coco | 2017')).toBe('coco');
|
||||
expect(normalizeTitle('Момо | Momo')).toBe('момо momo');
|
||||
expect(normalizeTitle('Мумия | The Mummy')).toBe('мумия the mummy');
|
||||
});
|
||||
|
||||
it('keeps bare 4-5 char words before a spaced dash (real titles)', () => {
|
||||
@@ -268,7 +282,6 @@ describe('provider tag stripping', () => {
|
||||
'EN │ Breaking Bad',
|
||||
'DE ¦ Breaking Bad',
|
||||
'FR|Breaking Bad',
|
||||
'en| Breaking Bad',
|
||||
];
|
||||
|
||||
it.each([
|
||||
|
||||
@@ -42,25 +42,13 @@ const QUALITY_TAGS = new Set([
|
||||
*/
|
||||
export const PROVIDER_PIPE_CLASS = '[|¦│┃❘∣⏐⎪︱︳丨|]';
|
||||
|
||||
/**
|
||||
* The alphabet a pipe-delimited tag may use: Latin or Cyrillic, either case
|
||||
* — the same set the language-prefix reader accepts, so the two files agree
|
||||
* on what a tag looks like. Every segment must still contain a letter, so a
|
||||
* numeric fragment is never read as one.
|
||||
*
|
||||
* Wider than the uppercase-Latin `SEG` below because it is only ever used
|
||||
* where a pipe is the separator. Nothing but a tag precedes a pipe: it is
|
||||
* not valid in a Windows filename and does not occur in real titles.
|
||||
*/
|
||||
const PIPE_LETTER = 'A-Za-zА-Яа-яЁё';
|
||||
const PIPE_SEG = `(?=[0-9+]*[${PIPE_LETTER}])[${PIPE_LETTER}0-9+]`;
|
||||
|
||||
/**
|
||||
* Wrapped tag at the very start of a provider title: "|DE| ARD",
|
||||
* "|MULTI| Fallout", "|РУС| Дюна".
|
||||
* "|MULTI| Fallout". The lookahead requires a letter in the tag so a
|
||||
* numeric fragment can never be treated as one.
|
||||
*/
|
||||
const WRAPPED_TAG_PREFIX = new RegExp(
|
||||
`^\\s*${PROVIDER_PIPE_CLASS}${PIPE_SEG}{2,5}${PROVIDER_PIPE_CLASS}\\s*`
|
||||
`^\\s*${PROVIDER_PIPE_CLASS}(?=[0-9+]*[A-Z])[A-Z0-9+]{2,5}${PROVIDER_PIPE_CLASS}\\s*`
|
||||
);
|
||||
|
||||
/**
|
||||
@@ -78,23 +66,26 @@ const WRAPPED_TAG_PREFIX = new RegExp(
|
||||
* - colon ("EN: "): 2–3 chars — longer acronyms are franchise titles
|
||||
* ("NCIS: LA")
|
||||
*
|
||||
* The pipe branch takes the wider `PIPE_SEG` alphabet (Cyrillic as well as
|
||||
* Latin, either case) and does not require a space after the separator.
|
||||
* "РУС | Фильм", "ru| Movie" and "EN|Movie" otherwise keep their tag welded
|
||||
* to the title, so the tagged copy never matches the bare one and
|
||||
* multi-source cannot offer the film at all.
|
||||
* The pipe branch alone does not require a space after the separator:
|
||||
* "|FR|VO|Le dernier empereur" welds the tag to the title, and no real title
|
||||
* begins with a short word immediately followed by a pipe.
|
||||
*
|
||||
* Neither widening is extended to dash and colon: those are ordinary title
|
||||
* punctuation, and the uppercase-Latin restriction plus the required space
|
||||
* are what keep "ОНО: Часть 2" and "Spider-Man" intact — the Cyrillic and
|
||||
* hyphenated cases of the "It: Chapter Two" hazard.
|
||||
* The UPPERCASE-only restriction stays on every branch, pipe included. It is
|
||||
* tempting to drop it there on the theory that nothing but a tag precedes a
|
||||
* pipe — 1.27M real catalog titles say otherwise, and in two ways at once:
|
||||
* "Akira | 1988" and "Coco | 2017" put the film's NAME before the pipe and
|
||||
* the year after it, and Russian catalogs write "Момо | Momo",
|
||||
* "Мумия (2026) | Lee Cronin's The Mummy" — the localized title, then the
|
||||
* original. Case-insensitivity (or a Cyrillic alphabet) turns every one of
|
||||
* those names into a "tag" and strips it; measured against that corpus it
|
||||
* corrupted 349 keys and rescued none.
|
||||
*/
|
||||
const SEG = '(?=[0-9+]*[A-Z])[A-Z0-9+]';
|
||||
const COMPOUND_TAG = `${SEG}{2,5}(?:-${SEG}{2,6}){1,2}`;
|
||||
const LANGUAGE_PREFIX = new RegExp(
|
||||
'^(?:' +
|
||||
`(?:${COMPOUND_TAG}|${SEG}{2,3})\\s*-\\s+` +
|
||||
`|(?:${COMPOUND_TAG}|${PIPE_SEG}{2,5})\\s*${PROVIDER_PIPE_CLASS}\\s*` +
|
||||
`|(?:${COMPOUND_TAG}|${SEG}{2,5})\\s*${PROVIDER_PIPE_CLASS}\\s*` +
|
||||
`|${SEG}{2,3}\\s*:\\s+` +
|
||||
')'
|
||||
);
|
||||
|
||||
Reference in new issue
Block a user