Files
iptvnator/libs/portal/shared/data-access
4grayandClaude Fable 5 36c2867d36 feat(xtream): recognize more language tags in VOD multi-source (#1417)
* feat(xtream): recognize more language tags in VOD multi-source

The sources popover's language filter and copy chips now read prefixes
with Unicode pipe lookalikes, brackets and spaced dashes, Cyrillic tags
and MULTI. When a stream title carries no tag, the language falls back
to what the stream's visible categories unambiguously state ("EN |
Netflix") — discovery aggregates category names per (playlist, stream)
in SQL, and category prefixes must pass a known-language gate because
everyday category words like new/top/hot are real ISO 639-3 codes.

Both signals stay parsed guesses: browse filter and chips only, never
ranking, failover or dub-warning inputs.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(xtream): address Codex review on multi-source language detection

Gate the new bracket and dash title forms through isKnownLanguageTag:
those positions carry quality/rip tags ([HD], [CAM], NEW -) whose
fabricated "language" would outrank and mask a real category-derived
one. The legacy pipe form stays permissive.

Overlay a late-arriving route category onto the existing route row in
the same-key refresh path — cold/direct routes load categories after
discovery, and the category is outside the movie key on purpose. The
mid-flight case is redelivered by the bind() effect re-running on the
controller's sources signal; that tracked read is now documented as
load-bearing and pinned by a session spec.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(xtream): pair brackets and strip the new tag forms when matching

Greptile: the bracket prefix chose its opening and closing delimiter
independently, so a malformed "[EN)" was read as a language tag.

Codex: recognizing a prefix is only half the job — normalizeTitleKeys
has to strip the same tag, or the tagged copy never matches the bare
one and multi-source cannot offer the film at all. Its leading-tag rule
now shares the pipe-lookalike set and, on the pipe branch only, takes
the same Latin+Cyrillic any-case alphabet with no required trailing
space. Dash and colon keep their uppercase-Latin spaced form: those are
ordinary punctuation, and loosening them would amputate "ОНО: Часть 2"
the way a case-insensitive rule amputates "It: Chapter Two".

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(xtream): keep normalization uppercase-only, measured on real catalogs

The previous commit widened the pipe branch of normalizeTitleKeys to any
case and to Cyrillic, on the theory that nothing but a tag precedes a
pipe. Checked against 1.27M real catalog titles that theory is wrong in
two ways at once: "Akira | 1988" and "Coco | 2017" put the film's name
before the pipe and the year after it, and Russian catalogs write
"Момо | Momo" — localized title, then original. The widening corrupted
349 keys and rescued none, so it is reverted.

What survives is what the data supports: the pipe-lookalike set (0
changed keys, and correct for panels that use them) and dropping the
required space after a pipe (35 changed keys, genuine welded tags like
"EN|Dark Shadows" and "|FR|VO|Le dernier empereur").

A leading-tag guard that refused to strip when no letter remained is
also dropped: it fixes "AKA | 2023" but breaks "IT - 65", so telling
those apart needs a tag vocabulary and belongs in its own change.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(xtream): cite the measured evidence for the category language gate

The gate's rationale named hypothetical category shapes. On a real
catalog the four it actually turns away are VOD (5,245 movies), KIDS
(1,010), SHOW and WWE — without it the language select offers "VOD" and
"KIDS" as languages. Comments, doc and one spec case only.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(xtream): restore the docblock currentSourceRow lost to an insertion

routeCategoryLanguage was added between currentSourceRow's docblock and
its signature, so the paragraph describing "the row standing for the
source the route is already playing" ended up introducing a function
that returns a language string. Moved below; no behavior change.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(xtream): stop grouping the scan tier, it can drop a matching source

content is unique per (category, type, stream), so one stream sitting in
several categories is several rows and nothing forces their titles to
agree. The GROUP BY added for group_concat let SQLite keep an arbitrary
row's title, and the normalized confirmation then rejected the whole
stream on a title a sibling row would have matched — the source vanished.

The FTS tier can afford that grouping because its window makes it
necessary; the scan tier takes no window at all, so it now returns a row
per category and their names are merged per stream in TypeScript, which
also keeps the rejected sibling's category in the language derivation.

Found by Codex. Latent rather than active on the catalog I measured (0
streams currently carry differing titles across categories), but the
schema permits it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(xtream): record why category names stay scoped to matched rows

Codex flagged that the FTS predicate runs before the aggregate, so a
sibling row under a localized title contributes no category. True, and
deliberate: the field is a guess feeding a chip and a browse filter, and
completing it costs measured latency — 0.74s to 2.0s for a correlated
subquery on a 3.9GB catalog, 19.7s for a second bounded lookup — to
correct a cosmetic guess in a shape that occurs 0 times in 2.7M rows.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-12 07:59:07 +02:00
..