mirror of
https://github.com/4gray/iptvnator.git
synced 2026-10-11 11:06:16 -08:00
Cross-playlist matching compares normalized titles ("Amélie" -> "amelie")
against an index built from the raw title, and the trigram tokenizer does not
fold diacritics by default. Every accented title was therefore invisible to
the FTS path: two identical `Amélie` entries produced no candidates at all.
That is the broadest of the Unicode gaps review found, and it predates the
short-title work.
The tokenizer is fixed at CREATE time, so existing databases recreate and
rebuild the index once behind a migration marker. `remove_diacritics` needs
SQLite 3.45+, so support is probed on a temp table first: an older runtime
keeps its working index untouched and the migration is not recorded as done,
leaving a later version free to upgrade it.
Case folding for non-ASCII remains impossible in stock SQLite — "ОН" cannot
find "Он" by any available predicate — and is documented as the known limit
rather than patched around again.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>