Commit Graph
6 Commits
Author SHA1 Message Date
4grayandClaude Fable 5.1 8ebb7e3424 perf(ci): run Tier A coverage concurrently with isolatedModules ts-jest (#1701)
Tier A coverage runs projects a few at a time (largest first, bounded Jest workers, buffered output, fail-fast kept) and ts-jest transpiles with isolatedModules instead of type-checking per process; five type re-exports become export type, two decorated inputs use import type. Unit Tests and Typechecks job: 26 min -> 9 min (Tier A step 23 min -> 6.5 min).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-26 23:05:26 +02:00
4grayandClaude Fable 5.1 01c423ac43 fix(portals): stop treating a slow panel as a dead host; IPv4 fallback budget in Electron (#1621)
* fix(portals): apply the IPv6->IPv4 fallback budget in the Electron process

The 2500 ms happy-eyeballs attempt timeout from #1404 only ever ran in the
web backend. The Electron main process and its playlist-refresh and EPG
workers kept Node's 250 ms default, so a dual-stack panel hostname behind a
VPN or a slow link failed every connection attempt in a row and tripped the
host connectivity guard. The module now lives in `@iptvnator/shared/host-health`
and every Node isolate that opens connections applies it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(portals): stop treating a slow panel as a dead one in the host guard

axios raises the same ECONNABORTED whether the SYN went unanswered or the
panel accepted the connection and then thought for longer than the request
budget. Two such timeouts opened the breaker and every request to the panel
was refused for 30 s with "portal is not responding" — the shape behind the
"connection keeps dropping" reports on 0.23 and nightly.

Both transports now report whether the TCP connection was established
(`onConnect`: Electron through a per-request observed agent instead of the
shared keep-alive globalAgent, the web backend through the transport that owns
the ClientRequest), and `classifyHostRequestFailure(error, { connected })`
downgrades a host-level code observed after the handshake to inconclusive.
Redirect attribution keeps precedence. A host that never accepts the
connection trips the guard exactly as before.

The Xtream mock gains a `silent:silent` scenario whose detail actions accept
and never answer, plus a real-socket regression spec for the guard.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(portals): let an accepted connection clear the host-failure streak

Review finding: an unanswered SYN, then an accepted-but-slow timeout, then
another unanswered SYN still reached the two-failure threshold, because the
middle request was merely not counted. An accepted TCP connection is the
reachability the guard measures, so it now reads as `responded` and clears
the streak like an HTTP response would. Regression coverage for the mixed
sequence on one flapping loopback origin (Electron) and through the proxy
route (web backend).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(portals): credit an accepted connection when it happens, not when the request settles

Review findings. A request that connected and then hung for 30 s cleared,
on its eventual timeout, the failures later requests had recorded while it
waited — reopening a host that had just died on evidence older than theirs.
The connect hook now reports the connection the moment it fires through a
new `HostConnectivityGuard.reportConnected`, which clears the failure streak
but closes no open or half-open breaker (the trial keeps its slot until it
settles), and the settled timeout is inconclusive.

Electron also skips the socket observer while an environment proxy
(`http_proxy` / `https_proxy` / `all_proxy`) applies to the request: through
a proxy the socket connects to the proxy, whose handshake proves nothing
about the portal, so those requests keep the pre-observer behaviour.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(portals): decide the proxy exemption with axios' own resolution

Review findings. The hand-rolled environment check ignored `no_proxy`, so a
LAN portal exempted from the proxy lost its connect observer and slow
requests to it still tripped the breaker; it also read the variables with
`??`, letting an empty lowercase one mask a populated uppercase one that
axios would honour. The decision now calls `proxy-from-env`'s
`getProxyForUrl`, the same pinned package axios' http adapter uses,
declared as a direct dependency so the packaged app carries it.

The validated-axios spec clears and restores every proxy variable around each
case, so a runner that exports a proxy cannot change what the cases prove.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 14:50:50 +02:00
4gray e9eca1c386 chore(deps): upgrade Angular to 22.1 and Nx to 23.2 (#1603)
* chore(deps): upgrade Angular to 22.1 and Nx to 23.2

* fix(deps): complete Angular migrations after rebasing on master

* fix(ci): use the Node pin for Windows runtime refresh

* docs(deps): synchronize the workspace-shell Node requirements
2026-09-14 19:02:40 +02:00
4gray e40f31db97 fix(host-health): retain trial ownership until requests settle (#1547) 2026-09-06 00:43:54 +02:00
4gray 1c56b3ea02 fix(portals): count concurrent connection failures once (#1537)
* fix(portals): count concurrent connection failures once

* docs(portals): clarify failure counting after guard eviction
2026-09-05 12:21:49 +02:00
4grayandClaude Opus 5 e94cc029eb fix(portals): time out and fast-fail PWA proxy requests to dead hosts (#1424)
* refactor(portals): hoist the connectivity guard into libs/shared/host-health

The breaker was written dependency-free so both processes that talk to
portals could share it. Move the part that has no Electron in it — the
state machine, the failure classification, the redirect-attribution
helpers and the fast-fail error — into `@iptvnator/shared/host-health`
(`scope:shared` / `domain:shared-runtime` / `type:util`).

What stays in `apps/electron-backend` is the genuinely main-process part:
one guard for the whole process, so both portal IPC handlers see each
other's evidence, and the console warning that announces it. Every call
site is unchanged; the wrapper re-exports the two types they import.

The spec splits the same way — the state machine moves with the class,
the singleton and its redirect attribution stay with the wrapper.

Register the new project in the coverage policy, which every project with
a test target must declare a tier for.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(portals): time out and fast-fail PWA proxy requests to dead hosts

The web backend's proxy routes were bare `axios.get()` calls with no
`timeout`, so a provider that accepted a connection and then went silent
held the request until the OS gave up on the TCP connection — minutes,
rather than the 15/30 s budget the Electron handlers use. Add the same
per-route timeouts (Xtream 30 s, Stalker 15 s / 30 s for `create_link`,
playlist and XMLTV 30 s).

Those numbers are safe for large downloads: on axios' default transport
`timeout` bounds the time to response headers and then continues as the
socket's inactivity timeout, so a multi-megabyte XMLTV file that keeps
delivering bytes is never cut off — only a stalled one is.

With requests bounded, run `/xtream` and `/stalker` through the shared
breaker, injected via `WebBackendAppOptions.hostGuard` so specs drive it
with a clock they own. Playlist and XMLTV downloads keep the timeout but
no breaker, matching Electron: a download is one request rather than a
catalog fan-out, it is usually the direct result of the user asking for
it, and it can outlive the half-open trial window.

The breaker is checked before the Xtream URL revalidation, which resolves
the hostname — a dead host is where DNS is slow too, and a request
admitted and then abandoned by the URL policy hands its token back rather
than holding the trial slot.

`resetHostConnectivityGuard()` no longer no-ops in the PWA: the breaker
lives in the backend process, so it travels to a new
`POST /connectivity-guard/reset`, which reads only the origin and never
logs the credential-bearing URL. `skipConnectionGuard` now survives the
PWA transport too, so Stalker endpoint discovery keeps the exemption it
has on the desktop instead of tripping the breaker with its own probes.

A fast-fail keeps each route's HTTP 200 `{message, status}` envelope. The
Stalker path needs one extra step: `forwardStalkerRequest` turns that
envelope into `HTTP Error <code>: …` with a numeric `status`, and the
renderer reads both as "the endpoint answered" — which would make
discovery walk every candidate and fire lazy repair at a host just
declared dead. A prior branch keyed on the shared
`isHostConnectivityFastFailMessage` rethrows it bare instead.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(portals): report exempt Stalker probe responses to the PWA breaker

Flagged by the author of #1421 as one of the twelve fixes that landed
there after this branch cherry-picked the pre-review commit: an exempt
discovery probe must still REPORT, it just must not COUNT.

The web backend was skipping the report entirely for a probe, which loses
the case that matters. A failure carrying an HTTP response proves the
endpoint answered, and this route sets no `validateStatus`, so axios
rejects every non-2xx with `error.response` attached — a probe answered
with 404 or 500 was therefore dropped instead of clearing the record.
Two counted failures either side of it then read as consecutive and
opened the breaker in the middle of discovery, which is exactly what the
exemption exists to prevent.

`reportProviderRequestFailure` now takes `countFailures`, matching the
Electron reporter, and both routes always report.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(portals): read the redirect hop from the transport that followed it

`failedAfterRedirect` decided whether a failure belonged to a redirect
destination by reading `error.config.url`. That is right for the Electron
transport, which sets `maxRedirects: 0` and reissues every hop as its own
request, so the hop IS the config URL. It is blind on the web backend,
which uses axios' default transport: follow-redirects walks the chain
inside one request and `config` is built once, so `config.url` stays the
URL we asked for.

Verified against the installed axios 1.19.0 with a live server that 302s
to a dead port:

    asked for          : http://127.0.0.1:63953/player_api.php
    config.url         : http://127.0.0.1:63953/player_api.php
    request._currentUrl: http://127.0.0.1:1/dead

So the comparison was original-vs-original, found no redirect, and
charged two dead destinations to the provider that had answered both
times with a 302 — then fast-failed it. Read `request._currentUrl` first
and fall back to `config.url`, which covers both transports; Electron's
native per-hop requests expose no `_currentUrl` and are unaffected.

Also check redirect attribution BEFORE suppressing failure counting for
an exempt probe. A 3xx from the guarded endpoint is an answer, so a probe
that observed one must clear the record; otherwise a timeout, a probe
redirected to a dead destination, and another timeout still read as two
consecutive failures.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(portals): stop axios query params reading as a redirect

Codex found that the Xtream breaker never opened at all, and it was
right. The web backend passes credentials and the action through axios'
`params`, so axios sends `…/player_api.php?username=…&action=…` while the
baseline handed to `failedAfterRedirect` is the query-less URL the route
built. Verified against axios 1.19.0 with a plain ECONNREFUSED and no
redirect anywhere in sight:

    baseline           : http://127.0.0.1:1/player_api.php
    request._currentUrl: http://127.0.0.1:1/player_api.php?username=demo&…

The two normalized URLs differ, so every ordinary failure looked like a
post-redirect failure, credited the endpoint, and the breaker could never
trip.

Compare origin and path, not the whole URL. That keeps what the check is
for — an endpoint that answered and sent us elsewhere, including the
same-origin `/player_api.php` → `/slow/player_api.php` case — and gives
up only a redirect that changes nothing but the query, which is then
counted as an ordinary failure. Erring towards counting is the safe
direction here.

The reason 57 tests passed over a dead feature is the real lesson:
`StubHttpClient` threw bare `Error`s, so the guard's redirect check saw
neither `config.url` nor `request._currentUrl` and quietly did nothing.
The stub now shapes its rejections like axios does, including the query
axios appends. With that alone, four existing tests fail against the old
comparison.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(portals): count a hostname that stops resolving, and fix a stale docblock

Two from review, one behavioural and one documentation.

A name that will not resolve is the host failing to answer — the same
evidence as the ENOTFOUND the transport would have raised a moment later.
But the SSRF validation turns a lookup failure into a 400 "host could not
be resolved", and the release path added earlier handed the token back as
inconclusive, so the breaker could never open for a host whose DNS died
and every request kept paying for the same dead lookup.

`ProviderUrlError` now carries the underlying lookup error internally.
A refusal that has one is counted; a genuine policy refusal — private
address, bad scheme, credentials in the URL — still only releases the
half-open slot, because that says nothing about reachability. The field
is internal: `providerUrlErrorBody()` strips it at both call sites, so
the client sees exactly the body it saw before, which the test asserts.

The docblock on `resetHostConnectivityGuard` still said the PWA channel
is unknown and the call no-ops. That stopped being true when this branch
implemented `CONNECTIVITY_GUARD_RESET` over HTTP, and a stale contract
there is how the next caller silently skips the PWA path. (The edit was
in an earlier commit and was lost when the branch was rebuilt on master.)

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(portals): record the transport-specific redirect contract

The redirect section still described one transport: hop-by-hop requests,
`error.config.url`, whole-URL comparison. Two of those three are now
wrong for the web backend, and this document is the canonical contract —
leaving it stale is how the attribution bugs fixed in the last two
commits get reintroduced.

Says what is actually true: which field holds the failed hop on each
transport and why the helper reads both, and that the comparison is
origin + path because the web backend's credentials ride in axios'
`params` and a whole-URL comparison therefore reported a redirect for
every ordinary failure.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(portals): scope the guard summary to both processes

The opening line still defined the breaker as an Electron main-process
concern, which contradicted the ownership section below it and is the
part a reader skims to decide whether the document applies to them.

Names both processes, and records that the web backend had the worse
version of the problem first — no request timeout at all — since that is
why the timeouts and the breaker had to land there together.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(portals): declare the shared-interfaces dependency of host-health

The new library's manifest listed only `tslib`, but its emitted JavaScript
does `require('@iptvnator/shared/interfaces')` — the guard builds its
fast-fail message with `buildHostConnectivityFastFailMessage`. Anything
resolving the built artifact from its own manifest would have failed with
MODULE_NOT_FOUND.

The manifest was copied from `shared/logging`, which imports nothing
across libraries and therefore needs nothing beyond `tslib`.
`shared/m3u-utils` is the right precedent: it imports the same library
and declares `"@iptvnator/shared/interfaces": "0.0.1"`.

Verified against the build output rather than by inspection — the emitted
`host-connectivity-guard.js` requires the module, and the generated
`dist/libs/shared/host-health/package.json` now declares it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(portals): bound the PWA connectivity-guard reset

The reset was a bare `fetch` with no timeout, which is the exact failure
this change exists to remove, reintroduced one layer up. Every caller
awaits the reset BEFORE issuing the request it is clearing the way for —
`retryContentInitialization` awaits it first by design — so a backend or
reverse proxy that accepts the POST and then goes quiet would leave
Retry doing nothing at all, for as long as the socket stayed open.

Bound it with an AbortController and a 5 s timer. The abort rejects,
`resetHostConnectivityGuard` swallows it as it already does for any
other failure, and the caller proceeds to its real request — which is
what "best effort" was supposed to mean. The timer is cleared in a
`finally`, and it covers the body read as well as the headers.

Five seconds because this talks to the user's own backend rather than a
provider: it should answer immediately, and a slow one must not hold up
the retry that asked for it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-13 22:27:05 +02:00