perf(portals): fast-fail requests to portal hosts that stopped answering (#1421)

This commit is contained in:
4gray authored and GitHub committed 2026-08-13 07:31:23 +02:00
1 parent 2d7811eb5f
commit e3f72f7dce
50 files changed
+3073 -174

No files matched your search

@@ -0,0 +1,245 @@
# Host Connectivity Guard
Per-host circuit breaker for portal requests in the Electron main process.
## The problem
Every request to an unreachable portal costs its full axios timeout — 30 s for
`XTREAM_REQUEST`, 15 s for `STALKER_REQUEST` (30 s for `create_link`). Browsing a
dead portal's catalog issues dozens of those back to back, which shows up as
30-second spinners and a main-process log full of identical failures. Once a host
has refused to answer twice in a row there is nothing left to learn from waiting
again.
## Where it lives
`apps/electron-backend/src/app/util/host-connectivity-guard.ts` — a pure module
(no Electron imports) wired into both IPC handlers.
The handlers are the choke point that sees _all_ traffic to a host. The
renderer's `executeStalkerRequest` is not: it has four documented bypasses
(auth, endpoint discovery, account info, the row-less stream resolver), and that
is exactly the traffic that hits dead hosts.
**The PWA is deliberately not covered yet.** `apps/web-backend` sets no
per-request timeout at all, so a dead host there hangs on OS-level TCP timeouts
rather than a 15/30 s budget — a timeout-driven breaker would rarely trip. The
guard is written dependency-free so it can move into a `domain:shared-runtime`
library and be shared with `web-backend` when that gap is closed.
## Rules
Being wrong here means refusing to talk to a portal that works, so every rule
errs towards contacting the host:
| | |
| ----------- | ---------------------------------------------------------------------- |
| Trip | 2 consecutive host-level failures within an inclusive 120 s window |
| Open for | 30 s (`OPEN_DURATION_MS`), matching the repo's other cooldowns |
| Half-open | exactly ONE trial request; the rest keep fast-failing until it settles |
| Reset | any HTTP response — 200, 404, even 502 — the host answered |
| Key | `URL.origin` — scheme, host **and** port (see below) |
| Kill switch | `IPTVNATOR_DISABLE_CONNECTIVITY_GUARD=1` (read per call) |
**Host-level failure** means an error with no HTTP response whose code is one of
`ETIMEDOUT`, `ECONNABORTED`, `ENOTFOUND`, `EAI_AGAIN`, `ECONNREFUSED`,
`EHOSTUNREACH`, `ENETUNREACH`. `ECONNRESET` is deliberately excluded: a reset
mid-transfer happens on hosts that are very much alive. Cancelled requests
(`ERR_CANCELED`) and SSRF-policy refusals are `inconclusive` — they say nothing
about reachability and only release the half-open slot.
**A failure is only charged to the endpoint that produced it.** Redirects are
followed hop by hop, each with its own config, so a failure on a later hop
carries that hop's URL in `error.config.url`. Reaching any later hop _proves_ the
guarded endpoint answered — the first hop is always the URL we asked for, and
only a redirect status advances the chain — so it CLEARS the guarded endpoint's
record, exactly like any other response. Merely declining to count it would leave
an earlier direct failure standing, and a single later timeout would then
fast-fail an endpoint that answered in between.
The comparison is against the whole request URL, not just its origin: a
same-origin redirect (`/player_api.php` → `/slow/player_api.php`) proves the
endpoint answered just as much as a cross-origin one, and charging it would
fast-fail every OTHER call to a portal that answers. Both handlers therefore pass
the URL they asked for. It requires positive evidence — anything unparseable or
unknown counts the failure as usual, because guessing "redirect" here would stop
the guard from ever tripping — and a failure that names no URL at all is still
counted. Round-tripping through `URL` is identity for both handlers' URL shapes
(including Stalker's hand-encoded `cmd`), and `requestWithValidatedRedirects`
normalizes hop 1 the same way, so the comparison is exact.
Known gap: the failing hop is not guarded either (it has no token of its own),
so a permanently broken redirect chain keeps costing a full timeout.
**The key is the origin, not the host.** `URL.host` omits a default port, so
`http://panel.example` and `https://panel.example` would share one record —
two genuinely different endpoints, and a panel whose TLS listener is broken
while plain HTTP works is a routine IPTV setup. Sharing state there would let
the dead one fast-fail the working one without ever contacting it.
`URL.origin` also leaves out any `user:pass@` userinfo, so no credential
reaches the key or the log line.
Two more rules exist because of specific failure modes:
- **Siblings are not a streak.** Catalog initialization fans out three category
requests at once; one network hiccup failing all three is one piece of
evidence, not a trip. A failure counts only if its request started at or after
the moment the previous failure was recorded. Timestamps are millisecond
coarse, so this only separates siblings once a request actually took time —
which is precisely the expensive case worth protecting. Only a failure that
was actually counted may trip the threshold: a sibling settling after the open
window elapsed would otherwise start a fresh one off the existing count and
push the half-open trial past the intended cooldown.
- **A reset invalidates reports already in flight.** `reset()` bumps a per-host
epoch instead of deleting the record, and a failure reported under an older
epoch is discarded. Without that, the 30-second stragglers a user was waiting
behind settle right after they press Retry and re-open the breaker underneath
the very retry that cleared it.
A half-open trial that never reports back expires after 45 s, so a leaked token
cannot leave the breaker open forever. That expiry is why the slot has an
identity: a trial can genuinely outlive its window — `requestWithValidatedRedirects`
gives each of up to five redirect hops its own 30 s budget — and once a
replacement has been admitted, the abandoned request's late report must not free
the replacement's slot and let a third request through. `trial: true` alone
cannot tell the two apart, so the token carries the slot id it owns; an
abandoned trial's failure is still counted as ordinary evidence.
## The fast-fail error is a renderer contract
`HostConnectivityGuardError` is a real `Error`: Electron serializes a rejected
plain object to `[object Object]`, which would destroy the renderer's
classification. It carries no `status` property, because
`getStalkerRequestErrorStatus` reads that field first.
The message comes from `buildHostConnectivityFastFailMessage()` in
`libs/shared/interfaces/src/lib/host-connectivity.util.ts`. It names the full
endpoint, scheme included, so a user who imported the same panel over both HTTP
and HTTPS can tell which one was skipped. Its wording is load-bearing — the
Stalker renderer classifies transport failures purely from message text:
| Must not contain | Otherwise |
| ---------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------- |
| `HTTP Error <code>` | reads as "endpoint absent, probe the next candidate"; 404/401/403 also fire lazy portal repair against a host we just declared dead |
| `timed out`, `timeout of Nms`, `ETIMEDOUT` | discovery walks every candidate instead of aborting early |
| `authorization`, `unauthorized`, `access denied`, `invalid token`, `auth failed`, `handshake failed` | fires lazy portal repair (a bare `authorization` matches) |
What remains is the "connection-level failure" slot the renderer already has for
ECONNREFUSED/ENOTFOUND: discovery stops probing and reports the host
unreachable, `shouldAttemptRepair` returns false, and the message reaches error
snackbars verbatim — which is why it reads like a sentence and names the
endpoint.
**The endpoint in that sentence is user data, so the marker outranks the
heuristics.** A portal at `https://authorization.example` would otherwise make
its own fast-fail message match the broad auth phrase set and send an
unreachable host into lazy portal repair. `isStalkerAuthFailureMessage` and
`isStalkerProbeTimeout` therefore both return false for a message
`isHostConnectivityFastFailMessage` recognises, before their phrase matching
runs. (`getStalkerRequestErrorStatus` needs no such guard: `HTTP Error <code>`
contains a space, which a hostname cannot.)
`stalker-portal-discovery.utils.spec.ts` pins all three properties for the bare
and the IPC-wrapped form.
## Exemption: endpoint discovery
`STALKER_REQUEST` accepts `skipConnectionGuard`, set **only** by
`StalkerPortalDiscoveryService.probeContent`. Semantics: bypass the check, never
count failures, **but still report successes.**
Discovery walks several candidate paths on one host and expects most of them to
fail; counting that would let it declare a slow-but-alive portal unreachable.
Reporting successes is equally load-bearing: `confirmFullPortal` runs the full
non-exempt authentication flow against auth-gated candidates, so without it two
hung handshakes could open the breaker mid-discovery and the next authenticate
would fast-fail into the "host unreachable" slot — abandoning a portal that
works. "Success" here means any observed response, including one attached to a
rejection: `validateStatus` lets 4xx through but rejects 5xx with
`error.response` set, and a 5xx proves the origin answered just as well as a
body does. That is why the exempt path reports through
`reportGuardedHostFailure(token, error, { countFailures: false })` rather than
skipping the report.
## Explicit reset
`CONNECTIVITY_GUARD_RESET` (`{ url }`) is handled by
`apps/electron-backend/src/app/events/connectivity-guard.events.ts`. One key
derivation is enough: both `normalizeXtreamServerUrl` and
`buildStalkerRequestUrl` rebuild their request URL from `URL.origin`, so the
origin a request ends up using is always the origin of the URL stored on the
playlist.
**The rule: every user-driven retry or refresh that issues portal requests must
reset the guard before its first request.** The failures that opened the breaker
are usually the very ones the user is retrying, so a reset placed after the
request — or missing — makes the affordance do nothing until the window expires.
Automatic and first-load paths deliberately do NOT reset: only a user action
means "contact this host now", and clearing evidence the guard just collected
would defeat it.
Call sites:
- `retryContentInitialization` (`with-content.feature.ts`) — the Xtream
content-gate Retry button. The reset is the **first** awaited statement, before
the portal status check: a tripped guard fast-fails that check, its
`unavailable` verdict returns early, and a reset placed any later would never
run.
- `StalkerPortalDiscoveryService.discover()` — one site covering import, the Edit
dialog and lazy repair, which also guarantees a freshly edited address never
inherits a refusal recorded for the previous one.
- `retryContentPage` (`with-stalker-content.feature.ts`) — the Stalker grid
tail's append retry.
- `StalkerSearchComponent`'s search-page retry — the search results have their
own append error and retry, separate from the catalog's.
- `StalkerItvCacheService.refresh()` — the Live TV refresh button, on the same
path that already clears the cache's own error cooldown. This also covers
`refreshChannels()` in the live layout.
- Both account-info dialogs' Retry buttons (`AccountInfoComponent.reload()` for
Xtream, `StalkerAccountInfoComponent.reload()` for Stalker). Their automatic
load on open goes through a private `load()` that does not reset.
- The destructive Xtream refresh — before anything is deleted. It removes the
cached catalog and then bootstraps a re-import whose status request an open
guard would fast-fail, leaving the user with no catalog at all until the
cooldown expires. There are **two independent implementations** of this flow
and both need the reset: `PlaylistRefreshActionService.refreshXtream()` and
`RecentPlaylistsComponent.refreshXtreamPlaylist()` (the Workspace sources
page). They are near-duplicates of each other, which is exactly why the second
one was missed first time round.
- `PortalStatusService.checkPortalStatusDetails` when `skipCache` is set — the
user-initiated "Test Connection".
Two retry paths deliberately have no reset: Xtream's `retryAppend()` is a no-op
because in-memory appends cannot fail, and the guard's own half-open trial is not
a user action.
Where a retry clears a UI error flag, that flag is cleared **synchronously**
before awaiting the reset — otherwise the retry branch stays re-enterable and
the next `nearEnd` event fires a second retry.
They all go through `resetHostConnectivityGuard()`
(`libs/services/src/lib/host-connectivity-reset.ts`), which holds the one rule
they share: the reset is best effort, because the guard only ever _delays_ a
request and a failed reset must not block the action that asked for it. In the
PWA the channel is unknown and `sendIpcEvent` no-ops, which is correct —
nothing there records per-host failures yet.
## Interaction with VOD multi-source
No exemption is needed. Multi-source resolution reaches `get_vod_info` over
`XTREAM_REQUEST`, but `VodSourceResolverService.loadVodDetails` already catches
that failure and falls back to its learned container cache or `null`, and
`probeSource` maps `null` to the verdict `unknown` — which
`VodSourceProbeCacheService` deliberately does not cache. A guard fast-fail
therefore degrades to a retryable "could not check", exactly matching the
module's contract that unreachable ≠ contacted-and-refused. The
`STREAM_PROBE_URL` reachability half never rejects and is untouched.
## Tests
- `apps/electron-backend/src/app/util/host-connectivity-guard.spec.ts` — the
state machine, with an injected clock.
- `apps/electron-backend/src/app/events/stalker.events.spec.ts` and
`xtream.events.spec.ts` — trip, fast-fail without contacting axios, reset,
exemption, and the absence of per-request log spam.
- `libs/portal/stalker/data-access/src/lib/stalker-portal-discovery.utils.spec.ts`
and `stalker-portal-repair.service.spec.ts` — the message contract.
+7
View File
@@ -840,6 +840,13 @@ are logged and never retried or escalated.
## Request Transport and `cmd` Encoding
Requests to an unreachable portal are short-circuited by the main process' host
connectivity guard rather than hanging their full 15/30 s timeout again. That
guard's refusal is classified by the same message-text rules discovery uses (it
lands in the "connection-level failure" slot below), and endpoint-discovery
probes are exempt from it via `skipConnectionGuard` — see
[`host-connectivity-guard.md`](./host-connectivity-guard.md).
A real MAG/STB sends `cmd` unencoded: the portal's client JS concatenates raw
`key=value` pairs, the browser URL layer escapes only what a URL cannot carry,
and PHP's `$_GET` applies exactly one form-urldecode. The portal therefore sees