mirror of
https://github.com/4gray/iptvnator.git
synced 2026-10-08 17:06:15 -08:00
perf(portals): fast-fail requests to portal hosts that stopped answering (#1421)
This commit is contained in:
1 parent
2d7811eb5f
commit
e3f72f7dce
50 files changed
+3073
-174
No files matched your search
@@ -0,0 +1,245 @@
|
||||
# Host Connectivity Guard
|
||||
|
||||
Per-host circuit breaker for portal requests in the Electron main process.
|
||||
|
||||
## The problem
|
||||
|
||||
Every request to an unreachable portal costs its full axios timeout — 30 s for
|
||||
`XTREAM_REQUEST`, 15 s for `STALKER_REQUEST` (30 s for `create_link`). Browsing a
|
||||
dead portal's catalog issues dozens of those back to back, which shows up as
|
||||
30-second spinners and a main-process log full of identical failures. Once a host
|
||||
has refused to answer twice in a row there is nothing left to learn from waiting
|
||||
again.
|
||||
|
||||
## Where it lives
|
||||
|
||||
`apps/electron-backend/src/app/util/host-connectivity-guard.ts` — a pure module
|
||||
(no Electron imports) wired into both IPC handlers.
|
||||
|
||||
The handlers are the choke point that sees _all_ traffic to a host. The
|
||||
renderer's `executeStalkerRequest` is not: it has four documented bypasses
|
||||
(auth, endpoint discovery, account info, the row-less stream resolver), and that
|
||||
is exactly the traffic that hits dead hosts.
|
||||
|
||||
**The PWA is deliberately not covered yet.** `apps/web-backend` sets no
|
||||
per-request timeout at all, so a dead host there hangs on OS-level TCP timeouts
|
||||
rather than a 15/30 s budget — a timeout-driven breaker would rarely trip. The
|
||||
guard is written dependency-free so it can move into a `domain:shared-runtime`
|
||||
library and be shared with `web-backend` when that gap is closed.
|
||||
|
||||
## Rules
|
||||
|
||||
Being wrong here means refusing to talk to a portal that works, so every rule
|
||||
errs towards contacting the host:
|
||||
|
||||
| | |
|
||||
| ----------- | ---------------------------------------------------------------------- |
|
||||
| Trip | 2 consecutive host-level failures within an inclusive 120 s window |
|
||||
| Open for | 30 s (`OPEN_DURATION_MS`), matching the repo's other cooldowns |
|
||||
| Half-open | exactly ONE trial request; the rest keep fast-failing until it settles |
|
||||
| Reset | any HTTP response — 200, 404, even 502 — the host answered |
|
||||
| Key | `URL.origin` — scheme, host **and** port (see below) |
|
||||
| Kill switch | `IPTVNATOR_DISABLE_CONNECTIVITY_GUARD=1` (read per call) |
|
||||
|
||||
**Host-level failure** means an error with no HTTP response whose code is one of
|
||||
`ETIMEDOUT`, `ECONNABORTED`, `ENOTFOUND`, `EAI_AGAIN`, `ECONNREFUSED`,
|
||||
`EHOSTUNREACH`, `ENETUNREACH`. `ECONNRESET` is deliberately excluded: a reset
|
||||
mid-transfer happens on hosts that are very much alive. Cancelled requests
|
||||
(`ERR_CANCELED`) and SSRF-policy refusals are `inconclusive` — they say nothing
|
||||
about reachability and only release the half-open slot.
|
||||
|
||||
**A failure is only charged to the endpoint that produced it.** Redirects are
|
||||
followed hop by hop, each with its own config, so a failure on a later hop
|
||||
carries that hop's URL in `error.config.url`. Reaching any later hop _proves_ the
|
||||
guarded endpoint answered — the first hop is always the URL we asked for, and
|
||||
only a redirect status advances the chain — so it CLEARS the guarded endpoint's
|
||||
record, exactly like any other response. Merely declining to count it would leave
|
||||
an earlier direct failure standing, and a single later timeout would then
|
||||
fast-fail an endpoint that answered in between.
|
||||
|
||||
The comparison is against the whole request URL, not just its origin: a
|
||||
same-origin redirect (`/player_api.php` → `/slow/player_api.php`) proves the
|
||||
endpoint answered just as much as a cross-origin one, and charging it would
|
||||
fast-fail every OTHER call to a portal that answers. Both handlers therefore pass
|
||||
the URL they asked for. It requires positive evidence — anything unparseable or
|
||||
unknown counts the failure as usual, because guessing "redirect" here would stop
|
||||
the guard from ever tripping — and a failure that names no URL at all is still
|
||||
counted. Round-tripping through `URL` is identity for both handlers' URL shapes
|
||||
(including Stalker's hand-encoded `cmd`), and `requestWithValidatedRedirects`
|
||||
normalizes hop 1 the same way, so the comparison is exact.
|
||||
|
||||
Known gap: the failing hop is not guarded either (it has no token of its own),
|
||||
so a permanently broken redirect chain keeps costing a full timeout.
|
||||
|
||||
**The key is the origin, not the host.** `URL.host` omits a default port, so
|
||||
`http://panel.example` and `https://panel.example` would share one record —
|
||||
two genuinely different endpoints, and a panel whose TLS listener is broken
|
||||
while plain HTTP works is a routine IPTV setup. Sharing state there would let
|
||||
the dead one fast-fail the working one without ever contacting it.
|
||||
`URL.origin` also leaves out any `user:pass@` userinfo, so no credential
|
||||
reaches the key or the log line.
|
||||
|
||||
Two more rules exist because of specific failure modes:
|
||||
|
||||
- **Siblings are not a streak.** Catalog initialization fans out three category
|
||||
requests at once; one network hiccup failing all three is one piece of
|
||||
evidence, not a trip. A failure counts only if its request started at or after
|
||||
the moment the previous failure was recorded. Timestamps are millisecond
|
||||
coarse, so this only separates siblings once a request actually took time —
|
||||
which is precisely the expensive case worth protecting. Only a failure that
|
||||
was actually counted may trip the threshold: a sibling settling after the open
|
||||
window elapsed would otherwise start a fresh one off the existing count and
|
||||
push the half-open trial past the intended cooldown.
|
||||
- **A reset invalidates reports already in flight.** `reset()` bumps a per-host
|
||||
epoch instead of deleting the record, and a failure reported under an older
|
||||
epoch is discarded. Without that, the 30-second stragglers a user was waiting
|
||||
behind settle right after they press Retry and re-open the breaker underneath
|
||||
the very retry that cleared it.
|
||||
|
||||
A half-open trial that never reports back expires after 45 s, so a leaked token
|
||||
cannot leave the breaker open forever. That expiry is why the slot has an
|
||||
identity: a trial can genuinely outlive its window — `requestWithValidatedRedirects`
|
||||
gives each of up to five redirect hops its own 30 s budget — and once a
|
||||
replacement has been admitted, the abandoned request's late report must not free
|
||||
the replacement's slot and let a third request through. `trial: true` alone
|
||||
cannot tell the two apart, so the token carries the slot id it owns; an
|
||||
abandoned trial's failure is still counted as ordinary evidence.
|
||||
|
||||
## The fast-fail error is a renderer contract
|
||||
|
||||
`HostConnectivityGuardError` is a real `Error`: Electron serializes a rejected
|
||||
plain object to `[object Object]`, which would destroy the renderer's
|
||||
classification. It carries no `status` property, because
|
||||
`getStalkerRequestErrorStatus` reads that field first.
|
||||
|
||||
The message comes from `buildHostConnectivityFastFailMessage()` in
|
||||
`libs/shared/interfaces/src/lib/host-connectivity.util.ts`. It names the full
|
||||
endpoint, scheme included, so a user who imported the same panel over both HTTP
|
||||
and HTTPS can tell which one was skipped. Its wording is load-bearing — the
|
||||
Stalker renderer classifies transport failures purely from message text:
|
||||
|
||||
| Must not contain | Otherwise |
|
||||
| ---------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `HTTP Error <code>` | reads as "endpoint absent, probe the next candidate"; 404/401/403 also fire lazy portal repair against a host we just declared dead |
|
||||
| `timed out`, `timeout of Nms`, `ETIMEDOUT` | discovery walks every candidate instead of aborting early |
|
||||
| `authorization`, `unauthorized`, `access denied`, `invalid token`, `auth failed`, `handshake failed` | fires lazy portal repair (a bare `authorization` matches) |
|
||||
|
||||
What remains is the "connection-level failure" slot the renderer already has for
|
||||
ECONNREFUSED/ENOTFOUND: discovery stops probing and reports the host
|
||||
unreachable, `shouldAttemptRepair` returns false, and the message reaches error
|
||||
snackbars verbatim — which is why it reads like a sentence and names the
|
||||
endpoint.
|
||||
|
||||
**The endpoint in that sentence is user data, so the marker outranks the
|
||||
heuristics.** A portal at `https://authorization.example` would otherwise make
|
||||
its own fast-fail message match the broad auth phrase set and send an
|
||||
unreachable host into lazy portal repair. `isStalkerAuthFailureMessage` and
|
||||
`isStalkerProbeTimeout` therefore both return false for a message
|
||||
`isHostConnectivityFastFailMessage` recognises, before their phrase matching
|
||||
runs. (`getStalkerRequestErrorStatus` needs no such guard: `HTTP Error <code>`
|
||||
contains a space, which a hostname cannot.)
|
||||
`stalker-portal-discovery.utils.spec.ts` pins all three properties for the bare
|
||||
and the IPC-wrapped form.
|
||||
|
||||
## Exemption: endpoint discovery
|
||||
|
||||
`STALKER_REQUEST` accepts `skipConnectionGuard`, set **only** by
|
||||
`StalkerPortalDiscoveryService.probeContent`. Semantics: bypass the check, never
|
||||
count failures, **but still report successes.**
|
||||
|
||||
Discovery walks several candidate paths on one host and expects most of them to
|
||||
fail; counting that would let it declare a slow-but-alive portal unreachable.
|
||||
Reporting successes is equally load-bearing: `confirmFullPortal` runs the full
|
||||
non-exempt authentication flow against auth-gated candidates, so without it two
|
||||
hung handshakes could open the breaker mid-discovery and the next authenticate
|
||||
would fast-fail into the "host unreachable" slot — abandoning a portal that
|
||||
works. "Success" here means any observed response, including one attached to a
|
||||
rejection: `validateStatus` lets 4xx through but rejects 5xx with
|
||||
`error.response` set, and a 5xx proves the origin answered just as well as a
|
||||
body does. That is why the exempt path reports through
|
||||
`reportGuardedHostFailure(token, error, { countFailures: false })` rather than
|
||||
skipping the report.
|
||||
|
||||
## Explicit reset
|
||||
|
||||
`CONNECTIVITY_GUARD_RESET` (`{ url }`) is handled by
|
||||
`apps/electron-backend/src/app/events/connectivity-guard.events.ts`. One key
|
||||
derivation is enough: both `normalizeXtreamServerUrl` and
|
||||
`buildStalkerRequestUrl` rebuild their request URL from `URL.origin`, so the
|
||||
origin a request ends up using is always the origin of the URL stored on the
|
||||
playlist.
|
||||
|
||||
**The rule: every user-driven retry or refresh that issues portal requests must
|
||||
reset the guard before its first request.** The failures that opened the breaker
|
||||
are usually the very ones the user is retrying, so a reset placed after the
|
||||
request — or missing — makes the affordance do nothing until the window expires.
|
||||
Automatic and first-load paths deliberately do NOT reset: only a user action
|
||||
means "contact this host now", and clearing evidence the guard just collected
|
||||
would defeat it.
|
||||
|
||||
Call sites:
|
||||
|
||||
- `retryContentInitialization` (`with-content.feature.ts`) — the Xtream
|
||||
content-gate Retry button. The reset is the **first** awaited statement, before
|
||||
the portal status check: a tripped guard fast-fails that check, its
|
||||
`unavailable` verdict returns early, and a reset placed any later would never
|
||||
run.
|
||||
- `StalkerPortalDiscoveryService.discover()` — one site covering import, the Edit
|
||||
dialog and lazy repair, which also guarantees a freshly edited address never
|
||||
inherits a refusal recorded for the previous one.
|
||||
- `retryContentPage` (`with-stalker-content.feature.ts`) — the Stalker grid
|
||||
tail's append retry.
|
||||
- `StalkerSearchComponent`'s search-page retry — the search results have their
|
||||
own append error and retry, separate from the catalog's.
|
||||
- `StalkerItvCacheService.refresh()` — the Live TV refresh button, on the same
|
||||
path that already clears the cache's own error cooldown. This also covers
|
||||
`refreshChannels()` in the live layout.
|
||||
- Both account-info dialogs' Retry buttons (`AccountInfoComponent.reload()` for
|
||||
Xtream, `StalkerAccountInfoComponent.reload()` for Stalker). Their automatic
|
||||
load on open goes through a private `load()` that does not reset.
|
||||
- The destructive Xtream refresh — before anything is deleted. It removes the
|
||||
cached catalog and then bootstraps a re-import whose status request an open
|
||||
guard would fast-fail, leaving the user with no catalog at all until the
|
||||
cooldown expires. There are **two independent implementations** of this flow
|
||||
and both need the reset: `PlaylistRefreshActionService.refreshXtream()` and
|
||||
`RecentPlaylistsComponent.refreshXtreamPlaylist()` (the Workspace sources
|
||||
page). They are near-duplicates of each other, which is exactly why the second
|
||||
one was missed first time round.
|
||||
- `PortalStatusService.checkPortalStatusDetails` when `skipCache` is set — the
|
||||
user-initiated "Test Connection".
|
||||
|
||||
Two retry paths deliberately have no reset: Xtream's `retryAppend()` is a no-op
|
||||
because in-memory appends cannot fail, and the guard's own half-open trial is not
|
||||
a user action.
|
||||
|
||||
Where a retry clears a UI error flag, that flag is cleared **synchronously**
|
||||
before awaiting the reset — otherwise the retry branch stays re-enterable and
|
||||
the next `nearEnd` event fires a second retry.
|
||||
|
||||
They all go through `resetHostConnectivityGuard()`
|
||||
(`libs/services/src/lib/host-connectivity-reset.ts`), which holds the one rule
|
||||
they share: the reset is best effort, because the guard only ever _delays_ a
|
||||
request and a failed reset must not block the action that asked for it. In the
|
||||
PWA the channel is unknown and `sendIpcEvent` no-ops, which is correct —
|
||||
nothing there records per-host failures yet.
|
||||
|
||||
## Interaction with VOD multi-source
|
||||
|
||||
No exemption is needed. Multi-source resolution reaches `get_vod_info` over
|
||||
`XTREAM_REQUEST`, but `VodSourceResolverService.loadVodDetails` already catches
|
||||
that failure and falls back to its learned container cache or `null`, and
|
||||
`probeSource` maps `null` to the verdict `unknown` — which
|
||||
`VodSourceProbeCacheService` deliberately does not cache. A guard fast-fail
|
||||
therefore degrades to a retryable "could not check", exactly matching the
|
||||
module's contract that unreachable ≠ contacted-and-refused. The
|
||||
`STREAM_PROBE_URL` reachability half never rejects and is untouched.
|
||||
|
||||
## Tests
|
||||
|
||||
- `apps/electron-backend/src/app/util/host-connectivity-guard.spec.ts` — the
|
||||
state machine, with an injected clock.
|
||||
- `apps/electron-backend/src/app/events/stalker.events.spec.ts` and
|
||||
`xtream.events.spec.ts` — trip, fast-fail without contacting axios, reset,
|
||||
exemption, and the absence of per-request log spam.
|
||||
- `libs/portal/stalker/data-access/src/lib/stalker-portal-discovery.utils.spec.ts`
|
||||
and `stalker-portal-repair.service.spec.ts` — the message contract.
|
||||
@@ -840,6 +840,13 @@ are logged and never retried or escalated.
|
||||
|
||||
## Request Transport and `cmd` Encoding
|
||||
|
||||
Requests to an unreachable portal are short-circuited by the main process' host
|
||||
connectivity guard rather than hanging their full 15/30 s timeout again. That
|
||||
guard's refusal is classified by the same message-text rules discovery uses (it
|
||||
lands in the "connection-level failure" slot below), and endpoint-discovery
|
||||
probes are exempt from it via `skipConnectionGuard` — see
|
||||
[`host-connectivity-guard.md`](./host-connectivity-guard.md).
|
||||
|
||||
A real MAG/STB sends `cmd` unencoded: the portal's client JS concatenates raw
|
||||
`key=value` pairs, the browser URL layer escapes only what a URL cannot carry,
|
||||
and PHP's `$_GET` applies exactly one form-urldecode. The portal therefore sees
|
||||
|
||||
Reference in new issue
Block a user