Files
iptvnator/docs/architecture/host-connectivity-guard.md
T
4grayandClaude Fable 5.1 01c423ac43 fix(portals): stop treating a slow panel as a dead host; IPv4 fallback budget in Electron (#1621)
* fix(portals): apply the IPv6->IPv4 fallback budget in the Electron process

The 2500 ms happy-eyeballs attempt timeout from #1404 only ever ran in the
web backend. The Electron main process and its playlist-refresh and EPG
workers kept Node's 250 ms default, so a dual-stack panel hostname behind a
VPN or a slow link failed every connection attempt in a row and tripped the
host connectivity guard. The module now lives in `@iptvnator/shared/host-health`
and every Node isolate that opens connections applies it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(portals): stop treating a slow panel as a dead one in the host guard

axios raises the same ECONNABORTED whether the SYN went unanswered or the
panel accepted the connection and then thought for longer than the request
budget. Two such timeouts opened the breaker and every request to the panel
was refused for 30 s with "portal is not responding" — the shape behind the
"connection keeps dropping" reports on 0.23 and nightly.

Both transports now report whether the TCP connection was established
(`onConnect`: Electron through a per-request observed agent instead of the
shared keep-alive globalAgent, the web backend through the transport that owns
the ClientRequest), and `classifyHostRequestFailure(error, { connected })`
downgrades a host-level code observed after the handshake to inconclusive.
Redirect attribution keeps precedence. A host that never accepts the
connection trips the guard exactly as before.

The Xtream mock gains a `silent:silent` scenario whose detail actions accept
and never answer, plus a real-socket regression spec for the guard.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(portals): let an accepted connection clear the host-failure streak

Review finding: an unanswered SYN, then an accepted-but-slow timeout, then
another unanswered SYN still reached the two-failure threshold, because the
middle request was merely not counted. An accepted TCP connection is the
reachability the guard measures, so it now reads as `responded` and clears
the streak like an HTTP response would. Regression coverage for the mixed
sequence on one flapping loopback origin (Electron) and through the proxy
route (web backend).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(portals): credit an accepted connection when it happens, not when the request settles

Review findings. A request that connected and then hung for 30 s cleared,
on its eventual timeout, the failures later requests had recorded while it
waited — reopening a host that had just died on evidence older than theirs.
The connect hook now reports the connection the moment it fires through a
new `HostConnectivityGuard.reportConnected`, which clears the failure streak
but closes no open or half-open breaker (the trial keeps its slot until it
settles), and the settled timeout is inconclusive.

Electron also skips the socket observer while an environment proxy
(`http_proxy` / `https_proxy` / `all_proxy`) applies to the request: through
a proxy the socket connects to the proxy, whose handshake proves nothing
about the portal, so those requests keep the pre-observer behaviour.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(portals): decide the proxy exemption with axios' own resolution

Review findings. The hand-rolled environment check ignored `no_proxy`, so a
LAN portal exempted from the proxy lost its connect observer and slow
requests to it still tripped the breaker; it also read the variables with
`??`, letting an empty lowercase one mask a populated uppercase one that
axios would honour. The decision now calls `proxy-from-env`'s
`getProxyForUrl`, the same pinned package axios' http adapter uses,
declared as a direct dependency so the packaged app carries it.

The validated-axios spec clears and restores every proxy variable around each
case, so a runner that exports a proxy cannot change what the cases prove.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 14:50:50 +02:00

428 lines
26 KiB
Markdown

# Host Connectivity Guard
Per-host circuit breaker for portal requests, in both processes that make them:
the Electron main process and the self-hosted web backend.
## The problem
Every request to an unreachable portal costs its full axios timeout — 30 s for
`XTREAM_REQUEST`, 15 s for `STALKER_REQUEST` (30 s for `create_link`), and the
same budgets on the web backend's `/xtream` and `/stalker` routes. Browsing a
dead portal's catalog issues dozens of those back to back, which shows up as
30-second spinners and a log full of identical failures. Once a host has refused
to answer twice in a row there is nothing left to learn from waiting again.
The web backend had a worse version of the same problem first: its proxy routes
passed no `timeout` at all, so a provider that accepted a connection and then
went silent held the request until the OS gave up on the TCP connection. A
breaker is only useful once _not answering_ is bounded, which is why the timeouts
and the breaker landed there together.
## Where it lives
`libs/shared/host-health` (`@iptvnator/shared/host-health`, tagged
`scope:shared` / `domain:shared-runtime` / `type:util`) — the breaker class, the
failure classification and the redirect-attribution helpers, with no transport,
logger or process singleton of its own. The owning app supplies the clock and
decides how many guards exist.
Each runtime owns its instance:
- **Electron** — `apps/electron-backend/src/app/util/host-connectivity-guard.ts`
holds one guard for the whole process, wired into both IPC handlers. The
handlers are the choke point that sees _all_ traffic to a host. The renderer's
`executeStalkerRequest` is not: it has four documented bypasses (auth,
endpoint discovery, account info, the row-less stream resolver), and that is
exactly the traffic that hits dead hosts.
- **PWA** — `apps/web-backend/src/app/host-guard.ts` guards the `/xtream` and
`/stalker` proxy routes. The guard instance is injected through
`WebBackendAppOptions.hostGuard`, alongside `now` and `guid`, so specs drive
it with a clock they own.
### Request timeouts are the precondition
A breaker keyed on "the endpoint did not answer" is only useful once _not
answering_ is bounded. `apps/web-backend` originally passed no `timeout` at all,
so a silent host hung on OS-level TCP timeouts. Both runtimes now use the same
budgets: Xtream 30 s, Stalker 15 s (30 s for `create_link`, which mints a stream
URL before answering), playlist and XMLTV downloads 30 s.
Those numbers are safe for large downloads. On axios 1.20's native HTTP
transport (`maxRedirects: 0`) each hop's `timeout` is **not** a wall-clock deadline for the
whole response: it bounds the time to response headers and then continues as the
socket's inactivity timeout for the body. A multi-megabyte XMLTV file that keeps
delivering bytes is never cut off mid-transfer — only a stalled one is. The web
backend explicitly restores socket inactivity handling while reading the final
stream response, because axios stream mode stops its timeout handler at headers.
Truncated, malformed and timed-out final bodies retain their received response
status for reachability classification; cancellation retains its own semantics.
### Scope: portal calls only
Both runtimes guard the portal API paths and nothing else. Playlist and XMLTV
downloads (`/parse`, `/parse-xml`, and their Electron equivalents) get the
timeouts but not the breaker, deliberately:
- The problem being solved is a catalog fan-out — dozens of requests to one
endpoint back to back. A download is a single request.
- A download is usually the direct result of the user asking for it (the
add-playlist dialog sends `PLAYLIST_PARSE_BY_URL`). Refusing an immediate
retry is a regression, not a protection, and there is no natural reset site on
that path the way portal Retry has one.
## Desktop preference and account feedback
Settings > General > Portal connections offers **Pause requests to unavailable
portals**, enabled by default. `Settings.portalConnectivityGuard` is default-on
for missing or non-boolean legacy values; only explicit `false` opts out. The
form stages edits until Save and persists the choice through the normal settings
store. `SETTINGS_UPDATE` mirrors it to Electron's `PORTAL_CONNECTIVITY_GUARD`
config key and applies it immediately. `SettingsEvents.bootstrapSettingsEvents`
loads that mirror before the renderer is loaded, so a saved opt-out already
applies to the first portal request after restarting.
The preference controls both Xtream and Stalker. A real preference transition
replaces the Electron guard and its weak set of admitted request tokens, clearing
existing cooldowns and ignoring completions from the old generation. Disabled
requests return no token and cannot contribute failures when the user re-enables
the guard. Saving an unchanged value preserves existing evidence.
`IPTVNATOR_DISABLE_CONNECTIVITY_GUARD=1` (or `true`) still overrides an enabled
preference. The UI is capability-gated by `supportsPortalConnectivityGuard`
(`updateSettings` plus `resetHostConnectivityGuard`); the PWA has no client-side
switch for its shared server-wide guard.
Both account-info dialogs classify the existing cross-process guard message and
show a localized **Requests temporarily paused** explanation with **Retry now**.
Stalker keeps cached account data visible and offers the same retry beside the
paused refresh notice. Retry resets before requesting again; actual network
failures retain the generic unavailable state. These notices describe the last
request outcome, not a live countdown. Same-millisecond sibling counting is
handled by monotonic admission ids (#1438).
## Rules
Being wrong here means refusing to talk to a portal that works, so every rule
errs towards contacting the host:
| | |
| ----------- | ---------------------------------------------------------------------- |
| Trip | 2 consecutive host-level failures within an inclusive 120 s window |
| Open for | 30 s (`OPEN_DURATION_MS`), matching the repo's other cooldowns |
| Half-open | exactly ONE trial request; the rest keep fast-failing until it settles |
| Reset | any HTTP response — 200, 404, even 502; an accepted TCP connection clears the streak but not an open breaker |
| Key | `URL.origin` — scheme, host **and** port (see below) |
| Kill switch | `IPTVNATOR_DISABLE_CONNECTIVITY_GUARD=1` (read per call) |
**Host-level failure** means an error with no HTTP response whose code is one of
`ETIMEDOUT`, `ECONNABORTED`, `ENOTFOUND`, `EAI_AGAIN`, `ECONNREFUSED`,
`EHOSTUNREACH`, `ENETUNREACH` — observed **before the TCP connection was
established**. `ECONNRESET` is deliberately excluded: a reset mid-transfer
happens on hosts that are very much alive. Cancelled requests (`ERR_CANCELED`)
and SSRF-policy refusals are `inconclusive` — they say nothing about
reachability and only release the half-open slot.
**A timeout after the handshake is a slow host, not a dead one.** axios raises
the same `ECONNABORTED` whether the SYN went unanswered or the panel accepted
the connection and then thought for longer than the request budget (a heavy
`get_vod_info` on a busy home server). Only the former is what the guard
exists for; tripping on the latter turned "slow" into thirty seconds of
"portal is not responding" for every request, which is how users described
it. Each transport therefore reports whether the connection was established —
Electron through `ValidatedAxiosRequestConfig.onConnect` (the request gets its
own agent, observed via `observeAgentSocketConnections`, instead of the shared
keep-alive `globalAgent`; a pooled socket that is already connected counts as
connected), the web backend through `WebBackendHttpGetOptions.onConnect`,
honoured by `ProviderAxiosTransport`, which owns the `ClientRequest` — and
the accepted connection is credited THE MOMENT IT HAPPENS
(`reportGuardedHostConnected` / `reportProviderRequestConnected` from the
connect hook call `HostConnectivityGuard.reportConnected`, which clears the
failure streak but deliberately closes no open or half-open breaker: whether a
host that accepts a connection also answers is exactly what the trial exists
to find out, so the trial keeps its slot until it settles), and
`classifyHostRequestFailure(error, { connected })` then reads the eventual
host-level code as `inconclusive`. Two things hang on that
ordering. Clearing at connect time is what keeps an unanswered SYN, an
accepted-but-slow request and another unanswered SYN from adding up to a
trip, since the middle request proved the host alive in between. Crediting it
at connect time rather than when the timeout settles is what keeps a request
that connected and then hung for 30 s from reopening a breaker that later
requests opened in the meantime — by then its evidence is older than theirs.
Such a request still costs its full timeout; the guard only stops charging it
to the host. Electron does not install the observer while an environment
proxy applies to the request — decided by the very resolution axios performs
(`proxy-from-env`, the same pinned package: `<protocol>_proxy` / `all_proxy`
in either case, `no_proxy` exemptions honoured, so a LAN portal listed there
keeps its observer): through a proxy the socket connects to the proxy, whose
handshake proves nothing about the portal, so those requests keep reporting
their timeouts as host-level exactly as before. Regression coverage:
`apps/electron-backend/src/app/util/host-connectivity-guard.slow-host.spec.ts`
(real loopback sockets) and the Xtream mock's `silent` scenario, whose
`get_vod_info` / `get_series_info` accept the connection and never answer.
**A failure is only charged to the endpoint that produced it.** Reaching any
later hop _proves_ the guarded endpoint answered — the first hop is always the
URL we asked for, and only a redirect status advances the chain — so a failure
there CLEARS the guarded endpoint's record, exactly like any other response.
Merely declining to count it would leave an earlier direct failure standing, and
a single later timeout would then fast-fail an endpoint that answered in between.
Every caller passes the URL it asked for as the baseline.
**Web backend records redirect evidence explicitly.** Its `ValidatedHttpClient`
disables automatic redirects and holds one admission for the entire chain.
`ProviderRequestError.initialResponded` becomes true only after receiving an
actual redirect response. A later transport or DNS failure, URL policy refusal,
invalid location, cycle or exhausted budget therefore clears the initial
endpoint's failure streak when the chain settles. This also handles redirects
that change only the query. It does not release the half-open slot early: the
route retains ownership through validation, all hops and the final body, and
releases in `finally` even when reporting throws. Before any redirect, initial
DNS failures carry their internal resolver cause into normal classification;
policy refusals remain inconclusive. Error bodies never serialize that cause.
**Electron retains URL-based attribution.** `failedRequestUrlOf` reads
`error.request._currentUrl` first (for a following transport) and falls back to
`error.config.url` (for a native per-hop transport). The comparison is origin +
path, not the query: axios `params` can append credentials absent from the
caller's baseline. Unknown/unparseable URLs conservatively count as ordinary
failures, and query-only redirects cannot be distinguished by this fallback.
Wrapped web-backend requests use their explicit chain evidence and never infer
redirects from the transport's fixed logical hostname. The shared helper and
Electron behavior are unchanged by the web-backend fix.
Known gap: the failing hop is not guarded either (it has no token of its own),
so a permanently broken redirect chain keeps costing a full timeout.
**The key is the origin, not the host.** `URL.host` omits a default port, so
`http://panel.example` and `https://panel.example` would share one record —
two genuinely different endpoints, and a panel whose TLS listener is broken
while plain HTTP works is a routine IPTV setup. Sharing state there would let
the dead one fast-fail the working one without ever contacting it.
`URL.origin` also leaves out any `user:pass@` userinfo, so no credential
reaches the key or the log line.
Two more rules exist because of specific failure modes:
- **Siblings are not a streak.** Catalog initialization fans out three category
requests at once; one network hiccup failing all three is one piece of
evidence, not a trip. Each admitted request receives a guard-wide monotonic
admission id. A counted failure saves the highest id admitted so far in that
endpoint's failure record, not just the id of the request that failed. During
that streak, only a request admitted beyond this boundary can add another
failure. This distinguishes
siblings from later attempts even if admission and completion all happen in
the same millisecond (#1438), regardless of completion order. Admission ids
are never reused when an endpoint's state is evicted. A recreated record has
no failure streak, so the first old in-flight failure can still count once;
its remaining siblings cannot add further links to that streak. Timestamps
still determine streak expiry and cooldown. Only a failure that
was actually counted may trip the threshold: a sibling settling after the open
window elapsed would otherwise start a fresh one off the existing count and
push the half-open trial past the intended cooldown.
- **A reset invalidates reports already in flight.** `reset()` bumps a per-host
epoch instead of deleting the record, and a failure reported under an older
epoch is discarded. Without that, the 30-second stragglers a user was waiting
behind settle right after they press Retry and re-open the breaker underneath
the very retry that cleared it.
### Trial ownership follows the request lifetime
A half-open slot has no wall-clock expiry (#1439). A redirect chain gives each
hop a separate timeout, and a body that keeps delivering bytes can exceed any
fixed deadline without reaching its inactivity timeout. Neither may admit a
second trial while the first request remains pending.
The four request owners (Electron Xtream/Stalker IPC and web-backend
`/xtream`/`/stalker`) acquire inside `try`, await the whole transport operation,
report its outcome, and unconditionally release any remaining token in `finally`.
`reportInconclusive` is the idempotent cleanup operation: it releases only the
matching `trialId` and epoch and records no reachability evidence. This is the
leak backstop even when debug logging, classification, or response formatting
throws before an outcome report. Cleanup also runs while the environment kill
switch is on, so switching it back off cannot revive a completed trial's slot.
Desktop preference transitions invalidate the old guard and tokens as before.
Cancellation releases after the awaited transport rejects, when it actually
settles, rather than merely when an abort is requested. Xtream's AbortSignal
continues through all Electron redirect hops. Stalker has no IPC cancellation
signal, and closing a PWA client connection does not cancel the backend's outbound
request; those slots remain owned until that transport settles. The transport's
existing timeout handles inactivity. No heartbeat, polling timer, or larger
trial deadline is needed. A future caller must preserve the `try`/`finally`
contract; a custom transport that never settles needs cancellation at its own
layer and must not be silently overlapped by a second trial.
The slot still has an identity: delayed cleanup from a released owner must not
free its replacement. A late failure still follows the ordinary streak and epoch
rules; a real HTTP response still proves reachability and clears the record.
Explicit reset, discovery success, preference transitions and bounded state
eviction retain their existing rules for forgetting evidence.
## The fast-fail error is a renderer contract
`HostConnectivityGuardError` is a real `Error`: Electron serializes a rejected
plain object to `[object Object]`, which would destroy the renderer's
classification. It carries no `status` property, because
`getStalkerRequestErrorStatus` reads that field first.
The message comes from `buildHostConnectivityFastFailMessage()` in
`libs/shared/interfaces/src/lib/host-connectivity.util.ts`. It names the full
endpoint, scheme included, so a user who imported the same panel over both HTTP
and HTTPS can tell which one was skipped. Its wording is load-bearing — the
Stalker renderer classifies transport failures purely from message text:
| Must not contain | Otherwise |
| ---------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------- |
| `HTTP Error <code>` | reads as "endpoint absent, probe the next candidate"; 404/401/403 also fire lazy portal repair against a host we just declared dead |
| `timed out`, `timeout of Nms`, `ETIMEDOUT` | discovery walks every candidate instead of aborting early |
| `authorization`, `unauthorized`, `access denied`, `invalid token`, `auth failed`, `handshake failed` | fires lazy portal repair (a bare `authorization` matches) |
What remains is the "connection-level failure" slot the renderer already has for
ECONNREFUSED/ENOTFOUND: discovery stops probing and reports the host
unreachable, `shouldAttemptRepair` returns false, and the message reaches error
snackbars verbatim — which is why it reads like a sentence and names the
endpoint.
**The endpoint in that sentence is user data, so the marker outranks the
heuristics.** A portal at `https://authorization.example` would otherwise make
its own fast-fail message match the broad auth phrase set and send an
unreachable host into lazy portal repair. `isStalkerAuthFailureMessage` and
`isStalkerProbeTimeout` therefore both return false for a message
`isHostConnectivityFastFailMessage` recognises, before their phrase matching
runs. (`getStalkerRequestErrorStatus` needs no such guard: `HTTP Error <code>`
contains a space, which a hostname cannot.)
`stalker-portal-discovery.utils.spec.ts` pins all three properties for the bare
and the IPC-wrapped form.
## Exemption: endpoint discovery
`STALKER_REQUEST` accepts `skipConnectionGuard`, set **only** by
`StalkerPortalDiscoveryService.probeContent`. Semantics: bypass the check, never
count failures, **but still report successes.**
Discovery walks several candidate paths on one host and expects most of them to
fail; counting that would let it declare a slow-but-alive portal unreachable.
Reporting successes is equally load-bearing: `confirmFullPortal` runs the full
non-exempt authentication flow against auth-gated candidates, so without it two
hung handshakes could open the breaker mid-discovery and the next authenticate
would fast-fail into the "host unreachable" slot — abandoning a portal that
works. "Success" here means any observed response, including one attached to a
rejection: `validateStatus` lets 4xx through but rejects 5xx with
`error.response` set, and a 5xx proves the origin answered just as well as a
body does. That is why the exempt path reports through
`reportGuardedHostFailure(token, error, { countFailures: false })` rather than
skipping the report.
The flag has to survive the PWA transport too, or discovery is exempt on the
desktop and policed on the web. `PwaService.forwardStalkerRequest` forwards it
as a `/stalker` control param and the route applies the same semantics — and,
like `macAddress`/`token`/`serialNumber`, it is stripped from the query
forwarded to the portal, because it is our control flag and not protocol
content.
## Explicit reset
`CONNECTIVITY_GUARD_RESET` (`{ url }`) is handled by
`apps/electron-backend/src/app/events/connectivity-guard.events.ts`. One key
derivation is enough: both `normalizeXtreamServerUrl` and
`buildStalkerRequestUrl` rebuild their request URL from `URL.origin`, so the
origin a request ends up using is always the origin of the URL stored on the
playlist.
**The rule: every user-driven retry or refresh that issues portal requests must
reset the guard before its first request.** The failures that opened the breaker
are usually the very ones the user is retrying, so a reset placed after the
request — or missing — makes the affordance do nothing until the window expires.
Automatic and first-load paths deliberately do NOT reset: only a user action
means "contact this host now", and clearing evidence the guard just collected
would defeat it.
Call sites:
- `retryContentInitialization` (`with-content.feature.ts`) — the Xtream
content-gate Retry button. The reset is the **first** awaited statement, before
the portal status check: a tripped guard fast-fails that check, its
`unavailable` verdict returns early, and a reset placed any later would never
run.
- `StalkerPortalDiscoveryService.discover()` — one site covering import, the Edit
dialog and lazy repair, which also guarantees a freshly edited address never
inherits a refusal recorded for the previous one.
- `retryContentPage` (`with-stalker-content.feature.ts`) — the Stalker grid
tail's append retry.
- `StalkerSearchComponent`'s search-page retry — the search results have their
own append error and retry, separate from the catalog's.
- `StalkerItvCacheService.refresh()` — the Live TV refresh button, on the same
path that already clears the cache's own error cooldown. This also covers
`refreshChannels()` in the live layout.
- Both account-info dialogs' Retry buttons (`AccountInfoComponent.reload()` for
Xtream, `StalkerAccountInfoComponent.reload()` for Stalker). Their automatic
load on open goes through a private `load()` that does not reset.
- The destructive Xtream refresh — before anything is deleted. It removes the
cached catalog and then bootstraps a re-import whose status request an open
guard would fast-fail, leaving the user with no catalog at all until the
cooldown expires. The reset lives in `XtreamRefreshFlowService.runRefresh()`
(`libs/playlist/shared/ui`), which owns the whole flow for both of its entry
points: `PlaylistRefreshActionService.refreshXtream()` and
`RecentPlaylistsComponent.refreshXtreamPlaylist()` (the Workspace sources
page). Those two used to be independent near-duplicates, which is exactly why
the second one was missed first time round; they now differ only in the
`XtreamRefreshProgressReporter` they hand over, and a reporter cannot reach
the reset. Keep it that way — a third entry point should pass a reporter, not
copy the sequence.
- `PortalStatusService.checkPortalStatusDetails` when `skipCache` is set — the
user-initiated "Test Connection".
Two retry paths deliberately have no reset: Xtream's `retryAppend()` is a no-op
because in-memory appends cannot fail, and the guard's own half-open trial is not
a user action.
Where a retry clears a UI error flag, that flag is cleared **synchronously**
before awaiting the reset — otherwise the retry branch stays re-enterable and
the next `nearEnd` event fires a second retry.
They all go through `resetHostConnectivityGuard()`
(`libs/services/src/lib/host-connectivity-reset.ts`), which holds the one rule
they share: the reset is best effort, because the guard only ever _delays_ a
request and a failed reset must not block the action that asked for it. Both
runtimes honour it, so neither keeps fast-failing an endpoint the user just
asked to retry — in the PWA `PwaService` forwards it to the web backend's
`POST /connectivity-guard/reset`, since the breaker lives in the backend
process and a local reset would clear nothing. That route takes the raw
provider URL rather than a registered `targetId`, because callers reset
precisely when the address may have changed, which is before any target exists
for it; it fetches nothing, reads only the origin, and never logs the URL.
## Interaction with VOD multi-source
No exemption is needed. Multi-source resolution reaches `get_vod_info` over
`XTREAM_REQUEST`, but `VodSourceResolverService.loadVodDetails` already catches
that failure and falls back to its learned container cache or `null`, and
`probeSource` maps `null` to the verdict `unknown` — which
`VodSourceProbeCacheService` deliberately does not cache. A guard fast-fail
therefore degrades to a retryable "could not check", exactly matching the
module's contract that unreachable ≠ contacted-and-refused. The
`STREAM_PROBE_URL` reachability half never rejects and is untouched.
## Tests
- `libs/shared/host-health/src/lib/host-connectivity-guard.spec.ts` — the state
machine, with an injected clock, long-lived trial ownership, late cleanup,
and settlement while the environment kill switch is on.
- `apps/electron-backend/src/app/util/host-connectivity-guard.spec.ts` — the
main-process singleton and redirect attribution through it.
- `apps/web-backend/src/app/web-backend-app.host-guard.spec.ts` — the proxy
routes: the HTTP 200 refusal shape, that the outbound request really is
skipped, that an open endpoint is refused before a DNS lookup is spent on it,
redirect attribution, the discovery exemption, half-open, the reset endpoint,
and that playlist/EPG downloads are never fast-failed. The per-route timeouts
are asserted in `web-backend-app.spec.ts`, which pins the whole outbound
request shape.
- `apps/web/src/app/services/pwa.service.spec.ts` — that a fast-fail reaches the
renderer with no numeric `status` and no `HTTP Error <code>` in its message,
that the discovery bypass is forwarded, and that a reset calls the backend.
- `apps/electron-backend/src/app/events/stalker.events.spec.ts` and
`xtream.events.spec.ts` — trip, fast-fail without contacting axios, reset,
exemption, long pending requests and redirect chains, cancellation, cleanup
when debug reporting throws, and the absence of per-request log spam.
- `libs/portal/stalker/data-access/src/lib/stalker-portal-discovery.utils.spec.ts`
and `stalker-portal-repair.service.spec.ts` — the message contract.