Files
iptvnator/libs/shared/host-health/src/lib/host-connectivity-guard.ts
T
4grayandClaude Fable 5.1 01c423ac43 fix(portals): stop treating a slow panel as a dead host; IPv4 fallback budget in Electron (#1621)
* fix(portals): apply the IPv6->IPv4 fallback budget in the Electron process

The 2500 ms happy-eyeballs attempt timeout from #1404 only ever ran in the
web backend. The Electron main process and its playlist-refresh and EPG
workers kept Node's 250 ms default, so a dual-stack panel hostname behind a
VPN or a slow link failed every connection attempt in a row and tripped the
host connectivity guard. The module now lives in `@iptvnator/shared/host-health`
and every Node isolate that opens connections applies it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(portals): stop treating a slow panel as a dead one in the host guard

axios raises the same ECONNABORTED whether the SYN went unanswered or the
panel accepted the connection and then thought for longer than the request
budget. Two such timeouts opened the breaker and every request to the panel
was refused for 30 s with "portal is not responding" — the shape behind the
"connection keeps dropping" reports on 0.23 and nightly.

Both transports now report whether the TCP connection was established
(`onConnect`: Electron through a per-request observed agent instead of the
shared keep-alive globalAgent, the web backend through the transport that owns
the ClientRequest), and `classifyHostRequestFailure(error, { connected })`
downgrades a host-level code observed after the handshake to inconclusive.
Redirect attribution keeps precedence. A host that never accepts the
connection trips the guard exactly as before.

The Xtream mock gains a `silent:silent` scenario whose detail actions accept
and never answer, plus a real-socket regression spec for the guard.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(portals): let an accepted connection clear the host-failure streak

Review finding: an unanswered SYN, then an accepted-but-slow timeout, then
another unanswered SYN still reached the two-failure threshold, because the
middle request was merely not counted. An accepted TCP connection is the
reachability the guard measures, so it now reads as `responded` and clears
the streak like an HTTP response would. Regression coverage for the mixed
sequence on one flapping loopback origin (Electron) and through the proxy
route (web backend).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(portals): credit an accepted connection when it happens, not when the request settles

Review findings. A request that connected and then hung for 30 s cleared,
on its eventual timeout, the failures later requests had recorded while it
waited — reopening a host that had just died on evidence older than theirs.
The connect hook now reports the connection the moment it fires through a
new `HostConnectivityGuard.reportConnected`, which clears the failure streak
but closes no open or half-open breaker (the trial keeps its slot until it
settles), and the settled timeout is inconclusive.

Electron also skips the socket observer while an environment proxy
(`http_proxy` / `https_proxy` / `all_proxy`) applies to the request: through
a proxy the socket connects to the proxy, whose handshake proves nothing
about the portal, so those requests keep the pre-observer behaviour.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(portals): decide the proxy exemption with axios' own resolution

Review findings. The hand-rolled environment check ignored `no_proxy`, so a
LAN portal exempted from the proxy lost its connect observer and slow
requests to it still tripped the breaker; it also read the variables with
`??`, letting an empty lowercase one mask a populated uppercase one that
axios would honour. The decision now calls `proxy-from-env`'s
`getProxyForUrl`, the same pinned package axios' http adapter uses,
declared as a direct dependency so the packaged app carries it.

The validated-axios spec clears and restores every proxy variable around each
case, so a runner that exports a proxy cannot change what the cases prove.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 14:50:50 +02:00

625 lines
23 KiB
TypeScript

/**
* Per-host circuit breaker for portal requests.
*
* Shared by the two processes that talk to portals on the user's behalf: the
* Electron main process and the self-hosted web backend. Neither the class nor
* anything below it reaches for a transport, a logger or a process singleton —
* the owning app supplies the clock and decides how many guards exist, which is
* what lets the web backend hand its Express app a per-instance guard in tests.
*
* A dead portal costs the full axios timeout on every call (30 s for Xtream,
* 15 s for Stalker), and browsing a dead portal's catalog issues dozens of
* those back to back — 30-second spinners and a flooded log. Once a host has
* refused to answer twice in a row there is nothing left to learn from waiting
* again, so subsequent requests fail immediately for a short while.
*
* The rules are deliberately timid, because being wrong here means refusing to
* talk to a portal that works:
*
* - Only connection-level evidence counts (see
* {@link classifyHostRequestFailure}). Any HTTP response — 200, 404, even
* 502 — proves the host is alive and clears the record.
* - The window is short (30 s), so a mistake costs one page of navigation, and
* a host that came back is retried on its own.
* - Half-open means exactly ONE request goes out, not a whole screenful.
* - An explicit reset (user retry, endpoint discovery) always wins, and
* failures from requests that started before it are discarded — otherwise a
* 30-second straggler settles right after the reset and re-opens the breaker
* underneath the retry that cleared it.
*
* Deliberately not persisted: process lifetime is the right scope for
* "unreachable right now".
*/
import { buildHostConnectivityFastFailMessage } from '@iptvnator/shared/interfaces';
/** Consecutive connection failures that trip the breaker. */
const FAILURE_THRESHOLD = 2;
/**
* How long requests fast-fail once the breaker is open. Exported because both
* apps say it out loud in their "skipping requests for Ns" log line, and a
* second copy of the number would drift from this one.
*/
export const OPEN_DURATION_MS = 30_000;
/**
* Failures further apart than this are unrelated, not a streak. Two sequential
* 30 s timeouts must fit inside it, hence comfortably above 60 s.
*/
const FAILURE_WINDOW_MS = 120_000;
/** Hosts tracked at once; portals per user are few, this is just a bound. */
const MAX_TRACKED_HOSTS = 256;
/** Idle records are forgotten; they hold nothing worth remembering. */
const IDLE_TTL_MS = 600_000;
const GUARD_DISABLED_ENV = 'IPTVNATOR_DISABLE_CONNECTIVITY_GUARD';
/**
* Error codes that prove the host itself did not answer.
*
* `ECONNRESET` is deliberately absent: a reset mid-transfer happens on hosts
* that are very much alive, and it is the one code a working stream can emit.
*/
const HOST_LEVEL_FAILURE_CODES = new Set([
'ECONNABORTED',
'ECONNREFUSED',
'EAI_AGAIN',
'EHOSTUNREACH',
'ENETUNREACH',
'ENOTFOUND',
'ETIMEDOUT',
]);
/**
* Thrown instead of making a request while the breaker is open.
*
* A real `Error`, because Electron serializes a rejected plain object to
* `[object Object]` and the renderer's classification would be lost. It
* carries NO `status` property on purpose — `getStalkerRequestErrorStatus`
* reads that field first, and a numeric status there would read as "the
* endpoint answered".
*/
export class HostConnectivityGuardError extends Error {
/** Scheme, host and port — see {@link portalEndpointKeyOf}. */
readonly endpoint: string;
constructor(endpoint: string) {
super(buildHostConnectivityFastFailMessage(endpoint));
this.name = 'HostConnectivityGuardError';
this.endpoint = endpoint;
}
}
/**
* Handed out by {@link HostConnectivityGuard.check} and passed back when the
* request settles. It records which attempt the report belongs to, so a reset
* or a parallel sibling cannot be mistaken for fresh evidence.
*/
export interface HostRequestToken {
/** Scheme, host and port — see {@link portalEndpointKeyOf}. */
readonly endpoint: string;
readonly epoch: number;
/** Admission order within this guard, independent of clock precision. */
readonly admissionId: number;
/** Whether this request is the single probe allowed while half-open. */
readonly trial: boolean;
/**
* Which half-open slot this request holds, when it holds one.
*
* Cleanup from a released owner can arrive after another trial has been
* admitted. Only this identity (plus epoch) may release the current slot.
*/
readonly trialId: number;
}
export type HostConnectivityCheck =
| { readonly allowed: true; readonly token: HostRequestToken }
| { readonly allowed: false; readonly retryAfterMs: number };
/**
* `responded`: the endpoint answered with an HTTP response of any status.
* `host-level`: the host never answered. `inconclusive`: the failure says
* nothing about reachability — including a timeout after an accepted
* connection, whose evidence the owner already reported at connect time (see
* {@link classifyHostRequestFailure}).
*/
export type HostRequestOutcome = 'responded' | 'host-level' | 'inconclusive';
interface HostState {
consecutiveFailures: number;
/** When the last COUNTED failure was recorded. */
lastFailureAt: number;
/** Latest guard-wide admission id when this host's last failure counted. */
lastFailureAdmissionId: number;
openUntil: number;
trialInFlight: boolean;
/** Monotonic id of the half-open slot; see `HostRequestToken.trialId`. */
trialId: number;
epoch: number;
lastTouchedAt: number;
}
export interface HostConnectivityGuardOptions {
now?: () => number;
onOpen?: (endpoint: string) => void;
}
/**
* Whether the guard is switched off. Read at call time, not at module load, so
* a debugging session can toggle it (matches `IPTVNATOR_ALLOW_INSECURE_TLS`).
*/
export function isHostConnectivityGuardDisabled(): boolean {
const value = process.env[GUARD_DISABLED_ENV]?.trim().toLowerCase();
return value === '1' || value === 'true';
}
/**
* What a failed request proves about the host.
*
* `responded` covers an error that still carries an HTTP response (5xx reaches
* the handlers as a rejection because `validateStatus` only tolerates <500) —
* the host is alive, so it clears the record. `inconclusive` covers everything
* that says nothing about reachability: a cancelled request, a URL rejected by
* the SSRF policy, a bug in our own code.
*/
export interface HostRequestFailureContext {
/**
* Whether the transport saw the TCP connection for this request get
* established (or handed the request a socket that already was).
*/
readonly connected?: boolean;
}
/**
* Reads what a failed request proves about its endpoint.
*
* `context.connected` matters for the timeout codes: axios raises the same
* `ECONNABORTED` whether the SYN went unanswered or the panel accepted the
* connection and then thought for longer than the request budget. Only the
* former is a host that stopped answering. A panel that is merely slow (a
* heavy `get_vod_info` on a busy home server) accepts every connection, and
* tripping the breaker on it turns "slow" into thirty seconds of "not
* responding" for every request — the reported symptom. So a host-level
* code observed after the handshake is `inconclusive`: it must not count.
*
* It does not clear the streak here either, and that is deliberate. The
* accepted connection IS reachability evidence and does clear the streak —
* but the owner reports it the moment the socket connects
* (`reportConnected` from the transport's connect hook), not when the
* timeout settles up to 30 s later. Ordering matters: while A sits connected and
* silent, B and C can fail to connect and open the breaker for a host that
* has just died; clearing at A's settle time would reopen it on evidence
* older than B's and C's failures. The remaining codes cannot occur once a
* connection exists, so the rule costs nothing for them.
*/
export function classifyHostRequestFailure(
error: unknown,
context: HostRequestFailureContext = {}
): HostRequestOutcome {
if (!error || typeof error !== 'object') {
return 'inconclusive';
}
if ((error as { response?: unknown }).response) {
return 'responded';
}
const code = (error as { code?: unknown }).code;
if (typeof code === 'string' && HOST_LEVEL_FAILURE_CODES.has(code)) {
return context.connected ? 'inconclusive' : 'host-level';
}
return 'inconclusive';
}
/**
* The guard key for a request URL: its ORIGIN, i.e. scheme, host and port.
*
* Not `URL.host`, which omits a default port and therefore gives
* `http://panel.example` and `https://panel.example` the same key — two
* genuinely different endpoints, and a panel whose TLS listener is broken while
* plain HTTP works is a routine IPTV setup. Sharing one record there would let
* the dead one fast-fail the working one without ever contacting it.
*
* `URL.origin` also leaves out any `user:pass@` userinfo, so no credential
* reaches the key or the log line.
*/
export function portalEndpointKeyOf(url: string): string | null {
try {
const origin = new URL(url).origin;
// Opaque origins serialize as "null" and are not a usable key.
return origin && origin !== 'null' ? origin : null;
} catch {
return null;
}
}
export class HostConnectivityGuard {
private readonly states = new Map<string, HostState>();
/** Never reused when an endpoint is evicted while requests are in flight. */
private admissionId = 0;
private readonly now: () => number;
private readonly onOpen?: (endpoint: string) => void;
constructor(options: HostConnectivityGuardOptions = {}) {
this.now = options.now ?? (() => Date.now());
this.onOpen = options.onOpen;
}
/**
* Whether a request to `endpoint` may go out, and the token to report with.
*
* Fast-fails while the breaker is open, and while a half-open trial is
* still in flight — the point of the trial is that a screenful of requests
* does not all hang again to learn the same thing.
*/
check(endpoint: string): HostConnectivityCheck {
const now = this.now();
if (isHostConnectivityGuardDisabled()) {
return {
allowed: true,
token: {
endpoint,
epoch: 0,
admissionId: 0,
trial: false,
trialId: 0,
},
};
}
const state = this.ensureState(endpoint, now);
state.lastTouchedAt = now;
if (state.openUntil > now) {
return { allowed: false, retryAfterMs: state.openUntil - now };
}
// Never opened, or the open window elapsed. The latter is half-open:
// let one request through and hold the rest back until it settles.
let trial = false;
if (state.openUntil > 0) {
if (state.trialInFlight) {
return { allowed: false, retryAfterMs: 0 };
}
state.trialInFlight = true;
state.trialId += 1;
trial = true;
}
return {
allowed: true,
token: {
endpoint,
epoch: state.epoch,
admissionId: ++this.admissionId,
trial,
trialId: trial ? state.trialId : 0,
},
};
}
/**
* The host answered. Always clears the record, whatever the status was and
* whichever attempt it belonged to — a reachable host is a reachable host.
*/
reportSuccess(token: HostRequestToken): void {
const state = this.states.get(token.endpoint);
if (!state) {
return;
}
if (isHostConnectivityGuardDisabled()) {
this.releaseTrial(state, token);
return;
}
state.consecutiveFailures = 0;
state.lastFailureAt = 0;
state.openUntil = 0;
state.trialInFlight = false;
state.lastTouchedAt = this.now();
}
/**
* The host accepted this request's TCP connection. Reported the moment it
* happens, so its place in the order of evidence is exact: it clears the
* failure streak recorded up to now, and later failures start a new one.
*
* It deliberately does NOT close an open or half-open breaker. Whether a
* host that has just accepted a connection also answers is exactly what
* the half-open trial exists to find out, so the trial keeps its slot
* until it settles, and a request admitted before the breaker opened does
* not reopen the host on the strength of a handshake alone.
*/
reportConnected(token: HostRequestToken): void {
const state = this.states.get(token.endpoint);
if (!state || isHostConnectivityGuardDisabled()) {
return;
}
state.consecutiveFailures = 0;
state.lastFailureAt = 0;
state.lastTouchedAt = this.now();
}
/** The host did not answer at all. */
reportFailure(token: HostRequestToken): void {
const state = this.states.get(token.endpoint);
if (!state) {
return;
}
if (isHostConnectivityGuardDisabled()) {
this.releaseTrial(state, token);
return;
}
const now = this.now();
// Superseded by an explicit reset: the user asked for a fresh attempt
// and this verdict predates it.
if (state.epoch !== token.epoch) {
this.releaseTrial(state, token);
return;
}
// Only the request that still holds the slot ends the half-open state.
// An abandoned trial's late failure is ordinary evidence — it goes
// through the streak rules below and leaves the replacement alone.
const wasTrial = this.ownsTrial(state, token);
if (wasTrial) {
state.trialInFlight = false;
}
state.lastTouchedAt = now;
if (
state.consecutiveFailures > 0 &&
now - state.lastFailureAt > FAILURE_WINDOW_MS
) {
state.consecutiveFailures = 0;
}
// A request already in flight when the previous failure was recorded is
// not the next link in a streak — it is a sibling of it. Catalog
// loading fans out several requests at once, and one hiccup failing all
// of them is one piece of evidence, not a trip. Admission order makes
// that boundary exact even when all events share one clock tick.
let counted = false;
if (
state.consecutiveFailures === 0 ||
token.admissionId > state.lastFailureAdmissionId
) {
state.consecutiveFailures += 1;
state.lastFailureAt = now;
state.lastFailureAdmissionId = this.admissionId;
counted = true;
}
// Only a failure this report actually counted may trip the threshold.
// A sibling arriving after the window elapsed would otherwise open a
// fresh one off the existing count, pushing the half-open trial past
// the intended cooldown. A failed trial still goes straight back to
// open: the host had its chance.
if (
wasTrial ||
(counted && state.consecutiveFailures >= FAILURE_THRESHOLD)
) {
const wasOpen = state.openUntil > now;
state.openUntil = now + OPEN_DURATION_MS;
if (!wasOpen) {
this.onOpen?.(token.endpoint);
}
}
}
/**
* The request failed for a reason that says nothing about the host. Only
* releases the half-open slot, so the next request can be the trial.
* Owners MUST also call this in finally, independently of outcome reporting.
* It is idempotent and still releases while the guard is disabled by the environment;
* no elapsed-time expiry can distinguish a leak from an active transfer.
*/
reportInconclusive(token: HostRequestToken): void {
const state = this.states.get(token.endpoint);
if (!state) {
return;
}
this.releaseTrial(state, token);
state.lastTouchedAt = this.now();
}
/**
* Forgets everything recorded for `endpoint`, and invalidates the reports of
* requests already in flight. Called when the user asks for a real attempt
* or when the URL may now point somewhere else.
*/
reset(endpoint: string): void {
const now = this.now();
const state = this.ensureState(endpoint, now);
state.consecutiveFailures = 0;
state.lastFailureAt = 0;
state.openUntil = 0;
state.trialInFlight = false;
state.epoch += 1;
state.lastTouchedAt = now;
}
/**
* A token for a request that must NOT be policed or counted, but whose
* success still clears the record.
*
* Stalker endpoint discovery probes several candidate paths on one host and
* expects most of them to fail; counting those would let it declare a
* slow-but-alive portal unreachable. It does not take the half-open slot
* either — a probe is not the trial the guard is waiting for.
*/
observe(endpoint: string): HostRequestToken {
return {
endpoint,
epoch: this.states.get(endpoint)?.epoch ?? 0,
admissionId: 0,
trial: false,
trialId: 0,
};
}
/** Test seam: drops all recorded state. */
clear(): void {
this.states.clear();
}
/** Whether `token` still holds the current half-open slot. */
private ownsTrial(state: HostState, token: HostRequestToken): boolean {
return (
token.trial &&
state.trialInFlight &&
state.epoch === token.epoch &&
state.trialId === token.trialId
);
}
private releaseTrial(state: HostState, token: HostRequestToken): void {
if (this.ownsTrial(state, token)) {
state.trialInFlight = false;
}
}
private ensureState(endpoint: string, now: number): HostState {
const existing = this.states.get(endpoint);
if (existing) {
return existing;
}
this.prune(now);
const state: HostState = {
consecutiveFailures: 0,
lastFailureAt: 0,
lastFailureAdmissionId: 0,
openUntil: 0,
trialInFlight: false,
trialId: 0,
epoch: 0,
lastTouchedAt: now,
};
this.states.set(endpoint, state);
return state;
}
private prune(now: number): void {
for (const [endpoint, state] of this.states) {
if (
now - state.lastTouchedAt > IDLE_TTL_MS &&
state.openUntil <= now &&
!state.trialInFlight
) {
this.states.delete(endpoint);
}
}
// Still at the cap: forget the oldest records. Dropping one means
// contacting that host again, which is the safe direction to err in.
while (this.states.size >= MAX_TRACKED_HOSTS) {
const oldest = this.states.keys().next();
if (oldest.done) {
return;
}
this.states.delete(oldest.value);
}
}
}
/**
* The URL a failed request was actually talking to, when the error says.
*
* Redirects are followed hop by hop, each with its own config, so a failure on
* a later hop carries THAT hop's URL rather than the one we asked for.
*/
function failedRequestUrlOf(error: unknown): string | null {
// Two transports, two places to look, and only one of them is ever right.
//
// The Electron transport follows redirects itself with `maxRedirects: 0`,
// reissuing each hop as its own request, so the hop that failed is the
// error's `config.url`.
//
// The web backend uses axios' default transport, where follow-redirects
// walks the chain inside a single request. `config` is built once and never
// rewritten, so `config.url` stays the URL we asked for — comparing it
// against itself would find no redirect and charge a dead destination to
// the provider that answered. follow-redirects tracks the hop it is on as
// `request._currentUrl`, so that is read first; a transport that does not
// expose it falls through to `config.url`.
const currentUrl = (
error as { request?: { _currentUrl?: unknown } | null } | null
)?.request?._currentUrl;
if (typeof currentUrl === 'string') {
return currentUrl;
}
const url = (error as { config?: { url?: unknown } } | null)?.config?.url;
return typeof url === 'string' ? url : null;
}
/**
* Origin and path only — deliberately without the query string.
*
* The baseline a caller can supply is the URL it handed the transport, but the
* transport is free to add to the query before sending: the web backend passes
* Xtream credentials through axios' `params`, so a request for
* `…/player_api.php` goes out as `…/player_api.php?username=…&action=…`.
* Comparing whole URLs then reports a redirect for every ordinary failure,
* which credits the endpoint instead of counting it and stops the breaker from
* ever opening.
*
* Dropping the query keeps what the comparison is actually for — an endpoint
* that answered and sent us somewhere else, including the same-origin
* `/player_api.php` → `/slow/player_api.php` case — and gives up only a
* redirect that changes nothing but the query. That one is then counted as an
* ordinary failure, which is the safe direction to be wrong in.
*/
function normalizedUrlOrNull(url: string): string | null {
try {
const parsed = new URL(url);
return `${parsed.origin}${parsed.pathname}`;
} catch {
return null;
}
}
/**
* Whether the failure happened on a hop the guarded endpoint redirected us to.
*
* Reaching any later hop proves the guarded endpoint answered: the first hop is
* always the URL we asked for, and only a redirect status advances the chain.
* That holds for a same-origin redirect too, so comparing the whole URL — not
* just its origin — is what catches `/player_api.php` → `/slow/player_api.php`.
*
* Requires positive evidence: anything unparseable or unknown returns false and
* the failure is counted as usual, because guessing "redirect" here would stop
* the guard from ever tripping.
*/
export function failedAfterRedirect(
error: unknown,
token: HostRequestToken,
requestUrl: string | undefined
): boolean {
const failedUrl = failedRequestUrlOf(error);
if (!failedUrl) {
return false;
}
const failedEndpoint = portalEndpointKeyOf(failedUrl);
if (failedEndpoint && failedEndpoint !== token.endpoint) {
return true;
}
if (!requestUrl) {
return false;
}
const failedNormalized = normalizedUrlOrNull(failedUrl);
const requestedNormalized = normalizedUrlOrNull(requestUrl);
return (
failedNormalized !== null &&
requestedNormalized !== null &&
failedNormalized !== requestedNormalized
);
}