fix(portals): time out and fast-fail PWA proxy requests to dead hosts (#1424)

* refactor(portals): hoist the connectivity guard into libs/shared/host-health

The breaker was written dependency-free so both processes that talk to
portals could share it. Move the part that has no Electron in it — the
state machine, the failure classification, the redirect-attribution
helpers and the fast-fail error — into `@iptvnator/shared/host-health`
(`scope:shared` / `domain:shared-runtime` / `type:util`).

What stays in `apps/electron-backend` is the genuinely main-process part:
one guard for the whole process, so both portal IPC handlers see each
other's evidence, and the console warning that announces it. Every call
site is unchanged; the wrapper re-exports the two types they import.

The spec splits the same way — the state machine moves with the class,
the singleton and its redirect attribution stay with the wrapper.

Register the new project in the coverage policy, which every project with
a test target must declare a tier for.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(portals): time out and fast-fail PWA proxy requests to dead hosts

The web backend's proxy routes were bare `axios.get()` calls with no
`timeout`, so a provider that accepted a connection and then went silent
held the request until the OS gave up on the TCP connection — minutes,
rather than the 15/30 s budget the Electron handlers use. Add the same
per-route timeouts (Xtream 30 s, Stalker 15 s / 30 s for `create_link`,
playlist and XMLTV 30 s).

Those numbers are safe for large downloads: on axios' default transport
`timeout` bounds the time to response headers and then continues as the
socket's inactivity timeout, so a multi-megabyte XMLTV file that keeps
delivering bytes is never cut off — only a stalled one is.

With requests bounded, run `/xtream` and `/stalker` through the shared
breaker, injected via `WebBackendAppOptions.hostGuard` so specs drive it
with a clock they own. Playlist and XMLTV downloads keep the timeout but
no breaker, matching Electron: a download is one request rather than a
catalog fan-out, it is usually the direct result of the user asking for
it, and it can outlive the half-open trial window.

The breaker is checked before the Xtream URL revalidation, which resolves
the hostname — a dead host is where DNS is slow too, and a request
admitted and then abandoned by the URL policy hands its token back rather
than holding the trial slot.

`resetHostConnectivityGuard()` no longer no-ops in the PWA: the breaker
lives in the backend process, so it travels to a new
`POST /connectivity-guard/reset`, which reads only the origin and never
logs the credential-bearing URL. `skipConnectionGuard` now survives the
PWA transport too, so Stalker endpoint discovery keeps the exemption it
has on the desktop instead of tripping the breaker with its own probes.

A fast-fail keeps each route's HTTP 200 `{message, status}` envelope. The
Stalker path needs one extra step: `forwardStalkerRequest` turns that
envelope into `HTTP Error <code>: …` with a numeric `status`, and the
renderer reads both as "the endpoint answered" — which would make
discovery walk every candidate and fire lazy repair at a host just
declared dead. A prior branch keyed on the shared
`isHostConnectivityFastFailMessage` rethrows it bare instead.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(portals): report exempt Stalker probe responses to the PWA breaker

Flagged by the author of #1421 as one of the twelve fixes that landed
there after this branch cherry-picked the pre-review commit: an exempt
discovery probe must still REPORT, it just must not COUNT.

The web backend was skipping the report entirely for a probe, which loses
the case that matters. A failure carrying an HTTP response proves the
endpoint answered, and this route sets no `validateStatus`, so axios
rejects every non-2xx with `error.response` attached — a probe answered
with 404 or 500 was therefore dropped instead of clearing the record.
Two counted failures either side of it then read as consecutive and
opened the breaker in the middle of discovery, which is exactly what the
exemption exists to prevent.

`reportProviderRequestFailure` now takes `countFailures`, matching the
Electron reporter, and both routes always report.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(portals): read the redirect hop from the transport that followed it

`failedAfterRedirect` decided whether a failure belonged to a redirect
destination by reading `error.config.url`. That is right for the Electron
transport, which sets `maxRedirects: 0` and reissues every hop as its own
request, so the hop IS the config URL. It is blind on the web backend,
which uses axios' default transport: follow-redirects walks the chain
inside one request and `config` is built once, so `config.url` stays the
URL we asked for.

Verified against the installed axios 1.19.0 with a live server that 302s
to a dead port:

    asked for          : http://127.0.0.1:63953/player_api.php
    config.url         : http://127.0.0.1:63953/player_api.php
    request._currentUrl: http://127.0.0.1:1/dead

So the comparison was original-vs-original, found no redirect, and
charged two dead destinations to the provider that had answered both
times with a 302 — then fast-failed it. Read `request._currentUrl` first
and fall back to `config.url`, which covers both transports; Electron's
native per-hop requests expose no `_currentUrl` and are unaffected.

Also check redirect attribution BEFORE suppressing failure counting for
an exempt probe. A 3xx from the guarded endpoint is an answer, so a probe
that observed one must clear the record; otherwise a timeout, a probe
redirected to a dead destination, and another timeout still read as two
consecutive failures.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(portals): stop axios query params reading as a redirect

Codex found that the Xtream breaker never opened at all, and it was
right. The web backend passes credentials and the action through axios'
`params`, so axios sends `…/player_api.php?username=…&action=…` while the
baseline handed to `failedAfterRedirect` is the query-less URL the route
built. Verified against axios 1.19.0 with a plain ECONNREFUSED and no
redirect anywhere in sight:

    baseline           : http://127.0.0.1:1/player_api.php
    request._currentUrl: http://127.0.0.1:1/player_api.php?username=demo&…

The two normalized URLs differ, so every ordinary failure looked like a
post-redirect failure, credited the endpoint, and the breaker could never
trip.

Compare origin and path, not the whole URL. That keeps what the check is
for — an endpoint that answered and sent us elsewhere, including the
same-origin `/player_api.php` → `/slow/player_api.php` case — and gives
up only a redirect that changes nothing but the query, which is then
counted as an ordinary failure. Erring towards counting is the safe
direction here.

The reason 57 tests passed over a dead feature is the real lesson:
`StubHttpClient` threw bare `Error`s, so the guard's redirect check saw
neither `config.url` nor `request._currentUrl` and quietly did nothing.
The stub now shapes its rejections like axios does, including the query
axios appends. With that alone, four existing tests fail against the old
comparison.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(portals): count a hostname that stops resolving, and fix a stale docblock

Two from review, one behavioural and one documentation.

A name that will not resolve is the host failing to answer — the same
evidence as the ENOTFOUND the transport would have raised a moment later.
But the SSRF validation turns a lookup failure into a 400 "host could not
be resolved", and the release path added earlier handed the token back as
inconclusive, so the breaker could never open for a host whose DNS died
and every request kept paying for the same dead lookup.

`ProviderUrlError` now carries the underlying lookup error internally.
A refusal that has one is counted; a genuine policy refusal — private
address, bad scheme, credentials in the URL — still only releases the
half-open slot, because that says nothing about reachability. The field
is internal: `providerUrlErrorBody()` strips it at both call sites, so
the client sees exactly the body it saw before, which the test asserts.

The docblock on `resetHostConnectivityGuard` still said the PWA channel
is unknown and the call no-ops. That stopped being true when this branch
implemented `CONNECTIVITY_GUARD_RESET` over HTTP, and a stale contract
there is how the next caller silently skips the PWA path. (The edit was
in an earlier commit and was lost when the branch was rebuilt on master.)

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(portals): record the transport-specific redirect contract

The redirect section still described one transport: hop-by-hop requests,
`error.config.url`, whole-URL comparison. Two of those three are now
wrong for the web backend, and this document is the canonical contract —
leaving it stale is how the attribution bugs fixed in the last two
commits get reintroduced.

Says what is actually true: which field holds the failed hop on each
transport and why the helper reads both, and that the comparison is
origin + path because the web backend's credentials ride in axios'
`params` and a whole-URL comparison therefore reported a redirect for
every ordinary failure.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(portals): scope the guard summary to both processes

The opening line still defined the breaker as an Electron main-process
concern, which contradicted the ownership section below it and is the
part a reader skims to decide whether the document applies to them.

Names both processes, and records that the web backend had the worse
version of the problem first — no request timeout at all — since that is
why the timeouts and the breaker had to land there together.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(portals): declare the shared-interfaces dependency of host-health

The new library's manifest listed only `tslib`, but its emitted JavaScript
does `require('@iptvnator/shared/interfaces')` — the guard builds its
fast-fail message with `buildHostConnectivityFastFailMessage`. Anything
resolving the built artifact from its own manifest would have failed with
MODULE_NOT_FOUND.

The manifest was copied from `shared/logging`, which imports nothing
across libraries and therefore needs nothing beyond `tslib`.
`shared/m3u-utils` is the right precedent: it imports the same library
and declares `"@iptvnator/shared/interfaces": "0.0.1"`.

Verified against the build output rather than by inspection — the emitted
`host-connectivity-guard.js` requires the module, and the generated
`dist/libs/shared/host-health/package.json` now declares it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(portals): bound the PWA connectivity-guard reset

The reset was a bare `fetch` with no timeout, which is the exact failure
this change exists to remove, reintroduced one layer up. Every caller
awaits the reset BEFORE issuing the request it is clearing the way for —
`retryContentInitialization` awaits it first by design — so a backend or
reverse proxy that accepts the POST and then goes quiet would leave
Retry doing nothing at all, for as long as the socket stayed open.

Bound it with an AbortController and a 5 s timer. The abort rejects,
`resetHostConnectivityGuard` swallows it as it already does for any
other failure, and the caller proceeds to its real request — which is
what "best effort" was supposed to mean. The timer is cleared in a
`finally`, and it covers the body read as well as the headers.

Five seconds because this talks to the user's own backend rather than a
provider: it should answer immediately, and a slow one must not hold up
the retry that asked for it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
4grayandClaude Opus 5 authored and GitHub committed 2026-08-13 22:27:05 +02:00
1 parent 3103eba083
commit e94cc029eb
24 files changed
+2895 -1171

No files matched your search

@@ -1,498 +1,16 @@
/**
* Main-process wrapper behaviour: the shared singleton and what a failure
* reported through it proves about the endpoint. The state machine itself is
* covered in `@iptvnator/shared/host-health`.
*/
import { HostConnectivityGuardError } from '@iptvnator/shared/host-health';
import {
buildHostConnectivityFastFailMessage,
isHostConnectivityFastFailMessage,
} from '@iptvnator/shared/interfaces';
import {
HostConnectivityGuard,
HostConnectivityGuardError,
HostRequestToken,
beginGuardedHostRequest,
classifyHostRequestFailure,
portalEndpointKeyOf,
reportGuardedHostFailure,
resetHostConnectivityGuardForTests,
} from './host-connectivity-guard';
const HOST = 'http://portal.example.com:8080';
const GUARD_DISABLED_ENV = 'IPTVNATOR_DISABLE_CONNECTIVITY_GUARD';
const OPEN_DURATION_MS = 30_000;
const FAILURE_WINDOW_MS = 120_000;
const TRIAL_TIMEOUT_MS = 45_000;
function timeoutError(code = 'ETIMEDOUT') {
return Object.assign(new Error('timeout of 30000ms exceeded'), { code });
}
describe('classifyHostRequestFailure', () => {
it('treats connection-level error codes as host-level failures', () => {
for (const code of [
'ECONNABORTED',
'ECONNREFUSED',
'EAI_AGAIN',
'EHOSTUNREACH',
'ENETUNREACH',
'ENOTFOUND',
'ETIMEDOUT',
]) {
expect(classifyHostRequestFailure(timeoutError(code))).toBe(
'host-level'
);
}
});
it('treats an error carrying an HTTP response as proof the host answered', () => {
// 5xx reaches the handlers as a rejection: validateStatus only
// tolerates < 500. The host still answered.
const serverError = Object.assign(new Error('Request failed'), {
code: 'ERR_BAD_RESPONSE',
response: { status: 502, statusText: 'Bad Gateway' },
});
expect(classifyHostRequestFailure(serverError)).toBe('responded');
});
it('does not read reachability into cancellations, SSRF refusals or plain errors', () => {
expect(
classifyHostRequestFailure(
Object.assign(new Error('canceled'), { code: 'ERR_CANCELED' })
)
).toBe('inconclusive');
expect(
classifyHostRequestFailure(
new Error('URL host could not be resolved')
)
).toBe('inconclusive');
expect(classifyHostRequestFailure(undefined)).toBe('inconclusive');
expect(classifyHostRequestFailure('ETIMEDOUT')).toBe('inconclusive');
});
it('never counts a mid-transfer connection reset', () => {
// A reset happens on hosts that are very much alive.
expect(classifyHostRequestFailure(timeoutError('ECONNRESET'))).toBe(
'inconclusive'
);
});
});
describe('portalEndpointKeyOf', () => {
it('keys on host and port so two panels on one machine stay separate', () => {
expect(
portalEndpointKeyOf('http://example.com:8080/player_api.php')
).toBe('http://example.com:8080');
expect(
portalEndpointKeyOf('http://example.com:9090/player_api.php')
).toBe('http://example.com:9090');
});
it('separates HTTP from HTTPS on their default ports', () => {
// `URL.host` omits a default port, so both would collapse onto
// `example.com` — and a panel whose TLS listener is broken while plain
// HTTP works is a routine IPTV setup.
expect(portalEndpointKeyOf('http://example.com/player_api.php')).toBe(
'http://example.com'
);
expect(portalEndpointKeyOf('https://example.com/player_api.php')).toBe(
'https://example.com'
);
});
it('leaves URL credentials out of the key', () => {
expect(portalEndpointKeyOf('http://user:pass@example.com/c')).toBe(
'http://example.com'
);
});
it('returns null for an unparseable URL instead of throwing', () => {
expect(portalEndpointKeyOf('not a url')).toBeNull();
});
});
describe('HostConnectivityGuardError', () => {
it('carries the shared fast-fail message and no status field', () => {
const error = new HostConnectivityGuardError(HOST);
expect(error).toBeInstanceOf(Error);
expect(error.message).toBe(buildHostConnectivityFastFailMessage(HOST));
// getStalkerRequestErrorStatus reads `status` first; a number there
// would read as "the endpoint answered".
expect(
(error as unknown as { status?: unknown }).status
).toBeUndefined();
expect(isHostConnectivityFastFailMessage(error.message)).toBe(true);
});
});
describe('HostConnectivityGuard', () => {
let clock = 1_000_000;
let opened: string[];
let guard: HostConnectivityGuard;
const advance = (ms: number) => {
clock += ms;
};
/** Runs a request that fails at the host level after `durationMs`. */
const failRequest = (durationMs = 0): void => {
const check = guard.check(HOST);
if (!check.allowed) {
throw new Error('expected the request to be allowed');
}
advance(durationMs);
guard.reportFailure(check.token);
};
const expectAllowed = (): HostRequestToken => {
const check = guard.check(HOST);
expect(check.allowed).toBe(true);
if (!check.allowed) {
throw new Error('unreachable');
}
return check.token;
};
const expectBlocked = (): void => {
expect(guard.check(HOST).allowed).toBe(false);
};
beforeEach(() => {
clock = 1_000_000;
opened = [];
delete process.env[GUARD_DISABLED_ENV];
guard = new HostConnectivityGuard({
now: () => clock,
onOpen: (host) => opened.push(host),
});
});
it('allows requests to a host it has never seen', () => {
expect(guard.check(HOST).allowed).toBe(true);
expect(opened).toEqual([]);
});
it('still allows the request after a single failure', () => {
failRequest(30_000);
expect(guard.check(HOST).allowed).toBe(true);
expect(opened).toEqual([]);
});
it('opens after two consecutive failures and reports the wait', () => {
failRequest(30_000);
failRequest(30_000);
expect(guard.check(HOST)).toEqual({
allowed: false,
retryAfterMs: OPEN_DURATION_MS,
});
expect(opened).toEqual([HOST]);
});
it('logs the open transition once, not per blocked request', () => {
failRequest(30_000);
failRequest(30_000);
expectBlocked();
expectBlocked();
expect(opened).toEqual([HOST]);
});
it('keeps other hosts untouched', () => {
failRequest(30_000);
failRequest(30_000);
expect(guard.check('http://other.example.com').allowed).toBe(true);
});
it('keeps the same host on another scheme reachable', () => {
// Regression: keying by `URL.host` dropped the default port, so a dead
// HTTPS panel fast-failed the working HTTP one on the same machine.
failRequest(30_000);
failRequest(30_000);
expectBlocked();
expect(guard.check('https://portal.example.com:8080').allowed).toBe(
true
);
});
it('counts a parallel fan-out that fails together as one failure', () => {
// Catalog init loads live/vod/series at once. One wifi hiccup failing
// all three is one piece of evidence, not a trip.
const first = expectAllowed();
const second = expectAllowed();
const third = expectAllowed();
advance(30_000);
guard.reportFailure(first);
guard.reportFailure(second);
guard.reportFailure(third);
expect(guard.check(HOST).allowed).toBe(true);
expect(opened).toEqual([]);
});
it('opens when a second fan-out fails after the first one did', () => {
const first = expectAllowed();
const second = expectAllowed();
advance(30_000);
guard.reportFailure(first);
guard.reportFailure(second);
failRequest(30_000);
expectBlocked();
expect(opened).toEqual([HOST]);
});
it('does not re-open on a sibling that settles after the window elapsed', () => {
// A and B start together; A plus a later request open the breaker. B is
// explicitly not counted as a strike, so it must not start a fresh
// 30-second window either — that would push the half-open trial past
// the intended cooldown.
const sibling = expectAllowed();
const first = expectAllowed();
advance(30_000);
guard.reportFailure(first);
failRequest(30_000);
expectBlocked();
expect(opened).toEqual([HOST]);
advance(OPEN_DURATION_MS);
guard.reportFailure(sibling);
const trial = expectAllowed();
expect(trial.trial).toBe(true);
expect(opened).toEqual([HOST]);
});
it('does not accumulate failures further apart than the streak window', () => {
failRequest(0);
advance(FAILURE_WINDOW_MS + 1);
failRequest(0);
expect(guard.check(HOST).allowed).toBe(true);
expect(opened).toEqual([]);
});
it('accumulates failures exactly at the streak window edge', () => {
failRequest(0);
advance(FAILURE_WINDOW_MS);
failRequest(0);
expectBlocked();
});
it('clears the record as soon as the host answers', () => {
failRequest(30_000);
const token = expectAllowed();
guard.reportSuccess(token);
failRequest(30_000);
expect(guard.check(HOST).allowed).toBe(true);
expect(opened).toEqual([]);
});
it('treats a success followed by an inconclusive report as a no-op', () => {
// The Stalker handler reports success on the response, then throws its
// own `HTTP Error 404` error, which is classified inconclusive.
failRequest(30_000);
const token = expectAllowed();
guard.reportSuccess(token);
guard.reportInconclusive(token);
failRequest(30_000);
expect(guard.check(HOST).allowed).toBe(true);
});
it('ignores an inconclusive failure entirely', () => {
const first = expectAllowed();
guard.reportInconclusive(first);
const second = expectAllowed();
guard.reportInconclusive(second);
expect(guard.check(HOST).allowed).toBe(true);
expect(opened).toEqual([]);
});
describe('half-open', () => {
beforeEach(() => {
failRequest(30_000);
failRequest(30_000);
expectBlocked();
advance(OPEN_DURATION_MS);
});
it('lets exactly one request through once the window elapses', () => {
const trial = expectAllowed();
expect(trial.trial).toBe(true);
expect(guard.check(HOST).allowed).toBe(false);
expect(guard.check(HOST).allowed).toBe(false);
});
it('closes the breaker when the trial succeeds', () => {
const trial = expectAllowed();
guard.reportSuccess(trial);
const next = expectAllowed();
expect(next.trial).toBe(false);
expect(guard.check(HOST).allowed).toBe(true);
});
it('re-opens immediately when the trial fails, without a second strike', () => {
const trial = expectAllowed();
advance(30_000);
guard.reportFailure(trial);
expect(guard.check(HOST)).toEqual({
allowed: false,
retryAfterMs: OPEN_DURATION_MS,
});
expect(opened).toEqual([HOST, HOST]);
});
it('offers the trial slot again when the trial is cancelled', () => {
const trial = expectAllowed();
guard.reportInconclusive(trial);
const retry = expectAllowed();
expect(retry.trial).toBe(true);
});
it('lets only the request holding the slot release it', () => {
// A trial can outlive its own 45 s window: the validated-redirect
// transport gives each of up to five hops its own 30 s budget. Once
// a replacement has been admitted, the abandoned request's late
// report must not hand the slot to a third request.
const abandoned = expectAllowed();
advance(TRIAL_TIMEOUT_MS);
const replacement = expectAllowed();
expect(replacement.trial).toBe(true);
guard.reportInconclusive(abandoned);
expectBlocked();
});
it('recovers when a trial never reports back', () => {
expectAllowed();
expectBlocked();
advance(TRIAL_TIMEOUT_MS);
const replacement = expectAllowed();
expect(replacement.trial).toBe(true);
});
});
describe('reset', () => {
it('reopens the host for requests immediately', () => {
failRequest(30_000);
failRequest(30_000);
expectBlocked();
guard.reset(HOST);
expect(expectAllowed().trial).toBe(false);
});
it('discards failures from requests that started before it', () => {
// The retry that cleared the breaker must not be poisoned by the
// 30-second stragglers it was waiting behind.
const straggler = expectAllowed();
const secondStraggler = expectAllowed();
advance(15_000);
guard.reset(HOST);
advance(15_000);
guard.reportFailure(straggler);
guard.reportFailure(secondStraggler);
expect(guard.check(HOST).allowed).toBe(true);
expect(opened).toEqual([]);
});
it('still counts failures from requests started after it', () => {
failRequest(30_000);
guard.reset(HOST);
failRequest(30_000);
failRequest(30_000);
expectBlocked();
});
it('supersedes in-flight requests for a host it has no record of', () => {
const token = expectAllowed();
guard.clear();
guard.reset(HOST);
advance(30_000);
guard.reportFailure(token);
failRequest(30_000);
expect(guard.check(HOST).allowed).toBe(true);
});
});
describe('bookkeeping bounds', () => {
it('forgets idle records instead of growing without bound', () => {
for (let index = 0; index < 300; index += 1) {
advance(1);
guard.check(`http://host-${index}.example.com`);
}
// The cap held, and a fresh host is still allowed through.
expect(guard.check(HOST).allowed).toBe(true);
});
it('keeps an open breaker while other hosts churn past the idle TTL', () => {
failRequest(30_000);
failRequest(30_000);
expectBlocked();
advance(1);
guard.check('http://noise.example.com');
expectBlocked();
});
});
describe('kill switch', () => {
afterEach(() => {
delete process.env[GUARD_DISABLED_ENV];
});
it('never blocks a request while disabled', () => {
process.env[GUARD_DISABLED_ENV] = '1';
failRequest(30_000);
failRequest(30_000);
failRequest(30_000);
expect(guard.check(HOST).allowed).toBe(true);
expect(opened).toEqual([]);
});
it('accepts the "true" spelling', () => {
process.env[GUARD_DISABLED_ENV] = 'true';
failRequest(30_000);
failRequest(30_000);
expect(guard.check(HOST).allowed).toBe(true);
});
it('is read per call, so an already open breaker stops blocking', () => {
failRequest(30_000);
failRequest(30_000);
expectBlocked();
process.env[GUARD_DISABLED_ENV] = '1';
expect(guard.check(HOST).allowed).toBe(true);
});
});
});
describe('reportGuardedHostFailure', () => {
const ENDPOINT = 'http://panel.example.com:8080';
const URL_ON_ENDPOINT = `${ENDPOINT}/player_api.php`;
@@ -1,453 +1,25 @@
/**
* Per-host circuit breaker for portal requests.
* Main-process ownership of the shared host connectivity guard.
*
* A dead portal costs the full axios timeout on every call (30 s for Xtream,
* 15 s for Stalker), and browsing a dead portal's catalog issues dozens of
* those back to back — 30-second spinners and a flooded main-process log. Once
* a host has refused to answer twice in a row there is nothing left to learn
* from waiting again, so subsequent requests fail immediately for a short
* while.
*
* The rules are deliberately timid, because being wrong here means refusing to
* talk to a portal that works:
*
* - Only connection-level evidence counts (see
* {@link classifyHostRequestFailure}). Any HTTP response — 200, 404, even
* 502 — proves the host is alive and clears the record.
* - The window is short (30 s), so a mistake costs one page of navigation, and
* a host that came back is retried on its own.
* - Half-open means exactly ONE request goes out, not a whole screenful.
* - An explicit reset (user retry, endpoint discovery) always wins, and
* failures from requests that started before it are discarded — otherwise a
* 30-second straggler settles right after the reset and re-opens the breaker
* underneath the retry that cleared it.
*
* Deliberately not persisted: process lifetime is the right scope for
* "unreachable right now".
* The breaker itself lives in `@iptvnator/shared/host-health` so the web
* backend can run the same rules; what stays here is the part that is
* genuinely main-process-specific: a single guard for the whole process (both
* portal IPC handlers must see each other's evidence, or a dead endpoint would
* be relearned once per protocol) and the console warning that announces it.
*/
import { buildHostConnectivityFastFailMessage } from '@iptvnator/shared/interfaces';
import {
classifyHostRequestFailure,
failedAfterRedirect,
HostConnectivityGuard,
HostConnectivityGuardError,
HostRequestToken,
OPEN_DURATION_MS,
portalEndpointKeyOf,
} from '@iptvnator/shared/host-health';
/** Consecutive connection failures that trip the breaker. */
const FAILURE_THRESHOLD = 2;
/** How long requests fast-fail once the breaker is open. */
const OPEN_DURATION_MS = 30_000;
/**
* Failures further apart than this are unrelated, not a streak. Two sequential
* 30 s timeouts must fit inside it, hence comfortably above 60 s.
*/
const FAILURE_WINDOW_MS = 120_000;
/**
* Safety net for a half-open trial that never reports back. Above the longest
* request timeout (30 s) plus margin, so it only fires if a caller leaked the
* token — without it a lost report would keep the breaker open forever.
*/
const TRIAL_TIMEOUT_MS = 45_000;
/** Hosts tracked at once; portals per user are few, this is just a bound. */
const MAX_TRACKED_HOSTS = 256;
/** Idle records are forgotten; they hold nothing worth remembering. */
const IDLE_TTL_MS = 600_000;
const GUARD_DISABLED_ENV = 'IPTVNATOR_DISABLE_CONNECTIVITY_GUARD';
/**
* Error codes that prove the host itself did not answer.
*
* `ECONNRESET` is deliberately absent: a reset mid-transfer happens on hosts
* that are very much alive, and it is the one code a working stream can emit.
*/
const HOST_LEVEL_FAILURE_CODES = new Set([
'ECONNABORTED',
'ECONNREFUSED',
'EAI_AGAIN',
'EHOSTUNREACH',
'ENETUNREACH',
'ENOTFOUND',
'ETIMEDOUT',
]);
/**
* Thrown instead of making a request while the breaker is open.
*
* A real `Error`, because Electron serializes a rejected plain object to
* `[object Object]` and the renderer's classification would be lost. It
* carries NO `status` property on purpose — `getStalkerRequestErrorStatus`
* reads that field first, and a numeric status there would read as "the
* endpoint answered".
*/
export class HostConnectivityGuardError extends Error {
/** Scheme, host and port — see {@link portalEndpointKeyOf}. */
readonly endpoint: string;
constructor(endpoint: string) {
super(buildHostConnectivityFastFailMessage(endpoint));
this.name = 'HostConnectivityGuardError';
this.endpoint = endpoint;
}
}
/**
* Handed out by {@link HostConnectivityGuard.check} and passed back when the
* request settles. It records which attempt the report belongs to, so a reset
* or a parallel sibling cannot be mistaken for fresh evidence.
*/
export interface HostRequestToken {
/** Scheme, host and port — see {@link portalEndpointKeyOf}. */
readonly endpoint: string;
readonly epoch: number;
readonly startedAt: number;
/** Whether this request is the single probe allowed while half-open. */
readonly trial: boolean;
/**
* Which half-open slot this request holds, when it holds one.
*
* A trial can outlive its own window — `requestWithValidatedRedirects`
* gives each of up to five redirect hops its own 30 s budget — and once a
* replacement has been admitted, the abandoned request must not be able to
* free the replacement's slot. `trial: true` alone cannot tell the two
* apart, so the owner is identified explicitly.
*/
readonly trialId: number;
}
export type HostConnectivityCheck =
| { readonly allowed: true; readonly token: HostRequestToken }
| { readonly allowed: false; readonly retryAfterMs: number };
export type HostRequestOutcome = 'responded' | 'host-level' | 'inconclusive';
interface HostState {
consecutiveFailures: number;
/** When the last COUNTED failure was recorded. */
lastFailureAt: number;
openUntil: number;
trialStartedAt: number | null;
/** Monotonic id of the half-open slot; see `HostRequestToken.trialId`. */
trialId: number;
epoch: number;
lastTouchedAt: number;
}
export interface HostConnectivityGuardOptions {
now?: () => number;
onOpen?: (endpoint: string) => void;
}
/**
* Whether the guard is switched off. Read at call time, not at module load, so
* a debugging session can toggle it (matches `IPTVNATOR_ALLOW_INSECURE_TLS`).
*/
export function isHostConnectivityGuardDisabled(): boolean {
const value = process.env[GUARD_DISABLED_ENV]?.trim().toLowerCase();
return value === '1' || value === 'true';
}
/**
* What a failed request proves about the host.
*
* `responded` covers an error that still carries an HTTP response (5xx reaches
* the handlers as a rejection because `validateStatus` only tolerates <500) —
* the host is alive, so it clears the record. `inconclusive` covers everything
* that says nothing about reachability: a cancelled request, a URL rejected by
* the SSRF policy, a bug in our own code.
*/
export function classifyHostRequestFailure(error: unknown): HostRequestOutcome {
if (!error || typeof error !== 'object') {
return 'inconclusive';
}
if ((error as { response?: unknown }).response) {
return 'responded';
}
const code = (error as { code?: unknown }).code;
if (typeof code === 'string' && HOST_LEVEL_FAILURE_CODES.has(code)) {
return 'host-level';
}
return 'inconclusive';
}
/**
* The guard key for a request URL: its ORIGIN, i.e. scheme, host and port.
*
* Not `URL.host`, which omits a default port and therefore gives
* `http://panel.example` and `https://panel.example` the same key — two
* genuinely different endpoints, and a panel whose TLS listener is broken while
* plain HTTP works is a routine IPTV setup. Sharing one record there would let
* the dead one fast-fail the working one without ever contacting it.
*
* `URL.origin` also leaves out any `user:pass@` userinfo, so no credential
* reaches the key or the log line.
*/
export function portalEndpointKeyOf(url: string): string | null {
try {
const origin = new URL(url).origin;
// Opaque origins serialize as "null" and are not a usable key.
return origin && origin !== 'null' ? origin : null;
} catch {
return null;
}
}
export class HostConnectivityGuard {
private readonly states = new Map<string, HostState>();
private readonly now: () => number;
private readonly onOpen?: (endpoint: string) => void;
constructor(options: HostConnectivityGuardOptions = {}) {
this.now = options.now ?? (() => Date.now());
this.onOpen = options.onOpen;
}
/**
* Whether a request to `endpoint` may go out, and the token to report with.
*
* Fast-fails while the breaker is open, and while a half-open trial is
* still in flight — the point of the trial is that a screenful of requests
* does not all hang again to learn the same thing.
*/
check(endpoint: string): HostConnectivityCheck {
const now = this.now();
if (isHostConnectivityGuardDisabled()) {
return {
allowed: true,
token: {
endpoint,
epoch: 0,
startedAt: now,
trial: false,
trialId: 0,
},
};
}
const state = this.ensureState(endpoint, now);
state.lastTouchedAt = now;
if (state.openUntil > now) {
return { allowed: false, retryAfterMs: state.openUntil - now };
}
// Never opened, or the open window elapsed. The latter is half-open:
// let one request through and hold the rest back until it settles.
let trial = false;
if (state.openUntil > 0) {
const trialStale =
state.trialStartedAt !== null &&
now - state.trialStartedAt >= TRIAL_TIMEOUT_MS;
if (state.trialStartedAt !== null && !trialStale) {
return { allowed: false, retryAfterMs: 0 };
}
state.trialStartedAt = now;
state.trialId += 1;
trial = true;
}
return {
allowed: true,
token: {
endpoint,
epoch: state.epoch,
startedAt: now,
trial,
trialId: trial ? state.trialId : 0,
},
};
}
/**
* The host answered. Always clears the record, whatever the status was and
* whichever attempt it belonged to — a reachable host is a reachable host.
*/
reportSuccess(token: HostRequestToken): void {
const state = this.states.get(token.endpoint);
if (!state || isHostConnectivityGuardDisabled()) {
return;
}
state.consecutiveFailures = 0;
state.lastFailureAt = 0;
state.openUntil = 0;
state.trialStartedAt = null;
state.lastTouchedAt = this.now();
}
/** The host did not answer at all. */
reportFailure(token: HostRequestToken): void {
const state = this.states.get(token.endpoint);
if (!state || isHostConnectivityGuardDisabled()) {
return;
}
const now = this.now();
// Superseded by an explicit reset: the user asked for a fresh attempt
// and this verdict predates it.
if (state.epoch !== token.epoch) {
this.releaseTrial(state, token);
return;
}
// Only the request that still holds the slot ends the half-open state.
// An abandoned trial's late failure is ordinary evidence — it goes
// through the streak rules below and leaves the replacement alone.
const wasTrial = this.ownsTrial(state, token);
if (wasTrial) {
state.trialStartedAt = null;
}
state.lastTouchedAt = now;
if (
state.consecutiveFailures > 0 &&
now - state.lastFailureAt > FAILURE_WINDOW_MS
) {
state.consecutiveFailures = 0;
}
// A request already in flight when the previous failure was recorded is
// not the next link in a streak — it is a sibling of it. Catalog
// loading fans out several requests at once, and one hiccup failing all
// of them is one piece of evidence, not a trip. A request that started
// at or after that moment is a genuine new attempt.
let counted = false;
if (
state.consecutiveFailures === 0 ||
token.startedAt >= state.lastFailureAt
) {
state.consecutiveFailures += 1;
state.lastFailureAt = now;
counted = true;
}
// Only a failure this report actually counted may trip the threshold.
// A sibling arriving after the window elapsed would otherwise open a
// fresh one off the existing count, pushing the half-open trial past
// the intended cooldown. A failed trial still goes straight back to
// open: the host had its chance.
if (
wasTrial ||
(counted && state.consecutiveFailures >= FAILURE_THRESHOLD)
) {
const wasOpen = state.openUntil > now;
state.openUntil = now + OPEN_DURATION_MS;
if (!wasOpen) {
this.onOpen?.(token.endpoint);
}
}
}
/**
* The request failed for a reason that says nothing about the host. Only
* releases the half-open slot, so the next request can be the trial.
*/
reportInconclusive(token: HostRequestToken): void {
const state = this.states.get(token.endpoint);
if (!state || isHostConnectivityGuardDisabled()) {
return;
}
this.releaseTrial(state, token);
state.lastTouchedAt = this.now();
}
/**
* Forgets everything recorded for `endpoint`, and invalidates the reports of
* requests already in flight. Called when the user asks for a real attempt
* or when the URL may now point somewhere else.
*/
reset(endpoint: string): void {
const now = this.now();
const state = this.ensureState(endpoint, now);
state.consecutiveFailures = 0;
state.lastFailureAt = 0;
state.openUntil = 0;
state.trialStartedAt = null;
state.epoch += 1;
state.lastTouchedAt = now;
}
/**
* A token for a request that must NOT be policed or counted, but whose
* success still clears the record.
*
* Stalker endpoint discovery probes several candidate paths on one host and
* expects most of them to fail; counting those would let it declare a
* slow-but-alive portal unreachable. It does not take the half-open slot
* either — a probe is not the trial the guard is waiting for.
*/
observe(endpoint: string): HostRequestToken {
return {
endpoint,
epoch: this.states.get(endpoint)?.epoch ?? 0,
startedAt: this.now(),
trial: false,
trialId: 0,
};
}
/** Test seam: drops all recorded state. */
clear(): void {
this.states.clear();
}
/** Whether `token` still holds the current half-open slot. */
private ownsTrial(state: HostState, token: HostRequestToken): boolean {
return (
token.trial &&
state.trialStartedAt !== null &&
state.epoch === token.epoch &&
state.trialId === token.trialId
);
}
private releaseTrial(state: HostState, token: HostRequestToken): void {
if (this.ownsTrial(state, token)) {
state.trialStartedAt = null;
}
}
private ensureState(endpoint: string, now: number): HostState {
const existing = this.states.get(endpoint);
if (existing) {
return existing;
}
this.prune(now);
const state: HostState = {
consecutiveFailures: 0,
lastFailureAt: 0,
openUntil: 0,
trialStartedAt: null,
trialId: 0,
epoch: 0,
lastTouchedAt: now,
};
this.states.set(endpoint, state);
return state;
}
private prune(now: number): void {
for (const [endpoint, state] of this.states) {
if (
now - state.lastTouchedAt > IDLE_TTL_MS &&
state.openUntil <= now &&
state.trialStartedAt === null
) {
this.states.delete(endpoint);
}
}
// Still at the cap: forget the oldest records. Dropping one means
// contacting that host again, which is the safe direction to err in.
while (this.states.size >= MAX_TRACKED_HOSTS) {
const oldest = this.states.keys().next();
if (oldest.done) {
return;
}
this.states.delete(oldest.value);
}
}
}
export { HostConnectivityGuardError };
export type { HostRequestToken };
let sharedGuard: HostConnectivityGuard | null = null;
@@ -506,66 +78,6 @@ export function reportGuardedHostSuccess(token: HostRequestToken | null): void {
getHostConnectivityGuard().reportSuccess(token);
}
}
/**
* The URL a failed request was actually talking to, when the error says.
*
* Redirects are followed hop by hop, each with its own config, so a failure on
* a later hop carries THAT hop's URL rather than the one we asked for.
*/
function failedRequestUrlOf(error: unknown): string | null {
const url = (error as { config?: { url?: unknown } } | null)?.config?.url;
return typeof url === 'string' ? url : null;
}
function normalizedUrlOrNull(url: string): string | null {
try {
return new URL(url).toString();
} catch {
return null;
}
}
/**
* Whether the failure happened on a hop the guarded endpoint redirected us to.
*
* Reaching any later hop proves the guarded endpoint answered: the first hop is
* always the URL we asked for, and only a redirect status advances the chain.
* That holds for a same-origin redirect too, so comparing the whole URL — not
* just its origin — is what catches `/player_api.php` → `/slow/player_api.php`.
*
* Requires positive evidence: anything unparseable or unknown returns false and
* the failure is counted as usual, because guessing "redirect" here would stop
* the guard from ever tripping.
*/
function failedAfterRedirect(
error: unknown,
token: HostRequestToken,
requestUrl: string | undefined
): boolean {
const failedUrl = failedRequestUrlOf(error);
if (!failedUrl) {
return false;
}
const failedEndpoint = portalEndpointKeyOf(failedUrl);
if (failedEndpoint && failedEndpoint !== token.endpoint) {
return true;
}
if (!requestUrl) {
return false;
}
const failedNormalized = normalizedUrlOrNull(failedUrl);
const requestedNormalized = normalizedUrlOrNull(requestUrl);
return (
failedNormalized !== null &&
requestedNormalized !== null &&
failedNormalized !== requestedNormalized
);
}
/**
* Records what a failed request proved about its endpoint.
*