From e94cc029eb61cd07ceb6052bcc7dd7402b3e0608 Mon Sep 17 00:00:00 2001 From: 4gray <4gray@users.noreply.github.com> Date: Thu, 13 Aug 2026 22:27:05 +0200 Subject: [PATCH] fix(portals): time out and fast-fail PWA proxy requests to dead hosts (#1424) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * refactor(portals): hoist the connectivity guard into libs/shared/host-health The breaker was written dependency-free so both processes that talk to portals could share it. Move the part that has no Electron in it — the state machine, the failure classification, the redirect-attribution helpers and the fast-fail error — into `@iptvnator/shared/host-health` (`scope:shared` / `domain:shared-runtime` / `type:util`). What stays in `apps/electron-backend` is the genuinely main-process part: one guard for the whole process, so both portal IPC handlers see each other's evidence, and the console warning that announces it. Every call site is unchanged; the wrapper re-exports the two types they import. The spec splits the same way — the state machine moves with the class, the singleton and its redirect attribution stay with the wrapper. Register the new project in the coverage policy, which every project with a test target must declare a tier for. Co-Authored-By: Claude Opus 5 * fix(portals): time out and fast-fail PWA proxy requests to dead hosts The web backend's proxy routes were bare `axios.get()` calls with no `timeout`, so a provider that accepted a connection and then went silent held the request until the OS gave up on the TCP connection — minutes, rather than the 15/30 s budget the Electron handlers use. Add the same per-route timeouts (Xtream 30 s, Stalker 15 s / 30 s for `create_link`, playlist and XMLTV 30 s). Those numbers are safe for large downloads: on axios' default transport `timeout` bounds the time to response headers and then continues as the socket's inactivity timeout, so a multi-megabyte XMLTV file that keeps delivering bytes is never cut off — only a stalled one is. With requests bounded, run `/xtream` and `/stalker` through the shared breaker, injected via `WebBackendAppOptions.hostGuard` so specs drive it with a clock they own. Playlist and XMLTV downloads keep the timeout but no breaker, matching Electron: a download is one request rather than a catalog fan-out, it is usually the direct result of the user asking for it, and it can outlive the half-open trial window. The breaker is checked before the Xtream URL revalidation, which resolves the hostname — a dead host is where DNS is slow too, and a request admitted and then abandoned by the URL policy hands its token back rather than holding the trial slot. `resetHostConnectivityGuard()` no longer no-ops in the PWA: the breaker lives in the backend process, so it travels to a new `POST /connectivity-guard/reset`, which reads only the origin and never logs the credential-bearing URL. `skipConnectionGuard` now survives the PWA transport too, so Stalker endpoint discovery keeps the exemption it has on the desktop instead of tripping the breaker with its own probes. A fast-fail keeps each route's HTTP 200 `{message, status}` envelope. The Stalker path needs one extra step: `forwardStalkerRequest` turns that envelope into `HTTP Error : …` with a numeric `status`, and the renderer reads both as "the endpoint answered" — which would make discovery walk every candidate and fire lazy repair at a host just declared dead. A prior branch keyed on the shared `isHostConnectivityFastFailMessage` rethrows it bare instead. Co-Authored-By: Claude Opus 5 * fix(portals): report exempt Stalker probe responses to the PWA breaker Flagged by the author of #1421 as one of the twelve fixes that landed there after this branch cherry-picked the pre-review commit: an exempt discovery probe must still REPORT, it just must not COUNT. The web backend was skipping the report entirely for a probe, which loses the case that matters. A failure carrying an HTTP response proves the endpoint answered, and this route sets no `validateStatus`, so axios rejects every non-2xx with `error.response` attached — a probe answered with 404 or 500 was therefore dropped instead of clearing the record. Two counted failures either side of it then read as consecutive and opened the breaker in the middle of discovery, which is exactly what the exemption exists to prevent. `reportProviderRequestFailure` now takes `countFailures`, matching the Electron reporter, and both routes always report. Co-Authored-By: Claude Opus 5 * fix(portals): read the redirect hop from the transport that followed it `failedAfterRedirect` decided whether a failure belonged to a redirect destination by reading `error.config.url`. That is right for the Electron transport, which sets `maxRedirects: 0` and reissues every hop as its own request, so the hop IS the config URL. It is blind on the web backend, which uses axios' default transport: follow-redirects walks the chain inside one request and `config` is built once, so `config.url` stays the URL we asked for. Verified against the installed axios 1.19.0 with a live server that 302s to a dead port: asked for : http://127.0.0.1:63953/player_api.php config.url : http://127.0.0.1:63953/player_api.php request._currentUrl: http://127.0.0.1:1/dead So the comparison was original-vs-original, found no redirect, and charged two dead destinations to the provider that had answered both times with a 302 — then fast-failed it. Read `request._currentUrl` first and fall back to `config.url`, which covers both transports; Electron's native per-hop requests expose no `_currentUrl` and are unaffected. Also check redirect attribution BEFORE suppressing failure counting for an exempt probe. A 3xx from the guarded endpoint is an answer, so a probe that observed one must clear the record; otherwise a timeout, a probe redirected to a dead destination, and another timeout still read as two consecutive failures. Co-Authored-By: Claude Opus 5 * fix(portals): stop axios query params reading as a redirect Codex found that the Xtream breaker never opened at all, and it was right. The web backend passes credentials and the action through axios' `params`, so axios sends `…/player_api.php?username=…&action=…` while the baseline handed to `failedAfterRedirect` is the query-less URL the route built. Verified against axios 1.19.0 with a plain ECONNREFUSED and no redirect anywhere in sight: baseline : http://127.0.0.1:1/player_api.php request._currentUrl: http://127.0.0.1:1/player_api.php?username=demo&… The two normalized URLs differ, so every ordinary failure looked like a post-redirect failure, credited the endpoint, and the breaker could never trip. Compare origin and path, not the whole URL. That keeps what the check is for — an endpoint that answered and sent us elsewhere, including the same-origin `/player_api.php` → `/slow/player_api.php` case — and gives up only a redirect that changes nothing but the query, which is then counted as an ordinary failure. Erring towards counting is the safe direction here. The reason 57 tests passed over a dead feature is the real lesson: `StubHttpClient` threw bare `Error`s, so the guard's redirect check saw neither `config.url` nor `request._currentUrl` and quietly did nothing. The stub now shapes its rejections like axios does, including the query axios appends. With that alone, four existing tests fail against the old comparison. Co-Authored-By: Claude Opus 5 * fix(portals): count a hostname that stops resolving, and fix a stale docblock Two from review, one behavioural and one documentation. A name that will not resolve is the host failing to answer — the same evidence as the ENOTFOUND the transport would have raised a moment later. But the SSRF validation turns a lookup failure into a 400 "host could not be resolved", and the release path added earlier handed the token back as inconclusive, so the breaker could never open for a host whose DNS died and every request kept paying for the same dead lookup. `ProviderUrlError` now carries the underlying lookup error internally. A refusal that has one is counted; a genuine policy refusal — private address, bad scheme, credentials in the URL — still only releases the half-open slot, because that says nothing about reachability. The field is internal: `providerUrlErrorBody()` strips it at both call sites, so the client sees exactly the body it saw before, which the test asserts. The docblock on `resetHostConnectivityGuard` still said the PWA channel is unknown and the call no-ops. That stopped being true when this branch implemented `CONNECTIVITY_GUARD_RESET` over HTTP, and a stale contract there is how the next caller silently skips the PWA path. (The edit was in an earlier commit and was lost when the branch was rebuilt on master.) Co-Authored-By: Claude Opus 5 * docs(portals): record the transport-specific redirect contract The redirect section still described one transport: hop-by-hop requests, `error.config.url`, whole-URL comparison. Two of those three are now wrong for the web backend, and this document is the canonical contract — leaving it stale is how the attribution bugs fixed in the last two commits get reintroduced. Says what is actually true: which field holds the failed hop on each transport and why the helper reads both, and that the comparison is origin + path because the web backend's credentials ride in axios' `params` and a whole-URL comparison therefore reported a redirect for every ordinary failure. Co-Authored-By: Claude Opus 5 * docs(portals): scope the guard summary to both processes The opening line still defined the breaker as an Electron main-process concern, which contradicted the ownership section below it and is the part a reader skims to decide whether the document applies to them. Names both processes, and records that the web backend had the worse version of the problem first — no request timeout at all — since that is why the timeouts and the breaker had to land there together. Co-Authored-By: Claude Opus 5 * fix(portals): declare the shared-interfaces dependency of host-health The new library's manifest listed only `tslib`, but its emitted JavaScript does `require('@iptvnator/shared/interfaces')` — the guard builds its fast-fail message with `buildHostConnectivityFastFailMessage`. Anything resolving the built artifact from its own manifest would have failed with MODULE_NOT_FOUND. The manifest was copied from `shared/logging`, which imports nothing across libraries and therefore needs nothing beyond `tslib`. `shared/m3u-utils` is the right precedent: it imports the same library and declares `"@iptvnator/shared/interfaces": "0.0.1"`. Verified against the build output rather than by inspection — the emitted `host-connectivity-guard.js` requires the module, and the generated `dist/libs/shared/host-health/package.json` now declares it. Co-Authored-By: Claude Opus 5 * fix(portals): bound the PWA connectivity-guard reset The reset was a bare `fetch` with no timeout, which is the exact failure this change exists to remove, reintroduced one layer up. Every caller awaits the reset BEFORE issuing the request it is clearing the way for — `retryContentInitialization` awaits it first by design — so a backend or reverse proxy that accepts the POST and then goes quiet would leave Retry doing nothing at all, for as long as the socket stayed open. Bound it with an AbortController and a 5 s timer. The abort rejects, `resetHostConnectivityGuard` swallows it as it already does for any other failure, and the caller proceeds to its real request — which is what "best effort" was supposed to mean. The timer is cleared in a `finally`, and it covers the body read as well as the headers. Five seconds because this talks to the user's own backend rather than a provider: it should answer immediately, and a slow one must not hold up the retry that asked for it. Co-Authored-By: Claude Opus 5 --------- Co-authored-by: Claude Opus 5 --- .changes/portals-pwa-request-timeouts.md | 10 + CLAUDE.md | 5 +- .../app/util/host-connectivity-guard.spec.ts | 496 +------------ .../src/app/util/host-connectivity-guard.ts | 522 +------------- apps/web-backend/src/app/host-guard.ts | 218 ++++++ .../app/web-backend-app.host-guard.spec.ts | 653 ++++++++++++++++++ .../src/app/web-backend-app.spec-helpers.ts | 154 +++++ .../src/app/web-backend-app.spec.ts | 134 +--- apps/web-backend/src/app/web-backend-app.ts | 210 +++++- apps/web/src/app/services/pwa.service.spec.ts | 193 ++++++ apps/web/src/app/services/pwa.service.ts | 99 +++ docs/architecture/host-connectivity-guard.md | 158 ++++- .../src/lib/host-connectivity-reset.ts | 8 +- libs/shared/host-health/jest.config.ts | 13 + libs/shared/host-health/package.json | 12 + libs/shared/host-health/project.json | 30 + libs/shared/host-health/src/index.ts | 1 + .../src/lib/host-connectivity-guard.spec.ts | 540 +++++++++++++++ .../src/lib/host-connectivity-guard.ts | 556 +++++++++++++++ libs/shared/host-health/tsconfig.json | 19 + libs/shared/host-health/tsconfig.lib.json | 10 + libs/shared/host-health/tsconfig.spec.json | 15 + tools/coverage/coverage-policy.json | 7 + tsconfig.base.json | 3 + 24 files changed, 2895 insertions(+), 1171 deletions(-) create mode 100644 .changes/portals-pwa-request-timeouts.md create mode 100644 apps/web-backend/src/app/host-guard.ts create mode 100644 apps/web-backend/src/app/web-backend-app.host-guard.spec.ts create mode 100644 apps/web-backend/src/app/web-backend-app.spec-helpers.ts create mode 100644 libs/shared/host-health/jest.config.ts create mode 100644 libs/shared/host-health/package.json create mode 100644 libs/shared/host-health/project.json create mode 100644 libs/shared/host-health/src/index.ts create mode 100644 libs/shared/host-health/src/lib/host-connectivity-guard.spec.ts create mode 100644 libs/shared/host-health/src/lib/host-connectivity-guard.ts create mode 100644 libs/shared/host-health/tsconfig.json create mode 100644 libs/shared/host-health/tsconfig.lib.json create mode 100644 libs/shared/host-health/tsconfig.spec.json diff --git a/.changes/portals-pwa-request-timeouts.md b/.changes/portals-pwa-request-timeouts.md new file mode 100644 index 000000000..c1e69e22b --- /dev/null +++ b/.changes/portals-pwa-request-timeouts.md @@ -0,0 +1,10 @@ +--- +type: fix +area: portals +--- + +In the self-hosted web version, a provider that accepted a connection and then +went silent could hang a request until the operating system gave up, with no +way to cancel it. Requests now time out on the same schedule as the desktop app, +and a host that fails to answer twice is skipped for a short while instead of +stalling every following request. diff --git a/CLAUDE.md b/CLAUDE.md index c2f1fd3e4..8dbcf27f7 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -290,7 +290,7 @@ This is an Nx monorepo with the following structure: - **apps/web** - Angular application (frontend, shared by Electron and PWA) - **apps/electron-backend** - Electron main process -- **apps/web-backend** - HTTP backend for the self-hosted PWA (`/parse`, `/parse-xml`, `/xtream`, `/stalker` CORS proxy endpoints). At startup it raises Node's happy-eyeballs per-attempt connection timeout to 2500 ms (`network-family-autoselection.ts`) so dual-stack provider hostnames fall back to IPv4 behind IPv6-less VPN/Docker networks; an explicit `--network-family-autoselection-attempt-timeout` passed via `NODE_OPTIONS`/CLI always wins. Outbound provider failures are logged hostname-only with the underlying Node error codes and return the primary code in the error body (`provider-error.ts`) — the proxied URL query carries credentials and must never be logged +- **apps/web-backend** - HTTP backend for the self-hosted PWA (`/parse`, `/parse-xml`, `/xtream`, `/stalker` CORS proxy endpoints). At startup it raises Node's happy-eyeballs per-attempt connection timeout to 2500 ms (`network-family-autoselection.ts`) so dual-stack provider hostnames fall back to IPv4 behind IPv6-less VPN/Docker networks; an explicit `--network-family-autoselection-attempt-timeout` passed via `NODE_OPTIONS`/CLI always wins. Outbound provider failures are logged hostname-only with the underlying Node error codes and return the primary code in the error body (`provider-error.ts`) — the proxied URL query carries credentials and must never be logged. Every proxied request carries the same timeout as its Electron counterpart (Xtream 30 s, Stalker 15 s / 30 s for `create_link`, playlist and XMLTV 30 s). The shared per-host circuit breaker (`host-guard.ts`, injected via `WebBackendAppOptions.hostGuard`) covers `/xtream` and `/stalker` only — playlist/XMLTV downloads keep the timeout but no breaker, matching Electron. A fast-fail keeps the route's normal failure shape (HTTP 200 with a `{message, status}` body), `skipConnectionGuard=true` carries the Stalker discovery exemption through the proxy, and `POST /connectivity-guard/reset` is the PWA's counterpart to the `CONNECTIVITY_GUARD_RESET` IPC - **apps/remote-control-web** - Mobile remote-control web app served by the Electron backend - **apps/web-e2e** - Playwright E2E tests against the web app - **apps/electron-backend-e2e** - Playwright E2E tests against the Electron app @@ -311,6 +311,7 @@ This is an Nx monorepo with the following structure: - **services** - Abstract DataService contract and shared app services (incl. the TMDB metadata enrichment module in `lib/tmdb/`) - **shared/interfaces** - TypeScript interfaces and types (incl. `ElectronBridgeApi`) - **shared/logging** - Dependency-free structured redaction for diagnostic logs + - **shared/host-health** - Per-host circuit breaker for portal requests (`HostConnectivityGuard`), shared by the Electron main process and the web backend; transport-free, the owning app supplies the clock and owns the instance - **shared/database** - Canonical Drizzle schema and DB connection (used by the Electron backend) - **shared/m3u-utils** - M3U playlist utilities - **shared/marketing-fixtures** - Provider-neutral fictional movie metadata shared by the Xtream and Stalker marketing mocks @@ -713,7 +714,7 @@ This project uses modern Angular signal-based APIs and patterns. **ALWAYS** use - `epg.events.ts` - EPG IPC registration; freshness/fetch orchestration lives in `epg-fetch.service.ts`, manual channel-mapping resolution and CRUD in `epg-mapping.service.ts`, worker lifecycle in `epg-worker.service.ts`, DB lookups in `epg-query.service.ts` - `xtream.events.ts` - Xtream Codes API - `stalker.events.ts` - Stalker portal API - - `connectivity-guard.events.ts` - `CONNECTIVITY_GUARD_RESET`: forgets the connection failures recorded for a portal host. Both portal handlers above run every request through the per-host circuit breaker in `util/host-connectivity-guard.ts` — after 2 consecutive connection-level failures (no HTTP response; `ETIMEDOUT`/`ENOTFOUND`/`ECONNREFUSED`/… but never `ECONNRESET`) requests to that endpoint fail immediately for 30 s. The key is `URL.origin`, not `URL.host`, which would give `http://panel` and `https://panel` one shared record and let a dead TLS listener fast-fail the working HTTP one instead of hanging the full 30 s/15 s axios timeout again, with one half-open trial request afterwards. Any HTTP response (4xx and 5xx included) clears the record. The refusal is a real `Error` whose wording is a renderer contract (`buildHostConnectivityFastFailMessage` in `libs/shared/interfaces`): it must carry no `HTTP Error `, no timeout wording and none of the auth phrases, or Stalker endpoint discovery misclassifies it and lazy portal repair fires against a host just declared dead. Discovery probes are exempt via the `skipConnectionGuard` payload flag (bypass + no failure counting, but successes still clear the record). Every user-driven retry/refresh that issues portal requests must reset BEFORE its first request, or the affordance fast-fails and looks broken; automatic and first-load paths deliberately do not reset. Current senders: Xtream content-gate Retry, Stalker catalog append retry (`retryContentPage`), Stalker search-page retry, `StalkerItvCacheService.refresh()` (Live TV refresh), both account-info dialogs' Retry, the destructive Xtream refresh (`XtreamRefreshFlowService`, before it deletes the cached catalog — one flow shared by both entry points, `PlaylistRefreshActionService.refreshXtream()` and `RecentPlaylistsComponent.refreshXtreamPlaylist()`, which supply only a progress reporter), `StalkerPortalDiscoveryService.discover()`, and `PortalStatusService` on `skipCache`. Kill switch: `IPTVNATOR_DISABLE_CONNECTIVITY_GUARD=1`. Contract: `docs/architecture/host-connectivity-guard.md` + - `connectivity-guard.events.ts` - `CONNECTIVITY_GUARD_RESET`: forgets the connection failures recorded for a portal host. Both portal handlers above run every request through the per-host circuit breaker (rules in `@iptvnator/shared/host-health`, process-wide instance in `util/host-connectivity-guard.ts`; the web backend runs the same breaker over its proxy routes) — after 2 consecutive connection-level failures (no HTTP response; `ETIMEDOUT`/`ENOTFOUND`/`ECONNREFUSED`/… but never `ECONNRESET`) requests to that endpoint fail immediately for 30 s. The key is `URL.origin`, not `URL.host`, which would give `http://panel` and `https://panel` one shared record and let a dead TLS listener fast-fail the working HTTP one instead of hanging the full 30 s/15 s axios timeout again, with one half-open trial request afterwards. Any HTTP response (4xx and 5xx included) clears the record. The refusal is a real `Error` whose wording is a renderer contract (`buildHostConnectivityFastFailMessage` in `libs/shared/interfaces`): it must carry no `HTTP Error `, no timeout wording and none of the auth phrases, or Stalker endpoint discovery misclassifies it and lazy portal repair fires against a host just declared dead. Discovery probes are exempt via the `skipConnectionGuard` payload flag (bypass + no failure counting, but successes still clear the record). Every user-driven retry/refresh that issues portal requests must reset BEFORE its first request, or the affordance fast-fails and looks broken; automatic and first-load paths deliberately do not reset. Current senders: Xtream content-gate Retry, Stalker catalog append retry (`retryContentPage`), Stalker search-page retry, `StalkerItvCacheService.refresh()` (Live TV refresh), both account-info dialogs' Retry, the destructive Xtream refresh (`XtreamRefreshFlowService`, before it deletes the cached catalog — one flow shared by both entry points, `PlaylistRefreshActionService.refreshXtream()` and `RecentPlaylistsComponent.refreshXtreamPlaylist()`, which supply only a progress reporter), `StalkerPortalDiscoveryService.discover()`, and `PortalStatusService` on `skipCache`. Kill switch: `IPTVNATOR_DISABLE_CONNECTIVITY_GUARD=1`. Contract: `docs/architecture/host-connectivity-guard.md` - `player.events.ts` - External player IPC registration; MPV/VLC lifecycle logic lives in `mpv-session.service.ts`, `vlc-session.service.ts`, and shared `external-player-*` helpers - `settings.events.ts` - App settings - `electron.events.ts` - App version, etc. diff --git a/apps/electron-backend/src/app/util/host-connectivity-guard.spec.ts b/apps/electron-backend/src/app/util/host-connectivity-guard.spec.ts index 3f61ff864..85181baa6 100644 --- a/apps/electron-backend/src/app/util/host-connectivity-guard.spec.ts +++ b/apps/electron-backend/src/app/util/host-connectivity-guard.spec.ts @@ -1,498 +1,16 @@ +/** + * Main-process wrapper behaviour: the shared singleton and what a failure + * reported through it proves about the endpoint. The state machine itself is + * covered in `@iptvnator/shared/host-health`. + */ + +import { HostConnectivityGuardError } from '@iptvnator/shared/host-health'; import { - buildHostConnectivityFastFailMessage, - isHostConnectivityFastFailMessage, -} from '@iptvnator/shared/interfaces'; -import { - HostConnectivityGuard, - HostConnectivityGuardError, - HostRequestToken, beginGuardedHostRequest, - classifyHostRequestFailure, - portalEndpointKeyOf, reportGuardedHostFailure, resetHostConnectivityGuardForTests, } from './host-connectivity-guard'; -const HOST = 'http://portal.example.com:8080'; -const GUARD_DISABLED_ENV = 'IPTVNATOR_DISABLE_CONNECTIVITY_GUARD'; -const OPEN_DURATION_MS = 30_000; -const FAILURE_WINDOW_MS = 120_000; -const TRIAL_TIMEOUT_MS = 45_000; - -function timeoutError(code = 'ETIMEDOUT') { - return Object.assign(new Error('timeout of 30000ms exceeded'), { code }); -} - -describe('classifyHostRequestFailure', () => { - it('treats connection-level error codes as host-level failures', () => { - for (const code of [ - 'ECONNABORTED', - 'ECONNREFUSED', - 'EAI_AGAIN', - 'EHOSTUNREACH', - 'ENETUNREACH', - 'ENOTFOUND', - 'ETIMEDOUT', - ]) { - expect(classifyHostRequestFailure(timeoutError(code))).toBe( - 'host-level' - ); - } - }); - - it('treats an error carrying an HTTP response as proof the host answered', () => { - // 5xx reaches the handlers as a rejection: validateStatus only - // tolerates < 500. The host still answered. - const serverError = Object.assign(new Error('Request failed'), { - code: 'ERR_BAD_RESPONSE', - response: { status: 502, statusText: 'Bad Gateway' }, - }); - - expect(classifyHostRequestFailure(serverError)).toBe('responded'); - }); - - it('does not read reachability into cancellations, SSRF refusals or plain errors', () => { - expect( - classifyHostRequestFailure( - Object.assign(new Error('canceled'), { code: 'ERR_CANCELED' }) - ) - ).toBe('inconclusive'); - expect( - classifyHostRequestFailure( - new Error('URL host could not be resolved') - ) - ).toBe('inconclusive'); - expect(classifyHostRequestFailure(undefined)).toBe('inconclusive'); - expect(classifyHostRequestFailure('ETIMEDOUT')).toBe('inconclusive'); - }); - - it('never counts a mid-transfer connection reset', () => { - // A reset happens on hosts that are very much alive. - expect(classifyHostRequestFailure(timeoutError('ECONNRESET'))).toBe( - 'inconclusive' - ); - }); -}); - -describe('portalEndpointKeyOf', () => { - it('keys on host and port so two panels on one machine stay separate', () => { - expect( - portalEndpointKeyOf('http://example.com:8080/player_api.php') - ).toBe('http://example.com:8080'); - expect( - portalEndpointKeyOf('http://example.com:9090/player_api.php') - ).toBe('http://example.com:9090'); - }); - - it('separates HTTP from HTTPS on their default ports', () => { - // `URL.host` omits a default port, so both would collapse onto - // `example.com` — and a panel whose TLS listener is broken while plain - // HTTP works is a routine IPTV setup. - expect(portalEndpointKeyOf('http://example.com/player_api.php')).toBe( - 'http://example.com' - ); - expect(portalEndpointKeyOf('https://example.com/player_api.php')).toBe( - 'https://example.com' - ); - }); - - it('leaves URL credentials out of the key', () => { - expect(portalEndpointKeyOf('http://user:pass@example.com/c')).toBe( - 'http://example.com' - ); - }); - - it('returns null for an unparseable URL instead of throwing', () => { - expect(portalEndpointKeyOf('not a url')).toBeNull(); - }); -}); - -describe('HostConnectivityGuardError', () => { - it('carries the shared fast-fail message and no status field', () => { - const error = new HostConnectivityGuardError(HOST); - - expect(error).toBeInstanceOf(Error); - expect(error.message).toBe(buildHostConnectivityFastFailMessage(HOST)); - // getStalkerRequestErrorStatus reads `status` first; a number there - // would read as "the endpoint answered". - expect( - (error as unknown as { status?: unknown }).status - ).toBeUndefined(); - expect(isHostConnectivityFastFailMessage(error.message)).toBe(true); - }); -}); - -describe('HostConnectivityGuard', () => { - let clock = 1_000_000; - let opened: string[]; - let guard: HostConnectivityGuard; - - const advance = (ms: number) => { - clock += ms; - }; - - /** Runs a request that fails at the host level after `durationMs`. */ - const failRequest = (durationMs = 0): void => { - const check = guard.check(HOST); - if (!check.allowed) { - throw new Error('expected the request to be allowed'); - } - advance(durationMs); - guard.reportFailure(check.token); - }; - - const expectAllowed = (): HostRequestToken => { - const check = guard.check(HOST); - expect(check.allowed).toBe(true); - if (!check.allowed) { - throw new Error('unreachable'); - } - return check.token; - }; - - const expectBlocked = (): void => { - expect(guard.check(HOST).allowed).toBe(false); - }; - - beforeEach(() => { - clock = 1_000_000; - opened = []; - delete process.env[GUARD_DISABLED_ENV]; - guard = new HostConnectivityGuard({ - now: () => clock, - onOpen: (host) => opened.push(host), - }); - }); - - it('allows requests to a host it has never seen', () => { - expect(guard.check(HOST).allowed).toBe(true); - expect(opened).toEqual([]); - }); - - it('still allows the request after a single failure', () => { - failRequest(30_000); - - expect(guard.check(HOST).allowed).toBe(true); - expect(opened).toEqual([]); - }); - - it('opens after two consecutive failures and reports the wait', () => { - failRequest(30_000); - failRequest(30_000); - - expect(guard.check(HOST)).toEqual({ - allowed: false, - retryAfterMs: OPEN_DURATION_MS, - }); - expect(opened).toEqual([HOST]); - }); - - it('logs the open transition once, not per blocked request', () => { - failRequest(30_000); - failRequest(30_000); - expectBlocked(); - expectBlocked(); - - expect(opened).toEqual([HOST]); - }); - - it('keeps other hosts untouched', () => { - failRequest(30_000); - failRequest(30_000); - - expect(guard.check('http://other.example.com').allowed).toBe(true); - }); - - it('keeps the same host on another scheme reachable', () => { - // Regression: keying by `URL.host` dropped the default port, so a dead - // HTTPS panel fast-failed the working HTTP one on the same machine. - failRequest(30_000); - failRequest(30_000); - expectBlocked(); - - expect(guard.check('https://portal.example.com:8080').allowed).toBe( - true - ); - }); - - it('counts a parallel fan-out that fails together as one failure', () => { - // Catalog init loads live/vod/series at once. One wifi hiccup failing - // all three is one piece of evidence, not a trip. - const first = expectAllowed(); - const second = expectAllowed(); - const third = expectAllowed(); - - advance(30_000); - guard.reportFailure(first); - guard.reportFailure(second); - guard.reportFailure(third); - - expect(guard.check(HOST).allowed).toBe(true); - expect(opened).toEqual([]); - }); - - it('opens when a second fan-out fails after the first one did', () => { - const first = expectAllowed(); - const second = expectAllowed(); - advance(30_000); - guard.reportFailure(first); - guard.reportFailure(second); - - failRequest(30_000); - - expectBlocked(); - expect(opened).toEqual([HOST]); - }); - - it('does not re-open on a sibling that settles after the window elapsed', () => { - // A and B start together; A plus a later request open the breaker. B is - // explicitly not counted as a strike, so it must not start a fresh - // 30-second window either — that would push the half-open trial past - // the intended cooldown. - const sibling = expectAllowed(); - const first = expectAllowed(); - advance(30_000); - guard.reportFailure(first); - failRequest(30_000); - expectBlocked(); - expect(opened).toEqual([HOST]); - - advance(OPEN_DURATION_MS); - guard.reportFailure(sibling); - - const trial = expectAllowed(); - expect(trial.trial).toBe(true); - expect(opened).toEqual([HOST]); - }); - - it('does not accumulate failures further apart than the streak window', () => { - failRequest(0); - advance(FAILURE_WINDOW_MS + 1); - failRequest(0); - - expect(guard.check(HOST).allowed).toBe(true); - expect(opened).toEqual([]); - }); - - it('accumulates failures exactly at the streak window edge', () => { - failRequest(0); - advance(FAILURE_WINDOW_MS); - failRequest(0); - - expectBlocked(); - }); - - it('clears the record as soon as the host answers', () => { - failRequest(30_000); - const token = expectAllowed(); - guard.reportSuccess(token); - failRequest(30_000); - - expect(guard.check(HOST).allowed).toBe(true); - expect(opened).toEqual([]); - }); - - it('treats a success followed by an inconclusive report as a no-op', () => { - // The Stalker handler reports success on the response, then throws its - // own `HTTP Error 404` error, which is classified inconclusive. - failRequest(30_000); - const token = expectAllowed(); - guard.reportSuccess(token); - guard.reportInconclusive(token); - failRequest(30_000); - - expect(guard.check(HOST).allowed).toBe(true); - }); - - it('ignores an inconclusive failure entirely', () => { - const first = expectAllowed(); - guard.reportInconclusive(first); - const second = expectAllowed(); - guard.reportInconclusive(second); - - expect(guard.check(HOST).allowed).toBe(true); - expect(opened).toEqual([]); - }); - - describe('half-open', () => { - beforeEach(() => { - failRequest(30_000); - failRequest(30_000); - expectBlocked(); - advance(OPEN_DURATION_MS); - }); - - it('lets exactly one request through once the window elapses', () => { - const trial = expectAllowed(); - expect(trial.trial).toBe(true); - - expect(guard.check(HOST).allowed).toBe(false); - expect(guard.check(HOST).allowed).toBe(false); - }); - - it('closes the breaker when the trial succeeds', () => { - const trial = expectAllowed(); - guard.reportSuccess(trial); - - const next = expectAllowed(); - expect(next.trial).toBe(false); - expect(guard.check(HOST).allowed).toBe(true); - }); - - it('re-opens immediately when the trial fails, without a second strike', () => { - const trial = expectAllowed(); - advance(30_000); - guard.reportFailure(trial); - - expect(guard.check(HOST)).toEqual({ - allowed: false, - retryAfterMs: OPEN_DURATION_MS, - }); - expect(opened).toEqual([HOST, HOST]); - }); - - it('offers the trial slot again when the trial is cancelled', () => { - const trial = expectAllowed(); - guard.reportInconclusive(trial); - - const retry = expectAllowed(); - expect(retry.trial).toBe(true); - }); - - it('lets only the request holding the slot release it', () => { - // A trial can outlive its own 45 s window: the validated-redirect - // transport gives each of up to five hops its own 30 s budget. Once - // a replacement has been admitted, the abandoned request's late - // report must not hand the slot to a third request. - const abandoned = expectAllowed(); - advance(TRIAL_TIMEOUT_MS); - const replacement = expectAllowed(); - expect(replacement.trial).toBe(true); - - guard.reportInconclusive(abandoned); - - expectBlocked(); - }); - - it('recovers when a trial never reports back', () => { - expectAllowed(); - expectBlocked(); - - advance(TRIAL_TIMEOUT_MS); - - const replacement = expectAllowed(); - expect(replacement.trial).toBe(true); - }); - }); - - describe('reset', () => { - it('reopens the host for requests immediately', () => { - failRequest(30_000); - failRequest(30_000); - expectBlocked(); - - guard.reset(HOST); - - expect(expectAllowed().trial).toBe(false); - }); - - it('discards failures from requests that started before it', () => { - // The retry that cleared the breaker must not be poisoned by the - // 30-second stragglers it was waiting behind. - const straggler = expectAllowed(); - const secondStraggler = expectAllowed(); - advance(15_000); - guard.reset(HOST); - - advance(15_000); - guard.reportFailure(straggler); - guard.reportFailure(secondStraggler); - - expect(guard.check(HOST).allowed).toBe(true); - expect(opened).toEqual([]); - }); - - it('still counts failures from requests started after it', () => { - failRequest(30_000); - guard.reset(HOST); - failRequest(30_000); - failRequest(30_000); - - expectBlocked(); - }); - - it('supersedes in-flight requests for a host it has no record of', () => { - const token = expectAllowed(); - guard.clear(); - guard.reset(HOST); - advance(30_000); - guard.reportFailure(token); - failRequest(30_000); - - expect(guard.check(HOST).allowed).toBe(true); - }); - }); - - describe('bookkeeping bounds', () => { - it('forgets idle records instead of growing without bound', () => { - for (let index = 0; index < 300; index += 1) { - advance(1); - guard.check(`http://host-${index}.example.com`); - } - - // The cap held, and a fresh host is still allowed through. - expect(guard.check(HOST).allowed).toBe(true); - }); - - it('keeps an open breaker while other hosts churn past the idle TTL', () => { - failRequest(30_000); - failRequest(30_000); - expectBlocked(); - - advance(1); - guard.check('http://noise.example.com'); - - expectBlocked(); - }); - }); - - describe('kill switch', () => { - afterEach(() => { - delete process.env[GUARD_DISABLED_ENV]; - }); - - it('never blocks a request while disabled', () => { - process.env[GUARD_DISABLED_ENV] = '1'; - - failRequest(30_000); - failRequest(30_000); - failRequest(30_000); - - expect(guard.check(HOST).allowed).toBe(true); - expect(opened).toEqual([]); - }); - - it('accepts the "true" spelling', () => { - process.env[GUARD_DISABLED_ENV] = 'true'; - - failRequest(30_000); - failRequest(30_000); - - expect(guard.check(HOST).allowed).toBe(true); - }); - - it('is read per call, so an already open breaker stops blocking', () => { - failRequest(30_000); - failRequest(30_000); - expectBlocked(); - - process.env[GUARD_DISABLED_ENV] = '1'; - - expect(guard.check(HOST).allowed).toBe(true); - }); - }); -}); - describe('reportGuardedHostFailure', () => { const ENDPOINT = 'http://panel.example.com:8080'; const URL_ON_ENDPOINT = `${ENDPOINT}/player_api.php`; diff --git a/apps/electron-backend/src/app/util/host-connectivity-guard.ts b/apps/electron-backend/src/app/util/host-connectivity-guard.ts index c0c08e186..52c713244 100644 --- a/apps/electron-backend/src/app/util/host-connectivity-guard.ts +++ b/apps/electron-backend/src/app/util/host-connectivity-guard.ts @@ -1,453 +1,25 @@ /** - * Per-host circuit breaker for portal requests. + * Main-process ownership of the shared host connectivity guard. * - * A dead portal costs the full axios timeout on every call (30 s for Xtream, - * 15 s for Stalker), and browsing a dead portal's catalog issues dozens of - * those back to back — 30-second spinners and a flooded main-process log. Once - * a host has refused to answer twice in a row there is nothing left to learn - * from waiting again, so subsequent requests fail immediately for a short - * while. - * - * The rules are deliberately timid, because being wrong here means refusing to - * talk to a portal that works: - * - * - Only connection-level evidence counts (see - * {@link classifyHostRequestFailure}). Any HTTP response — 200, 404, even - * 502 — proves the host is alive and clears the record. - * - The window is short (30 s), so a mistake costs one page of navigation, and - * a host that came back is retried on its own. - * - Half-open means exactly ONE request goes out, not a whole screenful. - * - An explicit reset (user retry, endpoint discovery) always wins, and - * failures from requests that started before it are discarded — otherwise a - * 30-second straggler settles right after the reset and re-opens the breaker - * underneath the retry that cleared it. - * - * Deliberately not persisted: process lifetime is the right scope for - * "unreachable right now". + * The breaker itself lives in `@iptvnator/shared/host-health` so the web + * backend can run the same rules; what stays here is the part that is + * genuinely main-process-specific: a single guard for the whole process (both + * portal IPC handlers must see each other's evidence, or a dead endpoint would + * be relearned once per protocol) and the console warning that announces it. */ -import { buildHostConnectivityFastFailMessage } from '@iptvnator/shared/interfaces'; +import { + classifyHostRequestFailure, + failedAfterRedirect, + HostConnectivityGuard, + HostConnectivityGuardError, + HostRequestToken, + OPEN_DURATION_MS, + portalEndpointKeyOf, +} from '@iptvnator/shared/host-health'; -/** Consecutive connection failures that trip the breaker. */ -const FAILURE_THRESHOLD = 2; -/** How long requests fast-fail once the breaker is open. */ -const OPEN_DURATION_MS = 30_000; -/** - * Failures further apart than this are unrelated, not a streak. Two sequential - * 30 s timeouts must fit inside it, hence comfortably above 60 s. - */ -const FAILURE_WINDOW_MS = 120_000; -/** - * Safety net for a half-open trial that never reports back. Above the longest - * request timeout (30 s) plus margin, so it only fires if a caller leaked the - * token — without it a lost report would keep the breaker open forever. - */ -const TRIAL_TIMEOUT_MS = 45_000; -/** Hosts tracked at once; portals per user are few, this is just a bound. */ -const MAX_TRACKED_HOSTS = 256; -/** Idle records are forgotten; they hold nothing worth remembering. */ -const IDLE_TTL_MS = 600_000; - -const GUARD_DISABLED_ENV = 'IPTVNATOR_DISABLE_CONNECTIVITY_GUARD'; - -/** - * Error codes that prove the host itself did not answer. - * - * `ECONNRESET` is deliberately absent: a reset mid-transfer happens on hosts - * that are very much alive, and it is the one code a working stream can emit. - */ -const HOST_LEVEL_FAILURE_CODES = new Set([ - 'ECONNABORTED', - 'ECONNREFUSED', - 'EAI_AGAIN', - 'EHOSTUNREACH', - 'ENETUNREACH', - 'ENOTFOUND', - 'ETIMEDOUT', -]); - -/** - * Thrown instead of making a request while the breaker is open. - * - * A real `Error`, because Electron serializes a rejected plain object to - * `[object Object]` and the renderer's classification would be lost. It - * carries NO `status` property on purpose — `getStalkerRequestErrorStatus` - * reads that field first, and a numeric status there would read as "the - * endpoint answered". - */ -export class HostConnectivityGuardError extends Error { - /** Scheme, host and port — see {@link portalEndpointKeyOf}. */ - readonly endpoint: string; - - constructor(endpoint: string) { - super(buildHostConnectivityFastFailMessage(endpoint)); - this.name = 'HostConnectivityGuardError'; - this.endpoint = endpoint; - } -} - -/** - * Handed out by {@link HostConnectivityGuard.check} and passed back when the - * request settles. It records which attempt the report belongs to, so a reset - * or a parallel sibling cannot be mistaken for fresh evidence. - */ -export interface HostRequestToken { - /** Scheme, host and port — see {@link portalEndpointKeyOf}. */ - readonly endpoint: string; - readonly epoch: number; - readonly startedAt: number; - /** Whether this request is the single probe allowed while half-open. */ - readonly trial: boolean; - /** - * Which half-open slot this request holds, when it holds one. - * - * A trial can outlive its own window — `requestWithValidatedRedirects` - * gives each of up to five redirect hops its own 30 s budget — and once a - * replacement has been admitted, the abandoned request must not be able to - * free the replacement's slot. `trial: true` alone cannot tell the two - * apart, so the owner is identified explicitly. - */ - readonly trialId: number; -} - -export type HostConnectivityCheck = - | { readonly allowed: true; readonly token: HostRequestToken } - | { readonly allowed: false; readonly retryAfterMs: number }; - -export type HostRequestOutcome = 'responded' | 'host-level' | 'inconclusive'; - -interface HostState { - consecutiveFailures: number; - /** When the last COUNTED failure was recorded. */ - lastFailureAt: number; - openUntil: number; - trialStartedAt: number | null; - /** Monotonic id of the half-open slot; see `HostRequestToken.trialId`. */ - trialId: number; - epoch: number; - lastTouchedAt: number; -} - -export interface HostConnectivityGuardOptions { - now?: () => number; - onOpen?: (endpoint: string) => void; -} - -/** - * Whether the guard is switched off. Read at call time, not at module load, so - * a debugging session can toggle it (matches `IPTVNATOR_ALLOW_INSECURE_TLS`). - */ -export function isHostConnectivityGuardDisabled(): boolean { - const value = process.env[GUARD_DISABLED_ENV]?.trim().toLowerCase(); - return value === '1' || value === 'true'; -} - -/** - * What a failed request proves about the host. - * - * `responded` covers an error that still carries an HTTP response (5xx reaches - * the handlers as a rejection because `validateStatus` only tolerates <500) — - * the host is alive, so it clears the record. `inconclusive` covers everything - * that says nothing about reachability: a cancelled request, a URL rejected by - * the SSRF policy, a bug in our own code. - */ -export function classifyHostRequestFailure(error: unknown): HostRequestOutcome { - if (!error || typeof error !== 'object') { - return 'inconclusive'; - } - - if ((error as { response?: unknown }).response) { - return 'responded'; - } - - const code = (error as { code?: unknown }).code; - if (typeof code === 'string' && HOST_LEVEL_FAILURE_CODES.has(code)) { - return 'host-level'; - } - - return 'inconclusive'; -} - -/** - * The guard key for a request URL: its ORIGIN, i.e. scheme, host and port. - * - * Not `URL.host`, which omits a default port and therefore gives - * `http://panel.example` and `https://panel.example` the same key — two - * genuinely different endpoints, and a panel whose TLS listener is broken while - * plain HTTP works is a routine IPTV setup. Sharing one record there would let - * the dead one fast-fail the working one without ever contacting it. - * - * `URL.origin` also leaves out any `user:pass@` userinfo, so no credential - * reaches the key or the log line. - */ -export function portalEndpointKeyOf(url: string): string | null { - try { - const origin = new URL(url).origin; - // Opaque origins serialize as "null" and are not a usable key. - return origin && origin !== 'null' ? origin : null; - } catch { - return null; - } -} - -export class HostConnectivityGuard { - private readonly states = new Map(); - private readonly now: () => number; - private readonly onOpen?: (endpoint: string) => void; - - constructor(options: HostConnectivityGuardOptions = {}) { - this.now = options.now ?? (() => Date.now()); - this.onOpen = options.onOpen; - } - - /** - * Whether a request to `endpoint` may go out, and the token to report with. - * - * Fast-fails while the breaker is open, and while a half-open trial is - * still in flight — the point of the trial is that a screenful of requests - * does not all hang again to learn the same thing. - */ - check(endpoint: string): HostConnectivityCheck { - const now = this.now(); - if (isHostConnectivityGuardDisabled()) { - return { - allowed: true, - token: { - endpoint, - epoch: 0, - startedAt: now, - trial: false, - trialId: 0, - }, - }; - } - - const state = this.ensureState(endpoint, now); - state.lastTouchedAt = now; - - if (state.openUntil > now) { - return { allowed: false, retryAfterMs: state.openUntil - now }; - } - - // Never opened, or the open window elapsed. The latter is half-open: - // let one request through and hold the rest back until it settles. - let trial = false; - if (state.openUntil > 0) { - const trialStale = - state.trialStartedAt !== null && - now - state.trialStartedAt >= TRIAL_TIMEOUT_MS; - if (state.trialStartedAt !== null && !trialStale) { - return { allowed: false, retryAfterMs: 0 }; - } - state.trialStartedAt = now; - state.trialId += 1; - trial = true; - } - - return { - allowed: true, - token: { - endpoint, - epoch: state.epoch, - startedAt: now, - trial, - trialId: trial ? state.trialId : 0, - }, - }; - } - - /** - * The host answered. Always clears the record, whatever the status was and - * whichever attempt it belonged to — a reachable host is a reachable host. - */ - reportSuccess(token: HostRequestToken): void { - const state = this.states.get(token.endpoint); - if (!state || isHostConnectivityGuardDisabled()) { - return; - } - - state.consecutiveFailures = 0; - state.lastFailureAt = 0; - state.openUntil = 0; - state.trialStartedAt = null; - state.lastTouchedAt = this.now(); - } - - /** The host did not answer at all. */ - reportFailure(token: HostRequestToken): void { - const state = this.states.get(token.endpoint); - if (!state || isHostConnectivityGuardDisabled()) { - return; - } - - const now = this.now(); - // Superseded by an explicit reset: the user asked for a fresh attempt - // and this verdict predates it. - if (state.epoch !== token.epoch) { - this.releaseTrial(state, token); - return; - } - - // Only the request that still holds the slot ends the half-open state. - // An abandoned trial's late failure is ordinary evidence — it goes - // through the streak rules below and leaves the replacement alone. - const wasTrial = this.ownsTrial(state, token); - if (wasTrial) { - state.trialStartedAt = null; - } - state.lastTouchedAt = now; - - if ( - state.consecutiveFailures > 0 && - now - state.lastFailureAt > FAILURE_WINDOW_MS - ) { - state.consecutiveFailures = 0; - } - - // A request already in flight when the previous failure was recorded is - // not the next link in a streak — it is a sibling of it. Catalog - // loading fans out several requests at once, and one hiccup failing all - // of them is one piece of evidence, not a trip. A request that started - // at or after that moment is a genuine new attempt. - let counted = false; - if ( - state.consecutiveFailures === 0 || - token.startedAt >= state.lastFailureAt - ) { - state.consecutiveFailures += 1; - state.lastFailureAt = now; - counted = true; - } - - // Only a failure this report actually counted may trip the threshold. - // A sibling arriving after the window elapsed would otherwise open a - // fresh one off the existing count, pushing the half-open trial past - // the intended cooldown. A failed trial still goes straight back to - // open: the host had its chance. - if ( - wasTrial || - (counted && state.consecutiveFailures >= FAILURE_THRESHOLD) - ) { - const wasOpen = state.openUntil > now; - state.openUntil = now + OPEN_DURATION_MS; - if (!wasOpen) { - this.onOpen?.(token.endpoint); - } - } - } - - /** - * The request failed for a reason that says nothing about the host. Only - * releases the half-open slot, so the next request can be the trial. - */ - reportInconclusive(token: HostRequestToken): void { - const state = this.states.get(token.endpoint); - if (!state || isHostConnectivityGuardDisabled()) { - return; - } - - this.releaseTrial(state, token); - state.lastTouchedAt = this.now(); - } - - /** - * Forgets everything recorded for `endpoint`, and invalidates the reports of - * requests already in flight. Called when the user asks for a real attempt - * or when the URL may now point somewhere else. - */ - reset(endpoint: string): void { - const now = this.now(); - const state = this.ensureState(endpoint, now); - state.consecutiveFailures = 0; - state.lastFailureAt = 0; - state.openUntil = 0; - state.trialStartedAt = null; - state.epoch += 1; - state.lastTouchedAt = now; - } - - /** - * A token for a request that must NOT be policed or counted, but whose - * success still clears the record. - * - * Stalker endpoint discovery probes several candidate paths on one host and - * expects most of them to fail; counting those would let it declare a - * slow-but-alive portal unreachable. It does not take the half-open slot - * either — a probe is not the trial the guard is waiting for. - */ - observe(endpoint: string): HostRequestToken { - return { - endpoint, - epoch: this.states.get(endpoint)?.epoch ?? 0, - startedAt: this.now(), - trial: false, - trialId: 0, - }; - } - - /** Test seam: drops all recorded state. */ - clear(): void { - this.states.clear(); - } - - /** Whether `token` still holds the current half-open slot. */ - private ownsTrial(state: HostState, token: HostRequestToken): boolean { - return ( - token.trial && - state.trialStartedAt !== null && - state.epoch === token.epoch && - state.trialId === token.trialId - ); - } - - private releaseTrial(state: HostState, token: HostRequestToken): void { - if (this.ownsTrial(state, token)) { - state.trialStartedAt = null; - } - } - - private ensureState(endpoint: string, now: number): HostState { - const existing = this.states.get(endpoint); - if (existing) { - return existing; - } - - this.prune(now); - const state: HostState = { - consecutiveFailures: 0, - lastFailureAt: 0, - openUntil: 0, - trialStartedAt: null, - trialId: 0, - epoch: 0, - lastTouchedAt: now, - }; - this.states.set(endpoint, state); - return state; - } - - private prune(now: number): void { - for (const [endpoint, state] of this.states) { - if ( - now - state.lastTouchedAt > IDLE_TTL_MS && - state.openUntil <= now && - state.trialStartedAt === null - ) { - this.states.delete(endpoint); - } - } - - // Still at the cap: forget the oldest records. Dropping one means - // contacting that host again, which is the safe direction to err in. - while (this.states.size >= MAX_TRACKED_HOSTS) { - const oldest = this.states.keys().next(); - if (oldest.done) { - return; - } - this.states.delete(oldest.value); - } - } -} +export { HostConnectivityGuardError }; +export type { HostRequestToken }; let sharedGuard: HostConnectivityGuard | null = null; @@ -506,66 +78,6 @@ export function reportGuardedHostSuccess(token: HostRequestToken | null): void { getHostConnectivityGuard().reportSuccess(token); } } - -/** - * The URL a failed request was actually talking to, when the error says. - * - * Redirects are followed hop by hop, each with its own config, so a failure on - * a later hop carries THAT hop's URL rather than the one we asked for. - */ -function failedRequestUrlOf(error: unknown): string | null { - const url = (error as { config?: { url?: unknown } } | null)?.config?.url; - return typeof url === 'string' ? url : null; -} - -function normalizedUrlOrNull(url: string): string | null { - try { - return new URL(url).toString(); - } catch { - return null; - } -} - -/** - * Whether the failure happened on a hop the guarded endpoint redirected us to. - * - * Reaching any later hop proves the guarded endpoint answered: the first hop is - * always the URL we asked for, and only a redirect status advances the chain. - * That holds for a same-origin redirect too, so comparing the whole URL — not - * just its origin — is what catches `/player_api.php` → `/slow/player_api.php`. - * - * Requires positive evidence: anything unparseable or unknown returns false and - * the failure is counted as usual, because guessing "redirect" here would stop - * the guard from ever tripping. - */ -function failedAfterRedirect( - error: unknown, - token: HostRequestToken, - requestUrl: string | undefined -): boolean { - const failedUrl = failedRequestUrlOf(error); - if (!failedUrl) { - return false; - } - - const failedEndpoint = portalEndpointKeyOf(failedUrl); - if (failedEndpoint && failedEndpoint !== token.endpoint) { - return true; - } - - if (!requestUrl) { - return false; - } - - const failedNormalized = normalizedUrlOrNull(failedUrl); - const requestedNormalized = normalizedUrlOrNull(requestUrl); - return ( - failedNormalized !== null && - requestedNormalized !== null && - failedNormalized !== requestedNormalized - ); -} - /** * Records what a failed request proved about its endpoint. * diff --git a/apps/web-backend/src/app/host-guard.ts b/apps/web-backend/src/app/host-guard.ts new file mode 100644 index 000000000..645295e7a --- /dev/null +++ b/apps/web-backend/src/app/host-guard.ts @@ -0,0 +1,218 @@ +/** + * Request timeouts and the per-host circuit breaker for the proxy routes. + * + * Both exist for the same reason. Every outbound call here used to be a bare + * `axios.get()` with no `timeout`, so a provider host that accepts a connection + * and then goes silent held the request until the OS gave up on the TCP + * connection — minutes, not seconds, and the browser tab waited for all of it. + * The Electron handlers have always passed an explicit timeout; these are the + * same numbers, so the PWA and the desktop app give up at the same point. + * + * A timeout alone only bounds a single request. Browsing a dead portal issues + * dozens back to back, so the breaker (shared with the Electron main process, + * see `@iptvnator/shared/host-health`) fast-fails the rest once a host has + * refused to answer twice in a row. + * + * On the axios timeout semantics: with the default (follow-redirects) + * transport, `timeout` is NOT a wall-clock deadline for the whole response. It + * bounds the time to response headers, and then continues as the socket's + * inactivity timeout for the body. A large XMLTV or M3U download that keeps + * delivering bytes is therefore never cut off mid-transfer — only a stalled one + * is, which is exactly the case these numbers are meant to catch. + */ + +import { + HostConnectivityGuard, + HostRequestToken, + classifyHostRequestFailure, + failedAfterRedirect, + portalEndpointKeyOf, +} from '@iptvnator/shared/host-health'; +import { buildHostConnectivityFastFailMessage } from '@iptvnator/shared/interfaces'; +import type { NormalizedProviderError } from './provider-error'; + +/** + * Per-route request timeouts, matching the Electron handlers so both runtimes + * give up at the same point: + * + * - `xtream` — `xtream.events.ts` (30 s for the Xtream API) + * - `stalker` / `stalkerCreateLink` — `stalker.events.ts`; `create_link` gets + * longer because the portal mints a stream URL before answering + * - `playlist` — `PLAYLIST_FETCH_TIMEOUT_MS` in `playlist-source.ts` + * - `epg` — the EPG worker streams its download with no explicit timeout, so + * there is no number to copy; it takes the playlist budget, which bounds a + * silent host without capping a healthy long transfer (see the note above). + */ +export const PROVIDER_REQUEST_TIMEOUT_MS = { + epg: 30_000, + playlist: 30_000, + stalker: 15_000, + stalkerCreateLink: 30_000, + xtream: 30_000, +} as const; + +/** Status used for a fast-fail, matching `normalizeProviderError`'s upstream-unreachable answer. */ +const HOST_UNREACHABLE_STATUS = 502; + +export type HostRequestAdmission = + | { readonly allowed: true; readonly token: HostRequestToken | null } + | { readonly allowed: false; readonly error: NormalizedProviderError }; + +/** + * Whether a request to `url` may go out. + * + * A refusal carries the ordinary `{ message, status }` provider-error body, so + * every route keeps answering in the shape its client already parses. The + * message itself is the shared one from `@iptvnator/shared/interfaces` — the + * Stalker renderer classifies transport failures from message text alone, and + * that wording is the only one it reads as "connection-level failure" rather + * than as a status, a timeout or an auth problem. + * + * A URL with no usable host is admitted untracked: there is nothing to key on, + * and refusing over it would be worse than letting the transport report the + * real problem. + */ +export function admitProviderRequest( + guard: HostConnectivityGuard, + url: string +): HostRequestAdmission { + const endpoint = portalEndpointKeyOf(url); + if (!endpoint) { + return { allowed: true, token: null }; + } + + const check = guard.check(endpoint); + if (!check.allowed) { + return { + allowed: false, + error: { + message: buildHostConnectivityFastFailMessage(endpoint), + status: HOST_UNREACHABLE_STATUS, + }, + }; + } + + return { allowed: true, token: check.token }; +} + +/** + * Gives a token back without recording anything about the host. + * + * For a request that was admitted and then abandoned before it went out — a URL + * the SSRF policy refused, for instance. That says nothing about reachability, + * so it must not count as a failure; but the slot a half-open trial reserved + * has to be released, or the breaker waits out the full trial timeout for a + * request that never happened. + */ +export function releaseProviderRequest( + guard: HostConnectivityGuard, + token: HostRequestToken | null +): void { + if (token) { + guard.reportInconclusive(token); + } +} + +/** The host answered — whatever the status was, it is reachable. */ +export function reportProviderRequestSuccess( + guard: HostConnectivityGuard, + token: HostRequestToken | null +): void { + if (token) { + guard.reportSuccess(token); + } +} + +/** + * A token for a request that must NOT be policed or counted, but whose success + * still clears the record. + * + * Stalker endpoint discovery probes several candidate paths on one host and + * expects most of them to fail; counting those would let it declare a + * slow-but-alive portal unreachable, and fast-failing them would abandon a + * portal mid-discovery. The probe does not take the half-open slot either — it + * is not the trial the breaker is waiting for. Mirrors the Electron handler's + * `skipConnectionGuard` path. + */ +export function observeProviderRequest( + guard: HostConnectivityGuard, + url: string +): HostRequestToken | null { + const endpoint = portalEndpointKeyOf(url); + return endpoint ? guard.observe(endpoint) : null; +} + +/** + * Forgets the failures recorded for the host `url` points at, so the next + * request contacts it for real. Sent by the renderer whenever the user asked + * for a genuine attempt (portal retry, "test connection") or handed over an + * address that may now point somewhere else (import, edited connection, lazy + * repair) — the counterpart of the Electron `CONNECTIVITY_GUARD_RESET` handler. + * + * Returns whether a host could be read from `url` at all. + */ +export function resetProviderHost( + guard: HostConnectivityGuard, + url: string +): boolean { + const endpoint = portalEndpointKeyOf(url); + if (!endpoint) { + return false; + } + + guard.reset(endpoint); + return true; +} + +/** + * Records what a failed request proved about its endpoint. + * + * `requestUrl` is the URL the request was actually issued against. It is what + * lets a failure on a redirect hop be told apart from a failure of the endpoint + * we guarded: a provider that answers `302` towards a dead CDN is demonstrably + * alive, and charging the destination's refusal to it would fast-fail a working + * portal. Same rule and same helper as the Electron handlers. + * + * `countFailures: false` is for requests exempt from the guard (endpoint + * discovery). Their failures are expected and must not count — but the report + * must still happen, because an error that carries an HTTP response proves the + * endpoint answered. This route sets no `validateStatus`, so axios rejects every + * non-2xx WITH `error.response`; dropping those instead of reporting them is + * what would let the breaker open in the middle of discovery. + */ +export function reportProviderRequestFailure( + guard: HostConnectivityGuard, + token: HostRequestToken | null, + error: unknown, + options: { countFailures?: boolean; requestUrl?: string } = {} +): void { + if (!token) { + return; + } + + const countFailures = options.countFailures ?? true; + switch (classifyHostRequestFailure(error)) { + case 'host-level': + // Redirect attribution is checked BEFORE the exemption, not after. + // A 3xx from the guarded endpoint is an answer, and an exempt probe + // observing one has to clear the record just as it does for any + // other response — otherwise an ordinary timeout, a probe that was + // redirected to a dead destination, and another ordinary timeout + // still read as two consecutive failures. + if (failedAfterRedirect(error, token, options.requestUrl)) { + guard.reportSuccess(token); + break; + } + if (!countFailures) { + break; + } + guard.reportFailure(token); + break; + case 'responded': + guard.reportSuccess(token); + break; + default: + guard.reportInconclusive(token); + break; + } +} diff --git a/apps/web-backend/src/app/web-backend-app.host-guard.spec.ts b/apps/web-backend/src/app/web-backend-app.host-guard.spec.ts new file mode 100644 index 000000000..0bb7e8af6 --- /dev/null +++ b/apps/web-backend/src/app/web-backend-app.host-guard.spec.ts @@ -0,0 +1,653 @@ +/** + * Per-host circuit breaker behaviour on the proxy routes. + * + * The rules themselves are covered by the guard's own spec in + * `@iptvnator/shared/host-health`; what matters here is the wiring: which + * routes consult it, what a refusal looks like on the wire, and that a refusal + * really does skip the outbound request. + */ + +import { + HostConnectivityGuard, + OPEN_DURATION_MS, +} from '@iptvnator/shared/host-health'; +import { isHostConnectivityFastFailMessage } from '@iptvnator/shared/interfaces'; +import { createWebBackendApp } from './web-backend-app'; +import { + registerProviderTarget, + resolvePublicHost, + StubHttpClient, + withServer, +} from './web-backend-app.spec-helpers'; + +/** A refusal the guard counts: the host itself never answered. */ +function hostLevelFailure(code = 'ECONNREFUSED'): Error { + return Object.assign(new Error(`connect ${code}`), { code }); +} + +/** Guard driven by a clock the test owns, so no timers are involved. */ +function createTestGuard(): { + guard: HostConnectivityGuard; + advance: (ms: number) => void; +} { + let clock = 1_000; + return { + guard: new HostConnectivityGuard({ now: () => clock }), + advance: (ms: number) => { + clock += ms; + }, + }; +} + +describe('web backend host connectivity guard', () => { + it('fast-fails Xtream requests with HTTP 200 and the provider-error body once the host stops answering', async () => { + const httpClient = new StubHttpClient(); + httpClient.queueNetworkError(hostLevelFailure()); + httpClient.queueNetworkError(hostLevelFailure()); + const { guard } = createTestGuard(); + + await withServer( + createWebBackendApp({ + hostGuard: guard, + httpClient, + resolveHostname: resolvePublicHost, + }), + async (baseUrl) => { + const targetId = await registerProviderTarget( + baseUrl, + 'http://xtream.example' + ); + const call = () => + fetch( + `${baseUrl}/xtream?targetId=${targetId}&username=demo&password=secret&action=get_account_info` + ); + + await call(); + await call(); + const refused = await call(); + + // The route answers provider failures with HTTP 200 and an + // error body; PwaService parses that shape, so a fast-fail + // must not suddenly become a real HTTP error status. + expect(refused.status).toBe(200); + const body = (await refused.json()) as { + message: string; + status: number; + }; + expect(body.status).toBe(502); + expect(isHostConnectivityFastFailMessage(body.message)).toBe( + true + ); + // Names the host so the snackbar it reaches says something + // useful, and never the credential-bearing query string. + expect(body.message).toContain('xtream.example'); + expect(body.message).not.toContain('secret'); + + // The whole point: the third call never went out. + expect(httpClient.requests).toHaveLength(2); + } + ); + }); + + it('fast-fails Stalker requests with HTTP 200 and a message carrying no HTTP status', async () => { + const httpClient = new StubHttpClient(); + httpClient.queueNetworkError(hostLevelFailure('ETIMEDOUT')); + httpClient.queueNetworkError(hostLevelFailure('ETIMEDOUT')); + const { guard } = createTestGuard(); + + await withServer( + createWebBackendApp({ + hostGuard: guard, + httpClient, + resolveHostname: resolvePublicHost, + }), + async (baseUrl) => { + const targetId = await registerProviderTarget( + baseUrl, + 'http://stalker.example/portal.php' + ); + const call = () => + fetch( + `${baseUrl}/stalker?targetId=${targetId}&macAddress=00:1A:79:00:00:01&action=get_categories&type=vod` + ); + + await call(); + await call(); + const refused = await call(); + + expect(refused.status).toBe(200); + const body = (await refused.json()) as { + message: string; + status: number; + }; + expect(isHostConnectivityFastFailMessage(body.message)).toBe( + true + ); + // The renderer classifies Stalker transport failures from the + // message alone. An `HTTP Error ` in it would read as + // "the endpoint answered" and make discovery walk every + // remaining candidate; timeout wording would do the same. + expect(body.message).not.toMatch(/HTTP Error \d{3}/); + expect(body.message).not.toMatch(/timed out|timeout of \d+/i); + expect(httpClient.requests).toHaveLength(2); + } + ); + }); + + it('exempts endpoint-discovery probes from the breaker but still credits their successes', async () => { + // Discovery walks several candidate paths on one host and expects most + // to fail. Counting those would declare a slow-but-alive portal + // unreachable, and fast-failing them would abandon a portal + // mid-discovery — the candidate that works may be the last one tried. + const httpClient = new StubHttpClient(); + httpClient.queueNetworkError(hostLevelFailure('ETIMEDOUT')); + httpClient.queueNetworkError(hostLevelFailure('ETIMEDOUT')); + httpClient.queueResponse({ js: [{ id: '1' }] }); + httpClient.queueResponse({ js: [{ id: '2' }] }); + const { guard } = createTestGuard(); + + await withServer( + createWebBackendApp({ + hostGuard: guard, + httpClient, + resolveHostname: resolvePublicHost, + }), + async (baseUrl) => { + const targetId = await registerProviderTarget( + baseUrl, + 'http://stalker.example/portal.php' + ); + const probe = () => + fetch( + `${baseUrl}/stalker?targetId=${targetId}&macAddress=00:1A:79:00:00:01&action=get_genres&skipConnectionGuard=true` + ); + + await probe(); + await probe(); + const third = await probe(); + + // Two probe failures must not have opened the breaker. + await expect(third.json()).resolves.toEqual({ + action: 'get_genres', + payload: { js: [{ id: '1' }] }, + }); + expect(httpClient.requests).toHaveLength(3); + + // The exemption is bypass-and-don't-count, not + // don't-report-at-all: a candidate that answers still clears + // whatever the other candidates recorded. + const normal = await fetch( + `${baseUrl}/stalker?targetId=${targetId}&macAddress=00:1A:79:00:00:01&action=get_categories` + ); + await expect(normal.json()).resolves.toEqual({ + action: 'get_categories', + payload: { js: [{ id: '2' }] }, + }); + expect(httpClient.requests).toHaveLength(4); + + // And the flag is a control param: it must never be forwarded + // into the portal's own query string. + for (const request of httpClient.requests) { + expect(request.url).not.toContain('skipConnectionGuard'); + } + } + ); + }); + + it('lets an exempt probe clear the record with a response it was rejected for', async () => { + // This route sets no `validateStatus`, so axios rejects EVERY non-2xx + // with `error.response` attached. A discovery probe answered with 500 + // has still proved the endpoint alive, so the report has to happen even + // though the probe's own failures are not counted — dropping it is what + // lets the breaker open in the middle of discovery. + const httpClient = new StubHttpClient(); + httpClient.queueNetworkError(hostLevelFailure('ETIMEDOUT')); + httpClient.queueFailure(500, 'Internal Server Error'); + httpClient.queueNetworkError(hostLevelFailure('ETIMEDOUT')); + httpClient.queueResponse({ js: [{ id: '1' }] }); + const { guard } = createTestGuard(); + + await withServer( + createWebBackendApp({ + hostGuard: guard, + httpClient, + resolveHostname: resolvePublicHost, + }), + async (baseUrl) => { + const targetId = await registerProviderTarget( + baseUrl, + 'http://stalker.example/portal.php' + ); + const base = `${baseUrl}/stalker?targetId=${targetId}&macAddress=00:1A:79:00:00:01`; + + // Counted failure, then an exempt probe answered with 500, then + // another counted failure. Without the probe's report the two + // counted failures are consecutive and open the breaker. + await fetch(`${base}&action=get_categories`); + await fetch( + `${base}&action=get_genres&skipConnectionGuard=true` + ); + await fetch(`${base}&action=get_categories`); + + const afterwards = await fetch(`${base}&action=get_categories`); + + await expect(afterwards.json()).resolves.toEqual({ + action: 'get_categories', + payload: { js: [{ id: '1' }] }, + }); + expect(httpClient.requests).toHaveLength(4); + } + ); + }); + + it('never fast-fails playlist or EPG downloads, however often they fail', async () => { + // The breaker covers the portal routes only, matching Electron, where + // it is wired into the two portal IPC handlers and not into the + // playlist or EPG download path. Guarding these was a mistake caught in + // review: a download is one request rather than a catalog fan-out, it + // is usually the direct result of the user asking for it — so refusing + // an immediate retry of an M3U import is a regression, not a + // protection — and a large XMLTV transfer can legitimately outlive the + // half-open trial window, which would let a second trial in behind it. + const httpClient = new StubHttpClient(); + for (let i = 0; i < 4; i++) { + httpClient.queueNetworkError(hostLevelFailure()); + } + const { guard } = createTestGuard(); + + await withServer( + createWebBackendApp({ + hostGuard: guard, + httpClient, + resolveHostname: resolvePublicHost, + }), + async (baseUrl) => { + const playlistTarget = await registerProviderTarget( + baseUrl, + 'https://provider.example/list.m3u' + ); + const epgTarget = await registerProviderTarget( + baseUrl, + 'https://provider.example/guide.xml' + ); + + for (let i = 0; i < 3; i++) { + await fetch(`${baseUrl}/parse?targetId=${playlistTarget}`); + } + const epg = await fetch( + `${baseUrl}/parse-xml?targetId=${epgTarget}` + ); + + // Every one of them reached the host: no admission check, and + // nothing recorded against it either. + expect(httpClient.requests).toHaveLength(4); + const body = (await epg.json()) as { message: string }; + expect(isHostConnectivityFastFailMessage(body.message)).toBe( + false + ); + } + ); + }); + + it('keeps talking to a host that answers, whatever the status says', async () => { + const httpClient = new StubHttpClient(); + // A 404 is an answer: the host is alive and the record must clear. + httpClient.queueNetworkError(hostLevelFailure()); + httpClient.queueFailure(404, 'Not Found'); + httpClient.queueNetworkError(hostLevelFailure()); + httpClient.queueResponse({ user_info: { username: 'demo' } }); + const { guard } = createTestGuard(); + + await withServer( + createWebBackendApp({ + hostGuard: guard, + httpClient, + resolveHostname: resolvePublicHost, + }), + async (baseUrl) => { + const targetId = await registerProviderTarget( + baseUrl, + 'http://xtream.example' + ); + const call = () => + fetch( + `${baseUrl}/xtream?targetId=${targetId}&username=demo&action=get_account_info` + ); + + await call(); + await call(); + await call(); + const answered = await call(); + + await expect(answered.json()).resolves.toEqual({ + action: 'get_account_info', + payload: { user_info: { username: 'demo' } }, + }); + expect(httpClient.requests).toHaveLength(4); + } + ); + }); + + it('fast-fails an open host before spending a DNS lookup on it', async () => { + // URL validation resolves the hostname, and a dead host is exactly + // where DNS is slow or failing too. Checking the breaker after that + // meant an open host still paid for a lookup and could answer "host + // could not be resolved" instead of the intended fast-fail. + const httpClient = new StubHttpClient(); + httpClient.queueNetworkError(hostLevelFailure()); + httpClient.queueNetworkError(hostLevelFailure()); + const { guard } = createTestGuard(); + const resolveHostname = jest + .fn, [string]>() + .mockResolvedValue(['93.184.216.34']); + + await withServer( + createWebBackendApp({ + hostGuard: guard, + httpClient, + resolveHostname, + }), + async (baseUrl) => { + const targetId = await registerProviderTarget( + baseUrl, + 'http://xtream.example' + ); + const call = () => + fetch( + `${baseUrl}/xtream?targetId=${targetId}&username=demo&action=get_account_info` + ); + + await call(); + await call(); + const lookupsBeforeRefusal = resolveHostname.mock.calls.length; + + const refused = await call(); + + expect(refused.status).toBe(200); + const body = (await refused.json()) as { message: string }; + expect(isHostConnectivityFastFailMessage(body.message)).toBe( + true + ); + expect(resolveHostname).toHaveBeenCalledTimes( + lookupsBeforeRefusal + ); + } + ); + }); + + it('returns the half-open slot when a URL policy refusal aborts the request', async () => { + // Admitted, then abandoned before anything went out. That is not + // evidence about the host, but the trial slot has to go back or the + // breaker waits out the full trial window for a request that never + // happened. + const httpClient = new StubHttpClient(); + httpClient.queueNetworkError(hostLevelFailure()); + httpClient.queueNetworkError(hostLevelFailure()); + httpClient.queueResponse({ user_info: { username: 'demo' } }); + const { advance, guard } = createTestGuard(); + let resolvable = true; + const resolveHostname = async () => { + if (!resolvable) { + throw new Error('EAI_AGAIN'); + } + return ['93.184.216.34']; + }; + + await withServer( + createWebBackendApp({ + hostGuard: guard, + httpClient, + resolveHostname, + }), + async (baseUrl) => { + const targetId = await registerProviderTarget( + baseUrl, + 'http://xtream.example' + ); + const call = () => + fetch( + `${baseUrl}/xtream?targetId=${targetId}&username=demo&action=get_account_info` + ); + + await call(); + await call(); + advance(OPEN_DURATION_MS + 1); + + // Half-open. This trial is abandoned by the policy, so it must + // not consume the single trial the breaker allows. + resolvable = false; + const refusedByPolicy = await call(); + expect(refusedByPolicy.status).toBe(400); + + resolvable = true; + const trial = await call(); + await expect(trial.json()).resolves.toEqual({ + action: 'get_account_info', + payload: { user_info: { username: 'demo' } }, + }); + expect(httpClient.requests).toHaveLength(3); + } + ); + }); + + it('counts a hostname that stops resolving, and never leaks the lookup error', async () => { + // A name that will not resolve is the host failing to answer, the same + // as the ENOTFOUND the transport would have raised a moment later. + // Releasing it as a policy refusal left the breaker permanently shut + // and every request paying for the same dead lookup. + const httpClient = new StubHttpClient(); + const { guard } = createTestGuard(); + let resolvable = true; + const resolveHostname = async () => { + if (!resolvable) { + throw Object.assign(new Error('getaddrinfo ENOTFOUND'), { + code: 'ENOTFOUND', + }); + } + return ['93.184.216.34']; + }; + + await withServer( + createWebBackendApp({ + hostGuard: guard, + httpClient, + resolveHostname, + }), + async (baseUrl) => { + // Registering the target resolves the name too, so do it while + // DNS still answers — then let the name die. + const targetId = await registerProviderTarget( + baseUrl, + 'http://xtream.example' + ); + resolvable = false; + + const call = () => + fetch( + `${baseUrl}/xtream?targetId=${targetId}&username=demo&action=get_account_info` + ); + + const first = await call(); + expect(first.status).toBe(400); + // The DNS error is internal: the client sees what it always saw. + await expect(first.json()).resolves.toEqual({ + message: 'Provider URL host could not be resolved', + status: 400, + }); + + await call(); + const refused = await call(); + + // Two counted lookup failures opened the breaker, so the third + // is refused outright instead of resolving again. + expect(refused.status).toBe(200); + const body = (await refused.json()) as { message: string }; + expect(isHostConnectivityFastFailMessage(body.message)).toBe( + true + ); + expect(httpClient.requests).toHaveLength(0); + } + ); + }); + + it('does not fast-fail a provider whose redirect destination is dead', async () => { + // The shape this route actually produces, verified against axios + // 1.19.0: follow-redirects walks the chain inside one `get()`, so + // `config` still holds the URL we asked for and only + // `request._currentUrl` names the hop that failed. Reading `config.url` + // alone would compare the original URL with itself, find no redirect, + // and charge the dead destination to the provider that answered. + const redirectedFailure = () => + Object.assign(new Error('connect ECONNREFUSED'), { + code: 'ECONNREFUSED', + config: { url: 'http://xtream.example/player_api.php' }, + request: { _currentUrl: 'http://cdn.dead.example/stream' }, + }); + const httpClient = new StubHttpClient(); + httpClient.queueNetworkError(redirectedFailure()); + httpClient.queueNetworkError(redirectedFailure()); + httpClient.queueResponse({ user_info: { username: 'demo' } }); + const { guard } = createTestGuard(); + + await withServer( + createWebBackendApp({ + hostGuard: guard, + httpClient, + resolveHostname: resolvePublicHost, + }), + async (baseUrl) => { + const targetId = await registerProviderTarget( + baseUrl, + 'http://xtream.example' + ); + const call = () => + fetch( + `${baseUrl}/xtream?targetId=${targetId}&username=demo&action=get_account_info` + ); + + await call(); + await call(); + const third = await call(); + + await expect(third.json()).resolves.toEqual({ + action: 'get_account_info', + payload: { user_info: { username: 'demo' } }, + }); + expect(httpClient.requests).toHaveLength(3); + } + ); + }); + + it('lets exactly one request through once the open window elapses', async () => { + const httpClient = new StubHttpClient(); + httpClient.queueNetworkError(hostLevelFailure()); + httpClient.queueNetworkError(hostLevelFailure()); + httpClient.queueResponse({ user_info: { username: 'demo' } }); + const { advance, guard } = createTestGuard(); + + await withServer( + createWebBackendApp({ + hostGuard: guard, + httpClient, + resolveHostname: resolvePublicHost, + }), + async (baseUrl) => { + const targetId = await registerProviderTarget( + baseUrl, + 'http://xtream.example' + ); + const call = () => + fetch( + `${baseUrl}/xtream?targetId=${targetId}&username=demo&action=get_account_info` + ); + + await call(); + await call(); + await call(); + expect(httpClient.requests).toHaveLength(2); + + advance(OPEN_DURATION_MS + 1); + const trial = await call(); + + await expect(trial.json()).resolves.toEqual({ + action: 'get_account_info', + payload: { user_info: { username: 'demo' } }, + }); + expect(httpClient.requests).toHaveLength(3); + } + ); + }); + + it('contacts the host for real again after an explicit reset', async () => { + const httpClient = new StubHttpClient(); + httpClient.queueNetworkError(hostLevelFailure()); + httpClient.queueNetworkError(hostLevelFailure()); + httpClient.queueResponse({ user_info: { username: 'demo' } }); + const { guard } = createTestGuard(); + + await withServer( + createWebBackendApp({ + hostGuard: guard, + httpClient, + resolveHostname: resolvePublicHost, + }), + async (baseUrl) => { + const targetId = await registerProviderTarget( + baseUrl, + 'http://xtream.example' + ); + const call = () => + fetch( + `${baseUrl}/xtream?targetId=${targetId}&username=demo&action=get_account_info` + ); + + await call(); + await call(); + await call(); + expect(httpClient.requests).toHaveLength(2); + + const reset = await fetch( + `${baseUrl}/connectivity-guard/reset`, + { + body: JSON.stringify({ + url: 'http://xtream.example/player_api.php', + }), + headers: { 'content-type': 'application/json' }, + method: 'POST', + } + ); + await expect(reset.json()).resolves.toEqual({ reset: true }); + + const retried = await call(); + await expect(retried.json()).resolves.toEqual({ + action: 'get_account_info', + payload: { user_info: { username: 'demo' } }, + }); + expect(httpClient.requests).toHaveLength(3); + } + ); + }); + + it('rejects a reset without a url and reports one it cannot read a host from', async () => { + await withServer( + createWebBackendApp({ resolveHostname: resolvePublicHost }), + async (baseUrl) => { + const post = (body: unknown) => + fetch(`${baseUrl}/connectivity-guard/reset`, { + body: JSON.stringify(body), + headers: { 'content-type': 'application/json' }, + method: 'POST', + }); + + const missing = await post({}); + expect(missing.status).toBe(400); + + const unusable = await post({ url: 'not a url' }); + expect(unusable.status).toBe(200); + await expect(unusable.json()).resolves.toEqual({ + reset: false, + }); + } + ); + }); +}); diff --git a/apps/web-backend/src/app/web-backend-app.spec-helpers.ts b/apps/web-backend/src/app/web-backend-app.spec-helpers.ts new file mode 100644 index 000000000..2266cedaf --- /dev/null +++ b/apps/web-backend/src/app/web-backend-app.spec-helpers.ts @@ -0,0 +1,154 @@ +/** + * Shared fixtures for the web backend specs. + * + * Extracted so the host-guard spec can drive the same stub transport and + * throwaway server as the main spec without a second copy of either — and so + * neither file has to grow past the 1200-line test limit to hold them. + */ + +import { AddressInfo } from 'node:net'; +import { Server } from 'node:http'; +import { STALKER_MAG_USER_AGENT } from '@iptvnator/shared/interfaces'; +import { + createWebBackendApp, + WebBackendHttpClient, + WebBackendHttpGetOptions, +} from './web-backend-app'; + +/** The transport-identity headers every portal-facing Stalker request carries. */ +export const STALKER_IDENTITY_HEADERS = { + 'User-Agent': STALKER_MAG_USER_AGENT, + 'X-User-Agent': STALKER_MAG_USER_AGENT, + Accept: '*/*', + Connection: 'keep-alive', + 'Accept-Language': 'en-US,en;q=0.9', +}; + +export interface HttpRequest { + readonly headers?: Record; + readonly params?: Record; + readonly responseData: unknown; + readonly responseStatus?: number; + readonly timeout?: number; + readonly url: string; +} + +export class StubHttpClient implements WebBackendHttpClient { + readonly requests: Omit[] = + []; + private readonly queuedResponses: Array<{ + readonly data: unknown; + readonly error?: Error; + readonly status?: number; + readonly statusText?: string; + }> = []; + + queueResponse(data: unknown): void { + this.queuedResponses.push({ data }); + } + + queueFailure(status: number, statusText = 'Provider failure'): void { + this.queuedResponses.push({ data: null, status, statusText }); + } + + queueNetworkFailure(message = 'connect ECONNREFUSED'): void { + this.queuedResponses.push({ data: null, error: new Error(message) }); + } + + queueNetworkError(error: Error): void { + this.queuedResponses.push({ data: null, error }); + } + + async get( + url: string, + options: WebBackendHttpGetOptions = {} + ): Promise<{ data: T }> { + this.requests.push({ + headers: options.headers, + params: options.params, + timeout: options.timeout, + url, + }); + + const response = this.queuedResponses.shift(); + if (!response) { + throw new Error(`No queued response for ${url}`); + } + + if (response.error) { + // Shaped like a real axios rejection, because the guard reads these + // fields to tell a redirect hop from the endpoint we asked for. A + // bare Error here hid a bug where every ordinary /xtream failure + // looked like a redirect — axios appends `params` to the sent URL, + // so `_currentUrl` carries a query the caller's baseline does not. + const error = response.error as Error & { + config?: { url: string }; + request?: { _currentUrl: string }; + }; + error.config ??= { url }; + error.request ??= { _currentUrl: withQuery(url, options.params) }; + throw error; + } + + if (response.status) { + const error = new Error(response.statusText) as Error & { + response: { status: number; statusText: string }; + }; + error.response = { + status: response.status, + statusText: response.statusText ?? 'Provider failure', + }; + throw error; + } + + return { data: response.data as T }; + } +} + +/** How axios renders `params` into the URL it actually sends. */ +function withQuery( + url: string, + params: Record | undefined +): string { + const entries = Object.entries(params ?? {}); + if (entries.length === 0) { + return url; + } + const query = new URLSearchParams(entries).toString(); + return url.includes('?') ? `${url}&${query}` : `${url}?${query}`; +} + +export const resolvePublicHost = async () => ['93.184.216.34']; + +export async function registerProviderTarget( + baseUrl: string, + url: string +): Promise { + const response = await fetch(`${baseUrl}/provider-targets`, { + body: JSON.stringify({ url }), + headers: { + 'content-type': 'application/json', + }, + method: 'POST', + }); + const body = (await response.json()) as { targetId: string }; + return body.targetId; +} + +export async function withServer( + app: ReturnType, + callback: (baseUrl: string) => Promise +): Promise { + const server = await new Promise((resolve) => { + const started = app.listen(0, '127.0.0.1', () => resolve(started)); + }); + + try { + const address = server.address() as AddressInfo; + return await callback(`http://127.0.0.1:${address.port}`); + } finally { + await new Promise((resolve, reject) => { + server.close((error) => (error ? reject(error) : resolve())); + }); + } +} diff --git a/apps/web-backend/src/app/web-backend-app.spec.ts b/apps/web-backend/src/app/web-backend-app.spec.ts index bd271f1b4..bd4ea7292 100644 --- a/apps/web-backend/src/app/web-backend-app.spec.ts +++ b/apps/web-backend/src/app/web-backend-app.spec.ts @@ -1,123 +1,11 @@ -import { AddressInfo } from 'node:net'; -import { Server } from 'node:http'; -import { STALKER_MAG_USER_AGENT } from '@iptvnator/shared/interfaces'; +import { createWebBackendApp } from './web-backend-app'; import { - createWebBackendApp, - WebBackendHttpClient, - WebBackendHttpGetOptions, -} from './web-backend-app'; - -/** The transport-identity headers every portal-facing Stalker request carries. */ -const STALKER_IDENTITY_HEADERS = { - 'User-Agent': STALKER_MAG_USER_AGENT, - 'X-User-Agent': STALKER_MAG_USER_AGENT, - Accept: '*/*', - Connection: 'keep-alive', - 'Accept-Language': 'en-US,en;q=0.9', -}; - -interface HttpRequest { - readonly headers?: Record; - readonly params?: Record; - readonly responseData: unknown; - readonly responseStatus?: number; - readonly url: string; -} - -class StubHttpClient implements WebBackendHttpClient { - readonly requests: Omit[] = - []; - private readonly queuedResponses: Array<{ - readonly data: unknown; - readonly error?: Error; - readonly status?: number; - readonly statusText?: string; - }> = []; - - queueResponse(data: unknown): void { - this.queuedResponses.push({ data }); - } - - queueFailure(status: number, statusText = 'Provider failure'): void { - this.queuedResponses.push({ data: null, status, statusText }); - } - - queueNetworkFailure(message = 'connect ECONNREFUSED'): void { - this.queuedResponses.push({ data: null, error: new Error(message) }); - } - - queueNetworkError(error: Error): void { - this.queuedResponses.push({ data: null, error }); - } - - async get( - url: string, - options: WebBackendHttpGetOptions = {} - ): Promise<{ data: T }> { - this.requests.push({ - headers: options.headers, - params: options.params, - url, - }); - - const response = this.queuedResponses.shift(); - if (!response) { - throw new Error(`No queued response for ${url}`); - } - - if (response.error) { - throw response.error; - } - - if (response.status) { - const error = new Error(response.statusText) as Error & { - response: { status: number; statusText: string }; - }; - error.response = { - status: response.status, - statusText: response.statusText ?? 'Provider failure', - }; - throw error; - } - - return { data: response.data as T }; - } -} - -const resolvePublicHost = async () => ['93.184.216.34']; - -async function registerProviderTarget( - baseUrl: string, - url: string -): Promise { - const response = await fetch(`${baseUrl}/provider-targets`, { - body: JSON.stringify({ url }), - headers: { - 'content-type': 'application/json', - }, - method: 'POST', - }); - const body = (await response.json()) as { targetId: string }; - return body.targetId; -} - -async function withServer( - app: ReturnType, - callback: (baseUrl: string) => Promise -): Promise { - const server = await new Promise((resolve) => { - const started = app.listen(0, '127.0.0.1', () => resolve(started)); - }); - - try { - const address = server.address() as AddressInfo; - return await callback(`http://127.0.0.1:${address.port}`); - } finally { - await new Promise((resolve, reject) => { - server.close((error) => (error ? reject(error) : resolve())); - }); - } -} + registerProviderTarget, + resolvePublicHost, + STALKER_IDENTITY_HEADERS, + StubHttpClient, + withServer, +} from './web-backend-app.spec-helpers'; describe('web backend app', () => { it('exposes a health endpoint', async () => { @@ -203,6 +91,7 @@ https://stream.example/news.m3u8`); { headers: undefined, params: undefined, + timeout: 30000, url: 'https://provider.example/list.m3u', }, ]); @@ -284,6 +173,7 @@ https://stream.example/news.m3u8`); { headers: undefined, params: undefined, + timeout: 30000, url: 'https://provider.example/guide.xml', }, ]); @@ -400,6 +290,7 @@ https://stream.example/live.m3u8`); password: 'secret', username: 'demo', }, + timeout: 30000, url: 'http://xtream.example/player_api.php', }, ]); @@ -437,6 +328,7 @@ https://stream.example/live.m3u8`); password: 'secret', username: 'demo', }, + timeout: 30000, url: 'http://xtream.example/panel/player_api.php', }, ]); @@ -477,6 +369,7 @@ https://stream.example/live.m3u8`); Cookie: 'mac=00:1A:79:00:00:01; stb_lang=en_US@rg=dezzzz; timezone=Europe/Berlin', }, params: undefined, + timeout: 15000, url: 'http://stalker.example/portal.php?action=get_categories&type=vod&JsHttpRequest=1-xml', }, ]); @@ -630,6 +523,9 @@ https://stream.example/live.m3u8`); Cookie: 'mac=00:1A:79:00:00:01; stb_lang=en_US@rg=dezzzz; timezone=Europe/Berlin', }, params: undefined, + // `create_link` takes the longer budget: the portal + // mints a stream URL before it answers. + timeout: 30000, url: 'http://stalker.example/portal.php' + '?action=create_link&type=itv' + diff --git a/apps/web-backend/src/app/web-backend-app.ts b/apps/web-backend/src/app/web-backend-app.ts index 90f50363b..75ebc8e61 100644 --- a/apps/web-backend/src/app/web-backend-app.ts +++ b/apps/web-backend/src/app/web-backend-app.ts @@ -7,12 +7,25 @@ import zlib from 'node:zlib'; import axios from 'axios'; import epgParser from 'epg-parser'; import parser from 'iptv-playlist-parser'; +import { + HostConnectivityGuard, + HostRequestToken, +} from '@iptvnator/shared/host-health'; import { buildStalkerIdentityRequestContext, buildStalkerRequestUrl, normalizeXtreamServerUrl, } from '@iptvnator/shared/interfaces'; import { extractDrmFromRaw } from '@iptvnator/shared/m3u-utils'; +import { + admitProviderRequest, + observeProviderRequest, + PROVIDER_REQUEST_TIMEOUT_MS, + releaseProviderRequest, + reportProviderRequestFailure, + reportProviderRequestSuccess, + resetProviderHost, +} from './host-guard'; import { collectProviderErrorCodes, logProviderRequestFailure, @@ -24,6 +37,7 @@ export interface WebBackendHttpGetOptions { readonly headers?: Record; readonly params?: Record; readonly responseType?: 'arraybuffer'; + readonly timeout?: number; } export interface WebBackendHttpClient { @@ -43,6 +57,11 @@ export interface WebBackendAppOptions { readonly allowPrivateNetworkTargets?: boolean; readonly clientOrigins?: string[]; readonly guid?: () => string; + /** + * Per-host circuit breaker for the proxy routes. One per app, so a test can + * drive it with a fake clock the same way `now` and `guid` are injected. + */ + readonly hostGuard?: HostConnectivityGuard; readonly httpClient?: WebBackendHttpClient; readonly now?: () => Date; readonly resolveHostname?: (hostname: string) => Promise; @@ -57,6 +76,25 @@ interface ProviderUrlPolicy { interface ProviderUrlError { readonly message: string; readonly status: number; + /** + * The DNS failure behind a "could not be resolved" refusal, when that is + * what this is. Internal only — {@link providerUrlErrorBody} strips it, + * because the client is told the same thing it always was. + * + * A name that does not resolve is evidence about reachability, exactly like + * the `ENOTFOUND` the transport would have raised a moment later. Without + * it, a host whose DNS died is refused by the policy on every request and + * the breaker never opens, so each call keeps paying for the lookup. + */ + readonly lookupError?: unknown; +} + +/** The client-facing half of a {@link ProviderUrlError}. */ +function providerUrlErrorBody(error: ProviderUrlError): { + message: string; + status: number; +} { + return { message: error.message, status: error.status }; } type ProviderTargetRegistry = Map; @@ -68,6 +106,16 @@ export function createWebBackendApp( const httpClient = (options.httpClient ?? axios) as WebBackendHttpClient; const guid = options.guid ?? createGuid; const now = options.now ?? (() => new Date()); + const hostGuard = + options.hostGuard ?? + new HostConnectivityGuard({ + // Host and port only — the provider URL's query string routinely + // carries Xtream credentials and must never reach a log. + onOpen: (host) => + console.warn( + `[web-backend] ${host} is not answering; skipping requests to it for a short while` + ), + }); const clientOrigins = options.clientOrigins ?? getClientOrigins(); const runtimeBackendUrl = options.runtimeBackendUrl ?? process.env['BACKEND_URL'] ?? '/api'; @@ -127,7 +175,7 @@ export function createWebBackendApp( const result = await validateProviderUrl(rawUrl, providerUrlPolicy); if ('message' in result) { - res.status(result.status).json(result); + res.status(result.status).json(providerUrlErrorBody(result)); return; } @@ -137,12 +185,47 @@ export function createWebBackendApp( } ); + app.options('/connectivity-guard/reset', corsMiddleware); + app.post( + '/connectivity-guard/reset', + corsMiddleware, + express.json({ limit: '4kb' }), + (req, res) => { + // Takes the raw provider URL rather than a registered targetId: + // callers reset a host precisely when its address may have changed, + // which is before any target exists for it. Nothing is fetched + // here — only the host is read, so there is no SSRF surface — and + // the URL is never logged, since its query string carries the + // Xtream credentials. + const rawUrl = + req.body && + typeof req.body === 'object' && + 'url' in req.body && + typeof req.body.url === 'string' + ? req.body.url + : undefined; + + if (!rawUrl) { + res.status(400).json({ message: 'Missing url', status: 400 }); + return; + } + + res.json({ reset: resetProviderHost(hostGuard, rawUrl) }); + } + ); + app.get('/parse', corsMiddleware, async (req, res) => { const url = getRegisteredProviderUrl(req, res, providerTargets); if (!url) { return; } + // Deliberately unguarded, matching Electron, where the breaker is + // wired into the two portal IPC handlers and not into the playlist or + // EPG download path. See "Scope" in the contract doc: a download is one + // request rather than a catalog fan-out, it is usually the direct + // result of the user asking for it, and it can legitimately run far + // longer than any portal call. const result = await handlePlaylistParse({ guid, httpClient, @@ -164,6 +247,7 @@ export function createWebBackendApp( return; } + // Unguarded for the same reasons as /parse above. try { const result = await fetchEpgDataFromUrl(httpClient, url); if (!result) { @@ -193,31 +277,75 @@ export function createWebBackendApp( } const url = new URL(registeredUrl.href); + let guardToken: HostRequestToken | null = null; + // The URL actually requested, so a failure on a redirect hop can be + // told apart from a failure of the endpoint we guarded. + let requestUrl: string | undefined; try { + // Before URL validation, not after: that step resolves the hostname + // over DNS, and a dead host is exactly where DNS is slow or failing + // too. Checking first means an open breaker answers immediately + // with the fast-fail the caller expects, instead of paying for a + // lookup and then reporting an unrelated "host could not be + // resolved". Safe to key on the registered URL because + // `normalizeXtreamServerUrl` rebuilds from `url.origin`, so + // normalization can change the path but never the host. + const admission = admitProviderRequest(hostGuard, url.href); + if (!admission.allowed) { + // This route answers provider failures with HTTP 200 and an + // error body; a fast-fail is one of them. + res.json(admission.error); + return; + } + guardToken = admission.token; + const providerUrlError = await normalizeAndValidateXtreamProviderUrl( url, providerUrlPolicy ); if (providerUrlError) { - res.status(providerUrlError.status).json(providerUrlError); + // Admitted, then abandoned before any request went out. A + // policy refusal — private address, bad scheme — says nothing + // about reachability, so it only hands the half-open slot back. + // A name that would not resolve is different: that IS the host + // failing to answer, and counting it is what lets the breaker + // stop paying for the same dead lookup on every request. + if (providerUrlError.lookupError !== undefined) { + reportProviderRequestFailure( + hostGuard, + guardToken, + providerUrlError.lookupError + ); + } else { + releaseProviderRequest(hostGuard, guardToken); + } + guardToken = null; + res.status(providerUrlError.status).json( + providerUrlErrorBody(providerUrlError) + ); return; } + requestUrl = appendPathSegment(url, 'player_api.php'); + // Provider URLs are validated by /provider-targets before they enter the registry. // codeql[js/request-forgery] - const response = await httpClient.get( - appendPathSegment(url, 'player_api.php'), - { - params: getProxyParams(req, ['targetId']), - } - ); + const response = await httpClient.get(requestUrl, { + params: getProxyParams(req, ['targetId']), + timeout: PROVIDER_REQUEST_TIMEOUT_MS.xtream, + }); + reportProviderRequestSuccess(hostGuard, guardToken); + guardToken = null; res.json({ action: getQueryString(req, 'action'), payload: response.data, }); } catch (error) { + reportProviderRequestFailure(hostGuard, guardToken, error, { + requestUrl, + }); logProviderRequestFailure({ error, route: '/xtream', url }); res.json(normalizeProviderError(error)); } @@ -232,6 +360,12 @@ export function createWebBackendApp( return; } + // Endpoint-discovery probes expect most candidates to fail; counting + // them would let discovery declare a slow-but-alive portal unreachable. + const countsTowardsGuard = + getQueryString(req, 'skipConnectionGuard') !== 'true'; + let guardToken: HostRequestToken | null = null; + let requestUrl: string | undefined; try { // `macAddress`, `token` and `serialNumber` are portal credentials, // not protocol content: they reach the portal only as the same @@ -242,7 +376,15 @@ export function createWebBackendApp( // a query param and is answered without authentication. const params: Record = getProxyParams( req, - ['targetId', 'macAddress', 'token', 'serialNumber'] + [ + 'targetId', + 'macAddress', + 'token', + 'serialNumber', + // Guard control flag, not protocol content — it must never + // reach the portal's query string. + 'skipConnectionGuard', + ] ); if (params['action'] === 'handshake' && token) { params['token'] = token; @@ -265,18 +407,52 @@ export function createWebBackendApp( delete headers['Cookie']; } + requestUrl = buildStalkerRequestUrl( + url.href, + identity.requestParams + ); + + if (countsTowardsGuard) { + const admission = admitProviderRequest(hostGuard, requestUrl); + if (!admission.allowed) { + // This route answers provider failures with HTTP 200 and + // an error body; a fast-fail is one of them. + res.json(admission.error); + return; + } + guardToken = admission.token; + } else { + // Endpoint discovery: never policed, never counted, but a + // candidate that answers still clears the record. + guardToken = observeProviderRequest(hostGuard, requestUrl); + } + // Provider URLs are validated by /provider-targets before they enter the registry. // codeql[js/request-forgery] - const response = await httpClient.get( - buildStalkerRequestUrl(url.href, identity.requestParams), - { headers } - ); + const response = await httpClient.get(requestUrl, { + headers, + // `create_link` gets the longer budget: the portal mints a + // stream URL before it answers. + timeout: + params['action'] === 'create_link' + ? PROVIDER_REQUEST_TIMEOUT_MS.stalkerCreateLink + : PROVIDER_REQUEST_TIMEOUT_MS.stalker, + }); + reportProviderRequestSuccess(hostGuard, guardToken); + guardToken = null; res.json({ action: getQueryString(req, 'action'), payload: response.data, }); } catch (error) { + // Always reported, even for exempt discovery probes: a failure + // carrying an HTTP response proves the endpoint answered, and + // dropping that is what lets the breaker open mid-discovery. + reportProviderRequestFailure(hostGuard, guardToken, error, { + countFailures: countsTowardsGuard, + requestUrl, + }); logProviderRequestFailure({ error, route: '/stalker', url }); res.json(normalizeProviderError(error)); } @@ -350,10 +526,11 @@ async function validateProviderUrl( let addresses: readonly string[]; try { addresses = await policy.resolveHostname(hostname); - } catch { + } catch (lookupError) { return { message: 'Provider URL host could not be resolved', status: 400, + lookupError, }; } @@ -480,7 +657,9 @@ async function handlePlaylistParse(options: { try { // Provider URLs are validated by /provider-targets before playlist parsing. // codeql[js/request-forgery] - const response = await options.httpClient.get(options.url); + const response = await options.httpClient.get(options.url, { + timeout: PROVIDER_REQUEST_TIMEOUT_MS.playlist, + }); const parsedPlaylist = parsePlaylist(response.data); const title = getLastUrlSegment(options.url); return createPlaylistObject({ @@ -518,6 +697,7 @@ async function fetchEpgDataFromUrl( // Provider URLs are validated by /provider-targets before XMLTV parsing. // codeql[js/request-forgery] const response = await httpClient.get(href, { + timeout: PROVIDER_REQUEST_TIMEOUT_MS.epg, ...(url.pathname.endsWith('.gz') ? { responseType: 'arraybuffer' } : {}), diff --git a/apps/web/src/app/services/pwa.service.spec.ts b/apps/web/src/app/services/pwa.service.spec.ts index cbb262674..c81cdff1e 100644 --- a/apps/web/src/app/services/pwa.service.spec.ts +++ b/apps/web/src/app/services/pwa.service.spec.ts @@ -9,10 +9,16 @@ import { Store } from '@ngrx/store'; import { TranslateService } from '@ngx-translate/core'; import { config as rxjsConfig, EMPTY } from 'rxjs'; import { + buildHostConnectivityFastFailMessage, + CONNECTIVITY_GUARD_RESET, PLAYLIST_PARSE_BY_URL, PLAYLIST_UPDATE, STALKER_REQUEST, } from '@iptvnator/shared/interfaces'; +import { + getStalkerRequestErrorStatus, + isStalkerProbeTimeout, +} from '@iptvnator/portal/stalker/data-access'; import { PwaService } from './pwa.service'; describe('PwaService', () => { @@ -238,4 +244,191 @@ describe('PwaService', () => { delete globalWithFetch.fetch; } }); + + it('surfaces a connectivity-guard fast-fail as a status-less connection failure', async () => { + // The backend refuses to contact a host that stopped answering, using + // the same HTTP 200 `{ message, status }` envelope as every other + // provider failure. It must NOT come out of here as the test above + // does: an `HTTP Error ` message or a numeric `status` both read + // as "the endpoint answered", so endpoint discovery would keep walking + // candidates, and a 4xx reading would fire lazy portal repair against + // a host just declared unreachable. + const message = buildHostConnectivityFastFailMessage( + 'portal.example:8080' + ); + const originalFetch = globalThis.fetch; + globalThis.fetch = jest.fn().mockResolvedValue({ + ok: true, + json: async () => ({ message, status: 502 }), + } as unknown as Response) as unknown as typeof fetch; + + const request = service.sendIpcEvent(STALKER_REQUEST, { + url: 'http://portal.example:8080/portal.php', + macAddress: '00:1A:79:AA:BB:CC', + params: { action: 'get_genres' }, + silent: true, + }) as Promise; + const outcome = request.then( + () => { + throw new Error('expected rejection'); + }, + (error: unknown) => error + ); + + await Promise.resolve(); + http.expectOne((candidate) => + candidate.url.endsWith('/provider-targets') + ).flush({ targetId: 'target-1' }); + + const error = (await outcome) as Error & { status?: number }; + expect(error.message).toBe(message); + expect(error.status).toBeUndefined(); + expect(getStalkerRequestErrorStatus(error)).toBeUndefined(); + expect(isStalkerProbeTimeout(error)).toBe(false); + + globalThis.fetch = originalFetch; + }); + + it('carries the discovery guard bypass through to the /stalker proxy', async () => { + // The Electron transport honours `skipConnectionGuard` on the payload. + // Dropping it here let two probe timeouts open the backend's breaker + // and fast-fail the healthy candidate discovery was looking for. + const fetchMock = jest.fn().mockResolvedValue({ + ok: true, + json: () => Promise.resolve({ payload: { js: [] } }), + } as unknown as Response); + const globalWithFetch = globalThis as { fetch?: typeof fetch }; + globalWithFetch.fetch = fetchMock as unknown as typeof fetch; + + try { + const request = service.forwardStalkerRequest({ + url: 'http://portal.example/portal.php', + macAddress: '00:1A:79:AA:BB:CC', + params: { action: 'get_genres' }, + silent: true, + skipConnectionGuard: true, + }); + + http.expectOne((req) => + req.url.endsWith('/provider-targets') + ).flush({ targetId: 'target-1' }); + + await expect(request).resolves.toEqual({ js: [] }); + + const requestUrl = new URL(String(fetchMock.mock.calls[0][0])); + expect(requestUrl.searchParams.get('skipConnectionGuard')).toBe( + 'true' + ); + } finally { + delete globalWithFetch.fetch; + } + }); + + it('omits the discovery guard bypass for ordinary Stalker requests', async () => { + const fetchMock = jest.fn().mockResolvedValue({ + ok: true, + json: () => Promise.resolve({ payload: { js: [] } }), + } as unknown as Response); + const globalWithFetch = globalThis as { fetch?: typeof fetch }; + globalWithFetch.fetch = fetchMock as unknown as typeof fetch; + + try { + const request = service.forwardStalkerRequest({ + url: 'http://portal.example/portal.php', + macAddress: '00:1A:79:AA:BB:CC', + params: { action: 'get_categories' }, + }); + + http.expectOne((req) => + req.url.endsWith('/provider-targets') + ).flush({ targetId: 'target-1' }); + + await expect(request).resolves.toEqual({ js: [] }); + + const requestUrl = new URL(String(fetchMock.mock.calls[0][0])); + expect(requestUrl.searchParams.has('skipConnectionGuard')).toBe( + false + ); + } finally { + delete globalWithFetch.fetch; + } + }); + + it('gives up on a reset the backend never answers', async () => { + // Every caller awaits the reset BEFORE the request it is clearing the + // way for, and `fetch` has no timeout of its own — so a backend that + // accepts the POST and goes quiet would leave Retry doing nothing at + // all, which is the failure this whole change exists to stop. + jest.useFakeTimers(); + const globalWithFetch = globalThis as { fetch?: typeof fetch }; + let capturedSignal: AbortSignal | undefined; + globalWithFetch.fetch = jest.fn((_url: unknown, init: RequestInit) => { + capturedSignal = init.signal ?? undefined; + // Never settles on its own; only the abort can end this. + return new Promise((_resolve, reject) => { + init.signal?.addEventListener('abort', () => + reject(new DOMException('Aborted', 'AbortError')) + ); + }); + }) as unknown as typeof fetch; + + try { + const request = service.sendIpcEvent(CONNECTIVITY_GUARD_RESET, { + url: 'http://portal.example:8080/portal.php', + }) as Promise; + const outcome = request.then( + () => 'resolved', + (error: unknown) => (error as Error).name + ); + + expect(capturedSignal).toBeDefined(); + expect(capturedSignal?.aborted).toBe(false); + + jest.advanceTimersByTime(5_000); + + await expect(outcome).resolves.toBe('AbortError'); + expect(capturedSignal?.aborted).toBe(true); + } finally { + jest.useRealTimers(); + delete globalWithFetch.fetch; + } + }); + + it('asks the backend to forget a host when the connectivity guard is reset', async () => { + const fetchMock = jest.fn().mockResolvedValue({ + ok: true, + json: async () => ({ reset: true }), + } as unknown as Response); + const globalWithFetch = globalThis as { fetch?: typeof fetch }; + globalWithFetch.fetch = fetchMock as unknown as typeof fetch; + + try { + // The breaker lives in the backend process, not the browser, so + // the reset has to travel over HTTP. Before this it was a no-op, + // and a PWA user's "retry" never cleared anything. + await expect( + service.sendIpcEvent(CONNECTIVITY_GUARD_RESET, { + url: 'http://portal.example:8080/portal.php', + }) as Promise<{ reset: boolean }> + ).resolves.toEqual({ reset: true }); + + expect(fetchMock).toHaveBeenCalledTimes(1); + const [url, init] = fetchMock.mock.calls[0]; + expect(String(url)).toContain('/connectivity-guard/reset'); + expect(init.method).toBe('POST'); + expect(JSON.parse(init.body)).toEqual({ + url: 'http://portal.example:8080/portal.php', + }); + + // Nothing to reset without a URL, and nothing sent either. + await expect( + service.sendIpcEvent(CONNECTIVITY_GUARD_RESET, {}) as Promise<{ + reset: boolean; + }> + ).resolves.toEqual({ reset: false }); + expect(fetchMock).toHaveBeenCalledTimes(1); + } finally { + delete globalWithFetch.fetch; + } + }); }); diff --git a/apps/web/src/app/services/pwa.service.ts b/apps/web/src/app/services/pwa.service.ts index c4a2c6f9f..4f7fed12b 100644 --- a/apps/web/src/app/services/pwa.service.ts +++ b/apps/web/src/app/services/pwa.service.ts @@ -15,7 +15,9 @@ import { } from 'rxjs'; import { DataService } from '@iptvnator/services'; import { + CONNECTIVITY_GUARD_RESET, ERROR, + isHostConnectivityFastFailMessage, Playlist, PLAYLIST_PARSE_BY_URL, PLAYLIST_UPDATE, @@ -35,6 +37,16 @@ import { } from '@iptvnator/portal/shared/util'; import { getRuntimeBackendUrl } from './runtime-config'; +/** + * How long to wait for the backend to forget a host before giving up. + * + * This talks to the user's own backend, not a provider, so it should answer + * immediately; the bound exists so a stuck one cannot hold up the retry that + * asked for the reset. Short on purpose — the reset is best effort, and the + * caller's own request reports the real state either way. + */ +const CONNECTIVITY_GUARD_RESET_TIMEOUT_MS = 5_000; + interface PwaXtreamResponse { readonly payload?: unknown; readonly status?: number; @@ -140,13 +152,69 @@ export class PwaService extends DataService { token?: string; serialNumber?: string; silent?: boolean; + skipConnectionGuard?: boolean; } ) as T; } + if (type === CONNECTIVITY_GUARD_RESET) { + return this.resetConnectivityGuard( + payload as { url?: string } + ) as T; + } + return undefined as T; } + /** + * Clears the web backend's per-host failure record, so the next request + * contacts the host for real instead of being fast-failed. + * + * The breaker lives in the backend process, not the browser, so this is an + * HTTP call rather than a local reset. Best effort by design — the caller's + * own request reports the real state, and `resetHostConnectivityGuard` + * swallows whatever this rejects with. + */ + private async resetConnectivityGuard(payload: { + url?: string; + }): Promise<{ reset: boolean }> { + if (!payload?.url) { + return { reset: false }; + } + + // Bounded, because every caller awaits this BEFORE issuing the request + // it is clearing the way for. `fetch` has no timeout of its own, so a + // backend or reverse proxy that accepts the POST and then goes quiet + // would leave Retry doing nothing at all — the failure this whole + // change exists to stop, reintroduced one layer up. The abort rejects, + // `resetHostConnectivityGuard` swallows it, and the caller proceeds. + const controller = new AbortController(); + const abortTimer = setTimeout( + () => controller.abort(), + CONNECTIVITY_GUARD_RESET_TIMEOUT_MS + ); + + try { + const response = await fetch( + `${this.corsProxyUrl}/connectivity-guard/reset`, + { + body: JSON.stringify({ url: payload.url }), + headers: { 'content-type': 'application/json' }, + method: 'POST', + signal: controller.signal, + } + ); + + if (!response.ok) { + return { reset: false }; + } + + return (await response.json()) as { reset: boolean }; + } finally { + clearTimeout(abortTimer); + } + } + refreshPlaylist(payload?: Partial) { if (!payload?.url || !payload?.id) { return; @@ -503,6 +571,13 @@ export class PwaService extends DataService { serialNumber?: string; /** Endpoint-discovery probes expect failures; no error snackbar. */ silent?: boolean; + /** + * Endpoint-discovery probes are exempt from the backend's per-host + * connectivity guard: they walk several candidates on one host and + * expect most to fail, so counting them would declare a working portal + * unreachable. Mirrors the Electron `STALKER_REQUEST` payload flag. + */ + skipConnectionGuard?: boolean; }) { let context = createPortalDebugRequestContext({ provider: 'stalker', @@ -530,6 +605,13 @@ export class PwaService extends DataService { ...(payload.serialNumber ? { serialNumber: payload.serialNumber } : {}), + // Endpoint-discovery probes are exempt from the backend's + // connectivity guard. Dropping this here would let two probes + // that time out open the breaker and fast-fail the healthy + // candidate discovery is looking for. + ...(payload.skipConnectionGuard + ? { skipConnectionGuard: 'true' } + : {}), }; const params = new URLSearchParams(requestParams); const requestUrl = `${this.corsProxyUrl}/stalker?${params.toString()}`; @@ -564,6 +646,23 @@ export class PwaService extends DataService { // can classify them — unwrapping `payload` here silently // returned `undefined`, making a dead endpoint look like an // empty answer and unreachable to the repair. + // The connectivity guard refused to contact the host. This one must + // NOT go through the branch below: an `HTTP Error ` prefix + // and a numeric `status` both read as "the endpoint answered", so + // endpoint discovery would keep walking candidates instead of + // aborting, and a 4xx reading would additionally fire lazy portal + // repair against a host just declared unreachable. Thrown bare, the + // message lands in the connection-level-failure slot — which is + // exactly what a tripped breaker means. + if ( + responseBody && + typeof responseBody === 'object' && + !('payload' in responseBody) && + isHostConnectivityFastFailMessage(responseBody.message) + ) { + throw new Error(String(responseBody.message)); + } + if ( responseBody && typeof responseBody === 'object' && diff --git a/docs/architecture/host-connectivity-guard.md b/docs/architecture/host-connectivity-guard.md index 857115f76..3ee0f8ba9 100644 --- a/docs/architecture/host-connectivity-guard.md +++ b/docs/architecture/host-connectivity-guard.md @@ -1,31 +1,73 @@ # Host Connectivity Guard -Per-host circuit breaker for portal requests in the Electron main process. +Per-host circuit breaker for portal requests, in both processes that make them: +the Electron main process and the self-hosted web backend. ## The problem Every request to an unreachable portal costs its full axios timeout — 30 s for -`XTREAM_REQUEST`, 15 s for `STALKER_REQUEST` (30 s for `create_link`). Browsing a +`XTREAM_REQUEST`, 15 s for `STALKER_REQUEST` (30 s for `create_link`), and the +same budgets on the web backend's `/xtream` and `/stalker` routes. Browsing a dead portal's catalog issues dozens of those back to back, which shows up as -30-second spinners and a main-process log full of identical failures. Once a host -has refused to answer twice in a row there is nothing left to learn from waiting -again. +30-second spinners and a log full of identical failures. Once a host has refused +to answer twice in a row there is nothing left to learn from waiting again. + +The web backend had a worse version of the same problem first: its proxy routes +passed no `timeout` at all, so a provider that accepted a connection and then +went silent held the request until the OS gave up on the TCP connection. A +breaker is only useful once _not answering_ is bounded, which is why the timeouts +and the breaker landed there together. ## Where it lives -`apps/electron-backend/src/app/util/host-connectivity-guard.ts` — a pure module -(no Electron imports) wired into both IPC handlers. +`libs/shared/host-health` (`@iptvnator/shared/host-health`, tagged +`scope:shared` / `domain:shared-runtime` / `type:util`) — the breaker class, the +failure classification and the redirect-attribution helpers, with no transport, +logger or process singleton of its own. The owning app supplies the clock and +decides how many guards exist. -The handlers are the choke point that sees _all_ traffic to a host. The -renderer's `executeStalkerRequest` is not: it has four documented bypasses -(auth, endpoint discovery, account info, the row-less stream resolver), and that -is exactly the traffic that hits dead hosts. +Each runtime owns its instance: -**The PWA is deliberately not covered yet.** `apps/web-backend` sets no -per-request timeout at all, so a dead host there hangs on OS-level TCP timeouts -rather than a 15/30 s budget — a timeout-driven breaker would rarely trip. The -guard is written dependency-free so it can move into a `domain:shared-runtime` -library and be shared with `web-backend` when that gap is closed. +- **Electron** — `apps/electron-backend/src/app/util/host-connectivity-guard.ts` + holds one guard for the whole process, wired into both IPC handlers. The + handlers are the choke point that sees _all_ traffic to a host. The renderer's + `executeStalkerRequest` is not: it has four documented bypasses (auth, + endpoint discovery, account info, the row-less stream resolver), and that is + exactly the traffic that hits dead hosts. +- **PWA** — `apps/web-backend/src/app/host-guard.ts` guards the `/xtream` and + `/stalker` proxy routes. The guard instance is injected through + `WebBackendAppOptions.hostGuard`, alongside `now` and `guid`, so specs drive + it with a clock they own. + +### Request timeouts are the precondition + +A breaker keyed on "the endpoint did not answer" is only useful once _not +answering_ is bounded. `apps/web-backend` originally passed no `timeout` at all, +so a silent host hung on OS-level TCP timeouts. Both runtimes now use the same +budgets: Xtream 30 s, Stalker 15 s (30 s for `create_link`, which mints a stream +URL before answering), playlist and XMLTV downloads 30 s. + +Those numbers are safe for large downloads. On axios' default +(follow-redirects) transport `timeout` is **not** a wall-clock deadline for the +whole response: it bounds the time to response headers and then continues as the +socket's inactivity timeout for the body. A multi-megabyte XMLTV file that keeps +delivering bytes is never cut off mid-transfer — only a stalled one is. + +### Scope: portal calls only + +Both runtimes guard the portal API paths and nothing else. Playlist and XMLTV +downloads (`/parse`, `/parse-xml`, and their Electron equivalents) get the +timeouts but not the breaker, deliberately: + +- The problem being solved is a catalog fan-out — dozens of requests to one + endpoint back to back. A download is a single request. +- A download is usually the direct result of the user asking for it (the + add-playlist dialog sends `PLAYLIST_PARSE_BY_URL`). Refusing an immediate + retry is a regression, not a protection, and there is no natural reset site on + that path the way portal Retry has one. +- A large XMLTV transfer can legitimately run for minutes — the timeout is + idle-based, see above — which outlives the 45 s half-open trial expiry and + would let a second trial in behind the first. ## Rules @@ -48,25 +90,41 @@ mid-transfer happens on hosts that are very much alive. Cancelled requests (`ERR_CANCELED`) and SSRF-policy refusals are `inconclusive` — they say nothing about reachability and only release the half-open slot. -**A failure is only charged to the endpoint that produced it.** Redirects are -followed hop by hop, each with its own config, so a failure on a later hop -carries that hop's URL in `error.config.url`. Reaching any later hop _proves_ the -guarded endpoint answered — the first hop is always the URL we asked for, and -only a redirect status advances the chain — so it CLEARS the guarded endpoint's -record, exactly like any other response. Merely declining to count it would leave -an earlier direct failure standing, and a single later timeout would then -fast-fail an endpoint that answered in between. +**A failure is only charged to the endpoint that produced it.** Reaching any +later hop _proves_ the guarded endpoint answered — the first hop is always the +URL we asked for, and only a redirect status advances the chain — so a failure +there CLEARS the guarded endpoint's record, exactly like any other response. +Merely declining to count it would leave an earlier direct failure standing, and +a single later timeout would then fast-fail an endpoint that answered in between. +Every caller passes the URL it asked for as the baseline. -The comparison is against the whole request URL, not just its origin: a -same-origin redirect (`/player_api.php` → `/slow/player_api.php`) proves the -endpoint answered just as much as a cross-origin one, and charging it would -fast-fail every OTHER call to a portal that answers. Both handlers therefore pass -the URL they asked for. It requires positive evidence — anything unparseable or -unknown counts the failure as usual, because guessing "redirect" here would stop -the guard from ever tripping — and a failure that names no URL at all is still -counted. Round-tripping through `URL` is identity for both handlers' URL shapes -(including Stalker's hand-encoded `cmd`), and `requestWithValidatedRedirects` -normalizes hop 1 the same way, so the comparison is exact. +**Where the failed hop is found depends on the transport, and both are in play.** + +| | Electron | Web backend | +| --- | --- | --- | +| Redirects | followed hop by hop (`maxRedirects: 0`), each its own request | followed inside one request by follow-redirects | +| Failed hop is in | `error.config.url` | `error.request._currentUrl` | + +`failedRequestUrlOf` reads `_currentUrl` first and falls back to `config.url`, +which is correct for both: a per-hop request exposes no `_currentUrl`, and on the +following transport `config` is built once and keeps the URL we asked for — so +reading `config.url` there would compare a URL with itself, find no redirect, and +charge a dead destination to the provider that answered. Anything added here must +work on both, because the same helper serves both. + +**The comparison is origin + path, not the whole URL.** A same-origin redirect +(`/player_api.php` → `/slow/player_api.php`) proves the endpoint answered just as +much as a cross-origin one, and charging it would fast-fail every OTHER call to a +portal that answers — so the path has to be part of it. The query must NOT be: +the web backend passes Xtream credentials through axios' `params`, so the sent +URL always carries a query the baseline does not, and comparing whole URLs made +every ordinary failure look like a redirect and stopped the breaker from ever +opening. What that gives up is a redirect that changes nothing but the query, +which is then counted as an ordinary failure — the safe direction. + +It requires positive evidence — anything unparseable or unknown counts the +failure as usual, because guessing "redirect" here would stop the guard from ever +tripping — and a failure that names no URL at all is still counted. Known gap: the failing hop is not guarded either (it has no token of its own), so a permanently broken redirect chain keeps costing a full timeout. @@ -160,6 +218,13 @@ body does. That is why the exempt path reports through `reportGuardedHostFailure(token, error, { countFailures: false })` rather than skipping the report. +The flag has to survive the PWA transport too, or discovery is exempt on the +desktop and policed on the web. `PwaService.forwardStalkerRequest` forwards it +as a `/stalker` control param and the route applies the same semantics — and, +like `macAddress`/`token`/`serialNumber`, it is stripped from the query +forwarded to the portal, because it is our control flag and not protocol +content. + ## Explicit reset `CONNECTIVITY_GUARD_RESET` (`{ url }`) is handled by @@ -223,9 +288,14 @@ the next `nearEnd` event fires a second retry. They all go through `resetHostConnectivityGuard()` (`libs/services/src/lib/host-connectivity-reset.ts`), which holds the one rule they share: the reset is best effort, because the guard only ever _delays_ a -request and a failed reset must not block the action that asked for it. In the -PWA the channel is unknown and `sendIpcEvent` no-ops, which is correct — -nothing there records per-host failures yet. +request and a failed reset must not block the action that asked for it. Both +runtimes honour it, so neither keeps fast-failing an endpoint the user just +asked to retry — in the PWA `PwaService` forwards it to the web backend's +`POST /connectivity-guard/reset`, since the breaker lives in the backend +process and a local reset would clear nothing. That route takes the raw +provider URL rather than a registered `targetId`, because callers reset +precisely when the address may have changed, which is before any target exists +for it; it fetches nothing, reads only the origin, and never logs the URL. ## Interaction with VOD multi-source @@ -240,8 +310,20 @@ module's contract that unreachable ≠ contacted-and-refused. The ## Tests +- `libs/shared/host-health/src/lib/host-connectivity-guard.spec.ts` — the state + machine, with an injected clock. - `apps/electron-backend/src/app/util/host-connectivity-guard.spec.ts` — the - state machine, with an injected clock. + main-process singleton and redirect attribution through it. +- `apps/web-backend/src/app/web-backend-app.host-guard.spec.ts` — the proxy + routes: the HTTP 200 refusal shape, that the outbound request really is + skipped, that an open endpoint is refused before a DNS lookup is spent on it, + redirect attribution, the discovery exemption, half-open, the reset endpoint, + and that playlist/EPG downloads are never fast-failed. The per-route timeouts + are asserted in `web-backend-app.spec.ts`, which pins the whole outbound + request shape. +- `apps/web/src/app/services/pwa.service.spec.ts` — that a fast-fail reaches the + renderer with no numeric `status` and no `HTTP Error ` in its message, + that the discovery bypass is forwarded, and that a reset calls the backend. - `apps/electron-backend/src/app/events/stalker.events.spec.ts` and `xtream.events.spec.ts` — trip, fast-fail without contacting axios, reset, exemption, and the absence of per-request log spam. diff --git a/libs/services/src/lib/host-connectivity-reset.ts b/libs/services/src/lib/host-connectivity-reset.ts index a05c61a35..678d46885 100644 --- a/libs/services/src/lib/host-connectivity-reset.ts +++ b/libs/services/src/lib/host-connectivity-reset.ts @@ -11,9 +11,11 @@ import type { DataService } from './data.service'; * (import, edited connection, lazy repair). * * Always best effort: the guard only ever delays a request, so a failure here - * must never block the action that asked for it. In the PWA the channel is - * unknown and `sendIpcEvent` no-ops, which is correct — nothing there records - * per-host failures yet. + * must never block the action that asked for it. Both runtimes honour it — the + * Electron main process over IPC, the PWA as an HTTP call to the web backend + * that owns its guard — so neither keeps fast-failing an endpoint the user just + * asked to retry. A new caller has to reach this helper on both paths; there is + * no longer a runtime where it quietly does nothing. */ export async function resetHostConnectivityGuard( dataService: Pick, diff --git a/libs/shared/host-health/jest.config.ts b/libs/shared/host-health/jest.config.ts new file mode 100644 index 000000000..8f3cc16ff --- /dev/null +++ b/libs/shared/host-health/jest.config.ts @@ -0,0 +1,13 @@ +export default { + displayName: 'shared-host-health', + preset: '../../../jest.preset.js', + testEnvironment: 'node', + transform: { + '^.+\\.[tj]s$': [ + 'ts-jest', + { tsconfig: '/tsconfig.spec.json' }, + ], + }, + moduleFileExtensions: ['ts', 'js'], + coverageDirectory: '../../../coverage/libs/shared/host-health', +}; diff --git a/libs/shared/host-health/package.json b/libs/shared/host-health/package.json new file mode 100644 index 000000000..182398860 --- /dev/null +++ b/libs/shared/host-health/package.json @@ -0,0 +1,12 @@ +{ + "name": "@iptvnator/shared/host-health", + "version": "0.0.1", + "private": true, + "type": "commonjs", + "main": "./src/index.js", + "types": "./src/index.d.ts", + "dependencies": { + "@iptvnator/shared/interfaces": "0.0.1", + "tslib": "^2.3.0" + } +} diff --git a/libs/shared/host-health/project.json b/libs/shared/host-health/project.json new file mode 100644 index 000000000..b0a17c4a0 --- /dev/null +++ b/libs/shared/host-health/project.json @@ -0,0 +1,30 @@ +{ + "name": "shared-host-health", + "$schema": "../../../node_modules/nx/schemas/project-schema.json", + "sourceRoot": "libs/shared/host-health/src", + "projectType": "library", + "tags": ["scope:shared", "domain:shared-runtime", "type:util"], + "targets": { + "build": { + "executor": "@nx/js:tsc", + "outputs": ["{options.outputPath}"], + "options": { + "outputPath": "dist/libs/shared/host-health", + "main": "libs/shared/host-health/src/index.ts", + "tsConfig": "libs/shared/host-health/tsconfig.lib.json", + "assets": [] + } + }, + "test": { + "executor": "@nx/jest:jest", + "outputs": ["{workspaceRoot}/coverage/{projectRoot}"], + "options": { + "jestConfig": "libs/shared/host-health/jest.config.ts", + "tsConfig": "libs/shared/host-health/tsconfig.spec.json" + } + }, + "lint": { + "executor": "@nx/eslint:lint" + } + } +} diff --git a/libs/shared/host-health/src/index.ts b/libs/shared/host-health/src/index.ts new file mode 100644 index 000000000..cc6c59bb7 --- /dev/null +++ b/libs/shared/host-health/src/index.ts @@ -0,0 +1 @@ +export * from './lib/host-connectivity-guard'; diff --git a/libs/shared/host-health/src/lib/host-connectivity-guard.spec.ts b/libs/shared/host-health/src/lib/host-connectivity-guard.spec.ts new file mode 100644 index 000000000..cd06f2786 --- /dev/null +++ b/libs/shared/host-health/src/lib/host-connectivity-guard.spec.ts @@ -0,0 +1,540 @@ +import { + buildHostConnectivityFastFailMessage, + isHostConnectivityFastFailMessage, +} from '@iptvnator/shared/interfaces'; +import { + HostConnectivityGuard, + HostConnectivityGuardError, + HostRequestToken, + classifyHostRequestFailure, + failedAfterRedirect, + portalEndpointKeyOf, +} from './host-connectivity-guard'; + +const HOST = 'http://portal.example.com:8080'; +const GUARD_DISABLED_ENV = 'IPTVNATOR_DISABLE_CONNECTIVITY_GUARD'; +const OPEN_DURATION_MS = 30_000; +const FAILURE_WINDOW_MS = 120_000; +const TRIAL_TIMEOUT_MS = 45_000; + +function timeoutError(code = 'ETIMEDOUT') { + return Object.assign(new Error('timeout of 30000ms exceeded'), { code }); +} + +describe('classifyHostRequestFailure', () => { + it('treats connection-level error codes as host-level failures', () => { + for (const code of [ + 'ECONNABORTED', + 'ECONNREFUSED', + 'EAI_AGAIN', + 'EHOSTUNREACH', + 'ENETUNREACH', + 'ENOTFOUND', + 'ETIMEDOUT', + ]) { + expect(classifyHostRequestFailure(timeoutError(code))).toBe( + 'host-level' + ); + } + }); + + it('treats an error carrying an HTTP response as proof the host answered', () => { + // 5xx reaches the handlers as a rejection: validateStatus only + // tolerates < 500. The host still answered. + const serverError = Object.assign(new Error('Request failed'), { + code: 'ERR_BAD_RESPONSE', + response: { status: 502, statusText: 'Bad Gateway' }, + }); + + expect(classifyHostRequestFailure(serverError)).toBe('responded'); + }); + + it('does not read reachability into cancellations, SSRF refusals or plain errors', () => { + expect( + classifyHostRequestFailure( + Object.assign(new Error('canceled'), { code: 'ERR_CANCELED' }) + ) + ).toBe('inconclusive'); + expect( + classifyHostRequestFailure( + new Error('URL host could not be resolved') + ) + ).toBe('inconclusive'); + expect(classifyHostRequestFailure(undefined)).toBe('inconclusive'); + expect(classifyHostRequestFailure('ETIMEDOUT')).toBe('inconclusive'); + }); + + it('never counts a mid-transfer connection reset', () => { + // A reset happens on hosts that are very much alive. + expect(classifyHostRequestFailure(timeoutError('ECONNRESET'))).toBe( + 'inconclusive' + ); + }); +}); + +describe('failedAfterRedirect', () => { + // Two transports put the failed hop in two different places, and reading + // only one of them silently mis-attributes on the other. + const ENDPOINT = 'http://panel.example:8080'; + const ASKED_FOR = `${ENDPOINT}/player_api.php`; + const token: HostRequestToken = { + endpoint: ENDPOINT, + epoch: 0, + startedAt: 0, + trial: false, + trialId: 0, + }; + + it('reads the hop from request._currentUrl when redirects were followed internally', () => { + // axios' default (follow-redirects) transport: `config` is built once + // and keeps the URL we asked for, whatever the chain did afterwards. + const error = { + code: 'ECONNREFUSED', + config: { url: ASKED_FOR }, + request: { _currentUrl: 'http://cdn.dead.example/stream' }, + }; + + expect(failedAfterRedirect(error, token, ASKED_FOR)).toBe(true); + }); + + it('falls back to config.url for a transport that reissues each hop', () => { + // The Electron transport uses `maxRedirects: 0` and follows redirects + // itself, so there is no `_currentUrl` and the hop is the config URL. + const error = { + code: 'ECONNREFUSED', + config: { url: 'http://cdn.dead.example/stream' }, + }; + + expect(failedAfterRedirect(error, token, ASKED_FOR)).toBe(true); + }); + + it('reports no redirect when the request failed against the URL we asked for', () => { + const error = { + code: 'ECONNREFUSED', + config: { url: ASKED_FOR }, + request: { _currentUrl: ASKED_FOR }, + }; + + expect(failedAfterRedirect(error, token, ASKED_FOR)).toBe(false); + }); +}); + +describe('portalEndpointKeyOf', () => { + it('keys on host and port so two panels on one machine stay separate', () => { + expect( + portalEndpointKeyOf('http://example.com:8080/player_api.php') + ).toBe('http://example.com:8080'); + expect( + portalEndpointKeyOf('http://example.com:9090/player_api.php') + ).toBe('http://example.com:9090'); + }); + + it('separates HTTP from HTTPS on their default ports', () => { + // `URL.host` omits a default port, so both would collapse onto + // `example.com` — and a panel whose TLS listener is broken while plain + // HTTP works is a routine IPTV setup. + expect(portalEndpointKeyOf('http://example.com/player_api.php')).toBe( + 'http://example.com' + ); + expect(portalEndpointKeyOf('https://example.com/player_api.php')).toBe( + 'https://example.com' + ); + }); + + it('leaves URL credentials out of the key', () => { + expect(portalEndpointKeyOf('http://user:pass@example.com/c')).toBe( + 'http://example.com' + ); + }); + + it('returns null for an unparseable URL instead of throwing', () => { + expect(portalEndpointKeyOf('not a url')).toBeNull(); + }); +}); + +describe('HostConnectivityGuardError', () => { + it('carries the shared fast-fail message and no status field', () => { + const error = new HostConnectivityGuardError(HOST); + + expect(error).toBeInstanceOf(Error); + expect(error.message).toBe(buildHostConnectivityFastFailMessage(HOST)); + // getStalkerRequestErrorStatus reads `status` first; a number there + // would read as "the endpoint answered". + expect( + (error as unknown as { status?: unknown }).status + ).toBeUndefined(); + expect(isHostConnectivityFastFailMessage(error.message)).toBe(true); + }); +}); + +describe('HostConnectivityGuard', () => { + let clock = 1_000_000; + let opened: string[]; + let guard: HostConnectivityGuard; + + const advance = (ms: number) => { + clock += ms; + }; + + /** Runs a request that fails at the host level after `durationMs`. */ + const failRequest = (durationMs = 0): void => { + const check = guard.check(HOST); + if (!check.allowed) { + throw new Error('expected the request to be allowed'); + } + advance(durationMs); + guard.reportFailure(check.token); + }; + + const expectAllowed = (): HostRequestToken => { + const check = guard.check(HOST); + expect(check.allowed).toBe(true); + if (!check.allowed) { + throw new Error('unreachable'); + } + return check.token; + }; + + const expectBlocked = (): void => { + expect(guard.check(HOST).allowed).toBe(false); + }; + + beforeEach(() => { + clock = 1_000_000; + opened = []; + delete process.env[GUARD_DISABLED_ENV]; + guard = new HostConnectivityGuard({ + now: () => clock, + onOpen: (host) => opened.push(host), + }); + }); + + it('allows requests to a host it has never seen', () => { + expect(guard.check(HOST).allowed).toBe(true); + expect(opened).toEqual([]); + }); + + it('still allows the request after a single failure', () => { + failRequest(30_000); + + expect(guard.check(HOST).allowed).toBe(true); + expect(opened).toEqual([]); + }); + + it('opens after two consecutive failures and reports the wait', () => { + failRequest(30_000); + failRequest(30_000); + + expect(guard.check(HOST)).toEqual({ + allowed: false, + retryAfterMs: OPEN_DURATION_MS, + }); + expect(opened).toEqual([HOST]); + }); + + it('logs the open transition once, not per blocked request', () => { + failRequest(30_000); + failRequest(30_000); + expectBlocked(); + expectBlocked(); + + expect(opened).toEqual([HOST]); + }); + + it('keeps other hosts untouched', () => { + failRequest(30_000); + failRequest(30_000); + + expect(guard.check('http://other.example.com').allowed).toBe(true); + }); + + it('keeps the same host on another scheme reachable', () => { + // Regression: keying by `URL.host` dropped the default port, so a dead + // HTTPS panel fast-failed the working HTTP one on the same machine. + failRequest(30_000); + failRequest(30_000); + expectBlocked(); + + expect(guard.check('https://portal.example.com:8080').allowed).toBe( + true + ); + }); + + it('counts a parallel fan-out that fails together as one failure', () => { + // Catalog init loads live/vod/series at once. One wifi hiccup failing + // all three is one piece of evidence, not a trip. + const first = expectAllowed(); + const second = expectAllowed(); + const third = expectAllowed(); + + advance(30_000); + guard.reportFailure(first); + guard.reportFailure(second); + guard.reportFailure(third); + + expect(guard.check(HOST).allowed).toBe(true); + expect(opened).toEqual([]); + }); + + it('opens when a second fan-out fails after the first one did', () => { + const first = expectAllowed(); + const second = expectAllowed(); + advance(30_000); + guard.reportFailure(first); + guard.reportFailure(second); + + failRequest(30_000); + + expectBlocked(); + expect(opened).toEqual([HOST]); + }); + + it('does not re-open on a sibling that settles after the window elapsed', () => { + // A and B start together; A plus a later request open the breaker. B is + // explicitly not counted as a strike, so it must not start a fresh + // 30-second window either — that would push the half-open trial past + // the intended cooldown. + const sibling = expectAllowed(); + const first = expectAllowed(); + advance(30_000); + guard.reportFailure(first); + failRequest(30_000); + expectBlocked(); + expect(opened).toEqual([HOST]); + + advance(OPEN_DURATION_MS); + guard.reportFailure(sibling); + + const trial = expectAllowed(); + expect(trial.trial).toBe(true); + expect(opened).toEqual([HOST]); + }); + + it('does not accumulate failures further apart than the streak window', () => { + failRequest(0); + advance(FAILURE_WINDOW_MS + 1); + failRequest(0); + + expect(guard.check(HOST).allowed).toBe(true); + expect(opened).toEqual([]); + }); + + it('accumulates failures exactly at the streak window edge', () => { + failRequest(0); + advance(FAILURE_WINDOW_MS); + failRequest(0); + + expectBlocked(); + }); + + it('clears the record as soon as the host answers', () => { + failRequest(30_000); + const token = expectAllowed(); + guard.reportSuccess(token); + failRequest(30_000); + + expect(guard.check(HOST).allowed).toBe(true); + expect(opened).toEqual([]); + }); + + it('treats a success followed by an inconclusive report as a no-op', () => { + // The Stalker handler reports success on the response, then throws its + // own `HTTP Error 404` error, which is classified inconclusive. + failRequest(30_000); + const token = expectAllowed(); + guard.reportSuccess(token); + guard.reportInconclusive(token); + failRequest(30_000); + + expect(guard.check(HOST).allowed).toBe(true); + }); + + it('ignores an inconclusive failure entirely', () => { + const first = expectAllowed(); + guard.reportInconclusive(first); + const second = expectAllowed(); + guard.reportInconclusive(second); + + expect(guard.check(HOST).allowed).toBe(true); + expect(opened).toEqual([]); + }); + + describe('half-open', () => { + beforeEach(() => { + failRequest(30_000); + failRequest(30_000); + expectBlocked(); + advance(OPEN_DURATION_MS); + }); + + it('lets exactly one request through once the window elapses', () => { + const trial = expectAllowed(); + expect(trial.trial).toBe(true); + + expect(guard.check(HOST).allowed).toBe(false); + expect(guard.check(HOST).allowed).toBe(false); + }); + + it('closes the breaker when the trial succeeds', () => { + const trial = expectAllowed(); + guard.reportSuccess(trial); + + const next = expectAllowed(); + expect(next.trial).toBe(false); + expect(guard.check(HOST).allowed).toBe(true); + }); + + it('re-opens immediately when the trial fails, without a second strike', () => { + const trial = expectAllowed(); + advance(30_000); + guard.reportFailure(trial); + + expect(guard.check(HOST)).toEqual({ + allowed: false, + retryAfterMs: OPEN_DURATION_MS, + }); + expect(opened).toEqual([HOST, HOST]); + }); + + it('offers the trial slot again when the trial is cancelled', () => { + const trial = expectAllowed(); + guard.reportInconclusive(trial); + + const retry = expectAllowed(); + expect(retry.trial).toBe(true); + }); + + it('lets only the request holding the slot release it', () => { + // A trial can outlive its own 45 s window: the validated-redirect + // transport gives each of up to five hops its own 30 s budget. Once + // a replacement has been admitted, the abandoned request's late + // report must not hand the slot to a third request. + const abandoned = expectAllowed(); + advance(TRIAL_TIMEOUT_MS); + const replacement = expectAllowed(); + expect(replacement.trial).toBe(true); + + guard.reportInconclusive(abandoned); + + expectBlocked(); + }); + + it('recovers when a trial never reports back', () => { + expectAllowed(); + expectBlocked(); + + advance(TRIAL_TIMEOUT_MS); + + const replacement = expectAllowed(); + expect(replacement.trial).toBe(true); + }); + }); + + describe('reset', () => { + it('reopens the host for requests immediately', () => { + failRequest(30_000); + failRequest(30_000); + expectBlocked(); + + guard.reset(HOST); + + expect(expectAllowed().trial).toBe(false); + }); + + it('discards failures from requests that started before it', () => { + // The retry that cleared the breaker must not be poisoned by the + // 30-second stragglers it was waiting behind. + const straggler = expectAllowed(); + const secondStraggler = expectAllowed(); + advance(15_000); + guard.reset(HOST); + + advance(15_000); + guard.reportFailure(straggler); + guard.reportFailure(secondStraggler); + + expect(guard.check(HOST).allowed).toBe(true); + expect(opened).toEqual([]); + }); + + it('still counts failures from requests started after it', () => { + failRequest(30_000); + guard.reset(HOST); + failRequest(30_000); + failRequest(30_000); + + expectBlocked(); + }); + + it('supersedes in-flight requests for a host it has no record of', () => { + const token = expectAllowed(); + guard.clear(); + guard.reset(HOST); + advance(30_000); + guard.reportFailure(token); + failRequest(30_000); + + expect(guard.check(HOST).allowed).toBe(true); + }); + }); + + describe('bookkeeping bounds', () => { + it('forgets idle records instead of growing without bound', () => { + for (let index = 0; index < 300; index += 1) { + advance(1); + guard.check(`http://host-${index}.example.com`); + } + + // The cap held, and a fresh host is still allowed through. + expect(guard.check(HOST).allowed).toBe(true); + }); + + it('keeps an open breaker while other hosts churn past the idle TTL', () => { + failRequest(30_000); + failRequest(30_000); + expectBlocked(); + + advance(1); + guard.check('http://noise.example.com'); + + expectBlocked(); + }); + }); + + describe('kill switch', () => { + afterEach(() => { + delete process.env[GUARD_DISABLED_ENV]; + }); + + it('never blocks a request while disabled', () => { + process.env[GUARD_DISABLED_ENV] = '1'; + + failRequest(30_000); + failRequest(30_000); + failRequest(30_000); + + expect(guard.check(HOST).allowed).toBe(true); + expect(opened).toEqual([]); + }); + + it('accepts the "true" spelling', () => { + process.env[GUARD_DISABLED_ENV] = 'true'; + + failRequest(30_000); + failRequest(30_000); + + expect(guard.check(HOST).allowed).toBe(true); + }); + + it('is read per call, so an already open breaker stops blocking', () => { + failRequest(30_000); + failRequest(30_000); + expectBlocked(); + + process.env[GUARD_DISABLED_ENV] = '1'; + + expect(guard.check(HOST).allowed).toBe(true); + }); + }); +}); + diff --git a/libs/shared/host-health/src/lib/host-connectivity-guard.ts b/libs/shared/host-health/src/lib/host-connectivity-guard.ts new file mode 100644 index 000000000..c2c85dae8 --- /dev/null +++ b/libs/shared/host-health/src/lib/host-connectivity-guard.ts @@ -0,0 +1,556 @@ +/** + * Per-host circuit breaker for portal requests. + * + * Shared by the two processes that talk to portals on the user's behalf: the + * Electron main process and the self-hosted web backend. Neither the class nor + * anything below it reaches for a transport, a logger or a process singleton — + * the owning app supplies the clock and decides how many guards exist, which is + * what lets the web backend hand its Express app a per-instance guard in tests. + * + * A dead portal costs the full axios timeout on every call (30 s for Xtream, + * 15 s for Stalker), and browsing a dead portal's catalog issues dozens of + * those back to back — 30-second spinners and a flooded log. Once a host has + * refused to answer twice in a row there is nothing left to learn from waiting + * again, so subsequent requests fail immediately for a short while. + * + * The rules are deliberately timid, because being wrong here means refusing to + * talk to a portal that works: + * + * - Only connection-level evidence counts (see + * {@link classifyHostRequestFailure}). Any HTTP response — 200, 404, even + * 502 — proves the host is alive and clears the record. + * - The window is short (30 s), so a mistake costs one page of navigation, and + * a host that came back is retried on its own. + * - Half-open means exactly ONE request goes out, not a whole screenful. + * - An explicit reset (user retry, endpoint discovery) always wins, and + * failures from requests that started before it are discarded — otherwise a + * 30-second straggler settles right after the reset and re-opens the breaker + * underneath the retry that cleared it. + * + * Deliberately not persisted: process lifetime is the right scope for + * "unreachable right now". + */ + +import { buildHostConnectivityFastFailMessage } from '@iptvnator/shared/interfaces'; + +/** Consecutive connection failures that trip the breaker. */ +const FAILURE_THRESHOLD = 2; +/** + * How long requests fast-fail once the breaker is open. Exported because both + * apps say it out loud in their "skipping requests for Ns" log line, and a + * second copy of the number would drift from this one. + */ +export const OPEN_DURATION_MS = 30_000; +/** + * Failures further apart than this are unrelated, not a streak. Two sequential + * 30 s timeouts must fit inside it, hence comfortably above 60 s. + */ +const FAILURE_WINDOW_MS = 120_000; +/** + * Safety net for a half-open trial that never reports back. Above the longest + * request timeout (30 s) plus margin, so it only fires if a caller leaked the + * token — without it a lost report would keep the breaker open forever. + */ +const TRIAL_TIMEOUT_MS = 45_000; +/** Hosts tracked at once; portals per user are few, this is just a bound. */ +const MAX_TRACKED_HOSTS = 256; +/** Idle records are forgotten; they hold nothing worth remembering. */ +const IDLE_TTL_MS = 600_000; + +const GUARD_DISABLED_ENV = 'IPTVNATOR_DISABLE_CONNECTIVITY_GUARD'; + +/** + * Error codes that prove the host itself did not answer. + * + * `ECONNRESET` is deliberately absent: a reset mid-transfer happens on hosts + * that are very much alive, and it is the one code a working stream can emit. + */ +const HOST_LEVEL_FAILURE_CODES = new Set([ + 'ECONNABORTED', + 'ECONNREFUSED', + 'EAI_AGAIN', + 'EHOSTUNREACH', + 'ENETUNREACH', + 'ENOTFOUND', + 'ETIMEDOUT', +]); + +/** + * Thrown instead of making a request while the breaker is open. + * + * A real `Error`, because Electron serializes a rejected plain object to + * `[object Object]` and the renderer's classification would be lost. It + * carries NO `status` property on purpose — `getStalkerRequestErrorStatus` + * reads that field first, and a numeric status there would read as "the + * endpoint answered". + */ +export class HostConnectivityGuardError extends Error { + /** Scheme, host and port — see {@link portalEndpointKeyOf}. */ + readonly endpoint: string; + + constructor(endpoint: string) { + super(buildHostConnectivityFastFailMessage(endpoint)); + this.name = 'HostConnectivityGuardError'; + this.endpoint = endpoint; + } +} + +/** + * Handed out by {@link HostConnectivityGuard.check} and passed back when the + * request settles. It records which attempt the report belongs to, so a reset + * or a parallel sibling cannot be mistaken for fresh evidence. + */ +export interface HostRequestToken { + /** Scheme, host and port — see {@link portalEndpointKeyOf}. */ + readonly endpoint: string; + readonly epoch: number; + readonly startedAt: number; + /** Whether this request is the single probe allowed while half-open. */ + readonly trial: boolean; + /** + * Which half-open slot this request holds, when it holds one. + * + * A trial can outlive its own window — `requestWithValidatedRedirects` + * gives each of up to five redirect hops its own 30 s budget — and once a + * replacement has been admitted, the abandoned request must not be able to + * free the replacement's slot. `trial: true` alone cannot tell the two + * apart, so the owner is identified explicitly. + */ + readonly trialId: number; +} + +export type HostConnectivityCheck = + | { readonly allowed: true; readonly token: HostRequestToken } + | { readonly allowed: false; readonly retryAfterMs: number }; + +export type HostRequestOutcome = 'responded' | 'host-level' | 'inconclusive'; + +interface HostState { + consecutiveFailures: number; + /** When the last COUNTED failure was recorded. */ + lastFailureAt: number; + openUntil: number; + trialStartedAt: number | null; + /** Monotonic id of the half-open slot; see `HostRequestToken.trialId`. */ + trialId: number; + epoch: number; + lastTouchedAt: number; +} + +export interface HostConnectivityGuardOptions { + now?: () => number; + onOpen?: (endpoint: string) => void; +} + +/** + * Whether the guard is switched off. Read at call time, not at module load, so + * a debugging session can toggle it (matches `IPTVNATOR_ALLOW_INSECURE_TLS`). + */ +export function isHostConnectivityGuardDisabled(): boolean { + const value = process.env[GUARD_DISABLED_ENV]?.trim().toLowerCase(); + return value === '1' || value === 'true'; +} + +/** + * What a failed request proves about the host. + * + * `responded` covers an error that still carries an HTTP response (5xx reaches + * the handlers as a rejection because `validateStatus` only tolerates <500) — + * the host is alive, so it clears the record. `inconclusive` covers everything + * that says nothing about reachability: a cancelled request, a URL rejected by + * the SSRF policy, a bug in our own code. + */ +export function classifyHostRequestFailure(error: unknown): HostRequestOutcome { + if (!error || typeof error !== 'object') { + return 'inconclusive'; + } + + if ((error as { response?: unknown }).response) { + return 'responded'; + } + + const code = (error as { code?: unknown }).code; + if (typeof code === 'string' && HOST_LEVEL_FAILURE_CODES.has(code)) { + return 'host-level'; + } + + return 'inconclusive'; +} + +/** + * The guard key for a request URL: its ORIGIN, i.e. scheme, host and port. + * + * Not `URL.host`, which omits a default port and therefore gives + * `http://panel.example` and `https://panel.example` the same key — two + * genuinely different endpoints, and a panel whose TLS listener is broken while + * plain HTTP works is a routine IPTV setup. Sharing one record there would let + * the dead one fast-fail the working one without ever contacting it. + * + * `URL.origin` also leaves out any `user:pass@` userinfo, so no credential + * reaches the key or the log line. + */ +export function portalEndpointKeyOf(url: string): string | null { + try { + const origin = new URL(url).origin; + // Opaque origins serialize as "null" and are not a usable key. + return origin && origin !== 'null' ? origin : null; + } catch { + return null; + } +} + +export class HostConnectivityGuard { + private readonly states = new Map(); + private readonly now: () => number; + private readonly onOpen?: (endpoint: string) => void; + + constructor(options: HostConnectivityGuardOptions = {}) { + this.now = options.now ?? (() => Date.now()); + this.onOpen = options.onOpen; + } + + /** + * Whether a request to `endpoint` may go out, and the token to report with. + * + * Fast-fails while the breaker is open, and while a half-open trial is + * still in flight — the point of the trial is that a screenful of requests + * does not all hang again to learn the same thing. + */ + check(endpoint: string): HostConnectivityCheck { + const now = this.now(); + if (isHostConnectivityGuardDisabled()) { + return { + allowed: true, + token: { + endpoint, + epoch: 0, + startedAt: now, + trial: false, + trialId: 0, + }, + }; + } + + const state = this.ensureState(endpoint, now); + state.lastTouchedAt = now; + + if (state.openUntil > now) { + return { allowed: false, retryAfterMs: state.openUntil - now }; + } + + // Never opened, or the open window elapsed. The latter is half-open: + // let one request through and hold the rest back until it settles. + let trial = false; + if (state.openUntil > 0) { + const trialStale = + state.trialStartedAt !== null && + now - state.trialStartedAt >= TRIAL_TIMEOUT_MS; + if (state.trialStartedAt !== null && !trialStale) { + return { allowed: false, retryAfterMs: 0 }; + } + state.trialStartedAt = now; + state.trialId += 1; + trial = true; + } + + return { + allowed: true, + token: { + endpoint, + epoch: state.epoch, + startedAt: now, + trial, + trialId: trial ? state.trialId : 0, + }, + }; + } + + /** + * The host answered. Always clears the record, whatever the status was and + * whichever attempt it belonged to — a reachable host is a reachable host. + */ + reportSuccess(token: HostRequestToken): void { + const state = this.states.get(token.endpoint); + if (!state || isHostConnectivityGuardDisabled()) { + return; + } + + state.consecutiveFailures = 0; + state.lastFailureAt = 0; + state.openUntil = 0; + state.trialStartedAt = null; + state.lastTouchedAt = this.now(); + } + + /** The host did not answer at all. */ + reportFailure(token: HostRequestToken): void { + const state = this.states.get(token.endpoint); + if (!state || isHostConnectivityGuardDisabled()) { + return; + } + + const now = this.now(); + // Superseded by an explicit reset: the user asked for a fresh attempt + // and this verdict predates it. + if (state.epoch !== token.epoch) { + this.releaseTrial(state, token); + return; + } + + // Only the request that still holds the slot ends the half-open state. + // An abandoned trial's late failure is ordinary evidence — it goes + // through the streak rules below and leaves the replacement alone. + const wasTrial = this.ownsTrial(state, token); + if (wasTrial) { + state.trialStartedAt = null; + } + state.lastTouchedAt = now; + + if ( + state.consecutiveFailures > 0 && + now - state.lastFailureAt > FAILURE_WINDOW_MS + ) { + state.consecutiveFailures = 0; + } + + // A request already in flight when the previous failure was recorded is + // not the next link in a streak — it is a sibling of it. Catalog + // loading fans out several requests at once, and one hiccup failing all + // of them is one piece of evidence, not a trip. A request that started + // at or after that moment is a genuine new attempt. + let counted = false; + if ( + state.consecutiveFailures === 0 || + token.startedAt >= state.lastFailureAt + ) { + state.consecutiveFailures += 1; + state.lastFailureAt = now; + counted = true; + } + + // Only a failure this report actually counted may trip the threshold. + // A sibling arriving after the window elapsed would otherwise open a + // fresh one off the existing count, pushing the half-open trial past + // the intended cooldown. A failed trial still goes straight back to + // open: the host had its chance. + if ( + wasTrial || + (counted && state.consecutiveFailures >= FAILURE_THRESHOLD) + ) { + const wasOpen = state.openUntil > now; + state.openUntil = now + OPEN_DURATION_MS; + if (!wasOpen) { + this.onOpen?.(token.endpoint); + } + } + } + + /** + * The request failed for a reason that says nothing about the host. Only + * releases the half-open slot, so the next request can be the trial. + */ + reportInconclusive(token: HostRequestToken): void { + const state = this.states.get(token.endpoint); + if (!state || isHostConnectivityGuardDisabled()) { + return; + } + + this.releaseTrial(state, token); + state.lastTouchedAt = this.now(); + } + + /** + * Forgets everything recorded for `endpoint`, and invalidates the reports of + * requests already in flight. Called when the user asks for a real attempt + * or when the URL may now point somewhere else. + */ + reset(endpoint: string): void { + const now = this.now(); + const state = this.ensureState(endpoint, now); + state.consecutiveFailures = 0; + state.lastFailureAt = 0; + state.openUntil = 0; + state.trialStartedAt = null; + state.epoch += 1; + state.lastTouchedAt = now; + } + + /** + * A token for a request that must NOT be policed or counted, but whose + * success still clears the record. + * + * Stalker endpoint discovery probes several candidate paths on one host and + * expects most of them to fail; counting those would let it declare a + * slow-but-alive portal unreachable. It does not take the half-open slot + * either — a probe is not the trial the guard is waiting for. + */ + observe(endpoint: string): HostRequestToken { + return { + endpoint, + epoch: this.states.get(endpoint)?.epoch ?? 0, + startedAt: this.now(), + trial: false, + trialId: 0, + }; + } + + /** Test seam: drops all recorded state. */ + clear(): void { + this.states.clear(); + } + + /** Whether `token` still holds the current half-open slot. */ + private ownsTrial(state: HostState, token: HostRequestToken): boolean { + return ( + token.trial && + state.trialStartedAt !== null && + state.epoch === token.epoch && + state.trialId === token.trialId + ); + } + + private releaseTrial(state: HostState, token: HostRequestToken): void { + if (this.ownsTrial(state, token)) { + state.trialStartedAt = null; + } + } + + private ensureState(endpoint: string, now: number): HostState { + const existing = this.states.get(endpoint); + if (existing) { + return existing; + } + + this.prune(now); + const state: HostState = { + consecutiveFailures: 0, + lastFailureAt: 0, + openUntil: 0, + trialStartedAt: null, + trialId: 0, + epoch: 0, + lastTouchedAt: now, + }; + this.states.set(endpoint, state); + return state; + } + + private prune(now: number): void { + for (const [endpoint, state] of this.states) { + if ( + now - state.lastTouchedAt > IDLE_TTL_MS && + state.openUntil <= now && + state.trialStartedAt === null + ) { + this.states.delete(endpoint); + } + } + + // Still at the cap: forget the oldest records. Dropping one means + // contacting that host again, which is the safe direction to err in. + while (this.states.size >= MAX_TRACKED_HOSTS) { + const oldest = this.states.keys().next(); + if (oldest.done) { + return; + } + this.states.delete(oldest.value); + } + } +} + +/** + * The URL a failed request was actually talking to, when the error says. + * + * Redirects are followed hop by hop, each with its own config, so a failure on + * a later hop carries THAT hop's URL rather than the one we asked for. + */ +function failedRequestUrlOf(error: unknown): string | null { + // Two transports, two places to look, and only one of them is ever right. + // + // The Electron transport follows redirects itself with `maxRedirects: 0`, + // reissuing each hop as its own request, so the hop that failed is the + // error's `config.url`. + // + // The web backend uses axios' default transport, where follow-redirects + // walks the chain inside a single request. `config` is built once and never + // rewritten, so `config.url` stays the URL we asked for — comparing it + // against itself would find no redirect and charge a dead destination to + // the provider that answered. follow-redirects tracks the hop it is on as + // `request._currentUrl`, so that is read first; a transport that does not + // expose it falls through to `config.url`. + const currentUrl = ( + error as { request?: { _currentUrl?: unknown } | null } | null + )?.request?._currentUrl; + if (typeof currentUrl === 'string') { + return currentUrl; + } + + const url = (error as { config?: { url?: unknown } } | null)?.config?.url; + return typeof url === 'string' ? url : null; +} + +/** + * Origin and path only — deliberately without the query string. + * + * The baseline a caller can supply is the URL it handed the transport, but the + * transport is free to add to the query before sending: the web backend passes + * Xtream credentials through axios' `params`, so a request for + * `…/player_api.php` goes out as `…/player_api.php?username=…&action=…`. + * Comparing whole URLs then reports a redirect for every ordinary failure, + * which credits the endpoint instead of counting it and stops the breaker from + * ever opening. + * + * Dropping the query keeps what the comparison is actually for — an endpoint + * that answered and sent us somewhere else, including the same-origin + * `/player_api.php` → `/slow/player_api.php` case — and gives up only a + * redirect that changes nothing but the query. That one is then counted as an + * ordinary failure, which is the safe direction to be wrong in. + */ +function normalizedUrlOrNull(url: string): string | null { + try { + const parsed = new URL(url); + return `${parsed.origin}${parsed.pathname}`; + } catch { + return null; + } +} + +/** + * Whether the failure happened on a hop the guarded endpoint redirected us to. + * + * Reaching any later hop proves the guarded endpoint answered: the first hop is + * always the URL we asked for, and only a redirect status advances the chain. + * That holds for a same-origin redirect too, so comparing the whole URL — not + * just its origin — is what catches `/player_api.php` → `/slow/player_api.php`. + * + * Requires positive evidence: anything unparseable or unknown returns false and + * the failure is counted as usual, because guessing "redirect" here would stop + * the guard from ever tripping. + */ +export function failedAfterRedirect( + error: unknown, + token: HostRequestToken, + requestUrl: string | undefined +): boolean { + const failedUrl = failedRequestUrlOf(error); + if (!failedUrl) { + return false; + } + + const failedEndpoint = portalEndpointKeyOf(failedUrl); + if (failedEndpoint && failedEndpoint !== token.endpoint) { + return true; + } + + if (!requestUrl) { + return false; + } + + const failedNormalized = normalizedUrlOrNull(failedUrl); + const requestedNormalized = normalizedUrlOrNull(requestUrl); + return ( + failedNormalized !== null && + requestedNormalized !== null && + failedNormalized !== requestedNormalized + ); +} diff --git a/libs/shared/host-health/tsconfig.json b/libs/shared/host-health/tsconfig.json new file mode 100644 index 000000000..ab8e5af25 --- /dev/null +++ b/libs/shared/host-health/tsconfig.json @@ -0,0 +1,19 @@ +{ + "extends": "../../../tsconfig.base.json", + "compilerOptions": { + "module": "commonjs", + "forceConsistentCasingInFileNames": true, + "strict": true, + "importHelpers": true, + "noImplicitOverride": true, + "noImplicitReturns": true, + "noFallthroughCasesInSwitch": true, + "noPropertyAccessFromIndexSignature": true + }, + "files": [], + "include": [], + "references": [ + { "path": "./tsconfig.lib.json" }, + { "path": "./tsconfig.spec.json" } + ] +} diff --git a/libs/shared/host-health/tsconfig.lib.json b/libs/shared/host-health/tsconfig.lib.json new file mode 100644 index 000000000..163d90724 --- /dev/null +++ b/libs/shared/host-health/tsconfig.lib.json @@ -0,0 +1,10 @@ +{ + "extends": "./tsconfig.json", + "compilerOptions": { + "outDir": "../../../dist/out-tsc", + "declaration": true, + "types": ["node"] + }, + "include": ["src/**/*.ts"], + "exclude": ["jest.config.ts", "src/**/*.spec.ts", "src/**/*.test.ts"] +} diff --git a/libs/shared/host-health/tsconfig.spec.json b/libs/shared/host-health/tsconfig.spec.json new file mode 100644 index 000000000..4b0383fc4 --- /dev/null +++ b/libs/shared/host-health/tsconfig.spec.json @@ -0,0 +1,15 @@ +{ + "extends": "./tsconfig.json", + "compilerOptions": { + "outDir": "../../../dist/out-tsc", + "module": "commonjs", + "moduleResolution": "node10", + "types": ["jest", "node"] + }, + "include": [ + "jest.config.ts", + "src/**/*.test.ts", + "src/**/*.spec.ts", + "src/**/*.d.ts" + ] +} diff --git a/tools/coverage/coverage-policy.json b/tools/coverage/coverage-policy.json index e9ee12961..f6303ffa1 100644 --- a/tools/coverage/coverage-policy.json +++ b/tools/coverage/coverage-policy.json @@ -208,6 +208,13 @@ "validationCommand": "pnpm nx test shared-logging", "e2eTags": ["@electron", "@xtream", "@stalker", "@settings"] }, + { + "name": "shared-host-health", + "root": "libs/shared/host-health", + "sourceRoot": "libs/shared/host-health/src", + "validationCommand": "pnpm nx test shared-host-health", + "e2eTags": ["@xtream", "@stalker"] + }, { "name": "m3u-utils", "root": "libs/shared/m3u-utils", diff --git a/tsconfig.base.json b/tsconfig.base.json index c934e2bb9..bc38b7d67 100644 --- a/tsconfig.base.json +++ b/tsconfig.base.json @@ -83,6 +83,9 @@ ], "@iptvnator/shared/testing": ["libs/shared/testing/src/index.ts"], "@iptvnator/shared/logging": ["libs/shared/logging/src/index.ts"], + "@iptvnator/shared/host-health": [ + "libs/shared/host-health/src/index.ts" + ], "@iptvnator/services": ["libs/services/src/index.ts"], "@iptvnator/shared/interfaces": [ "libs/shared/interfaces/src/index.ts"