Files
iptvnator/docs/architecture/performance-journeys.md
T
0f1cad4b16 test(performance): add the J1 launch-to-usable journey benchmark (#1698)
* test(performance): add the J1 launch-to-usable journey benchmark

Implements plan items A1, A3 (J1 only) and the minimal A4 from
.plans/2026-09-25-performance-journeys-ratchet.md.

- journey-renderer-probe.ts: init-script probe counting DOM mutations,
  layout shifts and long tasks until the first source card is visible on
  /workspace with the splash removed; unit-tested with jsdom fixtures.
- journey-main-ipc-capture.ts: counts bridge invocations from the preload's
  renderer-API trace channel up to a sentinel call the probe fires, so the
  IPC counter is exact without touching production code.
- launch.journey.ts + playwright.journeys.config.ts: seeded profile (one
  M3U source, one Xtream portal on the loopback mock), one warm-up and five
  measured iterations, fresh process and data directory each, writing
  dist/performance/journeys/<timestamp>/summary.json with exact counters
  and P50/P90 wall-clock.
- Nx target electron-backend-e2e:journeys and root script perf:journeys.
- docs/architecture/performance-journeys.md, README, context and
  validation map entries.

renderer.cdTicksToFirstCard and main.sqlStatementsBeforeReadyToShow are
reported as unavailable with the reason instead of being faked.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(performance): gate the renderer load so the probes never race startup

Review follow-up for #1698.

- journey-renderer-gate.cjs: a main-process hook loaded with `-r` (the
  mechanism Playwright uses for its own loader) makes the first
  loadFile/loadURL navigate to about:blank and holds the real load until
  the test releases it. Playwright reports no page before a navigation
  commits, so this is what lets the renderer probe be registered on the
  page before the real document exists; the IPC capture is installed
  before the release too. A safety timeout releases the gate on its own
  and marks the iteration invalid. Unit-tested with a fake BrowserWindow.
- launch-journey-app.ts: registers the probe on the parked page, releases
  the gate, waits for the real document to commit, fails fast when the
  probe is missing, and refuses an iteration whose gate timed out, saw a
  second load, or released before the probe was in place.
- journey-renderer-probe.ts: entries delivered live after the terminal
  batch are buffered and filtered by the same cutoff as queued ones, and
  the cutoff is sampled in a timer queued from the first rAF, i.e. after
  the card's frame is painted, so the render task's long task and layout
  shift are consistently included.
- launch-journey-record.ts: layoutShiftScore rounded to three decimals; a
  0.0001 shift flipped in and out of the cutoff between iterations.
- playwright.journeys.config.ts: reuse a mock server left on the journeys
  port locally (its fixtures are deterministic); CI still starts its own.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(performance): validate the journey gate after the probe completes

Codex follow-up on #1698: the gate state returned by release() cannot see a
reload or recovery navigation that happens before the first card. Re-read the
live state once the renderer probe has finished and validate that instead,
so an iteration spanning an extra navigation is rejected.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(performance): fail an iteration whose performance observers were unavailable

Codex follow-up on #1698: a renderer that cannot observe layout-shift or
longtask entries used to pass the probe with zero counters, which a ratchet
could not tell apart from a genuine zero. The probe assertion now rejects
such an iteration and names the missing observer.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: 4gray <fourgray@proton.me>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-26 19:18:19 +02:00

320 lines
19 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Performance journeys and the CI ratchet
IPTVnator measures performance through a small set of everyday user journeys.
Each journey has deterministic counters that are asserted exactly, and
wall-clock timings that are recorded as evidence. Counters are ratcheted in CI:
a committed baseline may only be lowered, and only with the measured output as
evidence. This document is the contract for that loop. The journey harness lives in
`apps/electron-backend-e2e/src/journeys` and
`apps/electron-backend-e2e/src/performance/journey-*.ts`; the ratchet scripts
live in `tools/performance/`.
## Journeys
| Journey | Start | End |
| ---------------- | -------------------------------------------- | ----------------------------------------------------------------------------- |
| J1 `launch` | Electron process spawn | first playlist or portal card rendered on `/workspace`, inline splash removed |
| J2 `open-source` | click on a portal card | live category list and first channel page painted |
| J3 `playback` | click on a channel | HTML5 `playing` event |
| J4 `search` | six-character query typed into global search | results list settled |
J1 is instrumented today: `renderer.initialBytes` from the built output, and
the runtime counters of the launch benchmark below. J2 to J4 follow the plan
in `.plans/` and are added one thread at a time; each thread names its journey
and counter in the PR description.
## Running the journeys
```bash
pnpm run perf:journeys
```
The script runs the Nx target `electron-backend-e2e:journeys`, which builds the
`electron-performance` configuration of the Electron app and the renderer
first, starts the Xtream mock server on the dedicated loopback port
`127.0.0.1:3231` (override with `IPTVNATOR_JOURNEY_XTREAM_MOCK_PORT`), and runs
`playwright.journeys.config.ts` with one worker. Each run writes one file:
```
dist/performance/journeys/<YYYYMMDDTHHMMSSZ>/summary.json
```
The file is never overwritten; a second run in the same second fails instead.
`IPTVNATOR_JOURNEY_MEASURED_ITERATIONS` lowers the five measured iterations
for a quick local check; the warm-up iteration always runs. Numbers from a
laptop are previews: the Linux CI runner is the canonical measurer for
baselines, as it is for `renderer.initialBytes`.
## J1 `launch`: launch to usable
The profile holds one M3U source and one Xtream portal, both served by the
Xtream mock (`/playlist.m3u` and `player_api.php` on the same origin). The
profile is seeded once per run through the app's own "Add playlist" dialogs,
then every iteration copies that seeded data directory into a fresh temporary
directory and spawns a fresh Electron process on it. One warm-up iteration is
recorded but excluded from the summary; five measured iterations follow. The
app lands on `/workspace/dashboard`, so the first card is a card of the
"Recent sources" rail; an `app-playlist-item` row on `/workspace/sources`
also ends the journey for profiles that disable the dashboard.
The journey ends at the first `MutationObserver` batch in which all of the
following hold: the location is below `/workspace`, `#initial-splash` is no
longer in the DOM, and a source card has a non-empty client rect. Counters are
frozen at that microtask checkpoint, so bridge calls and mutations issued
later in the same task are included and everything after it is not.
Three test-side pieces are injected; production code is not changed:
- `journey-renderer-gate.cjs` is loaded into the main process with `-r`, the
mechanism Playwright uses for its own loader. Playwright resolves
`electron.launch()` while the app is already creating its window, and
Electron reports no page until a navigation commits, so an init script
registered afterwards would race the first document. The gate makes the
first `loadFile` navigate to `about:blank` and holds the real load until
the test releases it. A 15 s safety timeout releases it on its own and the
iteration is then invalid.
- `journey-renderer-probe.ts` is registered with `addInitScript` on that
`about:blank` page, so it runs at the start of the real document. It
records that it ran while the document was still `loading` with zero
scripts and emits one JSON blob under `window.__iptvnatorJourneyProbe`.
- `journey-main-ipc-capture.ts` subscribes to the preload's renderer-API trace
channel (`IPTVNATOR_DEBUG_TRACE_EVENT`, enabled with
`IPTVNATOR_TRACE_IPC=1`) through `electronApp.evaluate`, also before the
release. The record refuses an iteration whose gate timed out, saw a second
load, or released before the probe was in place.
### Counters
| Counter | Source |
| ---------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `renderer.ipcCallsToFirstCard` | `start` trace events the preload emits for every bridge invocation (listener registrations `on*`/`remove*` excluded, as in `wrapElectronApi`). The renderer probe fires one sentinel `dbGetAppPlaylist('__iptvnator-journey-sentinel__')` at the terminal moment; renderer-to-main IPC is ordered, so events before the sentinel are the exact count. |
| `renderer.domMutationsToFirstCard` | `MutationRecord`s (not callback batches) from a `MutationObserver` on the document element with `childList`, `attributes`, `characterData` and `subtree`. When the init script runs before `<html>` exists the observer watches `document`, which the blob reports in `capabilities.observedTarget`. |
| `renderer.layoutShiftScore` | Sum of `layout-shift` entries with `hadRecentInput === false`, rounded to three decimals (a shift of 0.0001 flips in and out of the cutoff between runs; the CLS "good" threshold is 0.1, so three decimals keep the counter exact without hiding anything a user could see). The cutoff is sampled in a timer queued from the first `requestAnimationFrame` after the terminal batch, that is after the frame that paints the card has been committed; entries delivered live after the terminal batch are buffered and filtered by the same cutoff. |
| `renderer.longTasks` | `longtask` entries over 50 ms up to that same cutoff, which includes the task that rendered the card. The count depends on machine speed, so it is evidence until a run shows it is stable on the CI runner. |
Counters are exact: the summary carries the value shared by every measured
iteration. When iterations disagree, the summary reports the maximum and marks
the counter `stable: false` under `counterStability`; such a counter is not
promoted to a guardrail until it is deterministic.
Two counters from the plan are listed under `unavailable` with the reason
instead of being faked:
- `renderer.cdTicksToFirstCard`: the `electron-performance` build optimizes
scripts, which sets `ngDevMode` to false, so Angular does not publish
`window.ng` and `ɵsetProfiler` is unavailable. The probe checks this at the
terminal moment and the record refuses a build where the hook exists but was
not counted.
- `main.sqlStatementsBeforeReadyToShow`: SQL statements are only visible as
worker-thread trace lines on stdout, which Node forwards asynchronously, so
they cannot be ordered against `ready-to-show`. Plan item A2 adds a channel
that can be counted.
### Wall-clock
| Entry | Derivation |
| --------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `spawnToDidFinishLoadMs.p50/.p90` | `performance.timeOrigin + loadEventEnd` of the navigation entry (the main frame's `load`, which is what `did-finish-load` reports) minus the test-side timestamp taken just before `electron.launch`. |
| `spawnToFirstCardMs.p50/.p90` | Terminal epoch of the renderer probe minus the same spawn timestamp. |
Percentiles use linear interpolation over the five measured iterations. The
spawn timestamp includes Playwright's own launch overhead and the gate's
`about:blank` detour: Playwright holds `app.whenReady()` until its CDP session
is attached, and the real document loads only after the probes are in place,
so absolute values are larger than a bare launch. They are comparable between runs of the same
harness, which is what the ratchet needs. The main process start
(`Date.now() - process.uptime()`) is recorded per iteration under
`evidence.epochs` for cross-checks.
### Summary schema
```json
{
"schemaVersion": 1,
"generatedAt": "2026-09-26T11:02:14.318Z",
"harness": {
"platform": "darwin",
"electron": "43.3.0",
"measuredIterations": 5,
"warmupIterations": 1
},
"journeys": {
"launch": {
"counters": { "renderer.ipcCallsToFirstCard": 12 },
"counterStability": {
"renderer.ipcCallsToFirstCard": {
"stable": true,
"values": [12, 12, 12, 12, 12]
}
},
"wallClock": {
"spawnToFirstCardMs.p50": 1234.5,
"spawnToFirstCardMs.p90": 1300.1
},
"unavailable": { "renderer.cdTicksToFirstCard": "reason" },
"iterations": [
{
"index": 0,
"warmup": true,
"pid": 1,
"counters": {},
"wallClock": {},
"evidence": {}
}
]
}
}
}
```
`journeys.<id>.counters.<name>` and `journeys.<id>.wallClock.<name>` are plain
numbers so `tools/performance/check-journey-ratchet.mjs` can compare them with
`tools/performance/journey-baselines.json`. A J1 baseline is added once the
numbers are stable on the CI runner; until then the summary is evidence only.
## `renderer.initialBytes`
The bytes a browser fetches before Angular can bootstrap, read from the built
`dist/apps/web/index.html`:
- `index.html` itself,
- every same-origin `<script src>`, including `assets/app-config.js`,
- every `<link rel="stylesheet">`,
- every `<link rel="modulepreload">` chunk.
Manifest, icons, external URLs, commented-out tags and lazy chunks are not
counted, so the value is the same for every language: a non-English launch
additionally fetches that language's Angular locale chunk (about 2 KB), which
belongs to the per-profile J1 benchmark rather than to this counter. A file
that `index.html` references but the build did not emit is an error, never
zero bytes. The value is raw (uncompressed) size, which is what the renderer
parses. It is Angular's "Initial total" plus `index.html` and
`assets/app-config.js` (about 4 KB together), so it sits slightly above the
rounded figure the build prints; never copy that figure into a baseline, use
the script's output. The bundle embeds only the app version from
`package.json` (a named import, which esbuild tree-shakes), not the whole
file, so editing scripts or dependencies does not move the counter.
```bash
pnpm nx build web # production configuration
pnpm run perf:initial-bytes # human-readable breakdown
pnpm --silent run perf:initial-bytes -- --json # machine-readable; --silent keeps pnpm's headers out of stdout
node tools/performance/measure-initial-bytes.mjs --summary dist/performance/journey-summary.json
```
`--summary` writes the journey summary shape (`journeys.<journey>.counters`)
that the ratchet checker consumes. `--dist <dir>` points the script at another
build output, for example the `electron-performance` configuration.
The measurement script is `tools/performance/measure-initial-bytes.mjs`; its
Node tests run with `pnpm nx test performance-tools` (Tier B in the coverage
policy) and lint with `pnpm nx lint performance-tools`.
## Ratchet
`tools/performance/journey-baselines.json` holds one entry per journey and
counter:
```json
{
"journeys": {
"launch": {
"renderer.initialBytes": {
"value": 2739510,
"unit": "bytes",
"updatedAt": "2026-09-26",
"evidencePr": 1693,
"measuredWith": "pnpm nx build web && pnpm run perf:initial-bytes"
}
}
}
}
```
`tools/performance/check-journey-ratchet.mjs` compares a journey summary with
that file:
- a counter above its `value` fails; counters are exact, there is no slack;
- a wall-clock entry carries `toleranceRatio` and fails above
`value × toleranceRatio`;
- a baseline with no measurement in the summary fails, so dropping a
measurement cannot disable the ratchet; a counter is read only from
`journeys.<journey>.counters` and a wall-clock entry (one with
`toleranceRatio`) only from `journeys.<journey>.wallClock`, so a value in
the wrong section also counts as missing;
- a measurement below its baseline passes and prints a "tighten" hint;
- a measured counter without a baseline is noted, not failed;
- checking nothing fails: an empty baselines file, or `--only` naming an
entry that does not exist, cannot exit 0.
`--only <journey>/<counter>` (repeatable) restricts the check to the named
baselines. A script that measures one counter writes its own summary file
and checks only its counter, so it neither overwrites another measurement's
summary nor fails the other baselines as unmeasured.
```bash
pnpm run perf:initial-bytes:check # measure dist/apps/web into dist/performance/initial-bytes.summary.json, check only that counter
pnpm run perf:ratchet:check # check every baseline against dist/performance/journey-summary.json
```
CI runs `perf:initial-bytes:check` in the `Initial bytes ratchet` job of
`.github/workflows/ci.yml` after a production build of `apps/web`, and uploads
`dist/performance/` as the `performance-journey-summary` artifact. Like the
rest of that workflow it runs for pull requests that target `master` and for
pushes to `master`; a stacked PR that targets another branch gets no run until
it is retargeted, so dispatch one with `gh workflow run ci.yml --ref <branch>`
when you need the number. A PR that grows the counter fails that job.
That runner is the canonical measurer: take baseline values from its output,
not from a local build. A local macOS build of the code before #1695 is 2
bytes smaller in `main.js` (the eager locale imports); since #1695 the two
have been byte-identical. (An apparent 556-byte platform difference during
the first measurements was otherwise `package.json` text embedded in
`main.js`, which moved with every script edit; #1692 fixed that by importing
only the version.)
The job also refuses a weakened baselines file:
`tools/performance/check-baseline-direction.mjs` compares
`journey-baselines.json` with the revision the change is measured against
(the target branch of a pull request, the previous head of a `master` push,
`master` for a manual dispatch) and fails when any
entry's enforced limit (`value × toleranceRatio`) went up, a tolerance widened
or an entry disappeared, so a PR cannot grow the payload and raise the
baseline to match. Lowered limits and new entries pass.
Baselines only move down. Lower `value` in the same PR as the change that
earned it, set `updatedAt` and `evidencePr`, and paste the measurement output
into the PR. Never raise a value to make a PR pass: if growth is a deliberate
trade-off, say so in the PR and let the maintainer decide.
## Adding a counter
1. Produce the value from the built output or from a deterministic probe, not
from source heuristics. Missing inputs must fail the measurement.
2. Emit it under `journeys.<journey>.counters.<name>` in the summary JSON.
3. Cover the extraction and the failure modes with `node --test` and register
the test file in `tools/performance/project.json`.
4. Validate the counter before it becomes a guardrail: one PR must show that
lowering it moved wall-clock in the same journey.
## Adding a journey
1. Add `apps/electron-backend-e2e/src/journeys/<journey>.journey.ts`. Seed the
profile through the app's dialogs, spawn a fresh process per iteration
with `measureLaunchJourney` as the model, and drive the journey's start
action with Playwright.
2. Give the journey its own probe options (`cardSelector`, `routeFragment`,
terminal condition) or extend `journey-renderer-probe.ts` when the end
condition is not "an element became visible". Keep the probe
self-contained: Playwright serializes it with `toString()`.
3. Map the measurement to a `JourneyIterationRecord` in a
`<journey>-journey-record.ts` under `src/performance/`; name counters
`renderer.*` or `main.*`, and list counters you cannot measure under
`unavailable` with the reason.
4. Add the journey under `journeys.<id>` in the summary through
`summarizeJourneyIterations`; the schema needs no change.
5. Cover the probe with jsdom fixtures and the record and summary code with
`node:test` (`pnpm nx run electron-backend-e2e:test-performance-harness`).
6. Validate a counter before it becomes a guardrail: one PR must show that
lowering it moved wall-clock in the same journey.