Files
iptvnator/docs/architecture/performance-journeys.md
T
69056487bc ci(performance): run the performance journeys on the Linux runner (#1717)
* ci(performance): run the performance journeys on the Linux runner

Adds a warn-only performance-journeys job to ci.yml that builds the
electron-performance configuration, runs the journey benchmarks under
xvfb and uploads dist/performance/journeys/ as evidence. Pull requests
run it only when they touch journey-relevant paths.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(performance): document the journeys CI job and runner evidence

Describes the performance-journeys job and its path gate, and records why
the J1 runtime counters are not baselined yet: on the Linux runner the
launch journey is bimodal (13/576 vs 16/939 bridge calls/DOM mutations),
so the counters are not deterministic.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* ci(performance): run the journeys unless a PR changes only safe paths

The scope filter listed the paths that can move a journey, so a PR that
changed only a root build input (.nvmrc, nx.json, tsconfig.base.json)
skipped the measurement. List the paths that cannot instead: the E2E
workflow's ignore list plus release notes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: 4gray <fourgray@proton.me>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-27 20:53:26 +02:00

22 KiB
Raw Blame History

Performance journeys and the CI ratchet

IPTVnator measures performance through a small set of everyday user journeys. Each journey has deterministic counters that are asserted exactly, and wall-clock timings that are recorded as evidence. Counters are ratcheted in CI: a committed baseline may only be lowered, and only with the measured output as evidence. This document is the contract for that loop. The journey harness lives in apps/electron-backend-e2e/src/journeys and apps/electron-backend-e2e/src/performance/journey-*.ts; the ratchet scripts live in tools/performance/.

Journeys

Journey Start End
J1 launch Electron process spawn first playlist or portal card rendered on /workspace, inline splash removed
J2 open-source click on a portal card live category list and first channel page painted
J3 playback click on a channel HTML5 playing event
J4 search six-character query typed into global search results list settled

J1 is instrumented today: renderer.initialBytes from the built output, and the runtime counters of the launch benchmark below. J2 to J4 follow the plan in .plans/ and are added one thread at a time; each thread names its journey and counter in the PR description.

Running the journeys

pnpm run perf:journeys

The script runs the Nx target electron-backend-e2e:journeys, which builds the electron-performance configuration of the Electron app and the renderer first, starts the Xtream mock server on the dedicated loopback port 127.0.0.1:3231 (override with IPTVNATOR_JOURNEY_XTREAM_MOCK_PORT), and runs playwright.journeys.config.ts with one worker. Each run writes one file:

dist/performance/journeys/<YYYYMMDDTHHMMSSZ>/summary.json

The file is never overwritten; a second run in the same second fails instead. IPTVNATOR_JOURNEY_MEASURED_ITERATIONS lowers the five measured iterations for a quick local check; the warm-up iteration always runs. Numbers from a laptop are previews: the Linux CI runner is the canonical measurer for baselines, as it is for renderer.initialBytes.

J1 launch: launch to usable

The profile holds one M3U source and one Xtream portal, both served by the Xtream mock (/playlist.m3u and player_api.php on the same origin). The profile is seeded once per run through the app's own "Add playlist" dialogs, then every iteration copies that seeded data directory into a fresh temporary directory and spawns a fresh Electron process on it. One warm-up iteration is recorded but excluded from the summary; five measured iterations follow. The app lands on /workspace/dashboard, so the first card is a card of the "Recent sources" rail; an app-playlist-item row on /workspace/sources also ends the journey for profiles that disable the dashboard.

The journey ends at the first MutationObserver batch in which all of the following hold: the location is below /workspace, #initial-splash is no longer in the DOM, and a source card has a non-empty client rect. Counters are frozen at that microtask checkpoint, so bridge calls and mutations issued later in the same task are included and everything after it is not.

Three test-side pieces are injected; production code is not changed:

  • journey-renderer-gate.cjs is loaded into the main process with -r, the mechanism Playwright uses for its own loader. Playwright resolves electron.launch() while the app is already creating its window, and Electron reports no page until a navigation commits, so an init script registered afterwards would race the first document. The gate makes the first loadFile navigate to about:blank and holds the real load until the test releases it. A 15 s safety timeout releases it on its own and the iteration is then invalid.
  • journey-renderer-probe.ts is registered with addInitScript on that about:blank page, so it runs at the start of the real document. It records that it ran while the document was still loading with zero scripts and emits one JSON blob under window.__iptvnatorJourneyProbe.
  • journey-main-ipc-capture.ts subscribes to the preload's renderer-API trace channel (IPTVNATOR_DEBUG_TRACE_EVENT, enabled with IPTVNATOR_TRACE_IPC=1) through electronApp.evaluate, also before the release. The record refuses an iteration whose gate timed out, saw a second load, or released before the probe was in place.

Counters

Counter Source
renderer.ipcCallsToFirstCard start trace events the preload emits for every bridge invocation (listener registrations on*/remove* excluded, as in wrapElectronApi). The renderer probe fires one sentinel dbGetAppPlaylist('__iptvnator-journey-sentinel__') at the terminal moment; renderer-to-main IPC is ordered, so events before the sentinel are the exact count.
renderer.domMutationsToFirstCard MutationRecords (not callback batches) from a MutationObserver on the document element with childList, attributes, characterData and subtree. When the init script runs before <html> exists the observer watches document, which the blob reports in capabilities.observedTarget.
renderer.layoutShiftScore Sum of layout-shift entries with hadRecentInput === false, rounded to three decimals (a shift of 0.0001 flips in and out of the cutoff between runs; the CLS "good" threshold is 0.1, so three decimals keep the counter exact without hiding anything a user could see). The cutoff is sampled in a timer queued from the first requestAnimationFrame after the terminal batch, that is after the frame that paints the card has been committed; entries delivered live after the terminal batch are buffered and filtered by the same cutoff.
renderer.longTasks longtask entries over 50 ms up to that same cutoff, which includes the task that rendered the card. The count depends on machine speed, so it is evidence until a run shows it is stable on the CI runner.

Counters are exact: the summary carries the value shared by every measured iteration. When iterations disagree, the summary reports the maximum and marks the counter stable: false under counterStability; such a counter is not promoted to a guardrail until it is deterministic.

Two counters from the plan are listed under unavailable with the reason instead of being faked:

  • renderer.cdTicksToFirstCard: the electron-performance build optimizes scripts, which sets ngDevMode to false, so Angular does not publish window.ng and ɵsetProfiler is unavailable. The probe checks this at the terminal moment and the record refuses a build where the hook exists but was not counted.
  • main.sqlStatementsBeforeReadyToShow: SQL statements are only visible as worker-thread trace lines on stdout, which Node forwards asynchronously, so they cannot be ordered against ready-to-show. Plan item A2 adds a channel that can be counted.

Wall-clock

Entry Derivation
spawnToDidFinishLoadMs.p50/.p90 performance.timeOrigin + loadEventEnd of the navigation entry (the main frame's load, which is what did-finish-load reports) minus the test-side timestamp taken just before electron.launch.
spawnToFirstCardMs.p50/.p90 Terminal epoch of the renderer probe minus the same spawn timestamp.

Percentiles use linear interpolation over the five measured iterations. The spawn timestamp includes Playwright's own launch overhead and the gate's about:blank detour: Playwright holds app.whenReady() until its CDP session is attached, and the real document loads only after the probes are in place, so absolute values are larger than a bare launch. They are comparable between runs of the same harness, which is what the ratchet needs. The main process start (Date.now() - process.uptime()) is recorded per iteration under evidence.epochs for cross-checks.

Summary schema

{
  "schemaVersion": 1,
  "generatedAt": "2026-09-26T11:02:14.318Z",
  "harness": {
    "platform": "darwin",
    "electron": "43.3.0",
    "measuredIterations": 5,
    "warmupIterations": 1
  },
  "journeys": {
    "launch": {
      "counters": { "renderer.ipcCallsToFirstCard": 12 },
      "counterStability": {
        "renderer.ipcCallsToFirstCard": {
          "stable": true,
          "values": [12, 12, 12, 12, 12]
        }
      },
      "wallClock": {
        "spawnToFirstCardMs.p50": 1234.5,
        "spawnToFirstCardMs.p90": 1300.1
      },
      "unavailable": { "renderer.cdTicksToFirstCard": "reason" },
      "iterations": [
        {
          "index": 0,
          "warmup": true,
          "pid": 1,
          "counters": {},
          "wallClock": {},
          "evidence": {}
        }
      ]
    }
  }
}

journeys.<id>.counters.<name> and journeys.<id>.wallClock.<name> are plain numbers so tools/performance/check-journey-ratchet.mjs can compare them with tools/performance/journey-baselines.json. A J1 runtime baseline is added once its counter is deterministic on the CI runner; the launch counters are not yet (see Ratchet), so the summary is evidence only.

renderer.initialBytes

The bytes a browser fetches before Angular can bootstrap, read from the built dist/apps/web/index.html:

  • index.html itself,
  • every same-origin <script src>, including assets/app-config.js,
  • every <link rel="stylesheet">,
  • every <link rel="modulepreload"> chunk.

Manifest, icons, external URLs, commented-out tags and lazy chunks are not counted, so the value is the same for every language: a non-English launch additionally fetches that language's Angular locale chunk (about 2 KB), which belongs to the per-profile J1 benchmark rather than to this counter. A file that index.html references but the build did not emit is an error, never zero bytes. The value is raw (uncompressed) size, which is what the renderer parses. It is Angular's "Initial total" plus index.html and assets/app-config.js (about 4 KB together), so it sits slightly above the rounded figure the build prints; never copy that figure into a baseline, use the script's output. The bundle embeds only the app version from package.json (a named import, which esbuild tree-shakes), not the whole file, so editing scripts or dependencies does not move the counter.

pnpm nx build web                                # production configuration
pnpm run perf:initial-bytes                      # human-readable breakdown
pnpm --silent run perf:initial-bytes -- --json   # machine-readable; --silent keeps pnpm's headers out of stdout
node tools/performance/measure-initial-bytes.mjs --summary dist/performance/journey-summary.json

--summary writes the journey summary shape (journeys.<journey>.counters) that the ratchet checker consumes. --dist <dir> points the script at another build output, for example the electron-performance configuration.

The measurement script is tools/performance/measure-initial-bytes.mjs; its Node tests run with pnpm nx test performance-tools (Tier B in the coverage policy) and lint with pnpm nx lint performance-tools.

Ratchet

tools/performance/journey-baselines.json holds one entry per journey and counter:

{
  "journeys": {
    "launch": {
      "renderer.initialBytes": {
        "value": 2739510,
        "unit": "bytes",
        "updatedAt": "2026-09-26",
        "evidencePr": 1693,
        "measuredWith": "pnpm nx build web && pnpm run perf:initial-bytes"
      }
    }
  }
}

tools/performance/check-journey-ratchet.mjs compares a journey summary with that file:

  • a counter above its value fails; counters are exact, there is no slack;
  • a wall-clock entry carries toleranceRatio and fails above value × toleranceRatio;
  • a baseline with no measurement in the summary fails, so dropping a measurement cannot disable the ratchet; a counter is read only from journeys.<journey>.counters and a wall-clock entry (one with toleranceRatio) only from journeys.<journey>.wallClock, so a value in the wrong section also counts as missing;
  • a measurement below its baseline passes and prints a "tighten" hint;
  • a measured counter without a baseline is noted, not failed;
  • checking nothing fails: an empty baselines file, or --only naming an entry that does not exist, cannot exit 0.

--only <journey>/<counter> (repeatable) restricts the check to the named baselines. A script that measures one counter writes its own summary file and checks only its counter, so it neither overwrites another measurement's summary nor fails the other baselines as unmeasured.

pnpm run perf:initial-bytes:check   # measure dist/apps/web into dist/performance/initial-bytes.summary.json, check only that counter
pnpm run perf:ratchet:check         # check every baseline against dist/performance/journey-summary.json

CI runs perf:initial-bytes:check in the Initial bytes ratchet job of .github/workflows/ci.yml after a production build of apps/web, and uploads dist/performance/ as the performance-journey-summary artifact. Like the rest of that workflow it runs for pull requests that target master and for pushes to master; a stacked PR that targets another branch gets no run until it is retargeted, so dispatch one with gh workflow run ci.yml --ref <branch> when you need the number. A PR that grows the counter fails that job.

That runner is the canonical measurer: take baseline values from its output, not from a local build. A local macOS build of the code before #1695 is 2 bytes smaller in main.js (the eager locale imports); since #1695 the two have been byte-identical. (An apparent 556-byte platform difference during the first measurements was otherwise package.json text embedded in main.js, which moved with every script edit; #1692 fixed that by importing only the version.)

Two effects make the exact counter move for reasons outside a PR's own diff. A baseline lowered on a branch that predates a concurrent master merge can sit below what the merged code measures: #1712 lowered it on a branch without #1714, so master measured 108 bytes over and every later PR failed the job until a follow-up moved lazy-only modules out of main.js. Re-run the job on an up-to-date branch before merging a baseline change. And the bundler's chunk-level identifier renaming shifts when a module enters or leaves main.js: moving one service out once renamed an imported identifier at 162 call sites, eating about 320 of the bytes saved. Judge a small change by the --stats-json input sizes, not only by the counter.

The job also refuses a weakened baselines file: tools/performance/check-baseline-direction.mjs compares journey-baselines.json with the revision the change is measured against (the target branch of a pull request, the previous head of a master push, master for a manual dispatch) and fails when any entry's enforced limit (value × toleranceRatio) went up, a tolerance widened or an entry disappeared, so a PR cannot grow the payload and raise the baseline to match. Lowered limits and new entries pass.

Baselines only move down. Lower value in the same PR as the change that earned it, set updatedAt and evidencePr, and paste the measurement output into the PR. Never raise a value to make a PR pass: if growth is a deliberate trade-off, say so in the PR and let the maintainer decide.

The runtime counters come from the Performance journeys job of the same workflow, on ubuntu-latest only. It runs pnpm run perf:journeys under xvfb-run (the Nx target builds electron-backend:build-performance, the Playwright config starts the Xtream mock), writes the measurements to the job summary and uploads dist/performance/journeys/ as the performance-journeys artifact. The Performance journeys scope job skips it only for pull requests that change nothing but Markdown, docs/**, .plans/**, .codex/**, .claude/**, .changes/** or apps/website/** (the E2E workflow's ignore list plus release notes); any other file, including root build inputs such as .nvmrc, nx.json or tsconfig.base.json, runs it. Pushes to master and manual dispatches always run it. The job is warn-only (continue-on-error: true) for its first two weeks (plan item B3): a regression marks the job failed without failing the workflow. Making it required is a maintainer decision.

No J1 runtime counter is enforced yet. Three dispatched runs on 2026-09-27 (CI runs 36271875209, 36271879955 and 36271884616) reported the same summary values, renderer.ipcCallsToFirstCard 16 and renderer.domMutationsToFirstCard 939, but the third run marked both stable: false: its warm-up and one measured iteration reached the first card in about 750 ms with 13 bridge calls and 576 mutations, the others in about 1,400 ms with 16 and 939. The three extra calls (downloadsGetDefaultFolder and two dbGetGlobalRecentlyAdded) land before or after the first card depending on that race, so neither counter is promoted until the race is understood and the counters are deterministic. renderer.layoutShiftScore (0) and renderer.longTasks (2) were identical in all eighteen runner iterations; the spawnToFirstCardMs P50 ranged from 1,401 to 1,674 ms. All four stay evidence for now. Runner counters also differ from a Mac (12 and 571 there, the fast path without the Linux-only getWindowState call), so take J1 baseline values from the runner only.

Adding a counter

  1. Produce the value from the built output or from a deterministic probe, not from source heuristics. Missing inputs must fail the measurement.
  2. Emit it under journeys.<journey>.counters.<name> in the summary JSON.
  3. Cover the extraction and the failure modes with node --test and register the test file in tools/performance/project.json.
  4. Validate the counter before it becomes a guardrail: one PR must show that lowering it moved wall-clock in the same journey.

Adding a journey

  1. Add apps/electron-backend-e2e/src/journeys/<journey>.journey.ts. Seed the profile through the app's dialogs, spawn a fresh process per iteration with measureLaunchJourney as the model, and drive the journey's start action with Playwright.
  2. Give the journey its own probe options (cardSelector, routeFragment, terminal condition) or extend journey-renderer-probe.ts when the end condition is not "an element became visible". Keep the probe self-contained: Playwright serializes it with toString().
  3. Map the measurement to a JourneyIterationRecord in a <journey>-journey-record.ts under src/performance/; name counters renderer.* or main.*, and list counters you cannot measure under unavailable with the reason.
  4. Add the journey under journeys.<id> in the summary through summarizeJourneyIterations; the schema needs no change.
  5. Cover the probe with jsdom fixtures and the record and summary code with node:test (pnpm nx run electron-backend-e2e:test-performance-harness).
  6. Validate a counter before it becomes a guardrail: one PR must show that lowering it moved wall-clock in the same journey.