Commit Graph
32 Commits
Author SHA1 Message Date
58c54a3286 test(perf): add a settled layout-shift counter to J1 (#1756)
* test(perf): add a settled layout-shift counter to the J1 launch journey

renderer.layoutShiftScore stops at the first-card cutoff, so the dashboard
shift fixed in #1738 (about 0.23, roughly 15 ms after the first card) read
as 0. renderer.layoutShiftScoreSettled sums the same non-input layout-shift
entries until the workspace content has been quiet for 500 ms, capped at
3 s after the cutoff. The first-card counter is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(perf): attribute late layout shifts and record the first measurement

evidence.settle.lateShifts lists the counted shifts after the first-card
cutoff with the nodes the browser attributes them to. On master the settled
score is 0.236: the recent-sources rail moves up 316 px about 12 ms after
the first card and back down shortly after, a flicker #1738 did not cover.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(perf): end the settle window at its scheduled deadline

A cap or quiet timer delayed by a busy main thread used the moment it ran
as the settle point, so shifts after the 3 s cap could enter the settled
counter. The settle point is now the deadline the timer was scheduled for.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(perf): record the runner's settled layout-shift measurement

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: 4gray <fourgray@proton.me>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 11:27:57 +02:00
4gray 9631edf058 fix(e2e): run electron-backend-e2e targets without nested pnpm exec (#1752) 2026-09-30 07:04:29 +02:00
28b022ff48 test(perf): add the J2 open-source journey (#1730)
* test(perf): add the J2 open-source journey

Measure the click on the Xtream portal card until the category list and
the first page of the opened section are painted, in the same fresh
process as J1 after its counters are final and the app has settled.

- journey-renderer-probe: optional click start (capture-phase listener on
  window, start sentinel before the app sees the click, entries before the
  click dropped), companion selectors, recent-input layout shifts tallied
- journey-main-ipc-capture: optional start sentinel; counts calls between
  the two sentinels
- journey-mock-request-ledger: loopback proxy that counts every request
  the app sends to the mock without storing credentials
- open-source-journey-record: J2 counters and evidence
- journey-run / journey-summary: every journey spec of one perf:journeys
  run adds its entry to the same summary.json
- docs: J2 contract in performance-journeys.md

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(perf): stop echoing the request URL from the ledger spec's upstream

CodeQL flagged the fake upstream as reflected XSS. It now records what it
received server-side and answers with a fixed text/plain body.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(perf): address J2 review findings

- Sentinels use cancelSourceProbe: the preload traces the call before
  forwarding it and SOURCE_HEALTH_CANCEL is an in-memory map lookup, so a
  marker no longer runs a SQLite query on the worker ahead of the measured
  work (Codex P1). A spec pins that handler contract.
- A run is started only in the Playwright runner, replacing inherited
  values, and carries a random harness.runId; summaries from another
  invocation are never merged (Greptile P1, Codex P2).
- The mock ledger tracks in-flight requests; settling and the HTTP window
  require none in flight (Codex P2).
- clickToFirstPagePaintMs reports click to the committed paint next to
  the terminal-batch clickToFirstPageMs (Codex P1).
- The Playwright attachment carries the whole summary (Greptile P2).
- jsdom probe specs wait for the post-paint cutoff instead of a fixed
  40 ms, which flaked when the harness runs all files in parallel.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(validation): describe perf:journeys as running J1 and J2

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(perf): settle J2 on pending bridge calls and start HTTP at the click

- The journey IPC capture pairs every traced start with its success or
  error and exposes the calls still in flight. J2 settles only when J1's
  capture, installed before the document loaded, has none pending, so a
  slow startup call cannot resolve after the click and count as J2.
- The mock HTTP window starts at the renderer's click stamp instead of
  the test-side mark taken before Playwright's actionability checks.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(perf): restart the J2 quiet period when pending work completes

Both waits in the open-source journey (settling before the click, closing
the mock window after the terminal) now use one waitForJourneyQuiet
helper that compares whole samples, in-flight counts included. The poll
that first sees a request or bridge call complete restarts the quiet
period, so the window is never measured from a poll at which work was
still pending. A fake-clock spec covers the in-flight to zero case.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(perf): align J2 with the main-process counters from #1715

After rebasing on #1715, J1 measures main.sqlStatementsBeforeReadyToShow,
so J2's reason for listing main.sqlStatementsToFirstPage as unavailable
(no countable channel) was stale. State the actual limit: the running
total is read from the test process and cannot be bounded at the click
or the first-page batch. The performance-journeys CI job comment now
names both journeys.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(perf): launch J2 without SQL counting and stamp mock requests in sub-ms

- runLaunchJourney takes the launch instrumentation; only J1 turns on the
  main-process counters and IPTVNATOR_PERF_COUNT_SQL, so J2's click is not
  measured under the hook that wraps every SQLite statement. The flags are
  built in journey-launch-environment.ts, which the SQL opt-in guard now
  expects, and a launch record without main counters is rejected.
- The mock ledger stamps arrivals with performance.timeOrigin +
  performance.now(), the same sub-millisecond epoch as the renderer's
  click, so a request later in the click's millisecond is not counted
  before it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(perf): reject J2 iterations with activity after settling

- The open-source record compares the settle snapshot with what the probe
  and the IPC capture counted up to the click event, and with the mock
  requests between the snapshot and the click stamp. Any change means
  background work began during Playwright's actionability checks and
  could land in J2, so the iteration is rejected.
- The SQL opt-in guard also checks who passes mainCounters: true: only
  measureLaunchJourney may, and J2 must pass false.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(perf): count the long task that dispatches the J2 click

A long task's startTime precedes the click event's timestamp when the
listener runs inside it, so the start-time filter dropped the task that
performs the interaction. Long tasks now count when their range overlaps
the window: on one main thread only the dispatching task can overlap the
click. Layout shifts keep the start-time filter. J1 is unchanged (its
window starts at -Infinity).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(perf): bound J2's late-request check by the quiet sample's mark

The late-activity check compared requests against a fresh ledger mark
taken after waitForQuiet returned. A request that arrived while the final
quiet sample was still reading the IPC capture advanced that mark and
escaped the check. The boundary is now the ledger position read by the
accepted sample itself, like its DOM and IPC counts.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(perf): end J2's HTTP window at the accepted quiet sample

The post-terminal window read the ledger after waitForMockQuiet returned,
so a request arriving in between was counted although its completion was
never waited for. waitForMockQuiet now returns the ledger position its
accepted sample read; later requests are kept as evidence
(httpRequestsAfterSettledByRoute) instead of the counter.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(perf): observe late mock requests before reading J2's ledger

httpRequestsAfterSettledByRoute read the ledger right after the accepted
quiet sample, so late requests had no chance to appear in it. The ledger
is now read after another quiet interval; the counter stays bounded by
the quiet sample's mark and late traffic shows up in the evidence.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(perf): fail the J2 quiet wait when a sample stalls past its deadline

waitForJourneyQuiet accepted a sample that returned unchanged after a
stall longer than the timeout as the end of a quiet period, before the
deadline check ran. The deadline is now checked first, so a stalled
sample fails the wait instead of letting the click go ahead unobserved.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(perf): detach J1's IPC capture before the J2 click

J2 used J1's capture to see pending launch bridge calls while settling,
but its ipcMain listener stayed attached and ran for every bridge call of
the measured click. The capture can now be detached; J2 detaches J1's
right after settling, before the click.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(perf): sample both J2 settle captures in one main-process snapshot

The settle sample read J1's capture (pending calls) and J2's capture
(call count) in two evaluate calls, so a call starting in between was
counted with a stale zero in flight and its completion went unseen. Both
states are now read in one synchronous pass, where no ipcMain event can
be handled in between.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: 4gray <fourgray@proton.me>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 20:37:45 +02:00
4gray 8b570e3390 test(performance): measure two-byte string cost in M3U and XMLTV parsing (D1) (#1713) 2026-09-28 00:03:18 +02:00
1d9a563d1a test(performance): count startup phases and SQL statements for the J1 launch journey (#1715)
* test(performance): count startup phases and SQL statements for the J1 launch journey

Implements plan item A2. With IPTVNATOR_PERF_CAPTURE=1 the main process
keeps named counters and registers a main-only performance:read-counters
IPC handler; without the flag nothing is counted and the handler does not
exist.

- debug-trace.ts owns the registry; traceStartupPhase replaces the
  trace('startup', ...) sites and counts main.startupPhases.
- The database worker counts executed statements through better-sqlite3's
  Statement prototype (the verbose callback expands every statement and
  made bulk inserts 2-4x slower) and posts the count over its message
  port, flushed before every other worker message. The main-thread shared
  connection is counted through a new connection observer in the shared
  database library.
- The first main window freezes main.modulesRegisteredBeforeWindow at
  creation and main.sqlStatementsBeforeReadyToShow at ready-to-show.
- The journey gate drops the ready-to-show that Electron emits for the
  about:blank detour, so the app sees the real document's first paint,
  and taps the counters handler; the J1 record reads both counters after
  the renderer probe completes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(database): require one SQL statement per exec during initialization

The performance capture counts one exec call as one statement, because
SQL cannot be split reliably in the counter (trigger bodies contain
semicolons). The historical-upgrade driver now wraps exec on every
connection initDatabase opens and fails on a batch, so that counting
assumption holds for the fresh profile and all historical schemas.
Documents the definition in the counter and the architecture docs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(performance): count SQL statements only for the launch journey

Codex review: the M3U import, refresh-cancellation and Xtream benchmarks
also run with IPTVNATOR_PERF_CAPTURE=1, so the statement hook wrapped
every row of their bulk inserts and changed what they measure.

SQL counting now also needs IPTVNATOR_PERF_COUNT_SQL=1, which only the
launch journey sets; a harness test fails if another source sets it.
Startup phases, the window snapshot and the read handler stay on the
capture flag. Without SQL counting no ready-to-show listener is attached,
so a zero is never reported for statements nobody counted.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: 4gray <fourgray@proton.me>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-27 21:01:26 +02:00
4grayandClaude Fable 5.1 e39c854a41 perf(electron): load the main-process startup wiring after the window starts loading (#1702)
Registering IPC handlers costs 0.4 ms; evaluating the modules behind them (axios, drizzle-orm, better-sqlite3, electron-updater, fix-path) before the window could load was the real cost. main.ts now keeps only the pre-paint wiring and loads the rest as the deferred-events.js chunk inside the main window's did-start-loading listener, where the import and its registrations complete before any renderer invoke can arrive.

Interleaved A/B on the performance build: app.whenReady 323 -> 265 ms, did-finish-load 499 -> 447 ms; J1 journey spawnToDidFinishLoad ~405 -> ~365 ms with identical counters. Packaging ships the chunk explicitly, verify:package-layout requires it, and the benchmark build identity hashes it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-26 21:51:19 +02:00
4grayandClaude Fable 5.1 587e19541f perf(electron): compile the main-process bundle from a V8 code cache (#1696)
main.js is now a small entry that enables Node's on-disk V8 compile cache under userData/v8-compile-cache and then requires the application bundle main.app.js. Warm launches reach app.whenReady about 13 ms sooner at the median; IPTVNATOR_DISABLE_COMPILE_CACHE=1 turns it off and IPTVNATOR_COMPILE_CACHE_DIR relocates it.

Packaging now ships main.app.js explicitly and verify:package-layout requires the main-process entry files in app.asar, because nx-electron copies the backend through an allowlist. The Xtream benchmark build identity hashes both the launcher and the bundle.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-26 19:19:30 +02:00
0f1cad4b16 test(performance): add the J1 launch-to-usable journey benchmark (#1698)
* test(performance): add the J1 launch-to-usable journey benchmark

Implements plan items A1, A3 (J1 only) and the minimal A4 from
.plans/2026-09-25-performance-journeys-ratchet.md.

- journey-renderer-probe.ts: init-script probe counting DOM mutations,
  layout shifts and long tasks until the first source card is visible on
  /workspace with the splash removed; unit-tested with jsdom fixtures.
- journey-main-ipc-capture.ts: counts bridge invocations from the preload's
  renderer-API trace channel up to a sentinel call the probe fires, so the
  IPC counter is exact without touching production code.
- launch.journey.ts + playwright.journeys.config.ts: seeded profile (one
  M3U source, one Xtream portal on the loopback mock), one warm-up and five
  measured iterations, fresh process and data directory each, writing
  dist/performance/journeys/<timestamp>/summary.json with exact counters
  and P50/P90 wall-clock.
- Nx target electron-backend-e2e:journeys and root script perf:journeys.
- docs/architecture/performance-journeys.md, README, context and
  validation map entries.

renderer.cdTicksToFirstCard and main.sqlStatementsBeforeReadyToShow are
reported as unavailable with the reason instead of being faked.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(performance): gate the renderer load so the probes never race startup

Review follow-up for #1698.

- journey-renderer-gate.cjs: a main-process hook loaded with `-r` (the
  mechanism Playwright uses for its own loader) makes the first
  loadFile/loadURL navigate to about:blank and holds the real load until
  the test releases it. Playwright reports no page before a navigation
  commits, so this is what lets the renderer probe be registered on the
  page before the real document exists; the IPC capture is installed
  before the release too. A safety timeout releases the gate on its own
  and marks the iteration invalid. Unit-tested with a fake BrowserWindow.
- launch-journey-app.ts: registers the probe on the parked page, releases
  the gate, waits for the real document to commit, fails fast when the
  probe is missing, and refuses an iteration whose gate timed out, saw a
  second load, or released before the probe was in place.
- journey-renderer-probe.ts: entries delivered live after the terminal
  batch are buffered and filtered by the same cutoff as queued ones, and
  the cutoff is sampled in a timer queued from the first rAF, i.e. after
  the card's frame is painted, so the render task's long task and layout
  shift are consistently included.
- launch-journey-record.ts: layoutShiftScore rounded to three decimals; a
  0.0001 shift flipped in and out of the cutoff between iterations.
- playwright.journeys.config.ts: reuse a mock server left on the journeys
  port locally (its fixtures are deterministic); CI still starts its own.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(performance): validate the journey gate after the probe completes

Codex follow-up on #1698: the gate state returned by release() cannot see a
reload or recovery navigation that happens before the first card. Re-read the
live state once the renderer probe has finished and validate that instead,
so an iteration spanning an extra navigation is rejected.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(performance): fail an iteration whose performance observers were unavailable

Codex follow-up on #1698: a renderer that cannot observe layout-shift or
longtask entries used to pass the probe with zero counters, which a ratchet
could not tell apart from a genuine zero. The probe assertion now rejects
such an iteration and names the missing observer.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: 4gray <fourgray@proton.me>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-26 19:18:19 +02:00
4gray 0ecfc06494 test(electron): fix Windows cleanup and retain packaged smoke diagnostics (#1665)
* test(electron): reap Windows process trees and retain smoke diagnostics

* test(electron): require confirmed process exit before relaunch

* test(electron): retain process handle after application exit

* test(electron): bind captured processes to application instances
2026-09-23 22:55:48 +02:00
4gray b1f77c678e test(performance): prevent renderer heartbeat omission (#1308)
* test(performance): normalize sub-ms IPC clock skew

* test(performance): prevent heartbeat coordinated omission
2026-07-29 13:58:44 +02:00
4grayandClaude Opus 5 9b7776a901 chore(lint): hold tests to their own max-lines ceiling (#1306)
* chore(lint): hold tests to their own max-lines ceiling

The flat 400-line cap treated a spec like a component. A spec is a flat
list of independent cases, so hitting the cap there produces arbitrary
`-2.spec.ts` splits and hides coverage instead of surfacing design debt —
65 of the 138 files over the limit were tests.

Production code keeps 400. Tests (`**/*.spec.ts`, `**/*.e2e.ts`, and
everything under `apps/*-e2e/**`) get 1200. Blank lines and comments no
longer count, so a docblock can't be the reason a file must be split.

Both limits now live in tools/eslint/max-lines-config.mjs, imported by
eslint.config.mjs and the baseline generator alike. The generator decides
who belongs on the list by running ESLint's own max-lines rule instead of
counting lines itself — a private reimplementation would disagree with the
rule the moment either side changed (a `//` inside a template literal is
enough) and yield a baseline that turns CI red while looking correct.

The baseline drops 126 -> 68 entries with nothing added, and six now-dead
`eslint-disable max-lines` directives are removed. A new eslint-tools test
asserts the committed baseline still matches what the generator produces,
so a stale entry or a forgotten regeneration fails CI.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(lint): classify eslint-tools in the coverage policy

A project with a `test` target must be assigned a coverage tier, so
adding eslint-tools broke `coverage:policy:check` before the unit suite
even ran. Tier B alongside packaging and release-tools: these are Node
tests over lint tooling, and a coverage percentage across a generated
list would not mean anything.

CI runs Tier B/C through its own `--run-non-tier-a` step, so the
baseline-consistency test executes there rather than being skipped.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 08:08:04 +02:00
4gray 055170d188 test(performance): harden Xtream startup retry (#1307)
* test(performance): harden Xtream startup retry

* test(performance): preserve Xtream teardown failures

* test(performance): retry Xtream profile cleanup
2026-07-29 01:05:32 +02:00
4gray faec40ff7b test(performance): stabilize Xtream benchmark startup (#1304)
* test(performance): stabilize Xtream benchmark startup

* test(performance): bound cancellation clock skew
2026-07-28 22:55:49 +02:00
4gray bc4e3a2e2c test(performance): bound worker sampling finalization (#1302)
* test(performance): bound worker sampling finalization

* test(performance): reject timed-out worker captures

* test(performance): settle every worker sample
2026-07-28 21:30:35 +02:00
4gray a2fafcfc08 test(performance): add end-to-end Xtream benchmark harness (#1300)
* docs(performance): plan Xtream benchmark

* feat(xtream-mock-server): add deterministic 100k fixture

* style(xtream-mock-server): apply repository formatting

* fix(xtream-mock-server): harden performance fixture data

* feat(xtream-mock-server): add performance control plane

* docs(performance): correct Xtream capture plan

* fix(xtream-mock-server): harden performance controls

* fix(xtream-mock-server): harden control lifecycle

* feat(performance): add Xtream preload markers

* feat(performance): trace Xtream main phases

* feat(performance): mark Xtream store publications

* feat(performance): trace Xtream database phases

* feat(performance): trace Xtream delete cancellation

* feat(performance): capture Xtream phase attribution

* feat(performance): mark Sources Xtream refresh

* test(performance): define Xtream benchmark evidence contracts

* test(performance): add Xtream benchmark runner

* test(performance): surface failure evidence writes

* test(performance): align database read clock

* test(performance): preserve capture failure contracts
2026-07-28 08:08:07 +02:00
4gray 24f0dee6f0 test(performance): add formal M3U import benchmark (#1287)
* test(performance): add formal M3U import benchmark

* test(performance): harden formal capture validity

* test(performance): address benchmark review feedback
2026-07-27 10:34:22 +02:00
4gray 2fda6cea07 fix(perf): reject incomplete renderer RSS samples 2026-07-27 02:27:57 +02:00
4gray 554eecfee0 fix(perf): finalize exact worker capture safely 2026-07-27 02:15:59 +02:00
4gray d3fdbaa2ef fix(perf): aggregate every worker isolate 2026-07-27 02:15:59 +02:00
4gray b7ddb5e90a chore(perf): add database worker heap probe transport 2026-07-27 02:15:59 +02:00
4gray 1fe6f30e64 chore(perf): prepare database post-GC capture 2026-07-27 02:15:59 +02:00
4gray 918c8f2ded fix(perf): scope renderer RSS to exact window 2026-07-27 02:15:59 +02:00
4gray 5aab81e2cb refactor(perf): expose serializable renderer RSS capture 2026-07-27 02:15:58 +02:00
4gray e27685dc5d fix(perf): model exact renderer RSS capture 2026-07-27 02:15:58 +02:00
4gray 5533ac0fc1 fix(perf): preflight profile artifact capacity 2026-07-27 02:15:58 +02:00
4gray a3a8f6e90c fix(perf): optimize benchmark worker builds 2026-07-26 23:25:09 +02:00
4gray 8bafd62932 fix(perf): benchmark the performance build 2026-07-26 22:49:36 +02:00
4gray da3b657277 fix(perf): make worker profiling request scoped 2026-07-26 22:37:52 +02:00
4gray f9e71a6d8d fix(electron): isolate refresh performance correlation 2026-07-26 22:37:52 +02:00
4gray 2cc7d0acb1 fix(perf): declare renderer cache output path 2026-07-26 22:37:52 +02:00
4gray 28ed38906d test(perf): add production renderer benchmark build 2026-07-26 22:37:52 +02:00
4gray f147d4fe37 perf(m3u): stop cancelled refresh workers (#1268) 2026-07-26 09:25:51 +02:00