Commit Graph
3 Commits
Author SHA1 Message Date
28b022ff48 test(perf): add the J2 open-source journey (#1730)
* test(perf): add the J2 open-source journey

Measure the click on the Xtream portal card until the category list and
the first page of the opened section are painted, in the same fresh
process as J1 after its counters are final and the app has settled.

- journey-renderer-probe: optional click start (capture-phase listener on
  window, start sentinel before the app sees the click, entries before the
  click dropped), companion selectors, recent-input layout shifts tallied
- journey-main-ipc-capture: optional start sentinel; counts calls between
  the two sentinels
- journey-mock-request-ledger: loopback proxy that counts every request
  the app sends to the mock without storing credentials
- open-source-journey-record: J2 counters and evidence
- journey-run / journey-summary: every journey spec of one perf:journeys
  run adds its entry to the same summary.json
- docs: J2 contract in performance-journeys.md

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(perf): stop echoing the request URL from the ledger spec's upstream

CodeQL flagged the fake upstream as reflected XSS. It now records what it
received server-side and answers with a fixed text/plain body.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(perf): address J2 review findings

- Sentinels use cancelSourceProbe: the preload traces the call before
  forwarding it and SOURCE_HEALTH_CANCEL is an in-memory map lookup, so a
  marker no longer runs a SQLite query on the worker ahead of the measured
  work (Codex P1). A spec pins that handler contract.
- A run is started only in the Playwright runner, replacing inherited
  values, and carries a random harness.runId; summaries from another
  invocation are never merged (Greptile P1, Codex P2).
- The mock ledger tracks in-flight requests; settling and the HTTP window
  require none in flight (Codex P2).
- clickToFirstPagePaintMs reports click to the committed paint next to
  the terminal-batch clickToFirstPageMs (Codex P1).
- The Playwright attachment carries the whole summary (Greptile P2).
- jsdom probe specs wait for the post-paint cutoff instead of a fixed
  40 ms, which flaked when the harness runs all files in parallel.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(validation): describe perf:journeys as running J1 and J2

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(perf): settle J2 on pending bridge calls and start HTTP at the click

- The journey IPC capture pairs every traced start with its success or
  error and exposes the calls still in flight. J2 settles only when J1's
  capture, installed before the document loaded, has none pending, so a
  slow startup call cannot resolve after the click and count as J2.
- The mock HTTP window starts at the renderer's click stamp instead of
  the test-side mark taken before Playwright's actionability checks.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(perf): restart the J2 quiet period when pending work completes

Both waits in the open-source journey (settling before the click, closing
the mock window after the terminal) now use one waitForJourneyQuiet
helper that compares whole samples, in-flight counts included. The poll
that first sees a request or bridge call complete restarts the quiet
period, so the window is never measured from a poll at which work was
still pending. A fake-clock spec covers the in-flight to zero case.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(perf): align J2 with the main-process counters from #1715

After rebasing on #1715, J1 measures main.sqlStatementsBeforeReadyToShow,
so J2's reason for listing main.sqlStatementsToFirstPage as unavailable
(no countable channel) was stale. State the actual limit: the running
total is read from the test process and cannot be bounded at the click
or the first-page batch. The performance-journeys CI job comment now
names both journeys.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(perf): launch J2 without SQL counting and stamp mock requests in sub-ms

- runLaunchJourney takes the launch instrumentation; only J1 turns on the
  main-process counters and IPTVNATOR_PERF_COUNT_SQL, so J2's click is not
  measured under the hook that wraps every SQLite statement. The flags are
  built in journey-launch-environment.ts, which the SQL opt-in guard now
  expects, and a launch record without main counters is rejected.
- The mock ledger stamps arrivals with performance.timeOrigin +
  performance.now(), the same sub-millisecond epoch as the renderer's
  click, so a request later in the click's millisecond is not counted
  before it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(perf): reject J2 iterations with activity after settling

- The open-source record compares the settle snapshot with what the probe
  and the IPC capture counted up to the click event, and with the mock
  requests between the snapshot and the click stamp. Any change means
  background work began during Playwright's actionability checks and
  could land in J2, so the iteration is rejected.
- The SQL opt-in guard also checks who passes mainCounters: true: only
  measureLaunchJourney may, and J2 must pass false.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(perf): count the long task that dispatches the J2 click

A long task's startTime precedes the click event's timestamp when the
listener runs inside it, so the start-time filter dropped the task that
performs the interaction. Long tasks now count when their range overlaps
the window: on one main thread only the dispatching task can overlap the
click. Layout shifts keep the start-time filter. J1 is unchanged (its
window starts at -Infinity).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(perf): bound J2's late-request check by the quiet sample's mark

The late-activity check compared requests against a fresh ledger mark
taken after waitForQuiet returned. A request that arrived while the final
quiet sample was still reading the IPC capture advanced that mark and
escaped the check. The boundary is now the ledger position read by the
accepted sample itself, like its DOM and IPC counts.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(perf): end J2's HTTP window at the accepted quiet sample

The post-terminal window read the ledger after waitForMockQuiet returned,
so a request arriving in between was counted although its completion was
never waited for. waitForMockQuiet now returns the ledger position its
accepted sample read; later requests are kept as evidence
(httpRequestsAfterSettledByRoute) instead of the counter.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(perf): observe late mock requests before reading J2's ledger

httpRequestsAfterSettledByRoute read the ledger right after the accepted
quiet sample, so late requests had no chance to appear in it. The ledger
is now read after another quiet interval; the counter stays bounded by
the quiet sample's mark and late traffic shows up in the evidence.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(perf): fail the J2 quiet wait when a sample stalls past its deadline

waitForJourneyQuiet accepted a sample that returned unchanged after a
stall longer than the timeout as the end of a quiet period, before the
deadline check ran. The deadline is now checked first, so a stalled
sample fails the wait instead of letting the click go ahead unobserved.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(perf): detach J1's IPC capture before the J2 click

J2 used J1's capture to see pending launch bridge calls while settling,
but its ipcMain listener stayed attached and ran for every bridge call of
the measured click. The capture can now be detached; J2 detaches J1's
right after settling, before the click.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(perf): sample both J2 settle captures in one main-process snapshot

The settle sample read J1's capture (pending calls) and J2's capture
(call count) in two evaluate calls, so a call starting in between was
counted with a stale zero in flight and its completion went unseen. Both
states are now read in one synchronous pass, where no ipcMain event can
be handled in between.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: 4gray <fourgray@proton.me>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 20:37:45 +02:00
6b855eb73b fix(e2e): stop mock servers from outliving Playwright runs (#1710)
* fix(e2e): stop mock servers from outliving Playwright runs

Playwright stops a webServer with a SIGKILL to the process group it
spawned, but `nx run-commands` starts its command in a detached process
group of its own. Launching the Xtream/Stalker mocks through
`pnpm nx run *-mock-server:serve` therefore left the tsx server running
(reparented to PID 1) and holding its port after every run, so the next
run failed with "…/health is already used" or silently reused a stale
server.

Every Playwright config now starts the mocks as a single
`node --import tsx apps/<mock>/src/main.ts` process with
TSX_TSCONFIG_PATH=tsconfig.base.json, which stays in Playwright's group.
A project-config spec guards all playwright*.config.ts files against
regressing to the Nx launch.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(e2e): read sidebar categories atomically; tighten mock launch guard

- category-management: readVisibleSidebarCategoryNames read items one by
  one; when Save removed an item between isVisible() and textContent(),
  textContent() auto-waited for the gone label through the whole 15 s
  poll, so expect.poll never retried (ubuntu shard 1 failed 3/3 while the
  UI already showed "No categories available"). Take one snapshot with
  filter({ visible: true }).evaluateAll() instead.
- project-config.spec: pin which Playwright configs start which mock,
  reject any Nx form that mentions a mock server, and fail when a new
  config starts a mock without being listed (the old count check passed
  vacuously on zero matches).
- docs: state which configs start which mock instead of "every config
  starts both".

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* revert(e2e): leave the sidebar category read race to #1728

#1728 fixes the same readVisibleSidebarCategoryNames race with a shared
helper; keeping a second copy here would only conflict. This PR stays
about mock-server lifecycle.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: 4gray <fourgray@proton.me>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-27 22:57:46 +02:00
0f1cad4b16 test(performance): add the J1 launch-to-usable journey benchmark (#1698)
* test(performance): add the J1 launch-to-usable journey benchmark

Implements plan items A1, A3 (J1 only) and the minimal A4 from
.plans/2026-09-25-performance-journeys-ratchet.md.

- journey-renderer-probe.ts: init-script probe counting DOM mutations,
  layout shifts and long tasks until the first source card is visible on
  /workspace with the splash removed; unit-tested with jsdom fixtures.
- journey-main-ipc-capture.ts: counts bridge invocations from the preload's
  renderer-API trace channel up to a sentinel call the probe fires, so the
  IPC counter is exact without touching production code.
- launch.journey.ts + playwright.journeys.config.ts: seeded profile (one
  M3U source, one Xtream portal on the loopback mock), one warm-up and five
  measured iterations, fresh process and data directory each, writing
  dist/performance/journeys/<timestamp>/summary.json with exact counters
  and P50/P90 wall-clock.
- Nx target electron-backend-e2e:journeys and root script perf:journeys.
- docs/architecture/performance-journeys.md, README, context and
  validation map entries.

renderer.cdTicksToFirstCard and main.sqlStatementsBeforeReadyToShow are
reported as unavailable with the reason instead of being faked.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(performance): gate the renderer load so the probes never race startup

Review follow-up for #1698.

- journey-renderer-gate.cjs: a main-process hook loaded with `-r` (the
  mechanism Playwright uses for its own loader) makes the first
  loadFile/loadURL navigate to about:blank and holds the real load until
  the test releases it. Playwright reports no page before a navigation
  commits, so this is what lets the renderer probe be registered on the
  page before the real document exists; the IPC capture is installed
  before the release too. A safety timeout releases the gate on its own
  and marks the iteration invalid. Unit-tested with a fake BrowserWindow.
- launch-journey-app.ts: registers the probe on the parked page, releases
  the gate, waits for the real document to commit, fails fast when the
  probe is missing, and refuses an iteration whose gate timed out, saw a
  second load, or released before the probe was in place.
- journey-renderer-probe.ts: entries delivered live after the terminal
  batch are buffered and filtered by the same cutoff as queued ones, and
  the cutoff is sampled in a timer queued from the first rAF, i.e. after
  the card's frame is painted, so the render task's long task and layout
  shift are consistently included.
- launch-journey-record.ts: layoutShiftScore rounded to three decimals; a
  0.0001 shift flipped in and out of the cutoff between iterations.
- playwright.journeys.config.ts: reuse a mock server left on the journeys
  port locally (its fixtures are deterministic); CI still starts its own.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(performance): validate the journey gate after the probe completes

Codex follow-up on #1698: the gate state returned by release() cannot see a
reload or recovery navigation that happens before the first card. Re-read the
live state once the renderer probe has finished and validate that instead,
so an iteration spanning an extra navigation is rejected.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(performance): fail an iteration whose performance observers were unavailable

Codex follow-up on #1698: a renderer that cannot observe layout-shift or
longtask entries used to pass the probe with zero counters, which a ratchet
could not tell apart from a genuine zero. The probe assertion now rejects
such an iteration and names the missing observer.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: 4gray <fourgray@proton.me>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-26 19:18:19 +02:00