Commit Graph
16 Commits
Author SHA1 Message Date
6e229f85b6 test(perf): add the J3 playback journey (#1774)
* test(perf): add the J3 playback journey

J3 clicks a live channel of an Xtream portal and ends at the built-in
HTML5 player's first `playing` event, with `loadedmetadata` as a
secondary phase. It follows J2: every iteration is a fresh J1 launch on
a copy of a profile seeded through the app's dialogs, and the click
happens after the app has settled in the portal's first live category.

The portal is the mock's `live-fallback` account, whose `.ts` live URLs
serve the local H.264/AAC MPEG-TS fixture that mpegts.js plays through
MSE on every platform. The marketing accounts' local live bytes are
zero-filled and never reach `playing`. Seeding selects the HTML5 player
and the `ts` stream format; the catalog's picsum.photos logos are
cancelled from the test side so no request leaves the machine.

Counters: renderer.ipcCallsToPlaying, renderer.httpRequestsToPlaying,
renderer.domMutationsToPlaying, renderer.layoutShiftScore and
renderer.longTasks; wall-clock click->loadedmetadata and click->playing.
renderer.ipcSerialDepthToPlaying is listed as unavailable until the
serial-depth helper lands. No baseline yet.

The probe gains a media-event terminal; J2's pre-click settle moves to
journey-click-settle.ts so both journeys share it unchanged, and the
probe spec's jsdom fixtures move to a shared test helper.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(perf): show J3's HTTP boundary margins and watch 1 s after playing

Review follow-up. renderer.httpRequestsToPlaying compares the ledger's
arrival stamps with the renderer's click and playing stamps, which come
from different processes on the same host clock. Each iteration now
records the distance of the nearest request on either side of both
boundaries, so a count a clock difference could flip is visible.

Requests after playing were a single snapshot taken right after the
probe; the test now watches the ledger for a fixed 1 s after playing.
A live stream never leaves the mock quiet, so J2's quiet wait does not
apply.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: 4gray <fourgray@proton.me>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 08:37:02 +02:00
4gray ec8b931dbf ci(perf): add the weekly baseline tightening workflow (#1760) 2026-09-30 18:50:30 +02:00
58c54a3286 test(perf): add a settled layout-shift counter to J1 (#1756)
* test(perf): add a settled layout-shift counter to the J1 launch journey

renderer.layoutShiftScore stops at the first-card cutoff, so the dashboard
shift fixed in #1738 (about 0.23, roughly 15 ms after the first card) read
as 0. renderer.layoutShiftScoreSettled sums the same non-input layout-shift
entries until the workspace content has been quiet for 500 ms, capped at
3 s after the cutoff. The first-card counter is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(perf): attribute late layout shifts and record the first measurement

evidence.settle.lateShifts lists the counted shifts after the first-card
cutoff with the nodes the browser attributes them to. On master the settled
score is 0.236: the recent-sources rail moves up 316 px about 12 ms after
the first card and back down shortly after, a flicker #1738 did not cover.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(perf): end the settle window at its scheduled deadline

A cap or quiet timer delayed by a busy main thread used the moment it ran
as the settle point, so shifts after the 3 s cap could enter the settled
counter. The settle point is now the deadline the timer was scheduled for.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(perf): record the runner's settled layout-shift measurement

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: 4gray <fourgray@proton.me>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 11:27:57 +02:00
28b022ff48 test(perf): add the J2 open-source journey (#1730)
* test(perf): add the J2 open-source journey

Measure the click on the Xtream portal card until the category list and
the first page of the opened section are painted, in the same fresh
process as J1 after its counters are final and the app has settled.

- journey-renderer-probe: optional click start (capture-phase listener on
  window, start sentinel before the app sees the click, entries before the
  click dropped), companion selectors, recent-input layout shifts tallied
- journey-main-ipc-capture: optional start sentinel; counts calls between
  the two sentinels
- journey-mock-request-ledger: loopback proxy that counts every request
  the app sends to the mock without storing credentials
- open-source-journey-record: J2 counters and evidence
- journey-run / journey-summary: every journey spec of one perf:journeys
  run adds its entry to the same summary.json
- docs: J2 contract in performance-journeys.md

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(perf): stop echoing the request URL from the ledger spec's upstream

CodeQL flagged the fake upstream as reflected XSS. It now records what it
received server-side and answers with a fixed text/plain body.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(perf): address J2 review findings

- Sentinels use cancelSourceProbe: the preload traces the call before
  forwarding it and SOURCE_HEALTH_CANCEL is an in-memory map lookup, so a
  marker no longer runs a SQLite query on the worker ahead of the measured
  work (Codex P1). A spec pins that handler contract.
- A run is started only in the Playwright runner, replacing inherited
  values, and carries a random harness.runId; summaries from another
  invocation are never merged (Greptile P1, Codex P2).
- The mock ledger tracks in-flight requests; settling and the HTTP window
  require none in flight (Codex P2).
- clickToFirstPagePaintMs reports click to the committed paint next to
  the terminal-batch clickToFirstPageMs (Codex P1).
- The Playwright attachment carries the whole summary (Greptile P2).
- jsdom probe specs wait for the post-paint cutoff instead of a fixed
  40 ms, which flaked when the harness runs all files in parallel.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(validation): describe perf:journeys as running J1 and J2

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(perf): settle J2 on pending bridge calls and start HTTP at the click

- The journey IPC capture pairs every traced start with its success or
  error and exposes the calls still in flight. J2 settles only when J1's
  capture, installed before the document loaded, has none pending, so a
  slow startup call cannot resolve after the click and count as J2.
- The mock HTTP window starts at the renderer's click stamp instead of
  the test-side mark taken before Playwright's actionability checks.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(perf): restart the J2 quiet period when pending work completes

Both waits in the open-source journey (settling before the click, closing
the mock window after the terminal) now use one waitForJourneyQuiet
helper that compares whole samples, in-flight counts included. The poll
that first sees a request or bridge call complete restarts the quiet
period, so the window is never measured from a poll at which work was
still pending. A fake-clock spec covers the in-flight to zero case.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(perf): align J2 with the main-process counters from #1715

After rebasing on #1715, J1 measures main.sqlStatementsBeforeReadyToShow,
so J2's reason for listing main.sqlStatementsToFirstPage as unavailable
(no countable channel) was stale. State the actual limit: the running
total is read from the test process and cannot be bounded at the click
or the first-page batch. The performance-journeys CI job comment now
names both journeys.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(perf): launch J2 without SQL counting and stamp mock requests in sub-ms

- runLaunchJourney takes the launch instrumentation; only J1 turns on the
  main-process counters and IPTVNATOR_PERF_COUNT_SQL, so J2's click is not
  measured under the hook that wraps every SQLite statement. The flags are
  built in journey-launch-environment.ts, which the SQL opt-in guard now
  expects, and a launch record without main counters is rejected.
- The mock ledger stamps arrivals with performance.timeOrigin +
  performance.now(), the same sub-millisecond epoch as the renderer's
  click, so a request later in the click's millisecond is not counted
  before it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(perf): reject J2 iterations with activity after settling

- The open-source record compares the settle snapshot with what the probe
  and the IPC capture counted up to the click event, and with the mock
  requests between the snapshot and the click stamp. Any change means
  background work began during Playwright's actionability checks and
  could land in J2, so the iteration is rejected.
- The SQL opt-in guard also checks who passes mainCounters: true: only
  measureLaunchJourney may, and J2 must pass false.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(perf): count the long task that dispatches the J2 click

A long task's startTime precedes the click event's timestamp when the
listener runs inside it, so the start-time filter dropped the task that
performs the interaction. Long tasks now count when their range overlaps
the window: on one main thread only the dispatching task can overlap the
click. Layout shifts keep the start-time filter. J1 is unchanged (its
window starts at -Infinity).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(perf): bound J2's late-request check by the quiet sample's mark

The late-activity check compared requests against a fresh ledger mark
taken after waitForQuiet returned. A request that arrived while the final
quiet sample was still reading the IPC capture advanced that mark and
escaped the check. The boundary is now the ledger position read by the
accepted sample itself, like its DOM and IPC counts.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(perf): end J2's HTTP window at the accepted quiet sample

The post-terminal window read the ledger after waitForMockQuiet returned,
so a request arriving in between was counted although its completion was
never waited for. waitForMockQuiet now returns the ledger position its
accepted sample read; later requests are kept as evidence
(httpRequestsAfterSettledByRoute) instead of the counter.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(perf): observe late mock requests before reading J2's ledger

httpRequestsAfterSettledByRoute read the ledger right after the accepted
quiet sample, so late requests had no chance to appear in it. The ledger
is now read after another quiet interval; the counter stays bounded by
the quiet sample's mark and late traffic shows up in the evidence.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(perf): fail the J2 quiet wait when a sample stalls past its deadline

waitForJourneyQuiet accepted a sample that returned unchanged after a
stall longer than the timeout as the end of a quiet period, before the
deadline check ran. The deadline is now checked first, so a stalled
sample fails the wait instead of letting the click go ahead unobserved.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(perf): detach J1's IPC capture before the J2 click

J2 used J1's capture to see pending launch bridge calls while settling,
but its ipcMain listener stayed attached and ran for every bridge call of
the measured click. The capture can now be detached; J2 detaches J1's
right after settling, before the click.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(perf): sample both J2 settle captures in one main-process snapshot

The settle sample read J1's capture (pending calls) and J2's capture
(call count) in two evaluate calls, so a call starting in between was
counted with a stale zero in flight and its completion went unseen. Both
states are now read in one synchronous pass, where no ipcMain event can
be handled in between.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: 4gray <fourgray@proton.me>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 20:37:45 +02:00
d60c80746b perf(services): share the startup inventory read (J1) (#1716)
* perf(workspace): share the startup inventory read and defer non-first-card IPC (J1)

Journey J1 / renderer.ipcCallsToFirstCard: 12 -> 7.

- PlaylistsService.getAllPlaylists() shares one in-flight SQLite read
  between concurrent callers (the playlist effect and the XMLTV source
  reconciliation at startup). Settled reads are never reused and every
  SQLite write detaches the pending read.
- getM3uFavoriteChannels() stops re-reading the write-once IndexedDB ->
  SQLite migration receipt once it has been seen.
- StartupDeferralService holds the download list, app update status and
  the dashboard's recent items and favorites until one task after the
  render that reveals the routed content (5 s safety timeout).
  IPTVNATOR_DISABLE_STARTUP_DEFERRAL=1 is the kill switch.
- reconcileEpgSources and the two distinct migration-flag reads stay on
  the critical path: the first is the #1548 revision fence, the second
  are different keys, not duplicates.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(services): trust a migration receipt written in this session; correct J1 note

- Mark the IndexedDB -> SQLite receipt as confirmed after an empty-store
  receipt write or a committed dbMigrateAppPlaylists, not only when it was
  already present, so M3U favorites skip the per-playlist re-read on the
  first launch after an upgrade too.
- The release note no longer claims a wall-clock speedup the J1 benchmark
  did not show.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* perf(services): keep only the shared startup inventory read (J1)

Drop the startup deferral gate, its kill switch and the deferred loads:
they lowered renderer.ipcCallsToFirstCard but moved neither
spawnToFirstCardMs nor load->card beyond drift, and they grew
renderer.initialBytes. Keep the shared in-flight inventory read, inline
it in PlaylistsService with short property names, and copy the joiner's
result with structuredClone (the Electron renderer has it; the services
test setup polyfills it for jsdom).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(performance): describe the shared inventory read and the J1 validation result

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(services): drop the migration receipt memo for M3U favorites

Skipping the receipt read let the dashboard ask the database for M3U
favorite channels before a just-toggled favorite was written: the write
waits in the per-playlist queue and the cross-context lock, and the
dashboard does not reload when the store's favorites are unchanged.
The extra round trip had been masking that race, and the Windows E2E
run hit it (epg.e2e.ts "dashboard live rails find a programme that only
another playlist's XMLTV carries"). The memo was off the first card's
path, so it bought nothing measurable for J1; restore the per-call read.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(services): share only the startup inventory read

A pending read was shared with any concurrent caller, but not every
playlist write goes through PlaylistsService: the settings reset deletes
all playlists through DatabaseService, so a caller could join a read
taken before that deletion (Codex review). Share only the first read,
which the startup pair (playlist effect and XMLTV reconciliation) needs
while the startup screen still hides every writing action; sharing ends
when it settles or a PlaylistsService write starts.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: 4gray <fourgray@proton.me>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 17:41:42 +02:00
4gray de7b19aee2 docs(performance): record the idle work audit (#1719) 2026-09-29 07:03:01 +02:00
4grayandClaude Opus 5.5 af44e2368b ci(performance): give the initial-bytes ratchet slack and a labelled override (#1744)
* ci(performance): give the initial-bytes ratchet slack and a labelled override

The exact renderer.initialBytes counter failed PRs for reasons outside
their diff: two concurrent merges left master 108 bytes over the baseline
for hours, and bundler identifier renaming moves the counter by hundreds of
bytes. PRs growing it by 243 and 302 bytes had no way to pass at all,
because the direction check refuses any raised baseline.

- Counter entries accept `slack` (integer, entry unit): the ratchet enforces
  `value + slack`, reports how much slack a measurement uses, and still
  prints the tighten hint below `value`. renderer.initialBytes gets 4096
  bytes, so growth can accumulate at most 4 KiB past the last lowered
  baseline while regressions such as +35 KB still fail.
- check-baseline-direction.mjs compares `value + slack`, treats widened
  slack like a widened tolerance, and takes `--allow-increase`, which
  reports weakened entries as ALLOWED instead of failing.
- CI passes `--allow-increase` only when the pull request (or, for a master
  push, the pull request merged as the pushed commit) carries the
  perf-baseline-increase label, read from the API so a job re-run picks up
  a label added later.

This change widens the slack itself, so its own PR needs the label.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0166PWobUjpoiWt8E9bBdsCs

* fix(performance): report slack usage only for entries that have slack

A wall-clock measurement above `value` but within `value × toleranceRatio`
fell into the slack branch and was reported as using "slack", conflating
timing tolerance with counter slack. Only entries with `slack` report it now.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0166PWobUjpoiWt8E9bBdsCs

* fix(performance): scope the push-time label and refuse raised counter values

- A master push compares the whole push, but the label was read from the
  head commit's PR only, so one labelled PR could cover another commit's
  increase in the same push. The label now counts only when the push added
  exactly one first-parent commit (a squash or merge of one PR); any other
  push that weakens a baseline fails.
- Raising a counter's `value` while narrowing its `slack` lowered the
  enforced limit and was reported as "lowered". A counter's value is the
  measured evidence and only moves down, so that raise is now a weakening
  that needs the label.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0166PWobUjpoiWt8E9bBdsCs

* fix(performance): treat a counter/wall-clock type switch as a weakening

Turning `{ value: 100, slack: 10 }` into `{ value: 105, toleranceRatio: 1 }`
skipped the raised-counter-value rule and was reported as a lowered limit.
Switching an entry between counter and wall-clock now needs the
perf-baseline-increase label.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0166PWobUjpoiWt8E9bBdsCs

* ci(performance): paginate the label lookups of the direction check

The labels endpoint returns 30 entries per page by default, so a PR with
more labels could miss perf-baseline-increase. Both lookups now request
100 per page and paginate, like the release-note gate.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0166PWobUjpoiWt8E9bBdsCs

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-28 14:55:15 +02:00
4gray 8b570e3390 test(performance): measure two-byte string cost in M3U and XMLTV parsing (D1) (#1713) 2026-09-28 00:03:18 +02:00
1d9a563d1a test(performance): count startup phases and SQL statements for the J1 launch journey (#1715)
* test(performance): count startup phases and SQL statements for the J1 launch journey

Implements plan item A2. With IPTVNATOR_PERF_CAPTURE=1 the main process
keeps named counters and registers a main-only performance:read-counters
IPC handler; without the flag nothing is counted and the handler does not
exist.

- debug-trace.ts owns the registry; traceStartupPhase replaces the
  trace('startup', ...) sites and counts main.startupPhases.
- The database worker counts executed statements through better-sqlite3's
  Statement prototype (the verbose callback expands every statement and
  made bulk inserts 2-4x slower) and posts the count over its message
  port, flushed before every other worker message. The main-thread shared
  connection is counted through a new connection observer in the shared
  database library.
- The first main window freezes main.modulesRegisteredBeforeWindow at
  creation and main.sqlStatementsBeforeReadyToShow at ready-to-show.
- The journey gate drops the ready-to-show that Electron emits for the
  about:blank detour, so the app sees the real document's first paint,
  and taps the counters handler; the J1 record reads both counters after
  the renderer probe completes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(database): require one SQL statement per exec during initialization

The performance capture counts one exec call as one statement, because
SQL cannot be split reliably in the counter (trigger bodies contain
semicolons). The historical-upgrade driver now wraps exec on every
connection initDatabase opens and fails on a batch, so that counting
assumption holds for the fresh profile and all historical schemas.
Documents the definition in the counter and the architecture docs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(performance): count SQL statements only for the launch journey

Codex review: the M3U import, refresh-cancellation and Xtream benchmarks
also run with IPTVNATOR_PERF_CAPTURE=1, so the statement hook wrapped
every row of their bulk inserts and changed what they measure.

SQL counting now also needs IPTVNATOR_PERF_COUNT_SQL=1, which only the
launch journey sets; a harness test fails if another source sets it.
Startup phases, the window snapshot and the read handler stay on the
capture flag. Without SQL counting no ready-to-show listener is attached,
so a zero is never reported for statements nobody counted.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: 4gray <fourgray@proton.me>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-27 21:01:26 +02:00
69056487bc ci(performance): run the performance journeys on the Linux runner (#1717)
* ci(performance): run the performance journeys on the Linux runner

Adds a warn-only performance-journeys job to ci.yml that builds the
electron-performance configuration, runs the journey benchmarks under
xvfb and uploads dist/performance/journeys/ as evidence. Pull requests
run it only when they touch journey-relevant paths.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(performance): document the journeys CI job and runner evidence

Describes the performance-journeys job and its path gate, and records why
the J1 runtime counters are not baselined yet: on the Linux runner the
launch journey is bimodal (13/576 vs 16/939 bridge calls/DOM mutations),
so the counters are not deterministic.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* ci(performance): run the journeys unless a PR changes only safe paths

The scope filter listed the paths that can move a journey, so a PR that
changed only a root build input (.nvmrc, nx.json, tsconfig.base.json)
skipped the measurement. List the paths that cannot instead: the E2E
workflow's ignore list plus release notes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: 4gray <fourgray@proton.me>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-27 20:53:26 +02:00
f0e51d2806 perf(web): keep lazy-only services and SafePipe out of main.js (#1729)
* perf(web): keep lazy-only services and SafePipe out of main.js

The eager shell imported barrels that re-export Angular injectables and a
pipe it never uses, and their static definitions keep those modules in
main.js: PlaylistFileImportService came with PlaylistContextFacade,
normalizeDateLocale with SafePipe, and the workspace-shell-util barrel with
SettingsContextService, which #1714 grew with match counts. That growth put
master 108 bytes over the renderer.initialBytes baseline #1712 had measured
on a branch without #1714.

Add file-level entries for the three modules and use them from the eager
and settings code: renderer.initialBytes 1,626,127 -> 1,619,993 bytes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore(performance): lower the initial-bytes baseline to 1,619,993 bytes

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: 4gray <fourgray@proton.me>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-27 14:21:40 +02:00
0f1cad4b16 test(performance): add the J1 launch-to-usable journey benchmark (#1698)
* test(performance): add the J1 launch-to-usable journey benchmark

Implements plan items A1, A3 (J1 only) and the minimal A4 from
.plans/2026-09-25-performance-journeys-ratchet.md.

- journey-renderer-probe.ts: init-script probe counting DOM mutations,
  layout shifts and long tasks until the first source card is visible on
  /workspace with the splash removed; unit-tested with jsdom fixtures.
- journey-main-ipc-capture.ts: counts bridge invocations from the preload's
  renderer-API trace channel up to a sentinel call the probe fires, so the
  IPC counter is exact without touching production code.
- launch.journey.ts + playwright.journeys.config.ts: seeded profile (one
  M3U source, one Xtream portal on the loopback mock), one warm-up and five
  measured iterations, fresh process and data directory each, writing
  dist/performance/journeys/<timestamp>/summary.json with exact counters
  and P50/P90 wall-clock.
- Nx target electron-backend-e2e:journeys and root script perf:journeys.
- docs/architecture/performance-journeys.md, README, context and
  validation map entries.

renderer.cdTicksToFirstCard and main.sqlStatementsBeforeReadyToShow are
reported as unavailable with the reason instead of being faked.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(performance): gate the renderer load so the probes never race startup

Review follow-up for #1698.

- journey-renderer-gate.cjs: a main-process hook loaded with `-r` (the
  mechanism Playwright uses for its own loader) makes the first
  loadFile/loadURL navigate to about:blank and holds the real load until
  the test releases it. Playwright reports no page before a navigation
  commits, so this is what lets the renderer probe be registered on the
  page before the real document exists; the IPC capture is installed
  before the release too. A safety timeout releases the gate on its own
  and marks the iteration invalid. Unit-tested with a fake BrowserWindow.
- launch-journey-app.ts: registers the probe on the parked page, releases
  the gate, waits for the real document to commit, fails fast when the
  probe is missing, and refuses an iteration whose gate timed out, saw a
  second load, or released before the probe was in place.
- journey-renderer-probe.ts: entries delivered live after the terminal
  batch are buffered and filtered by the same cutoff as queued ones, and
  the cutoff is sampled in a timer queued from the first rAF, i.e. after
  the card's frame is painted, so the render task's long task and layout
  shift are consistently included.
- launch-journey-record.ts: layoutShiftScore rounded to three decimals; a
  0.0001 shift flipped in and out of the cutoff between iterations.
- playwright.journeys.config.ts: reuse a mock server left on the journeys
  port locally (its fixtures are deterministic); CI still starts its own.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(performance): validate the journey gate after the probe completes

Codex follow-up on #1698: the gate state returned by release() cannot see a
reload or recovery navigation that happens before the first card. Re-read the
live state once the renderer probe has finished and validate that instead,
so an iteration spanning an extra navigation is rejected.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(performance): fail an iteration whose performance observers were unavailable

Codex follow-up on #1698: a renderer that cannot observe layout-shift or
longtask entries used to pass the probe with zero counters, which a ratchet
could not tell apart from a genuine zero. The probe assertion now rejects
such an iteration and names the missing observer.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: 4gray <fourgray@proton.me>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-26 19:18:19 +02:00
4grayandClaude Fable 5.1 512e9787a8 perf(web): load Angular date locales lazily per language (#1695)
Plan thread C1, journey **J1 `launch`**, counter **`renderer.initialBytes`**. Stacked on #1694 → #1693 → #1692; merge in order.

`apps/web/src/app/app-date-locales.ts` imported the locale data of all 18 supported languages eagerly, so every user shipped and parsed all of it at startup. Each locale is now a dynamic import keyed by the Angular locale id that `normalizeDateLocale()` derives from the app language (`by` → `be`, `ary` → `ar-MA`, `zhtw` → `zh-Hant`); English needs no data.

Ordering is preserved so no template renders a locale whose data has not arrived (Angular throws in that case):
- `main.ts` awaits the initial language's data (from `getInitialLanguage()`) before `bootstrapApplication`.
- Both `TranslateService.use()` call sites, `AppComponent.initSettings()` and `SettingsFormFacade.applySavedSettings()`, register the data first through the new `AppDateLocaleService`.
- A failed load never leaves the locale without data: English formatting is registered under the requested id (eager 1.1 KB `@angular/common/locales/en`), so `DatePipe` renders instead of throwing; the locale is not marked registered, so the next call retries and a success replaces the fallback (review follow-up).
- `AppDateLocaleService.use()` orders switches by request, not by completion: a switch whose data arrives after a newer request is dropped, so the language chosen last wins (review follow-up).

No kill switch: behavior is identical once the locale resolves, and the only new failure mode (a same-origin chunk failing to load) is shared with every lazy route.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-26 13:44:29 +02:00
4grayandClaude Fable 5.1 d3b6e548cc ci(performance): fail when the web app's initial bytes grow (#1694)
Third step of the performance-journeys ratchet, stacked on #1693 (which is stacked on #1692; merge in order, GitHub retargets each to `master`).

- New `Initial bytes ratchet` job in `.github/workflows/ci.yml` (ubuntu-latest): install, `pnpm nx build web --skip-nx-cache` (production configuration, the one users download), then `pnpm run perf:initial-bytes:check`. The job fails when `renderer.initialBytes` exceeds `tools/performance/journey-baselines.json`.
- `dist/performance/` is uploaded as the `performance-journey-summary` artifact on every run, so a failing or tightenable run carries its evidence.
- After review: the job first runs the new `tools/performance/check-baseline-direction.mjs`, which compares `journey-baselines.json` with the revision the change is measured against (the target branch of a pull request, `github.event.before` for a `master` push, `master` for a manual dispatch) and fails on any raised enforced limit (`value × toleranceRatio`), any widened or newly added tolerance, or any removed entry, so a PR cannot grow the payload and raise the baseline to match (lowered limits and new entries pass; a target branch without the file has nothing to weaken). Node tests cover it.
- Docs: the performance-journeys contract and the validation map name the job, and the contract now states that this runner is the canonical measurer (take baseline values from its output, not from a local build).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-26 13:44:01 +02:00
4grayandClaude Fable 5.1 7abaad29e7 chore(performance): commit the initial-bytes baseline and ratchet checker (#1693)
Second step of the performance-journeys ratchet, stacked on #1692 (merge that first; this PR retargets to `master` automatically).

- `tools/performance/journey-baselines.json`: J1 `launch` / `renderer.initialBytes` = **2,739,510 bytes**, the ubuntu runner's production build of `apps/web` at this content (after #1692 stopped bundling `package.json` into `main.js`). A local macOS build of this pre-#1695 code is 2 bytes smaller in `main.js` (the eager locale imports); once #1695 removes them the two are byte-identical. Correction to an earlier version of this description: the "556-byte macOS vs Linux difference" was almost entirely `package.json` text embedded in `main.js`, which moved with every script edit in this stack, plus this 2-byte residue. The runner is the canonical measurer; the CI run on the stacked #1694 branch (this content plus the job) is where the number is confirmed.
- `tools/performance/check-journey-ratchet.mjs` compares a journey summary with the baselines: a counter above its value fails (exact, no slack), wall-clock entries fail above `value × toleranceRatio`, a baseline without a measurement fails so dropping a measurement cannot disable the ratchet, values below baseline print a "tighten" hint, and measured counters without a baseline are noted only. After review: checking nothing (empty file, or `--only` naming a missing entry) fails; a counter is read only from `counters` and a wall-clock entry only from `wallClock`; the repeatable `--only <journey>/<counter>` flag scopes a check.
- Root scripts: `perf:initial-bytes:check` (measures into its own `dist/performance/initial-bytes.summary.json`, then checks `--only launch/renderer.initialBytes`) and `perf:ratchet:check` (full check); `perf:tools:test` runs both test files, as does `pnpm nx test performance-tools`.
- `docs/architecture/performance-journeys.md` gains the Ratchet section (file format, rules, "baselines only move down"); the validation map lists the check.

The CI job that runs the check on every PR is #1694; C1 (lazy Angular date locales, #1695) then lowers the baseline with the measured output as evidence.

Note: `ci.yml` only triggers on pull requests targeting `master`, so this stacked PR shows no Actions runs until #1692 merges. The evidence runs above were dispatched with `gh workflow run ci.yml --ref <branch>`.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-26 13:43:33 +02:00
4grayandClaude Fable 5.1 8bc877b625 chore(performance): measure initial bytes of the built web app (#1692)
First step of the performance-journeys ratchet (plan thread: J1 `launch`, counter `renderer.initialBytes`).

- `tools/performance/measure-initial-bytes.mjs` reads the built `dist/apps/web/index.html` and sums `index.html` plus every same-origin `<script src>`, `<link rel="stylesheet">` and `<link rel="modulepreload">` it references. Manifest, icons, external URLs and lazy chunks are not counted. A referenced file missing from the build fails the measurement instead of counting as zero bytes.
- `--json` prints the breakdown; `--summary <file>` writes the `journeys.<journey>.counters` shape a ratchet checker will consume (next PR).
- New Nx project `performance-tools` (test + lint targets), Tier B in `tools/coverage/coverage-policy.json`, root scripts `perf:initial-bytes` and `perf:tools:test`.
- New contract `docs/architecture/performance-journeys.md`, linked from the validation map, the agent context map and the README.
- **Review follow-ups:** resources are deduplicated by request URL (query kept, fragment dropped); `index.html` is parsed with parse5 (already a repository dependency, scripting enabled), so comments, bogus comments, raw-text bodies (script/style/noscript/title/textarea), inert `<template>` contents and character references in attributes all follow the HTML5 algorithm instead of a hand-written scanner; the review's edge cases stay as regression tests; docs show the `pnpm --silent` form for JSON output and explain how the counter relates to Angular's rounded "Initial total".
- **Found while measuring:** the environment files and the playback diagnostic panel imported the whole `package.json` (`import packageJson from '@package'`), which esbuild cannot tree-shake, so `main.js` carried the complete file and the counter moved with every script or dependency edit. They now import `{ version }` only (eb662c485): `main.js` shrinks by **11,539 bytes** and the counter no longer depends on `package.json`. Jest's ESM loader exposes JSON only as a default export, so the two web Jest configs map `@package` to a stub that serves the real file's fields as named exports. Release note: `.changes/web-version-only-from-package-json.md` (`type: perf`).

Production build after this PR measures **2,739,508 bytes** (11 files + `index.html`); Angular's "Initial total" is this minus `index.html` and `assets/app-config.js`. The baseline file and the CI check land in the follow-up PRs; C1 (lazy Angular date locales) then lowers it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-26 13:31:27 +02:00