From 69056487bc641ad404b6826b5e583f00de82b38a Mon Sep 17 00:00:00 2001 From: 4gray <4gray@users.noreply.github.com> Date: Sun, 27 Sep 2026 20:53:26 +0200 Subject: [PATCH] ci(performance): run the performance journeys on the Linux runner (#1717) * ci(performance): run the performance journeys on the Linux runner Adds a warn-only performance-journeys job to ci.yml that builds the electron-performance configuration, runs the journey benchmarks under xvfb and uploads dist/performance/journeys/ as evidence. Pull requests run it only when they touch journey-relevant paths. Co-Authored-By: Claude Opus 5.5 * docs(performance): document the journeys CI job and runner evidence Describes the performance-journeys job and its path gate, and records why the J1 runtime counters are not baselined yet: on the Linux runner the launch journey is bimodal (13/576 vs 16/939 bridge calls/DOM mutations), so the counters are not deterministic. Co-Authored-By: Claude Opus 5.5 * ci(performance): run the journeys unless a PR changes only safe paths The scope filter listed the paths that can move a journey, so a PR that changed only a root build input (.nvmrc, nx.json, tsconfig.base.json) skipped the measurement. List the paths that cannot instead: the E2E workflow's ignore list plus release notes. Co-Authored-By: Claude Opus 5.5 --------- Co-authored-by: 4gray Co-authored-by: Claude Opus 5.5 --- .github/workflows/ci.yml | 138 ++++++++++++++++++++++ docs/architecture/performance-journeys.md | 35 +++++- 2 files changed, 171 insertions(+), 2 deletions(-) diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index 33550a61b..ce983514b 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -205,6 +205,144 @@ jobs: path: dist/performance/ retention-days: 14 + # Decides whether a pull request can move the performance journeys, so + # the journeys job below does not spend a runner on docs-only changes. + # Pushes to master and manual dispatches always run them. + performance-journeys-scope: + name: Performance journeys scope + runs-on: ubuntu-latest + timeout-minutes: 5 + outputs: + run: ${{ steps.scope.outputs.run }} + + steps: + # The PR checkout is the merge commit: its first parent is the + # target branch, so two commits are enough to diff the PR. + - name: Checkout code + if: github.event_name == 'pull_request' + uses: actions/checkout@v7 + with: + fetch-depth: 2 + + # Anything can move a journey (application code, the harness, root + # build inputs such as .nvmrc, nx.json or tsconfig.base.json), so + # the filter lists what cannot: the same paths the E2E workflow + # ignores, plus release notes. + - name: Detect journey-relevant changes + id: scope + env: + EVENT_NAME: ${{ github.event_name }} + run: | + set -euo pipefail + if [ "$EVENT_NAME" != "pull_request" ]; then + echo "run=true" >> "$GITHUB_OUTPUT" + exit 0 + fi + relevant="$(git diff --name-only HEAD^1 HEAD | + grep -vE '\.md$|^docs/|^\.plans/|^\.codex/|^\.claude/|^\.changes/|^apps/website/' || true)" + if [ -n "$relevant" ]; then + echo "Journey-relevant changes:" + echo "$relevant" + echo "run=true" >> "$GITHUB_OUTPUT" + else + echo "No journey-relevant changes; skipping the performance journeys." + echo "run=false" >> "$GITHUB_OUTPUT" + fi + + # Runs the journey benchmarks (docs/architecture/performance-journeys.md) + # on the canonical Linux runner and uploads the summary as evidence. + performance-journeys: + name: Performance journeys + needs: performance-journeys-scope + if: needs.performance-journeys-scope.outputs.run == 'true' + runs-on: ubuntu-latest + # The electron-performance build is the bulk of the time; the launch + # journey itself is six fresh Electron processes plus one seeding run. + timeout-minutes: 30 + # Warn-only for the first two weeks of plan item B3: a failure is + # visible on the run but does not fail the workflow. + continue-on-error: true + env: + NX_SKIP_NX_CACHE: true + + steps: + - name: Checkout code + uses: actions/checkout@v7 + + - name: Install pnpm + uses: pnpm/action-setup@v6.0.10 + + - name: Setup Node.js + uses: actions/setup-node@v7 + with: + node-version-file: '.nvmrc' + cache: 'pnpm' + + - name: Install dependencies + run: pnpm install --frozen-lockfile + + # The journeys drive Electron through Playwright's _electron API + # and never launch a Playwright browser, so no `playwright + # install`. The runner image ships Electron's shared libraries and + # xvfb; fail fast with a clear message if an image update drops one. + - name: Check Electron runtime dependencies + run: | + command -v xvfb-run || { echo "::error::xvfb-run is missing on the runner"; exit 1; } + missing="$(ldd node_modules/electron/dist/electron | grep 'not found' || true)" + if [ -n "$missing" ]; then + echo "::error::Electron is missing shared libraries:" + echo "$missing" + exit 1 + fi + + # The Nx target builds electron-backend:build-performance first + # and playwright.journeys.config.ts starts the Xtream mock server. + - name: Run the performance journeys + run: xvfb-run --auto-servernum --server-args="-screen 0 1280x960x24" pnpm run perf:journeys + env: + CI: true + NX_TASKS_RUNNER_DYNAMIC_OUTPUT: false + + # Every run writes a fresh timestamped directory, so a clean + # checkout must hold exactly one summary. + - name: Locate the journey summary + id: summary + run: | + set -euo pipefail + mapfile -t summaries < <(find dist/performance/journeys -mindepth 2 -maxdepth 2 -name summary.json) + if [ "${#summaries[@]}" -ne 1 ]; then + echo "::error::Expected one journey summary, found ${#summaries[@]}" + exit 1 + fi + echo "path=${summaries[0]}" >> "$GITHUB_OUTPUT" + + - name: Report journey measurements + env: + SUMMARY: ${{ steps.summary.outputs.path }} + run: | + set -euo pipefail + jq -r ' + "## Performance journeys (\(.harness.platform), \(.harness.measuredIterations) measured iterations)", "", + (.journeys | to_entries[] | .key as $journey | .value as $j | + "### `\($journey)`", "", + "| Measurement | Value | Iterations |", + "| --- | ---: | --- |", + (($j.counters // {}) | to_entries[] | + ($j.counterStability[.key] // {}) as $s | + "| `\(.key)` | \(.value) | \(($s.values // []) | map(tostring) | join(", "))\(if $s.stable == false then " (unstable)" else "" end) |"), + (($j.wallClock // {}) | to_entries[] | "| `\(.key)` | \(.value) | |"), + "") + ' "$SUMMARY" | tee -a "$GITHUB_STEP_SUMMARY" + + - name: Upload journey summaries + if: always() + uses: actions/upload-artifact@v7 + with: + name: performance-journeys + path: dist/performance/journeys/ + if-no-files-found: warn + retention-days: 14 + unit-and-typecheck: name: Unit Tests and Typechecks runs-on: ubuntu-latest diff --git a/docs/architecture/performance-journeys.md b/docs/architecture/performance-journeys.md index e15aea3bd..7a737f876 100644 --- a/docs/architecture/performance-journeys.md +++ b/docs/architecture/performance-journeys.md @@ -169,8 +169,9 @@ harness, which is what the ratchet needs. The main process start `journeys..counters.` and `journeys..wallClock.` are plain numbers so `tools/performance/check-journey-ratchet.mjs` can compare them with -`tools/performance/journey-baselines.json`. A J1 baseline is added once the -numbers are stable on the CI runner; until then the summary is evidence only. +`tools/performance/journey-baselines.json`. A J1 runtime baseline is added +once its counter is deterministic on the CI runner; the launch counters are +not yet (see [Ratchet](#ratchet)), so the summary is evidence only. ## `renderer.initialBytes` @@ -298,6 +299,36 @@ earned it, set `updatedAt` and `evidencePr`, and paste the measurement output into the PR. Never raise a value to make a PR pass: if growth is a deliberate trade-off, say so in the PR and let the maintainer decide. +The runtime counters come from the `Performance journeys` job of the same +workflow, on `ubuntu-latest` only. It runs `pnpm run perf:journeys` under +`xvfb-run` (the Nx target builds `electron-backend:build-performance`, the +Playwright config starts the Xtream mock), writes the measurements to the job +summary and uploads `dist/performance/journeys/` as the `performance-journeys` +artifact. The `Performance journeys scope` job skips it only for pull +requests that change nothing but Markdown, `docs/**`, `.plans/**`, +`.codex/**`, `.claude/**`, `.changes/**` or `apps/website/**` (the E2E +workflow's ignore list plus release notes); any other file, including root +build inputs such as `.nvmrc`, `nx.json` or `tsconfig.base.json`, runs it. +Pushes to `master` and manual dispatches always run it. The job is warn-only (`continue-on-error: true`) for its first two +weeks (plan item B3): a regression marks the job failed without failing the +workflow. Making it required is a maintainer decision. + +No J1 runtime counter is enforced yet. Three dispatched runs on 2026-09-27 +(CI runs 36271875209, 36271879955 and 36271884616) reported the same summary +values, `renderer.ipcCallsToFirstCard` 16 and +`renderer.domMutationsToFirstCard` 939, but the third run marked both +`stable: false`: its warm-up and one measured iteration reached the first +card in about 750 ms with 13 bridge calls and 576 mutations, the others in +about 1,400 ms with 16 and 939. The three extra calls +(`downloadsGetDefaultFolder` and two `dbGetGlobalRecentlyAdded`) land before +or after the first card depending on that race, so neither counter is +promoted until the race is understood and the counters are deterministic. +`renderer.layoutShiftScore` (0) and `renderer.longTasks` (2) were identical +in all eighteen runner iterations; the `spawnToFirstCardMs` P50 ranged from +1,401 to 1,674 ms. All four stay evidence for now. Runner counters also +differ from a Mac (12 and 571 there, the fast path without the Linux-only +`getWindowState` call), so take J1 baseline values from the runner only. + ## Adding a counter 1. Produce the value from the built output or from a deterministic probe, not