fix(perf): finalize exact worker capture safely

This commit is contained in:
4gray committed 2026-07-27 02:15:59 +02:00
1 parent d3fdbaa2ef
commit 554eecfee0
23 files changed
+2685 -181

No files matched your search

+53 -2
View File
@@ -137,8 +137,59 @@ still produces an outcome with its request identity, nullable metrics, and a
fixed capture-unavailable reason. A worker terminated before it can flush the
capture reports the metric as unavailable rather than zero.
The harness excludes workers created by pre-capture seed setup through an exact
capture-generation marker. Since the production M3U database payload
Database-worker post-GC heap is not inferred from a heap snapshot. The harness
selects exactly one current-generation database worker independently of its
sampling timer, waits for any in-flight sample plus one final sample, stops the
worker CPU profile, sends an explicit-GC one-shot probe over a transferred
`MessagePort`, and takes an optional diagnostic snapshot only afterward.
Capture stop first closes a synchronous request cutoff and stops the main CPU
profile, so worker profile writes and heap snapshots cannot appear as
application stacks in `main.cpuprofile`. A database request observed after that
cutoff cannot restart worker sampling or profiling; it records
`database-worker-activity-after-cutoff` and invalidates the post-GC value.
Worker artifacts include a stable isolate ordinal in their filename so an
invalid multiple-worker capture cannot overwrite another isolate's profile.
The raw heap and unavailability reason form a strict XOR; playlist workers
terminated during cancellation report `worker-force-terminated-before-gc`
instead of a snapshot-derived heap. Warm-up and measured cancellation invoke
the real worker termination before best-effort finalization so profiling cannot
delay headline acknowledgement metrics. The separate diagnostic iteration
stops its CPU profile before termination; its timings are excluded from
headline distributions. That pre-termination drain is bounded; timeout
invalidates the profile and worker termination still proceeds. Late
termination completions are generation-gated so they cannot mutate a reused
record or append timeline events after an atomic capture rollover. The
benchmark fails closed unless the diagnostic iteration observed cancellation,
its profile is inside the iteration directory and parses with non-empty nodes
and samples, and the dying worker has no heap snapshot.
Headline worker memory distributions include every matching isolate and
exclude only nullable unavailable values. This aggregation is not used as a
validity shortcut. Runs that reach database persistence must independently
contain exactly one database worker with a coherent numeric post-GC heap.
Parsing cancellation happens before the renderer dispatches persistence, so a
run with no database request is explicitly `N/A:
operation-cancelled-before-database-phase`, not a missing capture. Any database
activity in a run that observed the cancellation effect is unrelated
contamination and invalidates the run. Duplicate, busy, late, timed-out, or
malformed captures are listed under
`summary.validity.databaseWorkerPostGc` and invalidate the benchmark. The
benchmark writes `manifest.json` and `summary.json` first, then fails the formal
run so the raw evidence remains available without supporting a before/after
claim. A formal cancellation run also requires the cancellation effect in every
measured iteration. Database-worker values are comparable only when every
measured iteration reaches that phase and has one valid capture.
Seed setup runs inside its own capture generation. The harness requires an
observed seed database request, a completed playlist upsert, and no pending
request, then persists that generation as `seed-main-capture.json`. Main
capture finalization validates one idle database worker with a coherent
explicit-GC result and atomically rolls the clean cutoff into the measured
generation before yielding. A request before rollover remains late and rejects
the transition; a request after rollover belongs to the measured generation.
This prevents seed writes or an unobserved stop/start gap from crossing into
measured main RSS, worker peaks, or profiles; generation selection also
excludes the seed worker record. Since the production M3U database payload
deliberately omits the refresh operation ID, the benchmark correlates only a
valid preload `start` marker to the exact database operation and playlist. A
missing or ambiguous marker leaves the raw operation ID `null` with a fixed
+36 -1
View File
@@ -147,6 +147,40 @@ profiling queue. If captures overlap, every overlapping response carries
CPU, ELU, and event-loop-delay values are `null`. This avoids assigning shared
worker activity to one request while preserving normal worker concurrency.
### Opt-in post-GC heap capture
The same `IPTVNATOR_PERF_WORKER_PROFILING=1` opt-in enables a development/test
one-shot `performance:collect-post-gc-heap` request. The benchmark transfers a
dedicated `MessagePort`; the database worker accepts the request only while no
request performance capture is active, calls exposed `globalThis.gc()`, reads
its own V8 isolate through `v8.getHeapStatistics()`, posts one result, and
closes the port. Production launches do not expose GC or send this request.
The response is a strict XOR: either a non-negative
`postGcHeapUsedBytes` with a `null` reason, or a `null` heap with a fixed
unavailability reason. Disabled profiling, a busy worker, unavailable GC, and
capture failure all fail closed without changing the database operation or
terminating the worker. The performance launcher supplies
`--js-flags=--expose-gc` to Electron itself; worker `execArgv` remains
untouched.
The main-process benchmark selects exactly one current-generation database
worker, waits for its final sampling call, stops its CPU profile, performs the
explicit-GC probe, and only then takes an optional diagnostic heap snapshot.
Profile or snapshot failure cannot overwrite an already captured post-GC
value. Capture stop closes a synchronous cutoff before awaiting profile or
snapshot work. A database request after that cutoff cannot restart sampling;
it records `database-worker-activity-after-cutoff` and invalidates the result.
Missing, multiple, busy, timed-out, or malformed worker captures remain raw
nullable outcomes and make a database-applicable measured run invalid for
comparison. A scenario cancelled before its database phase reports the worker
metric as not applicable rather than manufacturing an idle database request.
Before a measured generation, the M3U benchmark persists and validates its seed
capture, then atomically rolls a clean stopping cutoff into the next active
generation. This closes the otherwise unobservable gap between separate
stop/start calls: pre-rollover requests remain late and fatal, while
post-rollover requests are attributed to the new generation.
### Progress event contract
The worker now emits request-scoped events with:
@@ -564,7 +598,8 @@ executed inside a synchronous transaction callback:
```ts
// favorites is playlist-scoped: filter by (contentId, playlistId), otherwise
// a same-contentId favorite in another playlist gets rewritten too.
const stmt = db.update(schema.favorites)
const stmt = db
.update(schema.favorites)
.set({ position: sql<number>`${sql.placeholder('position')}` })
.where(
and(