Port the embedded mpv frame-copy pipeline to Windows with WGL rendering and named shared memory. Includes packaging validation, platform gates, tests, and architecture documentation.
10 KiB
Frame-copy spike — measurement log
One section per machine. Append new runs here; do not overwrite old ones —
this file is the cross-hardware comparison baseline for the go/no-go
decision (gates in README.md and
.plans/2026-07-10-embedded-mpv-frame-copy-unification.md).
How to reproduce a row:
# synthetic (no decode load):
./run.sh 'av://lavfi:testsrc2=size=3840x2160:rate=60' 3840x2160
# real 4K60 HEVC with hw decode (generate once):
ffmpeg -y -f lavfi -i "testsrc2=size=3840x2160:rate=60" -t 12 \
-c:v hevc_videotoolbox -b:v 25M -tag:v hvc1 -pix_fmt yuv420p /tmp/spike-4k-hevc.mp4
HELPER_ARGS='--hwdec videotoolbox --no-audio --loop' ./run.sh /tmp/spike-4k-hevc.mp4 3840x2160
Read STATS lines from viewer stdout after they stabilize (skip the first
window); helper stats are on its stderr. CPU: ps -o %cpu -p <pid> for the
helper and the Electron Helper (Renderer) process during playback.
Column meanings: new fps — frames actually reaching the canvas; copy —
shm→ArrayBuffer memcpy in the addon; upload — texSubImage2D wall time;
age — produce→uploaded latency (helper memcpy done → texture updated,
same monotonic clock on both sides — CLOCK_MONOTONIC_RAW in the original
spike harness used for the M1 rows; the production engine uses
CLOCK_MONOTONIC via frame_shm_now_ns() since the Linux port).
MacBook Pro M1 Pro (arm64), macOS, 120 Hz internal display — 2026-07-10
Source: commit 7e39d2e5, libmpv 2.3.0 (Homebrew mpv 0.39.0), Electron 41.7.2.
| Scenario | Producer fps | New fps | copy ms avg/p95 | upload ms avg/p95 | age ms avg/p95 | torn | CPU helper / renderer |
|---|---|---|---|---|---|---|---|
| 1080p60 testsrc2, sw decode | 60.0 | 59.8 | 0.28 / 0.35 | 0.29 / 0.40 | 4.2 / 6.8 | 0 | — |
| 4K60 testsrc2, sw decode | 60.0 | 60.0 | 1.17 / 1.35 | 3.8 / 4.5 | 11.0 / 12.7 | 0 | — |
| 4K60 HEVC 25 Mbit, hwdec=videotoolbox | 60.0 | 59.9 | 1.2 / 1.6 | 3.3 / 4.1 | 9.8 / 11.7 | 0 | ~18 % / ~24 % |
Helper-side PBO map+copy at 4K: 1.0–1.8 ms avg. Only fps dip observed was at
the --loop file restart (decoder reinit), not in the copy path.
Pacing / judder and HDR (2026-07-10, same machine, viewer with interval instrumentation)
Metrics: present iv sd — stddev of intervals between texture uploads (viewer clock, rAF-quantized); src iv sd — stddev of intervals between produced frames (helper clock); late1.5x — present intervals > 1.5× the producer's median interval (a missed beat).
| Scenario | New fps | present iv sd | src iv sd | late (steady state) | Notes |
|---|---|---|---|---|---|
| 4K60 HEVC hwdec (steady state) | 60.0 | 0.5–1.4 ms | 0.7–1.2 ms | 0–1 per 2 s | worst intervals only at --loop restart |
| 1080p50 testsrc2 (50→120 Hz cadence) | 50.0 | ~4.1 ms | ~1.9 ms | 0; LONGRUN 30 s: 0.13 % (startup only) | sd is 120 Hz rAF grid quantization (16.7/25 ms alternation around 20 ms), not lost frames |
| 1080p25 testsrc2 | 25.0 | ~4.1 ms | ~2.7 ms | 0 | same grid effect |
| 4K25 HDR10 PQ/BT.2020 HEVC hwdec | 25.0 | ~5.5 ms | ~4.3 ms | 0 | mpv tonemaps to SDR before readback (PIXELPROBE spread 255); copy/upload costs unchanged |
Viewport-size scaling check (2026-07-10)
Design claim "render at viewport size, pay viewport price" confirmed: the same 4K60 HEVC clip rendered into a 1280×720 FBO costs helper map+copy 0.17 ms (vs 1.5 ms at 4K), viewer copy 0.16 ms (vs 1.2), upload 0.17 ms (vs 3.5), steady 60 fps. Full 4K price is only paid in 4K-sized viewports (i.e. fullscreen on a 4K display).
10-minute long run — 4K60 HEVC hwdec, --loop (2026-07-10)
Final cumulative line at t=574 s: 33 501 frames, avg 58.32 fps, late1.5x 0.475 %, late2.5x 0.20 %, worst interval 2306 ms, dropped 261, torn 0.
Reading the anomalies before quoting the headline numbers:
- All 261 dropped frames and the single 2.3 s worst-interval stall happened in the first ~60 s (startup/warmup); from t=93 s to the end — zero drops over ~8.5 minutes.
- The steady-state late frames (~102 × late1.5x, ~43 × late2.5x after
warmup ≈ 0.36 % / 0.15 %) track the
--looprestarts of the 12 s test clip (~48 restarts, each a decoder reinit hiccup) — a test-clip artifact, not a pipeline property. A real long-form stream should be cleaner. - torn=0 across the whole run: the seqlock ring never produced a torn read.
Cadence verdict on this hardware: the producer keeps a clean source cadence (mpv's own pacing survives); the only jitter is display-grid quantization in the rAF presenter, bounded by one 120 Hz tick (~8.3 ms). A future refinement could reduce it with display-rate matching, but nothing here blocks the gate.
HDR clip generation for repro:
ffmpeg -y -f lavfi -i "testsrc2=size=3840x2160:rate=25" -t 10 \
-c:v hevc_videotoolbox -profile:v main10 -pix_fmt p010le -b:v 30M -tag:v hvc1 /tmp/spike-4k-hdr.mp4
ffmpeg -y -i /tmp/spike-4k-hdr.mp4 -c:v copy \
-bsf:v "hevc_metadata=colour_primaries=9:transfer_characteristics=16:matrix_coefficients=9" \
/tmp/spike-4k-hdr10.mp4 # videotoolbox omits VUI color tags; inject HDR10 ones
Reference worst-case budget from the analysis doc, for orientation: 33 MB 4K memcpy est. 6–8 ms (measured ≈1.2 ms); end-to-end added latency est. 40–60 ms (measured ≈10 ms produce→uploaded, before compositor).
Intel Mac — SKIPPED by decision (2026-07-10)
Owner decision: the frame-copy engine targets Apple Silicon (M1+) only on macOS; support detection must gate on arm64. Rationale: Intel Macs able to run the app at all (Electron 41 ⇒ macOS 10.15+) are a shrinking 2015–2020 cohort and keep the existing docked/external/web player paths; the only Intel device sourced (iMac mid-2011, High Sierra) cannot run the app and would not have been representative. The macOS hardware gate is therefore closed by the M1 Pro numbers above; the remaining risk hardware is Windows/Linux.
Linux mid-range laptop (iGPU) — Ubuntu 25.04, i7-1165G7 / Iris Xe, x64 — 2026-07-11
Source: Linux port branch (headless-EGL frame_helper_gl.h backend), system
libmpv 2.5.0 (mpv 0.40), Mesa 25.0 iris. Measured with the production helper
binary + embedded_mpv_frame_reader.node in a Node probe loop
(linux-frame-probe.mjs in this directory — the spike viewer harness is
macOS-only), so age here is produce→reader-copy and excludes the renderer
texture upload. hwdec was NOT active — this machine
has no VAAPI driver installed (intel-media-va-driver), so HEVC rows are
software decode; treat them as a decode-limited floor, not a pipeline
ceiling.
| Scenario | New fps | copy ms avg/p95 | age ms avg/p95 | torn |
|---|---|---|---|---|
| 1080p60 testsrc2, sw | 60.1 | 1.16 / 1.37 | 2.26 / 3.21 | 0 |
| 4K60 testsrc2, sw | 50.0 | 5.75 / 6.92 | 7.22 / 8.00 | 0 |
| 4K60 HEVC 25 Mbit, sw decode | 39.9 | 7.62 / 14.6 | 9.06 / 17.0 | 0 |
| 4K60 HEVC 25 Mbit in a 1280×720 viewport | 53.1 | 0.94 / 2.57 | 2.42 / 5.28 | 0 |
Readings:
- 1080p60 — the realistic viewport class for this laptop's 1920×1200 screen — holds a clean 60 fps with ~1 ms copies.
- The 4K rows are stress rows: producers are limited by software decode/source generation on 4 cores, not by the copy path (copy stays well under one 60 Hz frame budget even at full 4K).
- The viewport-size claim reproduces on Linux: the same 4K60 HEVC clip in a 720p viewport drops the copy from 7.6 ms to 0.94 ms and lifts fps from ~40 to ~53 (remaining gap = software decode).
- torn=0 across every run; the aspect-fit generation bump was verified
separately (4:3 source in a 16:9 viewport →
-g2at 960×720). - EGL display tier used: Mesa surfaceless platform (first tier; no display server needed).
Windows mid-range laptop (iGPU) — Windows 11 Home 26200, i7-1165G7 / Iris Xe, x64 — 2026-07-12
Source: Windows port branch (WGL frame_helper_gl.h backend), vendored
libmpv from zhongfly/mpv-winbuild 2026-06-14 (git-7d245fd100,
libmpv-2.dll), Intel driver 30.0.101.1340. Same physical laptop as the
Linux section above (TUXEDO Book XP14 Gen12, dual boot), so the two
sections compare OS/driver stacks on identical hardware. Measured with the
production helper + embedded_mpv_frame_reader.node through
linux-frame-probe.mjs (now cross-platform; on Windows it polls via
setImmediate because setTimeout quantizes to the ~15.6 ms system timer,
which would dominate age), so age is produce→reader-copy and excludes
the renderer texture upload — same semantics as the Linux rows. Unlike the
Linux rows, hwdec IS available here: mpv's hwdec=auto engages d3d11va
(verified by helper CPU: 0.51 core-s/s vs 2.82 core-s/s with --hwdec no
on the 4K row — a 5.5× CPU drop). Test clip generated with hevc_qsv
(the Intel encoder), so decode complexity is not byte-identical to the
videotoolbox/x265 clips of the other sections.
| Scenario | New fps | copy ms avg/p95 | age ms avg/p95 | torn |
|---|---|---|---|---|
| 1080p60 testsrc2, sw | 60.0 | 1.80 / 2.10 | 1.83 / 2.14 | 0 |
| 4K60 testsrc2, sw | 56.0 | 8.81 / 9.88 | 8.83 / 9.86 | 0 |
| 4K60 HEVC 25 Mbit, hwdec=d3d11va | 41.1 | 8.57 / 10.4 | 8.71 / 10.4 | 0 |
| 4K60 HEVC 25 Mbit, sw decode | 45.7 | 10.3 / 15.1 | 10.4 / 15.4 | 0 |
| 4K60 HEVC 25 Mbit in a 1280×720 viewport, hwdec | 60.2 | 1.29 / 1.87 | 1.27 / 1.82 | 0 |
Readings:
- 1080p60 — the realistic viewport class for this laptop's 1920×1200 screen — holds a clean 60 fps with ~1.8 ms copies. A 60-second sustained run kept 60.0 fps over 3601 frames (copy 1.57 / 1.84 ms, torn 0).
- The 4K rows are stress rows, as on Linux: at a full-4K viewport the Iris Xe is saturated by mpv render + readback (56 fps ceiling with no decode at all), so d3d11va decode — which shares the same iGPU — buys CPU headroom (5.5×), not fps; sw decode trades ~2.3 cores for ~4 fps.
- The viewport-size claim reproduces on Windows: the same 4K60 HEVC clip in a 720p viewport runs 60 fps with 1.3 ms copies and hardware decode active.
- torn=0 across every run; the aspect-fit generation bump was verified
(4:3 960×720 source in a 1280×720 viewport →
-g2at 960×720). - WGL context:
gl renderer: Intel(R) Iris(R) Xe Graphics(hardware, 3.2 core via wglCreateContextAttribsARB). Machine prerequisite hit during bring-up: Windows 11 Smart App Control blocks locally-built unsigned executables (the helper) until turned off, and a fresh Windows install may run the iGPU on the Basic Display Adapter — the frame-copy engine needs the real Intel driver bound (WGL on the basic adapter has no 3.2 core context, and there is no d3d11va).