* fix(coverage): report retried-then-passing e2e tests as flaky
The semantic summary flattened every Playwright attempt of a spec and
checked for `failed` first, so a test that failed and then passed on a
retry was reported as `failed` and its critical journey as `failing`,
although Playwright counts it as flaky with zero unexpected results.
The `flaky` branch was unreachable.
Derive the status from the final attempt: only a failed or timed-out
final attempt is `failed`; a pass after earlier failures is `flaky`.
Skipped handling is unchanged. A journey with flaky tests is therefore
`covered`; the Statuses line already lists the flaky count.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix(coverage): keep non-passing retries failed and surface flaky over skipped
Review feedback on the final-attempt status: a failure followed by a
skipped or interrupted retry returned the final status and dropped the
failure, so the journey read as covered. Only a final pass now turns
earlier failures into flaky; any other ending after a failure stays
failed.
A spec runs once per Playwright project, and `skipped` outranked
`flaky`, so a skip in one browser hid a retried-then-passing test in
another. Flaky now outranks skipped.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: 4gray <fourgray@proton.me>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Run the sequential Electron E2E suite as three Playwright shards per OS
(one runner each) and summarize all shards of an OS in one follow-up job.
The semantic summary script accepts a directory of shard reports, merges
them and refuses to write when a shard is missing, duplicated or
malformed, or when an explicit input does not exist.
Slowest shard per OS in the final run: ubuntu 12.5 min (was 26),
macOS 13.7 min (was 34), Windows 24.5 min (was 35).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>