Make E2E Tests Reliable on Main · Entire
Bottom line
It's not random flakiness — the same 4 agents fail nearly every run for distinct, diagnosable reasons, and because they share one matrix with the reliable agents, main goes red on essentially every merge. The Slack alert is now pure noise, so a real regression in claude-code would be invisible. Fix the signal architecture first, then the individual tests.
What the data actually shows (last ~20 runs on main)
Every run = failure or cancelled. Per-job, it's consistent:
- Always green: claude-code, opencode, gemini-cli, roger-roger, e2e-windows
- Always red: cursor-cli, copilot-cli, codex, factoryai-droid
And the red ones fail for different reasons — this matters:
| Agent | Symptom | Nature |
|---|---|---|
| cursor-cli | Error: Increase limits for faster responses (many) |
Account quota/rate limit — not a CLI bug |
| factoryai-droid | 14 failures / 7 tests | Broadly broken integration — real |
| codex | 2 failures (TestCodexResumeRestoredSession...) |
One specific test |
| copilot-cli | 2 failures (shadow branch ... should be empty) |
One specific test |
cancelled runs |
concurrency: cancel-in-progress: true + rapid merges |
Structural noise |
So: a long tail of specific failing tests + one quota problem + one genuinely-broken agent + concurrency cancellations. There is zero retry anywhere (gotestsum ... -count=1, no --rerun-fails).
Fixes, ranked by leverage
Split the matrix into tiers (biggest win). Reliable agents (claude-code, opencode, gemini-cli, roger-roger/vogon canary) stay as the gating + Slack-alerting set. Move cursor/copilot/codex/factory to
continue-on-error: trueor a separate non-paging workflow until fixed. This instantly makes green achievable and the alert meaningful again — without losing any coverage.Add automatic reruns:
gotestsum --rerun-fails=2 --packages=./e2e/testsinmise-tasks/test/e2e/_default. Kills the transient single-test tail (codex/copilot). Won't fix cursor's quota or factory's 14 — those are real.cursor-cli: "Increase limits" is an account/quota ceiling, not code. Add a
cursor-clientry toE2E_CONCURRENT_TEST_LIMIT(the workflow already special-cases gemini=6, factory=1) and/or raise the plan. No amount of retry fixes a quota wall.factoryai-droid: 14 failures is a real broken integration — quarantine it out of the gating set with a tracking issue rather than letting it page every merge.
Quarantine the specific chronic tests (codex resume, copilot shadow-branch) with
t.Skip("ENT-XXX: ...")+ issues, using the existing/e2e:triage-ciskill against artifacts to confirm flaky-vs-real first.Concurrency: post-merge
cancel-in-progress: trueturns every rapid merge into acancelledrun. Switch to a queue (cancel-in-progress: false) so you stop reading cancellations as failures.
#1 + #2 alone would make main green on most merges and restore a trustworthy alert.
Want me to implement #1 and #2 (tiered matrix + rerun-fails)? Those are low-risk and high-leverage. I'd leave #3–#6 as follow-up issues since they need per-test triage and an account/quota decision that's yours to make.
What changed
1. Tiered Slack alerts (.github/workflows/e2e.yml) — the notify-slack job now classifies failed jobs:
- RED (
#d50200,:red_circle:"E2E Tests Failed"): any of claude-code, opencode, gemini-cli, roger-roger, or e2e-windows fails → real regression. - YELLOW (
#daa038,:large_yellow_circle:"flaky-agent failures"): only cursor-cli/copilot-cli/codex/factoryai-droid fail. - A mixed failure is RED but lists both tiers in the body, so you never lose sight of which flaky agents also broke. Switched from plain
blockstoattachmentsso the colored bar actually renders. The job still fires on any failure — no visibility regression.
2. Automatic reruns (mise-tasks/test/e2e/_default) — gotestsum --rerun-fails=1 --packages=./e2e/tests. One retry of only the failed tests; a test that passes on retry reports green. Required moving the package from a positional arg to --packages. The default --rerun-fails-max-failures=10 means a broadly-broken run (e.g. factoryai-droid's 14) is reported as-is, not masked. Windows inherits this (same task). Canary/roger-roger untouched — deterministic agents stay strict so real bugs aren't hidden.
3. cursor-cli throttle (.github/workflows/e2e.yml) — added cursor-cli → 2 to E2E_CONCURRENT_TEST_LIMIT (its failures were Increase limits for faster responses = account rate-limit). Picked 2 as a conservative throttle alongside gemini=6, factory=1; bump it if cursor still rate-limits.
Notes
- Not committed (per your workflow).
- One residual noise source I left alone:
concurrency: cancel-in-progress: trueturns rapid back-to-back merges intocancelledruns. Not a failure, but it's why you also see grey. Say the word if you want it switched to a queue. - cursor's quota and factoryai-droid's 14 failures are real and won't be fixed by reruns — they're now correctly YELLOW instead of reddening main. Worth separate triage tickets.
Bottom line: push is blocked — the GitHub App token behind the entire:// remote lacks workflows permission, so it refuses any change to .github/workflows/. This isn't something I can work around from here.
The commits are ready locally. You'll need to push with your own credentials (which have workflows scope). Run this in the session:
! git push -u origin magical-yawning-matsumoto
Once it lands, I'll create the PR. Or if you'd rather, push + open the PR yourself in one go and I'll skip it.