Make E2E Tests Reliable on Main · Entire

Bottom line

It's not random flakiness — the same 4 agents fail nearly every run for distinct, diagnosable reasons, and because they share one matrix with the reliable agents, main goes red on essentially every merge. The Slack alert is now pure noise, so a real regression in claude-code would be invisible. Fix the signal architecture first, then the individual tests.

What the data actually shows (last ~20 runs on main)

Every run = failure or cancelled. Per-job, it's consistent:

And the red ones fail for different reasons — this matters:

Agent Symptom Nature
cursor-cli Error: Increase limits for faster responses (many) Account quota/rate limit — not a CLI bug
factoryai-droid 14 failures / 7 tests Broadly broken integration — real
codex 2 failures (TestCodexResumeRestoredSession...) One specific test
copilot-cli 2 failures (shadow branch ... should be empty) One specific test
cancelled runs concurrency: cancel-in-progress: true + rapid merges Structural noise

So: a long tail of specific failing tests + one quota problem + one genuinely-broken agent + concurrency cancellations. There is zero retry anywhere (gotestsum ... -count=1, no --rerun-fails).

Fixes, ranked by leverage

  1. Split the matrix into tiers (biggest win). Reliable agents (claude-code, opencode, gemini-cli, roger-roger/vogon canary) stay as the gating + Slack-alerting set. Move cursor/copilot/codex/factory to continue-on-error: true or a separate non-paging workflow until fixed. This instantly makes green achievable and the alert meaningful again — without losing any coverage.

  2. Add automatic reruns: gotestsum --rerun-fails=2 --packages=./e2e/tests in mise-tasks/test/e2e/_default. Kills the transient single-test tail (codex/copilot). Won't fix cursor's quota or factory's 14 — those are real.

  3. cursor-cli: "Increase limits" is an account/quota ceiling, not code. Add a cursor-cli entry to E2E_CONCURRENT_TEST_LIMIT (the workflow already special-cases gemini=6, factory=1) and/or raise the plan. No amount of retry fixes a quota wall.

  4. factoryai-droid: 14 failures is a real broken integration — quarantine it out of the gating set with a tracking issue rather than letting it page every merge.

  5. Quarantine the specific chronic tests (codex resume, copilot shadow-branch) with t.Skip("ENT-XXX: ...") + issues, using the existing /e2e:triage-ci skill against artifacts to confirm flaky-vs-real first.

  6. Concurrency: post-merge cancel-in-progress: true turns every rapid merge into a cancelled run. Switch to a queue (cancel-in-progress: false) so you stop reading cancellations as failures.

#1 + #2 alone would make main green on most merges and restore a trustworthy alert.

Want me to implement #1 and #2 (tiered matrix + rerun-fails)? Those are low-risk and high-leverage. I'd leave #3–#6 as follow-up issues since they need per-test triage and an account/quota decision that's yours to make.

What changed

1. Tiered Slack alerts (.github/workflows/e2e.yml) — the notify-slack job now classifies failed jobs:

2. Automatic reruns (mise-tasks/test/e2e/_default) — gotestsum --rerun-fails=1 --packages=./e2e/tests. One retry of only the failed tests; a test that passes on retry reports green. Required moving the package from a positional arg to --packages. The default --rerun-fails-max-failures=10 means a broadly-broken run (e.g. factoryai-droid's 14) is reported as-is, not masked. Windows inherits this (same task). Canary/roger-roger untouched — deterministic agents stay strict so real bugs aren't hidden.

3. cursor-cli throttle (.github/workflows/e2e.yml) — added cursor-cli → 2 to E2E_CONCURRENT_TEST_LIMIT (its failures were Increase limits for faster responses = account rate-limit). Picked 2 as a conservative throttle alongside gemini=6, factory=1; bump it if cursor still rate-limits.

Notes

Bottom line: push is blocked — the GitHub App token behind the entire:// remote lacks workflows permission, so it refuses any change to .github/workflows/. This isn't something I can work around from here.

The commits are ready locally. You'll need to push with your own credentials (which have workflows scope). Run this in the session:

! git push -u origin magical-yawning-matsumoto

Once it lands, I'll create the PR. Or if you'd rather, push + open the PR yourself in one go and I'll skip it.