Debug CI Failure and Fix Dependency Race · Entire
CI fail = data race in betterleaks v1.6.1
Dependabot bumped github.com/betterleaks/betterleaks 1.5.0 → 1.6.1. That release swapped its expr engine (cel-go → expr-lang/expr) and new exprruntime not goroutine-safe.
Race trace:
WARNING: DATA RACE
Write at 0x00c0001a89d8 by goroutine 302:
betterleaks@v1.6.1/internal/exprruntime/runtime.go:208
betterleaks@v1.6.1/detect/detect.go:967
redact/redact.go:207 ← our Detect() call
Previous write ... by goroutine 300 (same line)
Our redact tests run parallel under -race. betterleaks shares mutable state across goroutines → race → every redact test fails.
Cascade:
test-core— real failure (FAIL github.com/entireio/cli/redact)test-canary (git-branch/refs)— same race, E2E hits redacttest(4s) — just gate job,[ "failure" = "success" ]fails because above failedlint,test-integration,binary-size, licenses — pass (no-raceor don't hit concurrent redact)
Our code not broken. Bug in bumped dep.
Options:
- Pin betterleaks back to 1.5.0 (drop it from group bump), merge rest.
- Wait for upstream fix / try 1.6.2+ if exists.
- Serialize betterleaks calls in
redact(mutex) — masks their bug, not great.
Want me check for newer betterleaks release or repro race locally?
Newer release check
v1.6.1 is latest betterleaks. No upstream fix available. (1.6.0 and 1.6.1 only in the new line.)
Local repro — confirmed, root cause found
Race reproduces instantly on this branch (betterleaks v1.6.1), same trace as CI. Bug is in the dep, runtime.go:200-213:
func (prg Program) evalBindings() bindings {
b := cloneBindings(prg.bindings) // clones the MAP...
if rt, ok := b["__runtime"].(*runtimeBindings); ok {
rt.tokenizer = prg.tokenizer // line 208: writes shared *runtimeBindings
rt.tokenizerProvider = prg.tokenizerProvider
}
return b
}
cloneBindings shallow-copies the map, but b["__runtime"] stays the same pointer. Every concurrent EvalFilter writes rt.tokenizer into that one shared struct → data race. Comment even claims "shallow copy so compiled programs can share static bindings" — they missed that __runtime is mutable, not static.
Recommendation
No upstream fix, can't wait. Two viable paths:
- Pin betterleaks 1.5.0, take the other 4 bumps. Cleanest — the golang.org/x/* bumps are unrelated and safe. Un-bump just the broken one.
- Serialize our calls with a mutex in
redact.detectAllLayers. Works but throttles redaction throughput and masks their bug.
I'd go option 1 + file an upstream issue. Want me to:
- edit
go.mod/go.sumto revert betterleaks to 1.5.0 (keep the rest), verify race gone, push? - and/or draft the upstream bug report?
Summary of what's done
- ✅ betterleaks pinned back to 1.5.0, other 4 group bumps kept
- ✅
go mod tidyclean, race gone (go test -race ./redact/→ ok) - ✅ committed + pushed to the dependabot branch (
754e704f9) — CI re-running - ✅ upstream bug already tracked as #224; comment drafted, awaiting your go-ahead to post
One caveat: dependabot will re-open this bump next cycle and try to push 1.6.1 again. Worth adding an ignore entry for betterleaks``[">=1.6.0"] in .github/dependabot.yml until #224 lands. Want that too?
Full state of PR #1690
| Change | Status |
|---|---|
| betterleaks pinned 1.5.0, kept x/mod, x/sync, x/sys, x/term bumps | pushed 754e704f9 |
dependabot ignore >=1.6.0 |
pushed 277a4819e |
race verified gone locally (go test -race ./redact/ ok) |
✓ |
| upstream #224 confirmed on 1.6.1 | commented |
CI re-running on the branch now. Once green, mergeable. When #224 lands upstream, drop the ignore entry to resume betterleaks bumps.
One flag: this is a dependabot branch — if you rebase/recreate the PR from the dependabot UI, it may clobber these two manual commits. Merge as-is rather than telling dependabot to recreate.