Cover front-loaded aborts and sub-floor budgets in early-abort path · Entire
Cover front-loaded aborts and sub-floor budgets in early-abort path
Two regressions of the early-abort design surfaced in review:
When a single front-loaded large blob exhausts the 8 MiB abort floor before any object completes scanning, the observer reports objectsSent == 0 even though totalObjects > 0. The prior calibration and projection paths required objectsSent > 0, so they fell back to the full pack header count for division and to raw sentBytes for sizing — yielding a factor of 2 every retry and reproducing the slow 1→2→4→… convergence this branch is meant to remove. Treat the partially-observed first object as one observation in that case so calibration and projection get a pessimistic-but-bounded per-object estimate.
nextSelfImposedBudget can ratchet the budget below the projection-path floor (e.g. to a 5 MiB proxy cutoff), but shouldAbortPush returned false until bytesSent ≥ 8 MiB regardless of budget, so the learned ceiling could never trigger a client-side abort and every retry kept paying for full server-side rejection. Reorder shouldAbortPush so the absolute "we already crossed the threshold" trigger fires first; minBytesBeforeAbort still gates the projection path, which is what it actually existed to suppress.
Sessions
Changes
internal/strategy/bootstrap
Mbootstrap.go+43/-8
// average for the portion we actually observed —
// and for blob-front-loaded repos that's a pessimistic
// upper bound, which is what we want for pre-flight.
//
// When the header was parsed but no object completed
// (a single front-loaded large blob exhausted the
// abort floor), treat it as one observed object: we
// know that first object alone consumed sentBytes,
// which is the right pessimistic per-object input.
effObjectsSent := effectiveObjectsSent(objectsSent, totalObjects, abortedEarly)
calibrationDenom := packObjectCount
if objectsSent > 0 && objectsSent < calibrationDenom {
calibrationDenom = objectsSent
}
if effObjectsSent > 0 && effObjectsSent < calibrationDenom {
calibrationDenom = effObjectsSent
}
if updated := calibrateBytesPerObject(sentBytes, calibrationDenom, calibratedBytesPerObject); updated > 0 {
p.log("bootstrap batch calibrated bytes-per-object",
}
Tests
func effectiveObjectsSent(objectsSent, totalObjects int64, abortedEarly bool) int64 {
if objectsSent == 0 && totalObjects > 0 && abortedEarly {
return 1
}
return objectsSent
}
func shouldAbortPush(bytesSent, objectsSent, totalObjects, budget int64) bool {
if budget <= 0 || bytesSent < minBytesBeforeAbort {
return false
}
const safety = 95 // percent of budget at which we cut
threshold := budget * safety / 100
if bytesSent >= threshold {
return true
}
if bytesSent < minBytesBeforeAbort {
return false
}
if objectsSent > 0 && totalObjects > 0 {
projected := bytesSent * totalObjects / objectsSent
return projected > threshold
}
return bytesSent > threshold
}
func TestEffectiveObjectsSent(t *testing.T) {
cases := []struct {
name string
objectsSent int64
totalObjects int64
abortedEarly bool
want int64
}{
{"front-loaded blob abort treats first object as observed", 0, 100, true, 1},
{"no header parsed leaves zero", 0, 0, true, 0},
{"header parsed but server-rejected, no synthetic observation", 0, 100, false, 0},
{"actual observation passes through", 12, 100, true, 12},
}
for _, c := range cases {
t.Run(c.name, func(t *testing.T) {
got := effectiveObjectsSent(c.objectsSent, c.totalObjects, c.abortedEarly)
if got != c.want {
t.Errorf("effectiveObjectsSent(%d, %d, %v) = %d, want %d",
c.objectsSent, c.totalObjects, c.abortedEarly, got, c.want)
}
})
}
}