Cover front-loaded aborts and sub-floor budgets in early-abort path · Entire

Cover front-loaded aborts and sub-floor budgets in early-abort path

Two regressions of the early-abort design surfaced in review:

  1. When a single front-loaded large blob exhausts the 8 MiB abort floor before any object completes scanning, the observer reports objectsSent == 0 even though totalObjects > 0. The prior calibration and projection paths required objectsSent > 0, so they fell back to the full pack header count for division and to raw sentBytes for sizing — yielding a factor of 2 every retry and reproducing the slow 1→2→4→… convergence this branch is meant to remove. Treat the partially-observed first object as one observation in that case so calibration and projection get a pessimistic-but-bounded per-object estimate.

  2. nextSelfImposedBudget can ratchet the budget below the projection-path floor (e.g. to a 5 MiB proxy cutoff), but shouldAbortPush returned false until bytesSent ≥ 8 MiB regardless of budget, so the learned ceiling could never trigger a client-side abort and every retry kept paying for full server-side rejection. Reorder shouldAbortPush so the absolute "we already crossed the threshold" trigger fires first; minBytesBeforeAbort still gates the projection path, which is what it actually existed to suppress.

Sessions

Changes

// average for the portion we actually observed —
// and for blob-front-loaded repos that's a pessimistic
// upper bound, which is what we want for pre-flight.
//
// When the header was parsed but no object completed
// (a single front-loaded large blob exhausted the
// abort floor), treat it as one observed object: we
// know that first object alone consumed sentBytes,
// which is the right pessimistic per-object input.
effObjectsSent := effectiveObjectsSent(objectsSent, totalObjects, abortedEarly)
calibrationDenom := packObjectCount
if objectsSent > 0 && objectsSent < calibrationDenom {
    calibrationDenom = objectsSent
}
if effObjectsSent > 0 && effObjectsSent < calibrationDenom {
    calibrationDenom = effObjectsSent
}
if updated := calibrateBytesPerObject(sentBytes, calibrationDenom, calibratedBytesPerObject); updated > 0 {
    p.log("bootstrap batch calibrated bytes-per-object",
}

Tests

func effectiveObjectsSent(objectsSent, totalObjects int64, abortedEarly bool) int64 {
    if objectsSent == 0 && totalObjects > 0 && abortedEarly {
        return 1
    }
    return objectsSent
}

func shouldAbortPush(bytesSent, objectsSent, totalObjects, budget int64) bool {
    if budget <= 0 || bytesSent < minBytesBeforeAbort {
        return false
    }
    const safety = 95 // percent of budget at which we cut
    threshold := budget * safety / 100
    if bytesSent >= threshold {
        return true
    }
    if bytesSent < minBytesBeforeAbort {
        return false
    }
    if objectsSent > 0 && totalObjects > 0 {
        projected := bytesSent * totalObjects / objectsSent
        return projected > threshold
    }
    return bytesSent > threshold
}

func TestEffectiveObjectsSent(t *testing.T) {
    cases := []struct {
        name         string
        objectsSent  int64
        totalObjects int64
        abortedEarly bool
        want         int64
    }{
        {"front-loaded blob abort treats first object as observed", 0, 100, true, 1},
        {"no header parsed leaves zero", 0, 0, true, 0},
        {"header parsed but server-rejected, no synthetic observation", 0, 100, false, 0},
        {"actual observation passes through", 12, 100, true, 12},
    }
    for _, c := range cases {
        t.Run(c.name, func(t *testing.T) {
            got := effectiveObjectsSent(c.objectsSent, c.totalObjects, c.abortedEarly)
            if got != c.want {
                t.Errorf("effectiveObjectsSent(%d, %d, %v) = %d, want %d",
                    c.objectsSent, c.totalObjects, c.abortedEarly, got, c.want)
            }
        })
    }
}