Replace probe-based bootstrap checkpoint planning with estimate · Entire

Replace probe-based bootstrap checkpoint planning with estimate

8e4c21c→main·

Soph·3mo ago·3 files·+117 added/-652 removed

The previous checkpoint planner did full FetchPack round-trips per probe candidate to measure actual pack sizes, then binary-searched for the boundary that fit under --batch-max-pack-bytes. For linux/master (75k commits) this took 13+ fetch-and-discard cycles — downloading gigabytes of throwaway data and burning minutes — before any real push started. The precision was rarely needed: the adaptive retry and resume mechanisms already handle batches that turn out too large.

Replace with estimate-based planning:

  1. Fetch the commit graph (tree:0 filter, one round-trip — unchanged).
  2. Walk the first-parent chain to get commit count (unchanged).
  3. Estimate total pack size as chainLen × 8 KiB/commit.
  4. Divide into ceil(estimated / batchMaxPack) evenly-spaced checkpoints.
  5. Done. No probe fetches.

For linux at 1 GiB batch limit: planning goes from ~4 minutes / 13 fetches to ~20 seconds / 1 fetch (just the commit graph). The estimate is intentionally conservative (8 KiB vs the old 4 KiB) so it errs toward more batches rather than fewer. If a batch still exceeds the target's limit, the push fails for that batch and bootstrap resume (via temp refs) ensures already-pushed batches aren't re-sent on the next run.

Deleted ~535 lines of probe infrastructure:

Added:

Co-Authored-By: Claude Opus 4.6 (1M context) noreply@anthropic.com

Sessions

d0974f7b54b6View transcript

Changes

3


package bootstrap

import (
    "bytes"
    "context"
    "encoding/json"
    "errors"
)

// Execute runs the bootstrap strategy (one-shot or batched).
func planBatches(ctx context.Context, p Params, desired []planner.DesiredRef) ([]plannedBatch, error) {
    out := make([]plannedBatch, 0, len(desired))
    for _, ref := range desired {
        checkpoints, prefetched, err := planCheckpointsWithCache(ctx, p, ref)
        if err != nil {
            return nil, err
        }
        out = append(out, plannedBatch{
            Planner: p,
            ResumeHash:  p.TargetRefs[planner.BootstrapTempRef(ref.TargetRef)],
            Checkpoints: checkpoints,
            PrefetchedPacks: prefetched,
        })
    }
    return out, nil
}

// PlanCheckpoints plans the checkpoint hashes for a single branch during batched bootstrap.
func PlanCheckpoints(ctx context.Context, p Params, ref planner.DesiredRef) ([]plumbing.Hash, error) {
    checkpoints, _, err := planCheckpointsWithCache(ctx, p, ref)
    return checkpoints, err
}

// EstimateBatchCount estimates the number of batches needed for a given chain length and pack limit.
func estimateBatchCount(chainLen int64, batchMaxPack int64) int {
    if batchMaxPack <= 0 || chainLen <= 0 {
        return 1
    }
    estimated := chainLen * estimatedBytesPerCommit
    n := int((estimated + batchMaxPack - 1) / batchMaxPack)
    if n < 1 {
        n = 1
    }
    return n
}

// ... (more function definitions here)
// Your structured test cases, test definitions, etc. continue here...

Conclusion

This new approach optimizes the checkpoint planning process significantly, reducing unnecessary fetches and improving performance.