Recombine checkpoints when consecutive packs underuse the target limit · Entire
Recombine checkpoints when consecutive packs underuse the target limit
7a03676→main·
Soph·2mo ago·2 files·+102 added/-0 removed
Subdivision is a one-way ratchet: the fine granularity needed to fit one heavy commit through the limit sticks for the rest of the chain, leading to thousands of tiny pushes after the heavy region has passed. cli-checkpoints reproduces this — one ~30 MB commit forces 928 → 7967 splits, but the commits behind it are 6-object deltas that comfortably fit dozens per pack.
After every successful push, drop enough upcoming checkpoints that the next pack should land near half the limit. Each dropped checkpoint roughly doubles the next pack's span, so the count is log2(target/2 / sent), capped to keep recovery cost bounded if a heavy commit shows up immediately after. If we overshoot, the existing abort-early plus subdivision path re-splits.
Sessions
8a50df172b3eView transcript
?\can you rebase soph/progress-indicators onto soph/smart-subdivisionClaude Code·Opus 4.7[1m]·1 step
Changes
2
internal/strategy/bootstrap
Mbootstrap.go+48
Mbootstrap_test.go+54
578 unmodified lines
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
378 unmodified lines
984
985
986
987
988
989
990
991
992
993
994
995
996
997
998
999
1000
1001
1002
1003
1004
1005
1006
1007
1008
1009
1010
1011
1012
1013
1014
1015
1016
578 unmodified lines
current = checkpoint
pushedCheckpoints = append(pushedCheckpoints, checkpoint)
result.BatchCount++
// Recombine: subdivision is a one-way ratchet, so the fine
// granularity needed for one heavy commit sticks around for
// the rest of the chain even when the deltas afterward are
// tiny. Aim for the next pack to be roughly half the target
// limit by dropping enough checkpoints that the span doubles
// approximately log2(target/2 / sent) times. If we
// overshoot, the abort-early + subdivision path re-splits.
// Leave at least the final checkpoint after idx (it carries
// the SourceHash cutover), so the cap is len-idx-2.
if dropCount := recombineDropCount(sentBytes, p.TargetMaxPack, len(batch.Checkpoints)-idx-2); dropCount > 0 {
dropped := batch.Checkpoints[idx+1]
batch.Checkpoints = append(batch.Checkpoints[:idx+1], batch.Checkpoints[idx+1+dropCount:]...)
p.log("bootstrap batch recombining after small push",
"branch", batch.Plan.TargetRef.String(),
"sent_bytes", sentBytes,
"target_limit_bytes", p.TargetMaxPack,
"dropped_count", dropCount,
"first_dropped_checkpoint", planner.ShortHash(dropped),
"remaining_checkpoints", len(batch.Checkpoints))
}
idx++
}
378 unmodified lines
return -1
}
// recombineDropCount picks how many of the upcoming checkpoints to
// drop after a small successful push. Each dropped checkpoint roughly
// doubles the span of the next pack — so doubling sentBytes until
// hitting target/2 gives the count. Capped by maxDrop (always leave
// at least one checkpoint ahead, including the final one) and by a
// hard ceiling that keeps any single overshoot's recovery cost
// bounded. Returns 0 when sentBytes already used at least half the
// limit, when we have no headroom to estimate, or when nothing can be
// dropped.
func recombineDropCount(sentBytes, targetLimit int64, maxDrop int) int {
const hardCap = 8
if sentBytes <= 0 || targetLimit <= 0 || maxDrop <= 0 {
return 0
}
target := targetLimit / 2
if sentBytes >= target {
return 0
}
count := 0
span := sentBytes
for span*2 <= target && count < maxDrop && count < hardCap {
span *= 2
count++
}
return count
}
// minBytesBeforeAbort is the floor below which the projection-based
// abort heuristic stays silent. The first few KB of a pack are header
// + small objects; their bytes/object ratio doesn't represent the