Fix subdivide retry loop skipping inserted midpoints, update docs · Entire

Fix subdivide retry loop skipping inserted midpoints, update docs

22f8778→main·Soph·3mo ago·2 files·+32 added/-16 removed

The subdivide-on-failure logic used continue inside a for idx++ loop, which incremented idx past the newly-inserted midpoint checkpoint on every retry. The result: infinite retries of the same oversized range (visible in the linux smoke test as repeated pushes of 6.5M objects at the same from/to hashes).

Fix: convert to a manual-index loop (idx := start; for idx < len; ... idx++ at the bottom). Now continue retries at the same idx, where the checkpoint slice has been rewritten to start with the midpoint. Both the PACK-header precheck path and the post-push target-rejection path benefit from the same fix.

Also updates docs/bootstrap-batching.md:

Co-Authored-By: Claude Opus 4.6 (1M context) noreply@anthropic.com

Sessions

e00bb1b6c6d8View transcript

Changes

2

Checkpoint Selection

The most practical first heuristic is first-parent checkpointing. Checkpoints are placed using a commit-count estimate, not measured pack sizes.

For a branch tip:

  1. Walk first-parent ancestry backward.
  2. Sample candidate checkpoint commits.
  3. Starting from the oldest candidate, estimate each batch by doing a source fetch against the previous checkpoint as have.
  4. Pick the largest checkpoint whose pack stays under --batch-max-pack-bytes.
  5. Repeat until the branch tip is reached.
    1. Fetch the commit graph (tree:0 filter, one round-trip — commits only, no blobs/trees).
    2. Walk first-parent ancestry backward to get the chain length.
    3. Estimate total pack size: chainLen × 8 KiB/commit.
    4. Compute number of batches: ceil(estimated / --batch-max-pack-bytes).
    5. Place checkpoints evenly along the first-parent chain.

This is only a heuristic: This is a heuristic — real bytes-per-commit varies widely (2–100+ KiB depending on blob churn). The estimate intentionally errs toward more batches.

Adaptive size correction

But it is much simpler than exact graph partitioning and good enough for a first implementation. If the estimate is too optimistic (fewer batches than needed), two safeguards catch it:

  1. PACK header pre-check: after starting a fetch, peek at the first 12 bytes of the pack to read the object count. Multiply by ~750 bytes/object. If the estimate exceeds --batch-max-pack-bytes, abort the fetch (12 bytes wasted, not gigabytes), insert a midpoint checkpoint, and retry. This avoids a full transfer for obviously-oversized batches.

  2. Target rejection retry: if the target's receive-pack rejects a push for exceeding its body-size limit, detect the error, insert a midpoint checkpoint from the stored chain, and retry. This catches cases where the PACK header estimate was close but the real pack was slightly over.

Both safeguards converge in O(log n) splits — each failure halves the commit range.

Why not probe (the previous design)

The previous implementation did full FetchPack round-trips per probe candidate to measure actual pack sizes. For linux/master (75k commits) this required 13+ fetch-and-discard cycles, downloading gigabytes of throwaway data and taking minutes before any real push started. The estimate approach reduces planning to one commit-graph fetch (~20 seconds) plus arithmetic.

Batch Flow For One Branch

Practical Risks

This is still likely worthwhile for very large initial migrations because it changes a single huge risky operation into several bounded ones.