[Home](/content/site-root.html)

Log in

# first pass at doc update

`c4549ae`→[main](/content/gh/entireio/git-sync/commits/main/index.html)·

Soph·2mo ago·11 files·+950 added/-318 removed

## Changes

11

- ACODE\_OF\_CONDUCT.md+94

- ACONTRIBUTING.md+196

- ALICENSE+21

- MREADME.md+32/-167

- ASECURITY.md+68

- docs

- Mbootstrap-batching.md+25/-40

- Mbootstrap.md+21/-60

- Aincremental-relay.md+63

- Aprotocol.md+357

- Areplicate.md+66

- Mtesting.md+7/-51

```
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94

# Code of Conduct

We're committed to providing a welcoming, respectful, and harassment-free environment for everyone, regardless of age, body size, visible or invisible disability, ethnicity, sex characteristics, gender identity and expression, level of experience, education, socio-economic status, nationality, personal appearance, race, caste, color, religion, or sexual identity and orientation.

## Scope

This Code of Conduct applies to all Entire community spaces, including:

- **GitHub** - repositories, issues, pull requests, discussions, and comments
- **Discord** - all channels in the Entire workspace
- **Events** - meetups, conferences, and online gatherings hosted by Entire
- **Public representation** - acting as a representative of Entire in public spaces, including social media, forums, and conferences

This Code of Conduct also applies when an individual is officially representing the community in public spaces.

## Expected Behavior

- **Be respectful** - Treat everyone professionally, listen actively, and be mindful of your words
- **Be inclusive** - Use inclusive language and make space for everyone to contribute
- **Be collaborative** - Help each other, share knowledge, and celebrate wins together
- **Be accountable** - Own your mistakes, follow through on commitments, and take responsibility for your impact
- **Be empathetic** - Try to understand different perspectives and experiences
- **Give and accept constructive feedback gracefully**

## Unacceptable Behavior

This includes but isn't limited to:

- Harassment, discrimination, or offensive comments related to personal characteristics
- Personal attacks, trolling, insulting or derogatory comments, or deliberate intimidation
- Unwelcome sexual attention, advances, or imagery
- Sharing private information (such as physical or email addresses) without explicit consent
- Sustained disruption of discussions or events
- Advocating for or encouraging any of the above behavior
- Conduct that could reasonably be considered inappropriate in a professional setting

## Reporting

To report a Code of Conduct violation, contact **[conduct@entire.io](mailto:conduct@entire.io)**.

For security vulnerabilities, please email **[security@entire.io](mailto:security@entire.io)** instead. See our [Security Policy](SECURITY.md).

### Confidentiality

**All reports will be kept confidential.** We will not share your identity or the details of your report with anyone outside the enforcement team without your consent, except as required by law or to protect safety.

### What to Include

- Your contact information (so we can follow up)
- Names of those involved (or identifying information)
- Description of the behavior and when/where it occurred
- Any additional context or evidence (screenshots, links)
- Whether you would like to remain anonymous to the reported party

## Enforcement

Community leaders are responsible for clarifying and enforcing standards of acceptable behavior. All complaints will be reviewed and investigated promptly and fairly.

Violations will be addressed according to their severity and frequency:

#### 1. Correction
**For:** First-time minor violations or misunderstandings

**Action:** A private, written notice explaining the violation and why the behavior was inappropriate. A public apology may be requested.

---

#### 2. Warning
**For:** A single incident or pattern of minor violations

**Action:** A formal warning with consequences for continued behavior. This includes no interaction with the people involved for a specified period. Violating these terms may lead to a temporary or permanent ban.

---

#### 3. Temporary Ban
**For:** Serious violations or sustained inappropriate behavior

**Action:** A temporary ban from all community interaction and public communication for a specified period (typically 30-90 days). No public or private interaction with the community is permitted during this time.

---

#### 4. Permanent Ban
**For:** Demonstrating a pattern of violations, severe harassment, or aggression toward individuals or groups

**Action:** Permanent removal from all community spaces. This decision is final.

---

## Attribution

This Code of Conduct is adapted from the [Contributor Covenant](https://www.contributor-covenant.org/version/2/1/code_of_conduct/), version 2.1.

For answers to common questions about the Contributor Covenant, see the [FAQ](https://www.contributor-covenant.org/faq).
```

ACODE\_OF\_CONDUCT.md+94

````
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196

# Contributing to git-sync

Thank you for your interest in contributing! `git-sync` is a remote-to-remote Git mirroring tool and library; we welcome contributions from everyone.

Please read our [Code of Conduct](CODE_OF_CONDUCT.md) before participating.

> **New here?** See the [README](README.md) for setup and usage, and [docs/architecture.md](docs/architecture.md) for the technical overview.

---

## Before You Code: Discuss First

The fastest way to get a contribution merged is to align with maintainers before writing code. Please **open an issue first** on [GitHub Issues](https://github.com/entireio/gitsync/issues) and wait for maintainer feedback before starting implementation.

### Contribution Workflow

1. **Open an issue** describing the problem or feature
2. **Wait for maintainer feedback** -- we may have relevant context or plans
3. **Get approval** before starting implementation
4. **Submit your PR** referencing the approved issue
5. **Address all feedback** including automated review comments
6. **Maintainer review and merge**

---

## First-Time Contributors

New to the project? Welcome! Good places to start:

- **Documentation improvements** - Fix typos, clarify explanations, add examples
- **Test contributions** - Add test cases, improve coverage of edge protocol behaviors
- **Small bug fixes** - Issues labeled `good-first-issue`

---

## Submitting Issues

All feature requests, bug reports, and general issues should be submitted through [GitHub Issues](https://github.com/entireio/gitsync/issues). Please search for existing issues before opening a new one.

For security-related issues, see [SECURITY.md](SECURITY.md) instead.

---

## How to Contribute

There are many ways to contribute:

- **Feature requests** - Open a [GitHub Issue](https://github.com/entireio/gitsync/issues) to discuss your idea
- **Bug reports** - Report issues via [GitHub Issues](https://github.com/entireio/gitsync/issues) (see [Reporting Bugs](#reporting-bugs))
- **Code contributions** - Fix bugs, add features, improve tests
- **Documentation** - Improve guides, fix typos, add examples
- **Community** - Help others, answer questions, share knowledge

## Reporting Bugs

Good bug reports help us fix issues quickly. When reporting a bug, please include:

### Required Information

1. **`git-sync` version or commit** - the binary you ran or `git rev-parse HEAD` if building from source
2. **Operating system**
3. **Go version** - run `go version`
4. **Source and target hosts** - what kind of remote (GitHub, GitLab, self-hosted, etc.) — this matters because protocol behavior differs

### What to Include

1. **What did you do?** - The exact `git-sync` command you ran (redact tokens)
2. **What did you expect to happen?**
3. **What actually happened?** - Full error message, and `--json` output if available
4. **Can you reproduce it?** - Every time, or intermittently?
5. **Any additional context?** - `--stats` output, `-v` verbose log, related issues

---

## Local Setup

### Prerequisites

- **Go 1.26.x** - Check with `go version`
- **mise** - Task runner and version manager. Install with `curl https://mise.run | sh`

### Clone and Build

```bash
git clone https://github.com/entireio/gitsync.git
cd gitsync

# Trust the mise configuration (required on first setup)
mise trust

# Install dependencies (mise will install the correct Go version)
mise install

# Download Go modules
go mod download

# Build the CLI
mise run build

# Verify setup by running tests
mise run test
```

---

## Making Changes

1. **Create a branch** for your changes:
   ```bash
   git checkout -b your-name/feature-name
   ```

2. **Make your changes** - follow the [Code Style](#code-style) guidelines.

3. **Test your changes** - see [Testing](#testing).

4. **Commit** with clear, descriptive messages.

---

## Code Style

Follow standard Go idioms and conventions.

### Key Points

- **Error handling**: Handle all errors explicitly - don't leave them unchecked
- **Formatting**: Code must pass `gofmt` (run `mise run fmt`)
- **Linting**: Code must pass `golangci-lint` (run `mise run lint`)
- **Naming**: Use meaningful, descriptive names following Go conventions
- **Public API**: `entire.io/entire/gitsync` is the stable embedding surface. Additions there should be reviewed carefully. `entire.io/entire/gitsync/unstable` is for advanced controls and may change. See [docs/embedding.md](docs/embedding.md).

---

## Testing

```bash
# Default suite (in-process smart HTTP, no listener required)
mise run test

# With race detection
mise run test:ci

# Optional: end-to-end against the system git-http-backend
mise run test:git-http-backend

# Optional: live linux bootstrap smoke (downloads from github)
mise run test:linux-smoke
mise run test:linux-smoke:batched
```

See [docs/testing.md](docs/testing.md) for the full list of suites and environment flags.

---

## Submitting a Pull Request

### Before You Submit

- **Related issue exists and is approved** -- Your PR references an issue where a maintainer has acknowledged the approach. (Exceptions: documentation fixes, typo corrections, and `good-first-issue` items.)
- **Linting passes** -- Run `mise run lint`
- **Tests pass** -- Run `mise run test`
- **Tests included** -- New Go code and behavior changes should have accompanying tests. Protocol-level changes should ideally be covered both in `internal/gitproto` unit tests and in an `internal/syncer` integration test.

PRs that skip these steps are likely to be closed without merge.

### Submitting

1. **Push** your branch to your fork
2. **Open a PR** against the `main` branch
3. **Describe your changes** -- Link the related issue, summarize what changed and what testing you did
4. **Address automated review feedback**
5. **Wait for maintainer review**

---

## Community

- **GitHub Issues** - bug reports, feature discussions
- **Discord** - [Join our server](https://discord.gg/jZJs3Tue4S) for questions and real-time conversation

---

## Additional Resources

- [README](README.md) - Setup and usage documentation
- [docs/architecture.md](docs/architecture.md) - Technical architecture and package layout
- [docs/embedding.md](docs/embedding.md) - Library embedding guide
- [docs/testing.md](docs/testing.md) - Test suites and integration coverage
- [Code of Conduct](CODE_OF_CONDUCT.md) - Community guidelines
- [Security Policy](SECURITY.md) - Reporting security vulnerabilities

---

Thank you for contributing!
````

ACONTRIBUTING.md+196

```
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21

MIT License

Copyright (c) 2026 Entire Inc.

Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:

The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.

THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.
```

ALICENSE+21

````
9 unmodified lines

10
11
12
13
13
14
15
15
16
17
18
19
20
17
18
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
19
20
21
22
23
57
58
24
25
26
62
27
28
29
30
31
32
33
68
34
35
36
37
72
38
39
40
41
46 unmodified lines

88
89
90
125
91
92
93
94
25 unmodified lines

120
121
122
157
123
124
125
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
126
127
128
129
1 unmodified line

131
132
133
210
134
135
136
213
137
138
139
140
141
142
143
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
144
145
253
254
255
256
257
146
147
148
149
53 unmodified lines

203
204
205
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
206
207
208
2 unmodified lines

211
212
213
346
214
215
216
217
218
219
220
353
354
355
221
222
223
224
23 unmodified lines

248
249
250
385
386
387
388
389
390
391
392
393
394
251
252
253
254
255
256
398
257
258
259
260
5 unmodified lines

266
267
268
410
269
270
412
271
272
273
274
275
276
277
278
279
280
414
281
282
416
417
418
283

9 unmodified lines

## Why This Exists

Git already has pieces of this problem, but not this exact tool shape.
The usual ways to mirror Git data between remotes are awkward at exactly the layer operators tend to need: a local `git clone --mirror` followed by `git push --mirror` turns a remote-to-remote movement into a local storage and bandwidth problem; host-specific migration features aren't portable across providers; and shell scripts around `git fetch` and `git push` usually lack planning, explicit policy, and machine-readable output.

What usually exists today:
`git-sync` is meant to be the missing middle layer: a provider-agnostic, remote-to-remote primitive that streams packs directly source-to-target when possible, front-loads validation, exposes typed JSON output, and covers both empty-target bootstrap and incremental sync with one tool. It's the right fit when relay is common enough to be the normal case rather than an exceptional optimization, and when avoiding persistent local repo storage is itself an operational advantage.

- a full local `git clone --mirror` followed by `git push --mirror`
- host-specific import or migration features
- CI jobs or shell scripts that glue fetch and push steps together
- one-off migration tooling tied to a specific platform
For when to use it (and when not), how it compares to local-clone services, and the operation-mode and transfer-mode model, see [docs/architecture.md](docs/architecture.md).

What those approaches usually do not give you:

- direct remote-to-remote relay behavior
- a small standalone CLI with explicit sync semantics
- front-loaded validation and planning
- machine-readable output for automation
- one tool that covers empty-target bootstrap, normal sync, and large-repo bootstrap fallback

That is the gap `git-sync` is trying to fill.

The main value is operational:

- avoid requiring a full local mirror checkout just to move refs between remotes
- make initial seeding of large repositories cheaper and more predictable
- keep incremental sync behavior explicit and safe
- give operators and automation a stable way to inspect, plan, execute, and benchmark the same workflows

This is especially useful when:

- the target is a new hosted Git service or internal Git endpoint
- bootstrap size matters more than local developer ergonomics
- you want a repeatable machine-oriented sync primitive rather than an ad hoc migration script
- you need clearer control over mapping, pruning, force rules, and relay behavior than generic shell glue usually provides

Compared to a service that keeps persistent local clones, `git-sync` is the better fit when:

- relay is common enough that streaming source-to-target is the normal case
- avoiding persistent local repo storage is an operational advantage
- remote-to-remote efficiency matters more than full local Git generality

If you need arbitrary complex reconciliation through one always-warm local full-state model, a local-clone service is still the more general tool.
## Commands

The command surface is:

- `git-sync probe`: inspect a source remote, and optionally a target remote
- `git-sync fetch`: exercise source-side fetch negotiation without pushing
- `git-sync bootstrap`: seed an empty target with create-only relay behavior
- `git-sync plan`: compute source-to-target ref actions without pushing, with `--mode sync|replicate`
- `git-sync sync`: execute the planned changes against the target
- `git-sync replicate`: execute source-authoritative relay-only replication against the target
- `git-sync-bench`: run repeatable benchmark scenarios against fresh empty targets

`sync` auto-selects the bootstrap relay path on empty targets, so the same command covers initial seeding and ongoing sync.

## Library API

`git-sync` now has a two-tier Go API:

- `gitsync`
- `entire.io/entire/gitsync`
  - stable embedding surface for queue workers and other external callers
  - typed `Probe`, `Plan`, `Sync`, and `Replicate` requests/results
  - injected auth and HTTP client support
- `unstable`
- `entire.io/entire/gitsync/unstable`
  - explicitly non-stable surface for first-party tooling and advanced controls
  - includes `Bootstrap`, `Fetch`, batching and measurement knobs, and CLI-oriented execution options

46 unmodified lines

https://github.com/target-org/target-repo.git
```

## Commands
## Examples

Plan a sync without pushing anything:

25 unmodified lines

If `replicate` cannot use relay against the target, it fails and tells you to rerun with `sync`.

Bootstrap an empty target without using the normal local object-store sync path:
For very large initial migrations, add `--target-max-pack-bytes` to split the initial pack into multiple relay batches with temporary refs. `sync` auto-bootstraps on empty targets, so the same flag works without invoking a separate command:

```bash
go run ./cmd/git-sync bootstrap \
  --stats \
  https://github.com/source-org/source-repo.git \
  https://github.com/target-org/target-repo.git
```

Add `--max-pack-bytes` to abort bootstrap if the streamed source pack grows past a safety threshold:

```bash
go run ./cmd/git-sync bootstrap \
  --max-pack-bytes 104857600 \
  <source-url> \
  <target-url>
```

Add `--target-max-pack-bytes` to split large branch bootstraps into multiple relay batches with temporary refs:

```bash
go run ./cmd/git-sync bootstrap \
  --target-max-pack-bytes 1073741824 \
  <source-url> \
  <target-url>
```

Current batching scope is intentionally narrow:

- protocol v2 only
- branch refs are batched
- optional create-only tags are pushed after branch batches complete
- temporary refs under `refs/gitsync/bootstrap/heads/`
- resume from existing temp refs is supported when they match a planned checkpoint

This mode is intended as an advanced large-repo fallback, not the default bootstrap path. Use plain `bootstrap` first when a single streamed initial sync is acceptable.

A practical starting point is:

- `--target-max-pack-bytes 536870912` for a conservative `512 MiB` target-side batch size
- `--target-max-pack-bytes 1073741824` when you want fewer, larger batches and the target has more headroom

For example:

```bash
go run ./cmd/git-sync bootstrap \
go run ./cmd/git-sync sync \
  --target-max-pack-bytes 536870912 \
  --protocol v2 \
  -v \
1 unmodified line

<target-url>
```

Add `--measure-memory` to `bootstrap`, `sync`, `plan`, `probe`, or `fetch` to sample elapsed time and Go heap usage:
Add `--measure-memory` to any command to sample elapsed time and Go heap usage:

```bash
go run ./cmd/git-sync bootstrap \
go run ./cmd/git-sync sync \
  --measure-memory \
  --json \
  <source-url> \
  <target-url>
```

That is useful for one-off measurements on the same fixture or test repo.

## Benchmarking

For repeated benchmark runs, prefer the dedicated benchmark command instead of manually wrapping `git-sync` invocations:

```bash
go run ./cmd/git-sync-bench \
  --scenario bootstrap \
  --source-url /tmp/git-sync-bench/kubernetes.git \
  --repeat 3 \
  --target-max-pack-bytes 104857600 \
  --stats \
  --json
```

`git-sync-bench` creates a fresh bare target repository for each run, executes the selected scenario in-process, and reports:

- per-run wall-clock time
- per-run `syncer.Result`
- aggregate min/avg/max wall time
- aggregate internal elapsed and heap metrics from `--measure-memory`
- relay modes observed across successful runs

If `--source-url` is a local path, it is converted to `file://...` automatically. The current scenarios are:

- `--scenario bootstrap`
- `--scenario sync`

For large-repo measurements, use a local bare mirror as the source so the benchmark reflects `git-sync` behavior rather than internet variance. See [docs/benchmarking.md](docs/benchmarking.md) for details.

## Sync Behavior

When `sync` sees that all managed target refs are absent and the run is compatible with bootstrap semantics, it automatically uses the bootstrap relay path instead of the normal decode-and-repack sync path.

`sync` also uses a narrow incremental relay path for fast-forward branch updates and tag creation when there is no prune/delete, no force, and the target does not advertise `no-thin`. This now includes multi-branch batches, branch-to-branch mappings, and create-only tags. Tag retargeting and other more complex updates still fall back to the normal local decode-and-repack path.

If `sync` falls back to the materialized path, `--materialized-max-objects` sets an explicit object-count safety bound for the in-memory object set. It is a conservative guardrail, not a precise heap-size limit.
`sync` auto-selects the bootstrap relay path when the target has no managed refs and the run matches bootstrap semantics. It also has a narrow incremental relay path for safe fast-forward updates that streams the source pack directly into target `receive-pack` without local materialization. Updates that aren't relay-eligible (force, prune, deletes, tag retargets) fall back to a materialized path bounded by `--materialized-max-objects`. See [docs/incremental-relay.md](docs/incremental-relay.md) and [docs/bootstrap.md](docs/bootstrap.md) for details.

Sync specific branches:

53 unmodified lines

<target-url>
```

Fetch from a source remote into memory without pushing anywhere:

```bash
go run ./cmd/git-sync fetch \
  --stats \
  --protocol auto \
  --branch main \
<source-url>
```

Advertise an existing source ref as a synthetic `have` to exercise incremental negotiation:

```bash
go run ./cmd/git-sync fetch \
  --stats \
  --protocol auto \
  --branch main \
  --have-ref main \
<source-url>
```

Dry run:

```bash
2 unmodified lines

## JSON Output

Add `--json` to `probe`, `fetch`, `bootstrap`, `plan`, or `sync` to emit machine-readable output instead of the default text format.
Add `--json` to `probe`, `plan`, or `sync` to emit machine-readable output instead of the default text format.

The JSON interface is intentionally stable:

- keys use `camelCase`
- refs and hashes are serialized as strings, not raw byte arrays
- `probe` returns top-level keys such as `sourceUrl`, `targetUrl`, `protocol`, `refPrefixes`, `sourceCapabilities`, `targetCapabilities`, `refs`, and `stats`
- `fetch` returns top-level keys such as `sourceUrl`, `protocol`, `wants`, `haves`, `fetchedObjects`, and `stats`
- `bootstrap`, `plan`, and `sync` return top-level keys such as `plans`, `pushed`, `skipped`, `blocked`, `deleted`, `dryRun`, `protocol`, and `stats`
- `bootstrap`, `plan`, and `sync` also expose `relay`, `relayMode`, `relayReason`, `batching`, `batchCount`, `plannedBatchCount`, and `tempRefs`
- `plan` and `sync` return top-level keys such as `plans`, `pushed`, `skipped`, `blocked`, `deleted`, `dryRun`, `protocol`, and `stats`, and also expose `relay`, `relayMode`, `relayReason`, `batching`, `batchCount`, `plannedBatchCount`, and `tempRefs`
- each item in `plans` includes stable string fields such as `branch`, `sourceRef`, `targetRef`, `sourceHash`, `targetHash`, `kind`, `action`, and `reason`

## Auth
23 unmodified lines

## Protocol Notes

- Source refs are listed with `GET /info/refs?service=git-upload-pack`.
- When the source supports it, the client can negotiate protocol v2 with `Git-Protocol: version=2`, then use `ls-refs` and `fetch`.
- Target refs are listed with `GET /info/refs?service=git-receive-pack`.
- The source fetch advertises current target tip hashes as `have`, so reruns download less when source and target already share history.
- Target push stays on the current `receive-pack` path.
- If a target ref does not exist, it is created.
- If a target ref already matches the source, it is skipped.
- Branches are updated only when the target tip is an ancestor of the source tip, unless `--force` is set.
- Tags are immutable by default. Retargeting an existing tag requires `--force`.
- If `--prune` is set, managed target refs that are absent on source are deleted.
- Source-side discovery and fetch can use protocol v2 when supported; push stays on the existing v1 `receive-pack` path. `--protocol auto` tries v2 first and falls back to v1; `--protocol v2` requires the source to negotiate v2.
- Source fetch advertises current target tip hashes as `have`, so reruns download less when source and target already share history.
- Branches are updated only when the target tip is an ancestor of the source tip, unless `--force` is set. Tags are immutable by default; retargeting an existing tag requires `--force`. If `--prune` is set, managed target refs that are absent on source are deleted.
- `plan` never pushes. If `sync` finds blocked refs, it exits non-zero before pushing anything.
- `--stats` adds per-service request, byte, want, have, and command counters to the output.

Push still uses the current low-level `receive-pack` path. Protocol v2 is used where it materially improves this tool: source-side ref discovery and source-side object download.
For the deeper protocol-level walkthrough (smart HTTP, pkt-line, capability negotiation, sideband stripping, relay framing), see [docs/protocol.md](docs/protocol.md).

## Testing

5 unmodified lines

Extended and environment-specific test instructions are in [docs/testing.md](docs/testing.md).

## Design Notes
## Documentation

`bootstrap` is the dedicated path for large initial syncs into an empty target. The goal is to relay a fetched source pack directly into target `receive-pack` instead of decoding the full object graph into local memory first.
- [docs/architecture.md](docs/architecture.md) — product rationale, package layout, operation modes vs transfer modes, memory model
- [docs/protocol.md](docs/protocol.md) — smart HTTP, pkt-line, capability negotiation, sideband, relay framing
- [docs/bootstrap.md](docs/bootstrap.md) — empty-target relay
- [docs/bootstrap-batching.md](docs/bootstrap-batching.md) — checkpoint batching for very large initial migrations
- [docs/incremental-relay.md](docs/incremental-relay.md) — narrow relay fast path inside `sync`
- [docs/replicate.md](docs/replicate.md) — source-authoritative relay-only overwrite mode
- [docs/embedding.md](docs/embedding.md) — using `git-sync` as a Go library
- [docs/benchmarking.md](docs/benchmarking.md) — `git-sync-bench` usage
- [docs/testing.md](docs/testing.md) — test suites and integration coverage

Current architectural summary and package boundaries are in [docs/architecture.md](docs/architecture.md).
## Contributing

The design note is in [docs/bootstrap.md](docs/bootstrap.md).

For very large single-branch repositories, there is also a batching design and initial implementation note in [docs/bootstrap-batching.md](docs/bootstrap-batching.md).
See [CONTRIBUTING.md](CONTRIBUTING.md), [SECURITY.md](SECURITY.md), and [CODE_OF_CONDUCT.md](CODE_OF_CONDUCT.md).
````

MREADME.md+32/-167

```
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68

# Security Policy

We take security seriously at Entire. We appreciate your efforts to responsibly disclose vulnerabilities and will make every effort to acknowledge your contributions.

## Reporting a Vulnerability

**Please do not report security vulnerabilities through public GitHub issues.**

Instead, please send security-related reports to **[security@entire.io](mailto:security@entire.io)**.

### What to Include

When reporting a vulnerability, please include:

1. **Description** - A clear description of the vulnerability
2. **Impact** - What an attacker could achieve by exploiting this issue
3. **Steps to reproduce** - Detailed steps to reproduce the vulnerability
4. **Affected versions** - Which versions of `git-sync` are affected (if known)
5. **Suggested fix** - If you have ideas on how to fix it (optional)

### What to Expect

- **Acknowledgment** - We will acknowledge receipt of your report within 48 hours
- **Updates** - We will keep you informed of our progress as we investigate
- **Resolution** - We aim to resolve critical vulnerabilities within 90 days

## Confidentiality

**All reports will be kept confidential.** We will not share your information with third parties without your consent, except as required by law.

## Supported Versions

We recommend always running the latest version of `git-sync`.

## Scope

This security policy applies to:

- The `git-sync` CLI and `git-sync-bench` benchmark command
- The `entire.io/entire/gitsync` and `entire.io/entire/gitsync/unstable` Go packages
- Official Entire GitHub repositories

### Out of Scope

The following are generally not considered security vulnerabilities:

- Issues in third-party dependencies (please report these upstream)
- Social engineering attacks
- Denial of service attacks against remotes you do not control
- Issues requiring physical access to a user's device

Because `git-sync` operates against Git remotes, please be especially careful when reporting issues that involve credentials, TLS verification, or remote-to-remote relay behavior — include the exact remote configuration that triggers the issue if it is reproducible.

---

## Security Advisories

Security advisories are issued when a confirmed vulnerability can be exploited by a remote or non-local actor. The following are generally treated as **bug reports rather than security advisories**:

- Regular expression performance issues (ReDoS) that only affect local execution
- Resource exhaustion that requires local access to trigger
- Issues that cannot be exploited without direct access to the user's machine or to credentials the user already controls

Use [GitHub Issues](https://github.com/entireio/gitsync/issues) to report bugs.

---

Thank you for helping keep `git-sync` and the Entire community safe!
```

ASECURITY.md+68

```
13 unmodified lines

14
15
16
17
17
18
19
20
100 unmodified lines

121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
38 unmodified lines

178
179
180
168
169
170
171
172
181
182
183
184
9 unmodified lines

194
195
196
188
189
190
197
198
199
200
54 unmodified lines

255
256
257
251
258
259
253
260
261
255
256
257
258
259
262
263
261
262
263
264
264
265
266
267
266
267
268
269
270
271
272
1 unmodified line

274
275
276
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
277

13 unmodified lines

## Non-Goals

V1 batching should not try to solve every large-migration problem.
Batching does not try to solve every large-migration problem.

Out of scope:

100 unmodified lines

Both safeguards converge in O(log n) splits — each failure halves the commit range.

### Trunk-aware planning

For multi-branch bootstraps, planning each branch in isolation re-fetches commit graph history that earlier branches have already reached. With one trunk and many feature branches that all descend from it, that becomes N full-history fetches and N independent first-parent walks over the same shared commits.

`git-sync` avoids this by:

1. Identifying the trunk via the source's HEAD symref (see [protocol.md](protocol.md#head-symref-discovery)) and ordering it first.
2. After each branch is planned, accumulating its first-parent commits into a `planStopSet` and its tip into `planHaves`.
3. For subsequent branches, passing `planHaves` as `have` lines on the commit-graph fetch, and stopping the first-parent walk when it hits a commit in `planStopSet`.
4. Skipping the pack push entirely when a branch tip is already in `planStopSet` — a *subsumed* branch. The only command emitted is a single ref-create to point the target ref at the tip, optionally combined with a temp-ref delete if a previous interrupted run left one behind.

Falls back to the per-branch behavior described above when HEAD is not advertised, or when the trunk ref is filtered out by `--branch` / `--map`.

### Why not probe (the previous design)

The previous implementation did full `FetchPack` round-trips per probe candidate to measure actual pack sizes. For linux/master (75k commits) this required 13+ fetch-and-discard cycles, downloading gigabytes of throwaway data and taking minutes before any real push started. The estimate approach reduces planning to one commit-graph fetch (~20 seconds) plus arithmetic.
38 unmodified lines

This avoids cases where a tag points at an object graph that is not yet fully present on target.

V1 batching should support:

- branch refs only

Tag batching can be added later.
Branch batches push first; create-only tags are pushed after all branch batches complete. Tag retargeting is not supported in the batched path.

## Restart and Recovery

9 unmodified lines

## Safety Model

V1 batching should remain strict.

Allow only:
Batching remains strict. It allows only:

- empty managed target refs
- branch-only bootstrap
54 unmodified lines

This is still likely worthwhile for very large initial migrations because it changes a single huge risky operation into several bounded ones.

## Recommended Phases
## Current Behavior

Phase A:
Batched bootstrap is invoked via `git-sync bootstrap --target-max-pack-bytes`.

- batch branch-only bootstrap
- no tags
- temp refs required
- no resume
- manual cleanup if interrupted
It:

Progress:

- implemented via `git-sync bootstrap --target-max-pack-bytes`
- currently requires source-side protocol v2 with fetch filter support
- batches branch refs only, with create-only tags pushed after all branch batches complete
- requires source-side protocol v2 with fetch filter support
- uses temporary target refs under `refs/gitsync/bootstrap/heads/`
- resumes from an existing temp ref when that temp ref matches a planned checkpoint
- exercised by `TestBootstrap_GitHTTPBackendBatchedBranch`
- validated against `torvalds/linux` as a large-source manual stress path
- plans the source's trunk first (when its HEAD symref is advertised) and reuses its commit-graph reachability to short-circuit later branches' walks and skip pack pushes for subsumed branches
- is exercised by `TestBootstrap_GitHTTPBackendBatchedBranch` and validated against `torvalds/linux` as a large-source manual stress path

Operator guidance:

1 unmodified line

- use batching when a single large bootstrap push is too risky, too large, or fails on the target side
- start with `--target-max-pack-bytes 536870912` and adjust upward only if the target has enough headroom

Phase B:

- add resume from existing temp refs
- add better progress reporting
- add batch-size estimation metrics

Phase C:

- consider tag creation after successful branch completion
- consider whether per-ref or per-branch parallelism is worth it

Progress:

- create-only tags are now pushed after successful branch batches complete

Phase D:

- only then consider using similar checkpoint batching ideas for non-empty target incremental relay
The current implementation does not parallelize across branches and does not extend the same checkpoint-batching idea to non-empty target incremental relay.
```

Mdocs/bootstrap-batching.md+25/-40

```
22 unmodified lines

23
24
25
26
26
27
28
28
29
30
31
60 unmodified lines

92
93
94
95
95
96
97
98
16 unmodified lines

115
116
117
118
118
119
120
121
2 unmodified lines

124
125
126
127
127
128
129
129
130
131
132
133
134
135
131
132
133
134
135
136
137
138
137
139
140
139
140
141
142
143
141
142
145
143
144
145
146
147
148
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
149

22 unmodified lines

- there is no need for fast-forward checks
- the main cost is moving a large pack from source to target efficiently

## V1 Scope
## Scope

`bootstrap` should be intentionally narrow:
`bootstrap` is intentionally narrow:

- create-only
- fail if any managed target ref already exists
60 unmodified lines

- push still depends on target `receive-pack` behavior and capabilities
- if a relay-safe path cannot be used, `bootstrap` should fail and tell the user to use `sync`

V1 should stay strict rather than trying to be clever.
`bootstrap` stays strict rather than trying to be clever.

## Implementation Notes

16 unmodified lines

## Failure Rules

V1 should fail when:
`bootstrap` fails when:

- any managed target ref already exists
- no source refs matched
2 unmodified lines

The error should explicitly recommend normal `sync` when the repository is no longer in bootstrap shape.

## Follow-Up Steps
## Current Behavior

Phase 1:
Bootstrap supports:

- implement `bootstrap` for create-only branch refs
- support optional tag creation
- add JSON and stats output
- add in-process integration tests
- add `git-http-backend` integration coverage for empty-target bootstrap
- create-only branch refs
- optional tag creation (non-batched path and after successful branch batches in the batched path)
- explicit mapped refs (`--map src:dst`)
- JSON and `--stats` output
- `--max-pack-bytes` as a safety threshold for the streamed source pack
- `--target-max-pack-bytes` for batched branch-only bootstrap on very large initial syncs
- in-process integration coverage and `git-http-backend` integration coverage

Progress:
`sync` auto-selects the bootstrap relay path when all managed target refs are absent and the run matches bootstrap semantics. `plan` surfaces a bootstrap suggestion for the same target shape.

- `bootstrap` is implemented
- optional tag creation is supported on the non-batched path
- JSON and stats output are supported
- in-process integration coverage exists
- `git-http-backend` integration coverage exists
The batched bootstrap path is intentionally narrow:

Phase 2:
- requires source-side protocol v2 with fetch filters
- batches branch refs, then optionally creates tags after the branch batches complete
- resumes from an existing temp ref when that temp ref matches a planned checkpoint
- uses temporary target refs under `refs/gitsync/bootstrap/heads/`
- treat as an advanced large-repo fallback when one-shot bootstrap is too risky or fails under target-side unpack/index pressure

- allow relay-safe create-only runs with explicit mapped refs
- add better operator output for large initial transfers
- add safety thresholds for advertised/fetched bytes

Progress:

- explicit mapped refs are supported
- `--max-pack-bytes` provides a first safety threshold for the streamed source pack during bootstrap
- `--target-max-pack-bytes` now enables a Phase A batched branch-only bootstrap mode for large initial syncs

Phase 3:

- investigate hybrid behavior: relay when the target is empty, otherwise fail fast into normal `sync`
- investigate whether target capability combinations require alternate pack handling
- measure source-to-target pack relay memory and CPU against current `sync`

Progress:

- `sync` now auto-selects the bootstrap relay path when all managed target refs are absent and the run matches bootstrap semantics
- dry-run `plan` surfaces a bootstrap suggestion for the same target shape

Batching note:

- the current batched bootstrap path is intentionally narrow
- it requires source-side protocol v2 with fetch filters
- it batches branch refs and then optionally creates tags after the branch batches complete
- it resumes from an existing temp ref when that temp ref matches a planned checkpoint
- it uses temporary target refs under `refs/gitsync/bootstrap/heads/`
- it should be treated as an advanced large-repo fallback when one-shot bootstrap is too risky or fails on target-side unpack/index pressure

Phase 4:

- consider a more advanced incremental relay mode for non-empty targets
- only pursue this if large migration workflows become important enough to justify the added protocol complexity

Progress:

- there is now a narrow incremental relay path in `sync`
- it now covers multi-branch fast-forward branch-only updates
- it now also covers branch-to-branch mappings
- it now also covers create-only tags
- tag retargets, deletes, force, and prune still use the normal path
Outside bootstrap, `sync` also has a narrow incremental relay path for non-empty targets that covers multi-branch fast-forward updates, branch-to-branch mappings, and create-only tags. Tag retargets, deletes, force, and prune still use the materialized path. See [incremental-relay.md](incremental-relay.md).
```

Mdocs/bootstrap.md+21/-60

```
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63

# Incremental Relay

`sync` has a narrow relay fast path for safe incremental updates. When eligible, it streams a fetched source pack directly into target `receive-pack` instead of decoding the object graph into the local in-memory store and re-encoding a push pack. This keeps the in-memory cost near zero for the common case where a sync run only needs to forward a small amount of new history.

This document describes when the fast path applies, what it covers, and when `sync` falls back to the materialized path.

## Eligibility

The incremental relay path is selected only when **all** of the following hold:

- no `--force`
- no `--prune`
- no managed-ref deletes
- no tag retargeting (creating a new tag at a new tip is allowed; moving an existing tag is not)
- target advertises `no-thin` on `receive-pack`

`no-thin` is the load-bearing capability. The relayed pack must be self-contained because git-sync's source fetch never requests the `thin-pack` capability, so the target can apply the pack directly without resolving deltas against existing target objects. See [protocol.md](protocol.md) for the framing details.

## What the fast path covers

When the eligibility conditions hold, the relay path covers:

- multi-branch fast-forward branch updates
- branch-to-branch ref mappings (`--map src:dst`)
- create-only tag pushes that fit alongside the branch updates

In other words: the everyday "mirror these branches and any new tags forward" case is fully covered.

## What still falls back to materialized

The materialized path (decode source objects into the local store, plan the push set, encode a target pack) still handles:

- `--force` and any non-fast-forward update
- `--prune` and managed-ref deletes
- tag retargeting (an existing tag pointing at a new object)
- runs against targets that don't advertise `no-thin`

The materialized path is bounded by `--materialized-max-objects` as a safety guardrail. See [architecture.md](architecture.md#memory-assumptions) for the memory model.

## Relationship with `bootstrap`

`sync` also auto-selects the bootstrap relay path when all managed target refs are absent and the run otherwise matches bootstrap semantics. That is a separate code path (in `internal/strategy/bootstrap`) and is documented in [bootstrap.md](bootstrap.md). The incremental relay path discussed here is for non-empty targets where the existing tips can be advertised as `have` lines during source fetch.

The decision flow inside `sync` is:

1. If target has none of the managed refs and the run is bootstrap-compatible → bootstrap relay path
2. Else if all eligibility conditions for incremental relay hold → incremental relay path
3. Else → materialized fallback (bounded by `--materialized-max-objects`)

## Why this matters

For repeat sync jobs against an actively used mirror, the common case is "a few new commits on a couple of branches plus maybe a new tag." Without the incremental relay path, every such run would decode the fetched objects into the in-memory store, then re-encode them into a push pack. With relay, the runner forwards a self-contained source pack directly, and the per-run memory and CPU cost stays roughly proportional to the size of the new history rather than the size of the touched repos.

## Implementation

The incremental relay strategy lives in `internal/strategy/incremental`. The shared relay framing, sideband stripping, and PACK header handling live in `internal/gitproto`. See [protocol.md](protocol.md) for protocol-level details and [architecture.md](architecture.md) for where this fits in the overall package layout.
