Merge pull request #13 from entireio/soph/docs-update-for-oss-release · Entire

Home

Log in

Merge pull request #13 from entireio/soph/docs-update-for-oss-release

1a41b08→main·

Soph·2mo ago·12 files·+893 added/-1,017 removed

README.md update and other docs

Changes

12

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94

# Code of Conduct

We're committed to providing a welcoming, respectful, and harassment-free environment for everyone, regardless of age, body size, visible or invisible disability, ethnicity, sex characteristics, gender identity and expression, level of experience, education, socio-economic status, nationality, personal appearance, race, caste, color, religion, or sexual identity and orientation.

## Scope

This Code of Conduct applies to all Entire community spaces, including:

- **GitHub** - repositories, issues, pull requests, discussions, and comments
- **Discord** - all channels in the Entire workspace
- **Events** - meetups, conferences, and online gatherings hosted by Entire
- **Public representation** - acting as a representative of Entire in public spaces, including social media, forums, and conferences

This Code of Conduct also applies when an individual is officially representing the community in public spaces.

## Expected Behavior

- **Be respectful** - Treat everyone professionally, listen actively, and be mindful of your words
- **Be inclusive** - Use inclusive language and make space for everyone to contribute
- **Be collaborative** - Help each other, share knowledge, and celebrate wins together
- **Be accountable** - Own your mistakes, follow through on commitments, and take responsibility for your impact
- **Be empathetic** - Try to understand different perspectives and experiences
- **Give and accept constructive feedback gracefully**

## Unacceptable Behavior

This includes but isn't limited to:

- Harassment, discrimination, or offensive comments related to personal characteristics
- Personal attacks, trolling, insulting or derogatory comments, or deliberate intimidation
- Unwelcome sexual attention, advances, or imagery
- Sharing private information (such as physical or email addresses) without explicit consent
- Sustained disruption of discussions or events
- Advocating for or encouraging any of the above behavior
- Conduct that could reasonably be considered inappropriate in a professional setting

## Reporting

To report a Code of Conduct violation, contact **[conduct@entire.io](mailto:conduct@entire.io)**.

For security vulnerabilities, please email **[security@entire.io](mailto:security@entire.io)** instead. See our [Security Policy](SECURITY.md).

### Confidentiality

**All reports will be kept confidential.** We will not share your identity or the details of your report with anyone outside the enforcement team without your consent, except as required by law or to protect safety.

### What to Include

- Your contact information (so we can follow up)
- Names of those involved (or identifying information)
- Description of the behavior and when/where it occurred
- Any additional context or evidence (screenshots, links)
- Whether you would like to remain anonymous to the reported party

## Enforcement

Community leaders are responsible for clarifying and enforcing standards of acceptable behavior. All complaints will be reviewed and investigated promptly and fairly.

Violations will be addressed according to their severity and frequency:

#### 1. Correction
**For:** First-time minor violations or misunderstandings

**Action:** A private, written notice explaining the violation and why the behavior was inappropriate. A public apology may be requested.

---

#### 2. Warning
**For:** A single incident or pattern of minor violations

**Action:** A formal warning with consequences for continued behavior. This includes no interaction with the people involved for a specified period. Violating these terms may lead to a temporary or permanent ban.

---

#### 3. Temporary Ban
**For:** Serious violations or sustained inappropriate behavior

**Action:** A temporary ban from all community interaction and public communication for a specified period (typically 30-90 days). No public or private interaction with the community is permitted during this time.

---

#### 4. Permanent Ban
**For:** Demonstrating a pattern of violations, severe harassment, or aggression toward individuals or groups

**Action:** Permanent removal from all community spaces. This decision is final.

---

## Attribution

This Code of Conduct is adapted from the [Contributor Covenant](https://www.contributor-covenant.org/version/2/1/code_of_conduct/), version 2.1.

For answers to common questions about the Contributor Covenant, see the [FAQ](https://www.contributor-covenant.org/faq).

ACODE_OF_CONDUCT.md+94

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286

# Contributing to git-sync

Thank you for your interest in contributing to Entire! We welcome contributions from everyone.

Please read our [Code of Conduct](CODE_OF_CONDUCT.md) before participating.

> **New to Entire?** See the [README](README.md) for setup and usage documentation.

---

## Before You Code: Discuss First

The fastest way to get a contribution merged is to align with maintainers before writing code. Please **open an issue first** using our [issue templates](https://github.com/entireio/gitsync/issues/new/choose) and wait for maintainer feedback before starting implementation.

### Contribution Workflow

1. **Open an issue** describing the problem or feature
2. **Wait for maintainer feedback** -- we may have relevant context or plans
3. **Get approval** before starting implementation
4. **Submit your PR** referencing the approved issue
5. **Address all feedback** including automated Copilot comments
6. **Maintainer review and merge**

---

## First-Time Contributors

New to the project? Welcome! Here's how to get started:

### Good First Issues

We recommend starting with:
- **Documentation improvements** - Fix typos, clarify explanations, add examples
- **Test contributions** - Add test cases, improve coverage
- **Small bug fixes** - Issues labeled `good-first-issue`

---

## Submitting Issues

All feature requests, bug reports, and general issues should be submitted through [GitHub Issues](https://github.com/entireio/gitsync/issues). Please search for existing issues before opening a new one.

For security-related issues, see the Security section below.

---

## Security

If you discover a security vulnerability, **do not report it through GitHub Issues**. Instead, please follow the instructions in our [SECURITY.md](SECURITY.md) file for responsible disclosure. All security reports are kept confidential as described in SECURITY.md.

---

## Contributions & Communication

Contributions and communications are expected to occur through:

- [GitHub Issues](https://github.com/entireio/gitsync/issues) - Bug reports and feature requests
- [Discord](https://discord.gg/jZJs3Tue4S) - Questions, general conversation, and real-time support

Please represent the project and community respectfully in all public and private interactions.

## How to Contribute

There are many ways to contribute:

- **Feature requests** - Open a [GitHub Issue](https://github.com/entireio/gitsync/issues) to discuss your idea
- **Bug reports** - Report issues via [GitHub Issues](https://github.com/entireio/gitsync/issues) (see [Reporting Bugs](#reporting-bugs))
- **Code contributions** - Fix bugs, add features, improve tests
- **Documentation** - Improve guides, fix typos, add examples
- **Community** - Help others, answer questions, share knowledge

## Reporting Bugs

Good bug reports help us fix issues quickly. When reporting a bug, please include:

### Required Information

1. **git-sync commit** - `git rev-parse HEAD` of the build you used (or release tag/version if applicable)
2. **Operating system**
3. **Go version** - run `go version`

### What to Include

Please answer these questions in your bug report:

1. **What did you do?** - Include the exact commands you ran
2. **What did you expect to happen?**
3. **What actually happened?** - Include the full error message or unexpected output
4. **Can you reproduce it?** - Does it happen every time or intermittently?
5. **Any additional context?** - Logs, screenshots, or related issues

---

## Local Setup

### Prerequisites

- **mise** - Task runner and version manager. Install with `curl https://mise.run | sh`

### Clone and Install

```bash
git clone https://github.com/entireio/gitsync.git
cd gitsync

# Trust the mise configuration (required on first setup)
mise trust

# Install dependencies (mise will install the correct Go version)
mise install

# Download Go modules
go mod download

# Build the CLI
mise run build

# Verify setup by running tests
mise run test
```

> See [docs/architecture.md](docs/architecture.md) for the architecture and package layout.

---

## Making Changes

1. **Create a branch** for your changes:
   ```bash
   git checkout -b feature/your-feature-name
   ```

2. **Make your changes** - follow the [Code Style](#code-style) guidelines

3. **Test your changes** - see [Testing](#testing)

4. **Commit** with clear, descriptive messages:
   ```bash
   git commit -m "Add feature: description of what you added"
   ```

---

## Code Style

Follow standard Go idioms and conventions.

### Key Points

- **Error handling**: Handle all errors explicitly - don't leave them unchecked
- **Formatting**: Code must pass `gofmt` (run `mise run fmt`)
- **Linting**: Code must pass `golangci-lint` (run `mise run lint`)
- **Naming**: Use meaningful, descriptive names following Go conventions
- **Public API**: `entire.io/entire/gitsync` is the stable embedding surface. Additions there should be reviewed carefully. `entire.io/entire/gitsync/unstable` is for advanced controls and may change.

---

## Testing

> See [docs/testing.md](docs/testing.md) for the full set of test suites and integration coverage.

```bash
# Default suite - always run before committing
mise run test

# With race detection
mise run test:ci

# Optional end-to-end against the system git-http-backend
mise run test:git-http-backend

# Optional live linux bootstrap smokes
mise run test:linux-smoke
mise run test:linux-smoke:batched
```

---

## Submitting a Pull Request

### Before You Submit

- **Related issue exists and is approved** -- Your PR references an issue where a maintainer has acknowledged the approach. (Exceptions: documentation fixes, typo corrections, and `good-first-issue` items.)
- **Linting passes** -- Run `mise run lint` (includes golangci-lint, gofmt, gomod, shellcheck)
- **Tests pass** -- Run `mise run test` to verify your changes
- **Tests included** -- New Go code and functionality should have accompanying tests
- **Entire checkpoint trailers included** -- See [Using Entire While Contributing](#using-entire-while-contributing) below

PRs that skip these steps are likely to be closed without merge.

### Submitting

1. **Push** your branch to your fork
2. **Open a PR** against the `main` branch
3. **Describe your changes** -- Link the related issue, summarize what changed and what testing you did
4. **Address Copilot feedback** -- See [Responding to Automated Review](#responding-to-automated-review)
5. **Wait for maintainer review**

---

## Responding to Automated Review

Copilot agent reviews every PR and provides feedback on code quality, potential bugs, and project conventions.

**Read and respond to every Copilot comment.** PRs with unaddressed Copilot feedback will not move to maintainer review.

- **Fixed** -- Push a commit addressing the issue.
- **Disagree** -- Reply explaining your reasoning. The Copilot isn't always right.
- **Question** -- Ask for clarification. We're happy to help.

Addressing Copilot feedback upfront is the fastest path to maintainer review.

---

## Using Entire While Contributing

We use Entire on Entire. When contributing, install the Entire CLI and let it capture your coding sessions -- this gives us valuable dogfooding data and helps improve the tool.

### Setup

Install the latest version of the Entire CLI (see [installation docs](https://docs.entire.io/cli/installation)) and verify with `entire version`. Entire is already configured in this repository, so there's no need to run `entire enable`.

### Checkpoint Trailers

All commits should include `Entire-Checkpoint` trailers from your sessions. These are added automatically by the `prepare-commit-msg` hook when Entire is enabled. The trailers link your commits to session metadata on the `entire/checkpoints/v1` branch.

### Sessions Branch

When you push your PR branch, Entire can automatically push the `entire/checkpoints/v1` branch alongside it (if `push_sessions` is enabled in your settings). Include this in your PR so maintainers can review the session context behind your changes.

---

## Troubleshooting

### Common Setup Issues

**`go mod download` fails with timeout**
```bash
# Try using direct mode
GOPROXY=direct go mod download
```

**`mise install` fails**
```bash
# Ensure mise is properly installed
curl https://mise.run | sh

# Reload your shell
source ~/.zshrc  # or ~/.bashrc
```

**Binary not updating after rebuild**
```bash
# Check which binary is being used
which git-sync
type -a git-sync

# You may have multiple installations - update the correct path
```

---

## Community

Join the Entire community:

- **Discord** - [Join our server][discord] for discussions and support

[discord]: https://discord.gg/jZJs3Tue4S

---

## Additional Resources

- [README](README.md) - Setup and usage documentation
- [docs/architecture.md](docs/architecture.md) - Architecture and package layout
- [docs/testing.md](docs/testing.md) - Test suites and integration coverage
- [Code of Conduct](CODE_OF_CONDUCT.md) - Community guidelines
- [Security Policy](SECURITY.md) - Reporting security vulnerabilities

---

Thank you for contributing!

ACONTRIBUTING.md+286

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21

MIT License

Copyright (c) 2026 Entire Inc.

Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:

The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.

THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.

ALICENSE+21

9 unmodified lines

10
11
12
13
13
14
15
15
16
17
18
19
20
17
18
22
19
20
24
25
26
27
28
21
22
30
23
24
25
32
26
27
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
28
29
30
31
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
32
33
34
35
5 unmodified lines

41
42
43
125
44
45
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
46
47
48
49
4 unmodified lines

54
55
56
157
57
58
59
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
60
61
62
63
1 unmodified line

65
66
67
210
68
69
70
213
71
72
73
74
75
76
77
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
78
79
253
254
255
256
257
80
81
82
83
34 unmodified lines

118
119
120
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
121
122
346
123
124
348
125
126
127
128
352
353
354
355
129
130
131
132
23 unmodified lines

156
157
158
385
386
387
388
389
390
391
392
393
394
395
159
160
161
162
163
164
398
165
166
167
168
5 unmodified lines

174
175
176
410
177
178
412
179
180
181
182
414
183
184
416
417
418
185

9 unmodified lines

## Why This Exists

Git already has pieces of this problem, but not this exact tool shape.
Mirroring Git data between remotes usually means a local mirror clone followed by a mirror push. That's fine for small repos but turns a remote-to-remote operation into a local storage problem at scale, and shell glue around `git fetch` / `git push` tends to skip planning and structured output.

What usually exists today:
`git-sync` fills that gap. It streams source packs directly into target `receive-pack` when it can, plans every action before pushing, and emits typed JSON for automation.

- a full local `git clone --mirror` followed by `git push --mirror`
- host-specific import or migration features
- CI jobs or shell scripts that glue fetch and push steps together
- one-off migration tooling tied to a specific platform
For when to use it (and when not), see [docs/architecture.md](docs/architecture.md).

What those approaches usually do not give you:
## Commands

- direct remote-to-remote relay behavior
- a small standalone CLI with explicit sync semantics
- front-loaded validation and planning
- machine-readable output for automation
- one tool that covers empty-target bootstrap, normal sync, and large-repo bootstrap fallback
The main commands are:

That is the gap `git-sync` is trying to fill.
- `git-sync sync`: mirror source refs into the target
- `git-sync replicate`: overwrite target refs to match source via relay, and fail rather than materialize locally

The main value is operational:
`sync` automatically bootstraps an empty target, so the same command covers initial seeding and ongoing sync. To preview what would happen without pushing, run `git-sync plan` — it takes the same flags as `sync`, and `--mode replicate` previews a `replicate` run.

- avoid requiring a full local mirror checkout just to move refs between remotes
- make initial seeding of large repositories cheaper and more predictable
- keep incremental sync behavior explicit and safe
- give operators and automation a stable way to inspect, plan, execute, and benchmark the same workflows

This is especially useful when:

- the target is a new hosted Git service or internal Git endpoint
- bootstrap size matters more than local developer ergonomics
- you want a repeatable machine-oriented sync primitive rather than an ad hoc migration script
- you need clearer control over mapping, pruning, force rules, and relay behavior than generic shell glue usually provides

Compared to a service that keeps persistent local clones, `git-sync` is the better fit when:

- relay is common enough that streaming source-to-target is the normal case
- avoiding persistent local repo storage is an operational advantage
- remote-to-remote efficiency matters more than full local Git generality

If you need arbitrary complex reconciliation through one always-warm local full-state model, a local-clone service is still the more general tool.

The command surface is:

- `git-sync probe`: inspect a source remote, and optionally a target remote
- `git-sync fetch`: exercise source-side fetch negotiation without pushing
- `git-sync bootstrap`: seed an empty target with create-only relay behavior
- `git-sync plan`: compute source-to-target ref actions without pushing, with `--mode sync|replicate`
- `git-sync sync`: execute the planned changes against the target
- `git-sync replicate`: execute source-authoritative relay-only replication against the target
- `git-sync-bench`: run repeatable benchmark scenarios against fresh empty targets
Additional commands (`bootstrap`, `probe`, `fetch`) and advanced flags are available through `git-sync --help` and the unstable library surface. They are not part of the recommended public surface.

## Library API

`git-sync` now has a two-tier Go API:

- `gitsync`
  - stable embedding surface for queue workers and other external callers
  - typed `Probe`, `Plan`, `Sync`, and `Replicate` requests/results
  - injected auth and HTTP client support
- `unstable`
  - explicitly non-stable surface for first-party tooling and advanced controls
  - includes `Bootstrap`, `Fetch`, batching and measurement knobs, and CLI-oriented execution options

If you are embedding `git-sync` outside this repo, prefer `gitsync`. The CLI and benchmark command use `unstable` because they still need direct access to advanced engine controls that are intentionally not part of the stable API.

The stable `gitsync` results are shaped for workers:

- `Refs`
  - per-ref outcomes
- `Counts`
  - aggregate applied/skipped/blocked/deleted counts
- `Execution`
  - execution mode, protocol, relay summary, and batch summary

See [docs/embedding.md](docs/embedding.md) for worker-oriented guidance.

## Current scope

- Smart HTTP only
- No local working tree
- Branch mirroring by default
- Optional tag mirroring with `--tags`
- Optional exact ref mapping with `--map`
- Fast-forward safety by default
- Optional forced retargeting with `--force`
- Optional source-authoritative relay-only replication with `replicate` / `plan --mode replicate`
- Optional managed-ref deletion with `--prune`
- Optional transfer stats output with `--stats`
- Optional machine-readable output with `--json`
- Optional source-side Git protocol v2 for `ls-refs` and `fetch`

## Limitations

- Push still uses the existing v1-style `receive-pack` path.
- Protocol v2 support currently covers source discovery and source fetch only.
- `--protocol auto` tries source-side v2 first and falls back to v1.
- `--protocol v2` requires the source remote to negotiate v2.
- Ref mapping is explicit, not wildcard-based.
- Only smart HTTP remotes are supported.
- Objects are kept in memory for the duration of the run.
- Non-relay materialized syncs are bounded by `--materialized-max-objects`, an object-count guardrail for the in-memory fallback path.
`git-sync` is also a Go library. Use `entire.io/entire/gitsync` for the stable embedding surface (`Probe`, `Plan`, `Sync`, `Replicate`, typed results, auth and HTTP injection). `entire.io/entire/gitsync/unstable` exposes advanced controls (`Bootstrap`, `Fetch`, batching knobs, heap measurement) and is not stable.

## Quick Start

5 unmodified lines

https://github.com/target-org/target-repo.git
```

## Commands
## Examples

Plan a sync without pushing anything:

```bash
go run ./cmd/git-sync plan \
  --stats \
  https://github.com/source-org/source-repo.git \
  https://github.com/target-org/target-repo.git
```

Plan a source-authoritative replication without pushing anything:

```bash
go run ./cmd/git-sync plan \
  --mode replicate \
  --stats \
  https://github.com/source-org/source-repo.git \
  https://github.com/target-org/target-repo.git
```

Execute relay-only replication that overwrites differing managed refs and fails instead of materializing locally:
Run a replication that overwrites differing target refs, and fail instead of falling back to local materialization:

```bash
go run ./cmd/git-sync replicate \
4 unmodified lines

If `replicate` cannot use relay against the target, it fails and tells you to rerun with `sync`.

Bootstrap an empty target without using the normal local object-store sync path:
For very large initial migrations, add `--target-max-pack-bytes` to split the initial pack into multiple smaller batches. The same flag works on `sync`, since `sync` auto-bootstraps on empty targets:

```bash
go run ./cmd/git-sync bootstrap \
  --stats \
https://github.com/source-org/source-repo.git \
https://github.com/target-org/target-repo.git
```

Add `--max-pack-bytes` to abort bootstrap if the streamed source pack grows past a safety threshold:

```bash
go run ./cmd/git-sync bootstrap \
  --max-pack-bytes 104857600 \
<source-url> \
<target-url>
```

Add `--target-max-pack-bytes` to split large branch bootstraps into multiple relay batches with temporary refs:

```bash
go run ./cmd/git-sync bootstrap \
  --target-max-pack-bytes 1073741824 \
<source-url> \
<target-url>
```

Current batching scope is intentionally narrow:

- protocol v2 only
- branch refs are batched
- optional create-only tags are pushed after branch batches complete
- temporary refs under `refs/gitsync/bootstrap/heads/`
- resume from existing temp refs is supported when they match a planned checkpoint

This mode is intended as an advanced large-repo fallback, not the default bootstrap path. Use plain `bootstrap` first when a single streamed initial sync is acceptable.

A practical starting point is:

- `--target-max-pack-bytes 536870912` for a conservative `512 MiB` target-side batch size
- `--target-max-pack-bytes 1073741824` when you want fewer, larger batches and the target has more headroom

For example:

```bash
go run ./cmd/git-sync bootstrap \
go run ./cmd/git-sync sync \
  --target-max-pack-bytes 536870912 \
  --protocol v2 \
  -v \
1 unmodified line

<target-url>
```

Add `--measure-memory` to `bootstrap`, `sync`, `plan`, `probe`, or `fetch` to sample elapsed time and Go heap usage:
Add `--measure-memory` to any command to sample elapsed time and Go heap usage:

```bash
go run ./cmd/git-sync bootstrap \
go run ./cmd/git-sync sync \
  --measure-memory \
  --json \
<source-url> \
<target-url>
```

That is useful for one-off measurements on the same fixture or test repo.

## Benchmarking

For repeated benchmark runs, prefer the dedicated benchmark command instead of manually wrapping `git-sync` invocations:

```bash
go run ./cmd/git-sync-bench \
  --scenario bootstrap \
  --source-url /tmp/git-sync-bench/kubernetes.git \
  --repeat 3 \
  --target-max-pack-bytes 104857600 \
  --stats \
  --json
```

`git-sync-bench` creates a fresh bare target repository for each run, executes the selected scenario in-process, and reports:

- per-run wall-clock time
- per-run `syncer.Result`
- aggregate min/avg/max wall time
- aggregate internal elapsed and heap metrics from `--measure-memory`
- relay modes observed across successful runs

If `--source-url` is a local path, it is converted to `file://...` automatically. The current scenarios are:

- `--scenario bootstrap`
- `--scenario sync`

For large-repo measurements, use a local bare mirror as the source so the benchmark reflects `git-sync` behavior rather than internet variance. See [docs/benchmarking.md](docs/benchmarking.md) for details.

## Sync Behavior

When `sync` sees that all managed target refs are absent and the run is compatible with bootstrap semantics, it automatically uses the bootstrap relay path instead of the normal decode-and-repack sync path.

`sync` also uses a narrow incremental relay path for fast-forward branch updates and tag creation when there is no prune/delete, no force, and the target does not advertise `no-thin`. This now includes multi-branch batches, branch-to-branch mappings, and create-only tags. Tag retargeting and other more complex updates still fall back to the normal local decode-and-repack path.

If `sync` falls back to the materialized path, `--materialized-max-objects` sets an explicit object-count safety bound for the in-memory object set. It is a conservative guardrail, not a precise heap-size limit.
`sync` picks the bootstrap relay path automatically when the target is empty. For non-empty targets, safe fast-forward updates also use a relay path that streams the source pack directly into target `receive-pack` without local materialization. Anything not relay-eligible (force, prune, deletes, tag retargets) falls back to a materialized path bounded by `--materialized-max-objects`.

Sync specific branches:

34 unmodified lines

<target-url>
```

Probe a source remote without pushing anything:

```bash
go run ./cmd/git-sync probe \
  --stats \
  --tags \
  --protocol auto \
  <source-url>
```

Probe both source and target remotes to inspect source fetch capabilities and target `receive-pack` capabilities:

```bash
go run ./cmd/git-sync probe \
  --stats \
  <source-url> \
  <target-url>
```

Fetch from a source remote into memory without pushing anywhere:

```bash
go run ./cmd/git-sync fetch \
  --stats \
  --protocol auto \
  --branch main \
  <source-url>
```

Advertise an existing source ref as a synthetic `have` to exercise incremental negotiation:

```bash
go run ./cmd/git-sync fetch \
  --stats \
  --protocol auto \
  --branch main \
  --have-ref main \
  <source-url>
```

Dry run:

```bash
go run ./cmd/git-sync plan --stats <source-url> <target-url>
```

## JSON Output

Add `--json` to `probe`, `fetch`, `bootstrap`, `plan`, or `sync` to emit machine-readable output instead of the default text format.
Add `--json` to any command to emit machine-readable output instead of the default text format.

The JSON interface is intentionally stable:
The JSON interface is stable:

- keys use `camelCase`
- refs and hashes are serialized as strings, not raw byte arrays
- `probe` returns top-level keys such as `sourceUrl`, `targetUrl`, `protocol`, `refPrefixes`, `sourceCapabilities`, `targetCapabilities`, `refs`, and `stats`
- `fetch` returns top-level keys such as `sourceUrl`, `protocol`, `wants`, `haves`, `fetchedObjects`, and `stats`
- `bootstrap`, `plan`, and `sync` return top-level keys such as `plans`, `pushed`, `skipped`, `blocked`, `deleted`, `dryRun`, `protocol`, and `stats`
- `bootstrap`, `plan`, and `sync` also expose `relay`, `relayMode`, `relayReason`, `batching`, `batchCount`, `plannedBatchCount`, and `tempRefs`
- top-level keys include `plans`, `pushed`, `skipped`, `blocked`, `deleted`, `dryRun`, `protocol`, and `stats`, plus `relay`, `relayMode`, `relayReason`, `batching`, `batchCount`, `plannedBatchCount`, and `tempRefs`
- each item in `plans` includes stable string fields such as `branch`, `sourceRef`, `targetRef`, `sourceHash`, `targetHash`, `kind`, `action`, and `reason`

## Auth
23 unmodified lines

## Protocol Notes

- Source refs are listed with `GET /info/refs?service=git-upload-pack`.
- When the source supports it, the client can negotiate protocol v2 with `Git-Protocol: version=2`, then use `ls-refs` and `fetch`.
- Target refs are listed with `GET /info/refs?service=git-receive-pack`.
- The source fetch advertises current target tip hashes as `have`, so reruns download less when source and target already share history.
- Target push stays on the current `receive-pack` path.
- If a target ref does not exist, it is created.
- If a target ref already matches the source, it is skipped.
- Branches are updated only when the target tip is an ancestor of the source tip, unless `--force` is set.
- Tags are immutable by default. Retargeting an existing tag requires `--force`.
- If `--prune` is set, managed target refs that are absent on source are deleted.
- `plan` never pushes. If `sync` finds blocked refs, it exits non-zero before pushing anything.
- Source-side discovery and fetch can use protocol v2 when supported. Push stays on the existing v1 `receive-pack` path. `--protocol auto` tries v2 first and falls back to v1. `--protocol v2` requires the source to negotiate v2.
- Source fetch advertises current target tip hashes as `have`, so reruns download less when source and target already share history.
- Branches are updated only when the target tip is an ancestor of the source tip, unless `--force` is set. Tags are immutable by default. Retargeting an existing tag requires `--force`. With `--prune`, managed target refs that are absent on source are deleted.
- If `sync` finds blocked refs, it exits non-zero before pushing anything.
- `--stats` adds per-service request, byte, want, have, and command counters to the output.

Push still uses the current low-level `receive-pack` path. Protocol v2 is used where it materially improves this tool: source-side ref discovery and source-side object download.
For the deeper protocol-level walkthrough (smart HTTP, pkt-line, capability negotiation, sideband stripping, relay framing), see [docs/protocol.md](docs/protocol.md).

## Testing

5 unmodified lines

Extended and environment-specific test instructions are in [docs/testing.md](docs/testing.md).

## Design Notes
## Documentation

`bootstrap` is the dedicated path for large initial syncs into an empty target. The goal is to relay a fetched source pack directly into target `receive-pack` instead of decoding the full object graph into local memory first.
- [docs/architecture.md](docs/architecture.md) — product rationale, package layout, operation modes vs transfer modes, memory model
- [docs/protocol.md](docs/protocol.md) — smart HTTP, pkt-line, capability negotiation, sideband, relay framing
- [docs/testing.md](docs/testing.md) — test suites and integration coverage

Current architectural summary and package boundaries are in [docs/architecture.md](docs/architecture.md).
## Contributing

The design note is in [docs/bootstrap.md](docs/bootstrap.md).

For very large single-branch repositories, there is also a batching design and initial implementation note in [docs/bootstrap-batching.md](docs/bootstrap-batching.md).
See [CONTRIBUTING.md](CONTRIBUTING.md), [SECURITY.md](SECURITY.md), and [CODE_OF_CONDUCT.md](CODE_OF_CONDUCT.md).

MREADME.md+31/-264

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68

# Security Policy

We take security seriously at Entire. We appreciate your efforts to responsibly disclose vulnerabilities and will make every effort to acknowledge your contributions.

## Reporting a Vulnerability

**Please do not report security vulnerabilities through public GitHub issues.**

Instead, please send security-related reports to **[security@entire.io](mailto:security@entire.io)**.

### What to Include

When reporting a vulnerability, please include:

1. **Description** - A clear description of the vulnerability
2. **Impact** - What an attacker could achieve by exploiting this issue
3. **Steps to reproduce** - Detailed steps to reproduce the vulnerability
4. **Affected versions** - Which versions of `git-sync` are affected (if known)
5. **Suggested fix** - If you have ideas on how to fix it (optional)

### What to Expect

- **Acknowledgment** - We will acknowledge receipt of your report within 48 hours
- **Updates** - We will keep you informed of our progress as we investigate
- **Resolution** - We aim to resolve critical vulnerabilities within 90 days

## Confidentiality

**All reports will be kept confidential.** We will not share your information with third parties without your consent, except as required by law.

## Supported Versions

We recommend always running the latest version of `git-sync`.

## Scope

This security policy applies to:

- The `git-sync` CLI and `git-sync-bench` benchmark command
- The `entire.io/entire/gitsync` and `entire.io/entire/gitsync/unstable` Go packages
- Official Entire GitHub repositories

### Out of Scope

The following are generally not considered security vulnerabilities:

- Issues in third-party dependencies (please report these upstream)
- Social engineering attacks
- Denial of service attacks against remotes you do not control
- Issues requiring physical access to a user's device

Because `git-sync` operates against Git remotes, please be especially careful when reporting issues that involve credentials, TLS verification, or remote-to-remote relay behavior — include the exact remote configuration that triggers the issue if it is reproducible.

---

## Security Advisories

Security advisories are issued when a confirmed vulnerability can be exploited by a remote or non-local actor. The following are generally treated as **bug reports rather than security advisories**:

- Regular expression performance issues (ReDoS) that only affect local execution
- Resource exhaustion that requires local access to trigger
- Issues that cannot be exploited without direct access to the user's machine or to credentials the user already controls

Use [GitHub Issues](https://github.com/entireio/gitsync/issues) to report bugs.

---

Thank you for helping keep `git-sync` and the Entire community safe!

ASECURITY.md+68

71 unmodified lines

72
73
74
75
76
77
75
76
77
78
79
80
112 unmodified lines

193
194
195
196
197
198
199
196
197

71 unmodified lines

- source-authoritative overwrite planning
  - relay-only execution
  - no materialized fallback
  - works against targets that advertise `no-thin` (the relayed pack is
    always self-contained because our upload-pack client does not request
    the `thin-pack` capability)
  - works against targets regardless of `no-thin` advertisement: the
    relayed pack is always self-contained because our upload-pack client
    does not request the `thin-pack` capability

The current transfer modes are:

112 unmodified lines

## Related Notes

- [bootstrap.md](bootstrap.md)
- [bootstrap-batching.md](bootstrap-batching.md)
- [benchmarking.md](benchmarking.md)
- [embedding.md](embedding.md)
- [protocol.md](protocol.md)
- [testing.md](testing.md)

Mdocs/architecture.md+5/-7

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51

# Benchmarking

`git-sync-bench` runs repeatable empty-target benchmarks against a source repository.

It currently supports:

- `bootstrap`: calls `syncer.Bootstrap` directly.
- `sync`: calls `syncer.Run`, which may choose bootstrap relay automatically on an empty target.

The tool creates a fresh bare target repository for each run and reports both wall-clock time and the internal `syncer` measurement data.

## Build

```bash
go build -o /tmp/git-sync-bench ./cmd/git-sync-bench
```

## Example

Against a local mirror:

```bash
/tmp/git-sync-bench \
  --scenario bootstrap \
  --source-url /tmp/git-sync-bench/kubernetes.git \
  --repeat 3 \
  --target-max-pack-bytes 104857600 \
  --stats \
  --json
```

If `--source-url` is a filesystem path, the tool converts it to `file://...` automatically.

## Output

The JSON report includes:

- per-run wall time
- per-run `syncer.Result`
- aggregate min/avg/max wall time
- aggregate min/avg/max internal elapsed time
- aggregate min/avg/max actual batch count for batched runs
- aggregate min/avg/max planned batch count for batched runs
- maximum observed alloc and heap-inuse peaks
- relay modes seen across successful runs

## Notes

- `bootstrap` benchmarks reject `--force` and `--prune`, matching `git-sync bootstrap`.
- `--keep-targets` retains the generated bare targets under `--work-dir` for inspection.
- For large real-repo runs, prefer using a local mirror rather than benchmarking directly against a hosted remote.

Ddocs/benchmarking.md-51

Bootstrap Batching Design

bootstrap currently streams one source pack into one target push. That is good for many initial syncs, but it is not enough for very large single-branch repositories where one initial pack is itself too large for comfortable target-side unpacking and indexing.

This note sketches a batching design for large bootstrap jobs.

Goal

Reduce per-push size and target-side receive-pack / index-pack pressure for very large initial syncs, while preserving the main benefit of bootstrap:

Non-Goals

V1 batching should not try to solve every large-migration problem.

Out of scope:

Why Per-Ref Batching Is Not Enough

Per-ref batching is easy:

That helps when there are many refs and each ref is moderate in size.

It does not help for repositories where a single branch is enormous. Linux master is the motivating example: even a single branch bootstrap can be too large for one target-side unpack/index step.

Preferred Model

Use branch checkpoint batching with temporary refs.

High-level idea:

  1. Choose a sequence of ancestor checkpoints for a source branch.
  2. Push them oldest to newest into a temporary target ref.
  3. Once the final tip is present, create the real target ref.
  4. Delete the temporary ref at the end.

This gives:

Command Shape

Possible CLI extension:

git-sync bootstrap \
  --target-max-pack-bytes 1073741824 \
  <source-url> \
  <target-url>

Possible related flags:

The first version should only need --target-max-pack-bytes.

Temporary Ref Strategy

For each target branch:

Flow:

  1. Create/update only the temp ref during intermediate batches.
  2. After the final tip batch succeeds:
    • create the real ref at the final tip
    • delete the temp ref

This keeps the target repository in a cleaner state:

If failure happens mid-run:

Checkpoint Selection

Checkpoints are placed using a commit-count estimate, not measured pack sizes.

For a branch tip:

  1. Fetch the commit graph (tree:0 filter, one round-trip — commits only, no blobs/trees).
  2. Walk first-parent ancestry backward to get the chain length.
  3. Estimate total pack size: chainLen × 8 KiB/commit.
  4. Compute number of batches: ceil(estimated / --target-max-pack-bytes).
  5. Place checkpoints evenly along the first-parent chain.

This is a heuristic — real bytes-per-commit varies widely (2–100+ KiB depending on blob churn). The estimate intentionally errs toward more batches.

Adaptive size correction

If the estimate is too optimistic (fewer batches than needed), two safeguards catch it:

  1. PACK header pre-check: after starting a fetch, peek at the first 12 bytes of the pack to read the object count. Multiply by ~750 bytes/object. If the estimate exceeds --target-max-pack-bytes, abort the fetch (12 bytes wasted, not gigabytes), insert a midpoint checkpoint, and retry. This avoids a full transfer for obviously-oversized batches.

  2. Target rejection retry: if the target's receive-pack rejects a push for exceeding its body-size limit, detect the error, insert a midpoint checkpoint from the stored chain, and retry. This catches cases where the PACK header estimate was close but the real pack was slightly over.

Both safeguards converge in O(log n) splits — each failure halves the commit range.

Why not probe (the previous design)

The previous implementation did full FetchPack round-trips per probe candidate to measure actual pack sizes. For linux/master (75k commits) this required 13+ fetch-and-discard cycles, downloading gigabytes of throwaway data and taking minutes before any real push started. The estimate approach reduces planning to one commit-graph fetch (~20 seconds) plus arithmetic.

Batch Flow For One Branch

Given checkpoints:

The flow is:

  1. Fetch source pack for C1 with no have.
  2. Push it to temp ref refs/gitsync/bootstrap/heads/main.
  3. Fetch source pack for C2 with have=C1.
  4. Push update of temp ref to C2.
  5. Fetch source pack for C3 with have=C2.
  6. Push update of temp ref to C3.
  7. Fetch source pack for tip with have=C3.
  8. Push update of temp ref to tip.
  9. Create real target ref refs/heads/main at tip.
  10. Delete temp ref.

The final ref creation can be:

The first version should prefer the separate final ref creation because it is easier to reason about.

Tags

Tags should not be pushed during intermediate branch batches.

Recommended rule:

This avoids cases where a tag points at an object graph that is not yet fully present on target.

V1 batching should support:

Tag batching can be added later.

Restart and Recovery

Batching is only worth doing if failures are restartable.

Minimum restart model:

  1. Detect existing temp refs on target.
  2. Resolve their current hashes.
  3. Resume from the latest completed checkpoint instead of starting from zero.

If --keep-temp-refs-on-failure is false, cleanup can still happen on clean failures, but default restartability is more valuable than aggressive cleanup.

Safety Model

V1 batching should remain strict.

Allow only:

Fail if:

Implementation Shape

Suggested pieces:

The estimator should reuse the existing relay mechanics:

But it will need one new planning pass to probe likely batch sizes before actual execution.

Operator Output

Batching should be explicit in output.

Text output should include:

JSON should include:

Practical Risks

This is still likely worthwhile for very large initial migrations because it changes a single huge risky operation into several bounded ones.

Recommended Phases

Phase A:

Progress:

Operator guidance:

Phase B:

Phase C:

Progress:

Phase D:


Ddocs/bootstrap-batching.md-292

# Bootstrap Design

`bootstrap` is the dedicated command path for initial remote-to-remote seeding when the target does not yet contain the managed refs.

The goal is to avoid decoding the fetched source objects into the local in-memory object store during an initial sync. Instead, `bootstrap` should fetch a pack from the source and relay it directly into target `receive-pack`.

## Why

The current `sync` path is optimized for general incremental reconciliation:

- it fetches from source with target tip hashes as `have`
- it builds plans locally
- it stores fetched source objects in a local object store
- it computes the object closure to push
- it encodes a new pack for target

That is a good general path, but it is a poor fit for very large initial syncs into an empty target because the missing object graph must fit in local memory.

`bootstrap` is meant to cover the opposite case:

- target refs are absent
- all actions are creates
- there is no need for fast-forward checks
- the main cost is moving a large pack from source to target efficiently

## V1 Scope

`bootstrap` should be intentionally narrow:

- create-only
- fail if any managed target ref already exists
- branch refs by default
- optional `--tags`
- optional explicit `--map`
- no `--force`
- no `--prune`
- no mixed create and update runs
- no automatic fallback to normal `sync`
- smart HTTP only

This command is for first-time seeding. After that, operators should use `sync`.

## Command Shape

Preferred CLI:

```bash
git-sync bootstrap [flags] <source-url> <target-url>
```

Expected v1 flags:

- `--branch`
- `--map`
- `--tags`
- `--max-pack-bytes`
- `--stats`
- `--json`
- `--protocol auto|v1|v2`
- existing source and target auth flags

## Intended Flow

1. List source refs.
2. List target refs.
3. Build the managed ref set from `--branch`, `--map`, and `--tags`.
4. Fail if any managed target ref already exists.
5. Build create commands for the target.
6. Ask source for a pack containing the selected source tips.
7. Strip protocol framing and sideband as needed.
8. Stream the resulting pack directly into target `receive-pack`.
9. Parse target report-status and return a create summary.

## Why This Helps

The large memory cost in the current implementation comes from storing fetched source objects locally before re-encoding them.

`bootstrap` should avoid that cost for initial syncs by not materializing the object graph in local storage unless a fallback path is explicitly chosen later.

The expected wins are:

- much lower RAM usage for empty-target syncs
- less local CPU spent decoding and re-encoding large object graphs
- better fit for large repo migrations

## Constraints

There are still some hard limits:

- source and target still need normal smart HTTP discovery
- target policy can still reject pushes
- push still depends on target `receive-pack` behavior and capabilities
- if a relay-safe path cannot be used, `bootstrap` should fail and tell the user to use `sync`

V1 should stay strict rather than trying to be clever.

## Implementation Notes

The cleanest implementation shape is a separate code path, not an optimization hidden inside `sync`.

Suggested pieces:

- `runBootstrap` in `cmd/git-sync/main.go`
- `syncer.Bootstrap(ctx, cfg)` in `internal/syncer`
- source fetch helper that returns a pack stream instead of writing objects into storage
- target receive-pack helper that accepts an externally supplied pack stream
- bootstrap-specific result type or reuse `Result` with only create actions

The initial implementation should prefer:

- one multi-ref source fetch
- one multi-command target push

That keeps it efficient and conceptually simple.

## Failure Rules

V1 should fail when:

- any managed target ref already exists
- no source refs matched
- the source fetch cannot be relayed cleanly
- target push fails

The error should explicitly recommend normal `sync` when the repository is no longer in bootstrap shape.

## Follow-Up Steps

Phase 1:

- implement `bootstrap` for create-only branch refs
- support optional tag creation
- add JSON and stats output
- add in-process integration tests
- add `git-http-backend` integration coverage for empty-target bootstrap

Progress:

- `bootstrap` is implemented
- optional tag creation is supported on the non-batched path
- JSON and stats output are supported
- in-process integration coverage exists
- `git-http-backend` integration coverage exists

Phase 2:

- allow relay-safe create-only runs with explicit mapped refs
- add better operator output for large initial transfers
- add safety thresholds for advertised/fetched bytes

Progress:

- explicit mapped refs are supported
- `--max-pack-bytes` provides a first safety threshold for the streamed source pack during bootstrap
- `--target-max-pack-bytes` now enables a Phase A batched branch-only bootstrap mode for large initial syncs

Phase 3:

- investigate hybrid behavior: relay when the target is empty, otherwise fail fast into normal `sync`
- investigate whether target capability combinations require alternate pack handling
- measure source-to-target pack relay memory and CPU against current `sync`

Progress:

- `sync` now auto-selects the bootstrap relay path when all managed target refs are absent and the run matches bootstrap semantics
- dry-run `plan` surfaces a bootstrap suggestion for the same target shape

Batching note:

- the current batched bootstrap path is intentionally narrow
- it requires source-side protocol v2 with fetch filters
- it batches branch refs and then optionally creates tags after the branch batches complete
- it resumes from an existing temp ref when that temp ref matches a planned checkpoint
- it uses temporary target refs under `refs/gitsync/bootstrap/heads/`
- it should be treated as an advanced large-repo fallback when one-shot bootstrap is too risky or fails on target-side unpack/index pressure

Phase 4:

- consider a more advanced incremental relay mode for non-empty targets
- only pursue this if large migration workflows become important enough to justify the added protocol complexity

Progress:

- there is now a narrow incremental relay path in `sync`
- it now covers multi-branch fast-forward branch-only updates
- it now also covers branch-to-branch mappings
- it now also covers create-only tags
- tag retargets, deletes, force, and prune still use the normal path

Ddocs/bootstrap.md-188

Embedding

git-sync can now be used as a library as well as a CLI.

For most embedders, there are two important rules:

Stable vs Unstable

Use gitsync when you want a durable worker-facing API:

Use unstable only when you need controls that are intentionally not yet stable:

The CLI and benchmark command use unstable because they still need those controls. External workers should generally not.

Worker Shape

A queue worker usually wants:

  1. deserialize a job into source, target, scope, and policy
  2. build a gitsync.Client
  3. inject auth and an http.Client
  4. call Plan or Sync
  5. persist structured result data
  6. decide success, retry, or escalation

Minimal example:

package worker

import (
    "context"
    "net/http"

"entire.io/entire/gitsync"
)

func runSync(ctx context.Context) error {
    client := gitsync.New(gitsync.Options{
        HTTPClient: &http.Client{},
        Auth: gitsync.StaticAuthProvider{
            Source: gitsync.EndpointAuth{Token: "source-token"},
            Target: gitsync.EndpointAuth{Token: "target-token"},
        },
    })

result, err := client.Sync(ctx, gitsync.SyncRequest{
        Source: gitsync.Endpoint{URL: "https://github.example/source/repo.git"},
        Target: gitsync.Endpoint{URL: "https://git.example/target/repo.git"},
        Scope: gitsync.RefScope{
            Branches: []string{"main"},
        },
        Policy: gitsync.SyncPolicy{
            IncludeTags: true,
            Protocol:    gitsync.ProtocolAuto,
        },
    })
    if err != nil {
        return err
    }

_ = result
    return nil
}

Auth Injection

gitsync uses one auth ownership model:

That avoids baking CLI-style precedence rules into request types.

Good uses of AuthProvider:

The simplest option is gitsync.StaticAuthProvider, but a real worker will usually implement AuthProvider itself.

HTTP Injection

Pass an *http.Client through gitsync.Options when you need:

git-sync clones and wraps the provided client internally so it can still collect transfer stats without mutating the caller's client directly.

Result Handling

The stable SyncResult is organized for worker consumption:

That gives a worker enough structure to:

Retry Guidance

Treat these differently:

For many workers, a useful pattern is:

What Not To Depend On

If you want stability, do not build external worker logic around:

Those are implementation details or advanced controls that currently belong in unstable, not the stable embedding contract.


Ddocs/embedding.md-164

# Protocol

This document is a code-grounded walkthrough of the Git wire protocol pieces that `git-sync` implements directly: smart HTTP transport, pkt-line framing, capability negotiation, sideband multiplexing, and the relay path that streams a source pack into target `receive-pack` without local materialization.

It is aimed at contributors and embedders who want to understand *why* the implementation looks the way it does. For higher-level operational behavior, see [architecture.md](architecture.md).

## Scope

Covered:

- Smart HTTP transport (`info/refs` discovery + RPC POST endpoints)
- Pkt-line framing
- Sideband-64k multiplexing
- Capability negotiation, both v1 (advertised on the first ref line) and v2 (`ls-refs` / `fetch` capability advertisement)
- Source fetch and target push request/response shape
- The relay path: how a fetched source pack is forwarded into target `receive-pack` without decoding object graphs locally
- Auth, TLS, redirect handling

Out of scope (with pointers):

- The pack format itself (object types, deltas, index format) — see [Git's pack-format docs](https://git-scm.com/docs/pack-format)
- Dumb HTTP — `git-sync` does not support it
- SSH transport — `git-sync` is HTTPS-only
- Bundle URI, partial clones, and other newer extensions

## Smart HTTP Overview

`git-sync` talks to two RPC services per remote:

| Service          | Discovery endpoint                            | RPC endpoint        | Direction             |
|------------------|-----------------------------------------------|---------------------|-----------------------|
| `git-upload-pack` | `GET /info/refs?service=git-upload-pack`     | `POST /git-upload-pack` | source: list refs, fetch objects |
| `git-receive-pack` | `GET /info/refs?service=git-receive-pack` | `POST /git-receive-pack` | target: list refs, push objects  |

The `/info/refs` GET serves two purposes: it advertises the available refs (in v1) or the v2 capability set, and it acts as the negotiation handshake — the response tells the client which protocol version the server speaks and which capabilities are available. The subsequent POST to the RPC endpoint carries the request body in `application/x-<service>-request` format and receives `application/x-<service>-result` back.

`git-sync` is smart-HTTP-only. Dumb HTTP is not supported because relay-friendly negotiation depends on capability discovery and `have`/`want` exchange that dumb HTTP does not provide.

The transport implementation lives in `internal/gitproto/smarthttp.go`. The two main entry points are `RequestInfoRefs` (the GET) and `PostRPC` / `PostRPCStream` / `PostRPCStreamBody` (the POST, with buffered or streaming response bodies).

## Pkt-Line Framing

Every smart HTTP request and response body is a sequence of pkt-line packets. Each packet begins with a 4-byte ASCII hex length prefix, followed by the payload.

Three special length values do not carry a payload:

| Marker | Hex     | Constant                 | Meaning           |
|--------|---------|--------------------------|-------------------|
| flush  | `0000`  | `PacketFlush`            | end of section / end of stream |
| delim  | `0001`  | `PacketDelim`            | section delimiter (v2) |
| end    | `0002`  | `PacketResponseEnd`      | end of v2 response |

Anything else is parsed as `0xXXXX` and indicates a packet of `n - 4` payload bytes.

Implementation: `internal/gitproto/pktline.go`. `PacketReader` reuses a fixed 4-byte header buffer and a growable payload buffer per call to `ReadPacket()` to keep allocations down on long pkt-line streams. `EncodeCommand` builds v2 command request bodies as `command=<name>` + capability args + optional delim + command args + flush. `FormatPktLine` is a small helper for building one-off payload packets.

### Sideband-64k

When the server advertises `side-band-64k`, the response payload is multiplexed across three logical channels. The first byte of each pkt-line payload is the channel selector:

| Channel | Meaning                  |
|---------|--------------------------|
| `0x01`  | pack data                |
| `0x02`  | progress text (stderr)   |
| `0x03`  | fatal error              |

`git-sync` always negotiates `side-band-64k` when available. `PreferredSideband` in `capability.go` enforces preferring `side-band-64k` over the older `side-band` capability. The demux happens in the fetch path; see "Source Fetch" below for how the pack stream is unwrapped.

## Capability Negotiation

### Source side (upload-pack)

For v1, capabilities are advertised at the end of the first ref line in the `info/refs` response, separated from the ref by a NUL byte. `git-sync` parses these via `go-git`'s `packp.AdvRefs`.

For v2, capabilities are advertised as their own pkt-line section after a `version 2\n` payload. `DecodeV2Capabilities` in `internal/gitproto/capability.go` parses this into `V2Capabilities{Caps: map[string]string}`. Each capability is either a bare name or `name=value`.

The capabilities `git-sync` cares about on the source side:

- `agent` — informational; `git-sync` echoes its own agent string back
- `side-band-64k` — multiplexed responses with progress
- `multi_ack_detailed`, `no-done` — fetch negotiation efficiency
- `ofs-delta` — pack encoding capability
- `fetch` (v2) — advertised features include `shallow`, `filter`, `wait-for-done`. `FetchSupports` parses the space-separated feature list.
- `ls-refs` (v2) — required for v2 ref discovery

### Target side (receive-pack)

`TargetFeaturesFromAdvRefs` in `internal/gitproto/target_features.go` summarizes the receive-pack advertisement into a typed `TargetFeatures` struct:

```go
type TargetFeatures struct {
    Known        bool
    DeleteRefs   bool
    NoThin       bool
    OFSDelta     bool
    ReportStatus bool
    Sideband     bool
    Sideband64k  bool
}
```

`git-sync` never requests `thin-pack` on the source side, so the source pack is always self-contained. That makes the relayed pack safe to push to a target regardless of its `no-thin` advertisement. The relay paths still differ on what they do when the target advertises `no-thin`: see "Why `no-thin` matters" under Target Push.

`DeleteRefs` gates `--prune` and explicit deletes. `ReportStatus` is required to learn which ref updates the target accepted. `OFSDelta` affects pack encoding compatibility. `Sideband` / `Sideband64k` allow the target to multiplex its own response.

## Protocol v1 vs v2

Protocol v2 was introduced to make ref discovery cheaper (you can ask for a ref prefix instead of receiving the full advertisement) and to make the RPC body more structured.

| Aspect                | v1                                          | v2                                              |
|-----------------------|---------------------------------------------|-------------------------------------------------|
| Discovery handshake   | full ref advertisement on `info/refs`       | capability advertisement only                    |
| Ref listing           | implicit in discovery                       | explicit `command=ls-refs` RPC                   |
| Fetch                 | `want`/`have`/`done` lines on `upload-pack` | `command=fetch` RPC with capability args         |
| Filtering (`tree:0`)  | not standardized                            | supported via `fetch` capability `filter` feature |
| Push                  | command list + pack on `receive-pack`       | (not yet standardized; not used)                 |

`git-sync` uses v2 on the source side when supported, and v1 on the target side always. The reasons:

- **Source-side v2** enables features `git-sync` actually uses: `ls-refs` for cheaper ref listing, and `fetch` with `filter=tree:0` for the commit-graph-only fetches that batched bootstrap planning relies on.
- **Target-side push** stays on v1 because v2 receive-pack is not yet adopted broadly enough to rely on, and because `git-sync` already needs explicit command construction and streaming control on the push path.

The CLI flag is `--protocol auto|v1|v2`. `auto` (the default) tries source-side v2 first via the `Git-Protocol: version=2` header, and falls back to v1 if the server doesn't acknowledge v2 in its response. `v2` requires the source to negotiate v2 and fails otherwise. `v1` skips the v2 attempt.

The negotiation is done in `internal/gitproto/refs.go` via `ListSourceRefs(ctx, conn, protocolMode, refPrefixes)`, which returns a `RefService` that records the negotiated protocol and either the v1 advertisement or the v2 capabilities. Subsequent fetches go through that `RefService`.

## HEAD Symref Discovery

When the source advertises HEAD as a symbolic ref, `git-sync` records what branch it points to. This is used by batched bootstrap planning as a *trunk hint*: the trunk branch is planned first so its commit graph becomes a stop-set for subsequent branches' first-parent walks, dramatically cutting per-branch graph fetches on multi-branch repos.

The hint is exposed as `RefService.HeadTarget plumbing.ReferenceName`. Empty value means "detached HEAD or no symref advertised", in which case the caller falls back to the original branch order.

### v1

In v1, HEAD's symref target is announced as a `symref=` capability on the first ref line:

```
<hash> HEAD\0... symref=HEAD:refs/heads/main side-band-64k ofs-delta ...
```

`headTargetFromAdv` walks the parsed `capability.SymRef` values and returns the right-hand side of `HEAD:<branch>`.

### v2

In v2, the symref-target attribute is delivered as part of the `ls-refs` response. To get it, `git-sync` always requests HEAD and the `symrefs` argument:

```
command=ls-refs
agent=git-sync/...
0001
peel
symrefs
ref-prefix HEAD
ref-prefix refs/heads/
ref-prefix refs/tags/
0000
```

The `peel` and `symrefs` arguments are ls-refs request features that ask the server to include peeled object IDs (for tags) and symref-target attributes. `ref-prefix HEAD` is unconditionally included so HEAD shows up in the response even when the caller only asked for `refs/heads/` or `refs/tags/`.

The server's HEAD line then looks like:

```
<hash> HEAD symref-target:refs/heads/main
```

`decodeV2LSRefs` parses the line, extracts the `symref-target:` attribute, and returns it as `headTarget` alongside the ref slice. HEAD itself is filtered out of the returned refs because it is a symbolic ref, not a real one — matching v1 behavior where symrefs are filtered out by the downstream `RefHashMap`.

### Consumers

The trunk hint is read from `RefService.HeadTarget` and passed into `bootstrap.Params.SourceHeadTarget` (the wiring lives in `internal/syncer/syncer.go`). `orderTrunkFirst` (in `internal/strategy/bootstrap/bootstrap.go`) reorders the desired-ref list so the trunk is planned first. If the trunk is not in the desired set (filtered out by `--branch` or `--map`), the original order is preserved and the trunk-first optimization is skipped.

## Source Fetch

### v1

Request body format (one pkt-line per element):

```
want <hash> <capabilities-on-first-line>\n
want <hash>\n
...
0000
have <hash>\n
...
done\n
```

Capabilities are appended to the first `want` line, separated from the hash by a space. `git-sync` requests `agent=...`, the preferred sideband (`side-band-64k` when supported, otherwise `side-band`), `ofs-delta` when supported, and `include-tag` when the request asks for tags. It adds `no-progress` unless `Verbose` is set on the `RefService`. `git-sync` deliberately does *not* request `thin-pack` — see "Why no-thin matters" under Target Push.

Response: a pkt-line stream that begins with NAK / ACK negotiation packets and transitions into pack data on sideband channel `0x01`. Channel `0x02` carries progress text; channel `0x03` is fatal.

### v2

Request body is built by `EncodeCommand("fetch", capabilityArgs, commandArgs)`:

```
command=fetch\n
agent=git-sync/...\n
0001
want <hash>\n
have <hash>\n
done\n
0000
```

The capability section is separated from the command-arg section by a delim packet (`0001`), and the whole body is terminated by a flush packet (`0000`). Filter features like `filter tree:0` go in the command-arg section.

Response: pkt-line sections delimited by `0001`. The interesting section is `packfile`, where each subsequent pkt-line carries one byte of channel selector + payload (same sideband-64k encoding as v1).

### Demux

`internal/gitproto/fetch.go` is responsible for unwrapping the response into a raw pack stream. It:

1. Reads pkt-lines from the buffered `PacketReader`.
2. Skips ACK/NAK negotiation packets (v1) or section delimiters until reaching `packfile` (v2).
3. For each subsequent payload pkt-line, dispatches the first byte to the right channel: `0x01` → forward to the pack consumer, `0x02` → progress (stderr if `Verbose`, otherwise discard), `0x03` → fail.
4. Stops on the next flush packet.

Once the stream switches to pack mode, subsequent reads bypass pkt-line framing where possible. The reader exposes `BufReader()` so the caller can hand the underlying buffered reader directly to a pack-consumer that reads raw bytes.

## Target Push

`receive-pack` POST body is, in order:

1. **Command list**, one pkt-line per ref update:
```
<old-hash> <new-hash> <refname>\0<capabilities>\n
```
Capabilities are appended to the first command after a NUL byte. `git-sync` requests `report-status` and the preferred sideband (`side-band-64k` when available, otherwise `side-band`); `delete-refs` is added when the command list contains a delete. `ofs-delta` is consulted from the target advertisement to choose pack-encoding behavior in `useRefDeltas`, but is not itself requested as a wire capability.
2. **Flush packet** (`0000`) terminating the command section.
3. **Raw pack bytes**. The pack itself is NOT pkt-line framed — the receive-pack server reads the pack directly off the request body after the flush.

### Why `no-thin` matters

A "thin pack" is one that contains delta objects whose base is not in the pack itself. The receiver is expected to resolve those bases against existing target objects. Thin packs are smaller on the wire but require the receiver to reach into the existing object database during indexing. When a `receive-pack` server advertises `no-thin`, it is signalling that it does not support thin packs — the client must send a self-contained pack.

`git-sync` always sends self-contained packs because it never requests the `thin-pack` capability on `upload-pack` (see the comment in `internal/gitproto/fetch.go`). That means the relayed pack is safe for both kinds of target. The relay paths still make different choices given the target's advertisement:

- **Replicate** explicitly tolerates `no-thin` and proceeds with relay. The comment in `internal/planner/relay.go` documents the reasoning: source never sends thin-pack, so the relayed pack is self-contained and safe.
- **Incremental relay** inside `sync` is conservative: it skips `no-thin` targets and falls back to the materialized path. This is a safety-margin choice; it could in principle be relaxed using the same argument as replicate.
- **Bootstrap** does not consult `no-thin` — empty-target relay is allowed regardless.

### report-status

After the pack is uploaded, the server responds with a `report-status` (or `report-status-v2`) section listing per-ref outcomes:

```
unpack ok\n
ok refs/heads/main\n
ng refs/heads/release pre-receive-hook-rejected\n
```

`git-sync` parses this in `internal/gitproto/push.go` and surfaces per-ref outcomes through the higher-level result types.

The push implementation lives in `internal/gitproto/push.go`.

## Relay Framing

The conceptual move that distinguishes `git-sync` from a `clone --mirror` + `push --mirror` workflow is **relay**: the source's pack bytes are forwarded into the target's `receive-pack` request body without going through a local object decode + repack cycle.

In framing terms:

```
source response stream                   target push request body
─────────────────────────────              ──────────────────────────
pkt-line(NAK/ACK or v2 sections)           pkt-line(command list)
pkt-line(0x02 progress) ─── stderr         pkt-line(flush)
pkt-line(0x01 [pack chunk]) ───────┐       raw pack bytes ◄───────────┐
pkt-line(0x01 [pack chunk]) ───────┼─►     raw pack bytes ◄───────────┤
pkt-line(0x01 [pack chunk]) ───────┘       raw pack bytes ◄───────────┘
pkt-line(flush)
```

The source's pack chunks come out one-per-pkt-line on sideband channel `0x01`. `git-sync` strips the pkt-line + sideband envelope and concatenates the raw pack bytes into the target push body, after the command list and flush. The pack itself is not re-encoded: byte for byte, the source-produced pack shows up on the target's wire.

This is why the relay paths (bootstrap, replicate, and the incremental relay path inside `sync`) avoid the in-memory cost that the materialized fallback pays. The local process never holds the object graph; it holds at most one pack chunk at a time.

### PACK header pre-check

The batched bootstrap path needs to make a sizing decision before committing to a full transfer. Every pack starts with a 12-byte header:

```
"PACK"  uint32 version  uint32 object_count
```

`checkPackSizeAndSubdivide` in `internal/strategy/bootstrap/bootstrap.go` peeks at the first 12 bytes of the unwrapped pack stream, multiplies the object count by a per-object byte estimate, and if the projected pack size exceeds `--target-max-pack-bytes`, aborts the fetch and inserts an additional checkpoint between the current `have` and the target tip — wasting 12 bytes of read instead of gigabytes of transfer.

The same code path also detects target-side body-size rejections after the fact: if the target rejects the push because the pack is too large, the planner inserts a midpoint checkpoint and retries. Both safeguards converge in O(log n) splits.

## Auth, TLS, Credential Helper

### Auth methods

`ApplyAuth` in `internal/gitproto/smarthttp.go` recognizes two auth method types from `go-git`:

- `*transporthttp.BasicAuth` — sets `Authorization: Basic <base64(user:pass)>`. For GitHub-style providers, the password slot carries the token (`username=git`, `password=$GITHUB_TOKEN`).
- `*transporthttp.TokenAuth` — sets `Authorization: Bearer <token>`. Used for providers that expect bearer tokens directly.

Resolution order (CLI side):

1. Explicit `--source-token`, `--source-bearer-token`, `--source-username` flags (and target equivalents)
2. `GITSYNC_*` environment variables (`GITSYNC_SOURCE_TOKEN`, `GITSYNC_TARGET_BEARER_TOKEN`, etc.)
3. `git credential fill` helper lookup, for `http://` and `https://` URLs
4. Anonymous (no `Authorization` header)

Library callers inject auth via the `AuthProvider` interface in `entire.io/entire/gitsync` instead.

### TLS

`NewHTTPTransport(skipTLS bool)` returns a clone of `http.DefaultTransport` with `InsecureSkipVerify` flipped on when requested. The CLI flags are `--source-insecure-skip-tls-verify` and `--target-insecure-skip-tls-verify`; environment overrides are `GITSYNC_*_INSECURE_SKIP_TLS_VERIFY=true|1|yes|on`.

This is intended for local self-signed targets. Do not use it against the public internet.

## /info/refs Redirect Behavior

Some hosting providers respond to `/info/refs` with a 307 redirect to a different host (typically a regional replica or a canonicalized URL). Vanilla `git` follows the redirect for the discovery GET but also retargets the subsequent RPC POSTs at the redirect target.

`git-sync` exposes this as the `FollowInfoRefsRedirect` field on `gitproto.Conn`, and as the CLI flags `--source-follow-info-refs-redirect` and `--target-follow-info-refs-redirect`:

- **Off (default)**: redirects are followed for the GET, but the subsequent RPC POSTs go to the original `Endpoint.Host`. This preserves stable behavior for callers that build URLs ahead of time.
- **On**: after `RequestInfoRefs` follows redirects, `Endpoint.Scheme` and `Endpoint.Host` are rewritten to the final URL's scheme and host. `Endpoint.Path` is never modified — it still contains the repo path.

Library callers can set this via `gitsync.Endpoint{FollowInfoRefsRedirect: true}`.

## Stats Counters and Common Failures

When `--stats` is set, every HTTP round-trip is annotated with the `X-Git-Sync-Stats-Phase` header (`StatsPhaseHeader` in smarthttp.go), e.g. `git-upload-pack info-refs`, `git-receive-pack push`, etc. The instrumented round-tripper accumulates per-phase counters:

- request count
- bytes sent / received
- `want` count, `have` count
- target push command count

These show up as the `stats` block in `--json` output and as a summary table in text output.

Bounded reads protect the local process:

- `info/refs` responses are capped at 64 MiB
- Buffered RPC responses (`PostRPC`) are capped at 128 MiB

Streaming RPC bodies (`PostRPCStreamBody`) bypass these caps because the consumer is the relay path, which reads pack chunks incrementally and forwards them on.

### Common failure surfaces

- **v2 negotiation failed**: source returned a v1 advertisement when `--protocol v2` was forced. `auto` falls back; `v2` errors out. Causes: older server, intermediate proxy stripping the `Git-Protocol` header.
- **Target rejected push (body too large)**: relay produced a pack larger than the target's request body limit. Solutions: lower `--target-max-pack-bytes` for batched bootstrap, or rerun on a target with a higher limit. Detected by `isTargetBodyLimitError` in `internal/strategy/bootstrap/bootstrap.go`.
- **Target advertises `no-thin`**: the incremental relay path inside `sync` skips it and falls back to the materialized path. `replicate` proceeds anyway.
- **Redirect chain mismatch**: the GET redirects to host X, but a subsequent RPC POST gets a different redirect or a 404 because the server expected sticky sessions. Resolution: set `FollowInfoRefsRedirect=true` so RPC POSTs go directly to the post-redirect host.
- **Auth 401 / 403**: 401 generally indicates missing or wrong credentials (will retry with credential helper if configured); 403 indicates the credentials were accepted but lack permission for the requested action.