chore(trail): tailor runner prompts to the CLI repo · Entire

chore(trail): tailor runner prompts to the CLI repo

fa4517b·

Soph·3w ago·3 files·+3 added/-3 removed

The shipped risk/confidence/review-focus runner templates were written for a generic web/backend app (payments, DB migrations, TypeScript). Rewrite them around what actually makes changes risky in this Go CLI:

Co-Authored-By: Claude Opus 4.8 (1M context) noreply@anthropic.com

Sessions

55f8c01f7de4View transcript

Changes

3

15 unmodified lines

16
17
18
19
19
20
21
22

15 unmodified lines

"kind": "trail_prompt"
  },
  "prompt": {
    "template": "You are a code confidence evaluator. Analyze the changes on branch \"{{branch}}\" compared to \"{{base_branch}}\".\n\nRun these commands to gather context:\n\n1. git diff origin/{{base_branch}}...HEAD --stat\n2. git diff origin/{{base_branch}}...HEAD\n3. Look for CI configuration files (.github/workflows/, .circleci/, etc.)\n4. Look for test files related to the changed code\n\nThen evaluate the **Confidence** of these changes (0-100). Confidence means: \"How sure can we be that these changes are correct and well-tested?\"\n\nConsider the following dimensions:\n\n1. **CI pipeline presence** — Does the project have CI configured? Are the workflows relevant to the changed code?\n2. **Test coverage on changed lines** — Are there tests that exercise the changed code paths? Are new features covered by new tests?\n3. **Test quality and relevance** — Do the tests actually assert meaningful behavior? Are edge cases covered? Are the tests testing the right things?\n4. **Type safety** — Does the project use TypeScript or similar? Are the changes well-typed?\n5. **Code review signals** — Are there obvious code smells, TODOs, or commented-out code in the changes?\n\nScore from 0 to 100:\n- 0-20: No tests, no CI, high uncertainty\n- 21-40: Minimal test coverage, basic CI\n- 41-60: Moderate coverage, some gaps in critical paths\n- 61-80: Good coverage, CI catches most issues\n- 81-100: Excellent coverage, comprehensive tests for all changes\n\nAfter your analysis, output ONLY this JSON object as the very last line of your response:\n\n{\"value\": <number 0-100>, \"rationale\": \"<1-2 sentence explanation>\"}"
    "template": "You are a confidence evaluator for the Entire CLI, a Go codebase (cobra/huh). Analyze branch \"{{branch}}\" vs \"{{base_branch}}\".\n\nGather context:\n1. git diff origin/{{base_branch}}...HEAD --stat\n2. git diff origin/{{base_branch}}...HEAD\n3. Look for tests next to the changed code (*_test.go, integration tests under cmd/entire/cli/integration_test/, e2e under e2e/).\n4. Look at CI config (.github/workflows/) and the mise tasks it runs (check, lint, test:ci).\n\nScore the **Confidence** (0-100): \"How sure can we be these changes are correct and well-tested?\"\n\nWeigh:\n1. **Test coverage on changed lines** — do unit/table tests exercise the changed paths, and do new features/branches get new tests? Strategy, hook, and git-mutating changes especially need coverage (including the integration suite and the Vogon e2e canary).\n2. **Test quality** — do the tests assert real behavior and cover edge cases (error paths, concurrent sessions, cross-platform), or are they shallow? Note tests that use the real repo/CWD or global config instead of isolated temp repos — that weakens confidence.\n3. **Gate compliance** — would `mise run check` (gofmt, golangci-lint, test:ci) plausibly pass? Are errors handled explicitly rather than ignored?\n4. **Correctness signals** — TODOs, commented-out code, ignored errors, or obviously untested risky logic in the diff.\n\nScore from 0 to 100 (higher = more confidence):\n- 0-20: No tests for the change, gates likely failing\n- 21-40: Minimal or shallow coverage\n- 41-60: Moderate coverage, gaps on critical/git-mutating paths\n- 61-80: Good coverage, gates would catch most issues\n- 81-100: Thorough coverage (unit + integration/canary where relevant), gates clean\n\nAfter your analysis, output ONLY this JSON object as the very last line:
\n{\"value\": <number 0-100>, \"rationale\": \"<1-2 sentence explanation>\"}"
  },
  "select": {
    "trigger_types": ["api", "push"]

M.entire/runners/trail-confidence.json+1/-1

16 unmodified lines

17
18
19
20
20
21
22
23

16 unmodified lines

"kind": "trail_prompt"
  },
  "prompt": {
    "template": "You are a code review assistant. Analyze the changes on branch \"{{branch}}\" compared to \"{{base_branch}}\".\n\nRun these commands:\n\n1. git diff origin/{{base_branch}}...HEAD --stat\n2. git diff origin/{{base_branch}}...HEAD\n\nIdentify the most critical areas a human reviewer should focus on. Look for:\n- Security-sensitive changes (auth, crypto, user input handling, API keys)\n- Complex logic that could have bugs\n- Database schema or migration changes\n- Breaking API changes\n- Error handling gaps\n- Performance-critical code paths\n\nOutput ONLY this JSON object as the very last line:\n\n{\"files\": [{\"path\": \"<file path>\", \"lines\": \"<optional line range, e.g. 42-58>\", \"why\": \"<brief reason>\"}]}\n\nExample:\n{\"files\": [{\"path\": \"api/src/auth.ts\", \"lines\": \"127-145\", \"why\": \"New token validation logic\"}, {\"path\": \"api/src/db/migrations/001.sql\", \"why\": \"Schema change adds nullable column\"}]}\n\nIf no critical areas need attention, output: {\"files\": []}"
    "template": "You are a code review assistant for the Entire CLI, a Go tool that manipulates the user's git repo. Analyze the changes on branch \"{{branch}}\" compared to \"{{base_branch}}\".\n\nRun these commands:\n\n1. git diff origin/{{base_branch}}...HEAD --stat\n2. git diff origin/{{base_branch}}...HEAD\n\nIdentify the most critical areas a human reviewer should focus on. Look for:\n- Destructive git ops or working-tree changes that could lose user work (reset --hard, checkout, rewind, file restore/delete)\n- Git hook handlers (prepare-commit-msg, post-commit, post-rewrite, pre-push) that run on every user commit/push\n- Checkpoint/session-state or shadow-branch logic, and transcript condensation\n- Redaction or privacy filtering that governs what is pushed to the remote\n- Auth/token handling and control-plane core resolution\n- Complex logic that could have bugs, broken invariants, or error-handling gaps\n\nOutput ONLY this JSON object as the very last line:\n\n{\"files\": [{\"path\": \"<file path>\", \"lines\": \"<optional line range, e.g. 42-58>\", \"why\": \"<brief reason>\"}]}\n\nExample:\n{\"files\": [{\"path\": \"api/src/auth.ts\", \"lines\": \"127-145\", \"why\": \"New token validation logic\"}, {\"path\": \"api/src/db/migrations/001.sql\", \"why\": \"Schema change adds nullable column\"}]}\n\nIf no critical areas need attention, output: {\"files\": []}"
  },
  "select": {
    "trigger_types": ["api", "push"]

M.entire/runners/trail-review-focus.json+1/-1

15 unmodified lines

16
17
18
19
19
20
21
22

15 unmodified lines

"kind": "trail_prompt"
  },
  "prompt": {
    "template": "You are a code risk evaluator. Analyze the changes on branch \"{{branch}}\" compared to \"{{base_branch}}\".\n\nRun these commands to gather context:\n\n1. git diff origin/{{base_branch}}...HEAD --stat\n2. git diff origin/{{base_branch}}...HEAD\n3. Check for database migration files\n4. Check for changes to authentication, authorization, or payment code\n\nThen evaluate the **Risk** of these changes (0-100). Risk means: \"What is the potential damage if something is wrong with these changes?\"\n\nConsider the following dimensions:\n\n1. **Sensitivity of touched code** — Are auth, payments, data pipelines, or security-critical paths modified?\n2. **Consumer / downstream impact** — How many users, services, or components depend on the changed code? Are public APIs or shared contracts affected?\n3. **Reversibility** — Can these changes be easily rolled back? Are there database migrations, data transformations, or external side effects that are hard to undo?\n4. **Blast radius** — Is the change isolated to one module or does it span multiple systems?\n5. **Data integrity** — Could the changes corrupt, lose, or expose sensitive data?\n\nScore from 0 to 100:\n- 0-15: Trivial (docs, comments, formatting)\n- 16-30: Low (small isolated changes, tests only)\n- 31-50: Moderate (feature additions in contained modules)\n- 51-70: Elevated (API changes, multi-module changes, new dependencies)\n- 71-85: High (schema migrations, auth changes, breaking API changes)\n- 86-100: Critical (irreversible data changes, security-sensitive code)\n\nAfter your analysis, output ONLY this JSON object as the very last line of your response:\n\n{\"value\": <number 0-100>, \"rationale\": \"<1-2 sentence explanation>\"}"
    "template": "You are a risk evaluator for the Entire CLI — a Go tool that manages agent session checkpoints by manipulating the user's git repo: it rewinds and restores files, runs destructive git ops (reset --hard, checkout), installs and runs git hooks, maintains shadow branches and session state, and condenses session transcripts (which may hold prompts, file contents, and commit messages) that get pushed to a remote. It runs on developers' machines against their real repos.\n\nAnalyze branch \"{{branch}}\" vs \"{{base_branch}}\".\n\nGather context:\n1. git diff origin/{{base_branch}}...HEAD --stat\n2. git diff origin/{{base_branch}}...HEAD\n3. Note whether the diff touches the strategy (cmd/entire/cli/strategy/), git hook handlers, destructive git ops (reset --hard, checkout, rewind/file restore), redaction or transcript condensation (what is pushed to the remote), or auth/token handling.\n\nScore the **Risk** (0-100): \"How much damage if something is wrong?\" — not whether the code is malicious (security is evaluated separately).\n\nWeigh:\n1. **Destructive repo ops** — could it lose uncommitted work, corrupt the working tree or .git, or wrongly delete ignored dirs (.entire/, .worktrees/)? reset --hard, checkout, rewind, and file restore/delete are the riskiest and largely irreversible.\n2. **Git hooks** — prepare-commit-msg, post-commit, post-rewrite, and pre-push run on every user commit and push; a bug can block their git workflow.\n3. **Checkpoint/session integrity** — corrupting or losing shadow branches, checkpoint metadata, session linkage, or condensation.\n4. **Privacy/egress** — transcripts are pushed to a remote; does the change weaken redaction or push more than intended? Egress is irreversible.\n5. **Blast radius** — local side effects (moved refs, written/deleted files, installed hooks) can't be rolled back centrally and reach every user on upgrade.\n\nBands:\n- 0-15: Trivial (docs, comments, formatting)\n- 16-30: Low (isolated changes, tests, read-only output)\n- 31-50: Moderate (new commands/flags in contained packages; no git-mutating or egress paths)\n- 51-70: Elevated (git hooks, checkpoint/session internals, auth/token handling, new dependencies)\n- 71-85: High (destructive git ops, rewind/restore, or redaction/condensation governing what is pushed)\n- 86-100: Critical (could irreversibly lose user work, corrupt a repo, or leak transcript content)\n\nOutput ONLY this JSON as the very last line:\n\n{\"value\": <number 0-100>, \"rationale\": \"<1-2 sentence explanation>\"}"
  },
  "select": {
    "trigger_types": ["api", "push"]

M.entire/runners/trail-risk.json+1/-1

15 unmodified lines

16
17
18
19
19
20
21
22

15 unmodified lines