feat(trail): scaffold default runners when tune runs in a repo with none · Entire

feat(trail): scaffold default runners when tune runs in a repo with none

6599ade·

Soph·3w ago·11 files·+456 added/-3 removed

entire trail tune now doubles as onboarding. In a repo with no .entire/runners/*.json, it offers to create the default set first (interactive confirmation, or --yes for non-interactive/CI runs) and then tailors them as usual — so a single tune --run bootstraps a repo end to end.

The defaults are embedded in the binary (new runnerdefaults package: the 7 canonical runners with correct contract fields + generic templates) rather than generated by the model, so the structural schema (output adapters, result types, runtime/automation) is always valid and only the templates get tailored. Non-interactive runs without --yes error rather than silently scaffolding.

Tests cover the embedded set's validity/completeness and the create/no-op paths of ensureRunnersPresent.

Sessions

Changes

11

// Package runnerdefaults embeds the canonical generic trail runner configs, so
// `entire trail tune` can scaffold them into a repository that has none yet.
// These are the structural contract (output adapters, result types, runtime)
// plus generic prompt templates; tune tailors the templates to the repo.
package runnerdefaults

import (
    "embed"
    "fmt"
    "io/fs"
    "path"
)

//go:embed runners/*.json
var runnersFS embed.FS

// File is one default runner config: its base filename and raw JSON bytes.
type File struct {
    Name string
    Data []byte
}

// Files returns the embedded default runner configs, sorted by name.
func Files() ([]File, error) {
    entries, err := fs.ReadDir(runnersFS, "runners")
    if err != nil {
        return nil, fmt.Errorf("reading embedded runner defaults: %w", err)
    }
    out := make([]File, 0, len(entries))
    for _, e := range entries {
        if e.IsDir() {
            continue
        }
        data, err := runnersFS.ReadFile(path.Join("runners", e.Name()))
        if err != nil {
            return nil, fmt.Errorf("reading embedded runner %s: %w", e.Name(), err)
        }
        out = append(out, File{Name: e.Name(), Data: data})
    }
    return out, nil
}

Changes in runnerdefaults

trail-confidence.json

{
  "id": "trail-confidence",
  "display_name": "Confidence Eval",
  "enabled": true,
  "scope": "trail",
  "runtime": {
    "kind": "prompt_runner",
    "agent": "claude",
    "timeout_ms": 300000,
    "sandbox": {
      "base_template": "claude",
      "repo_token": "read"
    }
  },
  "automation": {
    "kind": "trail_prompt"
  },
  "prompt": {
    "template": "You are a code confidence evaluator. Analyze the changes on branch \"{{branch}}\" compared to \"{{base_branch}}\".\n\nRun these commands to gather context:\n\n1. git diff origin/{{base_branch}}...HEAD --stat\n2. git diff origin/{{base_branch}}...HEAD\n3. Look for CI configuration files (.github/workflows/, .circleci/, etc.)\n4. Look for test files related to the changed code\n\nThen evaluate the **Confidence** of these changes (0-100). Confidence means: \"How sure can we be that these changes are correct and well-tested?\"\n\nScore from 0 to 100:\n- 0-20: No tests, no CI, high uncertainty\n- 21-40: Minimal test coverage, basic CI\n- 41-60: Moderate coverage, some gaps in critical paths\n- 61-80: Good coverage, CI catches most issues\n- 81-100: Excellent coverage, comprehensive tests for all changes\n\nAfter your analysis, output ONLY this JSON object as the very last line of your response:\n\n{"value": <number 0-100>, "rationale": \"<1-2 sentence explanation>\"}"
  },
  "select": {
    "trigger_types": ["api", "push"]
  },
  "output": {
    "adapter": "last_json_line",
    "result_type": "trail_monitor",
    "trail_monitor": {
      "key": "confidence",
      "label": "Confidence",
      "value_type": "percent",
      "polarity": "higher_is_better"
    }
  }
}

trail-drift.json

{
  "id": "trail-drift",
  "display_name": "Drift Eval",
  "enabled": true,
  "scope": "trail",
  "runtime": {
    "kind": "prompt_runner",
    "agent": "claude",
    "timeout_ms": 300000,
    "sandbox": {
      "base_template": "claude",
      "repo_token": "read"
    }
  },
  "automation": {
    "kind": "trail_prompt"
  },
  "prompt": {
    "template": "You are a code drift evaluator. Analyze the changes on branch \"{{branch}}\" compared to \"{{base_branch}}\".\n\nRun these commands to gather context:\n\n1. git diff origin/{{base_branch}}...HEAD --stat\n2. git diff origin/{{base_branch}}...HEAD\n3. Look at the project structure and README for architectural patterns\n4. Look at nearby files to understand existing conventions\n\nThen evaluate the **Drift** of these changes (0-100). Drift means: \"How much do these changes deviate from the project's established patterns and intended direction?\"\n\nScore from 0 to 100:\n- 0-15: No drift — changes are perfectly aligned with existing patterns\n- 16-30: Minimal drift — minor style inconsistencies\n- 31-50: Moderate drift — some new patterns introduced but justified\n- 51-70: Significant drift — multiple deviations from conventions\n- 71-85: High drift — fundamentally different approach from existing code\n- 86-100: Extreme drift — changes are inconsistent with the project's direction\n\nAfter your analysis, output ONLY this JSON object as the very last line of your response:\n\n{"value": <number 0-100>, "rationale": \"<1-2 sentence explanation>\"}"
  },
  "select": {
    "trigger_types": ["api", "push"]
  },
  "output": {
    "adapter": "last_json_line",
    "result_type": "trail_monitor",
    "trail_monitor": {
      "key": "drift",
      "label": "Drift",
      "value_type": "percent",
      "polarity": "lower_is_better"
    }
  }
}

trail-review-focus.json

{
  "id": "trail-review-focus",
  "display_name": "Review Focus",
  "enabled": true,
  "scope": "trail",
  "runtime": {
    "kind": "prompt_runner",
    "agent": "claude",
    "model": "haiku",
    "timeout_ms": 300000,
    "sandbox": {
      "base_template": "claude",
      "repo_token": "read"
    }
  },
  "automation": {
    "kind": "trail_prompt"
  },
  "prompt": {
    "template": "You are a code review assistant. Analyze the changes on branch \"{{branch}}\" compared to \"{{base_branch}}\".\n\nRun these commands:\n\n1. git diff origin/{{base_branch}}...HEAD --stat\n2. git diff origin/{{base_branch}}...HEAD\n\nIdentify the most critical areas a human reviewer should focus on. Look for:\n- Security-sensitive changes (auth, crypto, user input handling, API keys)\n- Complex logic that could have bugs\n- Database schema or migration changes\n- Breaking API changes\n- Error handling gaps\n- Performance-critical code paths\n\nOutput ONLY this JSON object as the very last line:\n\n{"files": [{"path": \"<file path>\", "lines": \"<optional line range, e.g. 42-58>\", "why": \"<brief reason>\"}]}

Example:
{"files": [{"path": "api/src/auth.ts", "lines": "127-145", "why": "New token validation logic"}, {"path": "api/src/db/migrations/001.sql", "why": "Schema change adds nullable column"}]}

If no critical areas need attention, output: {"files": []}"
  },
  "select": {
    "trigger_types": ["api", "push"]
  },
  "output": {
    "adapter": "last_json_line",
    "result_type": "trail_review_focus"
  }
}

trail-review.json

{
  "id": "trail-review",
  "display_name": "Trail Review",
  "enabled": true,
  "scope": "trail",
  "runtime": {
    "kind": "prompt_runner",
    "agent": "claude",
    "model": "sonnet",
    "timeout_ms": 900000,
    "sandbox": {
      "base_template": "claude",
      "repo_token": "read",
      "auto_stop_minutes": 20
    }
  },
  "automation": {
    "kind": "trail_prompt"
  },
  "prompt": {
    "template": "You are reviewing the changes on branch \"{{branch}}\" against \"{{base_branch}}\". There are most likely issues in the current implementation. Raise comments for real bugs, regressions, incorrect assumptions, broken invariants, security issues, data-loss risks, or missing guards. Tie them to concrete code in the diff if possible. Each finding must be classifiable as high, medium, or low severity. If you cannot honestly assign a severity, do not raise it.\n\nReturn findings as native Entire code review comments. Each comment must target a changed line on the RIGHT side of the diff. \n\nOutput ONLY this JSON object as the very last line:\n\n{"summary":"","comments":[{"severity":"<high|medium|low>","confidence":<0-1>,"body":"<concise comment>","location":{"granularity":"line","file_path":"<file path>","start_line":<final right-side line number>}}]}
\nIf there are no actionable findings, output: {"summary":"","comments":[]}"  
  },
  "select": {
    "trigger_types": ["push"]
  },
  "debounce_ms": 10000,
  "trails_review": {
    "enabled": true
  },
  "output": {
    "adapter": "last_json_line",
    "result_type": "code_review_comments"
  }
}

trail-risk.json

{
  "id": "trail-risk",
  "display_name": "Risk Eval",
  "enabled": true,
  "scope": "trail",
  "runtime": {
    "kind": "prompt_runner",
    "agent": "claude",
    "timeout_ms": 300000,
    "sandbox": {
      "base_template": "claude",
      "repo_token": "read"
    }
  },
  "automation": {
    "kind": "trail_prompt"
  },
  "prompt": {
    "template": "You are a code risk evaluator. Analyze the changes on branch \"{{branch}}\" compared to \"{{base_branch}}\".\n\nRun these commands to gather context:\n\n1. git diff origin/{{base_branch}}...HEAD --stat\n2. git diff origin/{{base_branch}}...HEAD\n3. Check for database migration files\n4. Check for changes to authentication, authorization, or payment code\n\nThen evaluate the **Risk** of these changes (0-100). Risk means: \"What is the potential damage if something is wrong with these changes?\"\n\nScore from 0 to 100:
- 0-15: Trivial (docs, comments, formatting)
- 16-30: Low (small isolated changes, tests only)
- 31-50: Moderate (feature additions in contained modules)
- 51-70: Elevated (API changes, multi-module changes, new dependencies)
- 71-85: High (schema migrations, auth changes, breaking API changes)
- 86-100: Critical (irreversible data changes, security-sensitive code)
\nAfter your analysis, output ONLY this JSON object as the very last line of your response:\n\n{"value": <number 0-100>, "rationale": \"<1-2 sentence explanation>\"}"
  },
  "select": {
    "trigger_types": ["api", "push"]
  },
  "output": {
    "adapter": "last_json_line",
    "result_type": "trail_monitor",
    "trail_monitor": {
      "key": "risk",
      "label": "Risk",
      "value_type": "percent",
      "polarity": "lower_is_better"
    }
  }
}

trail-security.json

{
  "id": "trail-security",
  "display_name": "Security Review",
  "enabled": true,
  "scope": "trail",
  "runtime": {
    "kind": "prompt_runner",
    "agent": "claude",
    "timeout_ms": 300000,
    "sandbox": {
      "base_template": "claude",
      "repo_token": "read"
    }
  },
  "automation": {
    "kind": "trail_prompt"
  },
  "prompt": {
    "template": "You are a security auditor performing adversarial code review. Your job is to assume the author of these changes may be acting maliciously and evaluate the changes on branch \"{{branch}}\" compared to \"{{base_branch}}\" for security threats.\n\nRun these commands to gather context:\n\n1. git diff origin/{{base_branch}}...HEAD --stat\n2. git diff origin/{{base_branch}}...HEAD\n3. Check for new or modified dependency files (package.json, pnpm-lock.yaml, go.mod, requirements.txt, Cargo.toml, etc.)\n4. Check for changes to CI/CD configuration files (.github/workflows/, Dockerfile, etc.)\n5. Look for any new network calls, URL references, or fetch/request invocations\n\nThen evaluate the **Security Risk** of these changes (0-100).\n\nScore from 0 to 100 (higher = more suspicious):\n- 0-10: Clean (no security-relevant changes detected)\n- 11-25: Low (minor changes near security boundaries but no clear risk)\n- 26-50: Moderate (security-relevant code touched, warrants careful review)\n- 51-70: Elevated (suspicious patterns detected, multiple security areas affected)\n- 71-85: High (strong indicators of intentionally malicious or dangerously insecure code)\n- 86-100: Critical (clear evidence of backdoors, exfiltration, or supply chain compromise)\n
After your analysis, output ONLY this JSON object as the very last line of your response:\n
{"value": <number 0-100>, "rationale": "<1-2 sentence explanation>"}
  },
  "select": {
    "trigger_types": ["api", "push"]
  },
  "output": {
    "adapter": "last_json_line",
    "result_type": "trail_monitor",
    "trail_monitor": {
      "key": "security",
      "label": "Security",
      "value_type": "percent",
      "polarity": "lower_is_better"
    }
  }
}

trail-summary.json

{
  "id": "trail-summary",
  "display_name": "Trail Summary",
  "enabled": true,
  "scope": "trail",
  "runtime": {
    "kind": "prompt_runner",
    "agent": "claude",
    "model": "haiku",
    "timeout_ms": 300000,
    "sandbox": {
      "base_template": "claude",
      "repo_token": "read"
    }
  },
  "automation": {
    "kind": "trail_prompt"
  },
  "prompt": {
    "template": "You summarize code changes for reviewers. Analyze branch \"{{branch}}\" compared to \"{{base_branch}}\".\n\nRun these commands to gather context (do not narrate that you are running them):\n\n1. git log origin/{{base_branch}}..HEAD --oneline\n2. git diff origin/{{base_branch}}...HEAD --stat\n3. git diff origin/{{base_branch}}...HEAD\n\nWrite a short Problem → Solution summary in plain language as Markdown.\n\nReturn Markdown only. Do not include a top-level title, JSON, or code fences."
  },
  "select": {
    "trigger_types": ["api", "push"]
  },
  "output": {
    "adapter": "markdown",
    "result_type": "trail_summary",
    "trail_field": "body"
  }
}