inspect: replace judge panel + chair with a single consolidating judge · Entire

Home

Log in

inspect: replace judge panel + chair with a single consolidating judge

7be889c→main·

dipree·1mo ago·12 files·+244 added/-576 removed

The multi-judge panel and chair-merge step added a lot of surface area (schema, panel provider, chair prompt, panel/chair pickers, scripted --set-chair, legacy master fields) without improving outcomes in testing. Step back to the simpler model the feature started from: inspectors fan out, then one judge consolidates their reports into the final verdict in a closing round.

- settings: replace Judges/Chair + legacy Master/MasterAgent/MasterModel with a single optional Judge *ReviewConfig (json "judge"). - profile.go: profileJudges/profileMasterIdentity -> profileJudge; defaultReviewMaster -> defaultJudge (judgeSpec); add resolveJudge (explicit judge, else auto-selected text-gen inspector). - delete synthesis_panel.go (PanelSynthesisProvider + composeChairPrompt); runMultiAgentPath always uses AgentSynthesisProvider for the one judge. - cmd.go: --set-judge is now single (drop --set-chair); update help, --list, catalog, validation. - picker.go: promptForJudges/chair picker -> single promptForJudge; advanced picker saves Judge instead of master. - update tests and docs (architecture doc, CLAUDE.md).

Note: settings parsing is strict, so profiles still using the old judges/chair/master keys (only created on this unreleased branch) must be re-saved; a multi-inspector profile with no judge now auto-selects one.

Sessions

06e9093212ecView transcript

?\ Checkout the hand off doc that I just added.Pi·Opus 4.8·5 steps

Changes

12

497 unmodified lines

498
499
500
501
501
502
503
503
504
505
506

497 unmodified lines

### `entire inspect` Command (aliased as `entire review`)

`entire inspect` runs a named review profile: a set of **inspector** agents run a shared task in parallel, then a panel of **judges** evaluates their reports (a **chair** merges a multi-judge panel into the final verdict). Inspector sessions are immutable facts attached to checkpoints; the final verdict is stored locally in the review manifest. On the next `git commit`, inspector sessions are condensed into the checkpoint metadata, permanently recording that the code was reviewed and which skills were run.
`entire inspect` runs a named review profile: a set of **inspector** agents run a shared task in parallel, then a single **judge** consolidates their reports into the final verdict in a closing round. Inspector sessions are immutable facts attached to checkpoints; the final verdict is stored locally in the review manifest. On the next `git commit`, inspector sessions are condensed into the checkpoint metadata, permanently recording that the code was reviewed and which skills were run.

Profiles live in `.entire/settings.json` under `review_profiles`; adapter-backed inspectors (claude-code, codex, gemini-cli) receive `ENTIRE_REVIEW_*` env vars that the `UserPromptSubmit` hook reads to tag the session as `Kind = "agent_review"`. Multi-agent runs use a TUI dashboard + automatic judge-panel synthesis.
Profiles live in `.entire/settings.json` under `review_profiles`; adapter-backed inspectors (claude-code, codex, gemini-cli) receive `ENTIRE_REVIEW_*` env vars that the `UserPromptSubmit` hook reads to tag the session as `Kind = "agent_review"`. Multi-agent runs use a TUI dashboard + automatic single-judge synthesis.

See [Review Command](docs/architecture/review-command.md) for the full command surface, settings schema, env-var handshake, multi-agent UI, anti-features (do NOT recreate), and key-file map.

MCLAUDE.md+2/-2

121 unmodified lines

122
123
124
125
125
126
127
128
118 unmodified lines

247
248
249
250
250
251
252
253

121 unmodified lines

"skills": []string{"/review"},
                    },
                },
                "master": "claude-code",
                "judge": map[string]any{"agent": "claude-code"},
            },
        },
    })
118 unmodified lines

"skills": []string{"/nonexistent:skill-xyz"},
                    },
                },
                "master": "claude-code",
                "judge": map[string]any{"agent": "claude-code"},
            },
        },
    })

Mcmd/entire/cli/integration_test/review_test.go+2/-2

78 unmodified lines

79
80
81
82
83
82
83
84
85
9 unmodified lines

95
96
97
99
100
101
102
98
99
100
101
102
103
104
105
106
107
109
110
108
109
110
111
1 unmodified line

113
114
115
118
116
117
118
119
55 unmodified lines

175
176
177
180
181
178
179
180
181
11 unmodified lines

193
194
195
199
200
196
197
198
199
1 unmodified line

201
202
203
208
204
205
206
207
9 unmodified lines

217
218
219
224
225
220
221
222
223
224
225
226
232
227
228
229
230
157 unmodified lines

388
389
390
396
397
398
399
400
401
402
403
404
391
392
393
394
395
101 unmodified lines

497
498
499
512
513
514
515
516
517
500
501
502
503
504
90 unmodified lines

595
596
597
614
615
598
599
600
601
618
619
620
621
622
623
624
625
626
602
603
604
605
606
607
628
629
630
631
632
633
634
635
636
637
638
639
640
608
609
610
611
612
613
643
614
615
616
617
166 unmodified lines

784
785
786
816
817
818
819
820
787
788
789
790
791
792
822
793
794
795
796
303 unmodified lines

1100
1101
1102
1132
1133
1134
1103
1104
1105
1106
1107
1137
1138
1139
1140
1141
1142
1143
1144
1145
1146
1147
1148
1149
1150
1151
1108
1109
1110
1111
1112
1113

78 unmodified lines

var listAgents bool
    var listProfiles bool
    var setAgents []string
    var setJudges []string
    var setChair string
    var setJudge string
    var setTask string
    var setModels []string
    var setSlots []string
9 unmodified lines

Hidden: true,
        Short:  "Run a multi-agent review against the current branch",
        Long: `Run a multi-agent review against the current branch: several inspector
agents inspect the change in parallel, then a panel of judges renders the final
verdict (a chair merges the panel when there is more than one judge). Reviews are
saved as named profiles in Entire settings and clone-local preferences. On first
run, guided setup writes a profile and asks before starting agents.
agents inspect the change in parallel, then a single judge consolidates their
reports into the final verdict in a closing round. Reviews are saved as named
profiles in Entire settings and clone-local preferences. On first run, guided
setup writes a profile and asks before starting agents.

Flags:
  --configure    set up a review profile (shows available agents + profiles).
                 With --set-* flags it writes the profile non-interactively;
                 otherwise it opens the wizard (interactive) without starting agents.
  --set-agents   with --configure: comma-separated inspector agents for the profile
  --set-judge    with --configure: a judge as agent[=model] (repeatable; >1 = panel)
  --set-chair    with --configure: the judge that merges a multi-judge panel
  --set-judge    with --configure: the consolidating judge as agent[=model]
  --set-task     with --configure: the profile's canonical task text
  --set-model    with --configure: per-inspector model as agent=model (repeatable)
  --set-slot     with --configure: an inspector slot as agent[=model] (repeatable;
1 unmodified line

--edit         re-open the advanced profile skill picker
  --findings     browse local findings
  --agent NAME   run only one inspector from the selected profile
  --list         list configured inspect profiles (their inspectors and judges)
  --list         list configured inspect profiles (their inspectors and judge)
  --agents       list the inspector agents you can pass to --agent for the profile
  --model NAME   override the model for the --agent inspector (requires --agent)
  --models       list the models each agent advertises (optionally --agent NAME)
55 unmodified lines

if configure {
                return runReviewConfigure(ctx, cmd, profileName, reviewConfigureOptions{
                    Agents: setAgents,
                    Judges: setJudges,
                    Chair:  setChair,
                    Judge:  setJudge,
                    Task:   setTask,
                    Models: setModels,
                    Slots:  setSlots,
11 unmodified lines

}
    cmd.Flags().BoolVar(&configure, "configure", false, "set up a review profile; shows available agents and accepts --set-* flags for non-interactive config")
    cmd.Flags().StringSliceVar(&setAgents, "set-agents", nil, "with --configure: inspector agents for the profile (comma-separated)")
    cmd.Flags().StringArrayVar(&setJudges, "set-judge", nil, "with --configure: a judge as agent[=model] (repeatable; multiple judges form a panel)")
    cmd.Flags().StringVar(&setChair, "set-chair", "", "with --configure: the judge (agent[=model]) that merges a multi-judge panel")
    cmd.Flags().StringVar(&setJudge, "set-judge", "", "with --configure: the consolidating judge as agent[=model]")
    cmd.Flags().StringVar(&setTask, "set-task", "", "with --configure: the profile's canonical task text")
    cmd.Flags().StringArrayVar(&setModels, "set-model", nil, "with --configure: per-inspector model as agent=model (repeatable)")
    cmd.Flags().StringArrayVar(&setSlots, "set-slot", nil, "with --configure: an inspector slot as agent[=model] (repeatable; same agent/model may repeat)")
1 unmodified line

cmd.Flags().BoolVar(&findings, "findings", false, "browse local review findings")
    cmd.Flags().BoolVar(&listAgents, "agents", false, "list the inspector agents you can pass to --agent for the selected profile")
    cmd.Flags().BoolVar(&listModels, "models", false, "list the models each review agent advertises (optionally filtered by --agent)")
    cmd.Flags().BoolVar(&listProfiles, "list", false, "list configured inspect profiles (inspectors and judges)")
    cmd.Flags().BoolVar(&listProfiles, "list", false, "list configured inspect profiles (inspectors and judge)")
    cmd.Flags().StringVar(&agentOverride, "agent", "", "run one configured inspector from the selected profile")
    cmd.Flags().StringVar(&modelOverride, "model", "", "override the model for the --agent inspector (requires --agent)")
    cmd.Flags().StringVar(&profileOverride, "profile", "", "review profile to run (default: review_default_profile or general)")
9 unmodified lines

// reviewConfigureOptions carries the non-interactive `--configure` inputs.
type reviewConfigureOptions struct {
    Agents []string // inspector agent names (--set-agents)
    Judges []string // judge slots as "agent[=model]" entries (--set-judge)
    Chair  string   // chair judge as agent[=model] (--set-chair)
    Judge  string   // consolidating judge as "agent[=model]" (--set-judge)
    Task   string   // profile task text (--set-task)
    Models []string // per-inspector "agent=model" entries (--set-model)
    Slots  []string // inspector slots as "agent[=model]" entries (--set-slot)
}

func (o reviewConfigureOptions) scripted() bool {
    return len(o.Agents) > 0 || len(o.Judges) > 0 || o.Chair != "" || o.Task != "" || len(o.Models) > 0 || len(o.Slots) > 0
    return len(o.Agents) > 0 || o.Judge != "" || o.Task != "" || len(o.Models) > 0 || len(o.Slots) > 0
}

func runReviewConfigure(ctx context.Context, cmd *cobra.Command, profileOverride string, opts reviewConfigureOptions, deps Deps) error {
157 unmodified lines

}
        fmt.Fprintf(out, "    inspectors: %s\n", strings.Join(inspectors, ", "))

if judges, chair := profileJudges(p); len(judges) > 0 {
            labels := make([]string, len(judges))
            for i, j := range judges {
                labels[i] = judgeLabel(j)
                if len(judges) > 1 && i == chair {
                    labels[i] += " (chair)"
                }
            }
            fmt.Fprintf(out, "    judges:     %s\n", strings.Join(labels, ", "))
        if j, ok := profileJudge(p); ok {
            fmt.Fprintf(out, "    judge:      %s\n", judgeLabel(j))
        }
    }
    fmt.Fprintln(out)
101 unmodified lines

marker = " (default)"
            }
            line := fmt.Sprintf("  %s%s: %s", name, marker, strings.Join(sortedProfileAgentNames(p), ", "))
            if judges, _ := profileJudges(p); len(judges) > 0 {
                judgeNames := make([]string, len(judges))
                for i, j := range judges {
                    judgeNames[i] = j.agent
                }
                line += "  judges=" + strings.Join(judgeNames, ",")
            if j, ok := profileJudge(p); ok {
                line += "  judge=" + j.agent
            }
            fmt.Fprintln(out, line)
        }
90 unmodified lines

profile.Task = profileTask(profileName, settings.ReviewProfileConfig{})
    }

// Judges: explicit --set-judge defines the panel; otherwise fall back to the
    // legacy single-master default picked from the inspectors.
    // Judge: explicit --set-judge wins; otherwise a multi-inspector profile gets
    // an auto-selected judge, and a single-inspector profile needs none.
    inspectorCount := len(nonZeroAgentConfigs(profile.Agents))
    switch {
    case len(opts.Judges) > 0:
        judges := make([]settings.ReviewConfig, 0, len(opts.Judges))
        for _, raw := range opts.Judges {
            rawName, model, _ := strings.Cut(raw, "=")
            name := strings.TrimSpace(rawName)
            if name == "" {
                continue
            }
            judges = append(judges, settings.ReviewConfig{Agent: name, Model: strings.TrimSpace(model)})
    case strings.TrimSpace(opts.Judge) != "":
        rawName, model, _ := strings.Cut(opts.Judge, "=")
        name := strings.TrimSpace(rawName)
        if name == "" {
            return settings.ReviewProfileConfig{}, errors.New("--set-judge needs an agent name")
        }
        if len(judges) == 0 {
            return settings.ReviewProfileConfig{}, errors.New("--set-judge listed no usable judges")
        }
        profile.Judges = judges
        profile.Chair = strings.TrimSpace(opts.Chair)
        profile.Master = ""
        profile.MasterAgent = ""
    case inspectorCount > 1 && len(profile.Judges) == 0 && strings.TrimSpace(profile.MasterAgent) == "":
        if strings.TrimSpace(profile.Master) == "" {
            profile.Master = defaultReviewMaster(ctx, profile.Agents)
        }
        if _, _, masterErr := selectProfileWorker(profile, profile.Master); masterErr != nil {
            return settings.ReviewProfileConfig{}, fmt.Errorf("judge %q is not one of the profile inspectors (%s)", profile.Master, strings.Join(sortedProfileAgentNames(profile), ", "))
        profile.Judge = &settings.ReviewConfig{Agent: name, Model: strings.TrimSpace(model)}
    case inspectorCount > 1 && (profile.Judge == nil || profile.Judge.IsZero()):
        if j, ok := defaultJudge(ctx, profile.Agents); ok {
            profile.Judge = &settings.ReviewConfig{Agent: j.agent, Model: j.model}
        }
    case inspectorCount <= 1:
        profile.Master = ""
        profile.Judge = nil
    }
    return profile, nil
}
166 unmodified lines

fmt.Fprintln(cmd.ErrOrStderr(), err.Error())
            return silentErr(err)
        }
        // Require at least one judge. Judges that can't actually write a verdict
        // (no text generation) are tolerated here and handled at synthesis time:
        // a lone judge fails gracefully ("final report unavailable") and a panel
        // simply drops the non-responding judge.
        if judges, _ := profileJudges(profile); len(judges) == 0 {
        // Require a consolidating judge (explicit or auto-selected). A judge that
        // can't actually write a verdict (no text generation) is tolerated here and
        // handled at synthesis time, where it fails gracefully ("final report
        // unavailable").
        if _, ok := resolveJudge(ctx, profile); !ok {
            cmd.SilenceUsage = true
            err := fmt.Errorf("review profile %q has multiple inspectors but no judge; set review_profiles.%s.judges", profileName, profileName)
            err := fmt.Errorf("review profile %q has multiple inspectors but no judge that can write a verdict; set review_profiles.%s.judge", profileName, profileName)
            fmt.Fprintln(cmd.ErrOrStderr(), err.Error())
            return silentErr(err)
        }
303 unmodified lines

}
    aggregateOutput := ""

// Build the judge panel. A single judge collapses to one verdict; multiple
    // judges each render a verdict in parallel and the chair merges them.
    judges, chairIdx := profileJudges(profile)
    // Resolve the single consolidating judge (explicit or auto-selected) that
    // turns the inspectors' reports into the final verdict. Validation upstream
    // guarantees one resolves; a nil provider would simply skip synthesis.
    var synthProvider SynthesisProvider
    masterLabel := ""
    switch len(judges) {
    case 0:
        // Validation upstream guarantees at least one judge; nil = no synthesis.
    case 1:
        synthProvider = AgentSynthesisProvider{AgentName: judges[0].agent, Model: judges[0].model}
        masterLabel = judgeLabel(judges[0])
    default:
        providers := make([]SynthesisProvider, len(judges))
        labels := make([]string, len(judges))
        for i, j := range judges {
            providers[i] = AgentSynthesisProvider{AgentName: j.agent, Model: j.model}
            labels[i] = judgeLabel(j)
        }
        synthProvider = PanelSynthesisProvider{Judges: providers, Labels: labels, ChairIdx: chairIdx}
        masterLabel = fmt.Sprintf("%s (chair of %d judges)", labels[chairIdx], len(judges))
    if judge, ok := resolveJudge(ctx, profile); ok {
        synthProvider = AgentSynthesisProvider{AgentName: judge.agent, Model: judge.model}
        masterLabel = judgeLabel(judge)
    }
    sinks := composeMultiAgentSinks(multiAgentSinkInputs{
        out:               out,

Mcmd/entire/cli/review/cmd.go+40/-81

53 unmodified lines

54
55
56
57
58
59
60
61
62
63
64
58
59
60
61
62
65
66
67
68
69
67
70
71
72
73

53 unmodified lines

prefs = &settings.ClonePreferences{}
    }
    prefs.ReviewDefaultProfile = review.DefaultProfileName
    profile := settings.ReviewProfileConfig{
        Task:   "Test review task.",
        Agents: cfg,
    }
    if judge := defaultTestJudge(cfg); judge != "" {
        profile.Judge = &settings.ReviewConfig{Agent: judge}
    }
    prefs.ReviewProfiles = map[string]settings.ReviewProfileConfig{
        review.DefaultProfileName: {
            Task:   "Test review task.",
            Agents: cfg,
            Master: defaultTestMaster(cfg),
        },
        review.DefaultProfileName: profile,
    }
    return settings.SaveClonePreferences(ctx, prefs)
}

func defaultTestMaster(cfg map[string]settings.ReviewConfig) string {
func defaultTestJudge(cfg map[string]settings.ReviewConfig) string {
    if _, ok := cfg[string(agent.AgentNameClaudeCode)]; ok {
        return string(agent.AgentNameClaudeCode)
    }

Mcmd/entire/cli/review/cmd_test.go+9/-6

37 unmodified lines

38
39
40
41
41
42
43
44
8 unmodified lines

53
54
55
56
57
58
59
60
56
57
58
59
60
31 unmodified lines

92
93
94
98
95
96
100
97
98
99
103
104
105
100
101
102
107
108
109
103
104
105
106
107
112
113
114
115
116
117
118
119
120
121
122
123
108
109
110
111
112
128
129
113
114
115
116
117
118
134
119
120
121
122
1 unmodified line

124
125
126
142
143
127
128
129
130
1 unmodified line

132
133
134
151
152
135
136
137
154
155
156
157
158
159
160
161
162
163
138
139
140
141
142
143
12 unmodified lines

156
157
158
182
159
160
161
162
163
164
188
189
165
166
167
168
193
169
170
171
197
172
173
174
175
7 unmodified lines

183
184
185
211
212
186
187
188
189
190
191

37 unmodified lines

"general",
        reviewConfigureOptions{
            Agents: []string{"claude-code", "codex"},
            Judges: []string{"codex"},
            Judge:  "codex",
            Models: []string{"claude-code=opus"},
        },
        &settings.EntireSettings{},
8 unmodified lines

if got := profile.Agents["claude-code"].Model; got != "opus" {
        t.Errorf("claude-code model = %q, want opus", got)
    }
    if len(profile.Judges) != 1 || profile.Judges[0].Agent != "codex" {
        t.Errorf("judges = %#v, want one judge codex", profile.Judges)
    }
    if profile.Master != "" {
        t.Errorf("legacy master should be cleared when judges are set, got %q", profile.Master)
    if profile.Judge == nil || profile.Judge.Agent != "codex" {
        t.Errorf("judge = %#v, want codex", profile.Judge)
    }
    if profile.Task == "" {
        t.Error("task should default to the built-in general task")
31 unmodified lines

}
}

func TestProfileMasterIdentity(t *testing.T) {
func TestProfileJudge(t *testing.T) {
    t.Parallel()
    t.Run("standalone master wins and need not be a worker", func(t *testing.T) {
    t.Run("explicit judge resolves with model", func(t *testing.T) {
        t.Parallel()
        profile := settings.ReviewProfileConfig{
            Agents:      map[string]settings.ReviewConfig{tAgentCodex: {Agent: tAgentCodex}},
            MasterAgent: tAgentClaude,
            MasterModel: tModelOpus,
            Agents: map[string]settings.ReviewConfig{tAgentCodex: {Agent: tAgentCodex}},
            Judge:  &settings.ReviewConfig{Agent: tAgentClaude, Model: tModelOpus},
        }
        name, model, ok := profileMasterIdentity(profile)
        if !ok || name != tAgentClaude || model != tModelOpus {
            t.Fatalf("got (%q,%q,%v), want (claude-code, opus, true)", name, model, ok)
        j, ok := profileJudge(profile)
        if !ok || j.agent != tAgentClaude || j.model != tModelOpus {
            t.Fatalf("got (%#v,%v), want claude-code/opus, true", j, ok)
        }
    })
    t.Run("legacy worker master resolves from Agents", func(t *testing.T) {
        t.Parallel()
        profile := settings.ReviewProfileConfig{
            Agents: map[string]settings.ReviewConfig{tAgentClaude: {Agent: tAgentClaude, Model: tModelSonnet}},
            Master: tAgentClaude,
        }
        name, model, ok := profileMasterIdentity(profile)
        if !ok || name != tAgentClaude || model != tModelSonnet {
            t.Fatalf("got (%q,%q,%v), want (claude-code, sonnet, true)", name, model, ok)
        }
    })
    t.Run("no master", func(t *testing.T) {
    t.Run("no judge", func(t *testing.T) {
        t.Parallel()
        profile := settings.ReviewProfileConfig{
            Agents: map[string]settings.ReviewConfig{tAgentCodex: {Agent: tAgentCodex}},
        }
        if _, _, ok := profileMasterIdentity(profile); ok {
            t.Fatal("expected ok=false when no master is set")
        if _, ok := profileJudge(profile); ok {
            t.Fatal("expected ok=false when no judge is set")
        }
    })
}

func TestBuildConfiguredProfile_JudgePanel(t *testing.T) {
func TestBuildConfiguredProfile_Judge(t *testing.T) {
    t.Parallel()
    deps := configureTestDeps("claude-code", "codex")
    profile, err := buildConfiguredProfile(
1 unmodified line

"general",
        reviewConfigureOptions{
            Agents: []string{tAgentClaude, tAgentCodex},
            Judges: []string{tAgentClaude + "=" + tModelOpus, tAgentCodex + "=gpt-5"},
            Chair:  tAgentClaude,
            Judge:  tAgentClaude + "=" + tModelOpus,
        },
        &settings.EntireSettings{},
        deps,
1 unmodified line

if err != nil {
        t.Fatalf("buildConfiguredProfile: %v", err)
    }
    if len(profile.Judges) != 2 {
        t.Fatalf("judges = %#v, want 2", profile.Judges)
    if profile.Judge == nil || profile.Judge.Agent != tAgentClaude || profile.Judge.Model != tModelOpus {
        t.Fatalf("judge = %#v, want claude-code/opus", profile.Judge)
    }
    if profile.Judges[0].Agent != tAgentClaude || profile.Judges[0].Model != tModelOpus {
        t.Errorf("judge[0] = %#v, want claude-code/opus", profile.Judges[0])
    }
    if profile.Chair != tAgentClaude {
        t.Errorf("chair = %q, want claude-code", profile.Chair)
    }
    // profileJudges resolves the panel + chair index.
    judges, chair := profileJudges(profile)
    if len(judges) != 2 || chair != 0 {
        t.Errorf("profileJudges = (%#v, %d), want 2 judges, chair 0", judges, chair)
    j, ok := profileJudge(profile)
    if !ok || j.agent != tAgentClaude || j.model != tModelOpus {
        t.Errorf("profileJudge = (%#v,%v), want claude-code/opus, true", j, ok)
    }
}

12 unmodified lines

}
}

func TestBuildConfiguredProfile_PreservesExistingTaskAndMasterModel(t *testing.T) {
func TestBuildConfiguredProfile_PreservesExistingTask(t *testing.T) {
    t.Parallel()
    deps := configureTestDeps("claude-code", "codex")
    s := &settings.EntireSettings{
        ReviewProfiles: map[string]settings.ReviewProfileConfig{
            "general": {
                Task:        "Custom task text.",
                MasterModel: "opus",
                Task: "Custom task text.",
                Agents: map[string]settings.ReviewConfig{
                    "claude-code": {Skills: []string{"/review"}},
                },
                Master: "claude-code",
            },
        },
    }
    // Only change the worker set; task + master_model must survive.
    // Only change the worker set; the custom task must survive.
    profile, err := buildConfiguredProfile(
        context.Background(),
        "general",
7 unmodified lines

if profile.Task != "Custom task text." {
        t.Errorf("task = %q, want preserved custom task", profile.Task)
    }
    if profile.MasterModel != "opus" {
        t.Errorf("master_model = %q, want preserved opus", profile.MasterModel)
    // Two inspectors with no explicit judge → one auto-selected.
    if _, ok := profileJudge(profile); !ok {
        t.Error("expected an auto-selected judge for a multi-inspector profile")
    }
}

Mcmd/entire/cli/review/configure_test.go+26/-50

134 unmodified lines

135
136
137
138
138
139
140
141
142
143
144
145
142
143
144
145
288 unmodified lines

434
435
436
440
437
438
439
440
441
442
443
444
445
446
159 unmodified lines

606
607
608
606
607
608
609
610
609
610
611
612
613
614
615
1 unmodified line

617
618
619
618
620
621
622
621
622
623
624
625
626
627
628
623
624
625
626
627
628
629
631
632
633
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
649
650
651
652
653
654
655
656
636
637
638
657
658
659
660
661
662
640
641
663
664
665
666
643
644
645
646
647
648
649
650
651
652
653
654
655
656
657
658
659
660
661
662
663
664
665
666
667
668
669
667
668
669
670
72 unmodified lines

743
744
745
748
746
747
748
749
750
751
754
752
753
754
755
756
757
86 unmodified lines

844
845
846
847
847
848
849
850
851
851
852
853
854
21 unmodified lines

876
877
878
879
879
880
881
882
5 unmodified lines

888
889
890
891
892
893
891
892
893
894
895
896
896
897
898
899
900
901
902
903
7 unmodified lines

911
912
913
910
911
914
915
916
917
918
919
920
921
918
922
923
924
925
926
923
927
928
929
930
13 unmodified lines

944
945
946
943
947
948
949
950
17 unmodified lines

968
969
970
967
971
972
973
974
1 unmodified line

976
977
978
975
976
979
980
981
982
983
984
985
982
986
987
988
989

134 unmodified lines

profile.Task = customTask
    }
    if len(profile.Agents) > 1 {
        judges, chair, err := promptForJudges(ctx, launchable, existing)
        judge, err := promptForJudge(ctx, launchable, existing)
        if err != nil {
            return "", settings.ReviewProfileConfig{}, err
        }
        profile.Judges = judges
        profile.Chair = chair
        profile.Master = ""
        profile.MasterAgent = ""
        profile.Judge = judge
    }
    fmt.Fprintf(out, "Saved %q review profile with %s.\n", profileName, strings.Join(sortedProfileAgentNames(profile), ", "))
    fmt.Fprintln(out)
288 unmodified lines

cfg.Model = s.model
        profile.Agents[workerIDForAgentModel(s.agent, s.model, profile.Agents)] = cfg
    }
    profile.Master = defaultReviewMaster(ctx, profile.Agents)
    // A default judge is only meaningful with more than one inspector;
    // RunReviewGuidedSetup re-asks for the judge in that case anyway.
    if len(profile.Agents) > 1 {
        if j, ok := defaultJudge(ctx, profile.Agents); ok {
            profile.Judge = &settings.ReviewConfig{Agent: j.agent, Model: j.model}
        }
    }
    return profile
}

159 unmodified lines

return models
}

// promptForJudges picks the panel of judges (each its own agent + model) that
// evaluate the inspectors' reports, and the chair that merges a multi-judge
// panel. Candidates are launchable agents that can write a verdict (text
// generation). Returns (judges, chair).
func promptForJudges(ctx context.Context, launchable []string, existing settings.ReviewProfileConfig) ([]settings.ReviewConfig, string, error) {
// promptForJudge picks the single judge (agent + model) that consolidates the
// inspectors' reports into the final verdict. Candidates are launchable agents
// that can write a verdict (text generation).
func promptForJudge(ctx context.Context, launchable []string, existing settings.ReviewProfileConfig) (*settings.ReviewConfig, error) {
    candidates := make([]string, 0, len(launchable))
    for _, name := range launchable {
        if agentSupportsTextGeneration(ctx, name) {
1 unmodified line

}
    }
    if len(candidates) == 0 {
        return nil, "", errors.New("no installed agent can write a verdict")
        return nil, errors.New("no installed agent can write a verdict")
    }

// Seed from existing judges, else a single default judge.
    var seed []crewSlot
    if existJudges, _ := profileJudges(existing); len(existJudges) > 0 {
        for _, j := range existJudges {
            seed = append(seed, crewSlot(j))
        }
    } else {
        seed = []crewSlot{{agent: candidates[0]}}
    seedAgent := candidates[0]
    seedModel := ""
    if j, ok := profileJudge(existing); ok {
        seedAgent = j.agent
        seedModel = j.model
    }

slots, err := pickSlotList(ctx, "Judges", "Each judge renders a verdict; with >1 a chair merges them. + Add to add a judge.", candidates, seed)
    if err != nil {
        return nil, "", err
    agentName := candidates[0]
    if len(candidates) > 1 {
        options := make([]huh.Option[string], 0, len(candidates))
        for _, name := range candidates {
            options = append(options, huh.NewOption(labelForSimpleAgent(name), name))
        }
        picked := candidates[0]
        for _, name := range candidates {
            if name == seedAgent {
                picked = seedAgent
                break
            }
        }
        form := newAccessibleForm(huh.NewGroup(
            huh.NewSelect[string]().
                Title("Judge (writes the final verdict)").
                Description("Consolidates the inspectors' reports into one verdict.").
                Options(options...).
                Height(reviewPickerHeight(len(options))).
                Value(&picked),
        ))
        if err := form.RunWithContext(ctx); err != nil {
            return nil, fmt.Errorf("judge picker: %w", err)
        }
        agentName = picked
    }

judges := make([]settings.ReviewConfig, len(slots))
    for i, s := range slots {
        judges[i] = settings.ReviewConfig{Agent: s.agent, Model: s.model}
    // Only carry the existing model forward when the judge agent is unchanged;
    // models are agent-specific.
    seededModel := ""
    if agentName == seedAgent {
        seededModel = seedModel
    }
    if len(judges) < 2 {
        return judges, "", nil
    model, err := promptCrewModel(ctx, agentName, seededModel)
    if err != nil {
        return nil, err
    }

// Pick the chair from the panel.
    options := make([]huh.Option[string], 0, len(judges))
    for _, j := range judges {
        options = append(options, huh.NewOption(judgeLabel(judgeSpec{agent: j.Agent, model: j.Model}), chairKey(j)))
    }
    picked := chairKey(judges[0])
    form := newAccessibleForm(huh.NewGroup(
        huh.NewSelect[string]().
            Title("Chair (writes the final verdict)").
            Description("Merges the panel's verdicts, surfacing agreement and dissent.").
            Options(options...).
            Height(reviewPickerHeight(len(options))).
            Value(&picked),
    ))
    if err := form.RunWithContext(ctx); err != nil {
        return nil, "", fmt.Errorf("chair picker: %w", err)
    }
    return judges, picked, nil
}

// chairKey is the Chair selector for a judge: "agent" or "agent:model".
func chairKey(j settings.ReviewConfig) string {
    if strings.TrimSpace(j.Model) != "" {
        return j.Agent + ":" + j.Model
    }
    return j.Agent
    return &settings.ReviewConfig{Agent: agentName, Model: model}, nil
}

func ConfirmRunReviewNow(ctx context.Context, out io.Writer) (bool, error) {
72 unmodified lines

// see why, but keep going with an empty prefill — runReview already
    // surfaces the same error distinctly when it's the first load.
    existing := map[string]settings.ReviewConfig{}
    existingMaster := ""
    existingJudge := ""
    if s, err := settings.Load(ctx); err != nil {
        logging.Warn(ctx, "settings.Load failed when pre-filling picker", slog.String("error", err.Error()))
    } else if s != nil {
        if profile, ok := s.ReviewProfiles[profileName]; ok {
            existing = profile.Agents
            existingMaster = profile.Master
            if j, jok := profileJudge(profile); jok {
                existingJudge = j.agent
            }
        }
    }

86 unmodified lines

return nil, errors.New("no review skills or prompt configured")
    }

masterAgent, err := pickReviewMasterAgentPreference(ctx, merged, existingMaster)
    judgeAgent, err := pickReviewJudgeAgentPreference(ctx, merged, existingJudge)
    if err != nil {
        return nil, err
    }
    if err := saveReviewProfileConfig(ctx, profileName, merged, masterAgent); err != nil {
    if err := saveReviewProfileConfig(ctx, profileName, merged, judgeAgent); err != nil {
        return nil, err
    }
    fmt.Fprintf(out, "Saved review profile %q to local review preferences. Edit later with `entire inspect --edit --profile %s`.\n", profileName, profileName)
21 unmodified lines

return merged
}

func saveReviewProfileConfig(ctx context.Context, profileName string, agents map[string]settings.ReviewConfig, master string) error {
func saveReviewProfileConfig(ctx context.Context, profileName string, agents map[string]settings.ReviewConfig, judgeAgent string) error {
    prefs, err := settings.LoadClonePreferences(ctx)
    if err != nil {
        return fmt.Errorf("load review preferences before save: %w", err)
5 unmodified lines

prefs.ReviewProfiles = map[string]settings.ReviewProfileConfig{}
    }
    // Merge into any existing profile so the advanced skills picker only
    // rewrites what it actually edits (agents + master). Profile-level fields
    // the picker never surfaces — custom `task` text and `master_model` — are
    // preserved instead of being clobbered with built-in defaults.
    // rewrites what it actually edits (agents + judge). Profile-level fields the
    // picker never surfaces — custom `task` text — are preserved instead of being
    // clobbered with built-in defaults.
    profile := prefs.ReviewProfiles[profileName]
    profile.Agents = agents
    profile.Master = master
    if strings.TrimSpace(judgeAgent) != "" {
        profile.Judge = &settings.ReviewConfig{Agent: strings.TrimSpace(judgeAgent)}
    } else {
        profile.Judge = nil
    }
    if strings.TrimSpace(profile.Task) == "" {
        profile.Task = profileTask(profileName, settings.ReviewProfileConfig{})
    }
7 unmodified lines

return nil
}

func pickReviewMasterAgentPreference(ctx context.Context, review map[string]settings.ReviewConfig, current string) (string, error) {
    choices := reviewMasterAgentChoices(review)
func pickReviewJudgeAgentPreference(ctx context.Context, review map[string]settings.ReviewConfig, current string) (string, error) {
    choices := reviewJudgeAgentChoices(review)
    switch len(choices) {
    case 0:
        return current, nil
    case 1:
        return choices[0].Name, nil
    default:
        return promptForReviewMasterAgent(ctx, choices, current)
        return promptForReviewJudgeAgent(ctx, choices, current)
    }
}

// defaultAgentPick returns the saved choice if it is still offered, otherwise
// the first choice. Shared by the master picker.
// the first choice. Shared by the judge picker.
func defaultAgentPick(choices []AgentChoice, saved string) string {
    if pick, ok := savedAgentPick(choices, saved); ok {
        return pick
13 unmodified lines

return "", false
}

func reviewMasterAgentChoices(configured map[string]settings.ReviewConfig) []AgentChoice {
func reviewJudgeAgentChoices(configured map[string]settings.ReviewConfig) []AgentChoice {
    choices := make([]AgentChoice, 0, len(configured))
    for name, cfg := range configured {
        if cfg.IsZero() {
17 unmodified lines

return choices
}

func promptForReviewMasterAgent(ctx context.Context, choices []AgentChoice, saved string) (string, error) {
func promptForReviewJudgeAgent(ctx context.Context, choices []AgentChoice, saved string) (string, error) {
    options := make([]huh.Option[string], 0, len(choices))
    for _, choice := range choices {
        options = append(options, huh.NewOption(choice.Label, choice.Name))
1 unmodified line

picked := defaultAgentPick(choices, saved)
    form := newAccessibleForm(huh.NewGroup(
        huh.NewSelect[string]().
            Title("Choose chair judge").
            Description("The chair critically evaluates the inspectors' reports and writes the final verdict.").
            Title("Choose judge").
            Description("The judge critically evaluates the inspectors' reports and writes the final verdict.").
            Options(options...).
            Height(reviewPickerHeight(len(options))).
            Value(&picked),
    ))
    if err := form.RunWithContext(ctx); err != nil {
        return "", fmt.Errorf("review master picker: %w", err)
        return "", fmt.Errorf("review judge picker: %w", err)
    }
    return picked, nil
}

Mcmd/entire/cli/review/picker.go+77/-73

145 unmodified lines

146
147
148
149
150
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
156
157
158
159
160
161
162
184
185
163
164
165
166
187
188
189
190
191
192
193
194
167
168
169
197
198
199
200
201
202
203
170
171
172
173
174
175
205
206
207
208
176
177
178
179
106 unmodified lines

286
287
288
321
289
290
291
324
325
292
293
294
295
296
297
298
299
28 unmodified lines

328
329
330
360
331
332
333
334
335
336
337
338
339
365
340
341
342
343
344
370
371
345
346
347
348
349
374
350
351
352
353

145 unmodified lines

return strings.Join(parts, "")
}

// judgeSpec is one resolved judge: the agent that renders a verdict plus its
// optional model.
// judgeSpec is the resolved consolidating judge: the agent that renders the
// final verdict plus its optional model.
type judgeSpec struct {
    agent string
    model string
}

// profileJudges resolves the panel of judges and the index of the chair (the
// judge that merges a multi-judge panel; 0 for a single judge). Resolution
// order: explicit Judges, then the legacy standalone MasterAgent, then the
// legacy worker Master. Returns an empty slice when the profile has no judge.
func profileJudges(profile settings.ReviewProfileConfig) ([]judgeSpec, int) {
    if len(profile.Judges) > 0 {
        judges := make([]judgeSpec, 0, len(profile.Judges))
        for _, cfg := range profile.Judges {
            name := strings.TrimSpace(cfg.Agent)
            if name == "" {
                continue
            }
            judges = append(judges, judgeSpec{agent: name, model: strings.TrimSpace(cfg.Model)})
        }
        if len(judges) == 0 {
            return nil, 0
        }
        chair := 0
        if sel := strings.TrimSpace(profile.Chair); sel != "" {
            for i, j := range judges {
                if j.agent == sel || j.agent+":"+j.model == sel {
                    chair = i
                    break
                }
            }
        }
        return judges, chair
// profileJudge resolves the configured consolidating judge. ok is false when
// the profile has no judge set (a single-inspector profile, or one left to the
// runtime default); callers fall back to resolveJudge for the default pick.
func profileJudge(profile settings.ReviewProfileConfig) (judgeSpec, bool) {
    if profile.Judge == nil {
        return judgeSpec{}, false
    }
    if ma := strings.TrimSpace(profile.MasterAgent); ma != "" {
        return []judgeSpec{{agent: ma, model: strings.TrimSpace(profile.MasterModel)}}, 0
    name := strings.TrimSpace(profile.Judge.Agent)
    if name == "" {
        return judgeSpec{}, false
    }
    if workerName, cfg, err := selectProfileWorker(profile, profile.Master); err == nil {
        model := strings.TrimSpace(profile.MasterModel)
        if model == "" {
            model = strings.TrimSpace(cfg.Model)
        }
        return []judgeSpec{{agent: reviewAgentName(workerName, cfg), model: model}}, 0
    }
    return nil, 0
    return judgeSpec{agent: name, model: strings.TrimSpace(profile.Judge.Model)}, true
}

// profileMasterIdentity reports the representative judge (the chair, or the
// single judge) and whether the profile has any judge. Kept for callers that
// need a single "who decides" identity (validation, labels, picker preselect).
func profileMasterIdentity(profile settings.ReviewProfileConfig) (string, string, bool) {
    judges, chair := profileJudges(profile)
    if len(judges) == 0 {
        return "", "", false
// resolveJudge returns the judge to use for a fan-out run: the explicitly
// configured judge, or an auto-selected text-gen inspector when none is set.
func resolveJudge(ctx context.Context, profile settings.ReviewProfileConfig) (judgeSpec, bool) {
    if j, ok := profileJudge(profile); ok {
        return j, true
    }
    if chair < 0 || chair >= len(judges) {
        chair = 0
    }
    return judges[chair].agent, judges[chair].model, true
    return defaultJudge(ctx, profile.Agents)
}

// judgeLabel renders a judge for UI output: "agent" or "agent · model".
106 unmodified lines

if len(agents) == 0 {
        return settings.ReviewProfileConfig{}, errors.New("no agents with review runner adapters and hooks installed; run `entire configure --agent claude-code`, `entire configure --agent codex`, or `entire configure --agent gemini`")
    }
    return settings.ReviewProfileConfig{
    profile := settings.ReviewProfileConfig{
        Task:   profileTask(profileName, settings.ReviewProfileConfig{}),
        Agents: agents,
        Master: defaultReviewMaster(ctx, agents),
    }, nil
    }
    if j, ok := defaultJudge(ctx, agents); ok {
        profile.Judge = &settings.ReviewConfig{Agent: j.agent, Model: j.model}
    }
    return profile, nil
}

func defaultReviewAgentConfig(profileName, agentName string) settings.ReviewConfig {
28 unmodified lines

}
}

func defaultReviewMaster(ctx context.Context, configured map[string]settings.ReviewConfig) string {
// defaultJudge auto-selects a consolidating judge from the configured
// inspectors: it prefers claude-code, then codex, then gemini, and otherwise
// takes the first inspector that can write a verdict (text generation). ok is
// false when no inspector can.
func defaultJudge(ctx context.Context, configured map[string]settings.ReviewConfig) (judgeSpec, bool) {
    for _, preferred := range []string{string(agent.AgentNameClaudeCode), string(agent.AgentNameCodex), string(agent.AgentNameGemini)} {
        for _, workerName := range sortedReviewConfigKeys(configured) {
            cfg := configured[workerName]
            if reviewAgentName(workerName, cfg) == preferred && agentSupportsTextGeneration(ctx, preferred) {
                return workerName
                return judgeSpec{agent: preferred, model: strings.TrimSpace(cfg.Model)}, true
            }
        }
    }
    for _, workerName := range sortedReviewConfigKeys(configured) {
        if agentSupportsTextGeneration(ctx, reviewAgentName(workerName, configured[workerName])) {
            return workerName
        cfg := configured[workerName]
        if name := reviewAgentName(workerName, cfg); agentSupportsTextGeneration(ctx, name) {
            return judgeSpec{agent: name, model: strings.TrimSpace(cfg.Model)}, true
        }
    }
    return ""
    return judgeSpec{}, false
}

func sortedReviewConfigKeys(configured map[string]settings.ReviewConfig) []string {

Mcmd/entire/cli/review/profile.go+34/-58

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125

package review

import (
    "context"
    "errors"
    "fmt"
    "strings"
    "sync"
)

// PanelSynthesisProvider runs a panel of judges over the inspectors' reports.
// Each judge independently produces a verdict from the same synthesis prompt;
// when two or more verdicts come back, the chair judge merges them into one
// final verdict and the individual verdicts are appended as a panel.
//
// It implements SynthesisProvider, so the existing SynthesisSink uses it
// unchanged — a panel is just a provider that happens to consult several
// judges. A single judge collapses to that judge's verdict (today's behavior).
type PanelSynthesisProvider struct {
    Judges   []SynthesisProvider // one per judge; index aligns with Labels
    Labels   []string            // display label per judge (e.g. "codex · gpt-5")
    ChairIdx int                 // index of the judge that merges the panel
}

// Synthesize fans out to each judge in parallel, then has the chair merge the
// verdicts. Failed judges are dropped; if only one verdict survives it is
// returned directly. If every judge fails, the error is returned so the caller
// can report "final report unavailable".
func (p PanelSynthesisProvider) Synthesize(ctx context.Context, prompt string) (string, error) {
    switch len(p.Judges) {
    case 0:
        return "", errors.New("no judges configured")
    case 1:
        return p.Judges[0].Synthesize(ctx, prompt) //nolint:wrapcheck // transparent single-judge passthrough
    }

verdicts := make([]string, len(p.Judges))
    errs := make([]error, len(p.Judges))
    var wg sync.WaitGroup
    for i := range p.Judges {
        wg.Add(1)
        go func(i int) {
            defer wg.Done()
            verdicts[i], errs[i] = p.Judges[i].Synthesize(ctx, prompt)
        }(i)
    }
    wg.Wait()

ok := make([]int, 0, len(p.Judges))
    for i := range verdicts {
        if errs[i] == nil && strings.TrimSpace(verdicts[i]) != "" {
            ok = append(ok, i)
        }
    }
    switch len(ok) {
    case 0:
        return "", fmt.Errorf("all judges failed: %w", firstNonNilErr(errs))
    case 1:
        return verdicts[ok[0]], nil
    }

// Pick the chair; fall back to the first successful judge if the configured
    // chair failed or is out of range. The chair participates as a full panel
    // judge (it produced its own verdict above) and then runs a second time to
    // merge the panel — these are deliberately distinct calls: an independent
    // verdict, then a reconciliation over all verdicts.
    chair := p.ChairIdx
    if chair < 0 || chair >= len(p.Judges) || errs[chair] != nil || strings.TrimSpace(verdicts[chair]) == "" {
        chair = ok[0]
    }

final, err := p.Judges[chair].Synthesize(ctx, composeChairPrompt(verdicts, p.Labels, ok))
    if err != nil || strings.TrimSpace(final) == "" {
        // Chair merge failed: surface the panel rather than nothing.
        final = "The judges could not be merged into a single verdict; see each judge's verdict below."
    }

var b strings.Builder
    b.WriteString(strings.TrimSpace(final))
    b.WriteString("\n\n## Panel\n\n")
    for _, i := range ok {
        fmt.Fprintf(&b, "### %s\n\n%s\n\n", p.labelAt(i), strings.TrimSpace(verdicts[i]))
    }
    return b.String(), nil
}

func (p PanelSynthesisProvider) labelAt(i int) string {
    if i >= 0 && i < len(p.Labels) && strings.TrimSpace(p.Labels[i]) != "" {
        return p.Labels[i]
    }
    return fmt.Sprintf("judge %d", i+1)
}

// composeChairPrompt instructs the chair judge to reconcile the panel's
// verdicts into one final verdict, explicitly surfacing agreement and dissent.
func composeChairPrompt(verdicts, labels []string, ok []int) string {
    var b strings.Builder
    b.WriteString("You are the presiding judge on a panel reviewing a code change. " +
        "The judges' verdicts are below. Write the single final verdict, nothing else:\n" +
        "  - One line: the decision and a one-sentence reason.\n" +
        "  - Then a short bullet list of the agreed actionable findings, most important first, one line each. " +
        "Note a disagreement only when it changes the decision, and resolve it (which judge is right and why).\n\n" +
        "Merge duplicates, drop unsupported claims, don't concatenate or repeat the verdicts. " +
        "No preamble or headings; be proportional to the change.\n\n")
    for _, i := range ok {
        label := fmt.Sprintf("judge %d", i+1)
        if i < len(labels) && strings.TrimSpace(labels[i]) != "" {
            label = labels[i]
        }
        fmt.Fprintf(&b, "## Verdict from %s\n\n%s\n\n", label, strings.TrimSpace(verdicts[i]))
    }
    return b.String()
}

func firstNonNilErr(errs []error) error {
    for _, e := range errs {
        if e != nil {
            return e
        }
    }
    return errors.New("unknown synthesis error")
}

// Compile-time check.
var _ SynthesisProvider = PanelSynthesisProvider{}

Dcmd/entire/cli/review/synthesis_panel.go-125

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97

package review

import (
    "context"
    "errors"
    "strings"
    "testing"
)

// stubJudge is a SynthesisProvider returning canned output (or an error).
type stubJudge struct {
    out      string
    err      error
    lastSeen string // the prompt it was asked to synthesize
    calls    int
}

func (s *stubJudge) Synthesize(_ context.Context, prompt string) (string, error) {
    s.calls++
    s.lastSeen = prompt
    return s.out, s.err
}

func TestPanel_SingleJudgePassesThrough(t *testing.T) {
    t.Parallel()
    j := &stubJudge{out: "verdict A"}
    p := PanelSynthesisProvider{Judges: []SynthesisProvider{j}, Labels: []string{"a"}}
    got, err := p.Synthesize(context.Background(), "PROMPT")
    if err != nil || got != "verdict A" {
        t.Fatalf("got (%q,%v), want (verdict A, nil)", got, err)
    }
    if j.lastSeen != "PROMPT" {
        t.Errorf("single judge should get the raw prompt, got %q", j.lastSeen)
    }
}

func TestPanel_MultiJudgeChairMerges(t *testing.T) {
    t.Parallel()
    j1 := &stubJudge{out: "ship it"}
    j2 := &stubJudge{out: "block: race in cache.go"}
    chair := &stubJudge{out: "FINAL: block — j2 found a real race"}
    // chair is judge index 2.
    p := PanelSynthesisProvider{
        Judges:   []SynthesisProvider{j1, j2, chair},
        Labels:   []string{"claude", "codex", "chair"},
        ChairIdx: 2,
    }
    got, err := p.Synthesize(context.Background(), "PROMPT")
    if err != nil {
        t.Fatalf("unexpected error: %v", err)
    }
    if !strings.Contains(got, "FINAL: block") {
        t.Errorf("final verdict should come from the chair, got:\n%s", got)
    }
    // Panel appendix shows each judge's verdict.
    for _, want := range []string{"## Panel", "ship it", "block: race in cache.go"} {
        if !strings.Contains(got, want) {
            t.Errorf("output missing %q:\n%s", want, got)
        }
    }
    // The chair was asked to merge (its prompt mentions the other verdicts).
    if !strings.Contains(chair.lastSeen, "ship it") || !strings.Contains(chair.lastSeen, "race in cache.go") {
        t.Errorf("chair prompt should include the panel verdicts, got:\n%s", chair.lastSeen)
    }
}

func TestPanel_DroppedFailuresCollapseToSingle(t *testing.T) {
    t.Parallel()
    good := &stubJudge{out: "only verdict"}
    bad := &stubJudge{err: errors.New("boom")}
    p := PanelSynthesisProvider{
        Judges:   []SynthesisProvider{bad, good},
        Labels:   []string{"bad", "good"},
        ChairIdx: 0, // chair failed → falls back to the surviving judge
    }
    got, err := p.Synthesize(context.Background(), "PROMPT")
    if err != nil {
        t.Fatalf("unexpected error: %v", err)
    }
    if got != "only verdict" {
        t.Errorf("one survivor should pass through without a panel, got:\n%s", got)
    }
}

func TestPanel_AllFailErrors(t *testing.T) {
    t.Parallel()
    p := PanelSynthesisProvider{
        Judges: []SynthesisProvider{
            &stubJudge{err: errors.New("e1")},
            &stubJudge{err: errors.New("e2")},
        },
        Labels: []string{"a", "b"},
    }
    if _, err := p.Synthesize(context.Background(), "PROMPT"); err == nil {
        t.Fatal("expected an error when every judge fails")
    }
}

Dcmd/entire/cli/review/synthesis_panel_test.go-97

500 unmodified lines

501
502
503
504
504
505
506
507

500 unmodified lines

if err := os.MkdirAll(entireDir, 0o750); err != nil {
        t.Fatalf("create .entire dir: %v", err)
    }
    settingsJSON := `{"enabled":true,"review_default_profile":"general","review_profiles":{"general":{"task":"Test review task.","agents":{"claude-code":{"skills":["/review"]}},"master":"claude-code"}}}` + "\n"
    settingsJSON := `{"enabled":true,"review_default_profile":"general","review_profiles":{"general":{"task":"Test review task.","agents":{"claude-code":{"skills":["/review"]}},"judge":{"agent":"claude-code"}}}}` + "\n"
    if err := os.WriteFile(filepath.Join(entireDir, "settings.json"), []byte(settingsJSON), 0o600); err != nil {
        t.Fatalf("write review settings: %v", err)
    }

Mcmd/entire/cli/review_context_test.go+1/-1

248 unmodified lines

249
250
251
252
252
253
254
255
254
255
256
257
258
2 unmodified lines

261
262
263
264
264
265
267
266
267
268
269
271
272
270
271
272
273
274
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
275
276
277
278
279
280
281
282
283
284
299
300
285
286
287
288

248 unmodified lines

}

// ReviewProfileConfig is a named review setup. The profile-level Task is the
// canonical task every worker agent is asked to run; per-agent ReviewConfig
// canonical task every inspector agent is asked to run; per-agent ReviewConfig
// entries adapt that task to agent-specific mechanics such as slash commands
// or additional instructions. Master names the agent that consolidates worker
// outputs into the final report.
// or additional instructions. Judge names the single agent that consolidates
// the inspectors' reports into the final verdict in a closing round.
//
// Example:
//
2 unmodified lines

//	    "task": "Review this change for auth, injection, secrets, and privilege-boundary bugs.",
//	    "agents": {
//	      "claude-sonnet": {"agent": "claude-code", "model": "sonnet", "skills": ["/security-review"]},
//	      "claude-opus": {"agent": "claude-code", "model": "opus", "skills": ["/security-review"]},
//	      "codex": {"model": "gpt-5-codex", "skills": ["/review"], "prompt": "Focus on security."}
//	    },
//	    "master": "claude-sonnet"
//	    "judge": {"agent": "claude-code", "model": "opus"}
//	  }
//	}
//
// MasterModel is an optional model hint passed to the master agent's text
// generation API.
// ReviewProfileConfig is intentionally small: the review package owns built-in
// default task text for conventional profile names like "general".
type ReviewProfileConfig struct {
    Task   string                  `json:"task,omitempty"`
    Agents map[string]ReviewConfig `json:"agents,omitempty"`
    // Judges is the panel of judges (each an agent + model) that independently
    // evaluate the inspectors' reports and render verdicts. Supersedes
    // MasterAgent/Master. One judge = a single verdict; multiple judges = a
    // panel reconciled by the Chair.
    Judges []ReviewConfig `json:"judges,omitempty"`
    // Chair, when set and there are >=2 judges, names the judge (by agent or
    // agent:model) that merges the panel verdicts into the final verdict.
    Chair string `json:"chair,omitempty"`
    // Master is the legacy single-master selector: a worker key in Agents that
    // also writes the final report. Superseded by Judges/MasterAgent; still
    // honored when those are empty.
    Master string `json:"master,omitempty"`
    // MasterAgent is the legacy standalone single master. Superseded by Judges;
    // still honored when Judges is empty.
    MasterAgent string `json:"master_agent,omitempty"`
    // MasterModel is the model for the legacy master (either form).
    MasterModel string `json:"master_model,omitempty"`
    // Judge is the single agent (plus optional model) that consolidates the
    // inspectors' reports into the final verdict. It is optional: a
    // one-inspector profile needs no judge (the lone report is the result),
    // and a multi-inspector profile with no judge set falls back to an
    // auto-selected inspector that can write a verdict.
    Judge *ReviewConfig `json:"judge,omitempty"`
}

// IsZero reports whether the profile is effectively unset.
func (c ReviewProfileConfig) IsZero() bool {
    return c.Task == "" && len(c.Agents) == 0 && len(c.Judges) == 0 &&
        c.Chair == "" && c.Master == "" && c.MasterAgent == "" && c.MasterModel == ""
    return c.Task == "" && len(c.Agents) == 0 && (c.Judge == nil || c.Judge.IsZero())
}

// ReviewConfig holds one worker's configuration within a review profile.

Mcmd/entire/cli/settings/settings.go+11/-26

2 unmodified lines

3
4
5
6
7
8
9
6
7
8
9
10
3 unmodified lines

14
15
16
19
17
18
19
20
23
21
22
23
24
17 unmodified lines

42
43
44
47
48
45
46
47
48
51
52
53
54
49
50
51
52
53
54
55
5 unmodified lines

61
62
63
66
67
68
64
65
66
67
68
69
10 unmodified lines

80
81
82
85
83
84
85
86
1 unmodified line

88
89
90
93
94
95
96
97
91
92
93
94
4 unmodified lines

99
100
101
108
109
110
111
112
113
114
102
103
104
105
106
107
108
109
23 unmodified lines

133
134
135
144
145
146
136
137
138
139
140
141
30 unmodified lines

172
173
174
183
184
185
186
187
188
189
190
191
192
175
176
177
178
179
180
181
182
183
184
185
10 unmodified lines

196
197
198
209
210
211
212
199
200
201
202
203
204
17 unmodified lines

222
223
224
236
237
225
226
227
228
229
230
231
242
232
233
244
245
246
247
248
249
234
235
236
237
238
239
13 unmodified lines

253
254
255
269
256

2 unmodified lines

`entire inspect` (aliased as `entire review`) runs a named review profile. A
profile defines one canonical task (for example `general`, `security`, or
`accessibility`), a set of **inspector** agents that all run that task, and a
panel of **judges** that evaluate the inspectors' reports. With one judge that
judge writes the verdict; with two or more, each judge renders an independent
verdict and the **chair** merges them into one final verdict, surfacing
agreement and dissent. Inspector sessions are immutable facts attached to
single **judge** that consolidates the inspectors' reports into the final
verdict in a closing round. Inspector sessions are immutable facts attached to
checkpoints; the final verdict is stored locally in the review manifest for
findings/fix workflows.

3 unmodified lines

entire inspect                          # Interactive: pick a profile to run. Non-interactive: list profiles + error
entire inspect security                 # Run a named profile
entire inspect --profile accessibility  # Same, flag form
entire inspect --list                   # List configured profiles (inspectors + judges), marking the default
entire inspect --list                   # List configured profiles (inspectors + judge), marking the default
entire inspect --configure                    # Interactive: guided wizard. Non-interactive: list agents + profiles
entire inspect --configure --profile general --set-agents claude-code,codex --set-judge claude-code
                                               # Configure a profile non-interactively (no TUI)
entire inspect --configure --profile sec --set-slot claude-code=opus --set-slot codex --set-judge claude-code=opus --set-judge codex=gpt-5 --set-chair claude-code=opus
entire inspect --configure --profile sec --set-slot claude-code=opus --set-slot codex --set-judge claude-code=opus
entire inspect --configure --profile general --set-model codex=gpt-5-codex --set-task "..."
entire inspect --edit --profile general       # Advanced skill-level config (skill picker)
entire inspect --agent <name>           # Run one inspector from the selected profile
17 unmodified lines

setup: choose a review focus (or `Custom…` to write the task), build the
inspector crew (a single-screen add/edit/remove slot list seeded with all
launchable agents — the same agent may appear more than once on different or
identical models), then choose the judges (another slot list) and, when there
are ≥2, a chair. It saves the profile and asks before starting agents.
identical models), then choose the judge that consolidates their reports. It
saves the profile and asks before starting agents.

`entire inspect --configure` is the configuration entry point:
- With `--set-agents` / `--set-slot` / `--set-judge` / `--set-chair` /
  `--set-task` / `--set-model agent=model`, it writes the profile
  non-interactively (no TUI). `--set-*` writes preserve profile-level fields the
  flags don't touch (custom `task`, etc.).
- With `--set-agents` / `--set-slot` / `--set-judge` / `--set-task` /
  `--set-model agent=model`, it writes the profile non-interactively (no TUI).
  `--set-*` writes preserve profile-level fields the flags don't touch (custom
  `task`, etc.).
- With no `--set-*` flags in an interactive terminal, it opens the guided
  wizard (which already lists the selectable agents).
- With no `--set-*` flags in a non-interactive context, it prints the discovery
5 unmodified lines

When two or more adapter-backed inspectors are configured and `--agent` is not
set, `entire inspect` fans out to all configured inspectors. There is no per-run
multi-picker: the profile is the fan-out contract. Multi-inspector profiles must
resolve at least one judge; the judges run after inspectors finish and produce
the final verdict.
multi-picker: the profile is the fan-out contract. Multi-inspector profiles
resolve one judge (explicit, or auto-selected from the inspectors); the judge
runs after the inspectors finish and produces the final verdict.

## Settings Schema

10 unmodified lines

"claude-code": {"skills": ["/review"]},
        "codex": {"skills": ["/review"]}
      },
      "judges": [{"agent": "claude-code", "model": "opus"}]
      "judge": {"agent": "claude-code", "model": "opus"}
    },
    "security": {
      "task": "Review this change for auth, injection, secrets, and privilege-boundary bugs.",
1 unmodified line

"claude-sonnet": {"agent": "claude-code", "model": "sonnet", "skills": ["/security-review"]},
        "codex": {"model": "gpt-5-codex", "skills": ["/review"], "prompt": "Focus on security."}
      },
      "judges": [\
        {"agent": "claude-code", "model": "opus"},\
        {"agent": "codex", "model": "gpt-5"}\
      ],
      "chair": "claude-code:opus"
      "judge": {"agent": "claude-code", "model": "opus"}
    }
  }
}
4 unmodified lines

the agent name; to run the same agent more than once, use aliases and set
  `agent` plus `model`. Per-inspector `skills`, `prompt`, and `model` adapt the
  task to agent-specific mechanics.
- `judges` is the panel: each entry is an agent (+ optional model) that renders
  a verdict. A judge need not be one of the inspectors. `chair` (an `agent` or
  `agent:model` selector) names the judge that merges a ≥2-judge panel; it
  defaults to the first judge.
- **Back-compat:** the legacy `master` (an inspector id) and `master_agent` /
  `master_model` fields are still honored as a single judge when `judges` is
  empty. `entire inspect --configure` writes `judges`/`chair` going forward.
- `judge` is the single agent (+ optional model) that consolidates the
  inspectors' reports into the final verdict. It need not be one of the
  inspectors. It is optional: a one-inspector profile needs none, and a
  multi-inspector profile with no judge set auto-selects a text-gen-capable
  inspector (preferring claude-code, then codex, then gemini).

`entire inspect --models` lists the models each agent advertises via the
optional `agent.ModelLister` capability (`cmd/entire/cli/agent/model_lister.go`).
23 unmodified lines

a `PendingReviewMarker` file and prints guidance — the user opens the agent
   themselves and runs the skills, then tags it with `entire attach --review`.
4. Inspectors run the selected profile's task; each session ends naturally.
5. In multi-inspector profiles, the judge panel runs after inspectors finish
   (see Multi-Agent UI). Each judge receives all inspector reports and renders a
   verdict; the chair merges a ≥2-judge panel into the final verdict.
5. In multi-inspector profiles, the judge runs after inspectors finish (see
   Multi-Agent UI). It receives all inspector reports and consolidates them into
   the final verdict.
6. On the next `git commit`, the PostCommit hook condenses inspector sessions
   into the checkpoint on `entire/checkpoints/v1`, with `Kind`, `ReviewSkills`,
   and `ReviewPrompt` recorded in `CommittedMetadata`.
30 unmodified lines

goroutine; events fan into a single dispatch loop so the serial-dispatch
  contract holds. Per-inspector skills/prompts are injected via
  `perAgentConfiguredReviewer`.
- **Judge resolution** (`profile.go`): `profileJudges` returns the resolved
  panel `[]judgeSpec` and the chair index (explicit `judges`, else the legacy
  `master_agent`, else the legacy worker `master`). `profileMasterIdentity`
  reports the chair/single judge for callers that need one identity.
- **Panel synthesis** (`synthesis_panel.go`): `PanelSynthesisProvider`
  implements `SynthesisProvider`, so `SynthesisSink` consumes it unchanged. It
  fans out to each judge in parallel; one judge collapses to that verdict; with
  ≥2 the chair merges the verdicts (`composeChairPrompt`) and the individual
  verdicts are appended as a `## Panel`. Failed judges are dropped; if all fail
  the error surfaces as "final report unavailable".
- **Judge resolution** (`profile.go`): `profileJudge` returns the explicitly
  configured judge (`judge`); `resolveJudge` falls back to `defaultJudge`, which
  auto-selects a text-gen-capable inspector (preferring claude-code, then codex,
  then gemini) when none is set.
- **Synthesis** (`synthesis_sink.go`): the single judge is an
  `AgentSynthesisProvider` consumed by `SynthesisSink`. It receives all
  inspector narratives and writes one verdict; provider failure surfaces as
  "final report unavailable".
- **Env-var contract** (`env.go`): single source of truth for `ENTIRE_REVIEW_*`.
- **Scope detection** (`scope.go`): first existing of
  `origin/HEAD → origin/main → origin/master → main → master`, overridable via
10 unmodified lines

below rather than overlapping.
- **`SynthesisSink`** (`synthesis_sink.go`): after the dump it composes an
  adjudication prompt from all inspector narratives + per-run prompt + profile
  task and calls its `SynthesisProvider`. For a single judge that's an
  `AgentSynthesisProvider`; for a panel it's a `PanelSynthesisProvider` (judges
  in parallel → chair merge). Skipped when cancelled or fewer than 2 inspectors
  produced usable output. Provider failures degrade gracefully.
  task and calls its `SynthesisProvider` — an `AgentSynthesisProvider` for the
  resolved judge. Skipped when cancelled or fewer than 2 inspectors produced
  usable output. Provider failures degrade gracefully.
- **Sink composition** (`composeMultiAgentSinks` in `cmd.go`): pure helper
  taking explicit `isTTY`/`canPrompt` so tests don't depend on real TTY
  detection.
17 unmodified lines

parsers own their format; shared code only sees `Event` variants)
- Fabricated "example" model lists for agents without an enumeration command
  (codex/gemini advertise nothing; Default + Custom only)
- A single "master" that both reviews and adjudicates as a worker slot (judges
  are a separate panel; a worker is never implicitly the judge)
- A "master" worker slot that both reviews and adjudicates in one pass (the
  judge is a separate consolidation round, even when auto-selected from the
  inspectors)

## Key Files

- `cmd/entire/cli/review/cmd.go` — `NewCommand()`, `runReview` dispatch fork,
  `runReviewListProfiles` (`--list`), judge-panel wiring, `composeMultiAgentSinks`
  `runReviewListProfiles` (`--list`), judge wiring, `composeMultiAgentSinks`
- `cmd/entire/cli/review/picker.go` — guided setup, focus picker (presets +
  custom task), `pickSlotList` (inspectors + judges), chair picker, profile
  chooser
- `cmd/entire/cli/review/profile.go` — profile resolution, `profileJudges`,
  default tasks
- `cmd/entire/cli/review/synthesis_panel.go` — `PanelSynthesisProvider` +
  `composeChairPrompt`
  custom task), `pickSlotList` (inspectors), `promptForJudge`, profile chooser
- `cmd/entire/cli/review/profile.go` — profile resolution, `profileJudge` /
  `resolveJudge` / `defaultJudge`, default tasks
- `cmd/entire/cli/review/synthesis_sink.go` / `synthesis_prompt.go` — final
  verdict sink + adjudication prompt
- `cmd/entire/cli/review/marker_fallback.go` — manual fallback for agents
13 unmodified lines

- `cmd/entire/cli/attach.go` — `entire attach --review` (post-hoc tagging;
  consumes a pending-review marker)
- `cmd/entire/cli/settings/settings.go` — `ReviewProfileConfig` (`Agents`,
  `Judges`, `Chair`, legacy `Master`/`MasterAgent`/`MasterModel`)
  `Judge`)

Mdocs/architecture/review-command.md+42/-55