Merge pull request #1666 from entireio/feat/review-live-tokens · Entire
Merge pull request #1666 from entireio/feat/review-live-tokens
380c265→main·
feat(review): live token counts for claude and codex reviewers
Changes
7
cmd/entire/cli/agent
claudecode
Mreviewer.go+51/-4
Mreviewer_test.go+100/-3
testdata
Astream_with_deltas.jsonl+7
codex
Areview_tokens.go+195
Areview_tokens_test.go+423
Mreviewer.go+97/-17
Mreviewer_test.go+58/-1
// Emits Started first, Finished{Success:...} last (success follows result.is_error).
// On a scanner error (torn stream), emits RunError then Finished{Success:false}.
//
// Tokens are emitted only at the terminal `result` envelope, not
// incrementally — claude's per-assistant `usage` fields aren't cumulative
// and summing them across messages would double-count.
// Live-token semantics: Claude's assistant envelopes carry a usage snapshot
// taken at the START of each API call — input_tokens/cache_* are populated
// but output_tokens is essentially zero (a 1–8 token "initial decision"
// count that does not update as text streams). The true output is only
// surfaced on `result` (aggregate across all calls in the run) or on the
// late `message_delta` event of --include-partial-messages mode.
//
// The Tokens contract (types/reviewer.go) is cumulative running totals, so
// the parser accumulates the input sum across unique message ids (the same
// usage block repeats verbatim on every content-block envelope of one API
// call — summing per envelope would multi-count) and emits
// `Tokens{In: <running sum>, Out: 0}` once per new message id. The running
// sum converges to the `result` aggregate, which is emitted last with the
// true {In, Out}. Out stays 0 mid-run because consumers render every Tokens
// event the same way — surfacing the 1–8 token stub would display a
// misleading real-looking output count.
//
// Package-private; called directly from this package's tests so they can
// drive raw stdout fixtures through the parser without going through the
```
## Type Definitions
type claudeMessage struct {
ID string `json:"id"`
Content []claudeBlock `json:"content"`
Usage messageUsage `json:"usage"`
}
type claudeBlock struct {
}
## Test Cases
- TestParseClaudeOutput_EmitsCumulativeInputDuringRun captures the live-token
contract for Claude. The `Tokens` type is documented as cumulative running
totals (each emission replaces the previous), so mid-run emissions must be
running sums, not per-call snapshots. Claude's assistant envelopes carry a
usage block per API call (repeated verbatim on every content-block envelope
of the same message id), where output_tokens is a 1–8 token "initial
decision" stub — so the parser accumulates input across unique message ids,
emits `Tokens{In: <running sum>, Out: 0}`, and lets the terminal `result`
envelope deliver the true {In, Out} aggregate.
- TestParseClaudeOutput_UsagelessResultDoesNotClobberCumulative pins the
terminal emission guard: a result envelope with no/zero usage must not
emit Tokens{0,0} — under the consumers' overwrite-not-sum semantics that
would erase the mid-run cumulative input total.