Fix Session Owner Tracking and Finalization · Entire

Here's trail 583:

Session liveness: detect crashed 'active' sessions via PID + start-time

Sev Location Issue
Medium resume_picker.go:264 buildCheckpointBranchIndex Walks up to 50 commits per branch with no dedup → O(branches × commits) on repos with hundreds of branches. Suggests a seen map / early-stop.
Low resume_picker.go branchCheckedOutElsewhere Silently returns ("", false) on git failure/timeout; a stale curPath from a prior iteration could cause a false match. Suggests clearing curPath on parse errors.
Low resume.go:174 resumeByCheckpointID Reads checkpoint metadata before verifying existence, triggering a possibly-unnecessary remote metadata check. Suggests a cheaper existence check first, or documenting it as intentional.

Want me to dig into any of these? I can pull up the actual diff on that branch and assess whether the medium finding is worth fixing, or just walk through the code at those locations. Let me know how deep you want to go.


Done. Summary of what happened:

Branch reset — soph/session-liveness-pid was a stale branch holding 3 early commits of the resume-picker work (and 196 commits behind). That work was already merged into main via PR #1445 (soph/resume-session-picker), with follow-up commits that addressed review feedback. I reset the branch to origin/main (f7fe32a3c).

Implementation check — all 3 findings are resolved in the merged code:

Findings dismissed — all 3 dismissed with reasons recording where each was fixed. Trail 583 now has 0 open findings.


The plan is revised with Codex's feedback incorporated. Here's a summary of what changed:

Codex review — accepted and folded in:


Session liveness — implemented on soph/session-liveness-pid

What it does: an ACTIVE session whose owning agent process is gone (clean /exit, crash, kill, closed terminal, reboot) is now detected immediately and finalized, instead of lingering as "active" for the 1-hour StuckActiveThreshold.

New cmd/entire/cli/proclive package (stdlib + x/sys/unix leaf, no import cycle):

Wiring:

Verification: mise run fmt && lint clean; mise run test:ci green (unit + integration + 59 Vogon + 4 external-agent canary); plus a live smoke test confirming entire status finalized a planted exited session (phase: ended, fully_condensed: true).


Trail 583 now has a body describing the feature and is moved from draft → open.


Four meaningful, dependency-ordered commits** (each builds independently) on soph/session-liveness-pid:

  1. proclive: add process-liveness package
  2. session: record owning process and detect exited sessions
  3. strategy: capture session owner at each turn start
  4. status, doctor: finalize sessions whose agent has exited

Codex re-review (second pass) confirmed the earlier fixes held and found 4 new issues — all fixed and folded into the relevant commits:

# Issue Fix
1 captureSessionOwner left a stale (dead) Owner when resolution failed → could finalize a now-live session Clear Owner first, set only on success (commit 3)
2 TOCTOU race: sweep acted on a stale list snapshot, ending a session a concurrent turn could have revived Threaded an optional guard through markSessionEnded/endSessionNow; the sweep re-checks OwnerExited() on the freshly-loaded state under the lock (commit 4)
3 Local-dev go run wrapper recorded go as owner (exits immediately → false finalize in dev) Added go to the transient-skip set (commit 1)
4 Check fell through to Alive when host/boot lookups failed Fail closed: unverifiable host/boot → Unknown (commit 1)

Added regression tests for the under-lock revalidation (TestFinalizeExitedSessions_RevalidatesUnderLock) and the go-transient case.


CI lint now passes. Here's everything I did:

PR review comments (PR #1488)

Two bots reviewed; all findings fixed:

Summary of what changed: