Conversation
# Conflicts: # scripts/file-size-allowlist.txt # tests/unit/lib/cloister/verification-worker-supervisor.test.ts
Resolves 46 conflicts between PAN-3668 (Prime Agent harness) and main's OpenCode harness work (PAN-3783) plus the ~50 commits that followed. Every harness enumeration is resolved as a union: opencode (main) and prime-agent (this branch) both land in Harness, RuntimeName, KNOWN_HARNESSES, HarnessMarker, the context-layer/telemetry/flywheel literal sets, CLI --harness help, and the Settings/ModelPicker lists. Where the two sides refactored the same code, main's newer structure wins and prime-agent is re-added in main's idiom: - sync.ts: keeps main's writeContextArtifactSync and the removal of project CLAUDE.md/AGENTS.md writing; prime-agent global render re-added in that idiom. - conversation-runtime.ts: keeps main's Kimi resume contract, codex prompt files and managedStateKey; prime-agent fields folded into each branch. - spawn-planning-session.ts: keeps main's delegation to claudeSystemPromptFiles, typed RuntimeName so prime-agent is accepted. - system-prerequisites/harness-binary: keeps main's ExecutableResolution and the WSL-interop guard (PAN-3827) alongside the prime-agent probe. - branding/ModelPicker tests: keep main's stronger icon-render assertion and HARNESS_OPTIONS-derived count, extended to prime-agent. Conflicts where this branch had already replaced an inline harness union with RuntimeName/Harness keep that refactor, which subsumes main's opencode addition via KNOWN_HARNESSES. normalizeResolution is exported so the prime-agent doctor consumes main's widened resolver result instead of duplicating the rule. Verified: typecheck, lint, build, slash-command regeneration, the Prime no-loss audit, contracts, and the branding/ModelPicker suites. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The origin/main merge resolution adopted main's writeContextArtifactSync for every rendered global layer, including the new prime-agent artifact. That left writeRenderedGlobalContext with no callers anywhere in the tree — it duplicated the same change-detect-then-write rule. Removing it keeps one writer for Overdeck-owned context artifacts. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The test-skip verification gate (PAN-3847) landed on main on 2026-09-17,
after this branch wrote the Prime Agent smoke on 2026-08-12. Its
`describe.skipIf(!runLive)` guard is exactly the shape the new gate
forbids, so merging main in turned a legitimate file into a required-gate
failure.
Rather than work around the gate, use the mechanism vitest.config.ts
already provides for opt-in suites: rename the file to
`prime-agent-smoke.slow.test.ts`. The slow lane is excluded from
`npm test`, from CI, and from verification runs unless
VITEST_INCLUDE_SLOW=1, so the suite needs no skip guard at all and the
gate has nothing to flag.
The env guard is replaced by a beforeAll that calls
requireHarnessBinary('prime-agent'). An enabled run without the binary
now fails immediately with the installation guidance instead of spawning
a shell that never produces a host and timing out 120s later inside
waitUntil().
docs/prime-agent-verification.md is updated to the new path, the new run
recipe, and the fail-fast behavior. It also drops
OVERDECK_PRIME_AGENT_PROVIDER, which the doc claimed but no source file
has ever read.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
src/lib/cloister/verification-artifact.ts writes each gate run to `.overdeck/verification/<timestamp>-<sha>.json`, but .gitignore only covered `.overdeck/verification-latest.json`. Every workspace that runs the verification gate is therefore left with an untracked directory that `pan done` sees as a dirty tree. The sibling artifact one line above is already ignored and the block's own comment says workspace runtime stays local, so this is the same rule applied to the directory form. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Running the suite in the slow lane surfaced the failure mode the skip guard used to hide: the host does not come up, waitUntil burns its full 120s, and throws "Timed out waiting for Prime Agent production host" with nothing to act on. The host was spawned with stdio: 'inherit', so whatever it printed about the cause went to a stream vitest discards. Spawn it through a helper that pipes and records stdout/stderr, and include the tail of that output in the timeout error. The spawned process is `dist/prime-agent-host.js`, whose RPC channel is its own child pipe and its HTTP interface, so piping the wrapper's own streams does not touch the protocol. Removing the skip guard is only an improvement if an enabled run explains itself; this is the other half of that change. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…s audit The PAN-3668 branch widened the `--changed origin/main` gate set, and two host-dependent failures came with it. `makeDbLive` opened `overdeck.db` without running the init migration, and `openDatabase` creates the file when it is missing. Building the layer against a not-yet-created path therefore left a table-less database behind; the next read-only open of that same path failed its schema audit with "overdeck.db schema is incompatible", surfacing inside whatever unrelated code read next (resources routes, `pan start`). It now runs the same idempotent `runOverdeckMigrationSync` as the sync door. The liveness no-loss audit already mocked `getDashboardApiUrlSync` so the no-resume probe could not reach a live dashboard, but the probe short-circuits on `OVERDECK_NO_RESUME` before it ever gets there. The verification gate inherits that variable from a dashboard booted with --no-resume, so every agent row carried gatingReason "Boot --no-resume" and the fixture comparison became machine-dependent. The test now clears and restores the variable alongside HOME. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Review CHANGES REQUESTED for PAN-3668Review — PAN-3668Verdict: CHANGES REQUESTED —
|
| Group | Files | Disposition |
|---|---|---|
New Prime core (src/lib/prime-agent/*: jsonl-framing, rpc-client, host, host-http, policy, provider-map, session-resume, session-controller, launch-command) |
9 | Reviewed line by line — sources of the blocker and N1–N5, N10 |
Runtime adapter + registry (src/lib/runtimes/prime-agent.ts, index.ts, types.ts) |
3 | Reviewed — source of the blocker, N7 |
Contracts (types.ts, harness-behavior.ts, flywheel.ts, artifacts.ts, telemetry.ts, context-layers.ts, composer-commands.generated.ts) |
7 | Reviewed — additive, every pre-existing literal preserved |
Config (schema.ts, defaults.ts, merge.ts, roles.ts) |
4 | Reviewed — primeAgent normalized with a validated positive-integer timeout and no model field |
| Delivery / conversation / transcript / cost wiring | 10 | Reviewed — source of N3, N8 |
| Context layers + sync | 4 | Reviewed — prime-agent-global.md rendered and injected via --append-system-prompt; nothing writes ~/.prime/agent |
| Dashboard server routes + services | 7 | Reviewed — source of N7 |
| Dashboard frontend | 13 | Reviewed — source of N6 |
CLI (index.ts, doctor.ts, doctor-prime-agent.ts, strike.ts, artifacts.ts, flywheel.ts) |
6 | Reviewed |
| Harness binary / prerequisites / launcher generator | 3 | Reviewed — source of N9 |
| Tests (unit, integration, contracts) | 31 | Reviewed — source of N1, N2, N8 |
Docs + evidence (harnesses.mdx, context-layers.mdx, MODEL_ROUTING.md, HARNESSES.md, TELEMETRY.md, no-loss audit, verification, screenshot) |
8 | Reviewed — source of N5 |
Build / infra (tsdown.config.ts, .gitignore, guard-agent-dir-removal.sh, allowlists, baselines) |
5 | Reviewed — prime-agent-host entry added; guard allowlist scoped to the two file removals |
Type-widening-only touches (RuntimeName substituted for inline harness unions across cloister, planning, specialists, lifecycle) |
12 | Reviewed — mechanical, no behavior change |
Four dimensions
- Correctness — one blocker (kill path) plus N1–N3, N10. Framing, id correlation, pending-map bound, process-exit rejection, and resume verification are otherwise sound.
resolveAllowedHarnesswidening toKNOWN_HARNESSESadds exactly one literal (prime-agent) over the old inline allowlist; no unintended harness became selectable. - Security — no blocking finding. The host binds a unix socket under
~/.overdeck/socketswith0700dirs and a per-launch 32-byte token in a0600file, checks the token before acting, caps request bodies at 1 MiB, and caps concurrency at 8. The child is spawned with an argument array, never a shell string (host.ts:86), satisfyingrpc-client.ac4.quoteShell/shellQuoteinlaunch-command.ts:13andprime-agent.ts:76implement the standard POSIX'\''escape correctly for the launcher script. Token comparison is non-constant-time, which is within Overdeck's documented trusted-local threat model and not a finding. No secret is logged: framing errors report byte counts, not record contents. - Performance — N7 only, advisory with measurements. The RPC client bounds its pending map (128) and record size (8 MiB);
host.tsdebounces stats writes at 250 ms and serializes them through a promise chain. NoexecSyncwas added to any server-reachable path. The spawn readiness loop polls with awaited 100 ms sleeps, not busy-waiting. - Requirements / UX — one AC failure (blocking), plus partial misses in N2 (
rpc-framing.ac3), N4 (provider-map.ac2), N6 (dashboard-frontend.ac1), N8 (no-loss-audit.ac3). Every other criterion I could check statically or by test is met.
AC-to-evidence matrix (by plan item)
| Item | Status | Evidence |
|---|---|---|
| contracts-harness-literal | met | types.ts union + KNOWN_HARNESSES; RuntimeName widened; typecheck green |
| contracts-behavior-discriminators | met | all six discriminator unions carry their Prime literal |
| contracts-prime-behavior-record | met | PRIME_AGENT_BEHAVIOR uses only Prime values; registered in BEHAVIORS and getHarnessBehavior |
| contracts-enumerations | met | flywheel, artifacts, telemetry, context-preview all include prime-agent; prior literals intact |
| cli-harness-flag-help | met | --harness help updated on plan/start/strike; composer manifest regenerated |
| config-schema-prime-agent | met | primeAgent section parses binaryPath/rpcStartupTimeoutMs, no model field; merge test covers default, override, and rejection |
| provider-map | partially met | ac1/ac3 met; ac2 unmet (no credential check names a variable or command) — N4 |
| harness-binary-resolution | met | HARNESS_BINARY_BY_RUNTIME + configuredHarnessBinaryPath + "Prime Agent" label in requireHarnessBinary |
| doctor-prime-checks | met | binary, semver range, --mode rpc probe, credential check; results are warn not error, consistent with the optional-harness precedent |
| rpc-framing | partially met | ac1 (real U+2028/U+2029 bytes in the fixture) and ac2 met; ac3 half-unmet — N2 |
| rpc-client | met | correlation, event channel, process-exit rejection, fake-timer timeouts, pending bound, argument-array spawn |
| runtime-adapter | ac3 unmet | registry + heartbeat met; killAgent — blocking finding |
| session-resume-recovery | met | missing/unreadable file and id mismatch both stop the process and raise PrimeAgentResumeError |
| managed-session-policy | met | all five policy rules present; 16 daemon commands denied at rpc-client.request; checkpoint recorded as policy-enforced |
| work-agent-launch | met | spawn-prime-agent.test.ts asserts explicit provider/model/session-dir and no claude-code fallback |
| conversation-launch | met | resolveAllowedHarness accepts it; preparePrimeAgentConversationLaunch injects the policy |
| message-delivery | ac3 unmet | ac1/ac2 met via get_state.isStreaming; ac3 tested only against the unreachable registry — blocker + N1 |
| transcript-adapter-prime | met | 'prime-agent': primeAgentAdapter registered; fixture renders user/assistant/thinking/tool/compaction/error |
| cost-usage-mapping | met | parser omits absent fields rather than substituting; reconcile inertness noted in N8 |
| dashboard-api-projection | met | both payloads project through the runtime read door; no route parses Prime session files directly |
| dashboard-frontend | partially met | picker, filters, labels, token/cost surfaces present; badge glyph duplicated — N6 |
| context-renderer-prime | met | prime-agent-global.md rendered by pan sync, injected via --append-system-prompt, preview enumerated; nothing touches ~/.prime/agent |
| docs-setup-reference | partially met | install, pinning, binaryPath, auth, mapping, recovery, security boundary, disabled automation, uninstall, subscription-billing caveat all present; one overstated resume claim — N5 |
| no-loss-audit | partially met | ac1/ac2 met; ac3 has a hole — N8 |
| verification-gauntlet | unverified at this HEAD | docs/prime-agent-verification.md records a full gate run, a throwaway Node 22 dashboard boot, a live smoke, and a Playwright screenshot — dated 2026-08-12, five commits before this HEAD. Not reused as exact-HEAD evidence. |
Verification
Run at HEAD 37430f70:
| Check | Result |
|---|---|
npm run typecheck |
Passed (exit 0). Dashboard guard: 26 known errors, none new. Frontend guard: 0 known, none new. |
npm run lint:effect-diagnostics |
Passed — 308 known findings, no NEW: diagnostics. |
Focused Prime suites (15 files: src/lib/prime-agent/__tests__/, tests/unit/lib/prime-agent/, no-loss audit, runtime adapter, spawn, conversation-runtime, transcript, cost parser, doctor, context-layers, runtime-session-projection) |
Passed — 68/68 in 115.8 s |
killAgent abort-failure probe (isolated, gitignored .tmp/, removed afterwards) |
Reproduced the blocker: ["abort"] with no terminate |
getSessionFilesSync timing probe against real ~/.claude/projects dirs |
0.46 ms (36 files) / 9.1 ms (510 files) per call; 0 of 187 dirs carry a sessions-index.json |
Not run, and why:
npm test(full suite) — the pipeline's verification stage owns it; the role's budget says not to repeat it. Focused suites covered every Prime-touching test file.npm run build,npm run lint— no exact-HEAD result available; deferred to the verification stage.tests/integration/prime-agent-smoke.slow.test.ts— opt-in slow lane (VITEST_INCLUDE_SLOW=1); it drives a real binary against a real provider, which spends credits. Recorded as unverified at this HEAD; the 2026-08-12 evidence indocs/prime-agent-verification.mdpredates the last five commits.- Live Prime session to observe the surviving tmux session for the blocker — same reason; the adapter probe plus the caller trace carry the finding.
Scope Note
Four REVIEWER_READY signals arrived mid-review (correctness, performance, requirements,
security) pointing at .pan/review/agent-pan-3668-review-be1bc0bb/ — a previous cycle's
run directory (commit be1bc0bb, 2026-09-09). This dispatch names run
agent-pan-3668-review-37430f70 and states there is no convoy and that I review every
dimension myself, so those lanes are stale and did not feed this verdict. I read the stale
correctness.md for context: it is a delta-only report that approves on the grounds that
everything outside the last few commits was "unchanged from the prior clean review" — it
never traced the kill path, which is how a blocker that has been present since
5c51e210b0d ("feat: register Prime Agent runtime adapter") survived several approving
cycles. Prior approval is not evidence of safety here; the verdict-routing drift between
be1bc0bb and 37430f70 is worth an operator look.
Source: /home/eltmon/Projects/overdeck/workspaces/feature-pan-3668/.pan/review/agent-pan-3668-review-37430f70/review.md
Required action
Fix every blocking review finding, commit the fixes, then re-request review with:
pan review request PAN-3668 -m "Fixed review issues"
…ew findings Blocking: `PrimeAgentRuntimeSync.killAgent` awaited `controller.abort` before `controller.terminate`, so any abort rejection skipped the terminate entirely. `postPrimeAgentHost` rejects when the host socket is missing, when the host answers non-2xx, and (previously, with no request deadline) never at all when the Prime child is wedged. In the wedged case the tmux session, the host process, and the Prime child all survived while `service-crash.ts` swallowed the rejection and emitted `killed_agent` — an agent recorded as killed that kept burning provider tokens. The kill now runs abort → bounded grace → terminate, with the abort best-effort and the terminate unconditional, and `postPrimeAgentHost` carries a request timeout so a wedged host cannot hang the kill it is the subject of. The in-process session registry that held the only correct abort → grace → terminate sequence was never populated in production, so `killPrimeAgentSession` was dead code and `hasPrimeAgentSession` was permanently false — and the test proving that AC exercised a path production never takes, which is how the kill defect survived several approving review cycles. The registry is gone; `deliverPrimeAgentMessage` is now the host POST it always was at runtime, and the tests drive `PrimeAgentRuntimeSync.killAgent` and a real host socket. Also from the review: - A malformed stdout record killed the host with an uncaught exception, and the framer reset its buffer *after* parsing, so every later record was corrupted. `push` now resets state before parsing and reports per-record errors, and `acceptStdout` never throws — the "later records still parse" half of rpc-framing.ac3. - A keyed delivery to a Prime agent fell past the dedup branch with the key silently dropped, so a retried keyed message was delivered twice. It now refuses, like the ACP tier, because one host POST cannot enforce at-most-once. - Nine mapped providers could never launch a work agent (`getProviderAuthMode` returning undefined was read as "no credentials") while the conversation path launched them with no check at all. Both paths now default to api-key and run one credential gate that names the missing variable or auth file and where to set it (provider-map.ac2). Its test file also moved into `__tests__/`, where vitest actually collects it — it had never run. - The stderr version fallback is now opt-in per prerequisite instead of applying to all thirteen. - Prime Agent gets its own badge glyph instead of oh-my-pi's. - The runtime projection no longer calls `getLastActivity` for harnesses without cached metrics, which put claude-code's readdir + statSync sort on the dashboard event loop once per agent row. - The no-loss audit's harness list was missing `opencode`; the docs claimed a resume check on model, provider, and workspace that the code does not perform; the file-size allowlist carried nine stale duplicate rows (the guard takes the last row per path, not the highest). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…tderr version Three gaps left by the review-fix commit: - The keyed-delivery refusal for a Prime target had no test. Added alongside the ACP and Channels refusals it is modelled on. - `delivery.test.ts` pointed OVERDECK_HOME at a directory its own afterEach had just deleted, leaving the shared per-worker home stale for every test file after it in that worker. It now saves and restores. - `doctor-prime-agent` passes its own probe, which read stdout only — so with the stderr fallback now opt-in per prerequisite, the doctor would have reported "did not return a semantic version" for a healthy Prime install. It honours the same option, and its credential fix text now names where a credential can live instead of saying "configure credentials". Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The opt-in stderr fallback passed an options object on every probe call, which changed the call shape for all thirteen prerequisites and broke the Kimi probe assertions in `src/lib/acp/__tests__/kimi.test.ts` and `tests/integration/cli/doctor.test.ts`. Only the prerequisite that needs the fallback passes it now; every other call is byte-for-byte what it was. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Review CHANGES REQUESTED for PAN-3668Review — PAN-3668Verdict: CHANGES REQUESTED — the new harness reaches
|
| File | Disposition |
|---|---|
src/lib/runtimes/prime-agent.ts |
Reviewed — prior blocker fixed: abort in try/catch, bounded grace, unconditional terminate; authMode ?? 'api-key' resolves the two-answers split (N4 of prior cycle) |
src/lib/prime-agent/session-controller.ts |
Reviewed — dead in-process registry deleted (prior N1); request timeout added; four removed exports have zero remaining references (grep across src/ + tests/) |
src/lib/agents/delivery.ts |
Reviewed — keyed deliveries to a Prime target now refuse, matching the ACP tier verbatim (delivery.ts:554); prior N3 closed |
src/lib/prime-agent/jsonl-framing.ts |
Reviewed — per-record errors, state reset before parse, bounded resynchronise; source of N8 |
src/lib/prime-agent/rpc-client.ts |
Reviewed — acceptStdout cannot throw; onRecordError channel |
src/lib/prime-agent/host.ts |
Reviewed — onRecordError → stderr, so a banner line no longer kills the host and orphans the child (prior N2); session id/path split into two expressions (N7) |
src/lib/prime-agent/provider-map.ts |
Reviewed — credential gate naming variable + settings key + auth file + login command (provider-map.ac2 now met); source of the blocker and N1, N3 |
src/lib/prime-agent/launch-command.ts |
Reviewed — single gate point shared by both launch paths |
src/lib/system-prerequisites.ts |
Reviewed — stderr version fallback is now opt-in per prerequisite (versionFromStderr), closing prior N9 for every other tool |
src/cli/commands/doctor-prime-agent.ts |
Reviewed — probe honours the opt-in; fix text names both credential homes |
src/dashboard/server/services/runtime-session-projection.ts |
Reviewed — gated on getSessionMetrics, so no harness but Prime pays a blocking directory scan (prior N7); no field regression versus main, which carried none of these fields |
src/dashboard/frontend/.../branding/index.tsx |
Reviewed — own lettermark, distinct from oh-my-pi (prior N6) |
src/lib/overdeck/infra.ts |
Reviewed — makeDbLive applies the schema like the sync door; idempotent (runOverdeckMigrationSync short-circuits on an existing agents table). Callers are terminal-issues.ts and one integration test, not the dashboard boot path. See Scope Note |
configuration/harnesses.mdx |
Reviewed — resume claim narrowed to the check that exists (prior N5). Its harness table and ToS section are what the blocker contradicts |
scripts/file-size-allowlist.txt |
Reviewed — six stale schema.ts rows and three launcher-generator.ts rows consolidated into one accurate row each (prior N11). Guard verified green at HEAD and in tree mode |
Tests (11 files: delivery, jsonl-framing, rpc-client, provider-map, deliver-agent-message, spawn-prime-agent, conversation-runtime-prime, infra, runtime-session-projection, liveness-no-loss-audit, prime-agent-no-loss-audit) |
Reviewed — fake timers used correctly for the grace period; the kill test now drives PrimeAgentRuntimeSync.killAgent (the production path) instead of the deleted registry; opencode added to the audit's ALL, which now matches KNOWN_HARNESSES exactly (8 literals); the two launch tests mock getOpenAIAuthStatus to loggedIn: false and pin providerAuth.openai: 'api-key', so the new gate takes the env-var branch and stays host-independent. Source of N4 |
Full-PR coverage
All 126 changed files are accounted for: the 27 above at this HEAD, and the remaining 99 in the prior cycle's file-by-file ledger (.pan/review/agent-pan-3668-review-37430f70/review.md), which I wrote and re-checked here for the paths the delta touches. Nothing in the delta invalidates those dispositions.
Four dimensions
- Correctness — no blocking correctness defect remains. The kill sequence, framing recovery, RPC routing, credential resolution, and the two launch paths' agreement all verified. Advisories N2, N8.
- Security / policy — the blocker is a policy-gate omission with a Terms of Service consequence, not a memory-safety or injection issue. The host's socket hygiene (0700 dirs, per-launch 32-byte 0600 token, 1 MiB body cap, concurrency 8, argument-array spawn, POSIX-correct
shellQuote) is unchanged from the prior cycle's clean assessment. Nothing in the delta logs a secret: framing errors report byte counts, and the new credential errors name variable names, never values. - Performance — improved by the delta: the projection no longer puts
readdirSync+ astatSync-comparator sort on the dashboard event loop for non-Prime harnesses (measured last cycle at 0.46 ms / 9.1 ms per call).makeDbLiveadds onesqlite_masterlookup per layer build. NoexecSyncon any server-reachable path. Advisory N2 is a latency-of-kill observation, not throughput. - Requirements / UX — one AC family fails (see matrix).
runtime-adapter.ac3,message-delivery.ac3,rpc-framing.ac3,provider-map.ac2,dashboard-frontend.ac1, andno-loss-audit.ac3all moved from unmet/partial to met this cycle.docs-setup-reference.ac3is met as written but the requirement itself is what needs correcting.
AC-to-evidence matrix (changes since prior cycle only; all others unchanged and met)
| AC | Prior | Now | Evidence |
|---|---|---|---|
runtime-adapter.ac3 |
unmet | met | prime-agent.ts:151-171; delivery.test.ts drives the production killAgent for abort-ok, abort-rejects, and host-gone |
message-delivery.ac3 |
unmet | met | Same; the grace period is a named constant advanced with fake timers |
rpc-framing.ac3 |
half-unmet | met | jsonl-framing.ts:40-95; tests cover malformed-then-later-records, split-chunk malformed, oversize completed and buffered, and resynchronise |
provider-map.ac2 |
unmet | met | assertPrimeAgentCredentialAvailable names env var + settings key (api-key) and auth file + login command (subscription); called from the one shared builder |
dashboard-frontend.ac1 |
partial | met (badge) | PrimeAgentHarnessIcon lettermark. Note: the picker offers the row unlocked in a cell policy should lock — that is the blocker, not this AC |
no-loss-audit.ac3 |
hole | met | ALL now equals KNOWN_HARNESSES (8 literals, opencode restored) |
docs-setup-reference.ac1 |
partial | met | Resume sentence now matches resumePrimeAgentSession's actual session-identity check |
docs-setup-reference.ac3 |
met | met, spec defect | The required caveat is present and accurate about billing; the requirement never reconciled the ToS gate — see blocker fix (4) |
verification-gauntlet |
unverified | unverified at this HEAD | docs/prime-agent-verification.md is dated 2026-08-12, nine commits back. Not reused as exact-HEAD evidence |
Verification
Run at HEAD 93cd6317:
| Check | Result |
|---|---|
npm run typecheck |
Passed (exit 0). Dashboard guard: 26 known errors, none new. Frontend guard: 0 known, none new. |
npm run lint:effect-diagnostics |
Passed — 308 known findings, no NEW: diagnostics |
Focused suites (20 files: src/lib/prime-agent/__tests__, tests/unit/lib/prime-agent, no-loss audits, runtime adapter, spawn, conversation-runtime, transcript, cost parser, doctor, context-layers, projection, deliver-agent-message, system-prerequisites, harness-binary, infra) |
Passed — 177/177 in 24.9 s |
bash scripts/lint-file-size.sh and --at HEAD |
Passed both modes — the allowlist consolidation is safe |
Harness-policy probe (isolated, gitignored .tmp/, removed; tree clean) |
Reproduced the blocker end to end — see the block above |
GitHub CI at exactly 93cd6317 (gh pr checks 3670) |
All green: build (22), lint, test, guard, Clean install + server smoke test, flake lane, prompt-trailer gate. Head SHA confirmed via gh pr view 3670 --json headRefOid |
Reused as exact-HEAD evidence rather than re-run: npm test, npm run build, npm run lint — CI's test, build (22), and lint jobs passed on this SHA, so the role's verification budget says not to repeat them locally.
Not run, and why:
tests/integration/prime-agent-smoke.slow.test.ts— opt-in slow lane (VITEST_INCLUDE_SLOW=1); it drives a real binary against a real provider and spends credits. Unverified at this HEAD.- A live Prime session — same reason, and for the blocker's cell specifically, launching it is the act under question.
- macOS keychain reproduction for N1 — no macOS host available.
Scope Note
- Commit
325bb2333b7is outside every xBRIEFfiles_scope. It changessrc/lib/overdeck/infra.ts(makeDbLivenow applies the schema) and de-hoststests/unit/lib/agents/liveness-no-loss-audit.test.ts(clearingOVERDECK_NO_RESUMEso the fixture comparison stops depending on how the host dashboard was booted). Both are sound, tested, and narrow — classic fix-forward for failures this work surfaced — but they belong to the substrate/liveness area, not Prime Agent. Flagging so the operator sees them rather than discovering them in a later archaeology pass. - The blocker's policy half needs an operator decision, not an agent one. Whether Prime Agent may run Anthropic models under a Claude Code OAuth subscription is a Terms of Service reading. The repo already answers it one way for
ohmypi; this PR answers it the other way forprime-agentby omission. An implementer should not resolve that silently in either direction — the mechanical fixes (explicit policy branch, picker decision key, typed decision record) are unambiguous and should land regardless. - The no-loss audit does not cover
harness-policy.tsorharness-policy-decisions.ts.no-loss-audit.ac2enumerates the surfaces it checks — Harness union, behavior map, config literal, transcript registry, telemetry, artifacts, flywheel, context previews — and the policy gate is not among them, which is why nine commits and several review cycles missed this. Adding those two files to the audit's enumeration list would make the next harness's omission a test failure instead of a review finding.
Source: /home/eltmon/Projects/overdeck/workspaces/feature-pan-3668/.pan/review/agent-pan-3668-review-93cd6317/review.md
Required action
Fix every blocking review finding, commit the fixes, then re-request review with:
pan review request PAN-3668 -m "Fixed review issues"
|
Closing this PR, not the issue. #3668 stays open. Why: main has moved about 534 commits since the last clean rebase on 9/18. The branch now conflicts in 22 files (roughly 30 hunks; 5 of them are tests main deleted), and 65 call sites use Where the work is: branch
It isn't pushed: the pre-push file-size guard rejects the stale branch. Worth porting regardless of Prime: type |
PAN-3668's merge was refused because merge-ops preferred the persisted merge_set_repos.artifact_url (#3670, the first, closed PR) over the open PR ensurePRExists resolved (#4251). Both the monorepo and remote paths now take the fresh URL and its number via freshMergeArtifact (new merge-artifact.ts, kept out of the capped merge-ops.ts), which rewrites a differing stored row and logs the replacement. The clean-direct and server-rebase tests' forge mocks now expose the real parseArtifactRef, which the new module needs. Item: merge-lands-fresh-pr Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* chore(plan): complete planning for PAN-4263 * chore(workspace): plan artifacts for PAN-4263 * fix(forge): gh/glab branch lookup throws on failure, null only on absence The gh path of getExistingGitHubArtifact swallowed every gh error with `2>/dev/null || true`, so a GraphQL rate limit read as "no PR" and review verdicts were refused. It now returns null only for gh's "no pull requests found for branch" answer or a non-OPEN PR, throws a timed-out message when the exec timeout kills gh, and throws gh's stderr otherwise. glab mr list gets the same split (empty list = absence). createReviewArtifact falls back to the number in the created URL when the follow-up lookup fails or finds nothing, so a created PR never loses its id. Item: lookup-failure-not-absence Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore(workspace): plan artifacts for PAN-4263 * chore(workspace): plan artifacts for PAN-4263 * fix(review): select the open PR for a head branch, not prs[0] The GitHub REST pulls list sorts by creation date, so prs[0] could be a newer closed PR while an older one is still open. selectPullRequestForHead ranks open (most recently updated) first and, when asked, then merged and closed. The forge App path now uses it open-only, so discoverArtifact never returns a closed PR. listPullRequestsForHead maps updated_at to updatedAt. Item: open-first-selector Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(review): move selectPullRequestForHead into github-pr-selection github-app.ts is a capped god file; the selector added in the previous commit pushed it to 1088 lines and the file-size guard refused the push. The selector and its tests now live in src/lib/github-pr-selection.ts and tests/unit/lib/github-pr-selection.test.ts (the PRD named github-app.ts). github-app.ts keeps only the updatedAt mapping (+3 lines, allowlisted). Item: open-first-selector Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore(workspace): plan artifacts for PAN-4263 * chore(workspace): plan artifacts for PAN-4263 * fix(review): retry transient forge failures in discoverArtifact New src/lib/forge-transient.ts: isTransientForgeError classifies rate limits, network errors, timeouts and 502/503/504 (anchored to the clients' "failed: 503" / "HTTP 502" formats so a /pull/504 URL never matches); retryTransientForgeOp retries only those, after 2 s then 8 s. Both adapters' discoverArtifact run through it. Tests use fake timers. Item: discover-transient-retry Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore(workspace): plan artifacts for PAN-4263 * chore(workspace): plan artifacts for PAN-4263 * fix(review): merge-gate branch lookup ranks open above merged above closed lookupPullRequestForBranch took prs[0] on the App path and `--limit 1` on the gh path; both lists sort by creation, so a newer closed PR beat an older open one. Both paths now list the branch's PRs (gh: --limit 20 with updatedAt) and pick with selectPullRequestForHead({ includeClosed: true }). Item: gate-lookup-selector Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore(workspace): plan artifacts for PAN-4263 * chore(workspace): plan artifacts for PAN-4263 * fix(review): merge lands the freshly resolved PR, not a stale stored URL PAN-3668's merge was refused because merge-ops preferred the persisted merge_set_repos.artifact_url (#3670, the first, closed PR) over the open PR ensurePRExists resolved (#4251). Both the monorepo and remote paths now take the fresh URL and its number via freshMergeArtifact (new merge-artifact.ts, kept out of the capped merge-ops.ts), which rewrites a differing stored row and logs the replacement. The clean-direct and server-rebase tests' forge mocks now expose the real parseArtifactRef, which the new module needs. Item: merge-lands-fresh-pr Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore(workspace): plan artifacts for PAN-4263 * chore(workspace): plan artifacts for PAN-4263 * fix(specialists): specialists done separates forge failure from absence A failed artifact lookup now prints "Couldn't reach <forge> to find the review artifact for <branch>: <reason>" and exits 1; only a true absence prints "No open review artifact". For review verdicts the caller identity is checked before any forge call, so a non-review agent is refused without touching the forge. A review verdict that hits a transient forge failure (lookup or post) is journaled as review.verdict-deferred with status, runId, byte-capped notes, callerId (null for an operator) and reason, through the new src/lib/cloister/deferred-verdict.ts. The journal type union gains review.verdict-deferred and review.verdict-replay-gave-up. The run-id fallback moved ahead of discovery so the deferral carries it. Item: specialists-done-honest-error Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore(workspace): plan artifacts for PAN-4263 * chore(workspace): plan artifacts for PAN-4263 * fix(cloister): deacon-lite replays deferred review verdicts New src/lib/cloister/deferred-verdict-replay.ts. When a review.verdict-deferred entry is the journal tail, recoverStalledReviews (after the issue-pause hold) replays it once the entry is >= 10 minutes old, through `pan admin specialists done review` under the original caller (OVERDECK_AGENT_ID = callerId; agent vars removed for an operator verdict), so reviewVerdictRefusal runs again in full. The child's own appended journal entries decide the outcome: review.verdict ends it, a fresh deferral is the next cooldown, anything else journals review.verdict-replay-gave-up {reason:'failed'}. A newer reviewRunId gives up as 'superseded'; seven deferrals of one run give up as 'cap' and warn in activity. An in-process set stops overlapping ticks. `pan show` summarizes the two new entry types. Deviation from D10: the outcome is read from entries appended during the replay, not from a deferral count keyed on runId, so a child that resolves a different runId (or records the verdict then fails later) is not mis-read as a failure. Implementation checkpoint: on this host the deacon child (pid 1868537) -> dashboard server (1867603) -> user systemd (1917, environ unreadable, walk stops) carry no OVERDECK_AGENT_ID, so readAncestorAgentIds() is [] in a replay child and the caller resolves from the env it is given. Item: deferred-verdict-replay Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore(workspace): plan artifacts for PAN-4263 * chore(workspace): plan artifacts for PAN-4263 * docs(review): document forge retry, deferred verdicts and the merged PR docs/PIPELINE-GATES.md: under "Verdict feedback routing", how discoverArtifact retries transient failures (2 s, 8 s), how a failed lookup is reported, and how a review verdict is deferred and replayed; entry-type rows for review.verdict-deferred and review.verdict-replay-gave-up; deacon-lite routine 5 names the replay; the merge gate section says which PR a branch lookup picks and that the merge lands the freshly resolved PR, overwriting a stale stored URL. Also commits the planner's concerns.md entry on forge lookup failure vs absence and stale stored artifact URLs, unchanged: its "Before PAN-4263" wording already reads correctly now that the fix has landed. Item: docs-pipeline-gates Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore(workspace): plan artifacts for PAN-4263 * test(review): merge leaves a stored row alone when it names the resolved PR Covers merge-lands-fresh-pr.ac4: a stored artifact_url equal to the resolved PR (trailing slash ignored) is not rewritten. Item: merge-lands-fresh-pr Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(review): say a stale stored artifact_url is never merged Matches docs-pipeline-gates.ac4 wording. Item: docs-pipeline-gates Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore(workspace): record PAN-4263 plan as running The pipeline stamped the xBRIEF status proposed -> running at work start. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test(review): acceptance criteria verified for PAN-4263 Each acceptance criterion is verified by its parent item's tests: forge.test.ts (lookup-failure-not-absence), github-pr-selection.test.ts and forge.test.ts (open-first-selector), forge-transient.test.ts (discover-transient-retry), github-pr-lookup.test.ts (gate-lookup-selector), merge-ops-stale-artifact.test.ts (merge-lands-fresh-pr), specialists-done-forge-lookup.test.ts (specialists-done-honest-error), deferred-verdict-replay.test.ts and deacon-lite-recover-stalled-reviews.test.ts (deferred-verdict-replay), and docs/PIPELINE-GATES.md (docs-pipeline-gates). Item: lookup-failure-not-absence.ac1 Item: lookup-failure-not-absence.ac2 Item: lookup-failure-not-absence.ac3 Item: lookup-failure-not-absence.ac4 Item: lookup-failure-not-absence.ac5 Item: lookup-failure-not-absence.ac6 Item: open-first-selector.ac1 Item: open-first-selector.ac2 Item: open-first-selector.ac3 Item: open-first-selector.ac4 Item: discover-transient-retry.ac1 Item: discover-transient-retry.ac2 Item: discover-transient-retry.ac3 Item: discover-transient-retry.ac4 Item: discover-transient-retry.ac5 Item: gate-lookup-selector.ac1 Item: gate-lookup-selector.ac2 Item: gate-lookup-selector.ac3 Item: merge-lands-fresh-pr.ac1 Item: merge-lands-fresh-pr.ac2 Item: merge-lands-fresh-pr.ac3 Item: merge-lands-fresh-pr.ac4 Item: specialists-done-honest-error.ac1 Item: specialists-done-honest-error.ac2 Item: specialists-done-honest-error.ac3 Item: specialists-done-honest-error.ac4 Item: specialists-done-honest-error.ac5 Item: deferred-verdict-replay.ac1 Item: deferred-verdict-replay.ac2 Item: deferred-verdict-replay.ac3 Item: deferred-verdict-replay.ac4 Item: deferred-verdict-replay.ac5 Item: docs-pipeline-gates.ac1 Item: docs-pipeline-gates.ac2 Item: docs-pipeline-gates.ac3 Item: docs-pipeline-gates.ac4 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore(workspace): plan artifacts for PAN-4263 * chore(workspace): plan artifacts for PAN-4263 * chore(workspace): plan artifacts for PAN-4263 * chore(workspace): plan artifacts for PAN-4263 * chore(workspace): plan artifacts for PAN-4263 * chore(workspace): plan artifacts for PAN-4263 * chore(workspace): plan artifacts for PAN-4263 * chore(workspace): plan artifacts for PAN-4263 * chore(workspace): plan artifacts for PAN-4263 * chore(workspace): plan artifacts for PAN-4263 * chore(workspace): plan artifacts for PAN-4263 * chore(workspace): plan artifacts for PAN-4263 * chore(workspace): plan artifacts for PAN-4263 * chore(workspace): plan artifacts for PAN-4263 * chore(workspace): plan artifacts for PAN-4263 * chore(workspace): plan artifacts for PAN-4263 * chore(workspace): plan artifacts for PAN-4263 * chore(workspace): plan artifacts for PAN-4263 * chore(workspace): plan artifacts for PAN-4263 * chore(workspace): plan artifacts for PAN-4263 * chore(workspace): plan artifacts for PAN-4263 * chore(workspace): plan artifacts for PAN-4263 * chore(workspace): plan artifacts for PAN-4263 * chore(workspace): plan artifacts for PAN-4263 * chore(workspace): plan artifacts for PAN-4263 * chore(workspace): plan artifacts for PAN-4263 * chore(workspace): plan artifacts for PAN-4263 * chore(workspace): plan artifacts for PAN-4263 * chore(workspace): plan artifacts for PAN-4263 * chore(workspace): plan artifacts for PAN-4263 * chore(workspace): plan artifacts for PAN-4263 * chore(workspace): plan artifacts for PAN-4263 * chore(workspace): plan artifacts for PAN-4263 * chore(workspace): plan artifacts for PAN-4263 * chore(workspace): plan artifacts for PAN-4263 * chore(workspace): plan artifacts for PAN-4263 --------- Co-authored-by: overdeck-agent[bot] <4205044+overdeck-agent[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Issue: #3668
Acceptance Criteria