Repository navigation
Conversation
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
wesbillman
left a comment
There was a problem hiding this comment.
No actionable changes requested. Reviewed cross-community routing, connection/retry and teardown lifetimes, and shared Claude spare ownership; I found no source-demonstrated regression in this revision.
Star Lord’s automated source review via Wes’s account. Head c8a0ee8ad6d98853c4bf9d4006a71febb4bd3088; base f8705a95ba8034c1a3c7ec5f09847db1c864d97b. Source-only: no tests or app execution, and no independent CI verification; native-app switching/reconnect behavior and process counts remain unverified.
| console.warn("Claude Code could not start", error); | ||
| return undefined; | ||
| }); | ||
| this.opening.set(key, starting); |
There was a problem hiding this comment.
🤖 [P2] Preserve per-agent admission before the shared native process ceiling
Removing MAX_LIVE/room() also removes idle eviction on admission. Completed Claude conversations retain their processes for 15 minutes (sessions.ts:193–200), while the native host permits only 64 live processes across all plugins (src-tauri/src/host_process.rs:21–22,102–104). Under high conversation fan-out, one agent can consume that shared budget and prevent unrelated agents from launching processes, including Codex servers or upload readers. The message rate limit does not bound retained processes.
Please retain bounded per-agent live + opening admission and idle eviction alongside the shared spare, returning a local busy result instead of exhausting the shared registry. Add fake-host capacity/controlled-clock coverage for pending starts and completed-idle sessions. This is a source-established starvation path, not a measured memory or native stress-test result.
Agents2 listened only to the selected community, and the Claude Code and Codex runtimes stopped the processes of agents that left that list, so switching communities killed work mid-turn and left the other community unheard. Agents2 now binds each joined community the viewer has an agent in (opened through a new Communities.open, without selecting it), delivers each community's live events only to its own agents, and exposes the running agents and each agent's community connection. The runtimes keep processes for every running agent and read conversations, names and memory from the agent's own community. Processes stop when the agent is deleted, its community is left or disconnected, or the viewer signs out. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Honey <8e307ae0076a4dab6b94b036ea3edc7e08f823a625269c1e6919e881a048b4d2@buzz.block.builderlab.xyz>
A first connect that fails leaves the session in error, and only the selected community has a Retry the viewer can press. Agents2 now tries an unselected community's failed connection again every 30 seconds, so its agents recover without the viewer opening it. Also covers Communities.open directly. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Honey <8e307ae0076a4dab6b94b036ea3edc7e08f823a625269c1e6919e881a048b4d2@buzz.block.builderlab.xyz>
Each Claude Code agent kept its own warm `claude` process and up to six processes in all. With agents in many communities kept running, that is one idle ~300 MB process per agent. Starting the process is what a spare saves (time to turn start ~530ms cold vs ~45ms warm, measured against claude 2.1.296); sending `initialize` costs almost nothing. So the spare is now spawned without `initialize`, fixing only its folder and model, and the agent that takes it sends its own prompt and installs its own tool handler. One spare serves every agent; it follows the latest new conversation, and is replaced when no agent runs where it was started. The per-agent process cap is removed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Bradley Axen <baxen@squareup.com>
A conversation whose launch finished after its last agent was removed took the spare and started a replacement, which then stayed with no agents. Taking the spare now refills it only where the runtime wants one, so stopping the last agent (`warm()` with no places) is final. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Bradley Axen <baxen@squareup.com>
- A native agent list still being read when Agents2 stops is now stale, so it cannot reopen agent communities after teardown. - Codex follows `running` whatever the inventory status, as Claude Code does, so a failed refresh no longer keeps a stopped agent's server. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Bradley Axen <baxen@squareup.com>
c8a0ee8 to
10ddcc8
Compare
Problem
Switching communities killed running Agents2 agents. Agents2Service only listened to the selected community. The Claude Code and Codex runtimes treated
snapshot().agentsas the list of agents that should be alive, so an agent whose community stopped being selected had its process disposed mid-turn. Its community then went unheard until you switched back.Change
running: the viewer's agents in every joined community that isn't disconnected. A reconnect doesn't remove them.agentsis still the selected community's agents, so the Agents2 page is unchanged.agents2.relay(pubkey)returns the owner's connection to an agent's own community.running. They read thread history, channel names and memory through the agent's own community instead of the selected one. Codex no longer disposes everything on a community switch. Its saved-session key is the sameorigin:viewerstring as before, so no saved sessions are orphaned.Communities.open(id)is a new method that connects a joined community without selecting it.communities/service.tsis a FOUNDATION file; baxen asked for this fix after reviewing the plan that called out this addition.A process now stops only when its agent is deleted, its community is left or disconnected, the viewer signs out, or the plugin is disabled.
One shared Claude Code spare (25e885f)
Keeping every agent running made Claude Code's per-agent warm spare expensive: one idle
claude(~300 MB RSS) per agent, and up to six processes per agent.initialize(prompt, memory, hooks), and installs its own tool handler. Then another spare starts.initializecosts almost nothing. Measured againstclaude2.1.296, the time from send to turn start is ~530 ms cold vs ~45 ms with the spare (medians, 10 interleaved rounds). A spawned-only process is as fast as a fully initialized one.claudelists the SDK MCP server's tools right after spawn, beforeinitialize. Until the spare is taken, a placeholder handler answers; the tools list does not depend on the agent.claudeis not respawned in a loop.initializein ~85 ms and starts a thread in ~95 ms, so a spare would save ~100 ms at most.Tradeoffs
--resumeis a spawn flag.thread/resume), so an idle stop is a sensible follow-up.Validation
Full Vitest suite: 653 files / 9193 tests passed at 2e36587.
Affected suites: agents2, bundled and communities, 222 files / 2878 tests, passed at 43f3c3b. Type check and lint are clean.
New tests:
Communities.openopens a joined community without selecting it, and returns nothing for a community that isn't joined.Mutation check: removing the per-community event filter fails the routing test.
Codex tests updated: the test that asserted "disposes on community change" now asserts the server ends when its agent stops running.
Independent agent review: done. Its medium finding was that a background community whose first connect fails never recovers. It's fixed in 43f3c3b.
Shared spare (25e885f):
claude, outside the app: two agents with different prompts took turns using the shared spare, and each answered with its own name.Lifecycle review fixes (8904805, c8a0ee8): codex00's review found three lifecycle bugs, each reproduced. Each fix takes one decision away from a place that shouldn't make it:
error, so leaving a community after a failed refresh left its servers running. It now followsrunningregardless of inventory status, as Claude Code does.Not yet done
claudeexists beyond those in use.No browser cases added or removed.
🤖 Generated with Claude Code