Repository navigation
Conversation
Codex agents each kept their own `codex app-server` (~300-400 MB with its MCP children) until the agent stopped running. With agents now running in every joined community, that cost grew with every Codex agent used. Everything agent-specific already travels with the thread: cwd, developer instructions and dynamic tools go in thread/start, and tool calls carry their threadId. So the runtime now owns one server, opened on first use in ~/.buzz, and finds a tool call's agent by thread and turn across all agents. Entries keep their conversations and saved bindings but no longer own a process. A stopped agent can no longer take its process down with it, so it interrupts its own turns, then cleans the background shells its threads left (a dev server from a finished turn, say). Starting the same agent again waits for that, so it never resumes a thread mid-teardown. The server ends with the last Codex agent, on dispose, or on a crash (the next mention reopens it). live.mjs passed the relay as an object since the per-agent relay change; it now passes a function, expands ~ like the desktop host, and checks two agents on one real server. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Bradley Axen <baxen@squareup.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #838 (base
honey/agents2-all-communities). Merge #838 first, then retarget this PR tomain.Why
Each Codex agent used to have its own
codex app-server, and it ran until that agent stopped. With #838, agents keep running in every joined community, so every Codex agent you've used holds one indefinitely. On a Mac, each one is about 300–400 MB:codexbinary: ~176 MBnodewrapper: ~49 MB~/.codex/config.toml: ~80 MBWhat changed
CodexRuntimenow owns one app-server for every agent. It opens on first use, in~/.buzz.This works because everything agent-specific already travels with the thread, not the process:
cwd,developerInstructionsanddynamicToolsgo inthread/start/thread/resume.item/tool/callcarries itsthreadIdandturnId.So a tool call finds its agent by thread and turn across all agents. An entry keeps its conversations and saved bindings but no longer owns a process.
Starting a Codex server cold is cheap: about 85 ms for
initializeand 95 ms for the firstthread/start, measured withcodex-cli0.162.0. A single shared server therefore needs no spare process, and no idle timer either: the most it holds is about 400 MB, however many agents you have.Stopping an agent no longer kills a process the others share. Instead:
thread/backgroundTerminals/cleanworks afterthread/unsubscribe.The server ends when the last Codex agent that has worked stops, when the plugin is turned off, or when it crashes. After a crash, the next mention reopens it.
Tradeoff
A crash, or a failed shell cleanup, now ends every Codex agent's current turn, not just one agent's. Each affected thread gets the usual "Codex could not finish" reply, and the next mention resumes the thread. A failed cleanup still closes the server on purpose, so a stopped agent's shells can't outlive it.
Tests
Unit: all 50 Codex tests pass. Four are new:
The new tests for one shared server, interrupting only the removed agent, and the cleanup/restart ordering fail without the change.
Full Vitest suite: 9207 pass.
tscand biome are clean.Real Codex (
bin/node src/bundled/codex/live.mjs): passes, including a new step where two agents with different workspaces and instructions run on one real server, and each replies as itself from its own workspace.live.mjshad been passing the relay as an object since fix(agents2): keep agents running when another community is selected #838 changed the constructor; it now passes a function and expands~like the desktop host does.Not yet run in the native app. To try it:
codex app-server, not two.🤖 Generated with Claude Code