Skip to content

perf(codex): one app-server for every agent - #852

Draft
baxen wants to merge 1 commit into
honey/agents2-all-communitiesfrom
baxen/codex-one-server
Draft

baxen wants to merge 1 commit into
honey/agents2-all-communitiesfrom
baxen/codex-one-server

Conversation

@baxen

@baxen baxen commented Oct 11, 2026

Copy link
Copy Markdown
Collaborator

Stacked on #838 (base honey/agents2-all-communities). Merge #838 first, then retarget this PR to main.

Why

Each Codex agent used to have its own codex app-server, and it ran until that agent stopped. With #838, agents keep running in every joined community, so every Codex agent you've used holds one indefinitely. On a Mac, each one is about 300–400 MB:

  • the codex binary: ~176 MB
  • the npm node wrapper: ~49 MB
  • the MCP servers from ~/.codex/config.toml: ~80 MB

What changed

CodexRuntime now owns one app-server for every agent. It opens on first use, in ~/.buzz.

This works because everything agent-specific already travels with the thread, not the process:

  • cwd, developerInstructions and dynamicTools go in thread/start / thread/resume.
  • Each item/tool/call carries its threadId and turnId.

So a tool call finds its agent by thread and turn across all agents. An entry keeps its conversations and saved bindings but no longer owns a process.

Starting a Codex server cold is cheap: about 85 ms for initialize and 95 ms for the first thread/start, measured with codex-cli 0.162.0. A single shared server therefore needs no spare process, and no idle timer either: the most it holds is about 400 MB, however many agents you have.

Stopping an agent no longer kills a process the others share. Instead:

  1. Its turns interrupt themselves.
  2. Once they've finished, its threads' background shells are cleaned up, e.g. a dev server left by a finished turn. Its own process exiting used to do this. Verified against real Codex: thread/backgroundTerminals/clean works after thread/unsubscribe.
  3. If the same agent starts again, it waits for step 2, so it never resumes a thread that's still being torn down.

The server ends when the last Codex agent that has worked stops, when the plugin is turned off, or when it crashes. After a crash, the next mention reopens it.

Tradeoff

A crash, or a failed shell cleanup, now ends every Codex agent's current turn, not just one agent's. Each affected thread gets the usual "Codex could not finish" reply, and the next mention resumes the thread. A failed cleanup still closes the server on purpose, so a stopped agent's shells can't outlive it.

Tests

  • Unit: all 50 Codex tests pass. Four are new:

    • two agents share one server, and each agent's tool calls publish as that agent
    • removing one agent interrupts only its turn
    • a stopped agent's leftover shells get cleaned up, and the agent starts again only after that
    • the server reopens after an exit

    The new tests for one shared server, interrupting only the removed agent, and the cleanup/restart ordering fail without the change.

  • Full Vitest suite: 9207 pass. tsc and biome are clean.

  • Real Codex (bin/node src/bundled/codex/live.mjs): passes, including a new step where two agents with different workspaces and instructions run on one real server, and each replies as itself from its own workspace. live.mjs had been passing the relay as an object since fix(agents2): keep agents running when another community is selected #838 changed the constructor; it now passes a function and expands ~ like the desktop host does.

  • Not yet run in the native app. To try it:

    1. Mention two different Codex agents.
    2. Check Activity Monitor: there should be one codex app-server, not two.
    3. Delete one agent mid-turn: the other agent should keep working.

🤖 Generated with Claude Code

Codex agents each kept their own `codex app-server` (~300-400 MB with its
MCP children) until the agent stopped running. With agents now running
in every joined community, that cost grew with every Codex agent used.

Everything agent-specific already travels with the thread: cwd,
developer instructions and dynamic tools go in thread/start, and tool
calls carry their threadId. So the runtime now owns one server, opened
on first use in ~/.buzz, and finds a tool call's agent by thread and
turn across all agents. Entries keep their conversations and saved
bindings but no longer own a process.

A stopped agent can no longer take its process down with it, so it
interrupts its own turns, then cleans the background shells its threads
left (a dev server from a finished turn, say). Starting the same agent
again waits for that, so it never resumes a thread mid-teardown. The
server ends with the last Codex agent, on dispose, or on a crash (the
next mention reopens it).

live.mjs passed the relay as an object since the per-agent relay
change; it now passes a function, expands ~ like the desktop host, and
checks two agents on one real server.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Bradley Axen <baxen@squareup.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant