Skip to content

fix(server): wait out transient EADDRINUSE with bounded bind retries - #5437

Open
chaucerj wants to merge 2 commits into
volcengine:mainfrom
chaucerj:fix/server-bind-retry
Open

chaucerj wants to merge 2 commits into
volcengine:mainfrom
chaucerj:fix/server-bind-retry

Conversation

@chaucerj

Copy link
Copy Markdown

Description

openviking-server bound its listen socket inside uvicorn.run, which logs a bind failure and exits gracefully — so a supervised restart (watchdog, systemd Restart=always, Docker healthcheck) racing the old process winding down killed the new process on the first EADDRINUSE.

The socket is now bound up front, retrying while the port is held (server.bind_retry_attempts, default 5; server.bind_retry_interval_seconds, default 1.0; 0 attempts keeps the die-immediately behavior), and then handed to uvicorn: Server.run(sockets=[...]) for the single-process path, and the inherited fd= for the multi-worker path, since uvicorn 0.41's run()/Config take no sock parameter and run() swallows the bind error instead of raising.

Human Involvement

  • A human participated in the implementation or review loop

Related Issue

Fixes #5401

Type of Change

  • Bug fix (non-breaking change that fixes an issue)

Changes Made

  • openviking/server/bootstrap.py: _bind_socket_with_retry (bounded EADDRINUSE retries with a stderr note per attempt) + _run_uvicorn handing the pre-bound socket to uvicorn per mode; both startup call sites routed through it.
  • openviking/server/config.py: server.bind_retry_attempts / server.bind_retry_interval_seconds with the rationale in the field description.
  • tests/server/test_bootstrap_bind_retry.py: real-socket tests for the waited-out holder, exhausted retries (original OSError semantics), zero-attempts fail-fast, free-port no-retry, and conservative defaults.

Testing

  • pytest tests/server/test_bootstrap_bind_retry.py: 6/6 pass; ruff check clean.
  • Live runs against a local server: port held → "still in use (attempt n/8)" retries → holder released → /health 200; port held forever → retries exhaust → original OSError; free port → unchanged startup.

🤖 Generated with Claude Code

openviking-server bound its listen socket inside uvicorn.run, which
logs a bind failure and exits gracefully — so a supervised restart
(watchdog, systemd Restart=always, Docker healthcheck) racing the old
process winding down killed the new process on the first EADDRINUSE.

Bind the socket up front instead, retrying while the port is held
(server.bind_retry_attempts, default 5; server.bind_retry_interval_
seconds, default 1.0; 0 attempts keeps the die-immediately behavior),
then hand it to uvicorn: Server.run(sockets=[...]) for the single
process path, and the inherited fd= for the multi-worker path, since
uvicorn 0.41's run()/Config take no sock parameter.

Verified live: port held -> 'still in use (attempt n/8)' retries ->
holder released -> /health 200; port held forever -> retries exhaust
-> original OSError; free port -> unchanged startup.

Fixes volcengine#5401

Co-Authored-By: Claude Code <noreply@anthropic.com>
Upstream's remote-restart recovery replaced the single-worker
_run_uvicorn call with _run_restartable_server. Keep both semantics:
bind the socket up front with the bounded EADDRINUSE retry, then hand
the pre-bound socket to the restartable server via uvicorn's sock=
config (dropping host/port, which a bound socket supersedes).

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Backlog

Development

Successfully merging this pull request may close these issues.

openviking-server dies on first EADDRINUSE at startup (no bind retry); --with-bot orphans the gateway child on failed starts

1 participant