Skip to content

Repository files navigation

rsloop logo

An event loop for asyncio written in Rust

PyPI - Version Tests PyPI Downloads

rsloop is a PyO3-based asyncio event loop implemented in Rust.

Each rsloop.Loop owns a dedicated Rust runtime thread for loop coordination and I/O work. That thread runs an rsloop-specialized vibeio runtime, using io_uring on Linux, IOCP on Windows, and native kqueue readiness on macOS. Native-stream TCP reads and Unix-domain socket reads run on that runtime. On Unix, generic TCP protocol readers use a second vibeio runtime on the Python loop thread (io_uring on Linux), avoiding cross-thread delivery of each read. Non-TLS accepts run on either runtime depending on where the server starts. Python callbacks, tasks, and coroutines run on the thread that calls run_forever() or run_until_complete() (usually the main Python thread).

The package exposes:

  • a native extension module at rsloop._loop
  • a Python wrapper in python/rsloop/__init__.py
  • rsloop.Loop, rsloop.EventLoopPolicy, rsloop.new_event_loop(), rsloop.run(...), rsloop.install(), rsloop.uninstall(), and rsloop.build_info()

Repository metadata currently targets Python >=3.10. The native runtime requires Linux 6.1+, macOS 13+, or Windows 10+ so its hot paths can rely on modern completion, timer, and scheduler primitives. Free-threaded CPython (3.14t) is supported: the extension declares gil_used = false, so importing it no longer re-enables the GIL. See Free-Threaded CPython for what that does and does not buy you.

Documentation

Project documentation now lives in docs/.

If you are new to the repository, start with:

To browse the docs locally with MkDocs:

uvx --from mkdocs mkdocs serve

Install

From PyPI:

pip install rsloop

With uv:

uv add rsloop

From conda-forge, using pixi:

pixi add rsloop

Usage

Simple entry point:

import rsloop


async def main(): ...


rsloop.run(main())

Install as the default asyncio event loop policy:

import asyncio
import rsloop

rsloop.install()
try:
    asyncio.run(main())
finally:
    rsloop.uninstall()

Manual loop creation also works:

import asyncio
import rsloop

loop = rsloop.new_event_loop()
asyncio.set_event_loop(loop)
try:
    loop.run_until_complete(...)
finally:
    asyncio.set_event_loop(None)
    loop.close()

Importing rsloop also patches asyncio.set_event_loop() so Python 3.10 can accept an rsloop.Loop instance, matching the behavior exercised by tests/test_run.py.

Custom Async Rust Extensions

rsloop now exposes a small Rust interop API for downstream PyO3 extensions. That lets you write your own async Rust code, return it to Python as an awaitable, and run it under the active rsloop event loop.

The public entry point is rsloop::rust_async:

  • get_current_locals(...)
  • future_into_py(...)
  • future_into_py_with_locals(...)
  • local_future_into_py(...)
  • local_future_into_py_with_locals(...)
  • re-exports of TaskLocals and into_future_with_locals(...)

See examples/rust/README.md for a complete extension example built with maturin.

Verified Surface Area

The current codebase implements these user-facing areas.

Loop lifecycle and scheduling:

  • run_forever, run_until_complete, stop, close
  • time, is_running, is_closed
  • get_debug, set_debug
  • call_soon, call_soon_threadsafe, call_later, call_at
  • returned Handle and TimerHandle objects with cancel() / cancelled()

Tasks, futures, and execution helpers:

  • create_future, create_task
  • set_task_factory, get_task_factory
  • set_exception_handler, get_exception_handler, call_exception_handler, default_exception_handler
  • set_default_executor, run_in_executor
  • shutdown_asyncgens, shutdown_default_executor
  • callback execution under captured contextvars.Context
  • asyncio.get_running_loop() support while running on rsloop
  • rsloop.run(...) helper, with asyncio.run(..., loop_factory=...) integration on Python 3.12+

I/O and networking:

  • add_reader, remove_reader, add_writer, remove_writer
  • sock_recv, sock_recv_into, sock_sendall, sock_accept, sock_connect
  • getaddrinfo, getnameinfo
  • create_server, create_connection
  • create_unix_server, create_unix_connection
  • connect_accepted_socket
  • returned Server objects with close(), is_serving(), get_loop(), and sockets()
  • returned StreamTransport objects with write(), writelines(), close(), abort(), is_closing(), write_eof(), can_write_eof(), get_extra_info(), get_protocol(), set_protocol(), pause_reading(), resume_reading(), is_reading()

Pipes, subprocesses, and signals:

  • connect_read_pipe, connect_write_pipe
  • subprocess_exec, subprocess_shell
  • returned ProcessTransport and ProcessPipeTransport objects
  • higher-level compatibility with asyncio.create_subprocess_exec() and asyncio.create_subprocess_shell()
  • Unix subprocess options including cwd, env, executable, pass_fds, start_new_session, process_group, user, group, extra_groups, umask, and restore_signals
  • add_signal_handler, remove_signal_handler

Profiling:

  • Python 3.15's external profiling.sampling profiler
  • opt-in transport counters through transport_stats() and reset_transport_stats()

Set RSLOOP_TRANSPORT_STATS=1 before importing rsloop to enable the transport counters. They report read completions and bytes, Python-thread read drains, wakeups, staged and direct writes, and Windows completion-to-poll rebinds. Counters remain disabled by default so diagnostics add only one predictable branch to transport hot paths.

Fast Streams

Importing rsloop patches asyncio.open_connection() and asyncio.start_server() by default.

That import-time behavior is controlled by RSLOOP_USE_FAST_STREAMS and can be disabled with:

export RSLOOP_USE_FAST_STREAMS=0

The native fast-stream path is used only when:

  • the running loop is an rsloop.Loop
  • ssl is unset or None

Otherwise rsloop falls back to the stdlib asyncio.streams helpers.

On that path the reader handed to your code is the native PyFastStreamReader rather than asyncio.StreamReader. It implements the reading surface protocols actually use:

  • read(n=-1), readexactly(n)
  • readline(), readuntil(separator=b"\n"), including the tuple-of-separators form CPython 3.13+ accepts
  • at_eof(), exception(), feed_data(), feed_eof(), set_exception()

These match asyncio.StreamReader down to the exception types and their attributes — IncompleteReadError.partial, LimitOverrunError.consumed, the ValueError that readline() raises on limit overrun — and down to what is left in the buffer afterwards. tests/test_stream_reader.py pins that by driving the same feed scripts through both readers and comparing the results.

The implementation lives in src/transport/stream/fast.rs and is backed by the lower level transport code in src/transport/stream/mod.rs.

Free-Threaded CPython

rsloop builds and runs on free-threaded CPython 3.14 (3.14t). The extension declares #[pymodule(gil_used = false)], which is what keeps CPython from silently switching the GIL back on for the whole process at import time:

import sys
import rsloop

assert not sys._is_gil_enabled()
assert rsloop.build_info()["free_threaded"]

What that buys you is that separate rsloop.Loop instances on separate threads run concurrently rather than taking turns. A loop is still single-threaded internally, and asyncio objects are still not thread-safe, so the model is one loop per thread — not one loop shared across threads. call_soon_threadsafe() remains the supported way to hand work to a loop from another thread, and it keeps its FIFO ordering guarantee.

The pieces that made this safe:

  • the generic stream-reader fast path writes into StreamReader._buffer through a raw pointer; the size read, resize, and copy now run inside a critical section on that bytearray, so a concurrent mutation cannot leave the copy writing into a freed allocation
  • the ready-queue refill preserves scheduling order when a drain slice leaves older callbacks in the batch. Under the GIL a cross-thread producer could only enqueue while the loop thread was parked, so the reordering was essentially unreachable; without the GIL producers append throughout the drain and it became routine

tests/test_free_threading.py covers this: parallel loops over both the native and stdlib stream reader paths, call_soon_threadsafe() fan-in from eight threads, and a check that importing rsloop leaves the GIL off.

Wheels are built for 3.14t alongside the GIL builds, and the test matrix runs it as its own entry.

Runtime Model

Each loop combines a coordination runtime with a loop-thread I/O runtime:

  • the coordination thread handles commands, timers, and cross-thread work
  • on Unix, generic TCP protocol readers run on the Python loop thread; native fast streams and Unix-domain readers retain coordination-thread I/O
  • non-TLS accept loops use vibeio on the thread that starts them
  • bounded ready-callback turns service loop-thread I/O even when Python tasks continually yield with sleep(0)
  • Windows TCP transports, including custom asyncio.Protocol implementations, start in IOCP completion mode and rebind to readiness mode before start_tls synchronously reclaims a socket
  • generic add_reader / add_writer descriptors use cancellable OS-poll workers because vibeio does not expose arbitrary raw-descriptor registration
  • some transport paths still fall back to helper threads, especially TLS I/O, TLS server accept, and parts of the legacy transport write path

The runtime dependency is now unified, but the codebase has not finished eliminating every helper thread yet.

Transport overload safeguards use conservative defaults: inbound reads pause at 1 MiB of pending data per connection, buffered writes are capped at 64 MiB, and a TLS server admits at most 256 simultaneous handshakes. The last two limits can be adjusted before importing rsloop with RSLOOP_MAX_WRITE_BUFFER_BYTES and RSLOOP_MAX_PENDING_TLS_HANDSHAKES.

Current Limitations

These gaps are visible in the current implementation.

  • TLS uses a rustls backend with a narrower compatibility surface than CPython's OpenSSL-backed ssl module. In particular, encrypted private keys are not supported yet, and the fast-stream monkeypatch still falls back to stdlib helpers whenever ssl is enabled. TLS transport internals also still use helper-thread paths instead of the runtime-thread vibeio socket path.
  • Subprocess support still has one notable gap: preexec_fn remains unsupported because running arbitrary Python between fork() and exec() is unsafe in this runtime model.
  • Unix-specific APIs remain Unix-specific: create_unix_server, create_unix_connection, add_signal_handler, remove_signal_handler.
  • Platform-specific limitations still apply: Unix socket APIs and Unix signal handlers remain Unix-only, and several subprocess options such as pass_fds, user, group, and umask are still specific to Unix process spawning.
  • The transport runtime model is still in transition: protocol readers on Unix avoid a coordination-thread hop, but native streams, generic descriptor watches, and TLS-heavy paths do not share one single-threaded I/O path.

Build

Local development uses Python 3.14.7, pinned in .python-version. Install that interpreter before running the uv commands below. This development pin does not change the package's Python 3.10+ support or the multi-version test matrix.

Local builds and build/test CI use Rust 1.98.1, pinned in rust-toolchain.toml. Rustup selects it automatically inside this repository. LLVM tools remain optional for PGO builds.

Quick check:

cargo check

Release build and editable install:

cargo build --release
uv run --with maturin maturin develop --release

Build release wheels into dist/wheels:

scripts/build-wheels.sh

Optionally build wheels with profile-guided optimization:

rustup component add llvm-tools-preview
scripts/build-pgo-wheels.sh

For each requested Python ABI, the PGO wrapper creates an instrumented wheel, trains it on sustained HTTP, TLS, WebSocket, mixed-stream, bulk-transfer, idle-connection, callback, task, and TCP workloads, merges the resulting LLVM profiles, and builds that ABI's final wheel with its matching profile. Per-ABI training avoids discarding counters when PyO3's generated control flow differs between Python versions or free-threaded builds. The target must be native because the instrumented extension runs during training.

Set RSLOOP_PGO_SCENARIOS to override the comma-separated network scenarios. The Wheels CI workflow disables PGO by default: tagged releases and ordinary manual runs use the normal release-wheel builder. To opt in, enable the pgo checkbox when manually running the workflow. LLVM tools are installed only for PGO runs; source-distribution and publishing steps are unchanged. When enabled, PGO is used on every supported platform except Windows ARM64. Rust profile-generation binaries currently crash on that target (rust-lang/rust#156675), so it temporarily falls back to the normal fat-LTO release build.

scripts/build-wheels.sh currently defaults to CPython 3.10 3.11 3.12 3.13 3.14 3.14t 3.15, and uses uv python install / uv python find to locate interpreters.

Profiling

Python 3.15 includes a low-overhead sampling profiler that can run rsloop without a special build or in-process instrumentation. Generate an interactive flame graph with:

uv run --python 3.15 --with maturin maturin develop --release
uv run --python 3.15 python -m profiling.sampling run \
  --all-threads --native --flamegraph \
  -o rsloop-profile.html examples/01_basics.py

--all-threads includes rsloop's runtime thread and --native marks time below the Python/native boundary. The profiler and target must use the same Python 3.15 interpreter. Python 3.15 does not allow these options together with --async-aware; use a separate async-aware pass when coroutine reconstruction is more important than native and multi-thread visibility.

Examples

Run the repository examples from the project root:

uv run python examples/01_basics.py
uv run python examples/02_fd_and_sockets.py
uv run python examples/03_streams.py
uv run python examples/04_unix_and_accepted_socket.py
uv run python examples/05_pipes_signals_subprocesses.py

Example files: examples/01_basics.py, examples/02_fd_and_sockets.py, examples/03_streams.py, examples/04_unix_and_accepted_socket.py, examples/05_pipes_signals_subprocesses.py.

The repository also includes:

Benchmark

uv run --with maturin maturin develop --release
uv run --with uvloop --with 'zuvloop; python_version >= "3.14"' python benches/compare_event_loops.py \
  --loops asyncio,uvloop,zuvloop,rsloop --repeat 7 --warmups 2

Four-loop comparison on Linux

Measured on September 7, 2026 at commit 6cc3444 on an Intel Core i9-9900K, Linux 7.0.0-31-generic (x86_64), and CPython 3.14.7, with rsloop 0.1.49 built in release mode, uvloop 0.22.1, and zuvloop 0.0.14. Each entry is the median of seven measured runs after two warmups, with each run in a fresh subprocess. Times are milliseconds; lower is better.

Workload asyncio uvloop zuvloop rsloop
200,000 callbacks 112.49 51.57 36.23 48.31
50,000 tasks 149.49 93.55 81.47 89.54
5,000 TCP roundtrips 150.87 126.20 109.64 84.60

The TCP workload uses 1,024-byte payloads and rsloop's native fast streams; the other loops use stdlib asyncio streams. Use --no-rsloop-fast-streams to compare all loops through the stdlib streams layer. Zuvloop led callbacks and tasks in this run, while rsloop led TCP roundtrips. These are local microbenchmarks, not isolated-lab measurements or general application performance claims; do not compare them directly with the historical macOS results below. See the full report and recorded results for the exact commands, build details, and process-run samples.

Historical macOS comparison

An earlier example output from the script on macOS (arm64) with CPython 3.14:

callbacks (200,000 ops)
loop           median_s       best_s      ops_per_s     peak_rss   vs_fastest    slower_by
rsloop         0.033083     0.032710      6,045,401     67.5 MiB        1.00x         0.0%
uvloop         0.040958     0.040721      4,883,026     72.8 MiB        1.24x        23.8%
asyncio        0.082233     0.082093      2,432,114     65.3 MiB        2.49x       148.6%

tasks (50,000 ops)
loop           median_s       best_s      ops_per_s     peak_rss   vs_fastest    slower_by
rsloop         0.063593     0.063286        786,247     37.6 MiB        1.00x         0.0%
uvloop         0.069614     0.069420        718,251     38.4 MiB        1.09x         9.5%
asyncio        0.108114     0.107502        462,473     36.1 MiB        1.70x        70.0%

tcp_streams (5,000 ops)
loop           median_s       best_s      ops_per_s     peak_rss   vs_fastest    slower_by
rsloop         0.090940     0.083355         54,981     32.2 MiB        1.00x         0.0%
uvloop         0.133182     0.127404         37,543     31.5 MiB        1.46x        46.5%
asyncio        0.302337     0.299813         16,538     29.6 MiB        3.32x       232.5%

Sustained network workloads

The production-shaped workload matrix exercises HTTP, WebSocket libraries, TLS, mixed message sizes, backpressure, and connection lifecycle behavior:

uv run --with uvloop --with zuvloop python benches/workload_matrix.py \
  --loops rsloop,uvloop,zuvloop \
  --sustained \
  --scenarios http_keepalive,tls_http,websocket_messages,websocket_tls,websockets_messages,websockets_tls,aiohttp_websocket_messages,aiohttp_websocket_tls,starlette_websocket_messages,starlette_websocket_tls,mixed_streams,bulk_transfer \
  --json-output target/matrix-zuvloop.json

Measured on September 7, 2026 with an Intel Core i9-9900K, Linux 7.0.0-31-generic (x86_64), CPython 3.14.7, rsloop 0.1.49 (release build, commit 6cc3444), uvloop 0.22.1, and zuvloop 0.0.14. Each row reports the median of seven measured runs after two warmups, using 16 concurrent connections. Request/response workloads send 500 requests per connection; bulk transfer sends 2 MiB per connection in 64 KiB chunks. Throughput is traffic-only operations per second, except for bulk_transfer, which reports traffic MiB/s. The p95 columns are the medians of each run's p95 latency. Higher throughput and lower latency are better. Each loop/scenario pair runs in its own subprocess, with warmups and measured runs sharing that process. Loops run sequentially in the order shown.

WebSocket library versions were websockets 17.0.1, aiohttp 3.14.3, Starlette 1.6.0, and uvicorn 0.52.3. The run used unrestricted CPU affinity, with other host services running but no concurrent builds or tests. These measurements are from a different host than the macOS microbenchmark example above.

Scenario rsloop uvloop zuvloop rsloop p95 uvloop p95 zuvloop p95
HTTP keep-alive 54,103 51,649 57,764 0.336 ms 0.346 ms 0.292 ms
TLS HTTP 72,496 26,226 23,408 0.241 ms 0.646 ms 0.717 ms
Raw WebSocket 4,870 4,890 4,898 4.454 ms 3.874 ms 3.357 ms
Raw WebSocket over TLS 4,933 4,455 4,397 4.355 ms 4.194 ms 3.800 ms
websockets 24,667 25,800 27,409 0.736 ms 0.631 ms 0.594 ms
websockets over TLS 26,781 14,872 14,555 0.644 ms 1.135 ms 1.167 ms
aiohttp WebSocket 31,636 32,765 35,425 0.605 ms 0.535 ms 0.484 ms
aiohttp WebSocket over TLS 34,165 19,110 18,400 0.504 ms 0.882 ms 0.909 ms
Starlette WebSocket 18,785 20,407 21,743 0.981 ms 0.829 ms 0.787 ms
Starlette WebSocket over TLS 18,509 13,104 13,038 0.948 ms 1.278 ms 1.291 ms
Mixed streams 43,795 34,196 37,116 0.463 ms 0.523 ms 0.473 ms
Bulk transfer (MiB/s) 2,009.9 1,264.9 1,313.3 15.058 ms 25.219 ms 24.317 ms

The former single-burst idle-activation row has been retired: its traffic phase lasted only a few milliseconds and produced unstable throughput rankings. Idle activation now has a separate, versioned latency benchmark described below. In this run, zuvloop had the highest plaintext HTTP and WebSocket throughput, while rsloop led TLS throughput, mixed streams, and bulk transfer. Throughput and tail latency do not always agree: rsloop's raw WebSocket p95 was higher than both alternatives, including over TLS. These results are not an across-the-board performance win or a before/after regression measurement. See the benchmark documentation for workload definitions and reproduction commands.

The ordinary matrix defaults are intentionally short enough for local smoke and CI runs. Even with --sustained, compare repeated runs before drawing performance conclusions for a deployment — competing desktop load matters more than it looks, because rsloop trades helper-thread CPU for loop-thread work and so has more to lose when cores are contended.

Idle activation latency

uv run --with uvloop --with zuvloop python benches/workload_matrix.py \
  --loops rsloop,uvloop,zuvloop --scenarios idle_connections --repeat 9 \
  --idle-cycles 100 --idle-warmup-cycles 5 --idle-seconds 0.2 \
  --json-output target/idle-v2-paired.json

Idle v2 reuses 200 established connections across repeated idle/wakeup cycles. It measures all replies from one shared activation timestamp, including task scheduling delay, and reports first/50%/95%/all-reply latency. Nine fresh-process blocks rotate loop order so each loop runs first three times; confidence intervals resample whole paired runs, not individual connections. Results are classified as improved, regressed, or inconclusive using a 5% practical threshold and an approximate 95% confidence interval. The command takes about ten minutes; use --idle-cycles 3 --idle-warmup-cycles 1 --idle-seconds 0.01 --repeat 1 for a smoke test only.

The new measurements cannot be compared with the retired ops/s row. See benchmark methodology and regression handling for timing definitions, host controls, raw distributions, and sample requirements.

Measured on September 7, 2026 at commit 6cc3444 on the Linux/i9-9900K host above with CPython 3.14.7, rsloop 0.1.49 (release), uvloop 0.22.1, and zuvloop 0.0.14. The run collected 900 measured cycles per loop in 27 distinct processes, with unrestricted affinity and no concurrent builds or test runs. These are medians across runs of each run's median cycle milestone, in milliseconds (lower is better):

Loop First reply 50% replied 95% replied All replied
rsloop 17.221 17.579 17.857 17.991
uvloop 18.193 18.722 19.174 19.221
zuvloop 16.441 16.869 17.248 17.288

Comparisons against uvloop use geometric mean paired process-run ratios, not ratios of the table medians:

  • rsloop: -23.5% p95 latency, approximate 95% interval [-48.6%, +5.9%]; inconclusive at the 5% threshold.
  • zuvloop: -15.2% p95 latency, approximate 95% interval [-31.0%, +4.5%]; inconclusive at the 5% threshold.

Individual cycle-p95 latencies span 3.374–28.184 ms for rsloop, 3.736–24.228 ms for uvloop, and 3.454–26.163 ms for zuvloop. Median ordering alone does not establish a latency win.

The full report links the recorded measurements and documents the settings used for all three benchmark suites.

See benches/README.md for workload details and extra flags, and examples/README.md for the FastAPI loop comparison example.

Acknowledgements

rsloop builds on the Python asyncio model and is implemented with PyO3 on the Rust side. Runtime and socket I/O are powered by vibeio.

License

This project is licensed under the Apache License, Version 2.0. See LICENSE for the full text.

Releases

Used by

Contributors

Languages