This project adds Python bindings to rust-pgn-reader. In addition, it also parses and extracts [%clk ..] and [%eval ..] tags from comments.
pip install rust_pgn_reader_python_binding
Three entry points are available:
parse_game(pgn)- Parse a single PGN stringparse_games(chunked_array)- Parse games from a PyArrow ChunkedArray (multithreaded)parse_games_from_strings(pgns)- Parse a list of PGN strings (multithreaded)
All return a ParsedGames container with flat NumPy arrays, supporting:
- Indexing (
result[i]), slicing (result[1:3]), and iteration (for game in result) - Per-game views (
PyGameView) with zero-copy array slices - Position-to-game and move-to-game mapping for ML workflows
- Optional comment storage (
store_comments=True) - Optional legal move storage (
store_legal_moves=True)
Below are some benchmarks on Lichess's 2013-07 chess games (293,459 games) on a 7800X3D.
| Parser | File format | Time |
|---|---|---|
| rust_pgn_reader_python_binding, parse_games (multithreaded) | parquet | 0.14s |
| rust_pgn_reader_python_binding, parse_games (singlethreaded) | parquet | 1.3s |
| rust_pgn_reader_python_binding, parse_games_from_strings (multithreaded) | PGN | 0.41s |
| chess-library | PGN | 2s |
| rust-pgn-reader | PGN | 1s |
| python-chess | PGN | 3+ min |
To replicate, download 2013-07-train-00000-of-00001.parquet and then run:
python src/bench_parse_games.py (recommended — multithreaded parse_games via Arrow)
python src/bench_parse_games_singlethreaded.py (singlethreaded parse_games via Arrow)
python src/bench_parse_pgn.py (multithreaded .pgn parsing)
python src/bench_data_access.py 2013-07-train-00000-of-00001.parquet (parsing + data access + memory)
maturin develop
maturin develop --release
For a more thorough tutorial, follow https://lukesalamone.github.io/posts/how-to-create-rust-python-bindings/
samply record --rate 10000 python src\bench_parse_games.py
py-spy record -s -F -f speedscope --output profile.speedscope -- python ./src/bench_parse_games.py
Linux/WSL-only:
py-spy record -s -F -n -f speedscope --output profile.speedscope -- python ./src/bench_parse_games.py
cargo test
python -m unittest src/test.py
We have a manual rust_pgn_reader_python_binding.pyi python type stub.
To keep it in sync with the Rust bindings, run:
python tools/check_stubs.py
This regenerates the stub via maturin generate-stubs (PyO3
experimental-inspect, behind the stubgen Cargo feature) and diffs the
symbol surface (module functions, classes, methods/properties, but not
annotations) against the hand-written file. CI runs the same check in the
stubs job.
- Interleaved parquet read + parse
- Opening cache. For positions that occur plenty of times, a
(zobrist(pos), san_token)→Movecache could skipsan_resolver+ legality on cache hits. - crude move-count estimator from movetext byte length for buffer preallocation (bytes-per-ply ≈ 6-6.5, over-estimate is relatively cheap, under-estimate forces realloc copies). We currently assume 70 moves per game. Since we have long chunks (not per-game), long game moves eat up the slack from short games. Wont' help on Lichess games (we're already doing quite well). But if going with a different / unknown corpus, 70 will be off (different game lengths, comments). But comments will also throw off our byte-based move count estimator.