Skip to content

Repository files navigation

rust_pgn_reader_python_binding

Fast PGN parsing bindings for Python

This project adds Python bindings to rust-pgn-reader. In addition, it also parses and extracts [%clk ..] and [%eval ..] tags from comments.

Installing

pip install rust_pgn_reader_python_binding

API

Three entry points are available:

  • parse_game(pgn) - Parse a single PGN string
  • parse_games(chunked_array) - Parse games from a PyArrow ChunkedArray (multithreaded)
  • parse_games_from_strings(pgns) - Parse a list of PGN strings (multithreaded)

All return a ParsedGames container with flat NumPy arrays, supporting:

  • Indexing (result[i]), slicing (result[1:3]), and iteration (for game in result)
  • Per-game views (PyGameView) with zero-copy array slices
  • Position-to-game and move-to-game mapping for ML workflows
  • Optional comment storage (store_comments=True)
  • Optional legal move storage (store_legal_moves=True)

Benchmarks

Below are some benchmarks on Lichess's 2013-07 chess games (293,459 games) on a 7800X3D.

Parser File format Time
rust_pgn_reader_python_binding, parse_games (multithreaded) parquet 0.14s
rust_pgn_reader_python_binding, parse_games (singlethreaded) parquet 1.3s
rust_pgn_reader_python_binding, parse_games_from_strings (multithreaded) PGN 0.41s
chess-library PGN 2s
rust-pgn-reader PGN 1s
python-chess PGN 3+ min

To replicate, download 2013-07-train-00000-of-00001.parquet and then run:

python src/bench_parse_games.py (recommended — multithreaded parse_games via Arrow)

python src/bench_parse_games_singlethreaded.py (singlethreaded parse_games via Arrow)

python src/bench_parse_pgn.py (multithreaded .pgn parsing)

python src/bench_data_access.py 2013-07-train-00000-of-00001.parquet (parsing + data access + memory)

Building

maturin develop

maturin develop --release

For a more thorough tutorial, follow https://lukesalamone.github.io/posts/how-to-create-rust-python-bindings/

Profiling

samply record --rate 10000 python src\bench_parse_games.py py-spy record -s -F -f speedscope --output profile.speedscope -- python ./src/bench_parse_games.py

Linux/WSL-only: py-spy record -s -F -n -f speedscope --output profile.speedscope -- python ./src/bench_parse_games.py

Testing

cargo test

python -m unittest src/test.py

Type stubs

We have a manual rust_pgn_reader_python_binding.pyi python type stub. To keep it in sync with the Rust bindings, run:

python tools/check_stubs.py

This regenerates the stub via maturin generate-stubs (PyO3 experimental-inspect, behind the stubgen Cargo feature) and diffs the symbol surface (module functions, classes, methods/properties, but not annotations) against the hand-written file. CI runs the same check in the stubs job.

Further performance squeezing:

  • Interleaved parquet read + parse
  • Opening cache. For positions that occur plenty of times, a (zobrist(pos), san_token) → Move cache could skip san_resolver + legality on cache hits.
  • crude move-count estimator from movetext byte length for buffer preallocation (bytes-per-ply ≈ 6-6.5, over-estimate is relatively cheap, under-estimate forces realloc copies). We currently assume 70 moves per game. Since we have long chunks (not per-game), long game moves eat up the slack from short games. Wont' help on Lichess games (we're already doing quite well). But if going with a different / unknown corpus, 70 will be off (different game lengths, comments). But comments will also throw off our byte-based move count estimator.

About

Fast PGN parsing bindings for Python

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages