GraphQL service for ibis tables. Ibis supports 20+ backends — DuckDB, PostgreSQL, Polars, BigQuery, etc. — so the same query API works across local files and remote databases. The schema is derived automatically.
Parquet datasets are also supported as a root source, with custom optimizations for partitions. As of version 2, execution is based on ibis (default backend: DuckDB).
There is an example app which reads a parquet dataset.
env PARQUET_PATH=... uvicorn graphique.service:appOpen http://localhost:8000/ to try out the API in GraphiQL. There is a test fixture at ./tests/fixtures/zipcodes.parquet.
env PARQUET_PATH=... strawberry export-schema graphique.service:app.schemaoutputs the graphql schema.
The example app uses Starlette's config: in environment variables or a .env file.
- PARQUET_PATH: path to the parquet directory or file
- NAME = '': GraphQL field on
Query; defaults to root type - COLUMNS = None: list of names, or mapping of aliases, of columns to select
Configuration options exist to provide a convenient no-code solution, but are subject to change in the future. Using a custom app is recommended for production usage.
For more options create a custom ASGI app. Call graphique's GraphQL on an ibis Table or parquet Dataset.
Use a Query type with dataset attributes for multiple roots, and to enable federation.
import ibis
from graphique import GraphQL, typed
# any ibis backend: DuckDB, PostgreSQL, Polars, BigQuery, ...
source = ibis.read_(...) # or `ibis.connect(...).table(...)` or `pyarrow.dataset.dataset(...)`
# apply initial projections or filters to `source`
app = GraphQL(source) # Table is root query type
# multiple named fields, with optional federation keys
class Query:
name = source # or `typed(source, name, keys=...)`
app = GraphQL(Query)Start like any ASGI app.
uvicorn <module>:appDataset: interface for an ibis table or parquet dataset.Table: implements theDatasetinterface. Adds typedrow,columns, andfilterfields from introspecting the schema.Column: interface for an ibis column. Each data type has a corresponding column implementation: Boolean, Int, BigInt, Float, Decimal, Date, Datetime, Time, Duration, Base64, String, Array, Struct. All columns have avaluesfield for their list of scalars. Additional fields vary by type.Row: scalar fields. Tables are column-oriented, and graphique encourages that usage for performance. A singlerowfield is provided for convenience, but a field for a list of rows is not. Requesting parallel columns is far more efficient.
slice: contiguous selection of rowsfilter: select rows by predicatesjoin,asofJoin,crossJoin: join tables by key columnsdifference,intersect,union: set operations on tablestake: rows by indexdropNull: remove rows with nulls
project: project columns with expressionscolumns: provides a field for everyColumnin the schemacolumn: access a column of any type by namerow: provides a field for each scalar of a single rowcast: cast column typesunpack: project struct fieldsfillNull: fill null values
group: group by given columns, and aggregate the othersdistinct: group with all columnsruns: group by adjacencyunnest: unnest an array columncount,any: number of rows
order: sort table by given columnsfirst: sort and filter by rank
type: type of data sourceschema: field names and typesoptional: nullable for errorstoSql: compiles SQL query
Performance is dependent on the Ibis backend, which defaults to DuckDB. There are no internal Python loops. Scalars do not become Python types until serialized. Table fields are lazily evaluated up until scalars are reached, and automatically cached as needed for multiple fields.
PyArrow is also used for partitioned dataset optimizations. python -m graphique.partition is a command-line script provided in graphique[cli], for out-of-core partitioning.
pip install graphique[server,cli]- ibis-framework (with duckdb or other backend)
- strawberry-graphql[asgi,cli]
- pyarrow
- isodate
- uvicorn (or other ASGI server)
100% branch coverage.
pytest [--cov]