PyArrow Tables + Parquet in Python (2026)
PyArrow Tables + Parquet in Python (2026) shows how to build an in-memory pa.table, run vectorized pyarrow.compute filters and aggregates, write a compressed Parquet file, and prove a lossless round-trip — without starting a database server.
Once the Arrow layer clicks, move data across engines with Narwhals for Pandas and Polars, query portable frames with Ibis + DuckDB, or stay in-process with DuckDB analytics.
TL;DR
- Apache Arrow is the columnar memory format;
pyarrowis the Python binding. - Build tables with
pa.table({...}), then usepyarrow.computefor filters and sums. - Parquet is the on-disk columnar twin — use
pq.write_table/pq.read_table. - Install with
pip install pyarrow(we tested25.0.1on Python 3.13.5). - Real run:
year >= 2021ranked D400 first at revenue2250.00;revenue_sum=6369.83; Parquet zstd round-tripequal True.
Why PyArrow in 2026?
Polars, pandas 2.x, DuckDB, and many ML loaders speak Arrow. Learning Tables + Parquet once lets you pass zero-copy batches between tools instead of rewriting CSV pipelines. If you only know DataFrames, Arrow is the shared substrate underneath.
| API | Role | Use in 2026 |
|---|---|---|
pa.table | In-memory columnar table | Build / transform batches |
pyarrow.compute | Vectorized kernels | Filter, sort, aggregate |
pyarrow.parquet | On-disk columnar storage | Interchange + compression |
Versions tested (2026-10-02)
- Python
3.13.5 pyarrow25.0.1- Compression:
zstdParquet
python -m venv .venv && source .venv/bin/activate
pip install pyarrow==25.0.1
python arrow_demo.py
1. Build a Table and add revenue
Save this as arrow_demo.py. Create six sales rows, multiply units by price with pc.multiply, and print the schema plus total revenue.
from pathlib import Path
import pyarrow as pa
import pyarrow.compute as pc
import pyarrow.parquet as pq
OUT = Path("sales.parquet")
if OUT.exists():
OUT.unlink()
print("pyarrow", pa.__version__)
table = pa.table(
{
"sku": ["A100", "B200", "C300", "A100", "B200", "D400"],
"year": [2020, 2021, 2022, 2023, 2021, 2024],
"units": [12, 40, 18, 55, 33, 9],
"price": [19.99, 8.50, 120.0, 19.99, 8.50, 250.0],
}
)
print("rows", table.num_rows, "cols", table.num_columns)
print("schema:", table.schema)
revenue = pc.multiply(table["units"], table["price"])
table = table.append_column("revenue", revenue)
print("revenue_sum", round(pc.sum(table["revenue"]).as_py(), 2))
Columns are typed Arrow arrays. pc.multiply and pc.sum stay in columnar land — no Python row loop required for the aggregate.
2. Filter, sort, write Parquet
Keep rows from 2021 onward, sort by revenue descending, write zstd Parquet, and read it back.
mask = pc.greater_equal(table["year"], 2021)
recent = table.filter(mask).sort_by([("revenue", "descending")])
print("filter year>=2021 order by revenue desc:")
for i in range(recent.num_rows):
row = {name: recent.column(name)[i].as_py() for name in recent.column_names}
print(
f" #{i+1} sku={row['sku']} year={row['year']} "
f"units={row['units']} revenue={row['revenue']:.2f}"
)
pq.write_table(table, OUT, compression="zstd")
loaded = pq.read_table(OUT)
print(f"parquet_bytes={OUT.stat().st_size} compression=zstd")
print("roundtrip_rows", loaded.num_rows, "equal", table.equals(loaded))
print("OK Table + compute + Parquet")
table.equals(loaded) confirms the Parquet round-trip kept every value. Swap zstd for snappy if a downstream tool prefers it.
Real output (this machine)
pyarrow 25.0.1
rows 6 cols 4
schema: sku: string
year: int64
units: int64
price: double
revenue_sum 6369.83
filter year>=2021 order by revenue desc:
#1 sku=D400 year=2024 units=9 revenue=2250.00
#2 sku=C300 year=2022 units=18 revenue=2160.00
#3 sku=A100 year=2023 units=55 revenue=1099.45
#4 sku=B200 year=2021 units=40 revenue=340.00
#5 sku=B200 year=2021 units=33 revenue=280.50
parquet_bytes=1771 compression=zstd
roundtrip_rows 6 equal True
OK Table + compute + Parquet
Five of six rows match year >= 2021. D400 leads on revenue despite only nine units. Full-table revenue sums to 6369.83; the zstd Parquet file is 1771 bytes and equals the in-memory table after reload.
Common upgrades
- Dataset scan:
ds.dataset("data/", format="parquet").to_table(filter=...)for folders of files. - Zero-copy to Polars:
pl.from_arrow(table)or write Parquet andpl.scan_parquet. - pandas:
table.to_pandas()when you need classic DataFrame APIs. - Flight / IPC: stream RecordBatches between processes without CSV.
When to use what
| Tool | Best for |
|---|---|
| PyArrow Table + compute | Typed columnar transforms, interchange |
| Parquet | Durable columnar storage and lakehouse files |
| Polars / pandas / Ibis | Higher-level DataFrame or SQL-ish APIs on top |
FAQ
Is PyArrow only for big data? No — even small tables benefit from typed columns and a portable Parquet file.
Do I need Spark? Not for this workflow. Local pyarrow covers Tables, compute, and Parquet.
How does this relate to Polars? Polars can ingest Arrow/Parquet directly; Arrow is the shared memory format underneath.
Next steps
Replace the toy sales table with your own columns, write partitioned Parquet under a lake path, then query it with DuckDB or Ibis without changing the file format.