Skip to content

v0.1.5

Choose a tag to compare

@github-actions github-actions released this 18 Feb 02:54
· 112 commits to main since this release

What's New in v0.1.5

scan_sav() — Lazy Reader for Polars

New scan_sav() returns a Polars LazyFrame with full pushdown support:

import ambers

lf, meta = ambers.scan_sav("survey.sav")
df = lf.select(["Q1", "Q2", "age"]).head(1000).collect()
  • Column projection pushdown — only reads selected columns
  • Row limit pushdown — stops reading after N rows
  • Predicate filtering — per-batch filter applied during scan
  • Powered by Polars register_io_source plugin API

Performance Overhaul

1. Direct-to-Columnar Builders

Replaced the 3-stage row-oriented pipeline with direct columnar construction:

Before:  Binary → Vec<Vec<SlotValue>> → Vec<Vec<CellValue>> → Arrow RecordBatch
After:   Binary → push directly into Arrow column builders → RecordBatch

This eliminates ~N×M intermediate heap allocations (e.g., 15M objects for a 22K×677 file). Decoded values are pushed straight into Arrow Float64Builder / StringBuilder with a reusable byte buffer for strings.

2. Arrow PyCapsule Interface (PyArrow dropped)

Replaced arrow::pyarrow::ToPyArrow with the Arrow PyCapsule Interface (__arrow_c_stream__):

Before:  RecordBatch → PyArrow (via arrow/pyarrow FFI) → pl.from_arrow()
After:   RecordBatch → PyCapsule (arrow_array_stream) → pl.from_arrow()

pyarrow is no longer a runtime dependency. Data transfers to Polars are zero-copy via the C Data Interface capsule protocol, supported in Polars >= 1.3.

3. Dead Code Cleanup

Removed the old row-oriented pipeline (data.rs, rows_to_record_batch, decompress_all_rows) — net reduction of 107 lines. Zero compiler warnings.

Benchmarks

Eager Read

All results return a Polars DataFrame. Average of 5 runs on Windows 11, Python 3.13, 24-core machine.

File Size Shape ambers polars_readstat pyreadstat pyreadstat mp (4w) ambers vs polars_readstat ambers vs pyreadstat
test_1 (bytecode) 0.2 MB 1,500 × 75 0.002s 0.012s 0.010s 0.409s 6.4x faster 5.5x faster
test_2 (bytecode) 147 MB 22,070 × 677 1.119s 1.091s 4.351s 1.773s ~tied 3.9x faster
test_3 (uncompressed) 1.1 GB 79,066 × 915 1.713s 1.532s 6.390s 2.635s ~tied 3.7x faster
test_4 (uncompressed) 0.6 MB 201 × 158 0.015s 0.023s 0.020s 0.424s 1.6x faster 1.3x faster
test_5 (uncompressed) 0.6 MB 203 × 136 0.003s 0.013s 0.014s 0.417s 4.3x faster 4.8x faster
  • vs pyreadstat: 4–6x faster across all file sizes
  • vs pyreadstat multiprocess (4 workers): ambers single-threaded still faster on every file
  • vs polars_readstat: tied on large files, 2–6x faster on small/medium files (lower startup overhead)

pyreadstat multiprocess returns pandas; timing includes pl.from_pandas() conversion.

Lazy Read with Pushdown

scan_sav() returns a Polars LazyFrame — it only reads the data you ask for:

File (size) Full collect Select 5 cols Head 1000 rows Select 5 + head 1000
test_2 (147 MB, 22K × 677) 1.282s 0.478s (2.7x) 0.205s (6.3x) 0.160s (8.0x)
test_3 (1.1 GB, 79K × 915) 1.668s 0.822s (2.0x) 0.031s (53.5x) 0.022s (75.8x)

On the 1.1 GB file, selecting 5 columns and 1000 rows completes in 22ms — 76x faster than reading the full dataset.

Full Changelog: v0.1.4...v0.1.5