Repository navigation
v0.1.5
What's New in v0.1.5
scan_sav() — Lazy Reader for Polars
New scan_sav() returns a Polars LazyFrame with full pushdown support:
import ambers
lf, meta = ambers.scan_sav("survey.sav")
df = lf.select(["Q1", "Q2", "age"]).head(1000).collect()- Column projection pushdown — only reads selected columns
- Row limit pushdown — stops reading after N rows
- Predicate filtering — per-batch filter applied during scan
- Powered by Polars
register_io_sourceplugin API
Performance Overhaul
1. Direct-to-Columnar Builders
Replaced the 3-stage row-oriented pipeline with direct columnar construction:
Before: Binary → Vec<Vec<SlotValue>> → Vec<Vec<CellValue>> → Arrow RecordBatch
After: Binary → push directly into Arrow column builders → RecordBatch
This eliminates ~N×M intermediate heap allocations (e.g., 15M objects for a 22K×677 file). Decoded values are pushed straight into Arrow Float64Builder / StringBuilder with a reusable byte buffer for strings.
2. Arrow PyCapsule Interface (PyArrow dropped)
Replaced arrow::pyarrow::ToPyArrow with the Arrow PyCapsule Interface (__arrow_c_stream__):
Before: RecordBatch → PyArrow (via arrow/pyarrow FFI) → pl.from_arrow()
After: RecordBatch → PyCapsule (arrow_array_stream) → pl.from_arrow()
pyarrow is no longer a runtime dependency. Data transfers to Polars are zero-copy via the C Data Interface capsule protocol, supported in Polars >= 1.3.
3. Dead Code Cleanup
Removed the old row-oriented pipeline (data.rs, rows_to_record_batch, decompress_all_rows) — net reduction of 107 lines. Zero compiler warnings.
Benchmarks
Eager Read
All results return a Polars DataFrame. Average of 5 runs on Windows 11, Python 3.13, 24-core machine.
| File | Size | Shape | ambers | polars_readstat | pyreadstat | pyreadstat mp (4w) | ambers vs polars_readstat | ambers vs pyreadstat |
|---|---|---|---|---|---|---|---|---|
| test_1 (bytecode) | 0.2 MB | 1,500 × 75 | 0.002s | 0.012s | 0.010s | 0.409s | 6.4x faster | 5.5x faster |
| test_2 (bytecode) | 147 MB | 22,070 × 677 | 1.119s | 1.091s | 4.351s | 1.773s | ~tied | 3.9x faster |
| test_3 (uncompressed) | 1.1 GB | 79,066 × 915 | 1.713s | 1.532s | 6.390s | 2.635s | ~tied | 3.7x faster |
| test_4 (uncompressed) | 0.6 MB | 201 × 158 | 0.015s | 0.023s | 0.020s | 0.424s | 1.6x faster | 1.3x faster |
| test_5 (uncompressed) | 0.6 MB | 203 × 136 | 0.003s | 0.013s | 0.014s | 0.417s | 4.3x faster | 4.8x faster |
- vs pyreadstat: 4–6x faster across all file sizes
- vs pyreadstat multiprocess (4 workers): ambers single-threaded still faster on every file
- vs polars_readstat: tied on large files, 2–6x faster on small/medium files (lower startup overhead)
pyreadstat multiprocess returns pandas; timing includes pl.from_pandas() conversion.
Lazy Read with Pushdown
scan_sav() returns a Polars LazyFrame — it only reads the data you ask for:
| File (size) | Full collect | Select 5 cols | Head 1000 rows | Select 5 + head 1000 |
|---|---|---|---|---|
| test_2 (147 MB, 22K × 677) | 1.282s | 0.478s (2.7x) | 0.205s (6.3x) | 0.160s (8.0x) |
| test_3 (1.1 GB, 79K × 915) | 1.668s | 0.822s (2.0x) | 0.031s (53.5x) | 0.022s (75.8x) |
On the 1.1 GB file, selecting 5 columns and 1000 rows completes in 22ms — 76x faster than reading the full dataset.
Full Changelog: v0.1.4...v0.1.5