Skip to content

v0.1.7 — Fastest SPSS Reader

Choose a tag to compare

@albertxli albertxli released this 18 Feb 06:34
· 105 commits to main since this release

Performance: 6 Optimizations

ambers is now faster than polars_readstat on every file, from small (0.2 MB) to large (1.1 GB).

Optimizations

  1. Bytecode match reorder — hot path 1..=251 first for better branch prediction
  2. Cow<str> for string decoding — eliminates heap allocation for UTF-8 files (most modern SPSS)
  3. Bulk I/O + buffer reuse — single read_exact per row for uncompressed data (was 72M individual reads on 1.1 GB file)
  4. Pre-computed VLS segment layout — avoids redundant per-row arithmetic for very long strings
  5. Smart capacity hints — StringBuilder uses actual variable width instead of hardcoded 32 bytes
  6. StringViewArray with deduplication — Utf8View + deduplicate_strings() for categorical SPSS data

Results (avg 5 runs, Polars DataFrame output)

File ambers polars_readstat pyreadstat vs polars_readstat vs pyreadstat
test_1 (0.2 MB, bytecode) 0.002s 0.004s 0.328s 2.0x faster 175x faster
test_2 (147 MB, bytecode) 0.880s 0.949s 3.618s 1.1x faster 4.1x faster
test_3 (1.1 GB, uncompressed) 1.094s 1.359s 5.002s 1.2x faster 4.6x faster
test_4 (0.6 MB, uncompressed) 0.013s 0.015s 0.022s 1.1x faster 1.7x faster
test_5 (0.6 MB, uncompressed) 0.002s 0.004s 0.016s 1.9x faster 8.2x faster

Lazy Pushdown

File Full collect Select 5 + head 1000 Speedup
test_2 (147 MB) 0.833s 0.084s 9.9x
test_3 (1.1 GB) 1.036s 0.006s 167x

On the 1.1 GB file, selecting 5 columns and 1000 rows completes in 6ms.

Full Changelog: v0.1.6...v0.1.7