Skip to content

ArrowMetal 0.5.0

Latest

Choose a tag to compare

@singhpratech singhpratech released this 08 Oct 18:30
· 1 commit to main since this release

Everything below is new in 0.5.0. The Polars MetalEngine runs on polars 2.0.0 (IR (15, 2)) as on 1.44, and
the default policy loads one measured crossover table per Polars major: on 2.0.0, whose own group-bys,
joins and sorts run at 0.72x of 1.44.1's time at the median, the default takes 52 of 107 benchmarked
in-memory cases at 50,000,000 rows, every one 1.35x to 7.94x faster than the faster Polars 2.0.0 engine and
none behind (Benchmarks/results/polars_engine_bench_2026-10-08-polars2-refit.csv); on 1.44.1 the 0.4.0
table and its numbers stand. A mean over an integer column runs in one pass from the integer column in
the Polars and plan engines, with the same bits as before (two int32 keys, 200 groups: 20.7 to 14.3 ms at
50,000,000 rows on polars 2.0.0). The group-count estimate cache holds on polars 2.0.0. The pip extra
installs polars>=1.44,<2.1. The DuckDB extensions are built and tested against DuckDB 1.5.6.

  • The default's crossover table for polars 2.0.0. python/arrowmetal/_engine_crossovers_pl2.py,
    fitted from a sweep on polars 2.0.0 (Benchmarks/results/polars_engine_crossover_2026-10-08-polars2.csv:
    250,000 to 50,000,000 rows in memory and the Parquet scan cases at 1,000,000 to 50,000,000 rows,
    one process per size, every one started at a 1-minute load below 3.5 on AC power; conditions in
    polars_engine_crossover_2026-10-08-polars2_conditions.txt) and the default benchmark
    (polars_engine_default_groupby_raw_2026-10-08-polars2.csv: 87 group-by, unique and sort cases at
    2,000,000 and 50,000,000 rows, and the sort, unique and 1,000- and 10,000-group cases at 5,000,000,
    10,000,000 and 20,000,000 rows). Polars 2.0.0 runs its group-bys, joins and sorts faster than 1.44.1
    (its own time is 0.72x of 1.44.1's at the median over the cases the 1.44 default takes, down to
    0.25x on unique over two keys), so on 2.0 the default leaves more to Polars: two-key and
    one-key sum and mean at 200 to 10,000 groups, the String-column sort (and (p) of the same
    class), the String-key join, semi join and top-k; the numeric sort crossover is 10,000,000 rows
    (1.44: 947,836), unique 2,943,034 (1.44: 5,494,090). Measured with the table in force
    (polars_engine_bench_2026-10-08-polars2-refit.csv, 50,000,000 rows, AC): the default takes 52 of
    107 in-memory cases, every one 1.35x to 7.94x faster than the faster Polars 2.0.0 engine, none
    behind; the 12 cases it stopped taking were 0.89x to 2.75x. Benchmarks/polars_engine_crossover.py --check regenerates both modules from their own inputs.
  • Group-count estimates cached on polars 2.0.0. The engine keys its group-count estimate cache
    on the key columns' buffer addresses through Series._get_buffer_info(), which polars 2.0 removed;
    on 2.0 every probe sampled again. A single-chunk column is now named through the pointer the
    Series exposes (PySeries.as_single_ptr()), so the cache holds on 2.0 as on 1.44; a chunked or
    String column still gets no entry and is probed each time.
  • One crossover table per Polars major. MetalEngine()'s default policy loads its crossover
    table by the running Polars' major (_engine_policy.TABLES, table_for), on its first decision
    rather than at import: Polars 1.x loads _engine_crossovers.py, the 1.44.1 table, unchanged;
    Polars 2.x loads _engine_crossovers_pl2.py, the module generated from the 2.0.0 sweep, or, in
    an install without it, the 1.44.1 table with a warning the first time. compatibility() gains
    crossover_table (the module, its SOURCE and the Polars version of its sweep), and
    python -m arrowmetal.polars_engine check prints it. Benchmarks/polars_engine_crossover.py
    gains --out (the module to write or check), writes a major's module only from a sweep run on
    that major, and --check checks every table module present. test_engine_policy.py runs the
    table tests and the --check test once per table module.
  • Integer means in the Polars and plan engines, one pass. A mean over an integer column
    (Int8 to Int64, UInt8 to UInt64) is computed from the integer column itself by the group-by's
    integer mean kernel: the exact sum of each group divided by its count, rounded once. Before, the
    plan materialised the column cast to Float64 (400 MB at 50,000,000 rows) and averaged that. The
    answer is the same: the engine checks in one min/max pass that every value converts to Float64
    exactly (below 2^53) and that no group's sum can wrap 64 bits, and otherwise takes the Float64
    path as before (Executor.integerMeanIfExact, test_integer_mean_matches_polars_for_ordinary_and_extreme_values).
    On Polars 2.0.0 at 50,000,000 rows (Benchmarks/results/polars_engine_bench_2026-10-07-polars2.csv
    before, polars_engine_bench_2026-10-08-polars2-after.csv after, both on AC power with their
    conditions files): two int32 keys, mean of an int64 column, 200 groups 20.7 -> 14.3 ms, 1,000
    groups 20.9 -> 14.8 ms, 100,000 groups 28.2 -> 18.9 ms, 1,000,000 groups 36.6 -> 25.5 ms; one
    key, 100,000 groups 24.5 -> 16.7 ms. Over the 62 cases the default takes on 2.0.0: 0.86x to
    7.72x with four behind before, 0.89x to 7.88x with one behind after (the sort with a String
    column, 0.89x).
  • Polars 2.0.0 and the engine. MetalEngine runs on polars 2.0.0 (IR (15, 2)) as on 1.44
    (IR (14, 7)). TESTED_IR_VERSION is now one version per IR major, ((14, 7), (15, 2)), and
    TESTED_POLARS is ("1.44.1", "1.44.2", "2.0.0"); an IR major outside these keeps every plan with
    Polars, and a newer minor of a tested major runs and warns once. KNOWN_NODE_KINDS keeps
    ExtContext, which polars 2.0.0 no longer has (2.0 removed LazyFrame.with_context).
    compatibility() gains ir_tested, and python -m arrowmetal.polars_engine check names the Polars
    the capability table was generated on when it differs. Polars 2.0 removed LazyFrame.profile;
    there MetalEngine.profile raises NotImplementedError. On polars 2.0.0 test_polars_engine.py,
    test_polars.py and test_polars_sort_order.py pass, with the profile tests and the
    capability-table comparison skipped (docs/ENGINE_CAPABILITIES.md is the 1.44.1 table;
    docs/ENGINE_CAPABILITIES_POLARS2.md is the same generator run on 2.0.0: the engine takes the same
    1,326 of 2,523 plans, and 63 plans Polars 1.44.1 accepted are rejected by Polars 2.0.0 itself). The
    in-memory engine of 2.0.0 lowers x * -1.0 to a negation and x / c to a multiply by the
    reciprocal as 1.44.1 does; the streaming engine, the default of collect() on 2.0.0, evaluated
    frames of 2 and 9 rows element-wise on both versions, so test_multiply_by_minus_one_is_a_negation_like_polars
    compares with engine="in-memory".
  • Polars bench: keep="any". Benchmarks/polars_engine_bench.py gains case (r2), unique over
    (region, sub) with keep="any", beside (r) with keep="first". A run of (r), (r2) and (m) at
    50,000,000 rows on 2026-10-04 (Benchmarks/results/polars_engine_bench_2026-10-04.csv, conditions in
    bench_conditions_2026-10-04-polars.txt): Polars in-memory 161.74 ms (keep first) and 150.31 ms (keep
    any) against MetalEngine() 15.71 and 15.72 ms, 10.3x and 9.6x; the sort 414.61 against 58.77 ms,
    7.06x. The README and docs/POLARS.md keep citing the 2026-10-02 run, the one every Polars figure
    comes from.
  • DuckDB 1.5.6. Both DuckDB extensions are built and tested against DuckDB 1.5.6 (source_id 069cc9f9b5), with no code change: duckdb-extension/build.sh fetches the v1.5.6 headers, and
    build_rewrite.sh builds the optimizer extension for the release of the installed duckdb module.
    On 1.5.6 the DuckDB tests pass as on 1.5.5 (test_duckdb.py and test_duckdb_rewrite.py, 252
    passed), and the conformance grid gives the same counts shape for shape: 33,376 queries, 21,844
    rewritten and identical, 11,532 left to DuckDB, 0 different
    (Benchmarks/results/engine_conformance_2026-10-04.csv). The rewrite benchmark rerun on 1.5.6 on
    AC power (Benchmarks/results/duckdb_rewrite_2026-10-04-ac.csv, conditions in
    duckdb_rewrite_2026-10-04-ac_conditions.txt) gives the same auto floors, and the floors now cite
    it: every query auto rewrites is 1.08x to 5.32x faster than DuckDB on the best run and 1.04x to
    4.54x on the median; DuckDB's own best times are 0.96 to 1.12 times those of the 1.5.5 run of
    2026-10-02. On the first run after 500 ms of idle, four of the thirteen rewritten pairs are behind
    DuckDB's, down to 0.38x. A run of the same day on battery power
    (Benchmarks/results/duckdb_rewrite_2026-10-04.csv) gives the same floors.