Skip to content

ArrowMetal 0.3.0

Latest

Choose a tag to compare

@singhpratech singhpratech released this 26 Sep 23:21
· 38 commits to main since this release

Everything below is new in 0.3.0. MetalEngine() decides per subtree from measured crossovers: in the
default benchmark (194 case-size pairs at 2,000,000 and 50,000,000 rows) it took 62 pairs, 42 of them
group-bys, every one ahead of the faster Polars engine, 1.21x to 9.42x
(Benchmarks/results/polars_engine_bench_2026-09-26-final3.csv). The engine takes a Polars
scan_parquet of one local file and reads it on the GPU. The C Data import and export take utf8_view
/ binary_view, the string kernels run on that layout, and Polars' String columns are handed over in
it. A whole-file Parquet read of the 50,000,000-row, 8-column files through a fresh open is 65 ms
(Snappy), 54 ms (LZ4) and 38 ms (uncompressed), against Polars' 95, 73 and 63 ms in the same run, with
Snappy and LZ4 pages decompressed by the host and the GPU at the same time
(Benchmarks/results/parquet_bench_2026-09-26-split.txt). python -m arrowmetal.router calibrate fits
the router's table on the Mac it runs on, and explain prints a decision with its reason and the
measured points behind it. The conformance grids compare the Polars engine with Polars (12,597 cases)
and the DuckDB rewrite with DuckDB (33,376 cases) bit for bit, with 0 unclassified differences. A
group-by over 2^24 groups or more returns the right Float64 and Float32 sums and means, products and
lists. The wheel carries the Polars expression plugin, so all four Polars tiers run from
pip install.

  • Parquet Snappy and LZ4 pages are decompressed by the host and the GPU at the same time, split page by
    page. A read's router (DecodeRouter) orders every Snappy and LZ4 page of the columns it stages by how
    token-dense its header says it is (its ratio: at or below 1.0 a literal, above 1.05 token-dense), and
    gives the host the densest ones until the host's share, spread over its cores, is predicted to finish
    with the GPU's, from measured per-byte costs on each side: a 160 KB token-dense page takes one CPU core
    41-81 µs and the GPU 12-15 ms on a SIMD group or 37-60 ms on a thread, while 1.2 GB of literal pages
    take the GPU 10 ms and 16 host threads 19 ms from a fresh mapping (docs/PARQUET.md, "Decompression").
    ZSTD, GZIP and BROTLI stay host-only in the same schedule. The GPU dispatch is committed first (slowest
    pages first), the host decodes its pages on every core while it runs, and the read waits once before
    the value kernels; the host's pages sit in their own page-aligned range of the staging buffer. A new
    bounds-checked LZ4Host decoder makes the kernels' checks with their outcomes; SnappyHost copies 8
    and 16 bytes at a time inside the page and the slot. This replaces the 16-page host rule for Snappy and
    the 2,048-page rule for the page-per-thread kernel. On the 50,000,000-row, 8-column files a whole-file
    read through a fresh open is 65 ms (Snappy) and 54 ms (LZ4), against 108 and 96 ms before and against
    Polars' 95 and 73 ms and pyarrow's 160 and 159 ms in the same run; it takes 674 and 508 CPU-ms, against
    80 and 78 before and Polars' 1,283 and 1,007 (Benchmarks/results/parquet_bench_2026-09-26-split.txt,
    parquet_cache_2026-09-26-split.csv; run conditions in bench_conditions_2026-09-26-split.txt). A
    10,000,000-row, 7-column pyarrow-default Snappy file reads in 25.4 ms against 67.5 ms.
    ARROWMETAL_PARQUET_DECODE=host|gpu|lane sends every Snappy and LZ4 page to one decoder. Tests: every
    fixture on every decoder against the split and against pyarrow, the 240 damaged files on each decoder,
    the host decoders against byte-at-a-time references and under damage between guard pages, the router
    and the staging layout (DecompressSplitTests, ParquetTests, test_parquet.py).
    Benchmarks/parquet_bench.py labels ArrowMetal's row arrowmetal.
  • MetalEngine()'s default (shapes="measured") decides per subtree from measured crossovers instead
    of taking large sorts only. Each translated subtree has shape classes (rowwise; aggregate,
    group_by and group_by_multi per aggregate family sum, count, mean, minmax; sort,
    sort_helper_keys, top_k; join:inner|left|semi|anti; distinct), a dtype class (string when a
    String column is among its inputs) and an input (in-memory frames, or a Parquet file judged by its
    footer's row count). It runs on Metal when its input rows are at or above the crossover of every
    class in it: the largest of the engine table (python/arrowmetal/_engine_crossovers.py, fitted by
    the new Benchmarks/polars_engine_crossover.py from Benchmarks/results/polars_engine_crossover_2026-09-26-final.csv,
    93 in-memory cases at eight sizes from 250,000 to 50,000,000 rows and the Parquet cases at six sizes
    from 1,000,000 to 50,000,000, best of 7, run conditions in bench_conditions_2026-09-26-final.txt),
    the router table in force for the kernels it routes, and the sort kernels' crossover against the
    fastest CPU library. Three rules sit on the fit, each kept in the table: a case counts as ahead when
    its time x 1.15, or x 1.35 for a shape with a String column, is at most the faster Polars engine's
    (MARGIN; a String shape's advantage grows slowly with size, the String sort being 0.98x at 1M rows
    and 1.37x at 5M); a shape with a String column is taken from 5,000,000 rows at the earliest
    (STRING_FLOOR; below it the String sorts are 1.2x and 1.38x in the sweep, inside the benchmark's
    noise band); and the default takes a shape from 1.5 times its fitted crossover (HEADROOM; the fit
    interpolates between sizes measured 2-2.5x apart, and shapes just past it were within run-to-run
    noise of Polars). The table keeps the fit as fit and what the default uses as rows. Taken:
    sorts from 1,026,501 rows and helper-key sorts from 1,000,000 (the sort kernels), numeric-key left
    joins from 1,875,000 input rows, inner joins from 2,029,827 and anti joins from 3,727,959, unique
    from 5,494,090, sorts of a Parquet file from 1,500,000, helper-key sorts and unique with a String
    column from 5,000,000 and sorts with one from 6,301,531; group-bys by their number of groups (next
    entry). Whole-frame aggregates, top-k, semi joins, row-wise shapes and the other String shapes are
    not taken. shapes="all", min_rows= and a new explicit set of class names
    (shapes={"sort", "join"}) override it. Each taken subtree's report entry carries rule, shape,
    dtype_class and input, and each node the policy leaves has a Kind#id: rule: ... line ("900,000
    input rows is below the 2,029,827-row crossover for join:inner (...)").
    arrowmetal.polars_engine.placement_rules() lists the table; SHAPE_CLASSES replaces
    MEASURED_SHAPES, and MetalEngine().min_rows is None unless given.
    Benchmarks/polars_engine_bench.py gains 85 cases and --crossover;
    python/tests/test_engine_policy.py tests the policy.
  • MetalEngine() judges a group-by by its number of groups. The crossover sweep has a group-by grid
    (each aggregate family over one int32 key and over two, keys drawn from 200, 1,000, 10,000, 100,000
    and 1,000,000 values and from half the rows) and records each group-by's group count; every
    group-by class is fitted per bucket of group counts (200, 1,000, 10,000, 100,000, 1,000,000, and at
    least a quarter of the input rows), each bucket from the worst of its cases at each size
    (_engine_crossovers.GROUPS, arrowmetal.polars_engine.group_placement_rules()), with the margin
    and headroom above. The bucket of at least a quarter of the rows is never taken (UNTAKEN_BUCKETS):
    the sweep is not monotone there, the one-key count over 0.43 times as many groups as rows running
    0.78x, 1.9x, 4.04x, 2.21x and 1.18x the faster Polars engine from 2M to 50M rows. At plan time the
    engine estimates the group count of a group-by whose keys are columns of one in-memory input frame:
    the distinct key tuples of a fixed-seed stratified sample, scaled by the bias-corrected Chao1
    estimator, from 512 sampled rows and four times more until every count in the estimate's range gets
    the same decision (at most 65,536 rows); the samples are cached per frame, key columns and size
    (clear_group_estimates()), so the same frame gets the same estimate and decision on every collect
    and in every process. A Parquet file's footer distinct counts are read when every row group states
    them. Taken now: over 10,000 and 100,000 groups every numeric group-by class, from 1,082,526 to
    16,235,764 rows by class and bucket; over 1,000,000 groups every class but the one-key mean, from
    7,500,000 (7,885,821 for the two-key mean); over two keys the count and min/max also at 200 groups
    and every family at 1,000; the (String, int32) sum over about 1,000,000 groups from 9,744,372. Not
    taken: one-key group-bys over 200 or 1,000 groups, and every group-by over a quarter of the rows or
    more; a group-by with no estimate (a computed key, two group-bys in one subtree) is judged by its
    class row, which takes the two-key count from 26,547,327 rows, the (String, int32) sum from
    9,744,372 and no other numeric group-by. The report names the estimate in the rule ("estimated 191
    groups over (region), a Chao1 estimate from a 512-row sample, 189 to 196: below the measured band
    for group_by:sum at 50,000,000 input rows (taken at 3,163 to 3,162,277 groups; ...)") and lists each
    probe with its time (last_report.groups). In Benchmarks/results/polars_engine_bench_2026-09-26-final3.csv
    (the 93 in-memory cases at 2M and 50M rows and the 50M Parquet cases, 194 case-size pairs; run
    conditions in bench_conditions_2026-09-26-final3.txt) the default took 62 pairs, 42 of them
    group-bys, every one ahead of the faster Polars engine, 1.21x to 9.42x; in the 132 pairs it left
    to Polars its time was 0.62x to 1.31x Polars' in-memory time where that is 5 ms or more.
    The probe takes 66 to 461 µs at 50,000,000 rows (0.10% to 1.19% of the group-by it decides) and 50
    to 393 µs at 2,000,000 (1.09% to 4.90%), median of 9 cold collects under the earlier group-count
    table (Benchmarks/results/group_probe_2026-09-26.csv, from the new Benchmarks/group_probe_bench.py);
    in the default benchmark's process, which holds every case's frames, its median is 99 µs at 2M and
    144 µs at 50M, and it reaches 13.5 ms on two-key frames with 0.43 times as many groups as rows at 50M.
  • The crossover sweep's (v3) case found ArrowMetal's group-by returning wrong Float64 sums and means from 16,777,216 groups (most groups null); fixed in the core (below), so the engine needs no guard for it.
  • Parquet reads, cold and warm. On the 50,000,000-row, 8-column benchmark files a whole-file read
    through a fresh open is 108 ms (Snappy), 96 ms (LZ4) and 41 ms (uncompressed), against 369, 340 and
    250 ms in parquet_bench_2026-09-25-quiet.txt and against Polars' 99, 74 and 67 ms in the same run,
    with 78-91 ms of process CPU against 760-1330 for pyarrow and Polars; time to first compute is
    9-13 ms (Polars 16-22); a one-column read through a fresh open is 6.9-12.2 ms, against 149-191 ms
    before (Benchmarks/results/parquet_bench_2026-09-26-quiet.txt, parquet_cache_2026-09-26-quiet.csv;
    run conditions in bench_conditions_2026-09-26-quiet.txt). What changed
    (docs/PARQUET.md, "The file's bytes are the GPU's bytes" and "Decompression"):
    • the file is mapped read-only and shared: the first kernel to read a wrapped range makes it resident
      for the GPU at 7-10 ms per GB, against 67-78 ms per GB for the private writable mapping, which stays
      as the fallback for a device that will not wrap read-only memory and on virtualised GPUs;
    • wrapping the reader's own mapping skips MetalArrowBuffer.wrapOrCopy's mincore probe, which took
      about 35 ms over 2 GB;
    • a projection whose chunks fill less than four fifths of the bytes they span maps them side by side
      (MappedView) and wraps only them; a column whose chunks span 4 GiB or more of a larger file now
      reads (the limit is 4 GiB of one column's chunks in one read);
    • page headers are parsed in parallel across the read's chunks and kept on the handle;
    • dense fixed-width values and all-dictionary code buffers are no longer zero-filled before kernels
      that write every slot (the padding still is), and dictionary codes get their row group's base as they
      are decoded instead of in a second pass;
    • ZSTD pages decode in runs sharing one ZSTD_DCtx instead of a fresh context per page;
    • token-dense Snappy and LZ4 pages (output at least 1.25x the input, in dispatches of at least 2,048
      such pages) decode one page per thread, and a read decompresses the pages of all its flat columns in
      one dispatch per codec before decoding them;
    • a closed file is unmapped on a background queue.
      ARROWMETAL_PARQUET_PROFILE=1 prints per-phase times and minor faults of each open and read.
      ParquetTests adds reads of every fixture through both mappings, one-column and row-group reads through
      column views, both decompression kernels on every Snappy and LZ4 fixture, and 240 damaged files through
      the page-per-thread kernel; test_parquet.py adds 2 to 400 pages a chunk in ZSTD, LZ4 and Snappy.
  • python -m arrowmetal.bench: one seeded 10,000,000-row dataset (drawn with pyarrow.compute from
    SplitMix64 streams, so nothing beyond pip install arrowmetal is needed), sum, filter, sort and group-by sum
    through pyarrow (and Polars when installed) and through ArrowMetal, each answer checked against
    pyarrow, one table for the Mac it runs on with a ready-to-paste block; --rows, --json, --quiet,
    --no-share, --no-polars. A GitHub issue form (.github/ISSUE_TEMPLATE/benchmark_result.yml)
    takes that block (docs/TESTING.md, python/tests/test_bench.py).
  • Router table per machine: python -m arrowmetal.router calibrate [--quick] [--out PATH] [--csv PATH]
    runs the router check sweep on the Mac it is run on (the sweep moved from Benchmarks/router_check.py
    into the package, which that script now calls), fits it with the code Benchmarks/router_table.py
    uses (python/arrowmetal/_router_fit.py) and writes ~/.arrowmetal/router/<chip>.json with the chip,
    core counts, Metal device, date, ArrowMetal version, grid and every measurement. At first use the
    router takes ARROWMETAL_ROUTER_TABLE (a path, or shipped), else that file for this chip, else the
    shipped table; an operation the sweep did not bring to a crossover keeps its shipped row.
    python -m arrowmetal.router explain <op> <rows> [--dtype] [--nulls] [--keys] [--json] prints the
    decision, its reason, the table row and the two measured points its crossover was fitted between, and
    whether the table is the shipped one or this machine's. router_table.py --json-out writes the shipped
    table in the same JSON format, and python -m arrowmetal.bench --calibrate runs the quick calibration
    after the benchmark. C: am_router_load_table, am_router_table_info, am_router_decide,
    am_router_explain; Swift: Router.table, Router.loadTable(path:), Router.useShippedTable(),
    Router.explain, RouterCrossovers; Python: am.router_table(), am.load_router_table(),
    am.route_decision(), am.explain_route(). am_router_crossover, Router.crossoverRows and
    am.router_crossovers() report the table in force. The Swift and Python test harnesses pin the
    shipped table unless ARROWMETAL_ROUTER_TABLE is set (docs/CROSSOVER.md).
  • Router determinism: Router.route, the rule every routed call goes through, is a pure function of the
    operation, value type, row count, mode, batch state and table row; nothing is timed at call time.
    python/tests/test_router_calibrate.py checks 200 (operation, size) pairs for the same decision
    across 1,000 calls and across two processes, and RouterTests.testRouteIsPure checks the Swift rule.
  • python -m arrowmetal.bench --parquet FILE: the same report on the user's own Parquet file. It reads
    the file's integer, floating-point and string columns with pyarrow.parquet.read_table,
    polars.read_parquet and am.read_parquet, then runs sum and filter > median on the largest numeric
    column and group-by sum keyed on the lowest-cardinality integer or string column, CPU against Metal,
    each ArrowMetal answer checked against pyarrow's. The Share it block and the prefilled issue link
    (title and the form's share field) carry the row count, column count, row groups, size, codecs and
    the timings, never the path, column names or values. A file whose columns would take more than a
    quarter of physical memory is refused with the limit printed; a file with no usable column gets one
    line saying so. On a generated 10,000,000-row Snappy file (M4 Max): read 134.58 ms against Polars'
    16.75 ms, filter 1.12 ms against 4.50 ms, sum 1.40 ms against 0.99 ms, group-by sum over a 5-value
    string key 18.25 ms against 6.82 ms (python/tests/test_bench.py).
  • The wheel carries the Polars expression plugin (tier 2): python/build_wheel.sh builds
    polars-plugin/ with cargo against the libArrowMetalC.dylib it bundles, and packages
    arrowmetal/_lib/libarrowmetal_polars.dylib with its rpath set to @loader_path and local symbols
    stripped. arrowmetal.polars_plugin.plugin_path() takes the packaged plugin first when Python loaded
    the packaged libArrowMetalC.dylib, and a cargo build first when $ARROWMETAL_LIB or a development
    build is loaded. scripts/check_wheel.sh installs a wheel with its polars extra into a fresh
    virtualenv with a scrubbed environment and no cargo, runs all four Polars tiers, bench and
    bench --parquet, and checks that one libArrowMetalC.dylib, the packaged one, is loaded; it passed
    with polars 1.44.2 and pyarrow 25.0.1, NumPy not installed. The wheel is 8.2 MB (3.2 MB without the plugin) (docs/POLARS.md, Install).
  • MetalEngine (tier 4) no longer imports NumPy: the scalar-divisor reciprocal is computed with Python
    floats, identical to the NumPy result on 800,046 checked values including zeros, infinities, NaN and
    subnormals. NumPy is not installed by the wheel, pyarrow or Polars, and the engine raised
    ModuleNotFoundError without it.
  • The polars extra is polars>=1.44,<1.45, the minor the packaged plugin's ABI matches (it was
    polars>=1.0). Tiers 1 and 3 run on any polars>=1.0; tier 4 is tested on 1.44.1 and 1.44.2
    (docs/POLARS.md, "Which Polars").
  • MetalEngine takes a Polars Scan of one local Parquet file: the file is read on the GPU with
    Polars' projection as the column list, the scan's predicate runs as a GPU filter over the read, and
    its comparisons that row-group and page statistics can judge the way Polars compares (integers;
    Strings ==/!=; floats <, <=, == only, so NaN rows under >, >=, != are never skipped)
    also go to the reader. Several files, hive partitions, URLs, row_index_name, n_rows,
    include_file_paths, schema=, unsupported dtypes, CSV and NDJSON scans stay with Polars, named in
    engine.last_report, whose entries now list each file read, its filters and the row groups and
    pages skipped. The defaults take a scan subtree on the same terms as an in-memory one. 50M rows:
    a sort of two columns is 4.4-7.6x the faster Polars engine through a freshly opened file and
    4.7-8.7x with the file open; Benchmarks/results/polars_engine_scan_2026-09-26-quiet.csv,
    polars_engine_bench.py --scan-only (docs/POLARS.md, "Parquet scans").
  • A plan with scan_ipc under MetalEngine collects on Polars instead of failing: polars 1.44.1
    raises NotImplementedError when an engine views that node, and the engine now leaves the node,
    and what is above it, to Polars with that reason.
  • Open-file cache for Parquet: read_parquet(..., cache=True) / read_parquet_table(..., cache=True)
    keep the opened, mapped file across reads, keyed by real path, inode, modification time and size
    (a changed or replaced file is opened afresh), LRU-bounded by 16 files and a quarter of physical
    memory (parquet_cache_limit), with parquet_cache_info and clear_parquet_cache. On the 50M-row,
    8-column files a one-column read is 6.9-12.2 ms through a fresh open and 3.2-7.6 ms through the
    cache (Benchmarks/results/parquet_cache_2026-09-26-quiet.csv, parquet_bench.py --cache;
    docs/PARQUET.md, "The open-file cache").
  • Parquet SNAPPY dictionary pages, and SNAPPY dispatches of at most 16 pages, are decompressed on the
    host by a bounds-checked decoder instead of one GPU SIMD group per page. A 1,000,000-row file with
    an int64 and a float64 column in pyarrow's defaults read in 102-106 ms before (94 ms of it one
    790 KB dictionary page) and 20-36 ms after; five columns of a 10,000,000-row, 7-column SNAPPY file
    with the file open, 221 ms before and 57-66 ms after (docs/PARQUET.md, "Decompression").
  • am_parquet_column_null_count / ParquetFile.column_null_count: a top-level column's null count from
    the footer (0 for a required column, else the sum of the row groups' statistics), without reading data.
  • Engine conformance grid: python/tests/engine_report.py runs the Polars engine against
    lf.collect() and the DuckDB rewrite against DuckDB with the rewrite off, over generated shapes x
    dtypes x null patterns x sizes (0 to 100,000 rows), compares bit for bit, classifies every
    difference against the documented divergences and exits 1 on an unclassified one; --csv-dir
    writes the summary and per-shape CSVs. Recorded run: Polars 12,597 cases, 12,392 pass, 32
    documented (float summation order), 0 unclassified, 173 not taken; DuckDB 33,376 cases, 21,844
    pass, 0 documented, 0 unclassified, 11,532 not taken
    (Benchmarks/results/engine_conformance_2026-09-25.csv, docs/COVERAGE.md "Engines").
    python/tests/test_engine_conformance.py runs the grid without its 100,000-row tables.
  • Polars engine, answers the grid found different from Polars' and now the same:
    a true division by a scalar and a float multiply by -1 over a column of one row, which Polars
    computes element-wise (the engine chooses by the input's row count, counting it when the plan runs
    if a filter or join decides it); min/max over a column or group holding both 0.0 and -0.0
    (Polars answers -0.0 and 0.0); a Float32 mean, which Polars accumulates in Float64 (docs/POLARS.md,
    Tier 4).
  • Polars engine: a sort by a nullable Date, Datetime, Duration or Time column with nulls first
    (Polars' default) runs on Metal; its validity key comes from the in-memory frame, where it used to
    be an expression ArrowMetal's compiler does not read, which left the sort to Polars.
  • Expression parser: a u64 literal above 2^63 - 1 is accepted (held as its bit pattern), so a
    Polars plan comparing a UInt64 column with such a value runs on Metal.
  • docs/DUCKDB.md states two shapes the extension leaves to DuckDB: an ungrouped query whose
    aggregates are all counts, and an input DuckDB has already replaced with an empty result.
  • Polars engine, every collect path explicit (docs/POLARS.md, "Collect paths"): collect,
    head(n).collect, a pl.Config engine affinity set to a MetalEngine, pl.collect_all,
    engine.profile(lf) and the new engine.explain(lf) (Polars' optimised plan followed by the
    placement report, with nothing run) run on Metal; collect_async, collect_all_async,
    collect_batches, collect(background=True) and every sink_* run the whole plan on Polars, say
    so in engine.last_report and issue a MetalEngineFallbackWarning once per process.
    last_report.path names the path. pl.collect_all(lfs, engine=MetalEngine()) collects each frame
    through the engine (Polars' own collect_all passes no engine callback), with one report per frame
    in last_reports. sink_batches through the engine used to fail with a ComputeError when
    cloudpickle is not installed (polars 1.44.1 cannot show that Sink node to an engine); it now runs on
    Polars. lf.explain(engine=), lf.profile(engine=) and eager DataFrame methods never call the
    engine, as the table states.
  • Polars engine capability table: docs/ENGINE_CAPABILITIES.md, generated by
    python/tests/engine_capabilities.py from 2,523 runs (29 plan shapes x 29 input dtypes x 3 null
    patterns) of MetalEngine(shapes="all", min_rows=0): 1,308 on Metal, 1,089 with Polars with the
    engine's reason, 126 that Polars itself rejects, none different from Polars' answer.
    test_the_capability_table_is_current fails when the committed file differs from a fresh run.
  • Polars engine version checks: TESTED_POLARS is a tuple, ("1.44.1", "1.44.2"); the full engine
    suite and the capability table are identical on both. KNOWN_NODE_KINDS holds the 20 IR node
    classes of those releases; a plan holding another kind, or a node Polars fails to show to the
    engine, stays with Polars whole with the report line The plan holds an unknown node <kind> in polars <version>, so the whole plan stays with Polars. instead of an error.
    python -m arrowmetal.polars_engine check prints the Polars version, the IR version, whether the
    callback API is present, the IR node kinds the engine does not know, and the capability table's
    header; it exits 1 when every plan would stay with Polars. The callback also accepts the
    one-argument call of polars 2.0.0rc2, whose IR version (15, 1) keeps every plan with Polars.
  • Polars engine messages: every translation reason in the report is a full sentence
    (test_every_unsupported_reason_is_a_full_sentence), and the engine's errors (raise_on_fail=True,
    a subtree failing on Metal, a result schema that is not Polars') name the node and the reason and
    end with how to run the plan on Polars instead. ARROWMETAL_METAL_ENGINE=off makes every
    MetalEngine leave every plan to Polars.
  • Strings in the view layout: the C Data import takes utf8_view / binary_view (vu / vz) and
    keeps the 16-byte views and the variadic data buffers as Metal shared buffers, without a copy at any
    alignment for the views and data buffers (MetalArrowBuffer.wrapCovering); export hands a view
    column back as vu / vz. The string kernels are written once against a row accessor and
    instantiated per layout: lengths, hash32, the four pattern predicates, str_eq (scalar and
    array), is_in / index_in, the gather behind filter / take, slice, the sort keys, the string
    hash table (group-by keys, dictionary_encode, unique, value_counts), the Unicode and ASCII case
    transforms, the trims, replace, repeat, slice_codeunits, the pads, str_reverse,
    count_substring / find_substring, the ascii_is_* / utf8_is_* predicates, utf8_center,
    utf8_replace_slice, the fused expression kernels, concatenation of view columns and the host-side
    row passes. The rest (str_concat, the splits, match_like, strptime, the byte-counting pads,
    utf8_zero_fill, the byte slices and reversals, casts from strings, the file writers) convert the
    column to offsets + bytes once, on the GPU, and keep it (docs/DESIGN.md, "Strings in two layouts").
    C: am_string_layout, am_string_view_info, am_string_view_conversions; Python:
    MetalArray.string_layout, MetalArray.string_view_import(), am.string_view_conversions(),
    am.STRING_VIEW_KERNELS, am.STRING_VIEW_CONVERTS; ARROWMETAL_TRACE_VIEW_CONVERSION=1 prints the
    call stack of each conversion. Swift: StringViewTests (8 tests).
  • Polars: from_polars, the .arrowmetal namespaces and MetalEngine hand String and Binary columns
    over in Polars' own Utf8View layout; from_polars(..., string_layout="offsets"),
    polars_bridge.DEFAULT_STRING_LAYOUT and polars_engine.STRING_LAYOUT select the large_string
    path. The tier-2 plugin is unchanged.
  • Differential matrix: every utf8 operation also runs with the values imported as utf8_view (the
    oracle stays on utf8): 40,824 cases over 46 column types, 0 unclassified; the 1,755 utf8_view
    cases give the same pass, fail and skip counts as their utf8 cells. engine_report.py takes
    --dtypes and --string-layout; the Polars grid's 459 String cases pass with 0 unclassified on
    both layouts, with no view column converted.
  • Group-by at 2^24 groups and more: the Float64 sum and mean and the Float32 sum and mean (in
    Float64) came back null for most groups once there were 2^24 groups or more, and product and list
    were wrong for those groups. These aggregates reduce one group per 256-thread threadgroup, and a grid
    dimension of 2^32 threads or more wraps on the GPU, so only groups mod 2^24 threadgroups ran. Every
    one-threadgroup-per-group kernel (also the segmented 64-bit min/max, the variance passes and the
    counting sort's run sort) now dispatches through Dispatch.perGroup, which folds the grid into rows
    of 65,536 threadgroups past that width; the 50M-row group-by timings are unchanged. Tests:
    GroupByGridFoldTests and python/tests/test_group_by_2_24.py, at 2^24 - 1, 2^24, 2^24 + 1 and
    2^24 + 2^20 groups (docs/FINDINGS.md, round 13).