Everything below is new in 0.3.0. MetalEngine() decides per subtree from measured crossovers: in the
default benchmark (194 case-size pairs at 2,000,000 and 50,000,000 rows) it took 62 pairs, 42 of them
group-bys, every one ahead of the faster Polars engine, 1.21x to 9.42x
(Benchmarks/results/polars_engine_bench_2026-09-26-final3.csv). The engine takes a Polars
scan_parquet of one local file and reads it on the GPU. The C Data import and export take utf8_view
/ binary_view, the string kernels run on that layout, and Polars' String columns are handed over in
it. A whole-file Parquet read of the 50,000,000-row, 8-column files through a fresh open is 65 ms
(Snappy), 54 ms (LZ4) and 38 ms (uncompressed), against Polars' 95, 73 and 63 ms in the same run, with
Snappy and LZ4 pages decompressed by the host and the GPU at the same time
(Benchmarks/results/parquet_bench_2026-09-26-split.txt). python -m arrowmetal.router calibrate fits
the router's table on the Mac it runs on, and explain prints a decision with its reason and the
measured points behind it. The conformance grids compare the Polars engine with Polars (12,597 cases)
and the DuckDB rewrite with DuckDB (33,376 cases) bit for bit, with 0 unclassified differences. A
group-by over 2^24 groups or more returns the right Float64 and Float32 sums and means, products and
lists. The wheel carries the Polars expression plugin, so all four Polars tiers run from
pip install.
- Parquet Snappy and LZ4 pages are decompressed by the host and the GPU at the same time, split page by
page. A read's router (DecodeRouter) orders every Snappy and LZ4 page of the columns it stages by how
token-dense its header says it is (its ratio: at or below 1.0 a literal, above 1.05 token-dense), and
gives the host the densest ones until the host's share, spread over its cores, is predicted to finish
with the GPU's, from measured per-byte costs on each side: a 160 KB token-dense page takes one CPU core
41-81 µs and the GPU 12-15 ms on a SIMD group or 37-60 ms on a thread, while 1.2 GB of literal pages
take the GPU 10 ms and 16 host threads 19 ms from a fresh mapping (docs/PARQUET.md, "Decompression").
ZSTD, GZIP and BROTLI stay host-only in the same schedule. The GPU dispatch is committed first (slowest
pages first), the host decodes its pages on every core while it runs, and the read waits once before
the value kernels; the host's pages sit in their own page-aligned range of the staging buffer. A new
bounds-checkedLZ4Hostdecoder makes the kernels' checks with their outcomes;SnappyHostcopies 8
and 16 bytes at a time inside the page and the slot. This replaces the 16-page host rule for Snappy and
the 2,048-page rule for the page-per-thread kernel. On the 50,000,000-row, 8-column files a whole-file
read through a fresh open is 65 ms (Snappy) and 54 ms (LZ4), against 108 and 96 ms before and against
Polars' 95 and 73 ms and pyarrow's 160 and 159 ms in the same run; it takes 674 and 508 CPU-ms, against
80 and 78 before and Polars' 1,283 and 1,007 (Benchmarks/results/parquet_bench_2026-09-26-split.txt,
parquet_cache_2026-09-26-split.csv; run conditions inbench_conditions_2026-09-26-split.txt). A
10,000,000-row, 7-column pyarrow-default Snappy file reads in 25.4 ms against 67.5 ms.
ARROWMETAL_PARQUET_DECODE=host|gpu|lanesends every Snappy and LZ4 page to one decoder. Tests: every
fixture on every decoder against the split and against pyarrow, the 240 damaged files on each decoder,
the host decoders against byte-at-a-time references and under damage between guard pages, the router
and the staging layout (DecompressSplitTests,ParquetTests,test_parquet.py).
Benchmarks/parquet_bench.pylabels ArrowMetal's rowarrowmetal. MetalEngine()'s default (shapes="measured") decides per subtree from measured crossovers instead
of taking large sorts only. Each translated subtree has shape classes (rowwise;aggregate,
group_byandgroup_by_multiper aggregate familysum,count,mean,minmax;sort,
sort_helper_keys,top_k;join:inner|left|semi|anti;distinct), a dtype class (stringwhen a
String column is among its inputs) and an input (in-memory frames, or a Parquet file judged by its
footer's row count). It runs on Metal when its input rows are at or above the crossover of every
class in it: the largest of the engine table (python/arrowmetal/_engine_crossovers.py, fitted by
the newBenchmarks/polars_engine_crossover.pyfromBenchmarks/results/polars_engine_crossover_2026-09-26-final.csv,
93 in-memory cases at eight sizes from 250,000 to 50,000,000 rows and the Parquet cases at six sizes
from 1,000,000 to 50,000,000, best of 7, run conditions inbench_conditions_2026-09-26-final.txt),
the router table in force for the kernels it routes, and the sort kernels' crossover against the
fastest CPU library. Three rules sit on the fit, each kept in the table: a case counts as ahead when
its time x 1.15, or x 1.35 for a shape with a String column, is at most the faster Polars engine's
(MARGIN; a String shape's advantage grows slowly with size, the String sort being 0.98x at 1M rows
and 1.37x at 5M); a shape with a String column is taken from 5,000,000 rows at the earliest
(STRING_FLOOR; below it the String sorts are 1.2x and 1.38x in the sweep, inside the benchmark's
noise band); and the default takes a shape from 1.5 times its fitted crossover (HEADROOM; the fit
interpolates between sizes measured 2-2.5x apart, and shapes just past it were within run-to-run
noise of Polars). The table keeps the fit asfitand what the default uses asrows. Taken:
sorts from 1,026,501 rows and helper-key sorts from 1,000,000 (the sort kernels), numeric-key left
joins from 1,875,000 input rows, inner joins from 2,029,827 and anti joins from 3,727,959,unique
from 5,494,090, sorts of a Parquet file from 1,500,000, helper-key sorts anduniquewith a String
column from 5,000,000 and sorts with one from 6,301,531; group-bys by their number of groups (next
entry). Whole-frame aggregates, top-k, semi joins, row-wise shapes and the other String shapes are
not taken.shapes="all",min_rows=and a new explicit set of class names
(shapes={"sort", "join"}) override it. Each taken subtree's report entry carriesrule,shape,
dtype_classandinput, and each node the policy leaves has aKind#id: rule: ...line ("900,000
input rows is below the 2,029,827-row crossover for join:inner (...)").
arrowmetal.polars_engine.placement_rules()lists the table;SHAPE_CLASSESreplaces
MEASURED_SHAPES, andMetalEngine().min_rowsisNoneunless given.
Benchmarks/polars_engine_bench.pygains 85 cases and--crossover;
python/tests/test_engine_policy.pytests the policy.MetalEngine()judges a group-by by its number of groups. The crossover sweep has a group-by grid
(each aggregate family over one int32 key and over two, keys drawn from 200, 1,000, 10,000, 100,000
and 1,000,000 values and from half the rows) and records each group-by's group count; every
group-by class is fitted per bucket of group counts (200, 1,000, 10,000, 100,000, 1,000,000, and at
least a quarter of the input rows), each bucket from the worst of its cases at each size
(_engine_crossovers.GROUPS,arrowmetal.polars_engine.group_placement_rules()), with the margin
and headroom above. The bucket of at least a quarter of the rows is never taken (UNTAKEN_BUCKETS):
the sweep is not monotone there, the one-key count over 0.43 times as many groups as rows running
0.78x, 1.9x, 4.04x, 2.21x and 1.18x the faster Polars engine from 2M to 50M rows. At plan time the
engine estimates the group count of a group-by whose keys are columns of one in-memory input frame:
the distinct key tuples of a fixed-seed stratified sample, scaled by the bias-corrected Chao1
estimator, from 512 sampled rows and four times more until every count in the estimate's range gets
the same decision (at most 65,536 rows); the samples are cached per frame, key columns and size
(clear_group_estimates()), so the same frame gets the same estimate and decision on every collect
and in every process. A Parquet file's footer distinct counts are read when every row group states
them. Taken now: over 10,000 and 100,000 groups every numeric group-by class, from 1,082,526 to
16,235,764 rows by class and bucket; over 1,000,000 groups every class but the one-key mean, from
7,500,000 (7,885,821 for the two-key mean); over two keys the count and min/max also at 200 groups
and every family at 1,000; the (String, int32) sum over about 1,000,000 groups from 9,744,372. Not
taken: one-key group-bys over 200 or 1,000 groups, and every group-by over a quarter of the rows or
more; a group-by with no estimate (a computed key, two group-bys in one subtree) is judged by its
class row, which takes the two-key count from 26,547,327 rows, the (String, int32) sum from
9,744,372 and no other numeric group-by. The report names the estimate in the rule ("estimated 191
groups over (region), a Chao1 estimate from a 512-row sample, 189 to 196: below the measured band
for group_by:sum at 50,000,000 input rows (taken at 3,163 to 3,162,277 groups; ...)") and lists each
probe with its time (last_report.groups). InBenchmarks/results/polars_engine_bench_2026-09-26-final3.csv
(the 93 in-memory cases at 2M and 50M rows and the 50M Parquet cases, 194 case-size pairs; run
conditions inbench_conditions_2026-09-26-final3.txt) the default took 62 pairs, 42 of them
group-bys, every one ahead of the faster Polars engine, 1.21x to 9.42x; in the 132 pairs it left
to Polars its time was 0.62x to 1.31x Polars' in-memory time where that is 5 ms or more.
The probe takes 66 to 461 µs at 50,000,000 rows (0.10% to 1.19% of the group-by it decides) and 50
to 393 µs at 2,000,000 (1.09% to 4.90%), median of 9 cold collects under the earlier group-count
table (Benchmarks/results/group_probe_2026-09-26.csv, from the newBenchmarks/group_probe_bench.py);
in the default benchmark's process, which holds every case's frames, its median is 99 µs at 2M and
144 µs at 50M, and it reaches 13.5 ms on two-key frames with 0.43 times as many groups as rows at 50M.- The crossover sweep's
(v3)case found ArrowMetal's group-by returning wrong Float64 sums and means from 16,777,216 groups (most groups null); fixed in the core (below), so the engine needs no guard for it. - Parquet reads, cold and warm. On the 50,000,000-row, 8-column benchmark files a whole-file read
through a fresh open is 108 ms (Snappy), 96 ms (LZ4) and 41 ms (uncompressed), against 369, 340 and
250 ms inparquet_bench_2026-09-25-quiet.txtand against Polars' 99, 74 and 67 ms in the same run,
with 78-91 ms of process CPU against 760-1330 for pyarrow and Polars; time to first compute is
9-13 ms (Polars 16-22); a one-column read through a fresh open is 6.9-12.2 ms, against 149-191 ms
before (Benchmarks/results/parquet_bench_2026-09-26-quiet.txt,parquet_cache_2026-09-26-quiet.csv;
run conditions inbench_conditions_2026-09-26-quiet.txt). What changed
(docs/PARQUET.md, "The file's bytes are the GPU's bytes" and "Decompression"):- the file is mapped read-only and shared: the first kernel to read a wrapped range makes it resident
for the GPU at 7-10 ms per GB, against 67-78 ms per GB for the private writable mapping, which stays
as the fallback for a device that will not wrap read-only memory and on virtualised GPUs; - wrapping the reader's own mapping skips
MetalArrowBuffer.wrapOrCopy'smincoreprobe, which took
about 35 ms over 2 GB; - a projection whose chunks fill less than four fifths of the bytes they span maps them side by side
(MappedView) and wraps only them; a column whose chunks span 4 GiB or more of a larger file now
reads (the limit is 4 GiB of one column's chunks in one read); - page headers are parsed in parallel across the read's chunks and kept on the handle;
- dense fixed-width values and all-dictionary code buffers are no longer zero-filled before kernels
that write every slot (the padding still is), and dictionary codes get their row group's base as they
are decoded instead of in a second pass; - ZSTD pages decode in runs sharing one
ZSTD_DCtxinstead of a fresh context per page; - token-dense Snappy and LZ4 pages (output at least 1.25x the input, in dispatches of at least 2,048
such pages) decode one page per thread, and a read decompresses the pages of all its flat columns in
one dispatch per codec before decoding them; - a closed file is unmapped on a background queue.
ARROWMETAL_PARQUET_PROFILE=1prints per-phase times and minor faults of each open and read.
ParquetTests adds reads of every fixture through both mappings, one-column and row-group reads through
column views, both decompression kernels on every Snappy and LZ4 fixture, and 240 damaged files through
the page-per-thread kernel;test_parquet.pyadds 2 to 400 pages a chunk in ZSTD, LZ4 and Snappy.
- the file is mapped read-only and shared: the first kernel to read a wrapped range makes it resident
python -m arrowmetal.bench: one seeded 10,000,000-row dataset (drawn withpyarrow.computefrom
SplitMix64 streams, so nothing beyondpip install arrowmetalis needed), sum, filter, sort and group-by sum
through pyarrow (and Polars when installed) and through ArrowMetal, each answer checked against
pyarrow, one table for the Mac it runs on with a ready-to-paste block;--rows,--json,--quiet,
--no-share,--no-polars. A GitHub issue form (.github/ISSUE_TEMPLATE/benchmark_result.yml)
takes that block (docs/TESTING.md,python/tests/test_bench.py).- Router table per machine:
python -m arrowmetal.router calibrate [--quick] [--out PATH] [--csv PATH]
runs the router check sweep on the Mac it is run on (the sweep moved fromBenchmarks/router_check.py
into the package, which that script now calls), fits it with the codeBenchmarks/router_table.py
uses (python/arrowmetal/_router_fit.py) and writes~/.arrowmetal/router/<chip>.jsonwith the chip,
core counts, Metal device, date, ArrowMetal version, grid and every measurement. At first use the
router takesARROWMETAL_ROUTER_TABLE(a path, orshipped), else that file for this chip, else the
shipped table; an operation the sweep did not bring to a crossover keeps its shipped row.
python -m arrowmetal.router explain <op> <rows> [--dtype] [--nulls] [--keys] [--json]prints the
decision, its reason, the table row and the two measured points its crossover was fitted between, and
whether the table is the shipped one or this machine's.router_table.py --json-outwrites the shipped
table in the same JSON format, andpython -m arrowmetal.bench --calibrateruns the quick calibration
after the benchmark. C:am_router_load_table,am_router_table_info,am_router_decide,
am_router_explain; Swift:Router.table,Router.loadTable(path:),Router.useShippedTable(),
Router.explain,RouterCrossovers; Python:am.router_table(),am.load_router_table(),
am.route_decision(),am.explain_route().am_router_crossover,Router.crossoverRowsand
am.router_crossovers()report the table in force. The Swift and Python test harnesses pin the
shipped table unlessARROWMETAL_ROUTER_TABLEis set (docs/CROSSOVER.md). - Router determinism:
Router.route, the rule every routed call goes through, is a pure function of the
operation, value type, row count, mode, batch state and table row; nothing is timed at call time.
python/tests/test_router_calibrate.pychecks 200 (operation, size) pairs for the same decision
across 1,000 calls and across two processes, andRouterTests.testRouteIsPurechecks the Swift rule. python -m arrowmetal.bench --parquet FILE: the same report on the user's own Parquet file. It reads
the file's integer, floating-point and string columns withpyarrow.parquet.read_table,
polars.read_parquetandam.read_parquet, then runs sum andfilter > medianon the largest numeric
column and group-by sum keyed on the lowest-cardinality integer or string column, CPU against Metal,
each ArrowMetal answer checked against pyarrow's. The Share it block and the prefilled issue link
(title and the form'ssharefield) carry the row count, column count, row groups, size, codecs and
the timings, never the path, column names or values. A file whose columns would take more than a
quarter of physical memory is refused with the limit printed; a file with no usable column gets one
line saying so. On a generated 10,000,000-row Snappy file (M4 Max): read 134.58 ms against Polars'
16.75 ms, filter 1.12 ms against 4.50 ms, sum 1.40 ms against 0.99 ms, group-by sum over a 5-value
string key 18.25 ms against 6.82 ms (python/tests/test_bench.py).- The wheel carries the Polars expression plugin (tier 2):
python/build_wheel.shbuilds
polars-plugin/with cargo against thelibArrowMetalC.dylibit bundles, and packages
arrowmetal/_lib/libarrowmetal_polars.dylibwith its rpath set to@loader_pathand local symbols
stripped.arrowmetal.polars_plugin.plugin_path()takes the packaged plugin first when Python loaded
the packagedlibArrowMetalC.dylib, and a cargo build first when$ARROWMETAL_LIBor a development
build is loaded.scripts/check_wheel.shinstalls a wheel with itspolarsextra into a fresh
virtualenv with a scrubbed environment and no cargo, runs all four Polars tiers,benchand
bench --parquet, and checks that onelibArrowMetalC.dylib, the packaged one, is loaded; it passed
with polars 1.44.2 and pyarrow 25.0.1, NumPy not installed. The wheel is 8.2 MB (3.2 MB without the plugin) (docs/POLARS.md, Install). MetalEngine(tier 4) no longer imports NumPy: the scalar-divisor reciprocal is computed with Python
floats, identical to the NumPy result on 800,046 checked values including zeros, infinities, NaN and
subnormals. NumPy is not installed by the wheel, pyarrow or Polars, and the engine raised
ModuleNotFoundErrorwithout it.- The
polarsextra ispolars>=1.44,<1.45, the minor the packaged plugin's ABI matches (it was
polars>=1.0). Tiers 1 and 3 run on anypolars>=1.0; tier 4 is tested on 1.44.1 and 1.44.2
(docs/POLARS.md, "Which Polars"). MetalEnginetakes a PolarsScanof one local Parquet file: the file is read on the GPU with
Polars' projection as the column list, the scan's predicate runs as a GPU filter over the read, and
its comparisons that row-group and page statistics can judge the way Polars compares (integers;
Strings==/!=; floats<,<=,==only, so NaN rows under>,>=,!=are never skipped)
also go to the reader. Several files, hive partitions, URLs,row_index_name,n_rows,
include_file_paths,schema=, unsupported dtypes, CSV and NDJSON scans stay with Polars, named in
engine.last_report, whose entries now list each file read, its filters and the row groups and
pages skipped. The defaults take a scan subtree on the same terms as an in-memory one. 50M rows:
a sort of two columns is 4.4-7.6x the faster Polars engine through a freshly opened file and
4.7-8.7x with the file open;Benchmarks/results/polars_engine_scan_2026-09-26-quiet.csv,
polars_engine_bench.py --scan-only(docs/POLARS.md, "Parquet scans").- A plan with
scan_ipcunderMetalEnginecollects on Polars instead of failing: polars 1.44.1
raisesNotImplementedErrorwhen an engine views that node, and the engine now leaves the node,
and what is above it, to Polars with that reason. - Open-file cache for Parquet:
read_parquet(..., cache=True)/read_parquet_table(..., cache=True)
keep the opened, mapped file across reads, keyed by real path, inode, modification time and size
(a changed or replaced file is opened afresh), LRU-bounded by 16 files and a quarter of physical
memory (parquet_cache_limit), withparquet_cache_infoandclear_parquet_cache. On the 50M-row,
8-column files a one-column read is 6.9-12.2 ms through a fresh open and 3.2-7.6 ms through the
cache (Benchmarks/results/parquet_cache_2026-09-26-quiet.csv,parquet_bench.py --cache;
docs/PARQUET.md, "The open-file cache"). - Parquet SNAPPY dictionary pages, and SNAPPY dispatches of at most 16 pages, are decompressed on the
host by a bounds-checked decoder instead of one GPU SIMD group per page. A 1,000,000-row file with
anint64and afloat64column in pyarrow's defaults read in 102-106 ms before (94 ms of it one
790 KB dictionary page) and 20-36 ms after; five columns of a 10,000,000-row, 7-column SNAPPY file
with the file open, 221 ms before and 57-66 ms after (docs/PARQUET.md, "Decompression"). am_parquet_column_null_count/ParquetFile.column_null_count: a top-level column's null count from
the footer (0 for a required column, else the sum of the row groups' statistics), without reading data.- Engine conformance grid:
python/tests/engine_report.pyruns the Polars engine against
lf.collect()and the DuckDB rewrite against DuckDB with the rewrite off, over generated shapes x
dtypes x null patterns x sizes (0 to 100,000 rows), compares bit for bit, classifies every
difference against the documented divergences and exits 1 on an unclassified one;--csv-dir
writes the summary and per-shape CSVs. Recorded run: Polars 12,597 cases, 12,392 pass, 32
documented (float summation order), 0 unclassified, 173 not taken; DuckDB 33,376 cases, 21,844
pass, 0 documented, 0 unclassified, 11,532 not taken
(Benchmarks/results/engine_conformance_2026-09-25.csv, docs/COVERAGE.md "Engines").
python/tests/test_engine_conformance.pyruns the grid without its 100,000-row tables. - Polars engine, answers the grid found different from Polars' and now the same:
a true division by a scalar and a float multiply by -1 over a column of one row, which Polars
computes element-wise (the engine chooses by the input's row count, counting it when the plan runs
if a filter or join decides it);min/maxover a column or group holding both 0.0 and -0.0
(Polars answers -0.0 and 0.0); a Float32mean, which Polars accumulates in Float64 (docs/POLARS.md,
Tier 4). - Polars engine: a sort by a nullable Date, Datetime, Duration or Time column with nulls first
(Polars' default) runs on Metal; its validity key comes from the in-memory frame, where it used to
be an expression ArrowMetal's compiler does not read, which left the sort to Polars. - Expression parser: a
u64literal above 2^63 - 1 is accepted (held as its bit pattern), so a
Polars plan comparing a UInt64 column with such a value runs on Metal. - docs/DUCKDB.md states two shapes the extension leaves to DuckDB: an ungrouped query whose
aggregates are all counts, and an input DuckDB has already replaced with an empty result. - Polars engine, every collect path explicit (docs/POLARS.md, "Collect paths"):
collect,
head(n).collect, apl.Configengine affinity set to aMetalEngine,pl.collect_all,
engine.profile(lf)and the newengine.explain(lf)(Polars' optimised plan followed by the
placement report, with nothing run) run on Metal;collect_async,collect_all_async,
collect_batches,collect(background=True)and everysink_*run the whole plan on Polars, say
so inengine.last_reportand issue aMetalEngineFallbackWarningonce per process.
last_report.pathnames the path.pl.collect_all(lfs, engine=MetalEngine())collects each frame
through the engine (Polars' owncollect_allpasses no engine callback), with one report per frame
inlast_reports.sink_batchesthrough the engine used to fail with aComputeErrorwhen
cloudpickle is not installed (polars 1.44.1 cannot show that Sink node to an engine); it now runs on
Polars.lf.explain(engine=),lf.profile(engine=)and eagerDataFramemethods never call the
engine, as the table states. - Polars engine capability table: docs/ENGINE_CAPABILITIES.md, generated by
python/tests/engine_capabilities.pyfrom 2,523 runs (29 plan shapes x 29 input dtypes x 3 null
patterns) ofMetalEngine(shapes="all", min_rows=0): 1,308 on Metal, 1,089 with Polars with the
engine's reason, 126 that Polars itself rejects, none different from Polars' answer.
test_the_capability_table_is_currentfails when the committed file differs from a fresh run. - Polars engine version checks:
TESTED_POLARSis a tuple,("1.44.1", "1.44.2"); the full engine
suite and the capability table are identical on both.KNOWN_NODE_KINDSholds the 20 IR node
classes of those releases; a plan holding another kind, or a node Polars fails to show to the
engine, stays with Polars whole with the report lineThe plan holds an unknown node <kind> in polars <version>, so the whole plan stays with Polars.instead of an error.
python -m arrowmetal.polars_engine checkprints the Polars version, the IR version, whether the
callback API is present, the IR node kinds the engine does not know, and the capability table's
header; it exits 1 when every plan would stay with Polars. The callback also accepts the
one-argument call of polars 2.0.0rc2, whose IR version (15, 1) keeps every plan with Polars. - Polars engine messages: every translation reason in the report is a full sentence
(test_every_unsupported_reason_is_a_full_sentence), and the engine's errors (raise_on_fail=True,
a subtree failing on Metal, a result schema that is not Polars') name the node and the reason and
end with how to run the plan on Polars instead.ARROWMETAL_METAL_ENGINE=offmakes every
MetalEngineleave every plan to Polars. - Strings in the view layout: the C Data import takes
utf8_view/binary_view(vu/vz) and
keeps the 16-byte views and the variadic data buffers as Metal shared buffers, without a copy at any
alignment for the views and data buffers (MetalArrowBuffer.wrapCovering); export hands a view
column back asvu/vz. The string kernels are written once against a row accessor and
instantiated per layout: lengths,hash32, the four pattern predicates,str_eq(scalar and
array),is_in/index_in, the gather behindfilter/take,slice, the sort keys, the string
hash table (group-by keys,dictionary_encode,unique,value_counts), the Unicode and ASCII case
transforms, the trims,replace,repeat,slice_codeunits, the pads,str_reverse,
count_substring/find_substring, theascii_is_*/utf8_is_*predicates,utf8_center,
utf8_replace_slice, the fused expression kernels, concatenation of view columns and the host-side
row passes. The rest (str_concat, the splits,match_like,strptime, the byte-counting pads,
utf8_zero_fill, the byte slices and reversals, casts from strings, the file writers) convert the
column to offsets + bytes once, on the GPU, and keep it (docs/DESIGN.md, "Strings in two layouts").
C:am_string_layout,am_string_view_info,am_string_view_conversions; Python:
MetalArray.string_layout,MetalArray.string_view_import(),am.string_view_conversions(),
am.STRING_VIEW_KERNELS,am.STRING_VIEW_CONVERTS;ARROWMETAL_TRACE_VIEW_CONVERSION=1prints the
call stack of each conversion. Swift:StringViewTests(8 tests). - Polars:
from_polars, the.arrowmetalnamespaces andMetalEnginehand String and Binary columns
over in Polars' ownUtf8Viewlayout;from_polars(..., string_layout="offsets"),
polars_bridge.DEFAULT_STRING_LAYOUTandpolars_engine.STRING_LAYOUTselect thelarge_string
path. The tier-2 plugin is unchanged. - Differential matrix: every utf8 operation also runs with the values imported as
utf8_view(the
oracle stays on utf8): 40,824 cases over 46 column types, 0 unclassified; the 1,755utf8_view
cases give the same pass, fail and skip counts as their utf8 cells.engine_report.pytakes
--dtypesand--string-layout; the Polars grid's 459 String cases pass with 0 unclassified on
both layouts, with no view column converted. - Group-by at 2^24 groups and more: the Float64 sum and mean and the Float32 sum and mean (in
Float64) came back null for most groups once there were 2^24 groups or more, and product and list
were wrong for those groups. These aggregates reduce one group per 256-thread threadgroup, and a grid
dimension of 2^32 threads or more wraps on the GPU, so onlygroups mod 2^24threadgroups ran. Every
one-threadgroup-per-group kernel (also the segmented 64-bit min/max, the variance passes and the
counting sort's run sort) now dispatches throughDispatch.perGroup, which folds the grid into rows
of 65,536 threadgroups past that width; the 50M-row group-by timings are unchanged. Tests:
GroupByGridFoldTestsandpython/tests/test_group_by_2_24.py, at 2^24 - 1, 2^24, 2^24 + 1 and
2^24 + 2^20 groups (docs/FINDINGS.md, round 13).