Everything below is in 0.1.0, the first public release.
Core
- Metal shared-memory Arrow buffers (page aligned, pooled) and primitive/boolean arrays.
- Kernels: sum/min/max/mean, compare, add/sub/mul/div (vectorised, defined integer division by zero),
filter (one command buffer, GPU scan) and fused filter(where:), take, cast, slice, boolean and/or/not/count/any/all. - Float64 compare/min/max/filter/take/slice on the GPU via order-preserving bit patterns; NaN semantics.
- GroupBy over dense integer keys: count, sum, mean, min, max (privatised and device-atomic paths), plus a
sort-based segmented path with no atomics — sumDouble/meanDouble/sumFloatAsDouble/meanFloat and
min64/max64 — which covers the Float64 and 64-bit min/max cases 32-bit atomics cannot express. - The key-value-to-dense-id mapping reads an integer key column as the type it is rather than widening it
to int64 first, and its three dispatches share one command buffer;hash_meanreuses the counts the
accumulation already produced and divides on the GPU, andhash_sumreturns the accumulator's own
buffer with a GPU-built validity bitmap instead of a host loop. Same ids, same answers to the bit;
sumby int32 key at 50M rows and a thousand groups went 8.75 -> 4.81 ms andmean9.78 -> 4.80 in
the A/B run recorded in docs/TO_IMPROVE.md; the published matrix reads 4.89 and 4.91 ms for those cells. - MetalRecordBatch with filter/take/slice/selecting; struct (+s) C Data import/export; ArrowArrayStream import.
- Batched execution (
MetalContext.batch { }) and its non-blocking form:batchAsync(Swiftasync
and completion-handler), withMetalArray.sumAsync/meanAsyncfor scalars, so the calling thread is
free while the GPU works. - Arrow C Data Interface and C Device Data Interface (ARROW_DEVICE_METAL) import and export.
Strings and sorting
MetalStringArray(utf8): byte/char length, equals/starts_with/ends_with/contains, MurmurHash3, GPU filter/take,
GPU dictionary encoding (hash, argsort, byte-comparing boundaries, rank scan and gather; collisions detected
and re-hashed, with the host path as the final fallback); import/export through the C Data Interface
(large_utf8 narrowed on import).- GPU LSD radix sort:
argsort,sorted,MetalRecordBatch.sorted(by:); stable, nulls last by default
(null_placementputs them at either end), IEEE total order.
The nulls are partitioned out before the sort rather than lifted out of the permutation afterwards, and
sorted()rebuilds the values from the sort's own keys instead of gathering them through it. topKfor any k: a GPU radix select (digit histogram over the order-preserving key, one compaction pass
keeping only the rows that can still win, a refinement round, then a bitonic or radix ordering) reads the
column twice whatever k is, and matches the full sort index for index. The per-threadgroup selection
(threshold plus a bitonic compaction in threadgroup memory) still serves small inputs at k <= 1024, and
the sort covers k near n and the case where fewer than k rows are non-null.kthElement(k, largest:): the exact k-th smallest or largest value, by the same selection with the winners
counted but never written.quantileandapproximateMedianrun on it instead of a full sort.
Temporal, binary and dictionary types
MetalTemporalArray: date32/date64, time32/time64, timestamp (unit + optional timezone) and duration,
forwarding compare/filter/take/slice/min/max/sort/argsort to the int32 or int64 array underneath.- GPU calendar fields in UTC (
year,month,day,dayOfWeek,hour,minute,second), plus
toDate32()andcastUnit(to:); civil-from-days arithmetic, correct for negative epochs. binaryandlarge_binaryshare the utf8 layout (MetalStringArray.isBinary, exported as "z").- Dictionary-encoded arrays import as int32 codes plus a value array; selection runs on the codes,
decode()materialises withtake, and export writesschema.dictionary/array.dictionary. - C ABI:
am_temporal_extract,am_temporal_cast_unit,am_dictionary_decode; Pythonyear()...
second(),cast_unit(),decode()and temporal/binary/dictionary types onMetalArray.type.
Execution model
MetalContext.batch { }: one command buffer per chain; kernels read lengths from device buffers so pending
filter results chain without a CPU sync; pool parks buffers while a batch is open.- Software IEEE-754 Float64 on the GPU:
add/sub/mul/divare correctly rounded and bit-exact
against Swift'sDouble(DoubleMathTests);sumuses the same adder but reassociates over
threadgroups, so it is checked to a tolerance rather than bit for bit.
Numerics
- Checked arithmetic:
add_checked…power_checked,negate/abs/sqrt/ln/log10/log2/log1p,
the checked shifts,logb_checked, and checkedcumulative_sum/cumulative_prod/pairwise_diff.
Each is the unchecked kernel plus a read-only check pass in the same command buffer, so the values are
bit-identical and the cost is one GPU round trip; a failure names the Arrow message and the first row. - Trigonometry: the twelve trigonometric and hyperbolic functions, their seven
_checkedtwins and
atan2— Metal's library functions for float32 (hyperbolics rewritten from well-conditioned
identities), a software binary64 implementation for float64. Measured within 5 ulp of the host libm
over 1,000,003 random arguments per function (float64 worst 5,tan; float32 worst 4 —TrigTests).
Above |x| ≈ 5e13 the argument reduction degrades in step with the argument's own ulp (7 ulp measured
at 5e13) and |x| ≥ 2^62 returns NaN — the one documented difference from libm. - The remaining element-wise math:
expm1,log1p,logb,hypot,round_to_multiple,round_binary
androundwith all ten Arrow round modes and anyndigits. - Float64
sqrt/exp/ln/log2/log10/power(and their_checkedtwins) in real software
binary64, not afloatevaluation widened back:sqrtis correctly rounded — bit-identical to
Foundation over 10^6 random bit patterns plus the extremes — and the rest measure 1 ulp over 10^6
inputs each across the whole domain, against a 2-ulp bound the tests assert.
poweron a float64 column is new; it used to raise. The trade is throughput, stated in
docs/BENCHMARKS.md: a binary64 logarithm costs about 25x what the seven-digit one did. - Software binary64 arithmetic made faster without losing a bit:
clznormalisation instead of shift
loops, four 32x32 partial products instead of an emulated 64x64, and a Newton reciprocal with an exact
128-bit remainder correction instead of a 57-step restoring division. Float64addandmultiply
now run at about 390 GB/s at 50M rows, the fastest these single-pass float64 rows reach, anddividewithin 15% of it
(331 GB/s) — docs/DESIGN.md. - Float classification (
is_nan,is_finite,is_inf) as raw bit-pattern tests, and the boolean
operators the bitmap family lacked:xor,and_not,and_not_kleene. - Statistical aggregates
skew,kurtosisandtdigest(GPU sort, host centroid merge), with grouped
forms of all three.
Grouped aggregation and windows
GroupByKeys: group-by over arbitrary key columns — sparse or negative integers, floats (with
-0.0 == 0.0and one NaN group), booleans, temporal values, utf8, binary, dictionary and decimal, and
several columns folded pairwise into an injective composite. A range path (mark, scan, rank) for narrow
integer-like columns and the dictionary-encode sort path for everything else.- All 24 Arrow
hash_*names against those ids: the fusedhash_min_max,hash_count_all,
hash_first/hash_last/hash_first_last,hash_one,hash_list,hash_distinct,
hash_approximate_medianandhash_quantile(exact, from a segmented sort),hash_product,
hash_variance/hash_stddev,hash_skew/hash_kurtosis,hash_tdigestandhash_pivot_wider. hash_count_distinctis a GPU hash set over the (key, value) pair — one insert pass over the rows,
then a histogram over the group ids of the occupied slots — instead of a dictionary encoding, a packed
int64 column and aunique()over it. 7.1 ms at 10M rows and 41.8 ms at 50M for a thousand groups,
and 7.3 ms at 10M for a hundred thousand — 60.7 ms at 50M rows and ten million groups is the worst
case (docs/TO_IMPROVE.md).- Several integer key columns whose ranges multiply out to at most 2^24 (and at most the row count) are
packed into one key in a single pass instead of folded pairwise through a range encoding each. The
dense ids are the fold's own, value for value, so the group order is unchanged. - Window and ordering functions:
rank,dense_rank,row_number,rank_quantile,rank_normal,
winsorize,shift, rollingsum/mean/min/max,cumulative_prod/cumulative_mean,
pairwise_diff, multi-columnlexsort_indices,partition_nth_indices,inverse_permutation,
scatter,unique,value_counts,count_all,true_unless_null,first_lastandrandom
(Philox4x32-10, seed-only stream).
Strings and Unicode
- The
utf8_is_*predicate family plusascii_is_printable,ascii_is_titleandstring_is_ascii. Each
answers every row on the GPU and re-decides on the host only the rows carrying a byte >= 0x80, so an
all-ASCII column never leaves the device. - Case and layout transforms:
ascii_title,utf8_capitalize,utf8_title,utf8_swapcase,
utf8_center,utf8_zero_fill,utf8_replace_slice,binary_replace_slice, theutf8_trim*family
against the full Unicode whitespace class, andutf8_normalize(NFC/NFKC/NFD/NFKD, host-side). extract_regex_span,binary_joinoverlist<utf8>, and string/binary value sets foris_inand
index_inthrough the GPU string hash table.
Temporal, timezones and the rest of the type matrix
weekwith all itsWeekOptions,day_of_weekwithDayOfWeekOptions,us_week,us_year,
iso_calendar,year_month_day,subsecond, and every*_betweendifference fromyears_between
down tonanoseconds_between.- The three interval layouts (
tiM,tiD,tin) withadd_intervaland the three interval-valued
differences;assume_timezone,local_timestampandis_dston the host, where the IANA database is. null,float16,decimal32/decimal64(GPU widening to decimal128 and narrowing back),
fixed_size_binary,list_view/large_list_viewimport,mapwithmap_lookup,list_slice,
list_parent_indices, run-end encoding and extension types.- Conditional transforms:
case_when,choose,replace_with_mask,fill_null_forward/_backward,
indices_nonzero,make_struct,pivot_wider, and a 64-bit element-wisehash64(am_hash64, plus
an FNV-1a form forfixed_size_binary) — an ArrowMetal extension, not one of the 307 Arrow names.
Arrow function coverage
arrowmetal.functions: a registry with one entry per Arrow v25 compute function name — all 307, the
283 in the C++ docs plus the 24hash_*— each carrying its status, Swift file, ArrowMetal call, note,
an executable call adapter under Arrow's own option names and apyarrow.computeoracle.list_functions()
andcall_function()mirror pyarrow's introspection API.docs/ARROW_FUNCTIONS.md, generated from that registry, is the by-name coverage page: 283 gpu, 17 cpu,
7 partial, 0 missing, 0 pending — all 307 names.
Fused expression compiler and lazy query engine
- Expression compiler (
Sources/ArrowMetal/Expr, docs/EXPR.md): a whole Arrow compute expression — arithmetic,
comparisons, null logic, casts, string predicates — lowered to one runtime-generated MSL kernel, so the
inputs are read once and the outputs written once;batch.query(query().filter(...).sum(...))in Swift,
am.query(table, am.filter(pred).sum(am.col("x")))in Python. Filter + aggregate, project and dense-key
group-by fuse into one dispatch; a five-operator expression at 50M rows goes from five kernels to one. - Lazy query engine (
Sources/ArrowMetal/Engine, docs/ENGINE.md):am.scan(table).filter().group_by().agg() .sort().limit().collect()and the SwiftLazyFrame; an optimizer (predicate pushdown, projection pruning,
filter fusion, constant folding, expression CSE, join type-check,explain()), inner/left/right/outer/
semi/anti joins on one or several keys incl. utf8,join_asofwith by-keys and tolerance, window
functions,unique,explode,concat; the plan runs inside one Metal command buffer. A JSON plan
grammar over the C ABI (am_plan_source_create,am_plan_run) for other front ends.
Parquet on the GPU
- A Parquet reader whose column data stays off the CPU for the encodings and codecs listed below
(docs/PARQUET.md): the host parses only
the Thrift footer and page headers; decompression (Snappy, LZ4, LZ4_RAW), definition levels, RLE /
dictionary indices, DELTA_BINARY_PACKED, DELTA_LENGTH_BYTE_ARRAY, DELTA_BYTE_ARRAY and BYTE_STREAM_SPLIT
decode as Metal kernels straight into shared-memory Arrow arrays. ZSTD/GZIP/BROTLI pages decompress on
the host. Row-group and column selection, nested lists; a small host-side writer for round trips.
am.read_parquet(path)in Python,ParquetReaderin Swift.
Out-of-core streaming
- A streaming executor for datasets larger than memory (docs/STREAMING.md): Arrow IPC files/directories
(parallel readers with readahead and backpressure) or any Arrow C Stream flow through the GPU one
record batch at a time; filter/project to a sink, sum/count/min/max/mean/variance, group-by with the
state resident on the GPU, top-k, sort with a bounded k-way merge, broadcast and grace hash joins,
a HyperLogLogcount_distinct_approx(0.0073 % error on 570M rows) and a t-digest quantile. On a
30.21 GB IPC directory: filter + sum in 0.53 s (ties Polars' 0.515), the approximate count-distinct
7.2x faster than Polars (0.532 s against 3.825) on 8.8 GB of peak RSS against Polars' 17.0; top-k,
sort + limit and joins are rows where Polars and DuckDB are ahead, and §9 says so.
Integrations
- Polars (docs/POLARS.md): a zero-copy bridge with
.arrowmetalnamespaces on Series/DataFrame/LazyFrame,
a Rust expression plugin (polars-plugin/) that runs inside a lazy plan, and a streaming hand-off. - DuckDB (docs/DUCKDB.md): a Python bridge that is copy-free where DuckDB returns one chunk and
assembles the DataChunks into one buffer otherwise (duckdb_aggregate,duckdb_group_by, streaming
duckdb_batches) and a loadable C-API extension (duckdb-extension/) exposing seven kernels as SQL
functions; its answers are checked against DuckDB's by the 11@extensiontests in
python/tests/test_duckdb.py, which run onceduckdb-extension/build.shhas built it, and it is not
yet a speedup (2048-row vectors). - pandas (docs/PANDAS.md): an
.amaccessor on Series/DataFrame, and an opt-in accel mode that patches a
documented set of pandas methods, routes to the GPU only when dtype, size and arguments qualify, and
restores the originals exactly onuninstall(). import arrowmetalimports none of the three; each bridge loads on first use through one chained
PEP 562 hook (_LAZY_HOOKS).- The public C header compiles as C, which
test_the_public_c_header_compilesnow holds it to (a
typedef/function name clash,am_plan_source, was a redefinition in both C and C++ and blocked every
C consumer including the DuckDB extension until the reviewer caught it).
Bindings
- libArrowMetalC C ABI (include/arrowmetal.h) and python/arrowmetal ctypes package (Arrow PyCapsule protocol).
- Four language bindings over that ABI, each with its own suite:
rust/(arrowmetal+arrowmetal-sys,
48 tests and 4no_rundoc-tests, docs/RUST.md),go/arrowmetal(45 test functions and one runnable example, docs/GO.md),
node/(N-API addon, 62 tests, docs/TYPESCRIPT.md) andr/arrowmetal(266 testthat expectations over
the 34 header entry points it wraps, docs/R.md). python/build_wheel.shpackages that ctypes package as a macOS arm64 wheel with the dylib bundled in
arrowmetal/_lib/, so an install needs no Swift toolchain; extraspolars,duckdb,pandas,test.
The publication steps are in docs/RELEASE.md (step 4 is the PyPI upload).
Fixed
GroupByKeys._aggin the Python package took the device handle of a temporaryMetalArraythat was
released before the C call read it, segfaulting every grouped aggregate whose values arrived as a
pyarrow array rather than aMetalArray.tdigestbuilt its digest by walking every value on the host. That walk never merged anything: the
weight limit is scaled by the weight seen so far rather than by the column's final weight, and the k1
scale function's inverse is bounded by 1, so every centroid holds exactly one value at any
compression. The digest of a sorted column is that column, so the quantile is read out of it directly
and the nulls are compacted away before the sort. Same answer to the last bit, 653 ms -> 29 ms at 10M
rows.list_value_lengthran at 70 GB/s: one thread per row read every offset twice and stored four bytes at
a time. Eight rows per thread through vector loads and stores, 1.05 ms -> 0.35 ms at 10M rows.sorted()wastake(argsort()), and the gather was most of what it cost above the sort. The radix
sort's own keys are an order-preserving map of the values that is a bijection except on -0.0 and NaN,
so the sorted values are inverted straight out of the sorted keys — and the sort then drops the
row-number payload it only carried for the gather. A column that does hold a -0.0 or a NaN copies back
only the runs they occupy.sort float649.585 -> 6.503 ms at 10M rows and 50.281 -> 31.948 ms at 50M;
sort int649.390 -> 6.142 and 49.409 -> 30.626. The PythonMetalArray.sortreaches it through the new
am_sort_ex; it used to betake(argsort())in the ctypes package and never called this path, and
it keeps returning the input's type, a dictionary column included.- Sorting a column with a validity bitmap lifted the nulls out of the finished permutation on the host:
three passes over the whole index array and a sort of the null row numbers, which cost an order of
magnitude more than the GPU sort it followed. A stable three-way partition (values, NaNs, nulls) now
runs before the sort instead, so the nulls never enter it and keep their input order by construction,
and the passes are planned around the value block rather than the column.sort float64with 10%
nulls 139.745 -> 6.837 ms at 10M rows and 725.935 -> 33.669 ms at 50M; with 50% nulls at 50M,
3,436.130 -> 22.879 ms. (Every figure in these two bullets and in the matching section of
docs/TO_IMPROVE.md comes from one set of runs, recorded with the scripts that produced it: the 10% rows
and the plain sorts from the matrix-conditions harness, the 50% row from the shape sweep. TO_IMPROVE.md
names the file each one is in.) partition_nth_indicessplit the column around the selected key with threecompare+filter
compactions and a concatenation: seven command buffers, and two of its steps ran on the host — the row
numbers it compacted were filled by a CPU loop and the output was allocated zeroed, a write and a
memset over 40 MB at 10M rows. One stable counting sort over five buckets in a single command buffer
instead, 5.13 ms -> 3.02 ms at 10M rows and 19.8 ms -> 12.3 ms at 50M.- Found by the pre-release review pass, each with a regression test:
- Expression compiler: an untyped literal that did not fit the other operand was truncated to it
(int8 > 200was true for every row,uint8 == -1matched 255); a float literal against an integer
column was truncated to an integer. Both now widen the pair, also inif_else/fill_null/coalesce/is_in. - Optimizer: constant folding absorbed a literal into the null-propagating
and/or(only the Kleene
forms may);if_else(c, a, a)dropped the condition's nulls; foldingInt64.min / -1and
abs(Int64.min)trapped the process; expression CSE deduplicatedselectoutputs by name;
join_reorderpermuted rows where the order is observable (now only below a sort, a reduction or a
group-by).am_plan_explain/am_plan_runleaked autoreleased objects when called from Python. - Parquet: nine ways a corrupt file could hang, trap or read out of bounds (unbounded BYTE_ARRAY
lengths, varint overflow traps, unbounded Thrift nesting, unchecked schema cursors, negative offsets
and sizes in footer and page headers, oversized dictionaries and FLBA lengths) now error; an explicit
empty projection returned every column. A projection naming a column twice (and a file with two
columns of one name) lost one of them, because the Python read returned a plain dict; it now returns
a positionalColumnSet. A dictionary-encoded column — whatread_parquetreturns by default —
was refused by every reduction, arithmetic, comparison, cast and sort entry point, so the documented
read_parquet(p)["price"].sum()raised; those entry points now decode the codes on the way in. The
argument-validation paths returned 2 without setting the error string, soam_last_error()handed
the caller an unrelated earlier failure; every non-zero return now names the function and the
argument, and a filter stringselected_row_groupscould not parse is reported instead of dropped. - Kernels:
lexsortignorednull_placementon a utf8/binary key; the regex pre-filter claimed a
literal no-match for three pattern shapes;partition_nth_indicesleft a NaN behind when nulls
moved to the front. - Integrations: an unknown column name in the DuckDB streaming helpers silently used the last column;
a Polars aggregate alias equal to a key replaced the keys;df.am.querylifted every column of the
frame;zero_copy_reportfailed on a Polars DataFrame;countover an empty DuckDB stream was
None; numpy scalars and strings were refused byarrow_table. - The public C header declared
am_plan_sourceas both a typedef and a function; no C consumer
(including the DuckDB extension) could compile it. - Polars tier-2 plugin (its tests had only ever skipped, because nobody had built the Rust
library): the scalar operand of.add/.sub/.mul/.truedivcrossed as anf64and was narrowed
with Rustas, which saturates and truncates instead of refusing —add(1000)on an Int8
column silently becameadd(127),add(-1)on a UInt8 one a no-op,add(1.5)on Int64
add(1), and any integer past 2^53 lost its low bits (add(2**60 + 1)added2**60). The
operand now travels exactly and is range-checked against the column dtype, which is what the
tier-1 bridge'sstruct.packdoes.
- Expression compiler: an untyped literal that did not fit the other operand was truncated to it
Quality
- A CPU oracle behind every kernel test —
Sources/ArrowMetal/CPUReference.swift, a hand-computed
vector, or a pyarrow 25.0.1 answer pinned as a literal; 769 XCTest cases in 61 files, run in release
(all 769 executed, 7 skipped, in the last gated run), including a scenario
matrix over every type, null density, size and sliced input; concurrency and pool tests (docs/TESTING.md).
CI on GitHub's hosted Apple silicon runs the suite in debug and release, but its GPU is virtual, so the
GPU tests skip there and only the build, interop and CPU paths are proved (CONTRIBUTING.md). python/tests/test_functions.pyexecutes the Arrow-name registry: every runnable row is called through
call_functionand compared topyarrow.compute, with a second input in a different Arrow type family
for the rows whose claim spans several, and float tolerances recorded per row infunctions.TOLERANCE.
It found thatbinary_length,binary_repeatandbinary_reverserefuse abinarycolumn — all
three now acceptbinaryandlarge_binary.- Differential matrix against
pyarrow.compute(docs/EVALUATION.md): 39,069 cases per run over 45 column types, 0 unclassified
divergences. Fixed from its findings: stable null order in argsort and top_k, the -0.0 tie in the sort
keys (top_k included, which had its own key mapping), NaN kept at the end of a descending sort, float32
sums in software double, unsigned group-by sums, exact float32 comparison, float32sign/ceil/floor/
trunc/roundand element-wisemin/maxon subnormals and signed zeros. - Benchmarks: Swift vs all-core CPU vs Accelerate; Polars/pyarrow/pandas; ArrowMetal from Python in-process; latency mode.
Benchmarks/full_matrix.pymeasures every CPU library twice: the plain eager idiom and the most parallel idiom
that library has for the same answer (polars-lazythroughpl.LazyFrameon the in-memory or streaming engine,
pyarrow-threadedthrough an Acero plan over one record batch per hardware thread), because the eager idioms use
about one core on the element-wise and reduction rows whatever the pool size.--verifyasserts the two answer
the same within 1e-9 relative, reporting the comparisons deliberately skipped (t-digest's partition-dependent
sketch, pyarrow's threaded groupedlist) separately;--coresandfull_matrix_<date>_cores.txtreport
cpu_ms/wall_ms per idiom; "fastest CPU" in the report is the best of every idiom, named.- 2026-09-07: the published baseline is now that parallel one (
Benchmarks/results/full_matrix_2026-09-07-parallel.csv,
cores per idiom infull_matrix_2026-09-07-parallel_cores.txt: Polars lazy a median of 11.5 cores, pyarrow through
Acero 11.1, the eager idioms 1.0 on most families). Of 339 measured rows: 145 at or above 3x, 102 between 1x and
3x, 77 where the fastest CPU idiom is ahead, 15 with no CPU equivalent — against the eager idioms alone the same
build read 247 / 62 / 15 / 15, and that CSV (full_matrix_2026-09-07.csv) is kept. Ten ArrowMetal rows a
regression check flagged were re-measured on a quieter machine and eight spliced in place, each saying so in its
note. docs/BENCHMARKS_MATRIX.md, docs/BENCHMARKS.md, docs/TO_IMPROVE.md and the README are written against the
parallel baseline; the 77 rows to improve are grouped by measured cause in docs/TO_IMPROVE.md, each with what would
change it. - Adversarial review pass before release (four independent reviewers over the integrations, the engine and
expression compiler, the GPU kernels, and the C ABI and Parquet reader): every finding carries a