Skip to content

1.1.0

Latest

Choose a tag to compare

@andygrove andygrove released this 06 Oct 17:37
· 245 commits to main since this release
992c806

DataFusion Comet 1.1.0 Changelog

This release consists of 416 commits from 41 contributors. See credits at the end of this changelog for more information.

Fixed bugs:

  • fix: skip null slots when checking overflow in unary negation #5162 (Smallfu666)
  • fix: surface next_day and make_date ANSI errors as Spark exceptions #5167 (peterxcli)
  • fix: normalize nested field nullability in ShuffleScanExec and ExpandExec #5138 (andygrove)
  • fix: propagate the Spark task ClassLoader to JVM UDF calls #5282 (andygrove)
  • fix: avoid duplicate CheckOverflow evaluation for decimal division #5225 (peterxcli)
  • fix: match Spark whitespace trimming in to_time and try_to_time #5364 (sunchao)
  • fix: preserve Catalyst nullability and field IDs in native Parquet writes #5369 (sunchao)
  • fix: guard against silent fail_on_error loss in scalar wiring (#5074) #5359 (sam-1112)
  • fix: preserve Spark semantics for dictionary-encoded Parquet inputs and reject dictionary targets #5234 (peterxcli)
  • fix: canonicalize NaN in flat arrays_overlap float keys #5376 (sunchao)
  • fix: report native shuffle write metrics accurately #5370 (sunchao)
  • fix: Native shuffle fails with a 2GB task serialization OOM on jobs with many partitions #5392 (parthchandra)
  • fix: format NativeMemoryConsumer id in toString #5398 (ywskycn)
  • fix: ignore reader-side parquet.hadoop.vectored.io.enabled in Iceberg native-write detection #5410 (snmvaughan)
  • fix: support map casts with NullType elements #5045 (peterxcli)
  • fix: distinguish reflection failure from absent accessor in IcebergReflection #5412 (unikdahal)
  • fix: use current shuffle config in aggregate test #5439 (peterxcli)
  • fix: support wide years in native make_date #5443 (peterxcli)
  • fix: report native child spill metrics in shuffle tasks #5445 (sunchao)
  • fix: expose native memory usage to Spark #5408 (ywskycn)
  • fix: release native shuffle reservation after spill failure #5461 (peterxcli)
  • fix: track peak memory before JVM shuffle spill #5463 (peterxcli)
  • fix: RAII for memory pool registration #5464 (peterxcli)
  • fix: tighten RSS writer visibility and JNI array limits #5475 (pingzh)
  • fix: report native operator spill metrics in Spark task metrics for non-shuffle stages #5497 (peterxcli)
  • fix: report bounded shuffle allocator memory usage #5516 (ywskycn)
  • fix: use Spark type names in ANSI abs overflow errors #5357 (Smallfu666)
  • fix: make collect_list/collect_set argument coercion a normalization barrier #5159 (andygrove)
  • fix: make task-shared memory pool as ref-counted RAII guard #5494 (peterxcli)
  • fix: accept UTC timezone aliases in Python Arrow input #5556 (sunchao)
  • fix: normalize scalar float sort and window rank keys #5469 (sunchao)
  • fix: make CometExplodeExec respect batch size #5362 (andygrove)
  • fix: align Spark 4.2 Python worker configuration #5561 (sunchao)
  • fix: prevent memory leak after failed Arrow vector import #5539 (1fanwang)
  • fix: preserve Arrow Field metadata across C Data exports #5552 (peterxcli)
  • fix: rebase map offsets in mapsort so sliced maps do not overrun entries #5630 (viirya)
  • fix: scope Celeborn bootstrap hooks to Comet clients #5627 (pingzh)
  • fix: skip codegen dispatcher null short-circuit when a foldable subtree can raise #5623 (andygrove)
  • fix: accept dictionary encodings in remote shuffle #5650 (pingzh)
  • fix: make CometDiskBlockWriter spill registry per-task instead of executor-global #5493 (peterxcli)
  • fix: Bump iceberg-rust so native Iceberg writes URL-escape partition paths #5651 (andygrove)
  • fix(celeborn): reject unsafe native push completion tracking #5665 (pingzh)
  • fix: match Spark's null short-circuiting in array_join and enable it natively #5558 (Visorgood)
  • fix: make native shuffle spill metrics independent of input batching #5628 (sunchao)
  • fix: expand object store option references, uniquify constant metadata names, drop dead parquet JNI #5653 (dwsmith1983)
  • fix: match Spark's ANSI bound check for float/double to integral casts #5683 (peterxcli)
  • fix: return NULL from rpad/lpad when the length column is NULL instead of panicking #5680 (peterxcli)
  • fix: fall back for concat_ws with array arguments instead of failing natively #5679 (peterxcli)
  • fix: prevent silent overflow when reading Parquet TIMESTAMP_MILLIS values #5177 (peterxcli)
  • fix: recover native Celeborn shuffle from oversized rows #5668 (pingzh)
  • fix: report native shuffle read metrics #5554 (peterxcli)
  • fix: correctly rounded decimal to double/float cast matching BigDecimal.doubleValue/floatValue #5684 (peterxcli)
  • fix: restore columnar transitions under the native Iceberg write #5696 (andygrove)
  • fix: make columnar-to-row benchmarks exercise Comet #5718 (rich7420)
  • fix: distinguish "nothing spilled" from a spill backend with no local path #5726 (andygrove)
  • fix: propagate Arrow array copy errors #5747 (rich7420)
  • fix: keep the dictionary hash fast path off nested and reseeded buffers #5757 (viirya)
  • fix: preserve ANSI errors for rejected TIMESTAMP_NTZ casts #5752 (peterxcli)
  • fix: apply the parent struct's null mask before hashing its fields #5754 (viirya)
  • fix: read Iceberg tables partitioned by an unknown transform #5759 (andygrove)
  • fix: native Iceberg write panics on an evolved partition spec and on a timestamptz partition path #5729 (andygrove)
  • fix: let AQE optimize queries over Comet caches #5733 (peterxcli)
  • fix: dispatch Iceberg system functions wrapped as ApplyFunctionExpression #5773 (andygrove)
  • fix: attach tokio runtime threads to the JVM as daemon threads #5748 (zhangfengcdt)
  • fix: check nested TIMESTAMP_MILLIS overflow in unfiltered scans #5740 (peterxcli)
  • fix: match iceberg-java's exception for unclustered input to a clustered Iceberg write #5779 (andygrove)
  • fix: enable FIRST/LAST partial merge #5041 (peterxcli)
  • fix: preserve aggregate result identity during exchange reuse #5470 (sunchao)
  • fix: ignore structural tags when lifting expression coverage #5471 (sunchao)
  • fix: align string to timestamp parsing with Spark's segment rules #5682 (peterxcli)
  • fix: roll native Iceberg data files on iceberg-java's 1000-row grid #5780 (andygrove)
  • fix: decide libhdfs routing from the scheme as written #5825 (comphead)
  • fix: list a fanout Iceberg write's data files in a stable order #5810 (andygrove)
  • fix: isolate object-store registration by backend and configuration #5503 (sunchao)
  • fix: Delete completed tasks' data files when an Iceberg write job fails #5663 (andygrove)
  • fix: apply Spark's Parquet conversion rules to nested struct/list/map fields #5681 (peterxcli)
  • fix: preserve nulls for Boolean/Byte/Short/Integer columns in FuzzDataGenerator #5855 (Smallfu666)
  • fix: Accept explicit positive years in timestamp casts #5858 (peterxcli)
  • fix: decline structs with duplicate field names before they reach Java Arrow #5866 (dwsmith1983)
  • fix: explain ObjectHashAggregate fallback when Comet shuffle is disabled #5746 (0lai0)
  • fix: preserve join and generator semantics in plan identity #5828 (ErikBPF)
  • fix: propagate Parquet field-name folding failures #5845 (sunchao)
  • fix: decode dictionary input for PyArrow UDFs #5560 (sunchao)
  • fix: size JVM shuffle pointer array growth from the array, not the data pages #5907 (andygrove)
  • fix: render float and double Iceberg partition values like iceberg-java #5840 (andygrove)
  • fix: align time parsing and native second extraction with Spark #5738 (peterxcli)
  • fix: normalize floating-point values in native collect_set #5166 (peterxcli)
  • fix: keep Iceberg complex null checks on native scans #5732 (ErikBPF)
  • fix: decode invalid UTF-8 at the JVM to native FFI import boundary #5310 (manuzhang, andygrove)
  • fix: fall back when a struct repeats a Parquet field id #6004 (comphead)
  • fix: normalize signed zero in nested float array comparisons #5235 (divyankshah)
  • fix: fall back to Spark for native Iceberg writes to gs:// through HadoopFileIO #5935 (zhangfengcdt)
  • fix: gate the regr_r2 degenerate-case swap on the Spark patch release #6042 (dwsmith1983)
  • fix: include ABFS container in object store cache key #5053 (peterxcli)
  • fix: preserve Spark row index read errors #6046 (liupoyi-1031)
  • fix: remove the ineffective spark.executor.memoryOverhead adjustment from the driver plugin #6054 (andygrove)
  • fix: count each memory pool once in analyze_trace #5991 (andygrove)
  • fix: revert unsafe partial aggregates after final fallback #5421 (sunchao)
  • fix: restore the site's mermaid diagrams and make a dropped one fail CI #6064 (andygrove)
  • fix: write cached batches to the schema width, not the batch width #6090 (andygrove)
  • fix(iceberg): guard native Iceberg scan driver-metric double-post, add metrics docs and tests #6085 (parthchandra)
  • fix: let decimal SUM recover from an intermediate overflow like Spark #6041 (dwsmith1983)
  • fix: Nested floating-point IN membership does not match Spark for signed zero #6073 (mizulun)
  • fix: let Comet memory pools overcommit on grow instead of panicking #6128 (andygrove)
  • fix: correct two nightly test failures on Spark 3.4 and 4.2 #6156 (andygrove)
  • fix: support null calendar interval literals #5133 (peterxcli)
  • fix: wrap Iceberg split-write failures the way Spark does when abort fails #6153 (andygrove)
  • fix: support Utf8/LargeUtf8/Utf8View in native RLike without panicking #5215 (sam-1112)
  • fix: normalize noncanonical NaN literals in comparisons #5472 (sunchao)
  • fix: preserve map field metadata and honor target sorted flag in cast_map_to_map #5227 (Smallfu666)
  • fix: reject duplicate Parquet field names before decoding #5786 (ErikBPF)
  • fix: preserve Parquet conversion errors during join filtering #6067 (pingzh)
  • fix: plan Iceberg writes with Spark's own operator when Comet is disabled #6151 (andygrove)
  • fix: read shuffle write buffer, spill limit and off-heap sizes in bytes #6191 (andygrove)
  • fix: bound shuffle schema cache retention and preserve eviction order #6098 (sunchao)
  • fix: do not run Comet in on-heap mode without spark.comet.exec.onHeap.enabled #6195 (andygrove)
  • fix: remove misleading native opt-in for dispatch-only datetime expressions #6182 (LinSimon-901101)
  • fix: support CalendarIntervalType hashing #5135 (peterxcli)
  • fix: report Input column when native Iceberg scan is enabled or native shuffle is enabled #5265 (hsiang-c)
  • fix: skip the executor memory overhead warning when a factor is set or in local mode #6198 (andygrove)
  • fix: keep operators above a cached relation native after the AQE re-plan #6208 (andygrove)
  • fix: preserve current AQE logical-stage links on Comet operators #5483 (sunchao)
  • fix: check each fair_unified reservation against its own share #6205 (andygrove)
  • fix(iceberg): don't push transform residuals as their source column, fail on residual errors #6154 (andygrove)
  • fix: match Spark's duplicate field and field id semantics in parquet field lookup #5654 (dwsmith1983)
  • fix: park the native scan loop instead of busy-polling while waiting on native I/O #6219 (andygrove, mixermt)
  • fix: refresh S3 policy locations when a location's credential fails #6223 (andygrove, snmvaughan)
  • fix: inject Comet session extension rules only once per session #6230 (andygrove)
  • fix: [branch-1.1] reject a file without field ids at any depth whether or not id matching is on (#6116) #6266 (andygrove, dwsmith1983)
  • fix: [branch-1.1] decline native Iceberg writes with a custom location provider (#6216) #6305 (andygrove, liupoyi-1031)
  • fix: [branch-1.1] Native S3 scan on EKS/IRSA turns a transient STS throttle into a hard 403 storm (#6025) #6323 (andygrove, parthchandra)
  • fix: [branch-1.1] log partial memory grants at DEBUG and drop the memory usage dump (#6269) #6346 (andygrove)
  • fix: [branch-1.1] zero sliced boolean offsets at every level before exporting to the JVM (#6339) #6449 (andygrove)
  • fix: [branch-1.1] build with Rust 1.99, which deprecates the legacy f64 constants and fetch_update (#6507) #6526 (andygrove)
  • fix: [branch-1.1] make adaptive aggregation skipping opt-in (#6474) #6488 (andygrove)
  • fix: [branch-1.1] fall back to Spark for _metadata.file_block_start and file_block_length (#6510) #6511 (andygrove)
  • fix: [branch-1.1] fall back for rank limits over nested float keys (#6468) #6487 (andygrove)
  • fix: [branch-1.1] fall back for incompatible regression aggregates (#6451) #6489 (andygrove)
  • fix: [branch-1.1] coerce native IF branches to a common type (#6458) #6491 (andygrove)
  • fix: [branch-1.1] match pre-epoch Iceberg temporal rounding (#6456) #6486 (andygrove)
  • fix: [branch-1.1] defer Parquet conversion errors until a row group is decoded, as Spark does (#6515) #6540 (andygrove)
  • fix: [branch-1.1] dispatch StaticInvoke and Invoke only into Spark's own classes (#6542) #6554 (andygrove)
  • fix: [branch-1.1] gate array distinct and union signed-zero semantics by Spark version (#5750) #6561 (andygrove, peterxcli)
  • fix: [branch-1.1] fall back from the native Iceberg scan when a nested field was added or renamed (#6543) #6560 (andygrove)

Performance related:

  • perf: optimize spark_base64 in spark-expr #4885 (andygrove)
  • perf: optimize spark_floor (up to 4x faster) #4911 (andygrove)
  • perf: cache Iceberg reflection lookups on the planning path #5222 (andygrove)
  • perf: vectorize integer-to-decimal cast #4939 (andygrove)
  • perf: compute spark_size list lengths with Arrow length kernel #5233 (0lai0)
  • perf: make CometShuffleBenchmark completable and fast #5388 (andygrove)
  • perf: vectorize Map in spark_size via offset buffer #5395 (0lai0)
  • perf: improve ArrowWriter performance for fixed-length vectors #5046 (peterxcli)
  • perf: reuse Arrow IPC compression context across shuffle blocks #5038 (peterxcli)
  • perf: use Arrow cast for decimal rescale check #5440 (peterxcli)
  • perf: bulk copy fixed-width columns in ArrowWriter #5442 (peterxcli)
  • perf: serialize Python input directly from Comet Arrow vectors #5368 (sunchao)
  • perf: reuse per-partition scratch in the shuffle write path #5568 (dwsmith1983)
  • perf: validate shuffle IPC context reuse savings #5727 (peterxcli)
  • perf: slice the child instead of gathering it when unnesting #5667 (andygrove)
  • perf: evaluate posexplode array expressions once per batch #5737 (rich7420)
  • perf: cache expected schemas for remote shuffle decoding #5722 (pingzh)
  • perf: use Arrow cast for date to timestamp NTZ #5735 (peterxcli)
  • perf: reduce allocations when collecting cache statistics #5734 (peterxcli)
  • perf: avoid repeated decimal promotion in expression serialization #5736 (peterxcli)
  • perf: give collect_list and collect_set a native GroupsAccumulator #5803 (andygrove)
  • perf: vectorize the native map lookup behind element_at and GetMapValue #5806 (andygrove)
  • perf: compile user regex patterns once per planned expression #5612 (dwsmith1983)
  • perf: batch the nested-element list hash for flat struct elements #5778 (viirya)
  • perf: optimize map_sort singleton normalization (18x faster) #5887 (viirya)
  • perf: optimize map_sort for multi-entry string maps (up to 3x faster) #5901 (viirya)
  • perf: skip calendar reconstruction in hour/minute/second and dayofweek/weekday #5771 (peterxcli)
  • perf: reuse zstd compression contexts across shuffle blocks #5565 (dwsmith1983)
  • perf: cache parsed plan data across a stage tasks #5615 (dwsmith1983)
  • perf: spill every shuffle partition of a task into one file #5916 (peterxcli)
  • perf: decode shuffle blocks against a cached schema instead of re-parsing per block #5809 (peterxcli)
  • perf: project cached batches by buffer selection, prune on collated strings #5543 (andygrove)
  • perf: make one thread-local access per tracked allocation #6166 (andygrove)
  • perf: reuse quantile summary buffers during merge #4932 (peterxcli)
  • perf: optimize list_extract without defaults using Arrow take #5174 (peterxcli)
  • perf: track decimal overflow without rescanning results #5044 (peterxcli)

Implemented enhancements:

  • feat: support explode_outer #5192 (comphead)
  • feat: support timestampadd and timestampdiff via codegen dispatch #5030 (andygrove)
  • feat: support _metadata constant columns in native Parquet scan #5237 (mbutrovich)
  • feat: add make_interval support (codegen dispatch + native) #5039 (peterxcli)
  • feat: Optionally split the Iceberg V2 write operator into distinct writer and committer operations #4658 (jordepic)
  • feat: detect Iceberg V2 writes and emit fall-back reasons #5298 (jordepic)
  • feat: remove native cast from boolean to decimal #5185 (andygrove)
  • feat: add micro benchmark runner and EC2 guide #5374 (andygrove)
  • feat: build gate + inert wiring for contrib Delta scans [Delta contrib split, part 2] #4952 (schenksj)
  • feat: Support native scans with unprojected Spark 4 VARIANT columns #5377 (sunchao)
  • feat: expose native aggregate spill and memory metrics #5423 (sunchao)
  • feat: support WindowGroupLimitExec #4870 (comphead)
  • feat: add RSS partition writer and task-owned JNI callback (1/n) #5473 (pingzh)
  • feat: native spark_unbase64 kernel #5451 (0lai0)
  • feat: add partition writer destinations to native shuffle plans (2/n) #5476 (pingzh)
  • feat: add destination-aware native shuffle execution (3/n) #5481 (pingzh)
  • feat: bind task-owned RSS callbacks to native shuffle plans (4/n) #5491 (pingzh)
  • feat: add Celeborn shuffle manager and partition pusher (5/n) #5501 (pingzh)
  • feat: complete Celeborn map-side shuffle push lifecycle (6/n) #5513 (pingzh)
  • feat: add experimental native support for in-memory cache, disabled by default #5051 (andygrove)
  • feat: add raw Celeborn native shuffle reader (7/n) #5531 (pingzh)
  • feat: support Map for CreateArray literal #5452 (comphead)
  • feat: route round on float/double through the codegen dispatcher #5600 (andygrove)
  • feat: enable native-only Celeborn shuffle planning (8/n) #5537 (pingzh)
  • feat: support unicode case sensitive field names for reading parquet #5602 (comphead)
  • feat: implement native Iceberg V2 writer via iceberg-rust #5361 (jordepic)
  • feat: Support Iceberg system functions (bucket, truncate, years/months/days/hours) natively #5638 (andygrove)
  • feat: implement regr_slope, regr_intercept, regr_r2, regr_sxx, regr_syy, regr_sxy aggregates #4775 (andygrove)
  • feat: carry VariantType identity through schema serialization #5631 (peterxcli)
  • feat: route unrecognized StaticInvoke and Invoke through the codegen dispatcher #5692 (andygrove)
  • feat: support S3 compliant filesystems #5314 (comphead)
  • feat: add native spark_sequence kernel for integral element types #5614 (0lai0)
  • feat: support nested types as native shuffle hash partitioning keys #5567 (viirya)
  • feat: route next_day and levenshtein collated input through the codegen dispatcher #5720 (0lai0)
  • chore: add_benches hash function aggregators #5730 (coderfender)
  • test: cover Spark-to-Arrow batch conversion edge cases #5713 (peterxcli)
  • feat: normalize marked Variant arrays at the native Parquet boundary #5715 (peterxcli)
  • feat: support native concat_ws with string arrays #5725 (peterxcli)
  • feat: native dynamic filter pushdown for hash join into Parquet scans #5699 (pingzh)
  • feat: enable codegen dispatch for lpad and rpad #5764 (Satyr09)
  • feat: address remaining issues for CreateArray #5766 (comphead)
  • ci: stop non-gating labels from skipping Preflight #5784 (comphead)
  • feat: route abs on interval types through the codegen dispatcher #5622 (kazantsev-maksim)
  • ci: shard Iceberg Spark tests across four runners #5459 (sunchao)
  • test: expand ANSI coverage for round, conv and elt #5799 (rich7420)
  • test: strengthen ANSI exception assertions #5800 (rich7420)
  • chore: Improve network retry configuration for maven and artifact upload #5782 (comphead)
  • ci: gate the Delta contrib build on symbols, not on libcomet size #5827 (andygrove)
  • test: add helpers to assert whether an expression ran natively or via codegen dispatch #5610 (andygrove)
  • feat: support Spark encode expression via codegen dispatch #5037 (andygrove)
  • feat: support native aggregate function mode #4782 (andygrove)
  • test: name the whole dispatched subtree in the decimal promotion assertion #5849 (andygrove)
  • feat: route translate and to_csv through codegen dispatch by default #5032 (andygrove)
  • feat: expose native Parquet scan I/O and read-amplification metrics #5453 (sunchao)
  • ci: move the job routing policy out of ci.yml expressions and into compute-changes.py #5850 (andygrove)
  • feat: support max_by and min_by aggregate expressions #4817 (andygrove)
  • ci: add a Required Checks aggregator job so main can have a required status check #5842 (andygrove)
  • feat: adapt Parquet storage for Variant projection #5794 (peterxcli)
  • ci: move the Spark 3.4/3.5/4.0, Iceberg, macOS and benchmark suites behind a merge queue #5843 (andygrove)
  • feat(iceberg): Iceberg table format V3: apply deletion vector on reads #5853 (mbutrovich)
  • test: restore ANSI array access error coverage #5798 (rich7420)
  • test: cover native memory accounting boundaries #5856 (rich7420)
  • test: cover date maps in Parquet temporal fuzz tests #5877 (rich7420)
  • feat: Enable adaptive partial aggregation for eligible native shuffle plans #5723 (peterxcli)
  • ci: move the Spark 4.1 sql_hive shards behind the merge queue #5871 (andygrove)
  • ci: retry the Maven wrapper bootstrap in every job that calls ./mvnw directly #5852 (andygrove)
  • test: cover slice over expression-produced non-null element arrays (#… #5839 (sam-1112)
  • test: run the libhdfs suite manually instead of in CI #5892 (andygrove)
  • feat: Add Lance contrib build gate #5728 (wirybeaver)
  • chore: drop support for JDK 11 #5897 (manuzhang)
  • test: cover regex expression routing configurations #5917 (rich7420)
  • test: cover round expression routing configurations #5878 (rich7420)
  • test: cover to_csv routing configurations #5921 (rich7420)
  • test: cover interval expression routing configurations #5920 (rich7420)
  • test: cover array and map expression routing configurations #5918 (rich7420)
  • test: make the cancelled Iceberg abort test yield deterministically #5919 (andygrove)
  • chore: deprecate Spark 3.4 rather than removing it in 1.1.0 #5885 (andygrove)
  • test: restore the TPC-H suite’s 2 GiB off-heap budget #5904 (ErikBPF)
  • test: cover string expression routing and native opt-in #5915 (rich7420)
  • chore: run only the cache-writing jobs on push to main #5930 (andygrove)
  • ci: publish nightly SNAPSHOT jars to repository.apache.org #5902 (andygrove)
  • ci: fold the Delta build gate and PyArrow UDF suite into the merge queue tiers #5926 (andygrove)
  • refactor: use Arrow decimal precision validation #5160 (Smallfu666)
  • ci: shrink the pull request tier to the Linux build on the default Spark profile #5939 (andygrove)
  • ci: run the non-default Spark and Iceberg suites nightly instead of in the merge queue #5963 (andygrove)
  • test: cover lpad and rpad routing configurations #5895 (rich7420)
  • test: preserve Spark SQL baselines for ordering-sensitive fixtures #5914 (rich7420)
  • test: cover json_array_length routing configurations #5940 (peterxcli)
  • test: cover SQL fixture metadata and statement parsing #5941 (peterxcli)
  • test: cover crc32 on binary inputs #5942 (peterxcli)
  • test: cover datetime and timezone routing configurations #5950 (rich7420)
  • test: cover collation and predicate routing configurations #5951 (rich7420)
  • test: cover array_intersect routing configurations #5952 (rich7420)
  • test: guard native plan equality against omitted parameters #5953 (rich7420)
  • test: expand replace compatibility regression coverage #5409 (sam-1112)
  • feat: add native allocation accounting for memory observability #5934 (andygrove)
  • feat: run rlike natively by default for Java-equivalent literal patterns #5415 (sam-1112)
  • ci: write large actions/cache entries only on push to main #5973 (andygrove)
  • refactor: extract shared runtime filter components #5937 (pingzh)
  • test: skip the Spark RocksDB state-store suites when Comet is enabled #5987 (comphead)
  • chore: bench nondeterministic and json kernels #5989 (coderfender)
  • feat: add BinaryType support for SortMergeJoin #5928 (xhumanoid)
  • refactor: replace remaining hand-rolled loops with Arrow kernels #5367 (0lai0)
  • feat: support Spark 4 EmptyRelationExec as a native input #5821 (jeffw13)
  • chore: Bench agg (welford) stats #5988 (coderfender)
  • ci: rename the umbrella workflow from CI to Comet CI #6003 (andygrove)
  • ci: bootstrap Maven from the setup actions so every mvnw job is covered #5881 (andygrove)
  • ci: add dev/local-ci.sh to run the Spark SQL and Iceberg suites locally #5974 (comphead)
  • feat: narrow strict floating-point admission for scalar sort keys #5981 (0lai0)
  • feat: route compatible xxhash64 args through SparkXxhash64 #5960 (sam-1112)
  • feat: support string arrays in Spark-to-Comet conversion #5954 (rich7420)
  • feat: run length on binary input natively #5874 (dwsmith1983)
  • chore: drop the redundant width_bucket shim registrations and guard serde uniqueness #5873 (dwsmith1983)
  • feat: hook native Parquet writes into Spark's WriteFilesExec seam on Spark 4.0+ #5763 (andygrove)
  • feat(iceberg): report native Iceberg scan planning metrics and scan time in the Spark UI #6027 (parthchandra)
  • test: guard page skipping in the native Iceberg scan #6040 (dwsmith1983)
  • test: stop claiming Spark agrees on signed-zero array literals #6055 (cestercian)
  • feat: trace Arrow memory held on the JVM side #6048 (andygrove)
  • feat: unix_timestamp codegen dispatch for strings, and fix pre-epoch fractional rounding #5789 (Satyr09)
  • chore: remove JaCoCo from the build #6084 (andygrove)
  • refactor: compose CometScanRule and CometExecRule into a single CometRule #6082 (andygrove)
  • chore: add a native Iceberg write benchmark #6038 (0lai0)
  • ci: retry the Spark test pre-compile step on resolution failures #6079 (andygrove)
  • ci: focus Miri on unsafe row and hash tests #6072 (rich7420)
  • test: run the Iceberg Spark tests with the native Iceberg writer enabled #5677 (andygrove)
  • feat: support listagg / string_agg aggregate (Spark 4.0+) #4816 (andygrove)
  • ci: run label-triggered CI as a separate workflow #6161 (andygrove)
  • feat: always count native allocations and log executor native memory usage #6162 (andygrove)
  • feat: use DataFusion unnest_outer instead of Comet ListEmptyToNullExpr #6132 (comphead)
  • chore: recommend --force-with-lease when updating PR branches #6171 (mizulun)
  • feat: positional round robin shuffle keyed on a row ordinal #6095 (andygrove)
  • chore: deprecate spark.comet.exec.memoryPool.fraction #6163 (andygrove)
  • feat: add spark.comet.explain.planOnly.enabled #5394 (andygrove)
  • feat: support direct Variant projection in native Parquet scans #5868 (peterxcli)
  • test: cover JSON and cast expression routing #6068 (rich7420)
  • feat: remove memory accounting from Comet's on-heap mode #6066 (andygrove)
  • test: accept CometHashAggregateExec in the Spark 4.0 CollationSuite hash agg check #6220 (andygrove)
  • feat: per-location credentials for the S3 credential SPI #6031 (snmvaughan)
  • refactor: centralize data type support predicates #5025 (peterxcli)
  • chore: [branch-1.1] change version from 1.1.0-SNAPSHOT to 1.1.0 #6236 (andygrove)
  • feat: [branch-1.1] built-in S3 credential provider adapters for the native Parquet scan (#6023) #6318 (andygrove, parthchandra)
  • chore: [branch-1.1] Add 1.1.0 changelog #6282 (andygrove)
  • test: [branch-1.1] fix flaky mixed field id directory test on Spark 3.4 and 3.5 (#6312) #6527 (andygrove)
  • test: [branch-1.1] cover sliced boolean arrays in explode (#6473) #6492 (andygrove)
  • test: [branch-1.1] backport the Iceberg write report and the mid-write retry test (#6155, #6111) #6307 (andygrove, sam-1112)
  • feat: [branch-1.1] reuse S3 credentials until shortly before their reported expiry (#6509) #6545 (andygrove, snmvaughan)

Documentation updates:

  • docs: add note about run-iceberg-tests label in CI #5247 (mbutrovich)
  • docs: add 1.0.0 TPC-DS benchmark results, remove TPC-H #5284 (mbutrovich)
  • docs: add 1.0.0 changelog to main #5263 (andygrove)
  • docs: correct Spark 4.2 version and CI test status in installation guide #5315 (andygrove)
  • docs: add suggest-native-expression skill for assessing native expression candidates #5348 (andygrove)
  • docs: document C2R cost for wide/nested schemas in tuning guide #5458 (DebadityaHait)
  • docs: update stale interval multiplication note in expressions.md #5518 (peterxcli)
  • docs: add contributing guide link #5521 (Dharan-K)
  • doc: fix benchmark examples for MacOS #5522 (xhumanoid)
  • docs: add comet meeting link #5621 (coderfender)
  • docs: add a contributor guide page for CI and the merge queue #5863 (andygrove)
  • docs: add contributor guide page on memory management #5933 (andygrove)
  • docs: how to check the scheduled CI runs are actually running #6000 (andygrove)
  • docs: explain allocator hazards and diagram where memory is allocated #6014 (andygrove)
  • docs: pre-render mermaid diagrams to SVG at build time #6021 (andygrove)
  • docs: split the PR review skill by area and correct the shuffle contributor docs #6018 (andygrove)
  • docs: correct the stale range-partitioning strict floating-point rule #6049 (andygrove)
  • docs: recommend setting spark.executor.memoryOverhead alongside off-heap memory #6051 (andygrove)
  • docs: exempt testing-category and internal configs from the versioning policy #6089 (andygrove)
  • docs: make doc changes weekly comet sync #6160 (coderfender)
  • docs: correct the plugin and shuffle sections of the plugin overview #6197 (andygrove)
  • docs: fix stale and missing native Iceberg write details #6150 (andygrove)
  • docs: add an Iceberg writes contributor guide and review skill #6149 (andygrove)
  • docs: correct which operators run separate native plans in a task #6194 (andygrove)
  • docs: update the user guide for the 1.1.0 release #6168 (andygrove)
  • docs: [branch-1.1] backport the 1.1.0 upgrade notes and tuning guide updates (#6237, #6244, #6248) #6265 (andygrove)
  • ci: [branch-1.1] run every tier on release-branch pull requests, and the full suite before an RC (#6218) #6284 (andygrove)
  • docs: [branch-1.1] generate release docs for 1.1.0 #6344 (andygrove)
  • docs: [branch-1.1] credit co-authors in the 1.1.0 changelog #6359 (andygrove)

Other:

  • chore: start 1.1.0 development #5242 (andygrove)
  • refactor: replace hand-coded rollup of expression fallback reasons onto operators #5236 (andygrove)
  • test: rename CometCastSuite to CometNativeCastSuite #5268 (andygrove)
  • refactor: use arity helper for Int to Decimal128 reinterpretation #5193 (0lai0)
  • test: add guard for Iceberg version on Variant fallback test #5278 (mbutrovich)
  • chore: update documentation links for 1.0.0 release #5290 (andygrove)
  • chore(deps): bump reqsign-core from 3.2.0 to 3.2.1 in /native in the all-other-cargo-deps group #5287 (dependabot[bot])
  • chore(deps): bump object_store_opendal from 0.57.0 to 0.58.0 in /native #5289 (dependabot[bot])
  • chore(deps): bump the codeql-actions group with 2 updates #5286 (dependabot[bot])
  • bug: Revert "chore(deps): bump object_store_opendal from 0.57.0 to 0.58.0 in /native" #5332 (coderfender)
  • test: cover the narrowing direction of cast_and_stamp_schema #5285 (andygrove)
  • chore(deps): bump opendal from 0.57.0 to 0.58.1 in /native #5324 (manuzhang)
  • chore: respect Cargo parallelism settings for native release builds #5344 (pingzh)
  • test: make expression benchmark harness fair and reproducible #5371 (andygrove)
  • chore(deps): bump the all-other-cargo-deps group in /native with 5 updates #5360 (dependabot[bot])
  • test: fix vacuous signed-zero coverage in SQL file tests #5393 (sam-1112)
  • chore: fix clippy warnings for Rust 1.98 #5400 (ywskycn)
  • chore(deps): bump the codeql-actions group with 2 updates #5405 (dependabot[bot])
  • test: strengthen signed-zero SQL assertions #5404 (sunchao)
  • chore(deps): bump actions/checkout from 6 to 7 #5406 (dependabot[bot])
  • test: add expression fallback-invariance suite #5329 (4ktLuffy)
  • refactor: delegate ANSI integer arithmetic to arrow checked kernels #5280 (kazantsev-maksim)
  • test: docs and test-coverage hardening for native collect_list / collect_set #5055 (andygrove)
  • test: deduplicate Throwable cause-chain traversal #5441 (peterxcli)
  • refactor: use Arrow timezone type #5129 (Hashim1999164)
  • chore: add md formatting to make format #5460 (comphead)
  • test: fail unexpected vacuous fallback-invariance checks #5417 (sunchao)
  • ci: cache Maven distributions and retry bootstrap downloads #5422 (sunchao)
  • chore: add math benches #5479 (coderfender)
  • chore: Bench string functions #5492 (coderfender)
  • chore(deps): bump actions/setup-java from 5 to 6 #5524 (dependabot[bot])
  • chore(deps): bump the codeql-actions group with 2 updates #5523 (dependabot[bot])
  • chore: bench additional math function #5520 (coderfender)
  • chore: remove redundant condition and stray println in shuffle code #5562 (viirya)
  • test: cover struct data columns in native shuffle #5564 (viirya)
  • chore: rename .claude directory to vendor-neutral .ai #5594 (andygrove)
  • chore: bench additional scalar functions #5598 (coderfender)
  • test: verify Celeborn reflection compatibility #5604 (pingzh)
  • test: enable native path in lower/upper_enabled sql fixtures #5619 (cestercian)
  • chore: add benches datetime #5620 (coderfender)
  • chore: move dead and defensive serde guards out of convert #5595 (andygrove)
  • test: exercise the Iceberg write split-operator plan in Iceberg's own suites #5640 (andygrove)
  • chore(deps): bump the codeql-actions group with 2 updates #5669 (dependabot[bot])
  • test: add explode operator microbenchmark #5381 (andygrove)
  • deps: bump DataFusion 55.0 and Arrow/Parquet 59.2 #5262 (mbutrovich)
  • chore: add benches array functions #5700 (coderfender)
  • test: restore Comet coverage for recursive HAVING and ORDER BY #5755 (rich7420)
  • test: restore Parquet V2 writer and delta encoding coverage #5760 (rich7420)
  • chore: Add benches for datetime funcs #5767 (coderfender)
  • bench: add a benchmark for the Spark hash kernels #5765 (viirya)
  • test: restore Spark 4.1 Variant shredding suites #5745 (rich7420)
  • refactor: share one helper for pushing a struct's null mask into its children #5769 (viirya)
  • ci: label pull requests by changed paths and title prefix #5762 (dwsmith1983)
  • test: cover ambiguous exact nested Parquet field matches #5751 (peterxcli)
  • bench: measure nested types as native shuffle hash partitioning keys #5788 (viirya)
  • deps: bump to datafusion 55.1.0 #5865 (comphead)
  • bench: isolate map normalization and nested key hashing #5822 (viirya)
  • bench: add a shuffle read benchmark covering the per-block schema parse #5805 (peterxcli)
  • chore(deps): bump actions/setup-java from 4 to 6 #6012 (dependabot[bot])
  • chore(deps): bump the codeql-actions group with 2 updates #6011 (dependabot[bot])
  • deps: bump the iceberg-rust pin to bb1e4a4 and document why it is pinned #6094 (andygrove)

Credits

Thank you to everyone who contributed to this release. Here is a breakdown of commits (PRs merged) per contributor. A PR with commits from more than one person counts for each of them.

   147	Andy Grove
    62	Peter Lee
    26	Chao Sun
    26	KUAN-HAO HUANG
    19	Ping Zhang
    15	Oleks V
    14	dustin
    13	Bhargava Vadlamani
    13	Liang-Chi Hsieh
    11	dependabot[bot]
    10	ChenChen Lai
     8	sam-1112
     6	Matt Butrovich
     5	Han-Yin Chang
     5	Parth Chandra
     4	Erik Bogado
     4	Steve Vaughan
     4	Wei Yan
     3	Jordan Epstein
     3	Manu Zhang
     2	Alexey
     2	Cestercian
     2	Daipayan Mukherjee
     2	Feng Zhang
     2	Kazantsev Maksim
     2	YuLun Mao
     2	liupoyi-1031
     1	Dharan-K
     1	Hashim Khan
     1	Henos D
     1	Jeff Wang
     1	LinSimon-901101
     1	Michael Taranov
     1	Scott Schenkein
     1	Stefan Wang
     1	Unik Dahal
     1	Viacheslav Inozemtsev
     1	Xuanyi Li
     1	debaditya
     1	divyank sameer shah
     1	hsiang-c

Thank you also to everyone who contributed in other ways such as filing issues, reviewing PRs, and providing feedback on this release.