Skip to content

perf: specialize native cache statistics collection - #5734

Open
peterxcli wants to merge 1 commit into
apache:mainfrom
peterxcli:perf/4781-native-cache-statistics
Open

perf: specialize native cache statistics collection#5734
peterxcli wants to merge 1 commit into
apache:mainfrom
peterxcli:perf/4781-native-cache-statistics

Conversation

@peterxcli

Copy link
Copy Markdown
Member

Which issue does this PR close?

Closes #4781.

Rationale for this change

Building an Arrow-backed cache currently dispatches on each column's data type for every row, boxes primitive values, and copies every non-null string while collecting statistics. Only the final lower and upper bounds need to survive the batch.

The issue's projection hoist and scan-wide deep-copy removal are already present. This PR implements the remaining statistics optimizations; the optional uncompressed cache format is deferred.

What changes are included in this PR?

  • Dispatch once per column and maintain primitive bounds in primitive locals, boxing only the final bounds.
  • Specialize decimal and string loops. Reuse Spark's string ordering and copy a string only when it replaces a bound, avoiding comparison-time byte-array copies.
  • Preserve null counts, all-null bounds, and the existing floating-point ordering, including NaN and signed zero.
  • Add typed-statistics regression coverage and a cache-materialization case to CometInMemoryCacheBenchmark.

How are these changes tested?

  • make core passed.
  • JVM compilation, Spotless and Scalastyle passed.
  • SPARK_LOCAL_IP=127.0.0.1 ./mvnw test -Dtest=none -Dsuites=org.apache.comet.exec.CometInMemoryCacheSuite: 34 tests passed on Spark 4.1.3.
  • The new regression compares stored bounds with Spark's uncached aggregates across primitive, decimal, string, date and timestamp types, covering leading nulls, all-null columns and repeated values. Existing NaN and batch-pruning tests also pass.
  • Inspected generated JVM bytecode to verify primitive boxing occurs after the row loops.

Benchmarks

Compared the original serializer at 75fdddc92 with this patch, using Spark 4.1.3 and JDK 21.0.6 on macOS. Both workloads use three long columns and three string columns. Results are medians; allocation units are decimal.

Measurement, 5 million rows Before After Result
Isolated statistics collection 508.00 ms 156.48 ms 3.25x faster
Statistics-loop allocated bytes 2.260 GB 0.724 GB 68% less
Full JVM cache materialization 1.926 s 1.935 s No reliable overall improvement

Statistics microbenchmark: invoked the actual before/after gatherColumnStats methods over a prebuilt 10,000-row OnHeapColumnVector batch, 500 times per sample. Both implementations ran in the same JVM with alternating order, five warmup rounds and 15 measured rounds. Bounds and null counts were checked for equality; results were consumed through a volatile sink. Allocation was measured using ThreadMXBean. These timings isolate statistics over Spark on-heap vectors, not decoded Arrow vectors.

Full materialization: used the same input expressions as CometInMemoryCacheBenchmark, with local[1], 16 partitions, 10,000-row batches and an 8 GiB JVM heap. Spark generated the rows and the selected Arrow serializer populated the cache; native execution was disabled for both versions. Timing included row generation, Arrow conversion, statistics, compression and storage. Each cache was unpersisted synchronously outside the timer. Four separate JVMs ran in before/after/after/before order, each with three warmups and seven measured samples. Pooled medians use 14 samples per implementation. The actual cached-batch classes were verified, and all runs produced exactly 168,804,118 serialized bytes.

Per-JVM materialization medians were 2.017 s / 2.172 s / 1.749 s / 1.826 s. This variation means the isolated-loop gain does not establish an end-to-end cache-build speedup. Measurements used temporary comparison harnesses; the committed materialization benchmark case supports subsequent whole-pipeline benchmarking.

@peterxcli
peterxcli force-pushed the perf/4781-native-cache-statistics branch from 180a154 to 21f1989 Compare September 6, 2026 06:20
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Performance optimizations for native in-memory cache (follow-on to #4591)

1 participant