Skip to content

fix(spark): reject MapType instead of emitting a childless Arrow List#8811

Merged
robert3005 merged 1 commit into
vortex-data:developfrom
jackylee-ch:spark-schema-map-fail-fast
Jul 17, 2026
Merged

fix(spark): reject MapType instead of emitting a childless Arrow List#8811
robert3005 merged 1 commit into
vortex-data:developfrom
jackylee-ch:spark-schema-map-fail-fast

Conversation

@jackylee-ch

Copy link
Copy Markdown
Contributor

SparkToArrowSchema mapped MapType to a childless ArrowType.List, which is an invalid Arrow list (a List must have exactly one child). Drop the branch so MapType falls through to the existing fail-fast UnsupportedOperationException, and add a test.

SparkToArrowSchema mapped MapType to a childless ArrowType.List, which is
an invalid Arrow list (List requires exactly one child). Drop the branch so
MapType falls through to the fail-fast UnsupportedOperationException, and
add a test.

Signed-off-by: jackylee <qcsd2011@gmail.com>
@jackylee-ch jackylee-ch reopened this Jul 17, 2026
@codspeed-hq

codspeed-hq Bot commented Jul 17, 2026

Copy link
Copy Markdown

Merging this PR will improve performance by 11.17%

⚡ 1 improved benchmark
✅ 1669 untouched benchmarks
⏩ 52 skipped benchmarks1

Performance Changes

Mode Benchmark BASE HEAD Efficiency
Simulation true_count_vortex_buffer[128] 580.6 ns 522.2 ns +11.17%

Tip

Curious why this is faster? Comment @codspeedbot explain why this is faster on this PR, or directly use the CodSpeed MCP with your agent.


Comparing jackylee-ch:spark-schema-map-fail-fast (1e47bbc) with develop (3975895)

Open in CodSpeed

Footnotes

  1. 52 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports.

@robert3005 robert3005 added the changelog/fix A bug fix label Jul 17, 2026
@robert3005
robert3005 enabled auto-merge (squash) July 17, 2026 12:53
@robert3005
robert3005 merged commit da76f5b into vortex-data:develop Jul 17, 2026
83 of 84 checks passed
connortsui20 added a commit that referenced this pull request Jul 20, 2026
## Rationale for this change

- Closes: #8807

Three benchmarks stayed flaky after #8742, flipping between the same two
values on PRs that can't affect them. Same root causes and fixes as
#8742:

| Benchmark | Seen flaky on | Why | Fix |
| --- | --- | --- | --- |
| `true_count_vortex_buffer[128]` | ±11.17% on 9 unrelated PRs (#8805,
#8811, #8812, #8820, #8843, #8803, …) | a 128-bit popcount measures
harness overhead and code layout, not the count | drop the 128 size |
| runend `compress[(100000, 4)]` | ±11.9% on #8805, #8750, #8856 |
allocates in the timed region; glibc malloc differs across runner images
| mimalloc as global allocator |
| `cast_decimal` `copy_*[65536]` | identical flags on #8838 and #8724 |
same glibc-malloc cause (512 KB alloc per iteration) | mimalloc as
global allocator |

Left alone: `compact_sliced[(4096, 90)]` (single sighting) and the CUDA
walltime benches (hosted-runner walltime noise, a runner config issue).

The allocator swap shifts every benchmark in the two touched binaries
once — see the comment below. Needs a one-time CodSpeed acknowledgment,
like #8742.

## What changes are included in this PR?

One commit per benchmark; bench files only. Ran `cargo check` + `clippy`
on the three bench targets, smoke-ran the binaries, `cargo +nightly
fmt`.

---------

Signed-off-by: Claude <noreply@anthropic.com>
Co-authored-by: Claude <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

changelog/fix A bug fix

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants