[improvement](parquet) Fuse dictionary selection planning - #66508
[improvement](parquet) Fuse dictionary selection planning#66508Gabriel39 wants to merge 6 commits into
Conversation
|
Thank you for your contribution to Apache Doris. Please clearly describe your PR:
|
|
run buildall |
|
/review |
|
Codex automated review failed and did not complete. Error: All Codex review accounts are usage-limited; earliest retry is 2026-08-08T03:32:00Z. Please trigger /review again after that time. |
|
run buildall |
|
/review |
|
Codex automated review failed and did not complete. Error: All Codex review accounts are usage-limited; earliest retry is 2026-08-08T03:32:00Z. Please trigger /review again after that time. |
BE UT Coverage ReportIncrement line coverage Increment coverage report
|
What problem does this PR solve?
Predicate-only dictionary filtering currently expands definition levels and row selection into
ColumnSelectVector, then scans that row-oriented state again to rebuild physical selection ranges and the selected NULL layout. Fragmented nullable filters spend significant CPU in these two planning passes before dictionary IDs are decoded.Same-column conjuncts are combined into one dictionary-id bitmap before row decoding. Their incoming row selection therefore remains identity, which previously kept this common shape on the two-pass planner.
What is changed?
ColumnSelectVector.This PR does not change predicate decomposition, supported data types, decode strategy, or merge-read behavior. It is stacked on #66504.
Performance
Release build, 65,536 rows, CPU-pinned runs with 7 repetitions. The table reports median CPU time; lower is better.
The last two negative controls are routed to the existing planner; their sub-1% differences are measurement noise from identical code paths.
Validation