[spark] Introduce reassign_row_id procedure - #9192
Merged
JingsongLi merged 2 commits intoAug 12, 2026
Merged
Conversation
Spark has no counterpart to Flink's reassign_row_id procedure, so row IDs of a data evolution table cannot be made partition-contiguous from Spark. Add sys.reassign_row_id(table, partitions), delegating to the engine-agnostic DataEvolutionRowIdReassigner in paimon-core.
Contributor
|
+1 |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
Flink provides
reassign_row_idboth as an action (ReassignRowIdAction) and as aprocedure, but Spark has no counterpart. As a result, users running Paimon on Spark
cannot make row IDs of a data evolution table partition-contiguous.
Row IDs are assigned in commit order, so writes that interleave across partitions
leave each partition with a scattered row-id range, which hurts column-store read
efficiency. Reassignment rewrites the row-id ranges so that each partition owns a
contiguous block.
This PR adds the Spark procedure:
Notes on the implementation:
DataEvolutionRowIdReassignerinpaimon-core,which is engine-agnostic and already used by the Flink side, so no reassignment
logic is duplicated here.
partitionsis parsed with the existingSparkProcedureUtils#convertPartitionsToPartitionPredicate, keeping the partitionspec syntax consistent with other Spark procedures; when omitted, all partitions
are reassigned.
row-tracking.enabled = trueanddata-evolution.enabled = true;both are validated by the core reassigner.
resultstring describing whether the reassignmenthappened, and for a skip, why it was skipped.
Docs for the new procedure are added to
docs/docs/spark/procedures.md.Tests
Added
ReassignRowIdProcedureTestinpaimon-spark-ut, covering:become contiguous per partition and the data itself is unchanged;
partitionsfilter, for both a spec that matches no partition (skipped) andone that matches (reassigned);
row-tracking.enabled=trueanddata-evolution.enabled=true;