perf(rdma): isolate SG writes and restore batch depth - #277
Merged
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
tp_rank, which is always 0 with DP attentionProblem
Multi-SGE CUDA writes shared the normal data connection pool. With batch depth greater than one, mlx5 could overlap writes from distinct registered CUDA ranges on one QP and report protection/completion errors. The HiCache connector worked around this by effectively running at depth one, which serialized batch GETs and left SSD/RDMA throughput underutilized.
The previous
rail_affinityimplementation also usedtp_rank. Under DP attention every process hastp_rank=0, collapsing all workers onto the same rail.Design
Lane::kSgDataowns a separate idle pool. Its QPs negotiate depth one, butCacheFromMultistill fans a batch across independent connections according to the configured transport depth. This preserves the one-CUDA-range-per-QP invariant without serializing the whole batch.When explicitly enabled, connector rail affinity now maps
pcp_rank % rail_countand disables additional NUMA rail selection for that process.Validation
Tested on one 8x B200 development node with eight active RDMA rails and two NVMe data devices. The inference container used unlimited memlock so the device pools could be registered once.
Runtime prerequisite: deployments using explicit large memory-pool registration must set an adequate
memlocklimit (for example, container--ulimit memlock=-1:-1).