You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
CUDA MPS under multi-GPU DP: evidence gaps and design constraints
#2063
#1800 proposes productizing CUDA MPS as training.cuda_process_sharing: null | mps and is now explicitly scoped to single-host, single-rank training (one rank = one learner process + one collector process on one physical GPU). Its measured evidence is single-GPU only.
This discussion tracks what changes when the same mode is requested under multi-GPU data-parallel (DP) training. It records the gaps and constraints so that:
DP support is claimed only after DP-specific evidence exists, per the evidence-driven rule both discussions share.
No code is proposed here either.
What the DP runtime actually looks like today
Facts from the current runtime (uni_rl.ipc + UniLab train_offpolicy path), because several of the gaps only make sense against them:
Launch: off-policy DP is not torchrun. Rank 0 spawns ranks 1..N−1 as subprocesses re-running the same entrypoint, each with exactly one entry in CUDA_VISIBLE_DEVICES (index, UUID, or MIG token); every in-process consumer uses rank-local cuda:0 (uni_rl/ipc/dp_launcher.py, DpRankSupervisor).
Cross-rank traffic is learner-side NCCL: one blocking dist.all_reduce(SUM) per optimizer step (uni_rl/ipc/dp_sync.py, DpParameterSync), plus a startup weight broadcast from rank 0. NCCL P2P/SHM are disabled by default (NCCL_P2P_DISABLE=1, NCCL_SHM_DISABLE=1); transport is TCP loopback. Intra-rank learner↔collector weight sync is POSIX shared memory, not NCCL.
Startup rendezvous: ranks join a FileStore at <log_dir>/.dp_rendezvous with a 120 s timeout. The rank-0 weight broadcast happens before any collector is spawned, and the supervisor watchdog kills the whole run if any rank exits non-zero. This is the de facto cross-rank startup barrier.
Memory guards: two fail-closed torch.cuda.mem_get_info checks (0.8 headroom threshold) run per rank before replay/ring allocation (double_buffer_runner.py CUDA preflight, gpu_resident.py).
Env propagation: rank env is fixed at Popen(env=...) time; spawn collectors inherit the rank process env. CUDA_MPS_PIPE_DIRECTORY can ride these seams — but it must be set before the first CUDA context in every process.
Gaps and open problems for DP
1. NCCL + MPS interaction is untested
#1800's evidence has no NCCL in the loop. Under DP, the learner process runs both its training kernels and the per-step blocking all-reduce on the same GPU that MPS now shares with the collector. Two distinct risks:
NCCL under MPS has a history of performance/transport caveats; with P2P/SHM already disabled, the remaining paths need to be validated under MPS, not assumed.
Straggler amplification: the all-reduce is a hard synchronization point. Under MPS, learner and collector genuinely compete inside each rank, so per-rank iteration jitter can grow relative to time-sliced execution — and blocking all-reduce propagates the slowest rank's jitter to every rank. The +18%/+45% single-rank gains do not automatically survive a world_size>1 sync boundary.
A DP support claim needs its own benchmark gate (world_size ≥ 2, NCCL active, MPS on/off arms), separate from #1800's single-GPU Phase 3 gate.
2. Fail-closed semantics must be cross-rank consistent and atomic
#1800 defines fail-closed for one rank. Under DP that is insufficient:
All ranks must agree on the configured mode; a config that differs across ranks is a run-level error, not a per-rank one.
If one rank fails MPS validation after others have passed, the run must abort cleanly — ranks already inside the FileStore rendezvous must not hang until the 120 s timeout.
Natural anchor: run the probe in the existing startup/dp_init phase and broadcast the validation verdict from rank 0 (same channel as the existing weight broadcast), before any collector spawn. The supervisor watchdog already escalates a dead rank to a full-run abort; validation failure should use that path deliberately rather than accidentally.
3. Per-host daemon vs per-rank validation layering
An MPS control daemon is scoped per (user, host, pipe directory) and serves all visible GPUs. With N ranks on one host, one daemon sees 2N+ client processes (learner + collector per rank, plus NCCL communicator contexts). Validation therefore splits into two layers:
Host level (check once): daemon reachable, pipe directory live, server list queryable.
Rank level (check per rank): learner and collector resolve to the same physical GPU.
#1800's manifest example is a single object; under DP the evidence must become one entry per rank sharing one host-level server record. #1800 should keep its evidence shape per-rank so this extension is additive rather than a schema change.
4. "Same physical GPU" must be UUID-based
Each rank's processes all report cuda:0 because the supervisor masks CUDA_VISIBLE_DEVICES down to one entry. Ordinal comparison proves nothing across ranks; the rank-local check must compare the resolved device UUID/PCI id of learner vs collector (the supervisor's single-entry contract already guarantees one candidate). This is mostly an implementation note for the Phase 1 probe in #1800, but it only becomes load-bearing under DP.
5. Daemon ownership and lifecycle under concurrent runs
#1800 Open Question 1 (validate-only vs tooling-started daemon) is sharper with multiple GPUs and multiple concurrent training runs on one host:
Two independent DP runs sharing one daemon: who is allowed to echo quit? A run that stops the shared daemon breaks the other run.
Conversely, per-run daemons need distinct pipe/log directories, and cleanup on abnormal exit (supervisor SIGKILL path) leaves orphan daemons holding GPU memory.
Recommendation candidate: UniLab validates only; daemon lifecycle stays host-owned, documented per deployment (bare metal, Slurm, container). DP raises the cost of getting this wrong but does not change the recommendation.
6. Memory-budget guards assume an exclusive-ish GPU
The 0.8-threshold mem_get_info preflights see the shared free-memory view under MPS: learner allocations are visible to the collector-side guard and vice versa. Under DP this is a per-rank concern (each rank has its own GPU), so it is not worse than single-rank — but the guards were calibrated on non-MPS behavior and may false-positive or false-negative once MPS merges the contexts' footprints. Also note CUDA_MPS_DEVICE_MEM_LIMIT / pinned-memory limits are daemon-global: they cannot be tuned per rank without per-rank daemons.
7. CUDA_MPS_ACTIVE_THREAD_PERCENTAGE conflicts with per-rank tuning
#1800 Phase 5 defers SM partitioning. Under DP the tension is structural: the percentage is a property of the MPS server, i.e. global across all clients on all GPUs it serves. Per-rank partitioning (e.g. different collector shares on different GPUs) requires multiple daemons with distinct pipe directories — a much heavier deployment contract. Any Phase 5 design should state explicitly whether it is single-daemon-global or multi-daemon, before DP support is claimed.
8. Multi-node DP is out of scope for this mode
MPS is strictly host-local. Externally launched multi-node DP (the UNILAB_DP_EXTERNAL direction explored on an unmerged branch) cannot share one daemon across nodes, and "same physical GPU" is meaningless across hosts. The mode's support statement should say single host even after DP lands; multi-node would be per-node daemon validation with the same cross-rank rules applied within each node — a separate evidence exercise.
9. Prerequisite: the DP launch contract itself is mid-refactor
On the current development head, rank 0's launch path rejects a multi-entry CUDA_VISIBLE_DEVICES (rank_local_cuda_device raises) while the user docs still show a multi-entry parent mask as the DP launch recipe, and selected_visible_entries is imported but unused in train_offpolicy.py. The rank-0 device-pinning contract needs to settle before MPS topology validation is layered onto it; otherwise DP + MPS validation would be specified against a moving target.
What #1800 should keep doing to stay DP-compatible
These are constraints on #1800's single-rank design so this follow-up remains additive:
Keep cuda_process_sharing semantics per-rank ("share GPU execution for the rank-local training process tree"), never host-global.
Keep the probe signature rank-local (requested, learner_device, collector_device); DP orchestration is a caller concern, not a probe concern.
Keep manifest evidence shaped per-rank so DP can record a list of per-rank entries plus one host-level server record without breaking single-rank readers.
Keep "fail closed before env/learner/collector construction" as the unit contract; the DP layer only adds cross-rank agreement on top.
Claim no DP support in docs/support matrix until the DP benchmark gate in this discussion passes.
Suggested DP acceptance definition (for later, not #1800)
DP support is claimable when:
Cross-rank config agreement is validated; any rank's validation failure aborts the run before any collector spawn, with a diagnostic naming the failing rank and prerequisite.
Manifest records per-rank evidence plus host-level server evidence.
A world_size ≥ 2 benchmark on the two reference MJWarp tasks (FlashSAC G1 Motion Tracking, FastSAC G1 Walk Flat) shows no throughput regression vs non-MPS DP, with per-rank jitter reported; thresholds set after a first measurement pass rather than copied from the single-rank gate.
NCCL all-reduce correctness and latency under MPS are explicitly measured, not inferred.
Docs state single-host scope and daemon lifecycle ownership for concurrent runs.
Open questions
Should DP validation broadcast only a boolean verdict, or the full per-rank evidence for rank-0 to aggregate into one manifest?
Is per-rank daemon isolation (distinct pipe dirs per rank) ever worth supporting, or is one daemon per host the only supported topology?
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
CUDA MPS under multi-GPU DP: evidence gaps and design constraints
Relationship to #1800
#1800 proposes productizing CUDA MPS as
training.cuda_process_sharing: null | mpsand is now explicitly scoped to single-host, single-rank training (one rank = one learner process + one collector process on one physical GPU). Its measured evidence is single-GPU only.This discussion tracks what changes when the same mode is requested under multi-GPU data-parallel (DP) training. It records the gaps and constraints so that:
No code is proposed here either.
What the DP runtime actually looks like today
Facts from the current runtime (
uni_rl.ipc+ UniLabtrain_offpolicypath), because several of the gaps only make sense against them:CUDA_VISIBLE_DEVICES(index, UUID, or MIG token); every in-process consumer uses rank-localcuda:0(uni_rl/ipc/dp_launcher.py,DpRankSupervisor).dist.all_reduce(SUM)per optimizer step (uni_rl/ipc/dp_sync.py,DpParameterSync), plus a startup weight broadcast from rank 0. NCCL P2P/SHM are disabled by default (NCCL_P2P_DISABLE=1,NCCL_SHM_DISABLE=1); transport is TCP loopback. Intra-rank learner↔collector weight sync is POSIX shared memory, not NCCL.<log_dir>/.dp_rendezvouswith a 120 s timeout. The rank-0 weight broadcast happens before any collector is spawned, and the supervisor watchdog kills the whole run if any rank exits non-zero. This is the de facto cross-rank startup barrier.torch.cuda.mem_get_infochecks (0.8 headroom threshold) run per rank before replay/ring allocation (double_buffer_runner.pyCUDA preflight,gpu_resident.py).Popen(env=...)time; spawn collectors inherit the rank process env.CUDA_MPS_PIPE_DIRECTORYcan ride these seams — but it must be set before the first CUDA context in every process.Gaps and open problems for DP
1. NCCL + MPS interaction is untested
#1800's evidence has no NCCL in the loop. Under DP, the learner process runs both its training kernels and the per-step blocking all-reduce on the same GPU that MPS now shares with the collector. Two distinct risks:
A DP support claim needs its own benchmark gate (world_size ≥ 2, NCCL active, MPS on/off arms), separate from #1800's single-GPU Phase 3 gate.
2. Fail-closed semantics must be cross-rank consistent and atomic
#1800 defines fail-closed for one rank. Under DP that is insufficient:
startup/dp_initphase and broadcast the validation verdict from rank 0 (same channel as the existing weight broadcast), before any collector spawn. The supervisor watchdog already escalates a dead rank to a full-run abort; validation failure should use that path deliberately rather than accidentally.3. Per-host daemon vs per-rank validation layering
An MPS control daemon is scoped per (user, host, pipe directory) and serves all visible GPUs. With N ranks on one host, one daemon sees 2N+ client processes (learner + collector per rank, plus NCCL communicator contexts). Validation therefore splits into two layers:
#1800's manifest example is a single object; under DP the evidence must become one entry per rank sharing one host-level server record. #1800 should keep its evidence shape per-rank so this extension is additive rather than a schema change.
4. "Same physical GPU" must be UUID-based
Each rank's processes all report
cuda:0because the supervisor masksCUDA_VISIBLE_DEVICESdown to one entry. Ordinal comparison proves nothing across ranks; the rank-local check must compare the resolved device UUID/PCI id of learner vs collector (the supervisor's single-entry contract already guarantees one candidate). This is mostly an implementation note for the Phase 1 probe in #1800, but it only becomes load-bearing under DP.5. Daemon ownership and lifecycle under concurrent runs
#1800 Open Question 1 (validate-only vs tooling-started daemon) is sharper with multiple GPUs and multiple concurrent training runs on one host:
echo quit? A run that stops the shared daemon breaks the other run.Recommendation candidate: UniLab validates only; daemon lifecycle stays host-owned, documented per deployment (bare metal, Slurm, container). DP raises the cost of getting this wrong but does not change the recommendation.
6. Memory-budget guards assume an exclusive-ish GPU
The 0.8-threshold
mem_get_infopreflights see the shared free-memory view under MPS: learner allocations are visible to the collector-side guard and vice versa. Under DP this is a per-rank concern (each rank has its own GPU), so it is not worse than single-rank — but the guards were calibrated on non-MPS behavior and may false-positive or false-negative once MPS merges the contexts' footprints. Also noteCUDA_MPS_DEVICE_MEM_LIMIT/ pinned-memory limits are daemon-global: they cannot be tuned per rank without per-rank daemons.7.
CUDA_MPS_ACTIVE_THREAD_PERCENTAGEconflicts with per-rank tuning#1800 Phase 5 defers SM partitioning. Under DP the tension is structural: the percentage is a property of the MPS server, i.e. global across all clients on all GPUs it serves. Per-rank partitioning (e.g. different collector shares on different GPUs) requires multiple daemons with distinct pipe directories — a much heavier deployment contract. Any Phase 5 design should state explicitly whether it is single-daemon-global or multi-daemon, before DP support is claimed.
8. Multi-node DP is out of scope for this mode
MPS is strictly host-local. Externally launched multi-node DP (the
UNILAB_DP_EXTERNALdirection explored on an unmerged branch) cannot share one daemon across nodes, and "same physical GPU" is meaningless across hosts. The mode's support statement should say single host even after DP lands; multi-node would be per-node daemon validation with the same cross-rank rules applied within each node — a separate evidence exercise.9. Prerequisite: the DP launch contract itself is mid-refactor
On the current development head, rank 0's launch path rejects a multi-entry
CUDA_VISIBLE_DEVICES(rank_local_cuda_deviceraises) while the user docs still show a multi-entry parent mask as the DP launch recipe, andselected_visible_entriesis imported but unused intrain_offpolicy.py. The rank-0 device-pinning contract needs to settle before MPS topology validation is layered onto it; otherwise DP + MPS validation would be specified against a moving target.What #1800 should keep doing to stay DP-compatible
These are constraints on #1800's single-rank design so this follow-up remains additive:
cuda_process_sharingsemantics per-rank ("share GPU execution for the rank-local training process tree"), never host-global.requested, learner_device, collector_device); DP orchestration is a caller concern, not a probe concern.Suggested DP acceptance definition (for later, not #1800)
DP support is claimable when:
Open questions
CUDA_VISIBLE_DEVICES)?All reactions