You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Productize CUDA MPS for rank-local MJWarp off-policy training
Summary
CUDA MPS should become an explicit supported runtime mode for MJWarp off-policy training rather than remaining an undocumented host-level setting.
UniLab already requires every rank to keep its learner and collector on the same rank-local GPU. Each side is still a separate process with its own CUDA context, so that required topology is arbitrated by coarse driver time-slicing. MPS lets the rank-local contexts share the GPU and overlap their kernels; it does not change placement or training semantics.
This discussion proposes the product direction, acceptance criteria, execution path, and open decisions. No code is proposed yet.
Evidence
All measurements below were taken on the same host, same code head, one physical GPU, training.no_play=true, training.export_onnx=false, with only the MPS environment changed. Values are the mean of the final 100 iterations. FlashSAC Motion Tracking used 150 iterations; FastSAC G1 Walk Flat used 200.
Host and topology
CPU: AMD Ryzen Threadripper 9980X 64-core
GPU: NVIDIA RTX 6000D
Backend: MJWarp
One measured CUDA rank: learner cuda:0, collector backend cuda:0
Inference ring / replay ingress device: cuda:0
Learner and collector are separate processes, hence separate CUDA contexts.
UniLab deliberately requires rank-local learner/collector placement. The proposal builds on that contract rather than replacing it.
FlashSAC + G1 Motion Tracking + MJWarp
Mode
Repeats
Iteration ms
Steps/s
Learner train ms
Learner wait for collector ms
Default multi-context CUDA
1
26.981
76,155
11.515
13.886
CUDA MPS
3
22.719
90,203
9.843
11.485
Relative improvement:
Iteration time: -15.8%
Throughput: +18.5%
Learner update time: -14.5%
FastSAC / SAC + G1 Walk Flat + MJWarp
Mode
Repeats
Iteration ms
Steps/s
Learner train ms
Learner wait for collector ms
Default multi-context CUDA
1
28.395
73,348
13.470
14.038
CUDA MPS
3
19.287
106,313
15.280
3.338
Relative improvement:
Iteration time: -32.1%
Throughput: +44.9%
Learner wait for collector: -76%
Why this works
The collector is not simply becoming faster. In FastSAC G1 Walk Flat with env_steps_per_sync=4, for example:
Metric
Default
MPS
Collector env step
14.678 ms
18.487 ms
Backend step
5.754 ms
9.270 ms
Replay write
0.215 ms
0.211 ms
Whole iteration
76.638 ms
19.493 ms
Under MPS, the collector step is slightly slower because learner kernels now genuinely compete with it on the GPU. The large win is end-to-end because learner work overlaps collector work instead of waiting through multi-context time-slicing.
This is the key distinction from the rejected CPU-priority approach:
Lever
Scope
Effect on two CUDA contexts
Linux nice
OS CPU scheduling
None
Linux CPU affinity
OS CPU placement
None
torch.cuda.Stream(priority=...)
One CUDA context
Does not arbitrate between learner and collector processes
CUDA MPS
GPU context sharing and kernel overlap
Directly addresses the issue
The OS-priority experiment was closed without merge because it did not match the bottleneck: the measured gain on motion tracking was noise-level, G1 Walk Flat regressed by about 3%, and CPU was not saturated.
Scope boundary
Two runtime facts are already fixed by design and are not under discussion:
Every rank keeps its learner and collector on the same rank-local GPU.
env_steps_per_sync is a training-semantics setting: it changes the number of collector ticks consumed per learner synchronization and therefore the algorithm’s update boundary.
MPS productization concerns only GPU execution sharing for the required rank-local topology. It does not introduce cross-GPU learner/collector placement and does not alter env_steps_per_sync.
Non-goals
Do not add generic learner/buffer/collector Linux nice or affinity configuration.
Do not replace the existing off-policy inference/replay protocol.
Do not silently force MPS for all users; it is a host and deployment dependent mode.
Do not modify env_steps_per_sync; it changes the learner update boundary and is already excluded by the scope boundary above.
Do not claim general support across all backends from this evidence alone. The measured cases are MJWarp FlashSAC G1 Motion Tracking and MJWarp FastSAC G1 Walk Flat.
Product direction
Principle
UniLab should expose an explicit runtime execution mode, not an implicit environment side channel.
Users should not need to know that MJWarp off-policy training is multi-process, that learner and collector have separate CUDA contexts, or that this topology can be improved by MPS. The owner configuration should describe the desired execution mode; UniLab should validate the host state and record the effective result in run evidence.
Proposed owner configuration
A single selection at the training owner level:
training:
cuda_process_sharing: null # null | mps
Semantics:
null: current behavior; the default remains unchanged.
mps: request shared GPU execution through CUDA MPS for the rank-local training process tree.
This is intentionally not a standalone backend switch and does not replace training.sim_backend. It is an execution-sharing mode for the configured runtime.
Requested behavior
When cuda_process_sharing: mps is selected:
Validate that the platform can support the mode:
Linux;
CUDA learner/collector;
same physical GPU for learner and collector;
CUDA_MPS_PIPE_DIRECTORY points to a live control pipe;
an MPS control daemon is reachable.
Validate rank-local learner/collector placement for every rank.
Fail closed before learner/collector construction if validation fails.
Emit actionable diagnostics, including the missing prerequisite and the exact command or service needed to start MPS.
Record the effective mode and server evidence in runtime_manifest.
Record the configured mode in run_config.json.
Honor an explicit request: there is no silent fallback to non-MPS execution.
validate rank-local learner/collector placement on the same physical GPU;
probe the MPS control pipe;
query the server list when available;
return structured evidence;
never silently repair the host;
never start/stop a system service itself.
Tests should use fake filesystem/process results and cover:
disabled default;
enabled and valid;
non-Linux rejection;
CPU learner rejection;
learner/collector GPU mismatch rejection;
missing pipe rejection;
control daemon unreachable rejection;
requested mps with no silent fallback.
Exit criteria:
all unit tests pass;
no environment construction occurs on invalid request;
error messages name the prerequisite and remediation.
Phase 2 - Owner config and builder wiring
Add the config field to the relevant off-policy owner YAML and resolve it in the shared off-policy assembly path before constructing the learner or collector.
The probe must run before:
env probe construction;
learner construction;
collector spawn.
Record the evidence into the runtime manifest supplied by the runner.
Tests should cover:
default owner value is null;
mps is resolved and passed to the builder path;
invalid topology fails before any env factory call;
run config snapshots the field;
runtime manifest contains structured evidence.
Exit criteria:
focused off-policy config and builder tests pass;
an invalid mps request cannot reach env creation.
Phase 3 - Short-run benchmark gate
Before merging Phase 2, rerun the short benchmark matrix on one reference host.
Required cases:
FlashSAC / G1 Motion Tracking / MJWarp
FastSAC / G1 Walk Flat / MJWarp
For each case:
non-MPS baseline;
explicit MPS mode.
Suggested shape:
300 iterations minimum;
ignore the first 50;
report final 100;
at least two repeats per arm.
Acceptance thresholds:
FastSAC G1 Walk Flat: at least 20% throughput improvement.
FlashSAC G1 Motion Tracking: at least 10% throughput improvement.
No completed-run status regression.
No replay ingress backpressure regression.
No NaN guard regression.
MPS must be explicitly observed in the runtime manifest.
Exit criteria:
benchmark table attached to the PR;
commands and environment recorded;
both thresholds met;
failures investigated rather than waived.
Phase 4 - Production documentation and runbook
Document:
supported platform and topology;
how to start and stop a control daemon;
required environment variables;
how UniLab validates the mode;
expected manifest fields;
supported tasks/algorithms;
unsupported scenarios;
known caveats.
Provide both English and Chinese pages or sections, following the repository's docs policy.
Only after explicit MPS mode is stable, investigate CUDA_MPS_ACTIVE_THREAD_PERCENTAGE.
This would allow SM resource partitioning between clients. It should be treated as a separate contract because it changes the performance tradeoff:
throughput;
collector latency;
learner update latency;
multi-tenant fairness.
Do not include it in the initial PR.
Phase 6 - Optional broader backend validation
After MJWarp is supported, evaluate whether other CUDA tensor backends benefit:
Genesis;
Newton;
IsaacGym;
IsaacSim.
This should be evidence-driven. Do not expand the support matrix without benchmark data.
Explicitly rejected alternative: CPU priority configuration
We should not productize learner_scheduling, buffer_scheduling, and collector_scheduling based on OS nice or affinity.
Reasons:
The measured gain on FlashSAC Motion Tracking was noise-level.
FastSAC G1 Walk Flat regressed by about 3%.
CPU was not saturated on the reference host.
Linux nice does not affect CUDA kernel arbitration.
Buffer and learner host-task boundaries were not sufficiently isolated.
The manifest reported requested deltas rather than final OS priority values.
The public contract complexity was not justified by measurable benefit.
The correct scheduling lever for this topology is GPU context sharing, not OS process priority.
Open questions
Daemon ownership.
Should UniLab only validate an existing MPS daemon, or should optional tooling start one under an explicitly user-owned pipe/log directory?
Default value.
Should null mean "do not use MPS", or should there be a third value such as auto that uses MPS when available?
Recommendation: ship null | mps first. Add auto only after a reliable host policy is defined.
DP behavior.
Every rank already keeps learner and collector on one rank-local GPU. mps applies to that rank-local topology; non-rank-local placement remains invalid.
Container policy.
How should the mode interact with Docker, Slurm, Kubernetes, or restricted runners where MPS control is unavailable?
Recommendation: fail closed with a clear diagnostic when explicitly requested.
Metric label.
Should throughput evidence record whether MPS was active in a canonical scalar or only in the runtime manifest?
Recommendation: runtime manifest only initially.
Scope of algorithms.
Should this be shared off-policy runtime behavior or only MJWarp owners?
Recommendation: shared off-policy assembly, but backend validation restricts effective support to MJWarp initially.
CI.
Can CI validate anything beyond the fail-closed path without an MPS-enabled GPU runner?
Recommendation: CI covers validation logic with fakes; the performance gate remains a recorded manual/reference-host benchmark.
Suggested acceptance definition
The feature is complete when:
A user can request MPS through one documented owner setting.
Invalid requests fail closed before environment construction.
Valid runs record structured evidence in runtime_manifest.
The two reference MJWarp tasks show at least the benchmark gains above.
Default behavior is unchanged when the setting is omitted.
English and Chinese documentation explain setup and limitations.
No unrelated backend support claim is added.
No CPU priority configuration is reintroduced.
Reproduction sketch
Start an MPS control daemon in an explicitly owned directory:
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Productize CUDA MPS for rank-local MJWarp off-policy training
Summary
CUDA MPS should become an explicit supported runtime mode for MJWarp off-policy training rather than remaining an undocumented host-level setting.
UniLab already requires every rank to keep its learner and collector on the same rank-local GPU. Each side is still a separate process with its own CUDA context, so that required topology is arbitrated by coarse driver time-slicing. MPS lets the rank-local contexts share the GPU and overlap their kernels; it does not change placement or training semantics.
This discussion proposes the product direction, acceptance criteria, execution path, and open decisions. No code is proposed yet.
Evidence
All measurements below were taken on the same host, same code head, one physical GPU,
training.no_play=true,training.export_onnx=false, with only the MPS environment changed. Values are the mean of the final 100 iterations. FlashSAC Motion Tracking used 150 iterations; FastSAC G1 Walk Flat used 200.Host and topology
cuda:0, collector backendcuda:0cuda:0UniLab deliberately requires rank-local learner/collector placement. The proposal builds on that contract rather than replacing it.
FlashSAC + G1 Motion Tracking + MJWarp
Relative improvement:
FastSAC / SAC + G1 Walk Flat + MJWarp
Relative improvement:
Why this works
The collector is not simply becoming faster. In FastSAC G1 Walk Flat with
env_steps_per_sync=4, for example:Under MPS, the collector step is slightly slower because learner kernels now genuinely compete with it on the GPU. The large win is end-to-end because learner work overlaps collector work instead of waiting through multi-context time-slicing.
This is the key distinction from the rejected CPU-priority approach:
torch.cuda.Stream(priority=...)The OS-priority experiment was closed without merge because it did not match the bottleneck: the measured gain on motion tracking was noise-level, G1 Walk Flat regressed by about 3%, and CPU was not saturated.
Scope boundary
Two runtime facts are already fixed by design and are not under discussion:
env_steps_per_syncis a training-semantics setting: it changes the number of collector ticks consumed per learner synchronization and therefore the algorithm’s update boundary.MPS productization concerns only GPU execution sharing for the required rank-local topology. It does not introduce cross-GPU learner/collector placement and does not alter
env_steps_per_sync.Non-goals
env_steps_per_sync; it changes the learner update boundary and is already excluded by the scope boundary above.Product direction
Principle
UniLab should expose an explicit runtime execution mode, not an implicit environment side channel.
Users should not need to know that MJWarp off-policy training is multi-process, that learner and collector have separate CUDA contexts, or that this topology can be improved by MPS. The owner configuration should describe the desired execution mode; UniLab should validate the host state and record the effective result in run evidence.
Proposed owner configuration
A single selection at the training owner level:
Semantics:
null: current behavior; the default remains unchanged.mps: request shared GPU execution through CUDA MPS for the rank-local training process tree.This is intentionally not a standalone backend switch and does not replace
training.sim_backend. It is an execution-sharing mode for the configured runtime.Requested behavior
When
cuda_process_sharing: mpsis selected:CUDA_MPS_PIPE_DIRECTORYpoints to a live control pipe;runtime_manifest.run_config.json.Runtime evidence
At minimum, the manifest should record:
{ "cuda_process_sharing": { "configured": "mps", "effective": "mps", "server_pid": 2768293, "control_pipe": "/tmp/mps-test/control", "learner_device": "cuda:0", "collector_device": "cuda:0", "validated": true } }If validation fails, there is no run; the failure should identify the first unmet prerequisite.
Documentation direction
Document the mode in the tensor-runtime production guide:
null;env_steps_per_syncremains a separate training-semantics owner setting and is not modified bycuda_process_sharing.Execution path
The work is split into reviewable phases. Each phase has concrete exit criteria; the next phase starts only after those criteria are met.
Phase 0 - Decision and scope record
Outcome: an issue or ADR records that
cuda_process_sharingis a supported runtime mode.The scope statement records UniLab's existing rank-local rule: each rank keeps its learner and collector on the same physical GPU.
Scope decision:
Suggested initial scope:
env_steps_per_syncremains a training-semantics owner decision and is not modified.Exit criteria:
Phase 1 - Probe and diagnostics library
Implement a small owner module, likely under
unilab/training/or an equivalent owner package, exposing:Responsibilities:
Tests should use fake filesystem/process results and cover:
mpswith no silent fallback.Exit criteria:
Phase 2 - Owner config and builder wiring
Add the config field to the relevant off-policy owner YAML and resolve it in the shared off-policy assembly path before constructing the learner or collector.
The probe must run before:
Record the evidence into the runtime manifest supplied by the runner.
Tests should cover:
null;mpsis resolved and passed to the builder path;Exit criteria:
mpsrequest cannot reach env creation.Phase 3 - Short-run benchmark gate
Before merging Phase 2, rerun the short benchmark matrix on one reference host.
Required cases:
For each case:
Suggested shape:
Acceptance thresholds:
Exit criteria:
Phase 4 - Production documentation and runbook
Document:
Provide both English and Chinese pages or sections, following the repository's docs policy.
Exit criteria:
Phase 5 - Optional resource partitioning investigation
Only after explicit MPS mode is stable, investigate
CUDA_MPS_ACTIVE_THREAD_PERCENTAGE.This would allow SM resource partitioning between clients. It should be treated as a separate contract because it changes the performance tradeoff:
Do not include it in the initial PR.
Phase 6 - Optional broader backend validation
After MJWarp is supported, evaluate whether other CUDA tensor backends benefit:
This should be evidence-driven. Do not expand the support matrix without benchmark data.
Explicitly rejected alternative: CPU priority configuration
We should not productize
learner_scheduling,buffer_scheduling, andcollector_schedulingbased on OS nice or affinity.Reasons:
The correct scheduling lever for this topology is GPU context sharing, not OS process priority.
Open questions
Daemon ownership.
Should UniLab only validate an existing MPS daemon, or should optional tooling start one under an explicitly user-owned pipe/log directory?
Default value.
Should
nullmean "do not use MPS", or should there be a third value such asautothat uses MPS when available?Recommendation: ship
null | mpsfirst. Addautoonly after a reliable host policy is defined.DP behavior.
Every rank already keeps learner and collector on one rank-local GPU.
mpsapplies to that rank-local topology; non-rank-local placement remains invalid.Container policy.
How should the mode interact with Docker, Slurm, Kubernetes, or restricted runners where MPS control is unavailable?
Recommendation: fail closed with a clear diagnostic when explicitly requested.
Metric label.
Should throughput evidence record whether MPS was active in a canonical scalar or only in the runtime manifest?
Recommendation: runtime manifest only initially.
Scope of algorithms.
Should this be shared off-policy runtime behavior or only MJWarp owners?
Recommendation: shared off-policy assembly, but backend validation restricts effective support to MJWarp initially.
CI.
Can CI validate anything beyond the fail-closed path without an MPS-enabled GPU runner?
Recommendation: CI covers validation logic with fakes; the performance gate remains a recorded manual/reference-host benchmark.
Suggested acceptance definition
The feature is complete when:
runtime_manifest.Reproduction sketch
Start an MPS control daemon in an explicitly owned directory:
Run the short benchmark:
Extract metrics:
Stop the daemon after testing:
All reactions