Replies: 1 comment
|
Thanks for opening up this discussion! Yeah I think this could benefit from some design work around the multi-device abstractions, how this impacts device selection as well as tuning, and how we can make this a generic foundation that still supports this initial use case. Ideally this would fit into our long-term ideas around automatically making use of any given hardware topology, that is, given a system with any configuration of GPUs, CPU, memory domains (GPU, unified, RAM), what is the best way to distribute the inference workload of a model on that system for optimal performance. I also think that the placement abstractions should be able to represent sharding, even if the initial implementation only supports pipeline parallelism. Tensor parallelism or even PP+TP should eventually be chosen dynamically based on the discovered topology and automatic placement to maximize throughput. Another consideration here that would support a generic foundation is interconnect modeling as part of the DeviceTopology abstraction. I think some of this can be discovered through device APIs, however it may also require on-device measurement. If the choice is simply "see multiple GPUs -> pipeline them" then this may not be immediately necessary, however to support any kind of well-informed automatic placement it would need to be available in some way. I would also like if we ensure that the abstractions do not couple themselves to CUDA specifically, even if the initial implementation focused on CUDA functionality. Let me know what you think of these ideas and what scope you'd be willing to tackle for this PR, appreciate the thought you've put into this! |
Uh oh!
There was an error while loading. Please reload this page.
TL;DR: I got multi-GPU model sharding working on my machine. I can now run Qwen3.6-35B-A3B Q8 (~40 GB) across two 24 GB RTX 3090s through Magnitude’s normal Chat Completions path, even though it cannot fit on either card individually. I’m wondering if this is a direction worth tidying up and upstreaming, and if so I’d appreciate feedback on the proposed shape before I spend time turning the spike into PRs or if this is irrelevant with the engine rewrite discussion.
Decision requested
I'd like feedback on whether contiguous pipeline placement across multiple GPUs is a direction Magnitude wants to support for models that cannot fit within one device's memory budget.
I have a working proof of concept on two RTX 3090s, including a model that cannot legally fit either card individually. Before cleaning this into upstream PRs, I'd mainly like an architectural yes/no:
This is related to #141/#128, but I'm treating larger-model sharding as a separate architectural question rather than assuming those reports requested this specific feature.
What I found
Magnitude's current single-device fit behavior appears internally consistent: a model is assessed against one allocation domain, including reserve, state, graph resources and workspace.
The limitation is that there is currently no execution strategy that can place different parts of one model in different GPU allocation domains.
I built a bounded two-stage CUDA implementation to test whether that limitation can be removed without treating VRAM as a flat pool.
The execution shape is:
Each stage owns its own:
Only the hidden activation crosses the device boundary.
A logical decode step is accepted only after both stages complete successfully.
There is no unified VRAM allocator.
Larger-than-single-card proof
Hardware:
Target:
Magnitude's independent legal device budgets on this machine were:
Ordinary single-device serving requires:
and returns the existing typed
DoesNotFitresult on either GPU.Weights alone are 38.140 GB, so the model physically cannot reside on either card individually.
The qualified placement was:
Weights are imported directly to their destination GPU. The model is never first materialized in full on GPU0.
The ordinary source-built Chat Completions API then successfully served the model.
Two fresh deterministic requests each generated the same 16 tokens:
Measured API timings:
The PCIe activation boundary measured approximately:
No NVLink or CUDA peer access was required.
Cancellation, fresh recovery, concurrent-request refusal and graceful shutdown were also exercised.
Host-memory inspection plus complete GPU weight accounting support that there is no CPU weight execution or model-weight streaming in this path. Host memory is still used normally for loading and activation staging.
Numerical correctness proof
For the 35B model there is intentionally no fabricated one-GPU control, because the model does not fit one card.
Before trying the large model, I qualified the exact same staged execution approach using Qwen3.5-4B, which does fit one GPU.
For that model, ordinary single-GPU execution and staged two-GPU execution matched bit-for-bit for:
across prompt ingestion and multiple generated tokens.
That was then integrated through Magnitude's existing owner/request/API lifecycle rather than exposed through a separate benchmark server.
Proposed architecture
I don't think the long-term abstraction should be specifically
PairedCuda*.The mechanism is an ordered pipeline, and two GPUs are just the first qualified case.
Conceptually:
The initial runtime can still explicitly require:
until additional configurations have real hardware qualification.
The important part is avoiding a two-device ownership model that would later need to be replaced to support 3/4/8 GPUs.
For N devices the execution remains conceptually:
with persistent state staying on the stage that owns its original model layers.
Heterogeneous GPUs
I also think placement should eventually be based on independent device budgets, not equal layer counts.
For example, a 24 GB + 12 GB configuration should be able to choose an asymmetric split if one exists.
A valid placement requires:
for every device independently.
It should never mean:
After satisfying capacity constraints, Magnitude's existing device-specific optimization/performance estimates could eventually help choose among valid cuts.
That preserves something I think is important about Magnitude: placement decides where blocks execute, while Magnitude's existing tuning machinery still decides how each stage executes efficiently on its actual hardware.
For a heterogeneous pair, each device should therefore keep its own local tuning/preparation results.
More than two GPUs
Machines with 3, 4, 6, or 8 GPUs exist, even if they are much less common than one- and two-GPU systems.
I think the internal representation should accommodate that now without claiming support before it has been qualified.
The natural extension is still an ordered pipeline:
Each stage keeps independent:
The planner should be able to reason about N available CUDA devices and heterogeneous capacities/speeds, but the first runtime implementation can still reject
stages.len() != 2.That lets us test placement logic for configurations such as:
without claiming those configurations execute successfully.
Longer-term, the planner should prefer the smallest number of devices that can legally fit the requested model/profile, then choose among legal placements using Magnitude's device-specific performance estimates and transfer/topology costs.
For a single autoregressive request, pipeline placement primarily increases capacity, not necessarily speed. Spreading a model across more GPUs than necessary may reduce latency performance because additional stage boundaries add synchronization and transfer.
Tensor parallelism and expert parallelism are different execution strategies and are intentionally out of scope for this proposal.
What I propose upstreaming
I do not want to submit the experimental branch/history as one giant PR.
I'd rather upstream this in small reviewable pieces.
PR 1 — execution primitive
Scope:
This PR would include:
The point of this PR is only to establish a correct execution mechanism.
PR 2 — certified placement planner
The planner would consume:
and return an immutable certified stage placement.
Conceptually:
The planner can be N-device-capable internally from the start, while the product executor initially accepts only the already-qualified two-stage CUDA case.
A placement is legal only if every stage independently fits its device.
The same certified plan should eventually be consumed by:
rather than assessment and execution independently reconstructing placement decisions.
PR 3 — automatic selection and catalog presentation
Only after the execution primitive and placement planner are accepted:
The UI/catalog could then say something like:
rather than presenting the machine as one fictitious larger GPU.
Explicit non-goals for the first implementation
This proposal does not attempt:
Those seem separable from the basic pipeline-placement primitive.
Why I'm asking before opening the PR
The experiment is complete enough that I think the mechanism is demonstrated, but I'd rather align on the abstraction before spending time reducing the spike into an upstream-quality patch.
The main question is therefore not "does this work on my two 3090s?" — it does.
The question is whether device-local pipeline stages, represented generally but initially qualified for two CUDA GPUs, fit the direction you want Magnitude's inference engine to take.
If yes, I'll rebase onto current
main/0.2.4, reduce the implementation to the first bounded execution PR above, and keep the larger experimental evidence outside the review diff.All reactions