Repository navigation
Release date: 2026-09-29
Compatible PyTorch: >=2.10,<2.11
GitHub: flagos-ai/PyTorch-Plugin-FL
torch-fl is a PyTorch device plugin for the FlagOS software stack. It exposes a
single flagos device that routes each operator, per-call, among native vendor
kernels, portable FlagGems kernels, CUDA-compatibility boxing, and an explicit
CPU fallback — so the same PyTorch program runs unchanged across accelerators
from nine different vendors.
Starting with this release, the torch-fl version tracks the PyTorch minor line
it binds to: wheels are named torch_fl-2.10.0+<sdk> (e.g.
torch_fl-2.10.0+cuda13.3, torch_fl-2.10.0+dtk2604,
torch_fl-2.10.0+cann9.0.0, …). One tag, one artifact per platform.
Highlights
- One
flagosdevice, nine accelerator platforms. NVIDIA (A100 / H800),
Ascend (910C), Hygon (BW1000), MetaX (C550), T-Head (ZW810E), Moore Threads
(S5000), Enflame (S60), Kunlunxin (P800), and D-Robotics-class BPU (s600)
are all reachable through the samedevice="flagos"API. No workload changes
when switching hardware. - Distributed on every mainstream card. Basic collectives, DDP, and FSDP2
are validated across NVIDIA, Ascend, Hygon, MetaX, T-Head, Moore Threads,
Enflame, and Kunlunxin — not just one or two reference boards. torch.compilefirst-class.flagosis registered as an inductor GPU
device, backed by FlagTree on NVIDIA / Ascend / Hygon / MetaX / T-Head /
Moore Threads and by vendor Triton on Enflame; the BPU path compiles through
the on-board BPU compiler.- Full profiler coverage. PrivateUse1 flows, device time, and kernel
metadata reach parity withtorch.cuda.profiler, with per-vendor tracers
(CUPTI, ROCtracer, MUPTI, MSPTI, TOPSPTI, MCPTI). Kunlunxin P800 is the
current gap. - Three new cross-platform capabilities.
- A new-chip onboarding skill that stands up a new accelerator backend
end-to-end; - FP8 software emulation on the boxing path (no hardware FP8 required);
- Automatic operator-selection tuning that picks the best kernel
implementation per operator per shape.
- A new-chip onboarding skill that stands up a new accelerator backend
- Self-describing wheels. Every wheel ships
torch_fl/compatibility.json
recording platform, kernel sets, bundled libtorch, build-time PyTorch/ABI,
and the pinned FlagTree/FlagGems/FlagCX builds. A new
torch-fl-preflightCLI inspects a wheel (or an installed env) without
importingtorch_fland can emit the release compatibility table as
Markdown.
Platform matrix
Numbers below are the validated operator counts for this release, taken from
the internal platform capability sheet. "FlagGems-Python" shows the route
counts before → after this release.
| Vendor | Board | Runtime | Vendor ops | FlagGems-Python | FlagGems-C++ | FlagCX | Distributed | Profiler | AMP | RNG | torch.compile |
|---|---|---|---|---|---|---|---|---|---|---|---|
| NVIDIA | A100 | ✅ | 2033 | 390 → 424 | 18 | ✅ | basic + DDP + FSDP2 | ✅ | ✅ | ✅ | FlagTree |
| NVIDIA | H800 | ✅ | (shares A100 config) | ||||||||
| Ascend | 910C | ✅ | 354 | 221 → 225 | 0 | ✅ | basic + DDP + FSDP2 | ✅ | ✅ | ✅ | FlagTree |
| Hygon | BW1000 | ✅ | 2033 | 388 → 470 | 0 | ✅ | basic + DDP + FSDP2 | ✅ | ✅ | ✅ | FlagTree |
| MetaX | C550 | ✅ | 2033 | 377 → 443 | 17 | ✅ | basic + DDP + FSDP2 | ✅ | ✅ | ✅ | FlagTree |
| T-Head | ZW810E | ✅ | 2033 | 390 → 435 | 0 | ✅ | basic + DDP + FSDP2 | ✅ | ✅ | ✅ | FlagTree |
| Moore Threads | S5000 | ✅ | 105 | 391 → 468 | 0 | ✅ | basic + DDP + FSDP2 | ✅ | ✅ | ✅ | FlagTree |
| Enflame | S60 | ✅ | 224 | 281 → 257 | 0 | ✅ | basic + DDP + FSDP2 | ✅ | ✅ | ✅ | vendor Triton |
| Kunlunxin | P800 | ✅ | 2022 | 380 | 0 | ✅ | basic + DDP + FSDP2 | ❌ | ✅ | ❌ | ❌ |
| BPU (edge) | s600 | ✅ | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ | FX → BPU Compiler |
Notes:
- H800 reuses the A100 wheel and route table; it is listed separately only
because it is a shipping SKU. - Enflame S60's FlagGems-Python count tightened from 281 to 257: routes
that were not shape-validated were dropped rather than advertised. - Kunlunxin P800 ships with runtime, vendor ops, FlagCX, AMP, and the
distributed path; profiler and RNG integration, plustorch.compile, are
on the roadmap. - BPU s600 is an edge target: eager operators run on CPU, and the only
execution path istorch.compile(backend="...")through the on-board BPU
compiler.
What's changed since 0.1.0
Device model & routing
- A single
flagosdevice (PrivateUse1) replaces per-vendor device names;
device="cuda"is aliased ontoflagosonly after asking PyTorch, not a
redirected probe, andtorch.cuda.devicecontext objects are accepted. - Per-operator dispatch chooses between native vendor kernels,
FlagGems C++ (kFlagOs) and Python paths, CUDA boxing via an external
libtorch_cuda.so, and CPU fallback. Routing is per-op, not per-device. - New optional operator library TileOPs is integrated.
- Sparse storage on
flagos: construct COO tensors and serve CSR/CSC/BSR/BSC
tensors onSparseCsrPrivateUse1. DataParallelworks on theflagosbackend.uint16/uint32/uint64dtype casts are supported.- Automatic operator-selection tuning picks the best kernel implementation
per operator per shape.
Backends
- Ascend 910C — 354 routed vendor ops (ACLNN), 221 → 225 FlagGems routes;
native ACL stream/event semantics for FSDP2; boolean SDPA-mask semantics and
indexboolean masks preserved; finite additive attention bias carried
through the Ascend SDPA realShift path; nearest-exact upsampling kept
on-device; Qwen3-0.6B training measured at 0.82×torch_npuon real 910C. - Hygon BW1000 — runs on upstream PyTorch rather than DTK's fork;
2033 vendor ops via CUDA boxing over hipified DTK; low-precision software
GEMM path; FlagGems routes grew 388 → 470. - MetaX C550 — 2033 vendor ops via
cu-bridge; FSDP2 full feature parity
with Qwen3 training match; real deviceEvent+pin_memoryin
_to_copy; scaled-dot-product-attention routed to FlagGems; FP8/FP4
software emulation on the boxing path. - T-Head ZW810E — same CUDA-boxing path as NVIDIA against the PPU
CUDA-13-compatible SDK, bundling its own libtorch; FlagGems routes grew
390 → 435; AMP integrated. - Moore Threads S5000 — native
mudnnbackend with 105 vendor kernels
and 391 → 468 FlagGems routes; its own stream class with explicit vendor
dispatch; fp64addmm/baddbmmrouted to mudnn; native RNG unified with
FlagGems. - Enflame S60 — native
libtopsaten.sobackend with 224 vendor kernels;
FlagGems Triton on S60; int64 pointwise routes escaped to the vendor kernel
before the compiler aborts; Qwen-Image-2512 cut to 7.1 s/it. - Kunlunxin P800 — 2022 vendor ops, 380 FlagGems-Python routes, FlagCX
and AMP in; profiler / RNG /torch.compilefollow-up. - NVIDIA A100/H800 — CUDA boxing over external
libtorch_cuda.so;
FlagGems-first routing with FlagTree 3.6 integration; FlagGems-C++ dispatch
path (18 routes); self-contained wheels that run on stocktorch+cpu. - BPU s600 — edge target; eager on CPU,
torch.compilegraph path via
hbdk4 / on-board BPU compiler.
torch.compile
flagosregistered as a first-class inductor GPU device.- FlagTree backends validated on NVIDIA, Ascend (via triton-ascend), Hygon,
MetaX, T-Head, and Moore Threads; Enflame uses its vendor Triton. - A parameterized
set_env.shentrypoint and per-platform hooks make the
compile env reproducible.
Distributed
- DDP and FSDP2 validated on every mainstream board in the matrix.
- Collectives: MUSA via FlagCX+MCCL, GCU routed through FlagCX, Hygon via
RCCL on DTK, Ascend HCCL fallback, NCCL-shaped route elsewhere. - gloo requests answered on the
flagosbackend; collectives staged over host
memory. - Cross-device copies ordered against peer-device D2D copies and staged in
MudnnCopy.
AMP, autocast & RNG
PrivateUse1autocast and dtype support added; FP16/BF16 autocast +
GradScalermeasured on every mainstream board.- FP8/FP4 software emulation on the MetaX boxing path — usable without
FP8 hardware. - RNG unified across native and FlagGems paths via shared generator injection.
Profiler
- Torch-CUPTI parity for PrivateUse1: flows, device time, kernel metadata.
- Per-vendor tracer integration: ROCtracer (Hygon), MUPTI (Moore Threads),
MSPTI (Ascend), TOPSPTI (Enflame), MCPTI (MetaX), CUPTI on NVIDIA / T-Head. - The CUPTI runtime callback-id table is generated rather than hand-maintained.
Performance
- Qwen-Image-2512 on Enflame S60 cut to 7.1 s/it (SDPA route, GEMM operand
handling, five host round trips removed). - Hygon: FlagGems launch, layout and autotune costs cut behind the
Qwen-Image-2.1 gap; Qwen-Image rotation registered from the plugin; inert
masks dropped. - MetaX:
scaled_dot_product_attentionrouted to FlagGems; fused SDPA on the
boxing path restored via a composite override. - FlagGems SDPA route widened in
head_dim, dtype, and query-length bounds. - CPU-fallback inference reached 96 % of CUDA baseline throughput by
removing PrivateUse1 training overhead.
Benchmark: Qwen-Image-2.1 across vendors
Single-card, bf16, one 1024×1024 image, 40 denoising steps; median latency
per image, measured out of the box:
| Board | s/image | Throughput vs. H100 native | Peak memory |
|---|---|---|---|
| NVIDIA H100, native | 6.52 | 100 % | 36.8 GiB |
| NVIDIA H100, FlagOS | 10.67 | 61 % | 37.8 GiB |
| Moore Threads S5000 | 14.87 | 44 % | 50.8 GiB |
| Hygon BW1000 | 23.59 | 28 % | 36.8 GiB |
| Ascend 910C | 24.94 | 26 % | 43.3 GiB |
| MetaX C550 | 31.61 | 21 % | 37.0 GiB |
| T-Head ZW810E | 37.44 | 17 % | 36.8 GiB |
| Enflame S60 | 58.04 | 11 % | 37.1 GiB |
- Image quality is aligned, not just latency: T2I-100 CLIP-Score stays
within ±1.2 % of H100 native and COCO CLIP-Score within ±1.6 % across the
boards above; FID differences are in the low single-digit percent range. - CPU-only Arm path (Mac M5 Pro, W8A8, same 1024² / 40-step workload):
FlagOS finishes in 785 s vs. 1365 s for stock PyTorch CPU — 1.74×
faster. - Kunlunxin P800 has not yet been benchmarked on this workload.
Packaging & release engineering
- Wheels are self-contained and run on a stock
torch+cpuinstall; the
installed PyTorch is left immutable against vendor libtorch. - Wheel local segment names the SDK it was built against; one interpreter per
platform (3.12 on CUDA/GCU/MetaX/PPU, 3.10 on DCU/MUSA, 3.11 on Ascend). torch_fl/compatibility.jsonis published and validated; the
torch-fl-preflightCLI checks dependencies, warns on build-env drift, and
fails a release table that is missing SDK/vendor provenance.- A tag push (
v2.10.0) builds wheels on each platform's own runner and
uploads to that vendor's PyPI lane; release wheels are uploaded concurrently
with a timeout and retry. - Required bootstrap is separated from optional framework hooks; backend
autoload and import-order are now a tested contract. - A new-chip onboarding skill automates standing up a new accelerator
backend, which is how T-Head ZW810E and Kunlunxin P800 were brought up in
this cycle.
Notable fixes
- Keep installed PyTorch immutable for vendor libtorch; include build
accelerator config in the wheel. - Keep the live allocation when resetting peak memory stats; keep the
flagoscurrent device across boxed calls. - Drop the global
Tensor.__getitem__patch inflagosdevice init. - Run each Ascend kernel on its operand's device, not the ambient one; route
Ascendtopkback to the ACLNN kernel. - On Moore Threads: materialize strided softmax operands before mudnn; keep
complex rotary ops on device; route integer division through mudnn. - On Hygon: reject out-of-range index values; realign FlagGems device name.
- On Enflame: serve
new_onesfrom a generated vendor kernel; escape int64 to
the vendor kernel on FlagGems pointwise routes. - Fail loud on an unrecognized
GEMS_VENDOR; report a missing
_flagos_ncclextension instead of aNoneTypeAttributeError.
Install
# After the 2.10.0 wheels land in the vendor PyPI lane (see
# docs/getting-started/installation.md for the index URL of your platform):
pip install "torch_fl==2.10.0+cuda13.3" # example for NVIDIA CUDA 13.3Before importing, sanity-check the wheel against the host:
torch-fl-preflight --wheel dist/torch_fl-2.10.0+cuda13.3.whl \
--platform cuda --sdk-version 13.3 --check-installedPython interpreters are pinned per platform — use the one matching your wheel
filename (cp312 / cp310 / cp311). FlagGems, FlagTree and FlagCX are pinned to
exact builds recorded in .github/version-pins.env; the routing tables were
generated against that cohort.
Known limitations
torch.compileis Experimental on every backend: it works on validated
models but has no CI step and may need per-model workarounds.- Profiling and RNG are not yet wired up on Kunlunxin P800.
torch.nn.attention.flex_attentionis not supported in this release;
support is planned for a future version.- The BPU s600 edge target runs eager ops on CPU; it is only a
torch.compilegraph device. - SDK/driver compatibility still requires platform-specific testing — the
compatibility manifest describes the artifact, not the host driver.
Full changelog
- Commits since
v0.1.0: 269 (62 feat, 101 fix, 11 perf, 25 ci, 18 docs,
16 test, 10 refactor, 4 build). - GitHub compare:
v0.1.0...main - Compatibility details:
docs/reference/compatibility.md - Operator support:
docs/reference/operator-support.md - HuggingFace coverage: 100+ HF models (BERT, Qwen3, Llama, BART, MPNet, …)
validated across the boards in the matrix above.
Credits
Thanks to every contributor who landed a backend, an operator route, a
kernel, a CI pipeline, or a docs fix in this cycle — see the
commit log
for the full list.