Repository navigation
Releases: flagos-ai/Torch-FL
Release list
torch-fl 2.10.0
Release date: 2026-09-29
Compatible PyTorch: >=2.10,<2.11
GitHub: flagos-ai/PyTorch-Plugin-FL
torch-fl is a PyTorch device plugin for the FlagOS software stack. It exposes a
single flagos device that routes each operator, per-call, among native vendor
kernels, portable FlagGems kernels, CUDA-compatibility boxing, and an explicit
CPU fallback — so the same PyTorch program runs unchanged across accelerators
from nine different vendors.
Starting with this release, the torch-fl version tracks the PyTorch minor line
it binds to: wheels are named torch_fl-2.10.0+<sdk> (e.g.
torch_fl-2.10.0+cuda13.3, torch_fl-2.10.0+dtk2604,
torch_fl-2.10.0+cann9.0.0, …). One tag, one artifact per platform.
Highlights
- One
flagosdevice, nine accelerator platforms. NVIDIA (A100 / H800),
Ascend (910C), Hygon (BW1000), MetaX (C550), T-Head (ZW810E), Moore Threads
(S5000), Enflame (S60), Kunlunxin (P800), and D-Robotics-class BPU (s600)
are all reachable through the samedevice="flagos"API. No workload changes
when switching hardware. - Distributed on every mainstream card. Basic collectives, DDP, and FSDP2
are validated across NVIDIA, Ascend, Hygon, MetaX, T-Head, Moore Threads,
Enflame, and Kunlunxin — not just one or two reference boards. torch.compilefirst-class.flagosis registered as an inductor GPU
device, backed by FlagTree on NVIDIA / Ascend / Hygon / MetaX / T-Head /
Moore Threads and by vendor Triton on Enflame; the BPU path compiles through
the on-board BPU compiler.- Full profiler coverage. PrivateUse1 flows, device time, and kernel
metadata reach parity withtorch.cuda.profiler, with per-vendor tracers
(CUPTI, ROCtracer, MUPTI, MSPTI, TOPSPTI, MCPTI). Kunlunxin P800 is the
current gap. - Three new cross-platform capabilities.
- A new-chip onboarding skill that stands up a new accelerator backend
end-to-end; - FP8 software emulation on the boxing path (no hardware FP8 required);
- Automatic operator-selection tuning that picks the best kernel
implementation per operator per shape.
- A new-chip onboarding skill that stands up a new accelerator backend
- Self-describing wheels. Every wheel ships
torch_fl/compatibility.json
recording platform, kernel sets, bundled libtorch, build-time PyTorch/ABI,
and the pinned FlagTree/FlagGems/FlagCX builds. A new
torch-fl-preflightCLI inspects a wheel (or an installed env) without
importingtorch_fland can emit the release compatibility table as
Markdown.
Platform matrix
Numbers below are the validated operator counts for this release, taken from
the internal platform capability sheet. "FlagGems-Python" shows the route
counts before → after this release.
| Vendor | Board | Runtime | Vendor ops | FlagGems-Python | FlagGems-C++ | FlagCX | Distributed | Profiler | AMP | RNG | torch.compile |
|---|---|---|---|---|---|---|---|---|---|---|---|
| NVIDIA | A100 | ✅ | 2033 | 390 → 424 | 18 | ✅ | basic + DDP + FSDP2 | ✅ | ✅ | ✅ | FlagTree |
| NVIDIA | H800 | ✅ | (shares A100 config) | ||||||||
| Ascend | 910C | ✅ | 354 | 221 → 225 | 0 | ✅ | basic + DDP + FSDP2 | ✅ | ✅ | ✅ | FlagTree |
| Hygon | BW1000 | ✅ | 2033 | 388 → 470 | 0 | ✅ | basic + DDP + FSDP2 | ✅ | ✅ | ✅ | FlagTree |
| MetaX | C550 | ✅ | 2033 | 377 → 443 | 17 | ✅ | basic + DDP + FSDP2 | ✅ | ✅ | ✅ | FlagTree |
| T-Head | ZW810E | ✅ | 2033 | 390 → 435 | 0 | ✅ | basic + DDP + FSDP2 | ✅ | ✅ | ✅ | FlagTree |
| Moore Threads | S5000 | ✅ | 105 | 391 → 468 | 0 | ✅ | basic + DDP + FSDP2 | ✅ | ✅ | ✅ | FlagTree |
| Enflame | S60 | ✅ | 224 | 281 → 257 | 0 | ✅ | basic + DDP + FSDP2 | ✅ | ✅ | ✅ | vendor Triton |
| Kunlunxin | P800 | ✅ | 2022 | 380 | 0 | ✅ | basic + DDP + FSDP2 | ❌ | ✅ | ❌ | ❌ |
| BPU (edge) | s600 | ✅ | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ | FX → BPU Compiler |
Notes:
- H800 reuses the A100 wheel and route table; it is listed separately only
because it is a shipping SKU. - Enflame S60's FlagGems-Python count tightened from 281 to 257: routes
that were not shape-validated were dropped rather than advertised. - Kunlunxin P800 ships with runtime, vendor ops, FlagCX, AMP, and the
distributed path; profiler and RNG integration, plustorch.compile, are
on the roadmap. - BPU s600 is an edge target: eager operators run on CPU, and the only
execution path istorch.compile(backend="...")through the on-board BPU
compiler.
What's changed since 0.1.0
Device model & routing
- A single
flagosdevice (PrivateUse1) replaces per-vendor device names;
device="cuda"is aliased ontoflagosonly after asking PyTorch, not a
redirected probe, andtorch.cuda.devicecontext objects are accepted. - Per-operator dispatch chooses between native vendor kernels,
FlagGems C++ (kFlagOs) and Python paths, CUDA boxing via an external
libtorch_cuda.so, and CPU fallback. Routing is per-op, not per-device. - New optional operator library TileOPs is integrated.
- Sparse storage on
flagos: construct COO tensors and serve CSR/CSC/BSR/BSC
tensors onSparseCsrPrivateUse1. DataParallelworks on theflagosbackend.uint16/uint32/uint64dtype casts are supported.- Automatic operator-selection tuning picks the best kernel implementation
per operator per shape.
Backends
- Ascend 910C — 354 routed vendor ops (ACLNN), 221 → 225 FlagGems routes;
native ACL stream/event semantics for FSDP2; boolean SDPA-mask semantics and
indexboolean masks preserved; finite additive attention bias carried
through the Ascend SDPA realShift path; nearest-exact upsampling kept
on-device; Qwen3-0.6B training measured at 0.82×torch_npuon real 910C. - Hygon BW1000 — runs on upstream PyTorch rather than DTK's fork;
2033 vendor ops via CUDA boxing over hipified DTK; low-precision software
GEMM path; FlagGems routes grew 388 → 470. - MetaX C550 — 2033 vendor ops via
cu-bridge; FSDP2 full feature parity
with Qwen3 training match; real deviceEvent+pin_memoryin
_to_copy; scaled-dot-product-attention routed to FlagGems; FP8/FP4
software emulation on the boxing path. - T-Head ZW810E — same CUDA-boxing path as NVIDIA against the PPU
CUDA-13-compatible SDK, bundling its own libtorch; FlagGems routes grew
390 → 435; AMP integrated. - Moore Threads S5000 — native
mudnnbackend with 105 vendor kernels
and 391 → 468 FlagGems routes; its own stream class with explicit vendor
dispatch; fp64addmm/baddbmmrouted to mudnn; native RNG unified with
FlagGems. - Enflame S60 — native
libtopsaten.sobackend with 224 vendor kernels;
FlagGems Triton on S60; int64 pointwise routes escaped to the vendor kernel
before the compiler aborts; Qwen-Image-2512 cut to 7.1 s/it. - Kunlunxin P800 — 2022 vendor ops, 380 FlagGems-Python routes, FlagCX
and AMP in; profiler / RNG /torch.compilefollow-up. - NVIDIA A100/H800 — CUDA boxing over external
libtorch_cuda.so;
FlagGems-first routing with FlagTree 3.6 integration; FlagGems-C++ dispatch
path (18 routes); self-contained wheels that run on stocktorch+cpu. - BPU s600 — edge target; eager on CPU,
torch.compilegraph path via
hbdk4 / on-board BPU compiler.
torch.compile
flagosregistered as a first-class inductor GPU device.- FlagTree backends validated on NVIDIA, Ascend (via triton-ascend), Hygon,
MetaX, T-Head, and Moore Threads; Enflame uses its vendor Triton. - A parameterized
set_env.shentrypoint and per-platform hooks make the
compile env reproducible.
Distributed
- DDP and FSDP2 validated on every mainstream board in the matrix.
- Collectives: MUSA via FlagCX+MCCL, GCU routed through FlagCX, Hygon via
RCCL on DTK, Ascend HCCL fallback, NCCL-shaped route elsewhere. - gloo requests answered on the
flagosbackend; collectives staged over host
memory. - Cross-device copies ordered against peer-device D2D copies and staged in
MudnnCopy.
AMP, autocast & RNG
PrivateUse1autocast and dtype support added; FP16/BF16 autocast +
GradScalermeasured on every mainstream board.- FP8/FP4 software emulation on the MetaX boxing path — usable without
FP8 hardware. - RNG unified across native and FlagGems paths via shared generator injection.
Profiler
- Torch-CUPTI parity for PrivateUse1: flows, device time, kernel metadata.
- Per-vendor tracer integration: ROCtracer (Hygon), MUPTI (Moore Threads),
MSPTI (Ascend), TOPSPTI (Enflame), MCPTI (MetaX), CUPTI on NVIDIA / T-Head. - The CUPTI runtime callback-id table is generated rather than hand-maintained.
Performance
- Qwen-Image-2512 on Enflame S60 cut to 7.1 s/it (SDPA route, GEMM operand
handling, five host round trips removed). - Hygon: FlagGems launch, layout and autotune costs cut behind the
Qwen-Image-2.1 gap; Qwen-Image rotation registered from the plugin; inert
masks dropped. - MetaX:
scaled_dot_product_attentionrouted to FlagGems; fused SDPA on the
boxing path restored via a composite override. - FlagGems SDPA route widened in
head_dim, dtype, and query-length bounds. - CPU-fallback inference reached 96 % of CUDA baseline throughput by
removing PrivateUse1 training overhead.
Benchmark: Qwen-Image-2.1 across vendors
Single-card, bf16, one 1024×1024 image, 40 denoising steps; median latency
per image, measured out of the box:
| Board | s/image | Throughput vs. H100 native | Peak memory |
|---|---|---|---|
| NVIDIA H100, native | 6.52 | 100 % | 36.8 GiB |
| NVIDIA H100, FlagOS | 10.67 | 61 % | 37.8 GiB |
| Moore Threads S5000 | 14.87 | 44 % | 50.8 GiB |
| Hygon BW1000 | 23.59 | 28 % | 36.8 GiB |
| Ascend 910C | 24.94 | 26 % | 43.3 GiB |
| MetaX C550 | 31.61 | 21 % | 37.0 GiB |
| T-Head ZW810E | 37.44 | 17 % | 36.8 GiB |
| Enflame S60 | 58.04 | 11 % | 37.1 GiB |
- Image quality is aligned, not just latency: T2I-100 CLIP-Score stays
within ±1.2 % of H100 native and CO...
FlagOS 2.1 — PyTorch-Plugin-FL v0.1.0
Release v0.1.0
Initial Release
Other
- Add the initial Implementation based on FlagOS (#1) by @Hchnr
- Support mm OP on Ascend NPU (#5) by @Hchnr
- Add Ascend Backend Support (#6) by @Hchnr
- Implemented inference and training capabilities for Qwen3-0.6B on the Metax platform. (#4) by @Chenyang409
- Dispatcher Infrastructure Cleanup and Naming Unification (#10) by @Hchnr
- Add triton-ascend patch script and update Ascend installation docs (#12) by @Hchnr
- Fix CPU fallback bottlenecks in inference, reaching 96% of CUDA baseline throughput (#14) by @Hchnr
- Feat(metax): MetaX hybrid backend config, abs kernel, and platform-aware ops tests (#13) by @Chenyang409
Contributors
Thanks to all contributors who made this release possible:
- @Hchnr (6 PRs)
- @Chenyang409 (2 PRs)
Generated by Release Notes Generator