Skip to content

torch-fl 2.10.0

Latest

Choose a tag to compare

@lvyufeng lvyufeng released this 29 Sep 10:20
· 4 commits to main since this release
6d7ee08

Release date: 2026-09-29
Compatible PyTorch: >=2.10,<2.11
GitHub: flagos-ai/PyTorch-Plugin-FL

torch-fl is a PyTorch device plugin for the FlagOS software stack. It exposes a
single flagos device that routes each operator, per-call, among native vendor
kernels, portable FlagGems kernels, CUDA-compatibility boxing, and an explicit
CPU fallback — so the same PyTorch program runs unchanged across accelerators
from nine different vendors.

Starting with this release, the torch-fl version tracks the PyTorch minor line
it binds to
: wheels are named torch_fl-2.10.0+<sdk> (e.g.
torch_fl-2.10.0+cuda13.3, torch_fl-2.10.0+dtk2604,
torch_fl-2.10.0+cann9.0.0, …). One tag, one artifact per platform.


Highlights

  • One flagos device, nine accelerator platforms. NVIDIA (A100 / H800),
    Ascend (910C), Hygon (BW1000), MetaX (C550), T-Head (ZW810E), Moore Threads
    (S5000), Enflame (S60), Kunlunxin (P800), and D-Robotics-class BPU (s600)
    are all reachable through the same device="flagos" API. No workload changes
    when switching hardware.
  • Distributed on every mainstream card. Basic collectives, DDP, and FSDP2
    are validated across NVIDIA, Ascend, Hygon, MetaX, T-Head, Moore Threads,
    Enflame, and Kunlunxin — not just one or two reference boards.
  • torch.compile first-class. flagos is registered as an inductor GPU
    device, backed by FlagTree on NVIDIA / Ascend / Hygon / MetaX / T-Head /
    Moore Threads and by vendor Triton on Enflame; the BPU path compiles through
    the on-board BPU compiler.
  • Full profiler coverage. PrivateUse1 flows, device time, and kernel
    metadata reach parity with torch.cuda.profiler, with per-vendor tracers
    (CUPTI, ROCtracer, MUPTI, MSPTI, TOPSPTI, MCPTI). Kunlunxin P800 is the
    current gap.
  • Three new cross-platform capabilities.
    1. A new-chip onboarding skill that stands up a new accelerator backend
      end-to-end;
    2. FP8 software emulation on the boxing path (no hardware FP8 required);
    3. Automatic operator-selection tuning that picks the best kernel
      implementation per operator per shape.
  • Self-describing wheels. Every wheel ships torch_fl/compatibility.json
    recording platform, kernel sets, bundled libtorch, build-time PyTorch/ABI,
    and the pinned FlagTree/FlagGems/FlagCX builds. A new
    torch-fl-preflight CLI inspects a wheel (or an installed env) without
    importing torch_fl
    and can emit the release compatibility table as
    Markdown.

Platform matrix

Numbers below are the validated operator counts for this release, taken from
the internal platform capability sheet. "FlagGems-Python" shows the route
counts before → after this release.

Vendor Board Runtime Vendor ops FlagGems-Python FlagGems-C++ FlagCX Distributed Profiler AMP RNG torch.compile
NVIDIA A100 ✅ 2033 390 → 424 18 ✅ basic + DDP + FSDP2 ✅ ✅ ✅ FlagTree
NVIDIA H800 ✅ (shares A100 config)
Ascend 910C ✅ 354 221 → 225 0 ✅ basic + DDP + FSDP2 ✅ ✅ ✅ FlagTree
Hygon BW1000 ✅ 2033 388 → 470 0 ✅ basic + DDP + FSDP2 ✅ ✅ ✅ FlagTree
MetaX C550 ✅ 2033 377 → 443 17 ✅ basic + DDP + FSDP2 ✅ ✅ ✅ FlagTree
T-Head ZW810E ✅ 2033 390 → 435 0 ✅ basic + DDP + FSDP2 ✅ ✅ ✅ FlagTree
Moore Threads S5000 ✅ 105 391 → 468 0 ✅ basic + DDP + FSDP2 ✅ ✅ ✅ FlagTree
Enflame S60 ✅ 224 281 → 257 0 ✅ basic + DDP + FSDP2 ✅ ✅ ✅ vendor Triton
Kunlunxin P800 ✅ 2022 380 0 ✅ basic + DDP + FSDP2 ❌ ✅ ❌ ❌
BPU (edge) s600 ✅ ❌ ❌ ❌ ❌ ❌ ❌ ❌ ❌ FX → BPU Compiler

Notes:

  • H800 reuses the A100 wheel and route table; it is listed separately only
    because it is a shipping SKU.
  • Enflame S60's FlagGems-Python count tightened from 281 to 257: routes
    that were not shape-validated were dropped rather than advertised.
  • Kunlunxin P800 ships with runtime, vendor ops, FlagCX, AMP, and the
    distributed path; profiler and RNG integration, plus torch.compile, are
    on the roadmap.
  • BPU s600 is an edge target: eager operators run on CPU, and the only
    execution path is torch.compile(backend="...") through the on-board BPU
    compiler.

What's changed since 0.1.0

Device model & routing

  • A single flagos device (PrivateUse1) replaces per-vendor device names;
    device="cuda" is aliased onto flagos only after asking PyTorch, not a
    redirected probe, and torch.cuda.device context objects are accepted.
  • Per-operator dispatch chooses between native vendor kernels,
    FlagGems C++ (kFlagOs) and Python paths, CUDA boxing via an external
    libtorch_cuda.so, and CPU fallback. Routing is per-op, not per-device.
  • New optional operator library TileOPs is integrated.
  • Sparse storage on flagos: construct COO tensors and serve CSR/CSC/BSR/BSC
    tensors on SparseCsrPrivateUse1.
  • DataParallel works on the flagos backend.
  • uint16/uint32/uint64 dtype casts are supported.
  • Automatic operator-selection tuning picks the best kernel implementation
    per operator per shape.

Backends

  • Ascend 910C — 354 routed vendor ops (ACLNN), 221 → 225 FlagGems routes;
    native ACL stream/event semantics for FSDP2; boolean SDPA-mask semantics and
    index boolean masks preserved; finite additive attention bias carried
    through the Ascend SDPA realShift path; nearest-exact upsampling kept
    on-device; Qwen3-0.6B training measured at 0.82× torch_npu on real 910C.
  • Hygon BW1000 — runs on upstream PyTorch rather than DTK's fork;
    2033 vendor ops via CUDA boxing over hipified DTK; low-precision software
    GEMM path; FlagGems routes grew 388 → 470.
  • MetaX C550 — 2033 vendor ops via cu-bridge; FSDP2 full feature parity
    with Qwen3 training match; real device Event + pin_memory in
    _to_copy; scaled-dot-product-attention routed to FlagGems; FP8/FP4
    software emulation on the boxing path.
  • T-Head ZW810E — same CUDA-boxing path as NVIDIA against the PPU
    CUDA-13-compatible SDK, bundling its own libtorch; FlagGems routes grew
    390 → 435; AMP integrated.
  • Moore Threads S5000 — native mudnn backend with 105 vendor kernels
    and 391 → 468 FlagGems routes; its own stream class with explicit vendor
    dispatch; fp64 addmm/baddbmm routed to mudnn; native RNG unified with
    FlagGems.
  • Enflame S60 — native libtopsaten.so backend with 224 vendor kernels;
    FlagGems Triton on S60; int64 pointwise routes escaped to the vendor kernel
    before the compiler aborts; Qwen-Image-2512 cut to 7.1 s/it.
  • Kunlunxin P800 — 2022 vendor ops, 380 FlagGems-Python routes, FlagCX
    and AMP in; profiler / RNG / torch.compile follow-up.
  • NVIDIA A100/H800 — CUDA boxing over external libtorch_cuda.so;
    FlagGems-first routing with FlagTree 3.6 integration; FlagGems-C++ dispatch
    path (18 routes); self-contained wheels that run on stock torch+cpu.
  • BPU s600 — edge target; eager on CPU, torch.compile graph path via
    hbdk4 / on-board BPU compiler.

torch.compile

  • flagos registered as a first-class inductor GPU device.
  • FlagTree backends validated on NVIDIA, Ascend (via triton-ascend), Hygon,
    MetaX, T-Head, and Moore Threads; Enflame uses its vendor Triton.
  • A parameterized set_env.sh entrypoint and per-platform hooks make the
    compile env reproducible.

Distributed

  • DDP and FSDP2 validated on every mainstream board in the matrix.
  • Collectives: MUSA via FlagCX+MCCL, GCU routed through FlagCX, Hygon via
    RCCL on DTK, Ascend HCCL fallback, NCCL-shaped route elsewhere.
  • gloo requests answered on the flagos backend; collectives staged over host
    memory.
  • Cross-device copies ordered against peer-device D2D copies and staged in
    MudnnCopy.

AMP, autocast & RNG

  • PrivateUse1 autocast and dtype support added; FP16/BF16 autocast +
    GradScaler measured on every mainstream board.
  • FP8/FP4 software emulation on the MetaX boxing path — usable without
    FP8 hardware.
  • RNG unified across native and FlagGems paths via shared generator injection.

Profiler

  • Torch-CUPTI parity for PrivateUse1: flows, device time, kernel metadata.
  • Per-vendor tracer integration: ROCtracer (Hygon), MUPTI (Moore Threads),
    MSPTI (Ascend), TOPSPTI (Enflame), MCPTI (MetaX), CUPTI on NVIDIA / T-Head.
  • The CUPTI runtime callback-id table is generated rather than hand-maintained.

Performance

  • Qwen-Image-2512 on Enflame S60 cut to 7.1 s/it (SDPA route, GEMM operand
    handling, five host round trips removed).
  • Hygon: FlagGems launch, layout and autotune costs cut behind the
    Qwen-Image-2.1 gap; Qwen-Image rotation registered from the plugin; inert
    masks dropped.
  • MetaX: scaled_dot_product_attention routed to FlagGems; fused SDPA on the
    boxing path restored via a composite override.
  • FlagGems SDPA route widened in head_dim, dtype, and query-length bounds.
  • CPU-fallback inference reached 96 % of CUDA baseline throughput by
    removing PrivateUse1 training overhead.

Benchmark: Qwen-Image-2.1 across vendors

Single-card, bf16, one 1024×1024 image, 40 denoising steps; median latency
per image, measured out of the box:

Board s/image Throughput vs. H100 native Peak memory
NVIDIA H100, native 6.52 100 % 36.8 GiB
NVIDIA H100, FlagOS 10.67 61 % 37.8 GiB
Moore Threads S5000 14.87 44 % 50.8 GiB
Hygon BW1000 23.59 28 % 36.8 GiB
Ascend 910C 24.94 26 % 43.3 GiB
MetaX C550 31.61 21 % 37.0 GiB
T-Head ZW810E 37.44 17 % 36.8 GiB
Enflame S60 58.04 11 % 37.1 GiB
  • Image quality is aligned, not just latency: T2I-100 CLIP-Score stays
    within ±1.2 % of H100 native and COCO CLIP-Score within ±1.6 % across the
    boards above; FID differences are in the low single-digit percent range.
  • CPU-only Arm path (Mac M5 Pro, W8A8, same 1024² / 40-step workload):
    FlagOS finishes in 785 s vs. 1365 s for stock PyTorch CPU — 1.74×
    faster
    .
  • Kunlunxin P800 has not yet been benchmarked on this workload.

Packaging & release engineering

  • Wheels are self-contained and run on a stock torch+cpu install; the
    installed PyTorch is left immutable against vendor libtorch.
  • Wheel local segment names the SDK it was built against; one interpreter per
    platform (3.12 on CUDA/GCU/MetaX/PPU, 3.10 on DCU/MUSA, 3.11 on Ascend).
  • torch_fl/compatibility.json is published and validated; the
    torch-fl-preflight CLI checks dependencies, warns on build-env drift, and
    fails a release table that is missing SDK/vendor provenance.
  • A tag push (v2.10.0) builds wheels on each platform's own runner and
    uploads to that vendor's PyPI lane; release wheels are uploaded concurrently
    with a timeout and retry.
  • Required bootstrap is separated from optional framework hooks; backend
    autoload and import-order are now a tested contract.
  • A new-chip onboarding skill automates standing up a new accelerator
    backend, which is how T-Head ZW810E and Kunlunxin P800 were brought up in
    this cycle.

Notable fixes

  • Keep installed PyTorch immutable for vendor libtorch; include build
    accelerator config in the wheel.
  • Keep the live allocation when resetting peak memory stats; keep the
    flagos current device across boxed calls.
  • Drop the global Tensor.__getitem__ patch in flagos device init.
  • Run each Ascend kernel on its operand's device, not the ambient one; route
    Ascend topk back to the ACLNN kernel.
  • On Moore Threads: materialize strided softmax operands before mudnn; keep
    complex rotary ops on device; route integer division through mudnn.
  • On Hygon: reject out-of-range index values; realign FlagGems device name.
  • On Enflame: serve new_ones from a generated vendor kernel; escape int64 to
    the vendor kernel on FlagGems pointwise routes.
  • Fail loud on an unrecognized GEMS_VENDOR; report a missing
    _flagos_nccl extension instead of a NoneType AttributeError.

Install

# After the 2.10.0 wheels land in the vendor PyPI lane (see
# docs/getting-started/installation.md for the index URL of your platform):
pip install "torch_fl==2.10.0+cuda13.3"   # example for NVIDIA CUDA 13.3

Before importing, sanity-check the wheel against the host:

torch-fl-preflight --wheel dist/torch_fl-2.10.0+cuda13.3.whl \
    --platform cuda --sdk-version 13.3 --check-installed

Python interpreters are pinned per platform — use the one matching your wheel
filename (cp312 / cp310 / cp311). FlagGems, FlagTree and FlagCX are pinned to
exact builds recorded in .github/version-pins.env; the routing tables were
generated against that cohort.


Known limitations

  • torch.compile is Experimental on every backend: it works on validated
    models but has no CI step and may need per-model workarounds.
  • Profiling and RNG are not yet wired up on Kunlunxin P800.
  • torch.nn.attention.flex_attention is not supported in this release;
    support is planned for a future version.
  • The BPU s600 edge target runs eager ops on CPU; it is only a
    torch.compile graph device.
  • SDK/driver compatibility still requires platform-specific testing — the
    compatibility manifest describes the artifact, not the host driver.

Full changelog

Credits

Thanks to every contributor who landed a backend, an operator route, a
kernel, a CI pipeline, or a docs fix in this cycle — see the
commit log
for the full list.