Skip to content

v5.0.6

Latest

Choose a tag to compare

@janelu9 janelu9 released this 18 Jun 03:07
· 10 commits to main since this release

Qwen3.5/3.6 dense with MTP supported.
Flash-Linear-Attention-NPU was leveraged to accelerate the training of Qwen3.5 and Qwen3.6 on NPUs, achieving a 3–10x speedup.
Tensor parallelism with Megatron and MindSpeed will no longer be required as dependencies on the NPU platform.
Upgrade the NPU platform environment to Python 3.11.