Skip to content

Releases: baidu-baige/LoongForge

v0.2.0

Choose a tag to compare

@Zachary-wW Zachary-wW released this 20 Aug 09:28
Immutable release. Only release title and notes can be modified.
c627566

Highlights

  • LoongForge-Embodied: New torch-native DDP/FSDP training stack for π0.5, GR00T-N1.6/N1.7, xVLA, LingBot-VA, FastWAM, DreamZero, and Cosmos3, delivering up to 2.67× training speedup
  • New model support: DeepSeek-V4 Flash/Pro, GLM-5.2, and MiniCPM-V 4.6; Qwen-Image-Edit-2511 and Wan2.1 diffusion training
  • Embodied evaluation: Unified evaluation workflows for LIBERO, CALVIN, SimplerEnv, RoboTwin, and ManiSkill
  • MoE & scaling: TAOT topology-aware expert load balancing; DeepSeek-V4 FP8 training and context parallelism
  • Checkpoint & data pipeline: Memory-efficient packed INT4/MXFP4 HF loading, WDS-native offline packing, and HuggingFace chat-template/tool-calling SFT support

What's Changed

  • LoongForge-Embodied: Added a torch-native DDP/FSDP training subsystem for π0.5, GR00T-N1.6/N1.7, xVLA, LingBot-VA, FastWAM, DreamZero, and Cosmos3 (#114)
  • Embodied evaluation: Added a unified evaluation framework with GR00T-N1.6 support, SimplerEnv WidowX tasks, and LingBot-VA online inference integration (#140, #154, #165)
  • DeepSeek-V4: Added training support, HuggingFace checkpoint loading, and Context Parallelism (#98, #141, #143)
  • New VLM support: Added MiniCPM-V 4.6 and GLM-5.2 + Kimi K2.6 multimodal LoRA (#146, #162)
  • Diffusion training: Added Qwen-Image-Edit-2511 with FSDP + TP and Wan2.1 with FSDP + Context Parallelism (#85, #101)
  • MoE load balancing: Added topology-aware dynamic expert replica placement through TAOT to improve Expert Parallel workload balance (paper)
  • Low-bit checkpoint loading: Added memory-efficient HuggingFace checkpoint bridge loading for packed INT4 and MXFP4 weights (#108)
  • Multimodal data pipeline: Added WDS-native offline packing, HuggingFace Jinja chat templates, Kimi K2.5 tool-calling SFT, and configurable Energon read order (#78, #87, #88, #139)
  • Packed sample support: Improved multi-turn and text-only PackedMultiMixQASample handling and reused standard encoders for packed VLM samples (#116, #150, #168)
  • FastWAM optimization: Added ZeRO-1 SFT examples, fused RMSNorm/RoPE/LayerNorm paths, and optimized VAE encoding (#123, #129, #136, #138, #159, #161)
  • Distributed training: Added configurable DDP communication hooks for gradient compression (#158)
  • Infrastructure: Unified Docker images across LLM, VLM, diffusion, and embodied training workloads (#126)

New Contributors

Full Changelog: v0.1.1...v0.2.0

v0.1.1

Choose a tag to compare

@Zachary-wW Zachary-wW released this 26 May 09:03
Immutable release. Only release title and notes can be modified.
c5c4205

Highlights

  • New model support: Kimi-K2.5 / K2.6 VLM training and MCore↔HF checkpoint conversion
  • Wan2.2 Diffusion: Data packing training support; FSDP + context parallelism (ring & ulysses)
  • GR00T-N1.6 VLA: Full-iteration / per-microbatch CUDA graph for accelerated training
  • INT4 checkpoint dequantization: Support HF compressed-tensors INT4 format

What's Changed

  • docs(readme): add LoongForge banner by @nullnonenilNULL in #38
  • Standard sft_dataset_config.yaml: Remove history from the Alpaca data… by @gsl322 in #52
  • [bridge] Enable fp8 force pow2on for bridge by @pengxiangyu in #50
  • [convert] Convert checkpoint support get hf checkpoint file names by expert_ids by @pengxiangyu in #53
  • [groot_n1.6] Support full-iteration / per-microbatch CUDA graph by @Dana57740 in #54
  • [Task] wan fsdp + cp(ring + ulysses) by @haichengv12 in #56
  • [ci, docs] feat: add initial GitHub Actions CI surface by @Zachary-wW in #51
  • [ckpt] feat: support HF INT4 compressed-tensors dequantization by @kaimo455 in #60
  • [vlm, data, train] fix: align packed FP8 padding and media broadcast by @kaimo455 in #61
  • [data] fix: log skipped VLM dataloader samples as warnings by @kaimo455 in #62
  • [data, vlm] fix: skip overlong Kimi VLM samples by @kaimo455 in #63
  • [Task] Wan2.2 data packing training support by @VEGA1215 in #64
  • [ckpt] feat: support Kimi K2.x checkpoint conversion by @kaimo455 in #67
  • [docs] docs: split bilingual docs into separate source trees by @Zachary-wW in #65
  • [docs] docs: add Kimi-K2.6 support model mapping by @Zachary-wW in #68
  • [docs] docs: fix diffusion docs structure and align ZH/EN toctree by @Zachary-wW in #69

New Contributors

Full Changelog: v0.1.0...v0.1.1

v0.1.0

Choose a tag to compare

@Zachary-wW Zachary-wW released this 09 May 14:06
Immutable release. Only release title and notes can be modified.
d209742

We are excited to announce the first public release of LoongForge, a modular, scalable, and highly efficient training framework for large-scale transformer models across diverse modalities and architectures. This initial release covers four modality types (LLM, VLM, VLA, Diffusion), with out-of-the-box support for Pretrain, MidTrain, SFT, and LoRA workflows.

📖 Documentation: Docs


Supported Models

LoongForge supports a wide range of model configurations out-of-the-box. Below are the highlighted models; for the complete list see Supported Models.

LLM

  • Qwen3 — 0.6B, 1.7B, 4B, 8B, 14B, 32B dense; 30B-A3B, 235B-A22B, 480B-A35B MoE; Coder-30B-A3B
  • Qwen3-Next — 80B-A3B (next-gen MoE)
  • DeepSeek-V3 — V3 (BF16 + FP8), V3.2 (FP8)
  • DeepSeek-V2 — V2, V2-Lite
  • Qwen2.5 — 0.5B to 72B (7 sizes)
  • MiniMax — M2.1 / M2.5 / M2.7 (230B)
  • GLM5 · MIMO-7B
  • Older families (LLaMA 2/3/3.1, Qwen, Qwen1.5, Qwen2)

VLM

  • Qwen3-VL — 30B-A3B, 235B-A22B MoE
  • Qwen3.5 — 0.8B–397B-A17B (8 sizes, dense + MoE)
  • Qwen3.6 — 27B, 35B-A3B
  • Qwen2.5-VL — 3B, 7B, 32B, 72B
  • InternVL-3.5 — 8B, 14B, 38B, 30B-A3B, 241B-A28B
  • InternVL-2.5 — 8B, 26B, 38B, 78B
  • ERNIE4.5-VL — 28B-A3B
  • LLaVA-OneVision-1.5 — 4B

Diffusion & VLA

  • WAN 2.2
  • GR00T N1.6
  • Pi0.5

Key Features

⚡ Adaptive FP8 Training

  • Adaptive FP8 — Benchmark-driven per-module FP8/BF16 dynamic precision selection. Enables FP8 only where profiling confirms a speedup; never slower than BF16. Supports blockwise FP8 (E4M3, DeepSeek-V3 scheme) and per-component granularity control for both LLMs and VLMs.

🔀 MoE Optimization

  • A2A Communication Overlap — 1F1B overlap strategy hides EP All-to-All behind forward/backward compute. FP8 compatible, lower memory than upstream Megatron-LM.
  • Fine-grained Activation Offload — Module-level selective recompute + tensor-level D2H/H2D overlap with compute streams, compatible with the 1F1B overlap schedule.

🖼️ VLM Training

  • Model Composition — YAML-driven ViT + Projector + LLM assembly via OmniCombinationModel. Zero-code model construction with interchangeable encoder/backbone combinations.
  • Heterogeneous Parallelism — Independent TP/DP for encoder vs decoder (Hetero-TP, Hetero-DP, Full Hetero-DP).
  • ViT DP Load Balancing — Redistributes image tokens across DP ranks based on quadratic ViT cost model before the encoder forward pass.
  • DP Load Balancing — Profiles O(n²) attention cost from fixed-length packing, reorders samples across DP ranks via greedy LPT to eliminate stragglers.

🧠 Memory & Compute

  • Fused Linear Cross Entropy — Fuses vocab-projection + cross-entropy into a single operation with chunked computation. The complete logits tensor is never instantiated.
  • Optimizer CPU Offload — Adam states offloaded to CPU; supports BF16/FP8 low-precision moments.

💾 Checkpoint & LoRA

  • Online HF Checkpoint Load/Save — Load HuggingFace checkpoints directly at startup with no offline conversion; export back with --save-hf true. Supports TP/PP/EP parallelism.
  • Offline Conversion — Bidirectional HF↔MCore for all models. FP8 + MoE expert merging included.
  • LoRA — YAML-configurable with wildcard module matching. VLM per-component toggles. Merge & export to HF supported.

🔧 Custom Operators

  • FusedDSA — FlashMLA + Lightning Indexer fused ops for DeepSeek-style Sparse Attention (DSA) training. Integrates forward and custom backward operators for end-to-end training acceleration. (CUDA kernels available on Baidu Baige Platform; not included in this release)
  • TileLang Ops — Open-sourced Sparse MLA Fwd/Bwd + Lightning Indexer Bwd, built on TileLang.

📦 Data Pipeline

  • Offline Packing — Groups and re-orders variable-length multimodal samples into fixed-length packed WebDataset shards, reducing padding and increasing training throughput.
  • Multi-Format Dataset Support — Natively supports ShareGPT and Alpaca conversation formats for SFT, with WebDataset for multimodal pretrain data.
  • E2E Test Suite — YAML-driven test framework covering all models.

🖥️ Heterogeneous Hardware

Same codebase runs on NVIDIA GPU and Baidu Kunlun XPU (P800). We have validated the following models on both platforms:

  • LLM Pretrain (DeepSeek-V3.1, MiniMax, Qwen2.5, Qwen3)
  • VLM Training (Qwen3-VL, InternVL3.5, Qwen3.5)
  • VLA Training (Pi0.5)

Contributors

Thanks to all contributors who made this release possible! In addition to the contributors listed below, we are grateful to 20+ internal developers and early contributors whose work laid the foundation for LoongForge.

@nullnonenilNULL @VVsssssk @XueSongTap @Zachary-wW @pengxiangyu @NeverlanD0829 @Dreamspr22 @Yangsx-1 @kaimo455


Feedback & Bug Reports

This is our first public release — we actively welcome community feedback:

  • Bug reports — Open a GitHub Issue with reproduction script, environment details, and full traceback
  • Feature requests — Use the feature request issue template
  • Contributing — PRs welcome! Please read the Contributing Guidelines first