Skip to content

Releases: flagos-ai/FlagGems-sglang

FlagGems-sglang v0.1.0

Choose a tag to compare

@wbavon wbavon released this 29 Sep 03:55
575d5bc

FlagGems-sglang v0.1.0

Part of FlagOS 2.2.

Initial release in FlagOS.

80 pull requests and 14 additional commits.

Source: v0.1.0, commit 575d5bc3f0ef.

New Features

  • [KernelGen][Nvidia]Add gemma_rms_norm operator with fused residual support (#3) by @yzw1128
  • Add CI for FlagGems-sglang (#5) by @liuhycs
  • [KernelGen][Nvidia] Add fused_recurrent_gated_delta_rule_packed_decode operator for sglang (#8) by @yzw1128
  • [KernelGen][Nvidia] Add mrotary_embedding operator for sglang (#10) by @yzw1128
  • [KernelGen][Nvidia] Add fused_moe operator with Triton kernel (#11) by @yzw1128
  • Add Apache 2.0 license header and KernelGen attribution to ops files (#17) by @yzw1128
  • Add _enflame backend support for Enflame chips (#28) by @liuhycs
  • [FlagOS Competition-Track1] Add silu_and_mul Triton Kernel for sglang (#33) by @yzw1128
  • [FlagOS Competition-Track1] Add causal_conv1d_fn Triton Kernel for sglang (#34) by @yzw1128
  • [FlagOS Competition-Track1] Add chunk_local_cumsum_scalar Triton Kernel for sglang (#37) by @xuanzhengdu-eng
  • [FlagOS Competition-Track1] Add merge_state Triton Kernel for sglang (#38) by @xuanzhengdu-eng
  • [FlagOS Competition-Track1] Add per_group_transpose Triton Kernel for sglang (#39) by @c2flowDS
  • [FlagOS Competition-Track1] Add mrope_fused Triton Kernel for sglang (#40) by @c2flowDS
  • [FlagOS Competition-Track1] Add fused_moe_gemm Triton Kernel for sglang (#41) by @c2flowDS
  • [FlagOS Competition-Track1] Add chunk_cumsum Triton Kernel for sglang (#42) by @ABan12
  • [FlagOS Competition-Track1] Add moe_sum_reduce Triton Kernel for sglang (#43) by @HAi-WORLD
  • [FlagOS Competition-Track1] Add chunk_local_cumsum_vector Triton Kernel for sglang (#45) by @yzw1128
  • [FlagOS Competition-Track1] Add chunk_state Triton Kernel for sglang (#46) by @yzw1128
  • [FlagOS Competition-Track1] Add context_attention Triton Kernel for sglang (#47) by @yzw1128
  • [FlagOS Competition-Track1] Add bmm_chunk Triton Kernel for sglang (#48) by @xuanzhengdu-eng
  • [FlagOS Competition-Track1] Add sgemm_lora_b Triton Kernel for sglang (#49) by @xuanzhengdu-eng
  • [FlagOS Competition-Track1] Add mamba_layernorm_gated Triton Kernel for sglang (#50) by @xuanzhengdu-eng
  • [FlagOS Competition-Track1] Add embedding_lora_a Triton Kernel for sglang (#51) by @xuanzhengdu-eng
  • [FlagOS Competition-Track1] Add decode_grouped_attention Triton Kernel for sglang (#52) by @xuanzhengdu-eng
  • [FlagOS Competition-Track1] Add decode_attention Triton Kernel for sglang (#53) by @xuanzhengdu-eng
  • [FlagOS Competition-Track1] Add apply_token_bitmask Triton Kernel for sglang (#56) by @Sawyer117
  • [FlagOS Competition-Track1] Add qkv_lora_b Triton Kernel for sglang (#58) by @c2flowDS
  • [FlagOS Competition-Track1] Add fused_rmsnorm Triton Kernel for sglang (#59) by @at0rp1d0i-cell
  • Add operators configuration file (#60) by @liuhycs
  • [FlagOS Competition-Track1] Add chunk_state_varlen Triton Kernel for sglang (#61) by @c2flowDS
25 more changes
  • [FlagOS Competition-Track1] Add interleaved_rope Triton Kernel for sglang (#63) by @HAi-WORLD
  • [FlagOS Competition-Track1] Add sigmoid_gate_topk_renorm Triton Kernel for sglang (#64) by @bagpipe1289
  • [FlagOS Competition-Track1] Add sgemm_lora_a Triton Kernel for sglang (#65) by @yzw1128
  • [FlagOS Competition-Track1] Add state_passing Triton Kernel for sglang (#66) by @yzw1128
  • [FlagOS Competition-Track1] Add fused_moe_router_tensorcore Triton Kernel for sglang (#67) by @yzw1128
  • [FlagOS Competition-Track1] Add selective_state_update Triton Kernel for sglang (#68) by @yzw1128
  • [FlagOS Competition-Track1] Add silu_and_mul_masked Triton Kernel for sglang (#69) by @yzw1128
  • [FlagOS Competition-Track1] Add gelu_and_mul Triton Kernel for sglang (#70) by @hangglider5
  • [FlagOS Competition-Track1] Add moe_fused_gate Triton Kernel for sglang (#71) by @xuanzhengdu-eng
  • [FlagOS Competition-Track1] Add gate_up_lora_b Triton Kernel for sglang (#72) by @xuanzhengdu-eng
  • [FlagOS Competition-Track1] Add rotary_embedding Triton Kernel for sglang (#73) by @c2flowDS
  • [FlagOS Competition-Track1] Add draft_topk1 Triton Kernel for sglang (#74) by @c2flowDS
  • [FlagOS Competition-Track1] Add per_token_quant_int8 Triton Kernel for sglang (#75) by @c2flowDS
  • [FlagOS Competition-Track1] Add per_token_group_quant_int8 Triton Kernel for sglang (#77) by @c2flowDS
  • [FlagOS Competition-Track1] Add softcap_inplace_logits Triton Kernel for sglang (#78) by @c2flowDS
  • [FlagOS Competition-Track1] Add fused_moe_router_cudacore Triton Kernel for sglang (#79) by @c2flowDS
  • Add missing operator entries to conf/operators.yaml (#86) by @liuhycs
  • [FlagOS Competition-Track1] Add moe_fused_mul_sum Triton Kernel for sglang (#87) by @yunyiliu
  • feat: integrate Kunlunxin and MUSA fused MoE routers (#106) by @liuhycs
  • Add README.md (4cefa23decf1; commit)
  • Add NORM_SHAPES for accuracy testing (f5ac6bb0d824; commit)
  • Add CI badge to README (720675c883ef; commit)
  • Add CI badge to README_cn.md (c9950eccc621; commit)
  • Add permissions for pull-requests in CI workflow (16f1052e3b26; commit)
  • Add statuses permission to CI workflow (51f1f8bf236e; commit)

Hardware Support

  • ascend: cap fused router K-loop pipelining at 2 stages (#96) by @liuhycs
  • enflame: register gcu vendor and codegen config (#99) by @liuhycs
  • ascend: compute fused router logits with tl.dot (#102) by @liuhycs
  • ascend: defer softmax divide in fused router top-k epilogue (#104) by @liuhycs
  • enflame: accumulate inside the MMA in fused router GEMMs (#105) by @liuhycs

Improvements

  • [Refactor][Backend] Backend reorganization + multi-level op registrar (#19) by @liuhycs

Documentation

CI/Infrastructure

  • ci(pr-template): expand with accuracy + speed test sections (#23) by @liuhycs

Testing

  • Test the mrope_fused op instead of the reference (#84) by @liuhycs
  • benchmark: ...
Read more