Skip to content

FlagGems v5.4.0

Latest

Choose a tag to compare

@wbavon wbavon released this 28 Sep 09:23
6ed2f39

FlagGems v5.4.0

Part of FlagOS 2.2.

Changes since v5.3.0 (FlagOS 2.1).

1247 pull requests and 39 additional commits.

Source: v5.4.0, commit 6ed2f39071ef.

New Features

536 more changes
  • [KernelGen][Nvidia] Add special_digamma operator with Triton kernel (#3520) by @LoserCheems
  • [KernelGen][Nvidia] Add gcd_ operator with Triton kernel (#3521) by @LoserCheems
  • [KernelGen][Nvidia] Add fmax operator with Triton kernel (#3522) by @LoserCheems
  • [KernelGen][Nvidia] Add permute_copy operator (#3535) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add resize operator with Triton kernel (#3537) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add resize_as operator with Triton kernel (#3538) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add special_chebyshev_polynomial_v operator (#3539) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add special_xlog1py operator with Triton kernel (#3540) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add true_divide standalone test and benchmark (#3542) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add upsample_trilinear3d operator with Triton kernel (#3555) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add rot90 operator with Triton kernel (#3556) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add rrelu_with_noise_functional operator with Triton kernel (#3557) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add rnn_relu operator with Triton kernel (#3558) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add logical_not_ operator with Triton kernel (#3561) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add test and benchmark for uniform_ operator (#3614) by @bwbwzzz
  • [SiliconFlow] Add unique_dim operator (#3642) by @ShawnXuan
  • [SiliconFlow] Add searchsorted operator (#3658) by @winfan1314
  • add flash_mla_with_kvcache op (#3659) by @Jennie1004
  • [KernelGen][Nvidia] Add special_modified_bessel_k1 operator with Triton kernel (#3661) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add special_shifted_chebyshev_polynomial_w operator with Triton kernel (#3662) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add BatchNorm operator with Muxi and Tianshu domestic backends (#3664) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add atanh operator with Triton kernel (#3669) by @LoserCheems
  • [KernelGen][Nvidia] Add view_copy operator with Triton kernel (#3670) by @LoserCheems
  • [KernelGen][Nvidia] Add tensor_split operator with Triton kernel (#3671) by @LoserCheems
  • [KernelGen][Nvidia] Add split_with_sizes_copy operator with Triton kernel (#3672) by @LoserCheems
  • [KernelGen][Nvidia] Add special_erfinv operator with Triton kernel (#3673) by @LoserCheems
  • [KernelGen][Nvidia] Add sinh operator with Triton kernel (#3674) by @LoserCheems
  • [KernelGen][Nvidia] Add mse_loss_backward operator with Triton kernel (#3675) by @LoserCheems
  • [KernelGen][Nvidia] Add lcm_ operator with Triton kernel (#3676) by @LoserCheems
  • [KernelGen][Nvidia] Add is_nonzero operator with Triton kernel (#3677) by @LoserCheems
  • [KernelGen][Nvidia] Add gt_scalar_ and gt_tensor_ operators with Triton kernel (#3678) by @LoserCheems
  • [KernelGen][Nvidia] Add frac and frac_ operators with Triton kernel (#3679) by @LoserCheems
  • [KernelGen][Nvidia] Add diagonal_scatter operator with Triton kernel (#3680) by @LoserCheems
  • [KernelGen][Nvidia] Add _upsample_nearest_exact3d operator with Triton kernel (#3682) by @LoserCheems
  • [KernelGen][Nvidia] Add _embedding_bag_per_sample_weights_backward operator with Triton kernel (#3683) by @LoserCheems
  • [KernelGen][Nvidia] Add _chunk_cat operator with Triton kernel (#3684) by @LoserCheems
  • [KernelGen][Nvidia] Add xor operator with Triton kernel (#3685) by @LoserCheems
  • [KernelGen][Nvidia] Add empty operator with Triton kernel (#3693) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add fmod_ operator with Triton kernel (#3694) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add BeamSearchScore operator with Triton kernel (#3695) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add im2col operator with Triton kernel (#3698) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add special_modified_bessel_k0 operator with Triton kernel (#3699) by @bwbwzzz
  • [SiliconFlow] Add index_reduce_ operator (#3709) by @winfan1314
  • [feature]Flagtune mm with GROUP_M (#3713) by @Fattysand
  • [KernelGen][Nvidia] Add soft_margin_loss_backward operator with Triton kernel (#3726) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add trunc_ operator with Triton kernel (#3729) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add ilshift operator with Triton kernel (#3730) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add _unsafe_masked_index operator with Triton kernel (#3731) by @XDYuanzhuLee
  • [Advanced Compiler] add fp64 dtype test for polar operator (#3746) by @AdvancedCompiler
  • Add benchmark test for cumulative sum operation (#3748) by @tianxiao-baai
  • [KernelGen][Nvidia] Add MatmulBiasActivation operator with Triton kernel (#3749) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add special_shifted_chebyshev_polynomial_v operator with Triton kernel (#3783) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add addmm_ operator with Triton kernel (#3786) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add _upsample_nearest_exact2d_backward operator with Triton kernel (#3787) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add unfold_copy operator with Triton kernel (#3788) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add log_normal_ operator with Triton kernel (#3789) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add _adaptive_avg_pool3d_backward operator with Triton kernel (#3790) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add linear operator with Triton kernel (#3797) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add _prelu_kernel_backward operator with Triton kernel (#3798) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add amp_foreach_non_finite_check_and_unscale operator with Triton kernel (#3799) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add channel_shuffle operator with Triton kernel (#3800) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add special_shifted_chebyshev_polynomial_u operator with Triton kernel (#3801) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add special_log_softmax operator with Triton kernel (#3803) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add broadcast_tensors operator with Triton kernel (#3804) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add special_hermite_polynomial_h operator with Triton kernel (#3807) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add _jagged_to_padded_dense_forward operator with Triton kernel (#3809) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add _fused_adam operator with Triton kernel (#3810) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add unbind_copy operator with Triton kernel (#3811) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add _upsample_bilinear2d_aa operator with Triton kernel (#3813) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add index_select_backward operator with Triton kernel (#3817) by @XDYuanzhuLee
  • [SUNRISE] add sunrise's constraint (#3822) by @Dayrker
  • [KernelGen][Nvidia] Add acosh operator with Triton kernel (#3852) by @bwbwzzz
  • [KernelGen][Nvidia] Add addcdiv_ operator with Triton kernel (#3857) by @bwbwzzz
  • [KernelGen][Nvidia] Add _resize_output operator with Triton kernel (#3860) by @Dingxingdi
  • [KernelGen][Nvidia] Add irshift operator with Triton kernel (#3862) by @bwbwzzz
  • [KernelGen][Nvidia] Add _upsample_nearest_exact2d operator with Triton kernel (#3866) by @Dingxingdi
  • [KernelGen][Nvidia] Add threshold_ operator with Triton kernel (#3867) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add subtract_ operator with Triton kernel (#3870) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add special_ndtr operator with Triton kernel (#3871) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add special_gammainc operator with Triton kernel (#3872) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add and operator with Triton kernel (#3873) by @Dingxingdi
  • [KernelGen][Nvidia] Add log_sigmoid_forward operator with Triton kernel (#3874) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add arctan and arctan_ operator with Triton kernel (#3875) by @bwbwzzz
  • [KernelGen][Nvidia] Add broadcast_to operator with Triton kernel (#3882) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add _add_relu operator with Triton kernel (#3883) by @Dingxingdi
  • [KernelGen][Nvidia] Add expand operator with Triton kernel (#3885) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add unsqueeze operator with Triton kernel (#3886) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add _sparse_semi_structured_addmm operator with Triton kernel (#3888) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add diagonal_copy operator with Triton kernel (#3889) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add _conj operator with Triton kernel (#3891) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add _unsafe_view operator with Triton kernel (#3892) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add _sparse_semi_structured_mm operator with Triton kernel (#3893) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add linalg_cholesky operator with Triton kernel (#3894) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add bucketize operator with Triton kernel (#3895) by @Dingxingdi
  • [KernelGen][Nvidia] Add squeeze_copy operator with Triton kernel (#3898) by @bwbwzzz
  • [KernelGen][Nvidia] Add _embedding_bag_dense_backward operator with Triton kernel (#3900) by @XDYuanzhuLee
  • Add containerfile for tsingmicro (#3901) by @tengqm
  • [KernelGen][Nvidia] Add _pdist_backward operator with Triton kernel (#3902) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add _thnn_fused_lstm_cell operator with Triton kernel (#3904) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add _prelu_kernel operator with Triton kernel (#3905) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add adaptive_max_pool3d_backward operator with Triton kernel (#3906) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add _thnn_fused_lstm_cell_backward_impl operator with Triton kernel (#3907) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add alpha_dropout operator with Triton kernel (#3908) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add amin operator with Triton kernel (#3909) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add arccos operator with Triton kernel (#3910) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add arcsin operator with Triton kernel (#3912) by @XDYuanzhuLee
  • Add containerfile for Kunlunxin backend (#3914) by @tengqm
  • [KernelGen][Nvidia] Add asin operator with Triton kernel (#3915) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add dequantize operator with Triton kernel (#3918) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add erfinv_ operator with Triton kernel (#3919) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add Concat operator with Triton kernel (#3920) by @XDYuanzhuLee
  • [SiliconFlow] Add mode operation (#3921) by @fpzh2011
  • [KernelGen][Nvidia] Add greater_equal operator with Triton kernel (#3922) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add isposinf operator with Triton kernel (#3923) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add copysign_ operator with Triton kernel (#3924) by @Dingxingdi
  • [KernelGen][Nvidia] Add log2 operator with Triton kernel (#3926) by @bwbwzzz
  • [KernelGen][Nvidia] Add kthvalue operator with Triton kernel (#3928) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add lgamma_ operator with Triton kernel (#3930) by @XDYuanzhuLee
  • [Performance Optimization] Add QUICK_MODE support to reduce test suite runtime (#3932) by @git-flyer
  • [KernelGen][Nvidia] Add less operator with Triton kernel (#3933) by @Dingxingdi
  • [KernelGen][Nvidia] Add miopen_batch_norm operator with Triton kernel (#3936) by @bwbwzzz
  • [c++ wrapper] Add gemma_rms_norm (#3942) by @chen98038
  • [SiliconFlow] Add and optimize the Ascend implementation of upsample_linear1d_backward operator (#3944) by @winfan1314
  • [KernelGen][Nvidia] Add _pdist_forward operator with Triton kernel (#3948) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add lift operator with Triton kernel (#3950) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add max_unpool2d operator with Triton kernel (#3952) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add new_ones operator with Triton kernel (#3954) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add nextafter operator with Triton kernel (#3955) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add Range operator with Triton kernel (#3963) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add sinc operator with Triton kernel (#3964) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add special_erfcx operator with Triton kernel (#3965) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add special_chebyshev_polynomial_w operator with Triton kernel (#3967) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add special_round operator with Triton kernel (#3969) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add special_scaled_modified_bessel_k1 operator with Triton kernel (#3971) by @XDYuanzhuLee
  • [GOSIM][Fused][Nvidia] Add mrope operator with Triton kernel (#3976) by @LoserCheems
  • [KernelGen][Nvidia] Add special_sinc operator with Triton kernel (#3977) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add lshift operator with Triton kernel (#3978) by @bwbwzzz
  • [GOSIM][Hygon] Add top_k_per_row_prefill fused kernel for sparse attention (#3981) by @LoserCheems
  • [KernelGen][Nvidia] Add acos_ operator with Triton kernel (#3985) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add multiply_ operator with Triton kernel (#3987) by @bwbwzzz
  • [KernelGen][Nvidia] Add narrow_copy operator with Triton kernel (#3992) by @bwbwzzz
  • [KernelGen][Nvidia] Add _masked_scale operator with Triton kernel (#3993) by @XDYuanzhuLee
  • [GOSIM][KernelGen][Iluvatar] Add top_k_per_row_prefill fused kernel for sparse attention (#3999) by @LoserCheems
  • [GOSIM][KernelGen][MetaX] Add top_k_per_row_prefill fused kernel for sparse attention (#4000) by @LoserCheems
  • [KernelGen][Nvidia] Add arccosh_ operator with Triton kernel (#4002) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add _native_batch_norm_legit_functional operator with Triton kernel (#4003) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add grid_sampler_3d operator with Triton kernel (#4004) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add arctan2 operator with Triton kernel (#4006) by @bwbwzzz
  • [KernelGen][Nvidia] Add special_log1p operator with Triton kernel (#4007) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add special_logit operator with Triton kernel (#4008) by @XDYuanzhuLee
  • [KernelGen][Tianshu] Add scatter_add_ backend specialization (#4018) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add _cdist_forward operator with Triton kernel (#4022) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add _convert_weight_to_int4pack operator with Triton kernel (#4025) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add _make_dep_token operator with Triton kernel (#4026) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add _nested_view_from_buffer_copy operator with Triton kernel (#4027) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add _scaled_dot_product_efficient_attention operator with Triton kernel (#4028) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add _scaled_dot_product_fused_attention_overrideable operator with Triton kernel (#4029) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add fft_irfftn operator with Triton kernel (#4031) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add linalg_svdvals operator using existing SVD kernel (#4033) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add special_bessel_j0 operator with Triton kernel (#4039) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add special_chebyshev_polynomial_u operator with Triton kernel (#4040) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add erfc and special_erfc operator with Triton kernel (#4041) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add special_legendre_polynomial_p operator with Triton kernel (#4042) by @XDYuanzhuLee
  • Add setup support for CANN9.0.0 for Ascend backend (#4046) by @tengqm
  • Add script for checking GPU availability on Enflame (#4078) by @tengqm
  • [KernelGen][Nvidia] Add bitwise_left_shift_ operator with Triton kernel (#4079) by @bwbwzzz
  • add fused_indexer_q_rope_quant op (#4081) by @huangyiqun
  • [KernelGen][Nvidia] Add binary_cross_entropy operator with Triton kernel (#4091) by @bwbwzzz
  • [KernelGen][Nvidia] Add fp8_fp4_mqa_logits Triton kernel (#4092) by @yzw1128
  • Add a new env shell script to source for Ascend (#4094) by @tengqm
  • Add script for GPU availability check on Sunrise (#4102) by @tengqm
  • Add code owners (#4109) by @0x45f
  • [KernelGen][Nvidia] Add float_power_ operator with Triton kernel (#4141) by @Dingxingdi
  • [KernelGen][Nvidia] Add gather_block_quantized operator with Triton kernel (#4142) by @Dingxingdi
  • [KernelGen][Nvidia] Add igamma_ operator with Triton kernel (#4145) by @Dingxingdi
  • [KernelGen][Nvidia] Add linalg_ldl_factor operator with Triton kernel (#4146) by @Dingxingdi
  • [KernelGen][Nvidia] Add linalg_ldl_factor_ex operator with Triton kernel (#4147) by @Dingxingdi
  • [KernelGen][Nvidia] Add bitwise_right_shift_ operator with Triton kernel (#4148) by @bwbwzzz
  • [KernelGen][Nvidia] Add fp8_fp4_paged_mqa_logits operator with Triton kernel VLLM (#4157) by @yzw1128
  • Add stage_deepseek_v4_mega_moe_inputs OP (#4158) by @huangyiqun
  • [KernelGen][Nvidia] Add _linalg_eigvals operator with Triton kernel (#4217) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add special_xlogy operator with Triton kernel (#4218) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add _functional_sym_constrain_range operator with Triton kernel (#4219) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add nextafter_ operator with Triton kernel (#4231) by @XDYuanzhuLee
  • Add FlagTune configs for compute global topk lens (#4234) by @Fattysand
  • Add benchmark test for addmv_out (#4240) by @tengqm
  • [KernelGen][Nvidia] Add cudnn_batch_norm_backward operator with Triton kernel (#4250) by @bwbwzzz
  • [KernelGen][Nvidia] Add linalg_ldl_solve operator with Triton kernel (#4251) by @Dingxingdi
  • [KernelGen][Nvidia] Add linalg_slogdet operator with Triton kernel (#4255) by @Dingxingdi
  • [KernelGen][Nvidia] Add special_expit operator with Triton kernel (#4263) by @XDYuanzhuLee
  • Add Iluvatar-specific CodeGenConfig for pointwise_dynamic ops (#4265) by @tengqm
  • Add i0 out tests (#4268) by @huangyiqun
  • Add gather backward tests (#4269) by @huangyiqun
  • Add hardsigmoid out tests (#4272) by @huangyiqun
  • Add div_out tests (#4274) by @huangyiqun
  • Add div_scalar inplace tests and benchmark (#4275) by @huangyiqun
  • [KernelGen][Nvidia] Add linear_backward operator with Triton kernel (#4277) by @Dingxingdi
  • [KernelGen][Nvidia] Add _unsafe_masked_index_put_accumulate operator with Triton kernel (#4278) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Add logcumsumexp operator with Triton kernel (#4283) by @Dingxingdi
  • [KernelGen][Nvidia] Add lstm operator with Triton kernel (#4288) by @Dingxingdi
  • [KernelGen][Nvidia] Add lt_ operator with Triton kernel (#4306) by @Dingxingdi
  • [KernelGen][Nvidia] Add MatMulAdd operator with Triton kernel (#4309) by @Dingxingdi
  • [KernelGen][Nvidia] Add sym_size operator with Triton kernel (#4312) by @bwbwzzz
  • Add containerfile for Cambricon 4.4.3 (#4314) by @tengqm
  • [KernelGen][Nvidia] Add sym_stride operator with Triton kernel (#4316) by @bwbwzzz
  • [KernelGen][Nvidia] Add special_gammaln operator with Triton kernel (replacement for #3968) (#4317) by @XDYuanzhuLee
  • Add cudagraph benchmark mode (#4321) by @0x45f
  • Add fp8_fp4_mega_moe OP (#4322) by @huangyiqun
  • [KernelGen][Nvidia] Add expand_copy operator with Triton kernel (#4324) by @bwbwzzz
  • Add GPU check script for Cambricon MLU backend (#4327) by @tengqm
  • Add env cambricon (#4329) by @tengqm
  • [KernelGen][Nvidia] Add max_unpool3d operator with Triton kernel (#4341) by @Dingxingdi
  • [KernelGen][Nvidia] Add median operator metadata (#4348) by @Dingxingdi
  • [KernelGen][Nvidia] Add miopen_batch_norm_backward operator with Triton kernel (#4353) by @Dingxingdi
  • Add FlagTune configs for _mxfp4_moe_gemm_kernel (#4354) by @Jennie1004
  • [KernelGen][Nvidia] Add mish operator with Triton kernel (#4355) by @Dingxingdi
  • [KernelGen][Nvidia] Add mish_backward operator with Triton kernel (#4362) by @Dingxingdi
  • [KernelGen][Nvidia] Add cudnn_convolution_transpose operator with Triton kernel (#4363) by @bwbwzzz
  • [KernelGen][Nvidia] Add cumsum_ operator with Triton kernel (#4365) by @bwbwzzz
  • [KernelGen][Nvidia] Add bitwise_xor operator with Triton kernel (#4366) by @LoserCheems
  • [HYGON] Add C++ wrapper support on HCU (#4402) by @fangzexian
  • Add Python LIBDIR to LD_LIBRARY_PATH for Iluvatar backend (#4414) by @tengqm
  • Add FlagTune configs for w8a8_block_fp8_bmm (#4423) by @Jennie1004
  • add special_i1_out testcase (#4424) by @huangyiqun
  • add remainder_scalar tests & benchmark (#4425) by @huangyiqun
  • Add missing KernelGen labels to 4 operators in operators.yaml (#4429) by @lukatao
  • Add project governance document and update maintainers (#4430) by @tengqm
  • [KernelGen][Nvidia] Add frexp operator with Triton kernel (#4433) by @bwbwzzz
  • add outplace_fused_experts op tests & benchmark (#4440) by @huangyiqun
  • add normed_cumsum op tests & benchmark (#4441) by @huangyiqun
  • add normal_float_float_ op tests & benchmark (#4442) by @huangyiqun
  • [KMCompiler][Nvidia] Add aten._unsafe_index operator (#4443) by @Caeruleann
  • add moe_align_block_size_triton op tests & benchmark (#4444) by @huangyiqun
  • add logit_out op tests & benchmark (#4445) by @huangyiqun
  • Add erfinv fallback for CPU backend compatibility (#4469) by @yzw1128
  • [KernelGen][Nvidia] Add binary_cross_entropy_with_logits operator with Triton kernel (#4475) by @bwbwzzz
  • [KernelGen][Nvidia] Add bf16_paged_mqa_logits fused operator with Triton kernel (#4480) by @Yangyang0906C
  • [KernelGen][Nvidia] Add fused_deepseek_v4_qnorm_rope_kv_rope_insert operator with Triton kernel (#4485) by @Yangyang0906C
  • [KernelGen][Nvidia] Add less_equal operator with Triton kernel (#4487) by @bwbwzzz
  • [KernelGen][Nvidia] Add mvlgamma_ operator with Triton kernel (#4493) by @Dingxingdi
  • [KMCompiler][Nvidia]Add Polygamma Operator (#4500) by @cheersluvs
  • [KernelGen][Nvidia] Add scatter_add operator with Triton kernel (#4502) by @chx7514
  • Add mul to parallel BLAS benchmark (#4504) by @liuxin2560
  • [KernelGen][Nvidia] Add subtract operator with Triton kernel (#4512) by @chx7514
  • add tests & benchmark of div_tensor_mode_,div_tensor_mode (#4544) by @huangyiqun
  • feat: add container/configs.yaml for unified container build config (#4545) by @tengqm
  • add tests & benchmark of div_scalar_mode_,div_scalar_mode (#4546) by @huangyiqun
  • feat: add tools/sync_pyproject_extras.py to detect extras drift (#4547) by @tengqm
  • [KernelGen][Nvidia] Add narrow operator with Triton kernel (#4561) by @Yukun-Cui
  • [KernelGen][Nvidia] Add unbind operator with Triton kernel (#4563) by @Yukun-Cui
  • Add logaddexp2 and xlogy pointwise operators (#4571) by @zheng1
  • [Iluvatar] Add version-aware triton_extra_name for libdevice selection. (#4590) by @awayzjj
  • [KernelGen][Nvidia] Add transpose operator with Triton kernel (#4591) by @chx7514
  • feat: record test start/end/total time in run_tests.py summary and log (#4609) by @wuwentao
  • feat: add unified Containerfile and build.py for all backends (#4619) by @tengqm
  • Add softplus_backward op (#4623) by @0x45f
  • Add MetaX mul implementation (#4630) by @Tiannn967
  • feat: integrate env vars into backends.yaml, auto-write to activate (#4633) by @tengqm
  • Add weekly test workflow (#4639) by @xmhubj
  • Add te rmsnorm_fwd and rmsnorm_bwd op (#4640) by @0x45f
  • Add floor fallback for Sunrise backend compatibility (#4650) by @kkkwb
  • [KMCompiler] Add Triton nonzero_static operator (#4661) by @chenkc2025
  • [KernelGen][T-Head] Add triton_unified_attention paged attention kernel (#4664) by @Yangyang0906C
  • Add mthreads 5.2.0 environment (#4668) by @tengqm
  • Add probe-node workflow to inspect self-hosted runner environment (#4733) by @tengqm
  • Add float64 support for vector_norm (#4739) by @douxetpur
  • [AUTOTUNER] Add (XGB+GA) FlagTune support (via FlagTree) (#4740) by @HenryRenYz
  • add dispatch_fused_moe_kernel testcase (#4741) by @huangyiqun
  • add cumsum_out testcase (#4742) by @huangyiqun
  • add alias_copy_out testcase (#4743) by @huangyiqun
  • [KernelGen][Metax] Add conv_depthwise2d operator for Metax backend (#4762) by @yzw1128
  • Add scalar tensor op (#4803) by @0x45f
  • [KMCompiler][NVIDIA][Hygon] Add nansum operator with triton kernel (#4804) by @Willing2329
  • [KMCompiler][Ascend] Add polygamma NPU backend support (#4809) by @cheersluvs
  • [KernelGen][Nvidia] Add linalg_svd operator with Triton kernel (#4817) by @Yukun-Cui
  • feat: test script support new vendor enflame (#4836) by @wuwentao
  • [KMCompiler]Add linalg_cross operator (#4850) by @elllzat
  • [KMCompiler] Add igammac operator with Triton kernel (#4860) by @LittleShun1214
  • feat: weekly ci job support vendors and ops arg (#4862) by @wuwentao
  • add Dao-style standard FHT accuracy tests (新增 Dao 风格标准 FHT 精度测试) (#4863) by @dawnyang050-debug
  • add var_dim benchmark testcases (#4883) by @huangyiqun
  • Add test pandas deps (#4919) by @0x45f
  • Add post_layer_norm_residual fused operator (#4930) by @ZYX223
  • [KMCompiler][Nvidia] add operator linalg_lu_factor (#4945) by @YangLong114514
  • [KMCompiler][Nvidia] Add pairwise_distance operator with triton kernel. (#4949) by @Caeruleann
  • [KMCompiler][NVIDIA][Ascend][Hygon][Thead][MetaX][iluvatar] Add sparse_sampled_addmm operator with Triton kernel (#4950) by @littlekid-zxt
  • [KMCompiler][Nvidia]Add masked_scatter_backend op (#4954) by @lelongmask-cmp
  • [Runtime][AMD] Add RDNA4 autotune configs and match arch by exact GPU target (#4991) by @WhatGhost
  • [KMCompiler][NVIDIA] Add NVIDIA cholesky solve operator (#5011) by @YangLong114514
  • [KMCompiler]Add log_sigmoid_backward operator (#5012) by @elllzat
  • [KernelGen][Nvidia] Add unsafe_chunk operator with Triton kernel (#5028) by @LoserCheems
  • [KernelGen][Nvidia] Add unflatten operator with Triton kernel (#5029) by @LoserCheems
  • [KernelGen][Nvidia] Add chunk operator (#5033) by @LoserCheems
  • [KernelGen][Nvidia] Add unfold operator (#5037) by @LoserCheems
  • [KernelGen][Nvidia] Add flatten operator (#5039) by @LoserCheems
  • [KernelGen][Nvidia] Add expand_as operator (#5041) by @LoserCheems
  • [KernelGen][Nvidia] Add alias operator (#5043) by @LoserCheems
  • Add wheel to build tools (#5047) by @tengqm
  • Add wheel to setup.sh for some backends (#5048) by @tengqm
  • [KernelGen][Nvidia] Add addbmm operator with Triton kernel (#5055) by @LoserCheems
  • Add logging.debug in common ops (#5057) by @0x45f
  • [KernelGen][Nvidia] Add block_diag operator with Triton kernel (#5068) by @LoserCheems
  • [KernelGen][Nvidia] Add lu_unpack operator with Triton kernel (#5072) by @LoserCheems
  • [KernelGen][Nvidia] Add _fused_rms_norm_backward operator with Triton kernel (#5074) by @Yukun-Cui
  • [KernelGen][Nvidia] Add rshift operator with Triton kernel (#5075) by @Yukun-Cui
  • [KernelGen][Nvidia] Add _dyn_quant_pack_4bit_weight operator with Triton kernel (#5076) by @Yukun-Cui
  • [KernelGen][Nvidia] Add adaptive_max_pool2d operator with Triton kernel (#5077) by @Yukun-Cui
  • [KernelGen][Nvidia] Add replication_pad3d_backward operator with Triton kernel (#5078) by @Yukun-Cui
  • [KernelGen][Nvidia] Add adaptive_max_pool2d_backward operator with Triton kernel (#5079) by @Yukun-Cui
  • [KernelGen][Nvidia] Add _weight_norm operator with Triton kernel (#5080) by @Yukun-Cui
  • [KernelGen][Nvidia] Add _cholesky_solve_helper operator with Triton kernel (#5081) by @Yukun-Cui
  • [KernelGen][Nvidia] Add as_strided_scatter operator with Triton kernel (#5082) by @Yukun-Cui
  • [KernelGen][Nvidia] Add _adaptive_avg_pool2d_backward operator with Triton kernel (#5084) by @Yukun-Cui
  • [KernelGen][Nvidia] Add binary_cross_entropy_backward operator with Triton kernel (#5090) by @LoserCheems
  • [Runtime][AMD] Add RDNA4 autotune configs for layer_norm and rms_norm (#5100) by @WhatGhost
  • [KernelGen][Nvidia] Add _native_batch_norm_legit operator with Triton kernel (#5104) by @xuanzhengdu-eng
  • [KernelGen][Nvidia] Add eq_ operator with Triton kernel (#5108) by @ShawnsYing
  • [KernelGen][Nvidia] Add special_modified_bessel_i0 operator with Triton kernel (#5110) by @bwbwzzz
  • [KernelGen][Nvidia] Add special_log_ndtr operator with Triton kernel (#5114) by @bwbwzzz
  • [KernelGen][Nvidia] Add hardsigmoid_backward operator with Triton kernel (#5116) by @bwbwzzz
  • [KernelGen][Nvidia] Add hardswish_backward operator with Triton kernel (#5117) by @bwbwzzz
  • [KernelGen][Nvidia] Add hardtanh_backward operator with Triton kernel (#5118) by @bwbwzzz
  • [KernelGen][Nvidia] Add heaviside operator with Triton kernel (#5120) by @bwbwzzz
  • [KernelGen][Nvidia] Add _reshape_alias operator with Triton kernel (#5121) by @xuanzhengdu-eng
  • [KernelGen][Nvidia] Add le_ operator with Triton kernel (#5122) by @ShawnsYing
  • [KernelGen][Nvidia] Add less_equal_ operator with Triton kernel (#5124) by @ShawnsYing
  • [KernelGen][Nvidia] Add less_ operator with Triton kernel (#5126) by @ShawnsYing
  • [KMCompiler][NVIDIA][Ascend][Hygon][Thead][MetaX][iluvatar] Add linalg_det operator with Triton kernel (#5136) by @littlekid-zxt
  • add Devcontainer support (#5137) by @Salvationhope
  • [KernelGen] Add leaky_relu_backward operator (#5145) by @ShawnsYing
  • [KMCompiler][NVIDIA][Hygon]Add replication_pad2d_backward operator with triton kernel (#5146) by @Willing2329
  • [KernelGen][Nvidia] Add lift_fresh operator with Triton kernel (#5148) by @ShawnsYing
  • [KernelGen][Nvidia] Add _fused_rms_norm operator with Triton kernel (#5152) by @chx7514
  • [KernelGen][Nvidia] Add max_pool2d_with_indices_backward operator with Triton kernel (#5159) by @chx7514
  • [KernelGen][Nvidia] Add _cudnn_attention_forward operator with Triton kernel (#5160) by @yzw1128
  • [KernelGen][Nvidia] Add cholesky_inverse operator with Triton kernel (#5161) by @LoserCheems
  • [KernelGen][Nvidia] Add divide operator with Triton kernel (#5162) by @CoyeCHEN
  • [KernelGen][Nvidia] Add _scaled_dot_product_flash_attention operator with Triton kernel (#5164) by @CoyeCHEN
  • [KernelGen][Nvidia] Add reflection_pad2d_backward operator with Triton kernel (#5165) by @chx7514
  • [KernelGen][Nvidia] Add upsample_bilinear2d operator with Triton kernel (#5167) by @chx7514
  • [KernelGen][Nvidia] Add _scaled_dot_product_attention_math operator with Triton kernel (#5168) by @chx7514
  • [KernelGen][Nvidia] Add replication_pad2d operator with Triton kernel (#5172) by @chx7514
  • [KMCompiler][Ascend]Add linalg_lstsq NPU backend support (#5173) by @cheersluvs
  • [KernelGen][Nvidia] Add _flash_attention_forward operator with Triton kernel (#5175) by @CoyeCHEN
  • [KernelGen][Nvidia] Add native_batch_norm operator with Triton kernel (#5177) by @CoyeCHEN
  • [KernelGen][Nvidia] Add native_group_norm operator with Triton kernel (#5178) by @CoyeCHEN
  • [KernelGen][Nvidia] Add linalg_householder_product operator with Triton kernel (#5179) by @LoserCheems
  • [KernelGen][Nvidia] Add native_layer_norm operator with Triton kernel (#5180) by @CoyeCHEN
  • [KernelGen][Nvidia] Add _scaled_dot_product_cudnn_attention operator (#5182) by @chx7514
  • [KernelGen][Nvidia] Add mvlgamma operator (#5190) by @NineAnnAnn
  • [KernelGen][Nvidia] Add multiply operator (#5191) by @NineAnnAnn
  • [KernelGen][Nvidia] Add negative_ operator (#5192) by @NineAnnAnn
  • [KernelGen][Nvidia] Add ne_ operator (#5193) by @NineAnnAnn
  • [KernelGen][Nvidia] Add resize_output operator (#5194) by @NineAnnAnn
  • [KernelGen][Nvidia] Add nan_to_num_ operator (#5195) by @NineAnnAnn
  • [KernelGen][Nvidia] Add addmv_ operator (#5196) by @NineAnnAnn
  • [KernelGen][Nvidia] Add xlogy_ operator (#5197) by @NineAnnAnn
  • [KernelGen][Nvidia] Add atan2_ operator (#5198) by @NineAnnAnn
  • [KernelGen][Nvidia] Add baddbmm_ operator (#5199) by @NineAnnAnn
  • [KernelGen][Nvidia] Add not_equal_ operator (#5200) by @NineAnnAnn
  • [KMCompiler] Add linalg_solve_triangular operator with Triton kernel (#5201) by @LittleShun1214
  • [KernelGen][Nvidia] Add true_divide operator with Triton kernel (#5205) by @CoyeCHEN
  • [KernelGen][Nvidia] Add true_divide_ operator with Triton kernel (#5206) by @CoyeCHEN
  • [KMCompiler] [Nvidia] add linalg_lu_factor_ex operator with Triton kernel (#5209) by @YangLong114514
  • add logger debug (#5210) by @douxetpur
  • [KernelGen][Nvidia] Add quantized_lstm operator with Triton kernel (#5211) by @LoserCheems
  • [KMCompiler][iluvatar] Add Iluvatar cholesky solve backend (#5242) by @YangLong114514
  • [KernelGen][Nvidia] Add _native_batch_norm_legit_no_training operator (#5243) by @ShawnsYing
  • [KMCompiler][MetaX] Add MetaX cholesky solve backend (#5247) by @YangLong114514
  • [KernelGen][Nvidia] Add _functional_assert_async operator with Triton kernel (#5249) by @xuanzhengdu-eng
  • [KernelGen][Nvidia] Add ormqr operator with Triton kernel (#5252) by @LoserCheems
  • [KernelGen][Nvidia] Add _fused_moving_avg_obs_fq_helper operator with Triton kernel (#5255) by @xuanzhengdu-eng
  • [KernelGen][Nvidia] Add max_pool3d_with_indices_backward operator with Triton kernel (#5264) by @LoserCheems
  • [KernelGen][Nvidia] Add _list_to_tensor operator with Triton kernel (#5273) by @xuanzhengdu-eng
  • [KernelGen][Nvidia] Add special_erf operator with Triton kernel (#5274) by @CoyeCHEN
  • [KernelGen][Nvidia] Add special_exp2 operator with Triton kernel (#5275) by @CoyeCHEN
  • [KernelGen][Nvidia] Add special_i1e operator with Triton kernel (#5276) by @CoyeCHEN
  • [KernelGen][Nvidia] Add sym_storage_offset operator with Triton kernel (#5279) by @CoyeCHEN
  • [KernelGen][Nvidia] Add _has_compatible_shallow_copy_type operator (#5280) by @xuanzhengdu-eng
  • [KernelGen][Nvidia] Add grid_sampler_3d_backward operator with Triton kernel (#5287) by @LoserCheems
  • [KernelGen][Nvidia] Add _upsample_bilinear2d_aa_backward operator wit… (#5289) by @NineAnnAnn
  • [KernelGen][Nvidia] add _cudnn_rnn_backward (single-layer unidirectional LSTM) (#5291) by @ShawnsYing
  • [KMCompiler][Ascend] Add Ascend cholesky solve backend (#5299) by @YangLong114514
  • [KernelGen][Nvidia] Add special_gammaincc operator with Triton kernel (#5315) by @CoyeCHEN
  • [KernelGen][Nvidia] Add special_multigammaln operator with Triton kernel (#5316) by @CoyeCHEN
  • [KernelGen][Nvidia] Add empty_permuted operator with Triton kernel (#5317) by @CoyeCHEN
  • [KernelGen][Nvidia] Add special_ndtri operator with Triton kernel (#5319) by @CoyeCHEN
  • [KernelGen][Nvidia] Add _weight_int4pack_mm_with_scales_and_zeros ope… (#5323) by @NineAnnAnn
  • [KernelGen][Nvidia] Add value_selecting_reduction_backward operator with Triton kernel (#5326) by @chx7514
  • [AMD] Add RDNA4 split softmax (#5327) by @WhatGhost
  • [KernelGen][Nvidia] Add _batch_norm_impl_index operator with Triton kernel (#5335) by @YufanPeter
  • [KernelGen][Nvidia] Add addbmm_ operator with Triton kernel (#5340) by @LoserCheems
  • [KernelGen][Nvidia] add mkldnn_rnn_layer (single-layer LSTM) Triton kernel (#5355) by @ShawnsYing
  • [KernelGen][Nvidia] Add sym_constrain_range operator (#5367) by @ShawnsYing
  • [KernelGen] Add weight_int8pack_mm operator (#5371) by @NineAnnAnn
  • [KernelGen][Nvidia] Add view_as_complex operator (view operation, zero-copy) (#5373) by @CoyeCHEN
  • [KMCompiler][Ascend][MetaX] Add nansum operator with triton kernel (#5377) by @Willing2329
  • [KMCompiler][Ascend] Add operator linalg_lu_factor for ascend backend. (#5382) by @YangLong114514
  • [KernelGen][Nvidia] Add slice operator with Triton kernel (#5393) by @ShawnsYing
  • [KMCompiler][Ascend][Iluvatar] Add ascend backend for pairwise_distance op, and fix iluvatar cpu reference error. (#5394) by @Caeruleann
  • [KernelGen][Nvidia] Add special_bessel_y0 operator with Triton kernel (#5398) by @CoyeCHEN
  • [KernelGen][Nvidia] Add special_shifted_chebyshev_polynomial_t operator with Triton kernel (#5400) by @CoyeCHEN
  • [KMCompiler][Nvidia][MetaX] Add dist operator support (#5406) by @Keriza
  • [KMCompiler][Nvidia] add linalg_vecdot operator with Triton kernel (#5416) by @rye985
  • [KernelGen][Nvidia] Add ctc_loss operator YAML registration (#5419) by @ShawnsYing
  • [KernelGen][Nvidia] Add cdist operator with Triton kernel (#5420) by @ShawnsYing
  • [KernelGen][Nvidia] Add hsplit operator with view implementation (#5421) by @ShawnsYing
  • [KernelGen][Nvidia] Add avg_pool1d operator with dimensionality reduction (#5422) by @ShawnsYing
  • [KernelGen] Add dsplit operator with view implementation (#5426) by @ShawnsYing
  • [KMCompiler]feat(conj_physical): add hygon runtime and ascend tune config (#5434) by @tspyc072
  • [KernelGen][Nvidia] Add addr_ operator with Triton kernel (#5440) by @LoserCheems
  • [KernelGen][Nvidia] Add fake_quantize_per_channel_affine operator with Triton kernel (#5450) by @xuanzhengdu-eng
  • [KernelGen][Nvidia] Add fake_quantize_per_channel_affine_cachemask operator with Triton kernel (#5455) by @xuanzhengdu-eng
  • [KernelGen][Nvidia] Add fake_quantize_per_channel_affine_cachemask_backward operator with Triton kernel (#5460) by @xuanzhengdu-eng
  • [KernelGen][Nvidia] Add fake_quantize_per_tensor_affine operator with Triton kernel (#5461) by @xuanzhengdu-eng
  • [KMCompiler][Ascend] add operator linalg_lu_factor_ex for ascend backend. (#5462) by @YangLong114514
  • [KMCompiler][NVIDIA][Ascend][Hygon][Thead][MetaX][iluvatar] add log_normal operator with Triton kernel (#5472) by @littlekid-zxt
  • [KernelGen][Nvidia] Add fake_quantize_per_tensor_affine_cachemask_backward operator with Triton kernel (#5473) by @xuanzhengdu-eng
  • [KernelGen][Nvidia] Add bilinear operator with Triton kernel (#5480) by @NineAnnAnn
  • [KernelGen][Nvidia] Add blackman_window operator with Triton kernel (#5485) by @NineAnnAnn
  • [KernelGen][Nvidia] Add binomial operator with Triton kernel (#5486) by @NineAnnAnn
  • [KernelGen][Nvidia] Add fill_diagonal_ operator with Triton Kernel (#5487) by @xuanzhengdu-eng
  • [KernelGen][Nvidia] Add igamma operator with Triton kernel (#5488) by @bwbwzzz
  • [KernelGen][Nvidia] Add fliplr operator with Triton Kernel (#5491) by @xuanzhengdu-eng
  • [KernelGen][Nvidia] Add chalf operator with Triton kernel (#5494) by @NineAnnAnn
  • [KernelGen][Nvidia] Add corrcoef operator with Triton kernel (#5497) by @ShawnsYing
  • [KernelGen][Nvidia] Add add_relu operator with Triton kernel (#5502) by @Yukun-Cui
  • [KernelGen][Nvidia] Add amp_update_scale operator with Triton kernel (#5503) by @Yukun-Cui
  • [KernelGen][Nvidia] Add _batch_norm_impl_index_backward operator with Triton kernel (#5504) by @Yukun-Cui
  • [KernelGen][Nvidia] Add _convolution_double_backward operator with Triton kernel (#5505) by @Yukun-Cui
  • [KMCompiler][Ascend] Add ascend backend for replication_pad2d_backward (#5508) by @Willing2329
  • [KernelGen][Nvidia] Add _convolution_mode operator with Triton kernel (#5512) by @Yukun-Cui
  • [KernelGen][Nvidia] Add _conj_copy operator with Triton kernel (#5513) by @Yukun-Cui
  • [KernelGen][Nvidia] Add coalesced operator with Triton kernel (#5514) by @Yukun-Cui
  • [KernelGen][Nvidia] Add _compute_linear_combination operator with Triton kernel (#5515) by @Yukun-Cui
  • [KernelGen][Nvidia] Add flipud operator with Triton Kernel (#5516) by @xuanzhengdu-eng
  • [KernelGen][Nvidia] Add float_power operator with Triton Kernel (#5517) by @xuanzhengdu-eng
  • [KMCompiler][Nvidia][Ascend]Add operator linalg_lu with triton kernel. (#5518) by @YangLong114514
  • [KernelGen][Nvidia] Add alpha_dropout_ operator with Triton kernel (#5519) by @LoserCheems
  • [KernelGen][Nvidia] Add cosine_embedding_loss operator with Triton kernel (#5521) by @ShawnsYing
  • [C++ Runtime MLU]Add Cambricon MLU support to C++ operators (#5535) by @chen98038
  • [KernelGen][Nvidia] Add column_stack operator with Triton kernel (#5536) by @NineAnnAnn
  • [KernelGen][Nvidia] Add choose_qparams_optimized operator with Triton… (#5537) by @NineAnnAnn
  • [KernelGen][Nvidia] Add _cummin_helper operator with Triton kernel (#5538) by @chx7514
  • [KernelGen][Nvidia] Add _fake_quantize_learnable_per_tensor_affine_backward operator with Triton kernel (#5540) by @chx7514
  • [KernelGen][Nvidia] Add _fake_quantize_learnable_per_channel_affine_backward operator with Triton kernel (#5543) by @chx7514
  • [KernelGen][Nvidia] Add _fake_quantize_learnable_per_tensor_affine operator with Triton kernel (#5544) by @chx7514
  • [KernelGen][Nvidia] Add adaptive_avg_pool1d operator with Triton kernel (#5550) by @CoyeCHEN
  • [KernelGen][Nvidia] Add fill_mem_eff_dropout_mask operator with Triton kernel (#5552) by @chx7514
  • [KernelGen][Nvidia] Add cov operator with Triton kernel (#5554) by @ShawnsYing
  • [KernelGen][MThreads] Add 20 operators for MThreads (#5556) by @CoyeCHEN
  • [KernelGen][Hygon] Add 25 operators for Hygon (#5557) by @CoyeCHEN
  • [KernelGen][Iluvatar] Add 28 operators for Iluvatar (#5558) by @CoyeCHEN
  • [KernelGen][MetaX] Add 26 operators for MetaX (#5559) by @CoyeCHEN
  • [KernelGen][T-Head] Add 38 operators for T-Head (#5560) by @CoyeCHEN
  • [KernelGen][Nvidia] Add cummaxmin_backward operator with Triton kernel (#5566) by @ShawnsYing
  • [KernelGen][Nvidia] Add special_softmax operator with Triton kernel (#5578) by @bwbwzzz
  • [KMCompiler][NVIDIA][Hygon][Thead][iluvatar] Add exponential operator with Triton kernel (#5579) by @littlekid-zxt
  • [KernelGen][Nvidia] Add unsafe_split_with_sizes operator (#5580) by @chx7514
  • [KMCompile][Nvidia][Ascend][Metax][Iluvatar] Add linalg.qr operator with triton kenrel (#5586) by @Caeruleann
  • [KernelGen][Nvidia] Add linalg_eig operator with Triton kernel (#5588) by @bwbwzzz
  • [KernelGen][Nvidia] Add split_with_sizes operator (#5590) by @chx7514
  • [KMCompiler][Nvidia][Hygon] Add linalg_matrix_norm operator (#5595) by @wuhansky
  • [KernelGen][Nvidia] Add embedding_renorm_ operator with Triton kernel (#5602) by @ShawnsYing
  • [KernelGen][Nvidia] Add conv_tbc_backward operator with Triton kernel (#5605) by @NineAnnAnn
  • [KernelGen][Nvidia] Add conv_transpose3d operator with Triton kernel (#5606) by @NineAnnAnn
  • [KernelGen][Nvidia] Add cumulative_trapezoid operator with Triton kernel (#5618) by @ShawnsYing
  • [KernelGen][Nvidia] Add _fake_quantize_per_tensor_affine_cachemask_tensor_qparams operator with Triton kernel (#5631) by @chx7514
  • [KernelGen][Nvidia] Add _cummax_helper operator with Triton kernel (#5632) by @chx7514
  • [KernelGen][Nvidia] Add _dirichlet_grad operator with Triton kernel (#5633) by @chx7514
  • [KernelGen][Nvidia] Add convolution_overrideable operator with Triton kernel (#5642) by @NineAnnAnn
  • [KMCompiler][Nvidia][Ascend] Add operator index_fill/index_fill_ with triton. (#5643) by @YangLong114514
  • [KMCompiler][Iluvatar][Hygon] Add linalg_matrix_norm operator with triton kernel. (#5648) by @wuhansky
  • [KernelGen][Nvidia] Add _upsample_lanczos2d_aa operator with Triton Kernel (#5651) by @xuanzhengdu-eng
  • [KernelGen][Nvidia] Add _upsample_nearest_exact1d_backward operator with Triton Kernel (#5653) by @xuanzhengdu-eng
  • [KernelGen][Nvidia] Add nll_loss2d operator with Triton kernel (#5660) by @ShawnsYing
  • [KernelGen][Nvidia] Add logit_backward operator with Triton kernel (#5661) by @kkkwb
  • [KernelGen][Nvidia] Add nuclear_norm operator (#5662) by @ShawnsYing
  • [Iluvatar] Add dedicated kernel for repeat_interleave_self_int (#5665) by @awayzjj
  • [KMCompiler][Nvidia][Ascend] Add adaptive_max_pool3d operator (#5671) by @wuhansky
  • [KernelGen][Nvidia] Add _batch_norm_with_update_functional operator with Triton Kernel (#5673) by @kkkwb
  • [KMCompiler][Ascend] add exponential operator with Triton kernel (#5682) by @littlekid-zxt
  • [KernelGen][Nvidia] Add _nested_from_padded_tensor operator with Triton kernel (#5688) by @chx7514
  • [KernelGen][Nvidia] Add _nested_sum_backward operator with Triton kernel (#5690) by @chx7514
  • [KernelGen][Nvidia] Add _nested_tensor_from_mask_left_aligned operator with Triton kernel (#5692) by @chx7514
  • [KernelGen][Nvidia] Add _nested_view_from_jagged operator (#5694) by @chx7514
  • [KernelGen][Nvidia] Add _nested_view_from_jagged_copy operator with Triton kernel (#5695) by @chx7514
  • [KMCompiler]Add stages to the linalg_lstsq operator entry (#5700) by @cheersluvs
  • [KernelGen] Add max_pool1d operator (#5704) by @NineAnnAnn
  • Add 'decorator' to the dependencies for Ascend CANN 8.5.0 backend. (#5707) by @wuhansky
  • Add attrs to the dependency list for Ascend CANN 8.5.0 (#5708) by @wuhansky
  • [KernelGen][Nvidia] Add matrix_exp_backward operator with Triton kernel (#5710) by @NineAnnAnn
  • Add psutil to dependency for CANN 8.5.0 (#5712) by @wuhansky
  • [KernelGen][NVIDIA] Add linalg_matrix_sqrth operator with Triton kernel (#5728) by @xuanzhengdu-eng
  • [KernelGen][NVIDIA] Add _thnn_differentiable_gru_cell_backward operator with Triton kernel (#5734) by @xuanzhengdu-eng
  • [KernelGen][Nvidia] Add sum_to_size operator with Triton kernel (#5761) by @NineAnnAnn
  • [FlagTune] Add multi-platform Mul cost model support (#5762) by @k6699k
  • [SiliconFlow] Add CI test aliases for approved and merged operators (#5788) by @winfan1314
  • Add conj_physical_ operator with contiguous-view Triton kernel (#5800) by @tspyc072
  • [KMCompiler][Ascend] Add ascend backend for linalg_solve_triangular (#5821) by @LittleShun1214
  • feat(tools): add quick mode (#5866) by @103yiran
  • [KMCompiler][Nvidia]Add operator linalg_norm with triton. (#5870) by @YangLong114514
  • add int8 tests for (un)pack_seq_triton (#5886) by @douxetpur
  • [ARM] Add Triton W4A8-G128 linear API for CPU (#5904) by @kevinzs2048
  • feat: enable and enforce _FULL_CONFIG sorting (#5908) by @gavin0x01
  • [KMCompiler][NVIDIA] Add matrix rank operator (#5942) by @YangLong114514
  • feat(setup.sh): install python-build-standalone by mirror (#5948) by @103yiran
  • [KMCompiler][Nvidia][Thead]Add linalg_matrix_power for triton (#5951) by @wuhansky
  • [QC][PPU] Add W8A16 RMSNorm operator (#5957) by @July-h5kf3
  • [KMCompiler][Ascend][Metax]Add linalg_matrix_power for triton (#5960) by @wuhansky
  • [k] add and fix fused kernels (fused) (#5965) by @RobertLuobo
  • [QC][PPU] Add W8A16 FP8 TopK operator (#5966) by @July-h5kf3
  • [k] add and fix math category ops (#5967) by @RobertLuobo
  • [QC][Ascend] Add INT8 W8A16 RMSNorm operator (#5968) by @July-h5kf3
  • [KMCompiler][Nvidia] Add gru operator with Triton kernel (#6002) by @Willing2329
  • [KMCompiler][MetaX] Add gru operator with Triton kernel (#6059) by @Willing2329
  • [KMCompiler][NVIDIA][Ascend][Hygon][Thead][MetaX][iluvatar] add linalg_matrix_exp operator with Triton kernel (#6069) by @littlekid-zxt
  • feat(hash_tensor): add Triton implementation of aten::hash_tensor (#6088) by @ShawnsYing
  • [KMCompiler][Iluvatar] Add gru operator with Triton kernel (#6094) by @Willing2329
  • [KMCompiler][Ascend] Add ascend backend for igammac (#6136) by @wuhansky
  • [KMCompiler][Ascend] Add ascend matrix_rank implement and fix some bug for general matrix_rank ops. (#6161) by @YangLong114514
  • [KernelGen][MThreads] Add fmod_ Moore Threads specialized operator (#6170) by @Yukun-Cui
  • [KernelGen][MThreads] Add conv_transpose1d Moore Threads specialized operator (#6174) by @Yukun-Cui
  • [KernelGen][MThreads] Add upsample_linear1d_backward Moore Threads specialized operator (#6175) by @Yukun-Cui
  • [KernelGen][MThreads] Add matmuladd Moore Threads specialized operator (#6180) by @Yukun-Cui
  • [QC][Hygon] Add topk_w8a16_fp8 (#6190) by @July-h5kf3
  • [KMCompiler][Ascend] Add gru operator with Triton kernel (#6201) by @Willing2329
  • [KMCompiler][Ascend] Add depends (#6205) by @Willing2329
  • [FlagTune] Add common MUL Cost Model support (#6232) (#6345) by @0x45f
  • [Ascend] Add specialized linear operator (#6455) by @modao1234
  • Add truncate option to command.yaml (2be4a2918ad1; commit)
  • Add permissions for test-operator job (2228de61dcd4; commit)
  • Add h100 to supported runner labels in command.yaml (d9d17349db5d; commit)
  • Add RUNNER_SSH_KEY secret to unittest workflow (769fe5fc8f71; commit)
  • Add workflow_dispatch trigger to cpp-op-test.yaml (99ef915e953d; commit)
  • Add concurrency control to unit test workflow (af9e34e433b7; commit)

Performance

  • [Iluvatar] optimize mm (#1988) by @awayzjj
  • Optimize flash varlen paged KV cache addressing (#3564) by @yysheng26
  • [KMCompiler]Perf/cp gather indexer k quant cache (#3686) by @chenkc2025
  • [KMCompiler]Perf/indexer k quant and cache (#3688) by @chenkc2025
  • [QC] Optimize fused_marlin_moe (w4a16) (#3778) by @monellz
  • [QC]Optimize GEMM(fp8) (#3821) by @ggbondbest
  • [SiliconFlow] Optimize upsample_bicubic2d_aa_backward on Ascend NPU and Fix Tests (#4101) by @winfan1314
  • [KMCompiler] Optimize compute_global_topk_indices_and_lens for DeepSeek V4 attention (#4113) by @LittleShun1214
  • [QC] optimize mxfp4 w4a16 fused_moe (#4117) by @monellz
  • [KMCompiler]Optimize pack_seq (#4125) by @chenkc2025
  • [KMCompiler]Optimize unpack_seq (#4126) by @chenkc2025
  • Optimize command workflow (#4161) by @tengqm
  • [KMCompiler]Optimize dequantize_and_gather_k_cache for DeepSeek V4 attention (#4267) by @littlekid-zxt
  • [KMCompiler] Optimize combine_topk_swa_indices for DeepSeek V4 attention (#4282) by @cheersluvs
  • Optimize std dim reduction (#4304) by @douxetpur
  • [KMCompiler]Optimize per-token group FP8 quantization (#4319) by @chenkc2025
  • Optimize run_tests.py script for marks collection (#4374) by @tengqm
  • [QC]Optimize fp8-w8a16 Rmsnorm (#4437) by @ggbondbest
  • optimize upsample_trilinear3d. (#4438) by @awayzjj
  • [Performance Optimization] Optimize split_with_sizes_copy launch overhead (#4456) by @stringl1l1l1l
  • [MTHREADS] Optimize one_hot (#4468) by @Oslomayor
  • Optimize MXFP4 fused Marlin MoE kernels (#4483) by @KingsleyLiu-NV
  • Optimize poisson sampling and fix new_full fill codegen (#4508) by @Salamanca001
  • [Iluvatar] Optimize linear with EVEN_K SME-friendly kernel (#4514) by @Salamanca001
  • [Iluvatar] optimize conv_transpose1d. (#4515) by @awayzjj
  • Optimize index for slice-before-index adjacent tensor indices (#4579) by @ZYX223
  • Optimize Hopper mul broadcast 2D tuning key (#4608) by @liuxin2560
  • Optimize prod single-dim reduction with inner/non-inner kernels (#4618) by @zheng1
  • Optimize Metax mul broadcast 2D tuning key (#4662) by @Tiannn967
  • [HCU] Optimize backend kernel operators for improved performance (#4663) by @wangxshuai
41 more changes
  • Optimize index_add for contiguous suffix layouts (#4670) by @ZYX223
  • Optimize w8a8_block_fp8_bmm (#4794) by @Jennie1004
  • [kunlunxin] optimize existing ops and add new op coverage (#4816) by @RobertLuobo
  • [QC]Optimize FP8 W8A8 bmm (#4880) by @ggbondbest
  • [Ascend] Optimize layer norm forward row scheduling (#4903) by @ZYX223
  • Optimize MXFP4 fused Marlin MoE dequant: fold scale + decode-opt (#4908) by @zeroherolin
  • Optimize MXFP4 Dequantization in Fused Marlin MoE (#4943) by @KingsleyLiu-NV
  • [QC] Optimize flash_mla_with_kvcache for model1 (#5010) by @monellz
  • Optimize wna16 MoE main loop: one mma per K tile instead of eight (#5140) by @zeroherolin
  • [MTHREADS] Optimize flip for multi-dim tensors (#5261) by @Oslomayor
  • [SiliconFlow] Optimize cauchy on Iluvatar and fix cross-backend tests (#5341) by @winfan1314
  • [MTHREADS] optimize conv1d_padding autotune configs (#5343) by @Oslomayor
  • [MTHREADS] optimize conv3d_padding autotune configs (#5353) by @Oslomayor
  • [Ascend] Optimize AddMM layouts and bias epilogue (#5383) by @ZYX223
  • [MThreads] Optimize AddMM layouts and bias epilogue (#5384) by @ZYX223
  • [MetaX] Optimize AddMM layouts and bias epilogue (#5386) by @ZYX223
  • [MTHREADS] Optimize constant_pad_nd: add mthreads-specific kernels (#5395) by @Oslomayor
  • Optimize MM kernels and autotuning (#5407) by @Tiannn967
  • [KMCompiler]Optimize index_copy and index_copy_ generic kernels (#5496) by @Onisen7
  • [MTHREADS] optimize matmul_bias_activation with SQMMA (#5551) by @Oslomayor
  • Optimize SiLU for Ascend and MetaX (#5565) by @YUANXIGUREN
  • [MTHREADS] optimize addmm_out with mthreads SQMMA/FMA override (#5604) by @Oslomayor
  • [Iluvatar] optimize matmul_bias_activation (#5656) by @awayzjj
  • [MTHREADS] optimize baddbmm with SQMMA (#5666) by @Oslomayor
  • [MTHREADS] optimize pad with mthreads constant_pad_nd override (#5675) by @Oslomayor
  • [MTHREADS] Optimize constant_pad_nd FillCopy path (#5684) by @shanyulu
  • [KMCompiler]Optimize log_sigmoid_backward Ascend performance (#5718) by @elllzat
  • [MTHREADS] Optimize feature_dropout_ with fused in-place path (#5742) by @gao0624
  • [KMCompiler] Optimization for operator linalg_solve_triangular. (#5753) by @YangLong114514
  • [MTHREADS] optimize upsample_linear1d_backward (#5779) by @Oslomayor
  • Optimize Hopper MM dispatch and tuning (#5780) by @Tiannn967
  • [MTHREADS] optimize matmuladd by reusing addmm SQMMA/FMA kernels (#5781) by @Oslomayor
  • [MTHREADS] Optimize conv_transpose1d with stride-phase gather (#5789) by @gao0624
  • [MTHREADS] optimize channel_shuffle with block-copy kernel (#5814) by @Oslomayor
  • [MTHREADS] optimize one_hot with fused validation and scatter (#5833) by @gao0624
  • [MTHREADS] optimize cudnn_convolution with no-bias conv2d path (#5834) by @gao0624
  • [MTHREADS] optimize conj performance (#5860) by @Oslomayor
  • [MTHREADS] optimize linear performance (#5863) by @Oslomayor
  • Optimize GLU kernel with TLE autotuning (#5871) by @lzllx123
  • [MTHREADS] optimize fp32 conv_transpose2d k3 (#5879) by @gao0624
  • [MTHREADS] Optimize isin Tensor_Scalar (#5977) by @gao0624

Hardware Support

  • [MTHREADS] restore pre_hook and libtuner for SQMMA (#3364) by @Oslomayor
  • [KernelGen][Nvidia] Implement renorm_ operator test functions (#3536) by @XDYuanzhuLee
  • Enflame to flagos 20260529 (#3602) by @chongzhouyang
  • [ARM]: add ARM64 CPU backend (NEON/SVE2) with INT8 quant and fuse… (#3616) by @kevinzs2048
  • [KernelGen][Nvidia] Migrate logical_xor_ inplace operator from experimental to ops (#3745) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Migrate deg2rad, fix, negative from experimental to ops (#3764) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Migrate addcmul_ operator from experimental to ops (#3766) by @XDYuanzhuLee
  • [KernelGen][Nvidia] Migrate atanh_operator from experimental to ops (#3768) by @XDYuanzhuLee
  • [KUNLUNXIN] all_dim (#3781) by @sh1653487844
  • [KernelGen][MetaX] Upgrade sparse_attention with multi-tier dispatch (#3998) by @LoserCheems
  • [Mthreads] Optimized baddbmm & w8a8_block_fp8_matmul with FlagOSTune (#4095) by @Fattysand
  • [metax] update masked_fill_ op (#4170) by @Alvin-YCHEN
  • [Backend] enable_nvidia_unused_ops (#4285) by @Galaxy1458
  • [metax] update layernorm backwar kernel (#4431) by @Alvin-YCHEN
  • cambricon: update and fix kernel (#4435) by @chenmiao1919
  • [Iluvatar] Opt tile repeat (#4458) by @awayzjj
  • [Iluvatar] Opt Iluvatar conv_depthwise2d. (#4513) by @huatuoli
  • gpu_check: report available GPUs instead of requiring all free (#4528) by @tengqm
  • [Kunlunxin] Register cumprod / cumprod_ in init (#4567) by @llaboon
  • [KMCompiler][Nvidia] opt for log_normal_ operator (#4621) by @littlekid-zxt
  • [ENFLAME]update 20260709 (#4646) by @chongzhouyang
  • [ENFLAME] update_ops_20260713 (#4793) by @chongzhouyang
  • [ENFLAME]workround for gcu300 to support sgl (#4814) by @chongzhouyang
  • [KernelGen][Nvidia] Move arctanh operator to ops (#4818) by @NineAnnAnn
  • [KMCompiler][Ascend] nonzero static ascend (#4921) by @chenkc2025
  • [KernelGen][Nvidia] Move sgn operator to ops (#4990) by @YufanPeter
  • [KernelGen][Nvidia] Move sign operator to ops (#5030) by @YufanPeter
  • [ENFLAME]enflame_to_FlagOS_20260729 (#5044) by @chongzhouyang
  • [KernelGen][Nvidia] Move log_ operator to ops (#5045) by @YufanPeter
  • [KernelGen][Nvidia] Move take operator to ops (#5070) by @YufanPeter
25 more changes
  • [KernelGen][Nvidia] Move hardtanh_ operator to ops (#5115) by @YufanPeter
  • [KernelGen][Nvidia] Move hardswish operator to ops (#5139) by @YufanPeter
  • [KernelGen][Nvidia] Move hardsigmoid operator to ops (#5149) by @YufanPeter
  • [KernelGen][Nvidia] Move huber_loss operator to ops (#5154) by @xuanzhengdu-eng
  • [KernelGen][Nvidia] Move hardshrink operator to ops (#5155) by @xuanzhengdu-eng
  • [KernelGen][Nvidia] Move heaviside_ operator to ops (#5156) by @xuanzhengdu-eng
  • [KernelGen][Nvidia] Move hypot_ operator to ops (#5157) by @xuanzhengdu-eng
  • [KernelGen][Nvidia] Move fix_ operator to ops (#5158) by @xuanzhengdu-eng
  • [KernelGen][Nvidia] Move arccosh operator to ops (#5163) by @YufanPeter
  • [KernelGen][Nvidia] Move absolute_ operator to ops (#5187) by @YufanPeter
  • Metax fix shared ops (#5334) by @CLYtou
  • Metax c550 operator fixes (#5337) by @CLYtou
  • [KernelGen][Nvidia] Remap native_dropout_backward to canonical wrapper (#5414) by @CoyeCHEN
  • [Ascend] Update flagtree ascend whl version to 0.6.1 in backends.yaml (#5433) by @zhzhcookie
  • [ENFLAME] enflame_to_flagos_20260818 (#5582) by @chongzhouyang
  • [KernelGen][Nvidia] Move hardtanh operator to ops (#5629) by @kkkwb
  • [KMCompiler][ASCEND] Stabilize Ascend cholesky solve layout test (#5698) by @YangLong114514
  • [KMCompiler][Nvidia] Refine dist operator impl and add full test coverage (#5958) by @Caeruleann
  • [KMCompiler][Ascend] Opt unsafe_index operator on Ascend. (#6007) by @Caeruleann
  • [KMCompiler][Ascend] Opt dist operator on Ascend. (#6014) by @Caeruleann
  • [KMCompiler][NVIDIA] Matrix rank test fixes. (#6076) by @YangLong114514
  • [kunlunxin] move the copy family onto tle.gpu (#6093) by @Ason93
  • [KMCompiler][Ascend] Opt replication pad2d backward (#6138) by @wuhansky
  • [KMCompiler][Metax] operator linalg_norm ord==0.5 scalar fix (#6162) by @YangLong114514
  • enflame to flagos 20260911 (#6200) by @chongzhouyang

Improvements

  • refactor type as Type to support old version Python (#1136) by @sgjzfzzf
  • Improve addmv_out test coverage (#3078) by @lukatao
  • Improve all_dims test coverage (#3086) by @lukatao
  • Improve baddbmm.out test coverage (#3099) by @lukatao
  • Improve cat_out test coverage (#3108) by @lukatao
  • Improve chunk_gated_delta_rule_fwd test coverage (#3109) by @lukatao
  • Improve sparse_mla_fwd_interface test coverage (#3201) by @lukatao
  • Improve atan2 test coverage (#3278) by @lukatao
  • Improve pixel test coverage (#3344) by @lukatao
  • Improve run_tests script for progress bar update (#3754) by @tengqm
  • Improve run_tests script for progress bar update (#3757) by @tengqm
  • [Performance Optimization] Improve any_dims parallel reduction (#4130) by @stringl1l1l1l
  • Improve the setup.sh script (#4168) by @tengqm
  • Improve mark collection logic (#4169) by @tengqm
  • Refactor workflows for a bootstrap checkout (#4174) by @tengqm
  • Improve W4A16 fused Marlin MoE performance (#4350) by @KingsleyLiu-NV
  • Improve setup logic (#4403) by @tengqm
  • Improve support to C++ operators (#4404) by @tengqm
  • Improve setup procedure for flagtree (#4459) by @tengqm
  • [Performance Optimization] Improve reflection_pad1d_backward (#4510) by @stringl1l1l1l
  • [Performance Optimization] Improve log_softmax_out non-inner reduction (#4511) by @stringl1l1l1l
  • [MTHREADS] improve bmm SQMMA tuning (#5909) by @Oslomayor
  • Refactor iluvatar installation commands in setup_vendor.sh (e7aadf568068; commit)
  • Refactor test-operator step in command.yaml (34626626550c; commit)

Dependencies

  • Bump torch version to 2.9.0 on Kunlunxin backends (#3773) by @tengqm
  • Bump metax versions (#3842) by @tengqm
  • Bump package versions for tsingmicro backend (#3846) by @tengqm
  • Bump enflame package dependencies (#3853) by @tengqm
  • Bump the actions-minor group with 2 updates (#4489) by @dependabot
  • Bump actions/labeler from 6.1.0 to 6.2.0 in the actions-minor group (#4634) by @dependabot
  • chore(deps): bump actions/setup-node from 6.4.0 to 7.0.0 (#4845) by @dependabot
  • chore(deps): bump actions/checkout from 6.0.2 to 7.0.0 (#4846) by @dependabot
  • chore(deps): bump actions/setup-go from 6.5.0 to 7.0.0 (#4847) by @dependabot
  • Bump numpy version to 2.3.5 (#4994) by @tengqm
  • Bump wheel version for CVE (#5271) by @tengqm
  • Bump flagtree version for NVIDIA (#5444) by @tengqm
  • Bump flagtree version on enflame 1.10.6 backend (#5649) by @tengqm
  • Bump python version for iluvatar (#5765) by @tengqm

Documentation

  • Auto-generate version from git tags using setuptools-scm (#4497) by @tengqm
  • docs: update installation guide for setup.sh workflow (#4647) by @tengqm
  • docs: update Chinese installation guide for setup.sh workflow (#4648) by @tengqm
  • docs: update install/cpp/packaging guides for cpp/ wheel split (#4868) by @tengqm
  • chore: add Apache 2.0 copyright header to all source files (#4873) by @tengqm
  • docs(release): place dev tag on empty commit after release, not on release tag (#4912) by @tengqm
  • Ship flaggems-setup as a top-level module and auto-install a compiler (#5207) by @tengqm
  • Contribution doc update (#5307) by @103yiran
  • Split multi-version vendors into versioned backend keys (#5350) by @tengqm
  • docs: drop Apache 2.0 headers added to third-party files (#6156) by @tengqm
  • docs(fused_marlin_moe): document weight layout is not vLLM Marlin layout (#6204) by @zeroherolin

CI/Infrastructure

  • ci(setup.sh): specify source when installing uv (#4106) by @103yiran
  • build(deps): bump the actions-minor group across 1 directory with 3 updates (#4356) by @dependabot
  • build(deps): bump actions/cache from 5.0.5 to 6.0.0 (#4357) by @dependabot
  • build(deps): bump actions/checkout from 6.0.2 to 7.0.0 (#4359) by @dependabot
  • ci: remove stale image-builder workflow (#4872) by @tengqm
  • ci: run weekly tests daily and use a per-vendor job-level timeout (#4887) by @wuwentao
  • ci: update schedule time and timeout value (#4889) by @wuwentao
  • [CI]: update weekly test output path and mthreads/hygon container (#4946) by @wuwentao
  • CI: improve weekly test gpu id and remove cache file (#4983) by @wuwentao
  • build(deps): bump the actions-minor group across 1 directory with 3 updates (#5071) by @dependabot
  • [CI]: Update op-monitor url and schedule time (#5546) by @wuwentao
  • ci: add rule-check workflow for PR validation (#5774) by @gavin0x01
  • Ci/rule check phase1 2 (#5787) by @gavin0x01
  • ci: add rule-check results reporting to Feishu Bitable (#5827) by @gavin0x01
  • Ci/rule check phase1 2 (#5853) by @gavin0x01
  • ci: add rule-check results reporting to Feishu Bitable (#5854) by @gavin0x01
  • Ci/rule check phase1 2 (#5868) by @gavin0x01
  • build: tolerate post-release tags in version derivation (#5885) by @wbavon
  • Ci/rule check phase1 2 (#5898) by @gavin0x01
  • ci(weekly): add workflow dispatch controls (#5919) by @wuwentao
  • [CI] Add rule check to forbid use_gems() in PR-changed test files (#5963) by @gavin0x01
  • cicd(.github): remove cann850 (#5970) by @103yiran
  • cicd(.github): remove metax-maca3720 (#6090) by @103yiran
  • [CI] Use three-dot diff to derive PR-changed files (#6097) by @gavin0x01
  • [CI/CD] Add rule-check-required job and cicd docs (#6119) by @103yiran
  • [Hygon] cicd(backends.yaml): update hygon compiler version (#6148) by @103yiran
  • Ci/extend init exports check (#6157) by @gavin0x01
  • [packaging] Port #6365 to 5.4.0-rc2 (Debian + RPM packaging) (#6394) by @shiptux

Testing

  • [Benchmark][Ascend] Update benchmark do_bench to do_bench_npu on ascend (#4857) by @zhzhcookie
  • TEST CASE FIX -- iluvatar (#5131) by @jia-heng
  • test(and_scalar): tag bitwise_and_scalar tests so pytest -m and_scalar collects (#6033) by @bin913
  • test: rename batch_norm_with_update_functional benchmark file to sing… (#6052) by @freaking-man-in
  • [QC][Benchmark] Use dedicated test_mm_w8a8_fp8 entry (#6084) by @ggbondbest
  • [Benchmark] Fix chunk gated delta rule forward baseline (#6177) by @lukatao
  • [Benchmark] Add standalone FP8 MM performance test with vLLM baseline (#6183) by @ggbondbest

Reverts

  • Revert arm vendor cmd (#3820) by @0x45f
  • Revert iluvatar LIBDIR hack and disable flagtree for iluvatar (#4420) by @tengqm
  • Revert "[FlagTune] Add multi-platform Mul cost model support" (#6043) by @Vincent-Xiao

Bug Fixes

  • fix(grid_sample): enable large output sizes (#3421) by @mygitljf
  • Fix scatter_reduce_two kernel (#3458) by @tengqm
  • [HYGON] Fix pow, mul, per_token_group_quant_fp8, allclose, isclose and reshape_and_cache (#3476) by @fangzexian
  • [MTHREADS] fix assert_async (#3504) by @Oslomayor
  • [KUNLUNXIN] fix cumprod_ (#3562) by @sh1653487844
  • Fix patch vllm util (#3657) by @0x45f
  • [KUNLUNXIN] fix vender problem (#3660) by @nianqi-tian
  • Fix environment setting for iluvatar backend (#3667) by @tengqm
  • Fix parameters used for benchmarking kunlunxin (#3690) by @tengqm
  • [KUNLUNXIN] fix cumprod_ boolean tensor in-place operation (#3747) by @sh1653487844
  • [KernelGen] Fix --ref cpu cross-device tensor comparison (#3750) by @lukatao
  • [KernelGen] Fix floor_divide for mixed int/float types and fp16/bf16 dtypes (#3751) by @lukatao
  • [KernelGen] Fix grouped_mm accuracy test and benchmark configuration (#3752) by @lukatao
  • [Advanced Compiler] fix dgeglu geglu dreglu reglu skipped test caused by invalid zero-dimension test shape (#3755) by @AdvancedCompiler
  • [KernelGen] Fix sparse_mla TLE kernel group/head-block index decomposition (#3759) by @lukatao
  • Fix cpp test (#3761) by @tengqm
  • [KernelGen] Fix addmm_dtype and addmm_dtype_out accuracy tests for --ref cpu mode (#3762) by @lukatao
  • [Backend ] Fix_replace_op_bugs (#3785) by @Galaxy1458
  • Fix github CI for script rename (#3795) by @tengqm
  • [ENFLAME] fix ops of gcu300 some bugs to enflame backend (#3814) by @chaaa-a
  • fix issue 3733 (#3815) by @huangyiqun
  • [KernelGen] Fix count_nonzero: replace deprecated logical operators and add sparse/negative dim support (#3816) by @yzw1128
  • fix svd & pointwise_dynamic (#3824) by @Dayrker
  • [Sunrise backend] fix scatter_reduce & mse_loss (#3829) by @Dayrker
  • fix issues:1012 (#3831) by @w1120029931-bit
  • Fix ix containerfile (#3841) by @tengqm
  • Fix env tsingmicro (#3844) by @tengqm
  • [Advanced Compiler] Fix conv_depthwise2d test with --ref cpu (#3859) by @AdvancedCompiler
  • [Backend] Fix replace op bugs (#3863) by @Galaxy1458
  • [Sunrise] fix fused_recurrent (#3878) by @Dayrker
231 more changes
  • Fix package dependency on klx (#3934) by @tengqm
  • Fix operator inventory (#3937) by @tengqm
  • [KUNLUNXIN] fix weight_norm, aminmax,l eaky_relu, mm, bmm, copysign, signbit, logsumexp,nonzero_numpy and tril (#3953) by @RobertLuobo
  • [KernelGen] Fix baddbmm.out: missing_test (#3958) by @lukatao
  • [KernelGen] Fix new_full: accuracy_fail (#3962) by @lukatao
  • fix issue 3007 (#3972) by @huangyiqun
  • fix issue 3585 (#3973) by @huangyiqun
  • Fix metadata for the tan_ operator (#3983) by @XDYuanzhuLee
  • Fix TsingMicro CI (#4010) by @tengqm
  • Fix GPU check script for Ascend (#4045) by @tengqm
  • Fix on-demand workflow for GPU check (#4047) by @tengqm
  • Fix operator inventory (#4048) by @tengqm
  • Fix incomplete operator entry (#4049) by @tengqm
  • fix issue 2896 (#4051) by @huangyiqun
  • fix issue 2829 (#4052) by @huangyiqun
  • fix issue 2850 (#4053) by @huangyiqun
  • Fix Kunlunxin setup (#4082) by @tengqm
  • [KernelGen][Iluvatar] Fix all_dim accuracy failure (float16/bfloat16 input) (#4093) by @lukatao
  • [Ascend] fix: fill_ ZeroDivisionError on empty tensor (#4096) by @JosephNew
  • Fix any_dim accuracy failure on Iluvatar BI-V150 (float16 input) (#4098) by @lukatao
  • Fix var accuracy failure on Iluvatar BI-V150 (Triton specialization bug) (#4099) by @lukatao
  • Fix thead setup for CI (#4103) by @tengqm
  • [SiliconFlow] Fix nanmedian cross-backend accuracy (#4116) by @winfan1314
  • [KernelGen][Iluvatar] Fix var_correction: accuracy_fail (#4118) by @lukatao
  • [SiliconFlow] Fix cumprod accuracy and inplace integer path (#4128) by @winfan1314
  • [enflame]Fix VendorDescriptor bug (#4134) by @Galaxy1458
  • Fix containerfile for Hygon (#4138) by @tengqm
  • Fix container file for KLX (#4139) by @tengqm
  • Fix PyYAML version in pyproject.toml (#4143) by @tengqm
  • Fix the vllm installation cache problem (#4164) by @tengqm
  • Fix ascend backend (#4165) by @tengqm
  • Fix renorm fp16 precision: allocate norms buffer in fp32 (#4237) by @tengqm
  • Fix median/nanmedian test crash on int16 in quick mode (#4239) by @tengqm
  • [SiliconFlow] Fix mode operator on Mthreads backend (#4248) by @winfan1314
  • [Mthreads] Fix heuristic fallback for log_softmax (#4258) by @mygitljf
  • Fix hardsigmoid_out benchmark (#4271) by @huangyiqun
  • [AdvancedCompiler] fix: avoid int32 overflow for large tensors in repeat and tile kernel (#4289) by @AdvancedCompiler
  • [KernelGen] Fix conv2d_padding: accuracy_fail (#4298) by @lukatao
  • [KernelGen] Fix multinomial: accuracy_fail (#4303) by @lukatao
  • fix mm tma (#4308) by @huangyiqun
  • Fix operator order in registry (#4313) by @XDYuanzhuLee
  • [AdvancedCompiler] fix: max_pool3d_backward test (#4328) by @AdvancedCompiler
  • Fix unbound variable errors when sourcing vendor env scripts (#4344) by @tengqm
  • Fix uv install failure when HOME is not writable (#4345) by @tengqm
  • Fix gh install failure in command.yaml when HOME is not writable (#4346) by @tengqm
  • [MTHREADS] fix soft_margin_loss performance with two-phase reduction (#4364) by @Oslomayor
  • Fix CMakeLists.txt for git cloning libtriton_jit (#4375) by @tengqm
  • Fix containerfile for NVIDIA 13.3 backend (#4378) by @tengqm
  • Fix weight_norm_interface norm dtype truncation (#4382) by @tengqm
  • Fix nanmedian kthvalue fallback for NaN-containing inputs (#4384) by @tengqm
  • Fix some nits in the operator source code (#4409) by @tengqm
  • Fix: defer yaml parsing in setup.sh to after venv creation (#4415) by @tengqm
  • Fix check-gpu step and failure-report duplicates in command.yaml (#4416) by @tengqm
  • Fix Iluvatar triton plugin loading in venv (#4417) by @tengqm
  • [MTHREADS] fix log10 perf (#4426) by @Oslomayor
  • Fix error when dim is empty tuple (#4428) by @0x45f
  • fix(sparse_attention): pad H dimension to satisfy tl.dot minimum size (#4447) by @Yangyang0906C
  • [MTHREADS] fix tile size config for unary elementwise ops (#4453) by @Oslomayor
  • [TSINGMICRO] fix bugs of test cases (#4463) by @tsingmicro-public-e
  • fix hcu import error (#4464) by @huangyiqun
  • [Ascend] fix argmax crash: remove duplicate @libentry() on argmax_kernel (#4467) by @JosephNew
  • fix(broadcast_to): use pin_memory to avoid illegal CPU->CUDA copy in CUDA graph capture (#4472) by @physics31415926
  • fix(container): remove garbled characters (#4486) by @103yiran
  • fix mark in test (#4492) by @103yiran
  • Fix setup.sh for setuptools-scm with --no-build-isolation (#4498) by @tengqm
  • fix(conf): remove div_tensor_mode, div_scalar_mode (#4509) by @103yiran
  • Fix the dispatch keys for some operators (#4522) by @tengqm
  • Fix psum_text comparison when before has no performance data (#4534) by @tengqm
  • fix: safe vllm install to protect torch/triton/flagtree (#4535) by @tengqm
  • Fix cumsum ZeroDivisionError on empty tensor (#4541) by @DannyP0
  • Fix randn error in multi card case (#4560) by @0x45f
  • Fix uv not found in subsequent CI steps (#4562) by @tengqm
  • Fix psum_text IndexError when before data is empty (#4569) by @tengqm
  • Fix code style error (#4585) by @0x45f
  • fix: affine_grid_generator float32 precision mismatch with PyTorch CUDA reference (#4586) by @yzw1128
  • fix(ascend): mask bmm boundary loads (#4631) by @yilongma0110
  • fix transformer engine import error (#4642) by @huangyiqun
  • Fix uv detection by adding ~/.local/bin to PATH before check (#4671) by @tengqm
  • Fix setup.sh and env.sh for set -u compatibility (#4672) by @tengqm
  • Fix ascend-cann900 post_install format in backends.yaml (#4674) by @tengqm
  • Fix post_install: use full shell commands with eval (#4675) by @tengqm
  • Fix Ascend env sourcing: unify path and add error tolerance (#4678) by @tengqm
  • Fix negative op on non-CUDA backends (e.g. MUSA) (#4682) by @tengqm
  • [KernelGen][Mthreads] Fix linear: accuracy_fail (#4792) by @xzliu-opt
  • Fix environment setting for Kunlunxin (#4795) by @tengqm
  • Fix illegal memory access in cluster_remote_gemm_kernel when cdiv(N, BN) is odd (#4797) by @yysheng26
  • [Hygon] fix mm float64 accumulation (#4799) by @douxetpur
  • fix fft's res_out (#4802) by @Dayrker
  • fix: test log output dir can't be access (#4805) by @wuwentao
  • Fix cpp op ci error (#4811) by @0x45f
  • fix: increase test job timeout and add label/csv/html to output dir (#4813) by @wuwentao
  • fix: resolve CI test TIMEOUT/FAILED/OOM issues for multiple operators (#4829) by @CLYtou
  • fix: container pip error and only run basic test (#4834) by @wuwentao
  • fix: filter None and mismatched out keys in pointwise_dynamic prepare_args (#4849) by @bwbwzzz
  • fix: nvidia/huawei/hygon/mthreads test env (#4851) by @wuwentao
  • fix: kunlunxin remove fg_mode and use default kernel mode (#4854) by @wuwentao
  • fix: nvidia container/mthreads env/huawei env for weekly ci job (#4859) by @wuwentao
  • fix: test script import datetime error in main func (#4861) by @wuwentao
  • fix(ci): correct setup-flaggems input name in random-test.yaml (#4869) by @tengqm
  • Fix sdpa head dim k (#4881) by @0x45f
  • fix(kunlunxin): remove --fg_mode operator from performance benchmark (#4886) by @wuwentao
  • fix(test): add try/except for fp8_einsum import avoiding false CI failures (#4888) by @CLYtou
  • fix: resolve complex conjugate recursion and vdot accuracy failures (#4890) by @CLYtou
  • fix sinc (#4892) by @llaboon
  • Fix tle error (#4894) by @0x45f
  • fix(tests): fix mm tma test (#4905) by @103yiran
  • Fix cpp import error (#4922) by @0x45f
  • fix: register empty as empty.memory_format so torch.empty uses FlagGems (#4929) by @awayzjj
  • fix: replace triton_src kernel references with ops/ paths in C++ wrappers (#4931) by @yysheng26
  • fix(config): lazy import c_operators — detect .so without loading (#4959) by @tengqm
  • fix(cmake): add global include_directories for NPU backend (#4960) by @tengqm
  • fix(cpp): disable aten dispatch + exclude PrivateUse1 from redispatch guard (#4962) by @tengqm
  • Fix setup for triton/flagtree combination (#4963) by @tengqm
  • Fix mul fallback error (#4999) by @0x45f
  • fix(rmsnorm): reorder multiply to (x * inv_rms) * weight for vendor precision alignment (#5009) by @CLYtou
  • Fix hygon (#5016) by @huangyiqun
  • fix(utils): add fallback_j1 (#5038) by @103yiran
  • Fix CI job to runner mapping (#5049) by @tengqm
  • fix: add tsingmicro/txda to device detection (#5056) by @tengqm
  • fix(apply_rotary_pos_emb): fix cos/sin error (#5073) by @103yiran
  • fix(utils): add pure-triton nextafter fallback for non-CUDA backends (#5088) by @tengqm
  • Fix import error in ascend (#5101) by @0x45f
  • fix: vector_norm accuracy test failure on mthreads backend (#5106) by @lyujheng
  • fix: resolve copy op accuracy failure on mthreads backend (#5107) by @lyujheng
  • fix: resolve_conj performance test failure on mthreads backend. (#5109) by @lyujheng
  • fix: resolve mul op accuracy and performance test failure on mthreads (#5111) by @lyujheng
  • [Ascend] Fix ascend ops when running Qwen3.6 vLLM with flagtree ascend3.5 (#5113) by @zhzhcookie
  • fix(mul): gate optimized path on active backend device, not hardcoded "cuda" (#5130) by @tengqm
  • fix: recognize vLLM stable ABI extensions (#5150) by @CherryLemon
  • Fix out-of-bounds read past the input tail in bucketize kernel (#5231) by @truong-v
  • Fix IX gate by pinning numpy version < 2 (#5244) by @tengqm
  • fix: index_add accuracy failure on mthreads (#5246) by @lyujheng
  • [KernelGen][metax] Fix num_warps exceeding hardware thread limit for softmax, logsumexp, dropout, exponential_, rand, randn, and uniform (#5258) by @yzw1128
  • fix: mul accuracy test failure on mthreads (#5294) by @lyujheng
  • fix: div accuracy test failures on mthreads backend (#5295) by @lyujheng
  • [Enflame] fix enflame backend operator bug (#5302) by @chaaa-a
  • fix(ascend): align full_like kernel API with updated full_kernel signature (#5308) by @yzw1128
  • [KernelGen][Cambricon] Fix nested_view_from_buffer_copy: accuracy_fail (#5324) by @xzliu-opt
  • fix add op test (#5325) by @huangyiqun
  • Fix code style (#5331) by @0x45f
  • fix(ops): skip triton kernels for complex dtypes in empty/lift_fresh/… (#5332) by @CLYtou
  • fix(conf): add missing info for special_log_softmax (#5333) by @103yiran
  • [KernelGen][Metax] Fix baddbmm: accuracy_fail (#5336) by @xzliu-opt
  • fix(utils): add pure-triton j0 and log2 fallbacks for non-CUDA backends (#5348) by @tengqm
  • fix(cambricon): tolerate triton forks lacking TRITON_MAX_TENSOR_NUMEL (#5351) by @tengqm
  • Fix special_chebyshev_polynomial_w ut (#5352) by @0x45f
  • fix: prevent run_tests.py deadlock when worker crashes (#5358) by @wuwentao
  • fix flip op accuracy test failure on mthreads backend (#5374) by @lyujheng
  • fix testcase iluvatar (#5375) by @jia-heng
  • [KernelGen][Ascend] Fix specialerfc-import-error-on-ascend: import_error (#5380) by @xzliu-opt
  • [KernelGen][Ascend] Fix any_dim: accuracy_fail (#5381) by @xzliu-opt
  • Fix cat cpp wrapper kernel path, cat dtype mismatch, and rwkv_mm_sparsity precision (#5388) by @yysheng26
  • fix: increase op test job timeout (#5391) by @wuwentao
  • fix: skip kernel launch for zero-element tensors in empty() (#5410) by @lukatao
  • Fix /test command parsing for space-separated params (#5425) by @tengqm
  • Fix empty (#5438) by @0x45f
  • [KernelGen][Mthreads] Fix conv2d_padding: accuracy_fail (#5445) by @xzliu-opt
  • [Ascend] Fix input param of do_bench_npu (#5451) by @zhzhcookie
  • fix(utils): add pure Triton lgamma fallback (#5464) by @huangyiqun
  • fix(cambricon): stop index autotune from killing the process on untileable expand_shape (#5510) by @tengqm
  • fix(utils): add pure-triton y0 fallback for non-CUDA backends (#5553) by @lukatao
  • Fix empty tensor sum error (#5567) by @0x45f
  • Fix ci import error (#5573) by @0x45f
  • [metax] Fix ops registration and update heuristics/tune configs (#5576) by @yzw1128
  • [metax] Fix upsample_linear1d autotune key mismatch (unblock metax CI) (#5591) by @yzw1128
  • fix(i0): avoid unnecessary x.to() temporary tensor (#5592) by @kkkwb
  • [KMCompiler][ASCEND] Fix tle import error for cholesky_solve op (#5616) by @YangLong114514
  • [KMCompiler][Ascend] Fix bug (#5641) by @Willing2329
  • [KMCompiler][Iluvatar] bug fix for operator nonzero_static (#5644) by @YangLong114514
  • Fix M-Threads image and backend name in upload log (#5650) by @wuwentao
  • [KMCompiler][Ascend] Fix replication pad2d backward bug (#5654) by @Willing2329
  • [PPU][OP Test][Bug Fix] Fix upsample and quant operator test (#5655) by @tianyi-dialect
  • [KMCompiler] fix linalg_vecdot: exclude fp64 benchmark for iluvatar and ascend platforms (#5658) by @rye985
  • fix(flagtune): fall back when platform model package is unavailable (#5678) by @k6699k
  • [KMComplier][NVIDIA][Thead][MetaX][iluvatar] Fix test_linalg_det hang with --ref cpu (#5701) by @littlekid-zxt
  • [KMCompiler][iluvatar] Fix sparse_sampled_addmm compilation failure for K < 16 (tl.dot K >= 16 constraint) (#5705) by @littlekid-zxt
  • fix(cambricon): use elementwise & | ~ instead of Python and/or/not on tensor masks (#5745) by @tengqm
  • [Triton] Fix tl.where input param in randperm op (#5769) by @zhzhcookie
  • fix dead ATen registrations and non-standard registration names (#5777) by @yysheng26
  • fix _weight_norm: split underscore op tests into a dedicated mark (#5778) by @Yukun-Cui
  • [Iluvatar] fix(pointwise): restore rank-1 vectorization on Triton 3.6 (#5783) by @awayzjj
  • fix(cambricon): broadcast scalar max_virtual_index in paged load_from_kvcache (#5786) by @tengqm
  • [kunlunxin] fix slice_scatter OOB src index on masked-out lanes (#5816) by @Cocytus486
  • fixbug: exponential_ operator accuracy failure (#5830) by @lyujheng
  • [KMCompiler][Nvidia]Fix igammac op register and its testcases related torch op name (#5831) by @lelongmask-cmp
  • fix(utils): add pure-triton normcdfinv fallback for non-CUDA backends (#5864) by @lukatao
  • [MTHREADS] Fix mthreads backend dispatch key (#5865) by @Oslomayor
  • [NVIDIA] Fix Hopper MM split-K accumulation precision (#5876) by @Tiannn967
  • [KernelGen][Nvidia] Fix index_put_impl (#5889) by @Yukun-Cui
  • fix(benchmark): isolate MM operator measurements (#5902) by @Tiannn967
  • fix:special_hermite_polynomial_h (#5930) by @douxetpur
  • fix(ascend): fix multinomial NaN/OOM and cumsum device bug on Ascend backend (#5980) by @CLYtou
  • fix(setup.sh): fix uv aarch64 unzip path error (#5999) by @103yiran
  • fix(slice): fix missing default args and broken view semantics (#6004) by @kkkwb
  • fix(broadcast_tensors): handle zero-size dims in broadcast shape computation (#6024) by @cccxxttt
  • fix(hygon): KeyError: 'BLOCK_M' triggered by flash attention operations when running ViT encoder inference (#6040) by @xmhubj
  • [KMCompiler][iluvatar] fix oom (#6070) by @YangLong114514
  • fix(attention): prevent _attn_bwd runtime crashes on Hopper sm90 (#6075) by @cccxxttt
  • fix(flash_attention): keep AABS from breaking tl.dot on decode backward (#6077) by @cccxxttt
  • fix(index): cast pid_p to int64 to prevent int32 offset overflow (#6087) by @JosephNew
  • Fix all and any zero div error (#6098) by @0x45f
  • fix(nextafter): propagate NaN when target operand is NaN (#6100) by @cccxxttt
  • fix(benchmark): write convert_weight_to_int4pack results as structured JSON (#6101) by @cccxxttt
  • fix(benchmark): write int4pack_mm_with_scales_and_zeros results as structured JSON (#6118) by @cccxxttt
  • [Fix] Respect FlagTune environment settings in MM tests and benchmarks (#6129) by @k6699k
  • fix benchmark naming for irshift (#6131) by @freaking-man-in
  • Fix expression syntax error in unittest.yaml concurrency group (#6137) by @103yiran
  • fix: add missing DSA/init.py so non-editable installs include the subpackage (#6146) by @JosephNew
  • Fix adaptive_avg_pool3d_backward dispatching to _adaptive_avg_pool3d_backward (#6147) by @kkkwb
  • [KMCompiler][MetaX] Fix gru bug (#6155) by @Willing2329
  • fix(.github): remove metax-maca3720 info (#6160) by @103yiran
  • fix(sunrise): guard the vendor cumsum against empty input, add an asin fallback (#6165) by @tengqm
  • fix(utils): add graph capture buffer reuse for pointwise_dynamic (#6171) by @CLYtou
  • fix(ci): update kunlunxin/iluvatar flagtree ci docker image (#6192) by @wuwentao
  • [Bug Fix] Enable SQLite WAL mode and busy_timeout to fix "database is locked" in autotune cache (#6203) by @modao1234
  • fix(ci): stop sort_exports from deleting imports on double all (#6213) by @gavin0x01
  • fix: gate Triton 3.6-only APIs (map_elementwise, triton.knobs) for the flagtree/Triton-3.5 line (#6236) by @JosephNew
  • fix: add Triton 3.5 compatibility for flash_attention_backward (issue… (#6253) by @CLYtou
  • [BugFix] Support boolean slice views (#6061) (#6350) by @0x45f
  • Fix iluvatar import (#6228) (#6377) by @0x45f
  • fix iluvatar randperm/sort (#5575) (#6380) by @0x45f
  • fix(sqlite): tolerate WAL journal-mode switch race on multi-process cold start (#6388) by @JosephNew
  • fix(autotune): reflect single table instead of whole db in build_sql_model_by_db (#6416) by @JosephNew
  • [MetaX] Fix masked_fill broadcast OOB and illegal num_warps=16 (#6443) by @zhangjieBaai
  • fix(enflame): chunk grid.x in index_put to respect the 65535 launch limit (#6538) by @JosephNew
  • [Ascend] fix: improve Ascend cat/concat dispatch and implementation (#6547) by @modao1234
  • fixbug: issue#6420 (#6566) by @lyujheng
  • Fix variable reference for pytest operator check (9479fcd1214d; commit)
  • Fix typo in pip install command in workflow (b4d5d12adb13; commit)
  • Fix activation of virtual environment in workflow (aff20564630e; commit)
  • Fix concurrency group expression in unittest.yaml (3896f48357ab; commit)

Other

  • [SiliconFlow] Operator Development adaptive_avg_pool2d (#2059) by @fpzh2011
  • [Advanced Compiler] topk (#3329) by @AdvancedCompiler
  • [SiliconFlow] Op: flash_attention_backward (#3386) by @LiangJiabaoY
  • [Performance Optimization] Reduce test cases in QUICK_MODE (#3423) by @git-flyer
  • Regist vllm ops in torch (#3497) by @0x45f
  • scaled_dot_product_attention_backward support non-square tuning, non-causal, and unequal lengths of q and kv sequences (#3607) by @103yiran
  • Prefer Triton over FlagTree on Iluvatar (#3666) by @tengqm
  • Skip flash_mla_sparse_fwd operator (#3692) by @tengqm
  • Be tolerant when running batch/daily tests (#3696) by @tengqm
  • Initial version of container file for mthreads (#3697) by @tengqm
  • Standardize Iluvatar backend debug log format (#3711) by @lukatao
  • Standardize MetaX backend debug log format (#3712) by @lukatao
  • Skip the fused_marlin_moe accuracy test due to assertion failures (#3734) by @tengqm
  • Standardize Cambricon backend debug log format (#3737) by @lukatao
  • Standardize Mthreads backend debug log format (#3738) by @lukatao
  • Standardize Kunlunxin backend debug log format (#3739) by @lukatao
  • Standardize Aipu backend debug log format (#3740) by @lukatao
  • Standardize Hygon backend debug log format (#3742) by @lukatao
  • Standardize Tsingmicro backend debug log format (#3743) by @lukatao
  • Pin triton to 3.6.0 in setup script (#3758) by @tengqm
  • Opname fix (#3780) by @tianxiao-baai
  • update _sunrise backend (#3791) by @Dayrker
  • Container file for the iluvatar backend (#3808) by @tengqm
  • Downgrade torch from 2.12 to 2.11 for Nvidia (#3825) by @tengqm
  • Re-enable tsingmicro backend test (#3827) by @tengqm
  • Enable backend tests for enflame, spacemit and sunrise (#3830) by @tengqm
  • Simplify pyproject.toml for extra dependency specification (#3843) by @tengqm
  • Allowing for specifying all GPUs when testing (#3845) by @tengqm
  • Try skip vendor.sh from the setup.sh script (#3848) by @tengqm
  • Containerfile for MetaX 3.7.2.1 (#3849) by @tengqm
216 more changes
  • Rename some containerfiles (#3850) by @tengqm
  • Update pyproject.toml (#3851) by @tengqm
  • Dockerfile for enflame (#3855) by @tengqm
  • Containerfile for Hygon 26.04 (#3858) by @tengqm
  • [Sunrise] Dev sunrise fix v4 (#3865) by @Dayrker
  • Tsingmicro chip count (#3887) by @tengqm
  • Update containerfile for mthreads backend (#3903) by @tengqm
  • Update operator registry (#3925) by @tengqm
  • No gpu op fix (#3935) by @tianxiao-baai
  • Sort operators by name (#3938) by @tengqm
  • Containerfile for Ascend backend (#3974) by @tengqm
  • Clarify operator stages (#3979) by @tengqm
  • Update package/environment setting for tsingmicro (#3982) by @tengqm
  • Update triton version for sunrise backend (#3989) by @tengqm
  • update tsingmicro container file to support flagtree (#4017) by @tengqm
  • Simplify ascend container file (#4021) by @tengqm
  • update containerfile for tsingmicro backend (#4035) by @tengqm
  • Update version string for the master branch (#4038) by @tengqm
  • Update CI setup for tsingmicro backend (#4044) by @tengqm
  • Debug Ascend chip check (#4062) by @tengqm
  • flaggems.ops to flaggems. (#4085) by @chongzhouyang
  • Remove buggy plugin dragged in by Triton on kunlunxin (#4090) by @tengqm
  • Amend the env.sh for environment setup on THead backend (#4107) by @tengqm
  • update from enflame until 20260616 (#4111) by @chongzhouyang
  • [SiliconFlow] Skip segment_reduce benchmark on MUSA (#4115) by @winfan1314
  • [KMCompiler] Reimplement top_k_per_row_prefill and top_k_per_row_decode (#4132) by @liuxiao0909c
  • Upgrade containerfile for nvidia 13.3 (#4135) by @tengqm
  • Install vllm=0.21.0 on H20 for operator testing (#4136) by @tengqm
  • Update containerfile for MetaX (#4140) by @tengqm
  • Update container file for mthreads (#4149) by @tengqm
  • Update containerfile for Iluvatar (#4150) by @tengqm
  • Update container file for Ascend 8.5.0 (#4151) by @tengqm
  • Update project setting for CANN 9.0.0 (#4152) by @tengqm
  • Cleanse the operator registry again (#4162) by @tengqm
  • Rename 'ilshift__' to 'ilshift' (#4163) by @tengqm
  • Allow cann850 and cann9 to be specified for workflow (#4166) by @tengqm
  • Speedup batch test run for non-existent test cases (#4171) by @tengqm
  • Reorg CI workflows (#4172) by @tengqm
  • Drop some useless tools that are no longer useful or maintained (#4173) by @tengqm
  • Drop experimental operator 'abs' and 'abs_' (#4177) by @tengqm
  • Drop experimental operator addcdiv (#4180) by @tengqm
  • Drop experimental operator celu and celu_ (#4182) by @tengqm
  • Drop experimental operator elu (#4184) by @tengqm
  • Drop experimental operator 'exp2' (#4186) by @tengqm
  • Drop experimental operator exp2_ (#4188) by @tengqm
  • Drop the experimental operator glu (#4190) by @tengqm
  • Drop experiemental operator 'logical_xor_' (#4192) by @tengqm
  • Drop experimental operator maximum (#4194) by @tengqm
  • Drop the experimental operator mse_loss (#4196) by @tengqm
  • Update containerfile for tsingmicro (#4197) by @tengqm
  • Drop experimental operator eye (#4204) by @tengqm
  • Drop experimental operator replication_pad3d (#4205) by @tengqm
  • Drop experimental operator slice_backward (#4206) by @tengqm
  • Drop experimental operator smooth_l1_loss (#4207) by @tengqm
  • Drop experimental operator triu (#4208) by @tengqm
  • Drop experimental operator leaky_relu (#4210) by @tengqm
  • Drop experimental operator 'mv' (#4212) by @tengqm
  • Drop experimental operator 'slice_scatter' (#4213) by @tengqm
  • Drop experimental operator 'deg2rad' (#4215) by @tengqm
  • Drop experimental operator 'diag' (#4221) by @tengqm
  • Drop experimental operator 'pixel_shuffle' (#4225) by @tengqm
  • Drop experimental operator 'trace' (#4229) by @tengqm
  • Remove skip annotation from trace benchmark (#4230) by @tengqm
  • Drop experimental operator 'upsample_nearest1d' (#4233) by @tengqm
  • Use correct vendor string for GPU check (#4235) by @tengqm
  • Cache gh CLI install on self-hosted runners (#4241) by @tengqm
  • Consolidate backend dispatch using matrix strategy (#4242) by @tengqm
  • Change logging warning to info (#4246) by @0x45f
  • Drop experimental operator fix (#4249) by @tengqm
  • Drop experimental masked_scatter implementation (#4257) by @tengqm
  • Drop experimental operator addcmul_ (#4259) by @tengqm
  • Drop experimental operator atanh_ (#4260) by @tengqm
  • Drop experimental masked_select operator (#4264) by @tengqm
  • Drop the experimental reciprocal operator (#4266) by @tengqm
  • [MACA] Support C++ wrappers with libtriton_jit (#4270) by @cjh9368
  • Change to debug level (#4273) by @0x45f
  • Setup script for Cambricon (#4307) by @tengqm
  • Use vendor name in backend-tests job names (#4330) by @tengqm
  • Change test count to time (#4338) by @0x45f
  • Use internal mirror for gh CLI download (#4347) by @tengqm
  • Persist corrected HOME to GITHUB_ENV in command.yaml (#4349) by @tengqm
  • update vendor's debug logger (#4361) by @huangyiqun
  • Make sure the preprocess step has source code checked out (#4373) by @tengqm
  • Drop vendor script (#4379) by @tengqm
  • Remove two operators from registry (#4380) by @tengqm
  • Widen normal distribution test tolerance to reduce flakiness (#4386) by @tengqm
  • Update configuration for enflame backend (#4406) by @tengqm
  • Separate compiler selection from deps in backends.yaml (#4407) by @tengqm
  • Drop the non-existent special.gammainc.out operator (#4408) by @tengqm
  • Merge some operator code for deduplication (#4410) by @tengqm
  • Post failure report when on-demand test fails to execute (#4413) by @tengqm
  • Split nvidia backend into nvidia-cuda128 and nvidia-cuda133 (#4418) by @tengqm
  • Update workflow files to use revised backend labels (#4421) by @tengqm
  • [QC] Enable swap_ab on H20 for small-tile fp8 blockwise GEMM (#4434) by @monellz
  • mul optimization from pointwise_dynamic to expert implementation (#4436) by @liuxin2560
  • Update package config and env var for TsingMicro (#4449) by @tengqm
  • [TSINGMICRO] update tsingmicro backend 0630 (#4450) by @tsingmicro-public-e
  • [TSINGMICRO] skip failed tests and long time benchmarks (#4452) by @tsingmicro-public-e
  • Use H20 backend for UT (#4462) by @tengqm
  • Update code owners file (#4470) by @tengqm
  • style: sort operator registrations alphabetically (#4477) by @lukatao
  • Rename the artifact from on-demand workflow (#4479) by @tengqm
  • Restore python-op CI job to jiuding machines (#4481) by @tengqm
  • Disable flagtree for sunrise (requires GLIBC_2.39) (#4495) by @tengqm
  • Simplify flag_gems version detection in run_tests.py (#4499) by @tengqm
  • Simplify flag_gems version detection in run_tests.py (#4505) by @tengqm
  • Remove extra LD_LIBRARY_PATH setting for metax (#4517) by @tengqm
  • Remove the extra post-install step for Kunlunxin (#4518) by @tengqm
  • run_tests: cap worker count to number of ops (#4529) by @tengqm
  • [C++ Wrapper] Align softmax and cat with Python wrapper (#4542) by @yysheng26
  • style: sort operator registrations alphabetically (#4559) by @lukatao
  • Tune Triton-TLE topk paths (#4572) by @zheng1
  • Update dependency and config for Sunrise (#4596) by @tengqm
  • [TSINGMICRO] update tsingmicro backend to 0708 (#4603) by @tsingmicro-public-e
  • Simplify NVIDIA backend config (#4606) by @tengqm
  • [TSINGMICRO] align grouped_topk with vllm grouped_topk (#4610) by @tsingmicro-public-e
  • [TSINGMICRO] align moe_align_block_size with vllm op (#4611) by @tsingmicro-public-e
  • [TSINGMICRO] session.mege will introduce dirty data while cocurrently… (#4612) by @tsingmicro-public-e
  • Skip affine_grid_generator accuracy tests (#4614) by @tengqm
  • Remove pre_install from tsingmicro backends (#4617) by @tengqm
  • chore: remove container template files — migrated to build-infra (#4620) by @tengqm
  • chore: remove container/ — legacy files migrated to build-infra (#4624) by @tengqm
  • chore: remove runtime_env from backends.yaml — migrated to build-infra (#4629) by @tengqm
  • chore: remove tools/env.sh from CI workflows (#4641) by @tengqm
  • [TSINGMICRO] update tsingmicro backend to 0709 (#4643) by @tsingmicro-public-e
  • Increase timeout in run_cmd function (#4645) by @tianxiao-baai
  • chore: remove requirements/ — superseded by backends.yaml (#4649) by @tengqm
  • Refine masked fill test (#4665) by @0x45f
  • Move optimized mul implementation to the general path (#4666) by @Tiannn967
  • Move compiler (flagtree/triton) management exclusively to setup.sh (#4673) by @tengqm
  • Debug ascend gpu_check failure; remove HOME override from command.yaml (#4677) by @tengqm
  • Rename post_install to triton_post_install; remove post_uninstall (#4680) by @tengqm
  • Remove debug logging from gpu_check_ascend.sh (#4681) by @tengqm
  • Drop experimental negative operator (#4683) by @tengqm
  • Always persist uv PATH to GITHUB_PATH (#4684) by @tengqm
  • Promote deg2rad_ (in-place) from experimental to stable (#4685) by @tengqm
  • Drop experimental _unsafe_view operator (#4723) by @tengqm
  • Drop experimental amin operator (#4724) by @tengqm
  • Drop experimental expand operator (#4725) by @tengqm
  • Drop experimental permute_copy operator (#4728) by @tengqm
  • Drop experimental relu operator (#4729) by @tengqm
  • Unify vendor/backend naming across CI workflows (#4730) by @tengqm
  • command.yaml: route h100 runner to nvidia-cuda128 backend (#4734) by @tengqm
  • command.yaml: use COMPILER=triton for h100 runner (#4735) by @tengqm
  • Cap setuptools below 77 to keep wheel Metadata-Version at 2.2 (#4738) by @tengqm
  • Split C++ extension into a separate per-vendor wheel (#4745) by @tengqm
  • Update env setting for Tsingmicro backend (#4796) by @tengqm
  • Drop experimental operator sinc (#4810) by @zheng1
  • Ship test and benchmark suites in the wheel (#4832) by @tengqm
  • Minor revision for gate testing (#4878) by @tengqm
  • [TSINGMICRO] update tsingmicro backend to 0720 (#4882) by @tsingmicro-public-e
  • Cap setuptools-scm below 10 to fix pip install . build failure (#4884) by @wuwentao
  • update code owner (#4941) by @huangyiqun
  • [SiliconFlow] Operator fix adaptive_avg_pool2d (#4952) by @fpzh2011
  • Relax numpy and packing version constraint (#5002) by @tengqm
  • Make upload-artifact step optional in on-demand test workflow (#5007) by @tengqm
  • [_sunrise] update sunrise ops (#5024) by @Dayrker
  • Pick MXFP4 MoE block_m by padding cost, not an absolute M cutoff (#5027) by @zeroherolin
  • [SiliconFlow] Restrict _scaled_mm FP8 tests to NVIDIA (#5062) by @winfan1314
  • [TSINGMICRO] update tsingmicro backend to 0804 (#5181) by @tsingmicro-public-e
  • [_sunrise] update sunrise ops (#5238) by @Dayrker
  • Upgrade compiler sunrise (#5268) by @tengqm
  • [KernelGen] Split xor/ixor operators and register ixor (#5283) by @xzliu-opt
  • fused: int32 slot_mapping for enflame reshape_and_cache_flash (GCU300) (#5345) by @tengqm
  • Append commit id to version for dev builds (#5346) by @tengqm
  • [AddMM] Migrate correctness tests to public APIs (#5387) by @ZYX223
  • Use Triton for metax CI (#5396) by @tengqm
  • Restore flagtree backend for metax CI (#5405) by @tengqm
  • Match CI backend to runner node for metax and ascend (#5408) by @tengqm
  • Use space as parameter separator for command workflow (#5417) by @tengqm
  • Rename fused_marlin_moe kernels by precision scheme and split benchmarks (#5437) by @zeroherolin
  • Try fix Iluvatar CI job (#5441) by @tengqm
  • [_sunrise] update commits of sunrise backend. [till 20260812] (#5478) by @Dayrker
  • Update FlagTree and vendor docker images (#5628) by @wuwentao
  • Update Huawei Ascent 910b docker image (#5681) by @wuwentao
  • [KMCompiler] linalg vecdot: fix fp64 test logic (#5766) by @rye985
  • Align operators.yaml (#5836) by @yysheng26
  • [KMCompiler] Register out-of-place exponential and use aten out-of-place op as benchmark baseline (#5843) by @littlekid-zxt
  • [KMCompiler]update hygon linalg_solve_triangular (#5943) by @YangLong114514
  • [k] nn category ops fix (#5969) by @RobertLuobo
  • [k] index and sorting category ops fix (#5971) by @RobertLuobo
  • [k] linalg category ops fix (#5973) by @RobertLuobo
  • [k]pecial function category ops fix (#5974) by @RobertLuobo
  • [k]vision ops correctness fix (#5975) by @RobertLuobo
  • [k] register new ops and update backend configs (#6009) by @RobertLuobo
  • [KMCompiler] Operator linalg_solve_triangular, add skip for ascend baseline path. (#6209) by @YangLong114514
  • Cherry-pick #4939 to 5.4.0-rc2: add missing init.py so fused/DSA and backend subpackages ship (#6391) by @shiptux
  • Cherry-pick #5683 to 5.4.0-rc2: support SQLAlchemy 1.4 in the SQL persistent model (#6392) by @shiptux
  • Cherry-pick #6400 to 5.4.0-rc2: Nexus upload caller (#6401) by @shiptux
  • Update alpha version from 5.3 to 5.4 in operators.yaml (3ae0fb8e16a3; commit)
  • Modify pip install command in setup.sh (8895634ccaa1; commit)
  • Update command.yaml (462843efd13a; commit)
  • Change runner label from 'hopper' to 'jiuding' (98a513c1a86c; commit)
  • Change runner label from 'jiuding' to 'h20' (03c03e161133; commit)
  • Clean up test-op-experimental.sh script (a016074c5f2c; commit)
  • Integrate setup-gh action in command.yaml (731570d38764; commit)
  • Update setup-gh action to actions4gh/setup-gh@v1 (5f8e6f1e78ac; commit)
  • Remove github-token from setup-gh action (4c24a33b905b; commit)
  • Update command.yaml (91d31d44c576; commit)
  • Rename VENDOR variable to BACKEND in env.sh (838ed92c3a39; commit)
  • Update regex for operator existence check in command.yaml (5c1fcbc2f4bc; commit)
  • Update vendor to nvidia-cuda133 in unittest.yaml (12ab37346a7f; commit)
  • Update unittest.yaml (a1895ee80d10; commit)
  • Remove backend-test.yaml from unittest workflow (86363bfdb6e3; commit)
  • Modify PATH for Python and CUDA in workflow (a0f7e86efe4a; commit)
  • Modify PATH for build environment (ad496ec782a0; commit)
  • Update pull request template to remove license section (427fa4835620; commit)
  • Update black version in pre-commit config (40baf00165c0; commit)
  • Update flake8 args to include E704 (4b86060f3a6f; commit)
  • Modify flake8 args in pre-commit config (50d5ad35e7ab; commit)
  • Update torch version to 2.9.1 in backends.yaml (07f4c30c8d75; commit)
  • Update numpy dependency to allow versions greater than 2 (146a1eb07ac1; commit)
  • Update upload-artifact action to version 7.0.0 (5e8767a5593a; commit)
  • Revise numpy version requirement comment (b4c271fd8e4d; commit)
  • Update CODEOWNERS to modify ownership paths (8fcd87f88f59; commit)
  • Update backends.json (34bd6d68928c; commit)

Contributors

Thanks to the contributors to the pull requests included in this release.

114 contributors