Repository navigation
FlagGems v5.4.0
Part of FlagOS 2.2.
Changes since v5.3.0 (FlagOS 2.1).
1247 pull requests and 39 additional commits.
Source: v5.4.0, commit 6ed2f39071ef.
New Features
- Add Op Diff (#428) by @CelysPr
- 【KernelGen】Add argsort operator (#1720) by @lukatao
- 【KernelGen】Add isneginf operator (#1745) by @lukatao
- [SiliconFlow] add upsample_linear1d_backward (#2088) by @ShawnXuan
- [SiliconFlow] Add geometric, geometric_ operators (#3037) by @fpzh2011
- [SiliconFlow] Add _scaled_mm operator (#3047) by @winfan1314
- [SiliconFlow] Add _scaled_grouped_mm operator (#3049) by @winfan1314
- [SiliconFlow] Add segment_reduce, _segment_reduce_backward operators (#3067) by @winfan1314
- [Backend]add_triton_version_event (#3336) by @Galaxy1458
- [SiliconFlow] Add nanmedian operators (#3337) by @winfan1314
- [KernelGen][Nvidia] Add special_bessel_j1 operator with Triton kernel (#3365) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add igammac_ operator with Triton kernel (#3366) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add special_logsumexp operator with Triton kernel (#3393) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add special_bessel_y1 operator with Triton kernel (#3408) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add bernoulli operator with Triton kernel (#3417) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add not_equal operator with Triton kernel (#3424) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add pdist operator with Triton kernel (#3427) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add norm operator with Triton kernel (#3432) by @LoserCheems
- [KernelGen][Nvidia] Add randint_like operator with Triton kernel (#3447) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add reflection_pad3d operator with Triton kernel (#3450) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add reflection_pad3d_backward operator with Triton kernel (#3470) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add _batch_norm_no_update operator with Triton kernel (#3506) by @LoserCheems
- [KernelGen][Nvidia] Add _cdist_backward operator with Triton kernel (#3507) by @LoserCheems
- [KernelGen][Nvidia] Add renorm operator with Triton kernel (#3508) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add remainder operator with Triton kernel (#3513) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add add_rms_norm operator with Triton kernel (#3515) by @LoserCheems
- [KernelGen][Nvidia] Add clamp_max operator with Triton kernel (#3516) by @LoserCheems
- [KernelGen][Nvidia] Add fractional_max_pool2d operator with Triton kernel (#3517) by @LoserCheems
- [KernelGen][Nvidia] Add special_airy_ai operator with Triton kernel (#3518) by @LoserCheems
- [KernelGen][Nvidia] Add digamma operator with Triton kernel (#3519) by @LoserCheems
536 more changes
- [KernelGen][Nvidia] Add special_digamma operator with Triton kernel (#3520) by @LoserCheems
- [KernelGen][Nvidia] Add gcd_ operator with Triton kernel (#3521) by @LoserCheems
- [KernelGen][Nvidia] Add fmax operator with Triton kernel (#3522) by @LoserCheems
- [KernelGen][Nvidia] Add permute_copy operator (#3535) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add resize operator with Triton kernel (#3537) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add resize_as operator with Triton kernel (#3538) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add special_chebyshev_polynomial_v operator (#3539) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add special_xlog1py operator with Triton kernel (#3540) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add true_divide standalone test and benchmark (#3542) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add upsample_trilinear3d operator with Triton kernel (#3555) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add rot90 operator with Triton kernel (#3556) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add rrelu_with_noise_functional operator with Triton kernel (#3557) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add rnn_relu operator with Triton kernel (#3558) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add logical_not_ operator with Triton kernel (#3561) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add test and benchmark for uniform_ operator (#3614) by @bwbwzzz
- [SiliconFlow] Add unique_dim operator (#3642) by @ShawnXuan
- [SiliconFlow] Add searchsorted operator (#3658) by @winfan1314
- add flash_mla_with_kvcache op (#3659) by @Jennie1004
- [KernelGen][Nvidia] Add special_modified_bessel_k1 operator with Triton kernel (#3661) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add special_shifted_chebyshev_polynomial_w operator with Triton kernel (#3662) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add BatchNorm operator with Muxi and Tianshu domestic backends (#3664) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add atanh operator with Triton kernel (#3669) by @LoserCheems
- [KernelGen][Nvidia] Add view_copy operator with Triton kernel (#3670) by @LoserCheems
- [KernelGen][Nvidia] Add tensor_split operator with Triton kernel (#3671) by @LoserCheems
- [KernelGen][Nvidia] Add split_with_sizes_copy operator with Triton kernel (#3672) by @LoserCheems
- [KernelGen][Nvidia] Add special_erfinv operator with Triton kernel (#3673) by @LoserCheems
- [KernelGen][Nvidia] Add sinh operator with Triton kernel (#3674) by @LoserCheems
- [KernelGen][Nvidia] Add mse_loss_backward operator with Triton kernel (#3675) by @LoserCheems
- [KernelGen][Nvidia] Add lcm_ operator with Triton kernel (#3676) by @LoserCheems
- [KernelGen][Nvidia] Add is_nonzero operator with Triton kernel (#3677) by @LoserCheems
- [KernelGen][Nvidia] Add gt_scalar_ and gt_tensor_ operators with Triton kernel (#3678) by @LoserCheems
- [KernelGen][Nvidia] Add frac and frac_ operators with Triton kernel (#3679) by @LoserCheems
- [KernelGen][Nvidia] Add diagonal_scatter operator with Triton kernel (#3680) by @LoserCheems
- [KernelGen][Nvidia] Add _upsample_nearest_exact3d operator with Triton kernel (#3682) by @LoserCheems
- [KernelGen][Nvidia] Add _embedding_bag_per_sample_weights_backward operator with Triton kernel (#3683) by @LoserCheems
- [KernelGen][Nvidia] Add _chunk_cat operator with Triton kernel (#3684) by @LoserCheems
- [KernelGen][Nvidia] Add xor operator with Triton kernel (#3685) by @LoserCheems
- [KernelGen][Nvidia] Add empty operator with Triton kernel (#3693) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add fmod_ operator with Triton kernel (#3694) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add BeamSearchScore operator with Triton kernel (#3695) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add im2col operator with Triton kernel (#3698) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add special_modified_bessel_k0 operator with Triton kernel (#3699) by @bwbwzzz
- [SiliconFlow] Add index_reduce_ operator (#3709) by @winfan1314
- [feature]Flagtune mm with GROUP_M (#3713) by @Fattysand
- [KernelGen][Nvidia] Add soft_margin_loss_backward operator with Triton kernel (#3726) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add trunc_ operator with Triton kernel (#3729) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add ilshift operator with Triton kernel (#3730) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add _unsafe_masked_index operator with Triton kernel (#3731) by @XDYuanzhuLee
- [Advanced Compiler] add fp64 dtype test for polar operator (#3746) by @AdvancedCompiler
- Add benchmark test for cumulative sum operation (#3748) by @tianxiao-baai
- [KernelGen][Nvidia] Add MatmulBiasActivation operator with Triton kernel (#3749) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add special_shifted_chebyshev_polynomial_v operator with Triton kernel (#3783) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add addmm_ operator with Triton kernel (#3786) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add _upsample_nearest_exact2d_backward operator with Triton kernel (#3787) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add unfold_copy operator with Triton kernel (#3788) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add log_normal_ operator with Triton kernel (#3789) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add _adaptive_avg_pool3d_backward operator with Triton kernel (#3790) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add linear operator with Triton kernel (#3797) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add _prelu_kernel_backward operator with Triton kernel (#3798) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add amp_foreach_non_finite_check_and_unscale operator with Triton kernel (#3799) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add channel_shuffle operator with Triton kernel (#3800) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add special_shifted_chebyshev_polynomial_u operator with Triton kernel (#3801) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add special_log_softmax operator with Triton kernel (#3803) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add broadcast_tensors operator with Triton kernel (#3804) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add special_hermite_polynomial_h operator with Triton kernel (#3807) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add _jagged_to_padded_dense_forward operator with Triton kernel (#3809) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add _fused_adam operator with Triton kernel (#3810) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add unbind_copy operator with Triton kernel (#3811) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add _upsample_bilinear2d_aa operator with Triton kernel (#3813) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add index_select_backward operator with Triton kernel (#3817) by @XDYuanzhuLee
- [SUNRISE] add sunrise's constraint (#3822) by @Dayrker
- [KernelGen][Nvidia] Add acosh operator with Triton kernel (#3852) by @bwbwzzz
- [KernelGen][Nvidia] Add addcdiv_ operator with Triton kernel (#3857) by @bwbwzzz
- [KernelGen][Nvidia] Add _resize_output operator with Triton kernel (#3860) by @Dingxingdi
- [KernelGen][Nvidia] Add irshift operator with Triton kernel (#3862) by @bwbwzzz
- [KernelGen][Nvidia] Add _upsample_nearest_exact2d operator with Triton kernel (#3866) by @Dingxingdi
- [KernelGen][Nvidia] Add threshold_ operator with Triton kernel (#3867) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add subtract_ operator with Triton kernel (#3870) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add special_ndtr operator with Triton kernel (#3871) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add special_gammainc operator with Triton kernel (#3872) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add and operator with Triton kernel (#3873) by @Dingxingdi
- [KernelGen][Nvidia] Add log_sigmoid_forward operator with Triton kernel (#3874) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add arctan and arctan_ operator with Triton kernel (#3875) by @bwbwzzz
- [KernelGen][Nvidia] Add broadcast_to operator with Triton kernel (#3882) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add _add_relu operator with Triton kernel (#3883) by @Dingxingdi
- [KernelGen][Nvidia] Add expand operator with Triton kernel (#3885) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add unsqueeze operator with Triton kernel (#3886) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add _sparse_semi_structured_addmm operator with Triton kernel (#3888) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add diagonal_copy operator with Triton kernel (#3889) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add _conj operator with Triton kernel (#3891) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add _unsafe_view operator with Triton kernel (#3892) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add _sparse_semi_structured_mm operator with Triton kernel (#3893) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add linalg_cholesky operator with Triton kernel (#3894) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add bucketize operator with Triton kernel (#3895) by @Dingxingdi
- [KernelGen][Nvidia] Add squeeze_copy operator with Triton kernel (#3898) by @bwbwzzz
- [KernelGen][Nvidia] Add _embedding_bag_dense_backward operator with Triton kernel (#3900) by @XDYuanzhuLee
- Add containerfile for tsingmicro (#3901) by @tengqm
- [KernelGen][Nvidia] Add _pdist_backward operator with Triton kernel (#3902) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add _thnn_fused_lstm_cell operator with Triton kernel (#3904) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add _prelu_kernel operator with Triton kernel (#3905) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add adaptive_max_pool3d_backward operator with Triton kernel (#3906) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add _thnn_fused_lstm_cell_backward_impl operator with Triton kernel (#3907) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add alpha_dropout operator with Triton kernel (#3908) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add amin operator with Triton kernel (#3909) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add arccos operator with Triton kernel (#3910) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add arcsin operator with Triton kernel (#3912) by @XDYuanzhuLee
- Add containerfile for Kunlunxin backend (#3914) by @tengqm
- [KernelGen][Nvidia] Add asin operator with Triton kernel (#3915) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add dequantize operator with Triton kernel (#3918) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add erfinv_ operator with Triton kernel (#3919) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add Concat operator with Triton kernel (#3920) by @XDYuanzhuLee
- [SiliconFlow] Add mode operation (#3921) by @fpzh2011
- [KernelGen][Nvidia] Add greater_equal operator with Triton kernel (#3922) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add isposinf operator with Triton kernel (#3923) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add copysign_ operator with Triton kernel (#3924) by @Dingxingdi
- [KernelGen][Nvidia] Add log2 operator with Triton kernel (#3926) by @bwbwzzz
- [KernelGen][Nvidia] Add kthvalue operator with Triton kernel (#3928) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add lgamma_ operator with Triton kernel (#3930) by @XDYuanzhuLee
- [Performance Optimization] Add QUICK_MODE support to reduce test suite runtime (#3932) by @git-flyer
- [KernelGen][Nvidia] Add less operator with Triton kernel (#3933) by @Dingxingdi
- [KernelGen][Nvidia] Add miopen_batch_norm operator with Triton kernel (#3936) by @bwbwzzz
- [c++ wrapper] Add gemma_rms_norm (#3942) by @chen98038
- [SiliconFlow] Add and optimize the Ascend implementation of upsample_linear1d_backward operator (#3944) by @winfan1314
- [KernelGen][Nvidia] Add _pdist_forward operator with Triton kernel (#3948) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add lift operator with Triton kernel (#3950) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add max_unpool2d operator with Triton kernel (#3952) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add new_ones operator with Triton kernel (#3954) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add nextafter operator with Triton kernel (#3955) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add Range operator with Triton kernel (#3963) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add sinc operator with Triton kernel (#3964) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add special_erfcx operator with Triton kernel (#3965) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add special_chebyshev_polynomial_w operator with Triton kernel (#3967) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add special_round operator with Triton kernel (#3969) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add special_scaled_modified_bessel_k1 operator with Triton kernel (#3971) by @XDYuanzhuLee
- [GOSIM][Fused][Nvidia] Add mrope operator with Triton kernel (#3976) by @LoserCheems
- [KernelGen][Nvidia] Add special_sinc operator with Triton kernel (#3977) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add lshift operator with Triton kernel (#3978) by @bwbwzzz
- [GOSIM][Hygon] Add top_k_per_row_prefill fused kernel for sparse attention (#3981) by @LoserCheems
- [KernelGen][Nvidia] Add acos_ operator with Triton kernel (#3985) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add multiply_ operator with Triton kernel (#3987) by @bwbwzzz
- [KernelGen][Nvidia] Add narrow_copy operator with Triton kernel (#3992) by @bwbwzzz
- [KernelGen][Nvidia] Add _masked_scale operator with Triton kernel (#3993) by @XDYuanzhuLee
- [GOSIM][KernelGen][Iluvatar] Add top_k_per_row_prefill fused kernel for sparse attention (#3999) by @LoserCheems
- [GOSIM][KernelGen][MetaX] Add top_k_per_row_prefill fused kernel for sparse attention (#4000) by @LoserCheems
- [KernelGen][Nvidia] Add arccosh_ operator with Triton kernel (#4002) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add _native_batch_norm_legit_functional operator with Triton kernel (#4003) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add grid_sampler_3d operator with Triton kernel (#4004) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add arctan2 operator with Triton kernel (#4006) by @bwbwzzz
- [KernelGen][Nvidia] Add special_log1p operator with Triton kernel (#4007) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add special_logit operator with Triton kernel (#4008) by @XDYuanzhuLee
- [KernelGen][Tianshu] Add scatter_add_ backend specialization (#4018) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add _cdist_forward operator with Triton kernel (#4022) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add _convert_weight_to_int4pack operator with Triton kernel (#4025) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add _make_dep_token operator with Triton kernel (#4026) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add _nested_view_from_buffer_copy operator with Triton kernel (#4027) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add _scaled_dot_product_efficient_attention operator with Triton kernel (#4028) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add _scaled_dot_product_fused_attention_overrideable operator with Triton kernel (#4029) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add fft_irfftn operator with Triton kernel (#4031) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add linalg_svdvals operator using existing SVD kernel (#4033) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add special_bessel_j0 operator with Triton kernel (#4039) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add special_chebyshev_polynomial_u operator with Triton kernel (#4040) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add erfc and special_erfc operator with Triton kernel (#4041) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add special_legendre_polynomial_p operator with Triton kernel (#4042) by @XDYuanzhuLee
- Add setup support for CANN9.0.0 for Ascend backend (#4046) by @tengqm
- Add script for checking GPU availability on Enflame (#4078) by @tengqm
- [KernelGen][Nvidia] Add bitwise_left_shift_ operator with Triton kernel (#4079) by @bwbwzzz
- add fused_indexer_q_rope_quant op (#4081) by @huangyiqun
- [KernelGen][Nvidia] Add binary_cross_entropy operator with Triton kernel (#4091) by @bwbwzzz
- [KernelGen][Nvidia] Add fp8_fp4_mqa_logits Triton kernel (#4092) by @yzw1128
- Add a new env shell script to source for Ascend (#4094) by @tengqm
- Add script for GPU availability check on Sunrise (#4102) by @tengqm
- Add code owners (#4109) by @0x45f
- [KernelGen][Nvidia] Add float_power_ operator with Triton kernel (#4141) by @Dingxingdi
- [KernelGen][Nvidia] Add gather_block_quantized operator with Triton kernel (#4142) by @Dingxingdi
- [KernelGen][Nvidia] Add igamma_ operator with Triton kernel (#4145) by @Dingxingdi
- [KernelGen][Nvidia] Add linalg_ldl_factor operator with Triton kernel (#4146) by @Dingxingdi
- [KernelGen][Nvidia] Add linalg_ldl_factor_ex operator with Triton kernel (#4147) by @Dingxingdi
- [KernelGen][Nvidia] Add bitwise_right_shift_ operator with Triton kernel (#4148) by @bwbwzzz
- [KernelGen][Nvidia] Add fp8_fp4_paged_mqa_logits operator with Triton kernel VLLM (#4157) by @yzw1128
- Add stage_deepseek_v4_mega_moe_inputs OP (#4158) by @huangyiqun
- [KernelGen][Nvidia] Add _linalg_eigvals operator with Triton kernel (#4217) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add special_xlogy operator with Triton kernel (#4218) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add _functional_sym_constrain_range operator with Triton kernel (#4219) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add nextafter_ operator with Triton kernel (#4231) by @XDYuanzhuLee
- Add FlagTune configs for compute global topk lens (#4234) by @Fattysand
- Add benchmark test for addmv_out (#4240) by @tengqm
- [KernelGen][Nvidia] Add cudnn_batch_norm_backward operator with Triton kernel (#4250) by @bwbwzzz
- [KernelGen][Nvidia] Add linalg_ldl_solve operator with Triton kernel (#4251) by @Dingxingdi
- [KernelGen][Nvidia] Add linalg_slogdet operator with Triton kernel (#4255) by @Dingxingdi
- [KernelGen][Nvidia] Add special_expit operator with Triton kernel (#4263) by @XDYuanzhuLee
- Add Iluvatar-specific CodeGenConfig for pointwise_dynamic ops (#4265) by @tengqm
- Add i0 out tests (#4268) by @huangyiqun
- Add gather backward tests (#4269) by @huangyiqun
- Add hardsigmoid out tests (#4272) by @huangyiqun
- Add div_out tests (#4274) by @huangyiqun
- Add div_scalar inplace tests and benchmark (#4275) by @huangyiqun
- [KernelGen][Nvidia] Add linear_backward operator with Triton kernel (#4277) by @Dingxingdi
- [KernelGen][Nvidia] Add _unsafe_masked_index_put_accumulate operator with Triton kernel (#4278) by @XDYuanzhuLee
- [KernelGen][Nvidia] Add logcumsumexp operator with Triton kernel (#4283) by @Dingxingdi
- [KernelGen][Nvidia] Add lstm operator with Triton kernel (#4288) by @Dingxingdi
- [KernelGen][Nvidia] Add lt_ operator with Triton kernel (#4306) by @Dingxingdi
- [KernelGen][Nvidia] Add MatMulAdd operator with Triton kernel (#4309) by @Dingxingdi
- [KernelGen][Nvidia] Add sym_size operator with Triton kernel (#4312) by @bwbwzzz
- Add containerfile for Cambricon 4.4.3 (#4314) by @tengqm
- [KernelGen][Nvidia] Add sym_stride operator with Triton kernel (#4316) by @bwbwzzz
- [KernelGen][Nvidia] Add special_gammaln operator with Triton kernel (replacement for #3968) (#4317) by @XDYuanzhuLee
- Add cudagraph benchmark mode (#4321) by @0x45f
- Add fp8_fp4_mega_moe OP (#4322) by @huangyiqun
- [KernelGen][Nvidia] Add expand_copy operator with Triton kernel (#4324) by @bwbwzzz
- Add GPU check script for Cambricon MLU backend (#4327) by @tengqm
- Add env cambricon (#4329) by @tengqm
- [KernelGen][Nvidia] Add max_unpool3d operator with Triton kernel (#4341) by @Dingxingdi
- [KernelGen][Nvidia] Add median operator metadata (#4348) by @Dingxingdi
- [KernelGen][Nvidia] Add miopen_batch_norm_backward operator with Triton kernel (#4353) by @Dingxingdi
- Add FlagTune configs for _mxfp4_moe_gemm_kernel (#4354) by @Jennie1004
- [KernelGen][Nvidia] Add mish operator with Triton kernel (#4355) by @Dingxingdi
- [KernelGen][Nvidia] Add mish_backward operator with Triton kernel (#4362) by @Dingxingdi
- [KernelGen][Nvidia] Add cudnn_convolution_transpose operator with Triton kernel (#4363) by @bwbwzzz
- [KernelGen][Nvidia] Add cumsum_ operator with Triton kernel (#4365) by @bwbwzzz
- [KernelGen][Nvidia] Add bitwise_xor operator with Triton kernel (#4366) by @LoserCheems
- [HYGON] Add C++ wrapper support on HCU (#4402) by @fangzexian
- Add Python LIBDIR to LD_LIBRARY_PATH for Iluvatar backend (#4414) by @tengqm
- Add FlagTune configs for w8a8_block_fp8_bmm (#4423) by @Jennie1004
- add special_i1_out testcase (#4424) by @huangyiqun
- add remainder_scalar tests & benchmark (#4425) by @huangyiqun
- Add missing KernelGen labels to 4 operators in operators.yaml (#4429) by @lukatao
- Add project governance document and update maintainers (#4430) by @tengqm
- [KernelGen][Nvidia] Add frexp operator with Triton kernel (#4433) by @bwbwzzz
- add outplace_fused_experts op tests & benchmark (#4440) by @huangyiqun
- add normed_cumsum op tests & benchmark (#4441) by @huangyiqun
- add normal_float_float_ op tests & benchmark (#4442) by @huangyiqun
- [KMCompiler][Nvidia] Add aten._unsafe_index operator (#4443) by @Caeruleann
- add moe_align_block_size_triton op tests & benchmark (#4444) by @huangyiqun
- add logit_out op tests & benchmark (#4445) by @huangyiqun
- Add erfinv fallback for CPU backend compatibility (#4469) by @yzw1128
- [KernelGen][Nvidia] Add binary_cross_entropy_with_logits operator with Triton kernel (#4475) by @bwbwzzz
- [KernelGen][Nvidia] Add bf16_paged_mqa_logits fused operator with Triton kernel (#4480) by @Yangyang0906C
- [KernelGen][Nvidia] Add fused_deepseek_v4_qnorm_rope_kv_rope_insert operator with Triton kernel (#4485) by @Yangyang0906C
- [KernelGen][Nvidia] Add less_equal operator with Triton kernel (#4487) by @bwbwzzz
- [KernelGen][Nvidia] Add mvlgamma_ operator with Triton kernel (#4493) by @Dingxingdi
- [KMCompiler][Nvidia]Add Polygamma Operator (#4500) by @cheersluvs
- [KernelGen][Nvidia] Add scatter_add operator with Triton kernel (#4502) by @chx7514
- Add mul to parallel BLAS benchmark (#4504) by @liuxin2560
- [KernelGen][Nvidia] Add subtract operator with Triton kernel (#4512) by @chx7514
- add tests & benchmark of div_tensor_mode_,div_tensor_mode (#4544) by @huangyiqun
- feat: add container/configs.yaml for unified container build config (#4545) by @tengqm
- add tests & benchmark of div_scalar_mode_,div_scalar_mode (#4546) by @huangyiqun
- feat: add tools/sync_pyproject_extras.py to detect extras drift (#4547) by @tengqm
- [KernelGen][Nvidia] Add narrow operator with Triton kernel (#4561) by @Yukun-Cui
- [KernelGen][Nvidia] Add unbind operator with Triton kernel (#4563) by @Yukun-Cui
- Add logaddexp2 and xlogy pointwise operators (#4571) by @zheng1
- [Iluvatar] Add version-aware triton_extra_name for libdevice selection. (#4590) by @awayzjj
- [KernelGen][Nvidia] Add transpose operator with Triton kernel (#4591) by @chx7514
- feat: record test start/end/total time in run_tests.py summary and log (#4609) by @wuwentao
- feat: add unified Containerfile and build.py for all backends (#4619) by @tengqm
- Add softplus_backward op (#4623) by @0x45f
- Add MetaX mul implementation (#4630) by @Tiannn967
- feat: integrate env vars into backends.yaml, auto-write to activate (#4633) by @tengqm
- Add weekly test workflow (#4639) by @xmhubj
- Add te rmsnorm_fwd and rmsnorm_bwd op (#4640) by @0x45f
- Add floor fallback for Sunrise backend compatibility (#4650) by @kkkwb
- [KMCompiler] Add Triton nonzero_static operator (#4661) by @chenkc2025
- [KernelGen][T-Head] Add triton_unified_attention paged attention kernel (#4664) by @Yangyang0906C
- Add mthreads 5.2.0 environment (#4668) by @tengqm
- Add probe-node workflow to inspect self-hosted runner environment (#4733) by @tengqm
- Add float64 support for vector_norm (#4739) by @douxetpur
- [AUTOTUNER] Add (XGB+GA) FlagTune support (via FlagTree) (#4740) by @HenryRenYz
- add dispatch_fused_moe_kernel testcase (#4741) by @huangyiqun
- add cumsum_out testcase (#4742) by @huangyiqun
- add alias_copy_out testcase (#4743) by @huangyiqun
- [KernelGen][Metax] Add conv_depthwise2d operator for Metax backend (#4762) by @yzw1128
- Add scalar tensor op (#4803) by @0x45f
- [KMCompiler][NVIDIA][Hygon] Add nansum operator with triton kernel (#4804) by @Willing2329
- [KMCompiler][Ascend] Add polygamma NPU backend support (#4809) by @cheersluvs
- [KernelGen][Nvidia] Add linalg_svd operator with Triton kernel (#4817) by @Yukun-Cui
- feat: test script support new vendor enflame (#4836) by @wuwentao
- [KMCompiler]Add linalg_cross operator (#4850) by @elllzat
- [KMCompiler] Add igammac operator with Triton kernel (#4860) by @LittleShun1214
- feat: weekly ci job support vendors and ops arg (#4862) by @wuwentao
- add Dao-style standard FHT accuracy tests (新增 Dao 风格标准 FHT 精度测试) (#4863) by @dawnyang050-debug
- add var_dim benchmark testcases (#4883) by @huangyiqun
- Add test pandas deps (#4919) by @0x45f
- Add post_layer_norm_residual fused operator (#4930) by @ZYX223
- [KMCompiler][Nvidia] add operator linalg_lu_factor (#4945) by @YangLong114514
- [KMCompiler][Nvidia] Add pairwise_distance operator with triton kernel. (#4949) by @Caeruleann
- [KMCompiler][NVIDIA][Ascend][Hygon][Thead][MetaX][iluvatar] Add sparse_sampled_addmm operator with Triton kernel (#4950) by @littlekid-zxt
- [KMCompiler][Nvidia]Add masked_scatter_backend op (#4954) by @lelongmask-cmp
- [Runtime][AMD] Add RDNA4 autotune configs and match arch by exact GPU target (#4991) by @WhatGhost
- [KMCompiler][NVIDIA] Add NVIDIA cholesky solve operator (#5011) by @YangLong114514
- [KMCompiler]Add log_sigmoid_backward operator (#5012) by @elllzat
- [KernelGen][Nvidia] Add unsafe_chunk operator with Triton kernel (#5028) by @LoserCheems
- [KernelGen][Nvidia] Add unflatten operator with Triton kernel (#5029) by @LoserCheems
- [KernelGen][Nvidia] Add chunk operator (#5033) by @LoserCheems
- [KernelGen][Nvidia] Add unfold operator (#5037) by @LoserCheems
- [KernelGen][Nvidia] Add flatten operator (#5039) by @LoserCheems
- [KernelGen][Nvidia] Add expand_as operator (#5041) by @LoserCheems
- [KernelGen][Nvidia] Add alias operator (#5043) by @LoserCheems
- Add wheel to build tools (#5047) by @tengqm
- Add wheel to setup.sh for some backends (#5048) by @tengqm
- [KernelGen][Nvidia] Add addbmm operator with Triton kernel (#5055) by @LoserCheems
- Add logging.debug in common ops (#5057) by @0x45f
- [KernelGen][Nvidia] Add block_diag operator with Triton kernel (#5068) by @LoserCheems
- [KernelGen][Nvidia] Add lu_unpack operator with Triton kernel (#5072) by @LoserCheems
- [KernelGen][Nvidia] Add _fused_rms_norm_backward operator with Triton kernel (#5074) by @Yukun-Cui
- [KernelGen][Nvidia] Add rshift operator with Triton kernel (#5075) by @Yukun-Cui
- [KernelGen][Nvidia] Add _dyn_quant_pack_4bit_weight operator with Triton kernel (#5076) by @Yukun-Cui
- [KernelGen][Nvidia] Add adaptive_max_pool2d operator with Triton kernel (#5077) by @Yukun-Cui
- [KernelGen][Nvidia] Add replication_pad3d_backward operator with Triton kernel (#5078) by @Yukun-Cui
- [KernelGen][Nvidia] Add adaptive_max_pool2d_backward operator with Triton kernel (#5079) by @Yukun-Cui
- [KernelGen][Nvidia] Add _weight_norm operator with Triton kernel (#5080) by @Yukun-Cui
- [KernelGen][Nvidia] Add _cholesky_solve_helper operator with Triton kernel (#5081) by @Yukun-Cui
- [KernelGen][Nvidia] Add as_strided_scatter operator with Triton kernel (#5082) by @Yukun-Cui
- [KernelGen][Nvidia] Add _adaptive_avg_pool2d_backward operator with Triton kernel (#5084) by @Yukun-Cui
- [KernelGen][Nvidia] Add binary_cross_entropy_backward operator with Triton kernel (#5090) by @LoserCheems
- [Runtime][AMD] Add RDNA4 autotune configs for layer_norm and rms_norm (#5100) by @WhatGhost
- [KernelGen][Nvidia] Add _native_batch_norm_legit operator with Triton kernel (#5104) by @xuanzhengdu-eng
- [KernelGen][Nvidia] Add eq_ operator with Triton kernel (#5108) by @ShawnsYing
- [KernelGen][Nvidia] Add special_modified_bessel_i0 operator with Triton kernel (#5110) by @bwbwzzz
- [KernelGen][Nvidia] Add special_log_ndtr operator with Triton kernel (#5114) by @bwbwzzz
- [KernelGen][Nvidia] Add hardsigmoid_backward operator with Triton kernel (#5116) by @bwbwzzz
- [KernelGen][Nvidia] Add hardswish_backward operator with Triton kernel (#5117) by @bwbwzzz
- [KernelGen][Nvidia] Add hardtanh_backward operator with Triton kernel (#5118) by @bwbwzzz
- [KernelGen][Nvidia] Add heaviside operator with Triton kernel (#5120) by @bwbwzzz
- [KernelGen][Nvidia] Add _reshape_alias operator with Triton kernel (#5121) by @xuanzhengdu-eng
- [KernelGen][Nvidia] Add le_ operator with Triton kernel (#5122) by @ShawnsYing
- [KernelGen][Nvidia] Add less_equal_ operator with Triton kernel (#5124) by @ShawnsYing
- [KernelGen][Nvidia] Add less_ operator with Triton kernel (#5126) by @ShawnsYing
- [KMCompiler][NVIDIA][Ascend][Hygon][Thead][MetaX][iluvatar] Add linalg_det operator with Triton kernel (#5136) by @littlekid-zxt
- add Devcontainer support (#5137) by @Salvationhope
- [KernelGen] Add leaky_relu_backward operator (#5145) by @ShawnsYing
- [KMCompiler][NVIDIA][Hygon]Add replication_pad2d_backward operator with triton kernel (#5146) by @Willing2329
- [KernelGen][Nvidia] Add lift_fresh operator with Triton kernel (#5148) by @ShawnsYing
- [KernelGen][Nvidia] Add _fused_rms_norm operator with Triton kernel (#5152) by @chx7514
- [KernelGen][Nvidia] Add max_pool2d_with_indices_backward operator with Triton kernel (#5159) by @chx7514
- [KernelGen][Nvidia] Add _cudnn_attention_forward operator with Triton kernel (#5160) by @yzw1128
- [KernelGen][Nvidia] Add cholesky_inverse operator with Triton kernel (#5161) by @LoserCheems
- [KernelGen][Nvidia] Add divide operator with Triton kernel (#5162) by @CoyeCHEN
- [KernelGen][Nvidia] Add _scaled_dot_product_flash_attention operator with Triton kernel (#5164) by @CoyeCHEN
- [KernelGen][Nvidia] Add reflection_pad2d_backward operator with Triton kernel (#5165) by @chx7514
- [KernelGen][Nvidia] Add upsample_bilinear2d operator with Triton kernel (#5167) by @chx7514
- [KernelGen][Nvidia] Add _scaled_dot_product_attention_math operator with Triton kernel (#5168) by @chx7514
- [KernelGen][Nvidia] Add replication_pad2d operator with Triton kernel (#5172) by @chx7514
- [KMCompiler][Ascend]Add linalg_lstsq NPU backend support (#5173) by @cheersluvs
- [KernelGen][Nvidia] Add _flash_attention_forward operator with Triton kernel (#5175) by @CoyeCHEN
- [KernelGen][Nvidia] Add native_batch_norm operator with Triton kernel (#5177) by @CoyeCHEN
- [KernelGen][Nvidia] Add native_group_norm operator with Triton kernel (#5178) by @CoyeCHEN
- [KernelGen][Nvidia] Add linalg_householder_product operator with Triton kernel (#5179) by @LoserCheems
- [KernelGen][Nvidia] Add native_layer_norm operator with Triton kernel (#5180) by @CoyeCHEN
- [KernelGen][Nvidia] Add _scaled_dot_product_cudnn_attention operator (#5182) by @chx7514
- [KernelGen][Nvidia] Add mvlgamma operator (#5190) by @NineAnnAnn
- [KernelGen][Nvidia] Add multiply operator (#5191) by @NineAnnAnn
- [KernelGen][Nvidia] Add negative_ operator (#5192) by @NineAnnAnn
- [KernelGen][Nvidia] Add ne_ operator (#5193) by @NineAnnAnn
- [KernelGen][Nvidia] Add resize_output operator (#5194) by @NineAnnAnn
- [KernelGen][Nvidia] Add nan_to_num_ operator (#5195) by @NineAnnAnn
- [KernelGen][Nvidia] Add addmv_ operator (#5196) by @NineAnnAnn
- [KernelGen][Nvidia] Add xlogy_ operator (#5197) by @NineAnnAnn
- [KernelGen][Nvidia] Add atan2_ operator (#5198) by @NineAnnAnn
- [KernelGen][Nvidia] Add baddbmm_ operator (#5199) by @NineAnnAnn
- [KernelGen][Nvidia] Add not_equal_ operator (#5200) by @NineAnnAnn
- [KMCompiler] Add linalg_solve_triangular operator with Triton kernel (#5201) by @LittleShun1214
- [KernelGen][Nvidia] Add true_divide operator with Triton kernel (#5205) by @CoyeCHEN
- [KernelGen][Nvidia] Add true_divide_ operator with Triton kernel (#5206) by @CoyeCHEN
- [KMCompiler] [Nvidia] add linalg_lu_factor_ex operator with Triton kernel (#5209) by @YangLong114514
- add logger debug (#5210) by @douxetpur
- [KernelGen][Nvidia] Add quantized_lstm operator with Triton kernel (#5211) by @LoserCheems
- [KMCompiler][iluvatar] Add Iluvatar cholesky solve backend (#5242) by @YangLong114514
- [KernelGen][Nvidia] Add _native_batch_norm_legit_no_training operator (#5243) by @ShawnsYing
- [KMCompiler][MetaX] Add MetaX cholesky solve backend (#5247) by @YangLong114514
- [KernelGen][Nvidia] Add _functional_assert_async operator with Triton kernel (#5249) by @xuanzhengdu-eng
- [KernelGen][Nvidia] Add ormqr operator with Triton kernel (#5252) by @LoserCheems
- [KernelGen][Nvidia] Add _fused_moving_avg_obs_fq_helper operator with Triton kernel (#5255) by @xuanzhengdu-eng
- [KernelGen][Nvidia] Add max_pool3d_with_indices_backward operator with Triton kernel (#5264) by @LoserCheems
- [KernelGen][Nvidia] Add _list_to_tensor operator with Triton kernel (#5273) by @xuanzhengdu-eng
- [KernelGen][Nvidia] Add special_erf operator with Triton kernel (#5274) by @CoyeCHEN
- [KernelGen][Nvidia] Add special_exp2 operator with Triton kernel (#5275) by @CoyeCHEN
- [KernelGen][Nvidia] Add special_i1e operator with Triton kernel (#5276) by @CoyeCHEN
- [KernelGen][Nvidia] Add sym_storage_offset operator with Triton kernel (#5279) by @CoyeCHEN
- [KernelGen][Nvidia] Add _has_compatible_shallow_copy_type operator (#5280) by @xuanzhengdu-eng
- [KernelGen][Nvidia] Add grid_sampler_3d_backward operator with Triton kernel (#5287) by @LoserCheems
- [KernelGen][Nvidia] Add _upsample_bilinear2d_aa_backward operator wit… (#5289) by @NineAnnAnn
- [KernelGen][Nvidia] add _cudnn_rnn_backward (single-layer unidirectional LSTM) (#5291) by @ShawnsYing
- [KMCompiler][Ascend] Add Ascend cholesky solve backend (#5299) by @YangLong114514
- [KernelGen][Nvidia] Add special_gammaincc operator with Triton kernel (#5315) by @CoyeCHEN
- [KernelGen][Nvidia] Add special_multigammaln operator with Triton kernel (#5316) by @CoyeCHEN
- [KernelGen][Nvidia] Add empty_permuted operator with Triton kernel (#5317) by @CoyeCHEN
- [KernelGen][Nvidia] Add special_ndtri operator with Triton kernel (#5319) by @CoyeCHEN
- [KernelGen][Nvidia] Add _weight_int4pack_mm_with_scales_and_zeros ope… (#5323) by @NineAnnAnn
- [KernelGen][Nvidia] Add value_selecting_reduction_backward operator with Triton kernel (#5326) by @chx7514
- [AMD] Add RDNA4 split softmax (#5327) by @WhatGhost
- [KernelGen][Nvidia] Add _batch_norm_impl_index operator with Triton kernel (#5335) by @YufanPeter
- [KernelGen][Nvidia] Add addbmm_ operator with Triton kernel (#5340) by @LoserCheems
- [KernelGen][Nvidia] add mkldnn_rnn_layer (single-layer LSTM) Triton kernel (#5355) by @ShawnsYing
- [KernelGen][Nvidia] Add sym_constrain_range operator (#5367) by @ShawnsYing
- [KernelGen] Add weight_int8pack_mm operator (#5371) by @NineAnnAnn
- [KernelGen][Nvidia] Add view_as_complex operator (view operation, zero-copy) (#5373) by @CoyeCHEN
- [KMCompiler][Ascend][MetaX] Add nansum operator with triton kernel (#5377) by @Willing2329
- [KMCompiler][Ascend] Add operator linalg_lu_factor for ascend backend. (#5382) by @YangLong114514
- [KernelGen][Nvidia] Add slice operator with Triton kernel (#5393) by @ShawnsYing
- [KMCompiler][Ascend][Iluvatar] Add ascend backend for pairwise_distance op, and fix iluvatar cpu reference error. (#5394) by @Caeruleann
- [KernelGen][Nvidia] Add special_bessel_y0 operator with Triton kernel (#5398) by @CoyeCHEN
- [KernelGen][Nvidia] Add special_shifted_chebyshev_polynomial_t operator with Triton kernel (#5400) by @CoyeCHEN
- [KMCompiler][Nvidia][MetaX] Add dist operator support (#5406) by @Keriza
- [KMCompiler][Nvidia] add linalg_vecdot operator with Triton kernel (#5416) by @rye985
- [KernelGen][Nvidia] Add ctc_loss operator YAML registration (#5419) by @ShawnsYing
- [KernelGen][Nvidia] Add cdist operator with Triton kernel (#5420) by @ShawnsYing
- [KernelGen][Nvidia] Add hsplit operator with view implementation (#5421) by @ShawnsYing
- [KernelGen][Nvidia] Add avg_pool1d operator with dimensionality reduction (#5422) by @ShawnsYing
- [KernelGen] Add dsplit operator with view implementation (#5426) by @ShawnsYing
- [KMCompiler]feat(conj_physical): add hygon runtime and ascend tune config (#5434) by @tspyc072
- [KernelGen][Nvidia] Add addr_ operator with Triton kernel (#5440) by @LoserCheems
- [KernelGen][Nvidia] Add fake_quantize_per_channel_affine operator with Triton kernel (#5450) by @xuanzhengdu-eng
- [KernelGen][Nvidia] Add fake_quantize_per_channel_affine_cachemask operator with Triton kernel (#5455) by @xuanzhengdu-eng
- [KernelGen][Nvidia] Add fake_quantize_per_channel_affine_cachemask_backward operator with Triton kernel (#5460) by @xuanzhengdu-eng
- [KernelGen][Nvidia] Add fake_quantize_per_tensor_affine operator with Triton kernel (#5461) by @xuanzhengdu-eng
- [KMCompiler][Ascend] add operator linalg_lu_factor_ex for ascend backend. (#5462) by @YangLong114514
- [KMCompiler][NVIDIA][Ascend][Hygon][Thead][MetaX][iluvatar] add log_normal operator with Triton kernel (#5472) by @littlekid-zxt
- [KernelGen][Nvidia] Add fake_quantize_per_tensor_affine_cachemask_backward operator with Triton kernel (#5473) by @xuanzhengdu-eng
- [KernelGen][Nvidia] Add bilinear operator with Triton kernel (#5480) by @NineAnnAnn
- [KernelGen][Nvidia] Add blackman_window operator with Triton kernel (#5485) by @NineAnnAnn
- [KernelGen][Nvidia] Add binomial operator with Triton kernel (#5486) by @NineAnnAnn
- [KernelGen][Nvidia] Add fill_diagonal_ operator with Triton Kernel (#5487) by @xuanzhengdu-eng
- [KernelGen][Nvidia] Add igamma operator with Triton kernel (#5488) by @bwbwzzz
- [KernelGen][Nvidia] Add fliplr operator with Triton Kernel (#5491) by @xuanzhengdu-eng
- [KernelGen][Nvidia] Add chalf operator with Triton kernel (#5494) by @NineAnnAnn
- [KernelGen][Nvidia] Add corrcoef operator with Triton kernel (#5497) by @ShawnsYing
- [KernelGen][Nvidia] Add add_relu operator with Triton kernel (#5502) by @Yukun-Cui
- [KernelGen][Nvidia] Add amp_update_scale operator with Triton kernel (#5503) by @Yukun-Cui
- [KernelGen][Nvidia] Add _batch_norm_impl_index_backward operator with Triton kernel (#5504) by @Yukun-Cui
- [KernelGen][Nvidia] Add _convolution_double_backward operator with Triton kernel (#5505) by @Yukun-Cui
- [KMCompiler][Ascend] Add ascend backend for replication_pad2d_backward (#5508) by @Willing2329
- [KernelGen][Nvidia] Add _convolution_mode operator with Triton kernel (#5512) by @Yukun-Cui
- [KernelGen][Nvidia] Add _conj_copy operator with Triton kernel (#5513) by @Yukun-Cui
- [KernelGen][Nvidia] Add coalesced operator with Triton kernel (#5514) by @Yukun-Cui
- [KernelGen][Nvidia] Add _compute_linear_combination operator with Triton kernel (#5515) by @Yukun-Cui
- [KernelGen][Nvidia] Add flipud operator with Triton Kernel (#5516) by @xuanzhengdu-eng
- [KernelGen][Nvidia] Add float_power operator with Triton Kernel (#5517) by @xuanzhengdu-eng
- [KMCompiler][Nvidia][Ascend]Add operator linalg_lu with triton kernel. (#5518) by @YangLong114514
- [KernelGen][Nvidia] Add alpha_dropout_ operator with Triton kernel (#5519) by @LoserCheems
- [KernelGen][Nvidia] Add cosine_embedding_loss operator with Triton kernel (#5521) by @ShawnsYing
- [C++ Runtime MLU]Add Cambricon MLU support to C++ operators (#5535) by @chen98038
- [KernelGen][Nvidia] Add column_stack operator with Triton kernel (#5536) by @NineAnnAnn
- [KernelGen][Nvidia] Add choose_qparams_optimized operator with Triton… (#5537) by @NineAnnAnn
- [KernelGen][Nvidia] Add _cummin_helper operator with Triton kernel (#5538) by @chx7514
- [KernelGen][Nvidia] Add _fake_quantize_learnable_per_tensor_affine_backward operator with Triton kernel (#5540) by @chx7514
- [KernelGen][Nvidia] Add _fake_quantize_learnable_per_channel_affine_backward operator with Triton kernel (#5543) by @chx7514
- [KernelGen][Nvidia] Add _fake_quantize_learnable_per_tensor_affine operator with Triton kernel (#5544) by @chx7514
- [KernelGen][Nvidia] Add adaptive_avg_pool1d operator with Triton kernel (#5550) by @CoyeCHEN
- [KernelGen][Nvidia] Add fill_mem_eff_dropout_mask operator with Triton kernel (#5552) by @chx7514
- [KernelGen][Nvidia] Add cov operator with Triton kernel (#5554) by @ShawnsYing
- [KernelGen][MThreads] Add 20 operators for MThreads (#5556) by @CoyeCHEN
- [KernelGen][Hygon] Add 25 operators for Hygon (#5557) by @CoyeCHEN
- [KernelGen][Iluvatar] Add 28 operators for Iluvatar (#5558) by @CoyeCHEN
- [KernelGen][MetaX] Add 26 operators for MetaX (#5559) by @CoyeCHEN
- [KernelGen][T-Head] Add 38 operators for T-Head (#5560) by @CoyeCHEN
- [KernelGen][Nvidia] Add cummaxmin_backward operator with Triton kernel (#5566) by @ShawnsYing
- [KernelGen][Nvidia] Add special_softmax operator with Triton kernel (#5578) by @bwbwzzz
- [KMCompiler][NVIDIA][Hygon][Thead][iluvatar] Add exponential operator with Triton kernel (#5579) by @littlekid-zxt
- [KernelGen][Nvidia] Add unsafe_split_with_sizes operator (#5580) by @chx7514
- [KMCompile][Nvidia][Ascend][Metax][Iluvatar] Add linalg.qr operator with triton kenrel (#5586) by @Caeruleann
- [KernelGen][Nvidia] Add linalg_eig operator with Triton kernel (#5588) by @bwbwzzz
- [KernelGen][Nvidia] Add split_with_sizes operator (#5590) by @chx7514
- [KMCompiler][Nvidia][Hygon] Add linalg_matrix_norm operator (#5595) by @wuhansky
- [KernelGen][Nvidia] Add embedding_renorm_ operator with Triton kernel (#5602) by @ShawnsYing
- [KernelGen][Nvidia] Add conv_tbc_backward operator with Triton kernel (#5605) by @NineAnnAnn
- [KernelGen][Nvidia] Add conv_transpose3d operator with Triton kernel (#5606) by @NineAnnAnn
- [KernelGen][Nvidia] Add cumulative_trapezoid operator with Triton kernel (#5618) by @ShawnsYing
- [KernelGen][Nvidia] Add _fake_quantize_per_tensor_affine_cachemask_tensor_qparams operator with Triton kernel (#5631) by @chx7514
- [KernelGen][Nvidia] Add _cummax_helper operator with Triton kernel (#5632) by @chx7514
- [KernelGen][Nvidia] Add _dirichlet_grad operator with Triton kernel (#5633) by @chx7514
- [KernelGen][Nvidia] Add convolution_overrideable operator with Triton kernel (#5642) by @NineAnnAnn
- [KMCompiler][Nvidia][Ascend] Add operator index_fill/index_fill_ with triton. (#5643) by @YangLong114514
- [KMCompiler][Iluvatar][Hygon] Add linalg_matrix_norm operator with triton kernel. (#5648) by @wuhansky
- [KernelGen][Nvidia] Add _upsample_lanczos2d_aa operator with Triton Kernel (#5651) by @xuanzhengdu-eng
- [KernelGen][Nvidia] Add _upsample_nearest_exact1d_backward operator with Triton Kernel (#5653) by @xuanzhengdu-eng
- [KernelGen][Nvidia] Add nll_loss2d operator with Triton kernel (#5660) by @ShawnsYing
- [KernelGen][Nvidia] Add logit_backward operator with Triton kernel (#5661) by @kkkwb
- [KernelGen][Nvidia] Add nuclear_norm operator (#5662) by @ShawnsYing
- [Iluvatar] Add dedicated kernel for repeat_interleave_self_int (#5665) by @awayzjj
- [KMCompiler][Nvidia][Ascend] Add adaptive_max_pool3d operator (#5671) by @wuhansky
- [KernelGen][Nvidia] Add _batch_norm_with_update_functional operator with Triton Kernel (#5673) by @kkkwb
- [KMCompiler][Ascend] add exponential operator with Triton kernel (#5682) by @littlekid-zxt
- [KernelGen][Nvidia] Add _nested_from_padded_tensor operator with Triton kernel (#5688) by @chx7514
- [KernelGen][Nvidia] Add _nested_sum_backward operator with Triton kernel (#5690) by @chx7514
- [KernelGen][Nvidia] Add _nested_tensor_from_mask_left_aligned operator with Triton kernel (#5692) by @chx7514
- [KernelGen][Nvidia] Add _nested_view_from_jagged operator (#5694) by @chx7514
- [KernelGen][Nvidia] Add _nested_view_from_jagged_copy operator with Triton kernel (#5695) by @chx7514
- [KMCompiler]Add stages to the linalg_lstsq operator entry (#5700) by @cheersluvs
- [KernelGen] Add max_pool1d operator (#5704) by @NineAnnAnn
- Add 'decorator' to the dependencies for Ascend CANN 8.5.0 backend. (#5707) by @wuhansky
- Add attrs to the dependency list for Ascend CANN 8.5.0 (#5708) by @wuhansky
- [KernelGen][Nvidia] Add matrix_exp_backward operator with Triton kernel (#5710) by @NineAnnAnn
- Add psutil to dependency for CANN 8.5.0 (#5712) by @wuhansky
- [KernelGen][NVIDIA] Add linalg_matrix_sqrth operator with Triton kernel (#5728) by @xuanzhengdu-eng
- [KernelGen][NVIDIA] Add _thnn_differentiable_gru_cell_backward operator with Triton kernel (#5734) by @xuanzhengdu-eng
- [KernelGen][Nvidia] Add sum_to_size operator with Triton kernel (#5761) by @NineAnnAnn
- [FlagTune] Add multi-platform Mul cost model support (#5762) by @k6699k
- [SiliconFlow] Add CI test aliases for approved and merged operators (#5788) by @winfan1314
- Add conj_physical_ operator with contiguous-view Triton kernel (#5800) by @tspyc072
- [KMCompiler][Ascend] Add ascend backend for linalg_solve_triangular (#5821) by @LittleShun1214
- feat(tools): add quick mode (#5866) by @103yiran
- [KMCompiler][Nvidia]Add operator linalg_norm with triton. (#5870) by @YangLong114514
- add int8 tests for (un)pack_seq_triton (#5886) by @douxetpur
- [ARM] Add Triton W4A8-G128 linear API for CPU (#5904) by @kevinzs2048
- feat: enable and enforce _FULL_CONFIG sorting (#5908) by @gavin0x01
- [KMCompiler][NVIDIA] Add matrix rank operator (#5942) by @YangLong114514
- feat(setup.sh): install python-build-standalone by mirror (#5948) by @103yiran
- [KMCompiler][Nvidia][Thead]Add linalg_matrix_power for triton (#5951) by @wuhansky
- [QC][PPU] Add W8A16 RMSNorm operator (#5957) by @July-h5kf3
- [KMCompiler][Ascend][Metax]Add linalg_matrix_power for triton (#5960) by @wuhansky
- [k] add and fix fused kernels (fused) (#5965) by @RobertLuobo
- [QC][PPU] Add W8A16 FP8 TopK operator (#5966) by @July-h5kf3
- [k] add and fix math category ops (#5967) by @RobertLuobo
- [QC][Ascend] Add INT8 W8A16 RMSNorm operator (#5968) by @July-h5kf3
- [KMCompiler][Nvidia] Add gru operator with Triton kernel (#6002) by @Willing2329
- [KMCompiler][MetaX] Add gru operator with Triton kernel (#6059) by @Willing2329
- [KMCompiler][NVIDIA][Ascend][Hygon][Thead][MetaX][iluvatar] add linalg_matrix_exp operator with Triton kernel (#6069) by @littlekid-zxt
- feat(hash_tensor): add Triton implementation of aten::hash_tensor (#6088) by @ShawnsYing
- [KMCompiler][Iluvatar] Add gru operator with Triton kernel (#6094) by @Willing2329
- [KMCompiler][Ascend] Add ascend backend for igammac (#6136) by @wuhansky
- [KMCompiler][Ascend] Add ascend matrix_rank implement and fix some bug for general matrix_rank ops. (#6161) by @YangLong114514
- [KernelGen][MThreads] Add fmod_ Moore Threads specialized operator (#6170) by @Yukun-Cui
- [KernelGen][MThreads] Add conv_transpose1d Moore Threads specialized operator (#6174) by @Yukun-Cui
- [KernelGen][MThreads] Add upsample_linear1d_backward Moore Threads specialized operator (#6175) by @Yukun-Cui
- [KernelGen][MThreads] Add matmuladd Moore Threads specialized operator (#6180) by @Yukun-Cui
- [QC][Hygon] Add topk_w8a16_fp8 (#6190) by @July-h5kf3
- [KMCompiler][Ascend] Add gru operator with Triton kernel (#6201) by @Willing2329
- [KMCompiler][Ascend] Add depends (#6205) by @Willing2329
- [FlagTune] Add common MUL Cost Model support (#6232) (#6345) by @0x45f
- [Ascend] Add specialized linear operator (#6455) by @modao1234
- Add truncate option to command.yaml (
2be4a2918ad1; commit) - Add permissions for test-operator job (
2228de61dcd4; commit) - Add h100 to supported runner labels in command.yaml (
d9d17349db5d; commit) - Add RUNNER_SSH_KEY secret to unittest workflow (
769fe5fc8f71; commit) - Add workflow_dispatch trigger to cpp-op-test.yaml (
99ef915e953d; commit) - Add concurrency control to unit test workflow (
af9e34e433b7; commit)
Performance
- [Iluvatar] optimize mm (#1988) by @awayzjj
- Optimize flash varlen paged KV cache addressing (#3564) by @yysheng26
- [KMCompiler]Perf/cp gather indexer k quant cache (#3686) by @chenkc2025
- [KMCompiler]Perf/indexer k quant and cache (#3688) by @chenkc2025
- [QC] Optimize fused_marlin_moe (w4a16) (#3778) by @monellz
- [QC]Optimize GEMM(fp8) (#3821) by @ggbondbest
- [SiliconFlow] Optimize upsample_bicubic2d_aa_backward on Ascend NPU and Fix Tests (#4101) by @winfan1314
- [KMCompiler] Optimize compute_global_topk_indices_and_lens for DeepSeek V4 attention (#4113) by @LittleShun1214
- [QC] optimize mxfp4 w4a16 fused_moe (#4117) by @monellz
- [KMCompiler]Optimize pack_seq (#4125) by @chenkc2025
- [KMCompiler]Optimize unpack_seq (#4126) by @chenkc2025
- Optimize command workflow (#4161) by @tengqm
- [KMCompiler]Optimize dequantize_and_gather_k_cache for DeepSeek V4 attention (#4267) by @littlekid-zxt
- [KMCompiler] Optimize combine_topk_swa_indices for DeepSeek V4 attention (#4282) by @cheersluvs
- Optimize std dim reduction (#4304) by @douxetpur
- [KMCompiler]Optimize per-token group FP8 quantization (#4319) by @chenkc2025
- Optimize run_tests.py script for marks collection (#4374) by @tengqm
- [QC]Optimize fp8-w8a16 Rmsnorm (#4437) by @ggbondbest
- optimize upsample_trilinear3d. (#4438) by @awayzjj
- [Performance Optimization] Optimize split_with_sizes_copy launch overhead (#4456) by @stringl1l1l1l
- [MTHREADS] Optimize one_hot (#4468) by @Oslomayor
- Optimize MXFP4 fused Marlin MoE kernels (#4483) by @KingsleyLiu-NV
- Optimize poisson sampling and fix new_full fill codegen (#4508) by @Salamanca001
- [Iluvatar] Optimize linear with EVEN_K SME-friendly kernel (#4514) by @Salamanca001
- [Iluvatar] optimize conv_transpose1d. (#4515) by @awayzjj
- Optimize index for slice-before-index adjacent tensor indices (#4579) by @ZYX223
- Optimize Hopper mul broadcast 2D tuning key (#4608) by @liuxin2560
- Optimize prod single-dim reduction with inner/non-inner kernels (#4618) by @zheng1
- Optimize Metax mul broadcast 2D tuning key (#4662) by @Tiannn967
- [HCU] Optimize backend kernel operators for improved performance (#4663) by @wangxshuai
41 more changes
- Optimize index_add for contiguous suffix layouts (#4670) by @ZYX223
- Optimize w8a8_block_fp8_bmm (#4794) by @Jennie1004
- [kunlunxin] optimize existing ops and add new op coverage (#4816) by @RobertLuobo
- [QC]Optimize FP8 W8A8 bmm (#4880) by @ggbondbest
- [Ascend] Optimize layer norm forward row scheduling (#4903) by @ZYX223
- Optimize MXFP4 fused Marlin MoE dequant: fold scale + decode-opt (#4908) by @zeroherolin
- Optimize MXFP4 Dequantization in Fused Marlin MoE (#4943) by @KingsleyLiu-NV
- [QC] Optimize flash_mla_with_kvcache for model1 (#5010) by @monellz
- Optimize wna16 MoE main loop: one mma per K tile instead of eight (#5140) by @zeroherolin
- [MTHREADS] Optimize flip for multi-dim tensors (#5261) by @Oslomayor
- [SiliconFlow] Optimize cauchy on Iluvatar and fix cross-backend tests (#5341) by @winfan1314
- [MTHREADS] optimize conv1d_padding autotune configs (#5343) by @Oslomayor
- [MTHREADS] optimize conv3d_padding autotune configs (#5353) by @Oslomayor
- [Ascend] Optimize AddMM layouts and bias epilogue (#5383) by @ZYX223
- [MThreads] Optimize AddMM layouts and bias epilogue (#5384) by @ZYX223
- [MetaX] Optimize AddMM layouts and bias epilogue (#5386) by @ZYX223
- [MTHREADS] Optimize constant_pad_nd: add mthreads-specific kernels (#5395) by @Oslomayor
- Optimize MM kernels and autotuning (#5407) by @Tiannn967
- [KMCompiler]Optimize index_copy and index_copy_ generic kernels (#5496) by @Onisen7
- [MTHREADS] optimize matmul_bias_activation with SQMMA (#5551) by @Oslomayor
- Optimize SiLU for Ascend and MetaX (#5565) by @YUANXIGUREN
- [MTHREADS] optimize addmm_out with mthreads SQMMA/FMA override (#5604) by @Oslomayor
- [Iluvatar] optimize matmul_bias_activation (#5656) by @awayzjj
- [MTHREADS] optimize baddbmm with SQMMA (#5666) by @Oslomayor
- [MTHREADS] optimize pad with mthreads constant_pad_nd override (#5675) by @Oslomayor
- [MTHREADS] Optimize constant_pad_nd FillCopy path (#5684) by @shanyulu
- [KMCompiler]Optimize log_sigmoid_backward Ascend performance (#5718) by @elllzat
- [MTHREADS] Optimize feature_dropout_ with fused in-place path (#5742) by @gao0624
- [KMCompiler] Optimization for operator linalg_solve_triangular. (#5753) by @YangLong114514
- [MTHREADS] optimize upsample_linear1d_backward (#5779) by @Oslomayor
- Optimize Hopper MM dispatch and tuning (#5780) by @Tiannn967
- [MTHREADS] optimize matmuladd by reusing addmm SQMMA/FMA kernels (#5781) by @Oslomayor
- [MTHREADS] Optimize conv_transpose1d with stride-phase gather (#5789) by @gao0624
- [MTHREADS] optimize channel_shuffle with block-copy kernel (#5814) by @Oslomayor
- [MTHREADS] optimize one_hot with fused validation and scatter (#5833) by @gao0624
- [MTHREADS] optimize cudnn_convolution with no-bias conv2d path (#5834) by @gao0624
- [MTHREADS] optimize conj performance (#5860) by @Oslomayor
- [MTHREADS] optimize linear performance (#5863) by @Oslomayor
- Optimize GLU kernel with TLE autotuning (#5871) by @lzllx123
- [MTHREADS] optimize fp32 conv_transpose2d k3 (#5879) by @gao0624
- [MTHREADS] Optimize isin Tensor_Scalar (#5977) by @gao0624
Hardware Support
- [MTHREADS] restore pre_hook and libtuner for SQMMA (#3364) by @Oslomayor
- [KernelGen][Nvidia] Implement renorm_ operator test functions (#3536) by @XDYuanzhuLee
- Enflame to flagos 20260529 (#3602) by @chongzhouyang
- [ARM]: add ARM64 CPU backend (NEON/SVE2) with INT8 quant and fuse… (#3616) by @kevinzs2048
- [KernelGen][Nvidia] Migrate logical_xor_ inplace operator from experimental to ops (#3745) by @XDYuanzhuLee
- [KernelGen][Nvidia] Migrate deg2rad, fix, negative from experimental to ops (#3764) by @XDYuanzhuLee
- [KernelGen][Nvidia] Migrate addcmul_ operator from experimental to ops (#3766) by @XDYuanzhuLee
- [KernelGen][Nvidia] Migrate atanh_operator from experimental to ops (#3768) by @XDYuanzhuLee
- [KUNLUNXIN] all_dim (#3781) by @sh1653487844
- [KernelGen][MetaX] Upgrade sparse_attention with multi-tier dispatch (#3998) by @LoserCheems
- [Mthreads] Optimized baddbmm & w8a8_block_fp8_matmul with FlagOSTune (#4095) by @Fattysand
- [metax] update masked_fill_ op (#4170) by @Alvin-YCHEN
- [Backend] enable_nvidia_unused_ops (#4285) by @Galaxy1458
- [metax] update layernorm backwar kernel (#4431) by @Alvin-YCHEN
- cambricon: update and fix kernel (#4435) by @chenmiao1919
- [Iluvatar] Opt tile repeat (#4458) by @awayzjj
- [Iluvatar] Opt Iluvatar conv_depthwise2d. (#4513) by @huatuoli
- gpu_check: report available GPUs instead of requiring all free (#4528) by @tengqm
- [Kunlunxin] Register cumprod / cumprod_ in init (#4567) by @llaboon
- [KMCompiler][Nvidia] opt for log_normal_ operator (#4621) by @littlekid-zxt
- [ENFLAME]update 20260709 (#4646) by @chongzhouyang
- [ENFLAME] update_ops_20260713 (#4793) by @chongzhouyang
- [ENFLAME]workround for gcu300 to support sgl (#4814) by @chongzhouyang
- [KernelGen][Nvidia] Move arctanh operator to ops (#4818) by @NineAnnAnn
- [KMCompiler][Ascend] nonzero static ascend (#4921) by @chenkc2025
- [KernelGen][Nvidia] Move sgn operator to ops (#4990) by @YufanPeter
- [KernelGen][Nvidia] Move sign operator to ops (#5030) by @YufanPeter
- [ENFLAME]enflame_to_FlagOS_20260729 (#5044) by @chongzhouyang
- [KernelGen][Nvidia] Move log_ operator to ops (#5045) by @YufanPeter
- [KernelGen][Nvidia] Move take operator to ops (#5070) by @YufanPeter
25 more changes
- [KernelGen][Nvidia] Move hardtanh_ operator to ops (#5115) by @YufanPeter
- [KernelGen][Nvidia] Move hardswish operator to ops (#5139) by @YufanPeter
- [KernelGen][Nvidia] Move hardsigmoid operator to ops (#5149) by @YufanPeter
- [KernelGen][Nvidia] Move huber_loss operator to ops (#5154) by @xuanzhengdu-eng
- [KernelGen][Nvidia] Move hardshrink operator to ops (#5155) by @xuanzhengdu-eng
- [KernelGen][Nvidia] Move heaviside_ operator to ops (#5156) by @xuanzhengdu-eng
- [KernelGen][Nvidia] Move hypot_ operator to ops (#5157) by @xuanzhengdu-eng
- [KernelGen][Nvidia] Move fix_ operator to ops (#5158) by @xuanzhengdu-eng
- [KernelGen][Nvidia] Move arccosh operator to ops (#5163) by @YufanPeter
- [KernelGen][Nvidia] Move absolute_ operator to ops (#5187) by @YufanPeter
- Metax fix shared ops (#5334) by @CLYtou
- Metax c550 operator fixes (#5337) by @CLYtou
- [KernelGen][Nvidia] Remap native_dropout_backward to canonical wrapper (#5414) by @CoyeCHEN
- [Ascend] Update flagtree ascend whl version to 0.6.1 in backends.yaml (#5433) by @zhzhcookie
- [ENFLAME] enflame_to_flagos_20260818 (#5582) by @chongzhouyang
- [KernelGen][Nvidia] Move hardtanh operator to ops (#5629) by @kkkwb
- [KMCompiler][ASCEND] Stabilize Ascend cholesky solve layout test (#5698) by @YangLong114514
- [KMCompiler][Nvidia] Refine dist operator impl and add full test coverage (#5958) by @Caeruleann
- [KMCompiler][Ascend] Opt unsafe_index operator on Ascend. (#6007) by @Caeruleann
- [KMCompiler][Ascend] Opt dist operator on Ascend. (#6014) by @Caeruleann
- [KMCompiler][NVIDIA] Matrix rank test fixes. (#6076) by @YangLong114514
- [kunlunxin] move the copy family onto tle.gpu (#6093) by @Ason93
- [KMCompiler][Ascend] Opt replication pad2d backward (#6138) by @wuhansky
- [KMCompiler][Metax] operator linalg_norm ord==0.5 scalar fix (#6162) by @YangLong114514
- enflame to flagos 20260911 (#6200) by @chongzhouyang
Improvements
- refactor type as Type to support old version Python (#1136) by @sgjzfzzf
- Improve addmv_out test coverage (#3078) by @lukatao
- Improve all_dims test coverage (#3086) by @lukatao
- Improve baddbmm.out test coverage (#3099) by @lukatao
- Improve cat_out test coverage (#3108) by @lukatao
- Improve chunk_gated_delta_rule_fwd test coverage (#3109) by @lukatao
- Improve sparse_mla_fwd_interface test coverage (#3201) by @lukatao
- Improve atan2 test coverage (#3278) by @lukatao
- Improve pixel test coverage (#3344) by @lukatao
- Improve run_tests script for progress bar update (#3754) by @tengqm
- Improve run_tests script for progress bar update (#3757) by @tengqm
- [Performance Optimization] Improve any_dims parallel reduction (#4130) by @stringl1l1l1l
- Improve the setup.sh script (#4168) by @tengqm
- Improve mark collection logic (#4169) by @tengqm
- Refactor workflows for a bootstrap checkout (#4174) by @tengqm
- Improve W4A16 fused Marlin MoE performance (#4350) by @KingsleyLiu-NV
- Improve setup logic (#4403) by @tengqm
- Improve support to C++ operators (#4404) by @tengqm
- Improve setup procedure for flagtree (#4459) by @tengqm
- [Performance Optimization] Improve reflection_pad1d_backward (#4510) by @stringl1l1l1l
- [Performance Optimization] Improve log_softmax_out non-inner reduction (#4511) by @stringl1l1l1l
- [MTHREADS] improve bmm SQMMA tuning (#5909) by @Oslomayor
- Refactor iluvatar installation commands in setup_vendor.sh (
e7aadf568068; commit) - Refactor test-operator step in command.yaml (
34626626550c; commit)
Dependencies
- Bump torch version to 2.9.0 on Kunlunxin backends (#3773) by @tengqm
- Bump metax versions (#3842) by @tengqm
- Bump package versions for tsingmicro backend (#3846) by @tengqm
- Bump enflame package dependencies (#3853) by @tengqm
- Bump the actions-minor group with 2 updates (#4489) by @dependabot
- Bump actions/labeler from 6.1.0 to 6.2.0 in the actions-minor group (#4634) by @dependabot
- chore(deps): bump actions/setup-node from 6.4.0 to 7.0.0 (#4845) by @dependabot
- chore(deps): bump actions/checkout from 6.0.2 to 7.0.0 (#4846) by @dependabot
- chore(deps): bump actions/setup-go from 6.5.0 to 7.0.0 (#4847) by @dependabot
- Bump numpy version to 2.3.5 (#4994) by @tengqm
- Bump wheel version for CVE (#5271) by @tengqm
- Bump flagtree version for NVIDIA (#5444) by @tengqm
- Bump flagtree version on enflame 1.10.6 backend (#5649) by @tengqm
- Bump python version for iluvatar (#5765) by @tengqm
Documentation
- Auto-generate version from git tags using setuptools-scm (#4497) by @tengqm
- docs: update installation guide for setup.sh workflow (#4647) by @tengqm
- docs: update Chinese installation guide for setup.sh workflow (#4648) by @tengqm
- docs: update install/cpp/packaging guides for cpp/ wheel split (#4868) by @tengqm
- chore: add Apache 2.0 copyright header to all source files (#4873) by @tengqm
- docs(release): place dev tag on empty commit after release, not on release tag (#4912) by @tengqm
- Ship flaggems-setup as a top-level module and auto-install a compiler (#5207) by @tengqm
- Contribution doc update (#5307) by @103yiran
- Split multi-version vendors into versioned backend keys (#5350) by @tengqm
- docs: drop Apache 2.0 headers added to third-party files (#6156) by @tengqm
- docs(fused_marlin_moe): document weight layout is not vLLM Marlin layout (#6204) by @zeroherolin
CI/Infrastructure
- ci(setup.sh): specify source when installing uv (#4106) by @103yiran
- build(deps): bump the actions-minor group across 1 directory with 3 updates (#4356) by @dependabot
- build(deps): bump actions/cache from 5.0.5 to 6.0.0 (#4357) by @dependabot
- build(deps): bump actions/checkout from 6.0.2 to 7.0.0 (#4359) by @dependabot
- ci: remove stale image-builder workflow (#4872) by @tengqm
- ci: run weekly tests daily and use a per-vendor job-level timeout (#4887) by @wuwentao
- ci: update schedule time and timeout value (#4889) by @wuwentao
- [CI]: update weekly test output path and mthreads/hygon container (#4946) by @wuwentao
- CI: improve weekly test gpu id and remove cache file (#4983) by @wuwentao
- build(deps): bump the actions-minor group across 1 directory with 3 updates (#5071) by @dependabot
- [CI]: Update op-monitor url and schedule time (#5546) by @wuwentao
- ci: add rule-check workflow for PR validation (#5774) by @gavin0x01
- Ci/rule check phase1 2 (#5787) by @gavin0x01
- ci: add rule-check results reporting to Feishu Bitable (#5827) by @gavin0x01
- Ci/rule check phase1 2 (#5853) by @gavin0x01
- ci: add rule-check results reporting to Feishu Bitable (#5854) by @gavin0x01
- Ci/rule check phase1 2 (#5868) by @gavin0x01
- build: tolerate post-release tags in version derivation (#5885) by @wbavon
- Ci/rule check phase1 2 (#5898) by @gavin0x01
- ci(weekly): add workflow dispatch controls (#5919) by @wuwentao
- [CI] Add rule check to forbid use_gems() in PR-changed test files (#5963) by @gavin0x01
- cicd(.github): remove cann850 (#5970) by @103yiran
- cicd(.github): remove metax-maca3720 (#6090) by @103yiran
- [CI] Use three-dot diff to derive PR-changed files (#6097) by @gavin0x01
- [CI/CD] Add rule-check-required job and cicd docs (#6119) by @103yiran
- [Hygon] cicd(backends.yaml): update hygon compiler version (#6148) by @103yiran
- Ci/extend init exports check (#6157) by @gavin0x01
- [packaging] Port #6365 to 5.4.0-rc2 (Debian + RPM packaging) (#6394) by @shiptux
Testing
- [Benchmark][Ascend] Update benchmark do_bench to do_bench_npu on ascend (#4857) by @zhzhcookie
- TEST CASE FIX -- iluvatar (#5131) by @jia-heng
- test(and_scalar): tag bitwise_and_scalar tests so pytest -m and_scalar collects (#6033) by @bin913
- test: rename batch_norm_with_update_functional benchmark file to sing… (#6052) by @freaking-man-in
- [QC][Benchmark] Use dedicated test_mm_w8a8_fp8 entry (#6084) by @ggbondbest
- [Benchmark] Fix chunk gated delta rule forward baseline (#6177) by @lukatao
- [Benchmark] Add standalone FP8 MM performance test with vLLM baseline (#6183) by @ggbondbest
Reverts
- Revert arm vendor cmd (#3820) by @0x45f
- Revert iluvatar LIBDIR hack and disable flagtree for iluvatar (#4420) by @tengqm
- Revert "[FlagTune] Add multi-platform Mul cost model support" (#6043) by @Vincent-Xiao
Bug Fixes
- fix(grid_sample): enable large output sizes (#3421) by @mygitljf
- Fix scatter_reduce_two kernel (#3458) by @tengqm
- [HYGON] Fix pow, mul, per_token_group_quant_fp8, allclose, isclose and reshape_and_cache (#3476) by @fangzexian
- [MTHREADS] fix assert_async (#3504) by @Oslomayor
- [KUNLUNXIN] fix cumprod_ (#3562) by @sh1653487844
- Fix patch vllm util (#3657) by @0x45f
- [KUNLUNXIN] fix vender problem (#3660) by @nianqi-tian
- Fix environment setting for iluvatar backend (#3667) by @tengqm
- Fix parameters used for benchmarking kunlunxin (#3690) by @tengqm
- [KUNLUNXIN] fix cumprod_ boolean tensor in-place operation (#3747) by @sh1653487844
- [KernelGen] Fix --ref cpu cross-device tensor comparison (#3750) by @lukatao
- [KernelGen] Fix floor_divide for mixed int/float types and fp16/bf16 dtypes (#3751) by @lukatao
- [KernelGen] Fix grouped_mm accuracy test and benchmark configuration (#3752) by @lukatao
- [Advanced Compiler] fix dgeglu geglu dreglu reglu skipped test caused by invalid zero-dimension test shape (#3755) by @AdvancedCompiler
- [KernelGen] Fix sparse_mla TLE kernel group/head-block index decomposition (#3759) by @lukatao
- Fix cpp test (#3761) by @tengqm
- [KernelGen] Fix addmm_dtype and addmm_dtype_out accuracy tests for --ref cpu mode (#3762) by @lukatao
- [Backend ] Fix_replace_op_bugs (#3785) by @Galaxy1458
- Fix github CI for script rename (#3795) by @tengqm
- [ENFLAME] fix ops of gcu300 some bugs to enflame backend (#3814) by @chaaa-a
- fix issue 3733 (#3815) by @huangyiqun
- [KernelGen] Fix count_nonzero: replace deprecated logical operators and add sparse/negative dim support (#3816) by @yzw1128
- fix svd & pointwise_dynamic (#3824) by @Dayrker
- [Sunrise backend] fix scatter_reduce & mse_loss (#3829) by @Dayrker
- fix issues:1012 (#3831) by @w1120029931-bit
- Fix ix containerfile (#3841) by @tengqm
- Fix env tsingmicro (#3844) by @tengqm
- [Advanced Compiler] Fix conv_depthwise2d test with --ref cpu (#3859) by @AdvancedCompiler
- [Backend] Fix replace op bugs (#3863) by @Galaxy1458
- [Sunrise] fix fused_recurrent (#3878) by @Dayrker
231 more changes
- Fix package dependency on klx (#3934) by @tengqm
- Fix operator inventory (#3937) by @tengqm
- [KUNLUNXIN] fix weight_norm, aminmax,l eaky_relu, mm, bmm, copysign, signbit, logsumexp,nonzero_numpy and tril (#3953) by @RobertLuobo
- [KernelGen] Fix baddbmm.out: missing_test (#3958) by @lukatao
- [KernelGen] Fix new_full: accuracy_fail (#3962) by @lukatao
- fix issue 3007 (#3972) by @huangyiqun
- fix issue 3585 (#3973) by @huangyiqun
- Fix metadata for the tan_ operator (#3983) by @XDYuanzhuLee
- Fix TsingMicro CI (#4010) by @tengqm
- Fix GPU check script for Ascend (#4045) by @tengqm
- Fix on-demand workflow for GPU check (#4047) by @tengqm
- Fix operator inventory (#4048) by @tengqm
- Fix incomplete operator entry (#4049) by @tengqm
- fix issue 2896 (#4051) by @huangyiqun
- fix issue 2829 (#4052) by @huangyiqun
- fix issue 2850 (#4053) by @huangyiqun
- Fix Kunlunxin setup (#4082) by @tengqm
- [KernelGen][Iluvatar] Fix all_dim accuracy failure (float16/bfloat16 input) (#4093) by @lukatao
- [Ascend] fix: fill_ ZeroDivisionError on empty tensor (#4096) by @JosephNew
- Fix any_dim accuracy failure on Iluvatar BI-V150 (float16 input) (#4098) by @lukatao
- Fix var accuracy failure on Iluvatar BI-V150 (Triton specialization bug) (#4099) by @lukatao
- Fix thead setup for CI (#4103) by @tengqm
- [SiliconFlow] Fix nanmedian cross-backend accuracy (#4116) by @winfan1314
- [KernelGen][Iluvatar] Fix var_correction: accuracy_fail (#4118) by @lukatao
- [SiliconFlow] Fix cumprod accuracy and inplace integer path (#4128) by @winfan1314
- [enflame]Fix VendorDescriptor bug (#4134) by @Galaxy1458
- Fix containerfile for Hygon (#4138) by @tengqm
- Fix container file for KLX (#4139) by @tengqm
- Fix PyYAML version in pyproject.toml (#4143) by @tengqm
- Fix the vllm installation cache problem (#4164) by @tengqm
- Fix ascend backend (#4165) by @tengqm
- Fix renorm fp16 precision: allocate norms buffer in fp32 (#4237) by @tengqm
- Fix median/nanmedian test crash on int16 in quick mode (#4239) by @tengqm
- [SiliconFlow] Fix mode operator on Mthreads backend (#4248) by @winfan1314
- [Mthreads] Fix heuristic fallback for log_softmax (#4258) by @mygitljf
- Fix hardsigmoid_out benchmark (#4271) by @huangyiqun
- [AdvancedCompiler] fix: avoid int32 overflow for large tensors in repeat and tile kernel (#4289) by @AdvancedCompiler
- [KernelGen] Fix conv2d_padding: accuracy_fail (#4298) by @lukatao
- [KernelGen] Fix multinomial: accuracy_fail (#4303) by @lukatao
- fix mm tma (#4308) by @huangyiqun
- Fix operator order in registry (#4313) by @XDYuanzhuLee
- [AdvancedCompiler] fix: max_pool3d_backward test (#4328) by @AdvancedCompiler
- Fix unbound variable errors when sourcing vendor env scripts (#4344) by @tengqm
- Fix uv install failure when HOME is not writable (#4345) by @tengqm
- Fix gh install failure in command.yaml when HOME is not writable (#4346) by @tengqm
- [MTHREADS] fix soft_margin_loss performance with two-phase reduction (#4364) by @Oslomayor
- Fix CMakeLists.txt for git cloning libtriton_jit (#4375) by @tengqm
- Fix containerfile for NVIDIA 13.3 backend (#4378) by @tengqm
- Fix weight_norm_interface norm dtype truncation (#4382) by @tengqm
- Fix nanmedian kthvalue fallback for NaN-containing inputs (#4384) by @tengqm
- Fix some nits in the operator source code (#4409) by @tengqm
- Fix: defer yaml parsing in setup.sh to after venv creation (#4415) by @tengqm
- Fix check-gpu step and failure-report duplicates in command.yaml (#4416) by @tengqm
- Fix Iluvatar triton plugin loading in venv (#4417) by @tengqm
- [MTHREADS] fix log10 perf (#4426) by @Oslomayor
- Fix error when dim is empty tuple (#4428) by @0x45f
- fix(sparse_attention): pad H dimension to satisfy tl.dot minimum size (#4447) by @Yangyang0906C
- [MTHREADS] fix tile size config for unary elementwise ops (#4453) by @Oslomayor
- [TSINGMICRO] fix bugs of test cases (#4463) by @tsingmicro-public-e
- fix hcu import error (#4464) by @huangyiqun
- [Ascend] fix argmax crash: remove duplicate @libentry() on argmax_kernel (#4467) by @JosephNew
- fix(broadcast_to): use pin_memory to avoid illegal CPU->CUDA copy in CUDA graph capture (#4472) by @physics31415926
- fix(container): remove garbled characters (#4486) by @103yiran
- fix mark in test (#4492) by @103yiran
- Fix setup.sh for setuptools-scm with --no-build-isolation (#4498) by @tengqm
- fix(conf): remove div_tensor_mode, div_scalar_mode (#4509) by @103yiran
- Fix the dispatch keys for some operators (#4522) by @tengqm
- Fix psum_text comparison when before has no performance data (#4534) by @tengqm
- fix: safe vllm install to protect torch/triton/flagtree (#4535) by @tengqm
- Fix cumsum ZeroDivisionError on empty tensor (#4541) by @DannyP0
- Fix randn error in multi card case (#4560) by @0x45f
- Fix uv not found in subsequent CI steps (#4562) by @tengqm
- Fix psum_text IndexError when before data is empty (#4569) by @tengqm
- Fix code style error (#4585) by @0x45f
- fix: affine_grid_generator float32 precision mismatch with PyTorch CUDA reference (#4586) by @yzw1128
- fix(ascend): mask bmm boundary loads (#4631) by @yilongma0110
- fix transformer engine import error (#4642) by @huangyiqun
- Fix uv detection by adding ~/.local/bin to PATH before check (#4671) by @tengqm
- Fix setup.sh and env.sh for set -u compatibility (#4672) by @tengqm
- Fix ascend-cann900 post_install format in backends.yaml (#4674) by @tengqm
- Fix post_install: use full shell commands with eval (#4675) by @tengqm
- Fix Ascend env sourcing: unify path and add error tolerance (#4678) by @tengqm
- Fix negative op on non-CUDA backends (e.g. MUSA) (#4682) by @tengqm
- [KernelGen][Mthreads] Fix linear: accuracy_fail (#4792) by @xzliu-opt
- Fix environment setting for Kunlunxin (#4795) by @tengqm
- Fix illegal memory access in cluster_remote_gemm_kernel when cdiv(N, BN) is odd (#4797) by @yysheng26
- [Hygon] fix mm float64 accumulation (#4799) by @douxetpur
- fix fft's res_out (#4802) by @Dayrker
- fix: test log output dir can't be access (#4805) by @wuwentao
- Fix cpp op ci error (#4811) by @0x45f
- fix: increase test job timeout and add label/csv/html to output dir (#4813) by @wuwentao
- fix: resolve CI test TIMEOUT/FAILED/OOM issues for multiple operators (#4829) by @CLYtou
- fix: container pip error and only run basic test (#4834) by @wuwentao
- fix: filter None and mismatched out keys in pointwise_dynamic prepare_args (#4849) by @bwbwzzz
- fix: nvidia/huawei/hygon/mthreads test env (#4851) by @wuwentao
- fix: kunlunxin remove fg_mode and use default kernel mode (#4854) by @wuwentao
- fix: nvidia container/mthreads env/huawei env for weekly ci job (#4859) by @wuwentao
- fix: test script import datetime error in main func (#4861) by @wuwentao
- fix(ci): correct setup-flaggems input name in random-test.yaml (#4869) by @tengqm
- Fix sdpa head dim k (#4881) by @0x45f
- fix(kunlunxin): remove --fg_mode operator from performance benchmark (#4886) by @wuwentao
- fix(test): add try/except for fp8_einsum import avoiding false CI failures (#4888) by @CLYtou
- fix: resolve complex conjugate recursion and vdot accuracy failures (#4890) by @CLYtou
- fix sinc (#4892) by @llaboon
- Fix tle error (#4894) by @0x45f
- fix(tests): fix mm tma test (#4905) by @103yiran
- Fix cpp import error (#4922) by @0x45f
- fix: register empty as empty.memory_format so torch.empty uses FlagGems (#4929) by @awayzjj
- fix: replace triton_src kernel references with ops/ paths in C++ wrappers (#4931) by @yysheng26
- fix(config): lazy import c_operators — detect .so without loading (#4959) by @tengqm
- fix(cmake): add global include_directories for NPU backend (#4960) by @tengqm
- fix(cpp): disable aten dispatch + exclude PrivateUse1 from redispatch guard (#4962) by @tengqm
- Fix setup for triton/flagtree combination (#4963) by @tengqm
- Fix mul fallback error (#4999) by @0x45f
- fix(rmsnorm): reorder multiply to (x * inv_rms) * weight for vendor precision alignment (#5009) by @CLYtou
- Fix hygon (#5016) by @huangyiqun
- fix(utils): add fallback_j1 (#5038) by @103yiran
- Fix CI job to runner mapping (#5049) by @tengqm
- fix: add tsingmicro/txda to device detection (#5056) by @tengqm
- fix(apply_rotary_pos_emb): fix cos/sin error (#5073) by @103yiran
- fix(utils): add pure-triton nextafter fallback for non-CUDA backends (#5088) by @tengqm
- Fix import error in ascend (#5101) by @0x45f
- fix: vector_norm accuracy test failure on mthreads backend (#5106) by @lyujheng
- fix: resolve copy op accuracy failure on mthreads backend (#5107) by @lyujheng
- fix: resolve_conj performance test failure on mthreads backend. (#5109) by @lyujheng
- fix: resolve mul op accuracy and performance test failure on mthreads (#5111) by @lyujheng
- [Ascend] Fix ascend ops when running Qwen3.6 vLLM with flagtree ascend3.5 (#5113) by @zhzhcookie
- fix(mul): gate optimized path on active backend device, not hardcoded "cuda" (#5130) by @tengqm
- fix: recognize vLLM stable ABI extensions (#5150) by @CherryLemon
- Fix out-of-bounds read past the input tail in bucketize kernel (#5231) by @truong-v
- Fix IX gate by pinning numpy version < 2 (#5244) by @tengqm
- fix: index_add accuracy failure on mthreads (#5246) by @lyujheng
- [KernelGen][metax] Fix num_warps exceeding hardware thread limit for softmax, logsumexp, dropout, exponential_, rand, randn, and uniform (#5258) by @yzw1128
- fix: mul accuracy test failure on mthreads (#5294) by @lyujheng
- fix: div accuracy test failures on mthreads backend (#5295) by @lyujheng
- [Enflame] fix enflame backend operator bug (#5302) by @chaaa-a
- fix(ascend): align full_like kernel API with updated full_kernel signature (#5308) by @yzw1128
- [KernelGen][Cambricon] Fix nested_view_from_buffer_copy: accuracy_fail (#5324) by @xzliu-opt
- fix add op test (#5325) by @huangyiqun
- Fix code style (#5331) by @0x45f
- fix(ops): skip triton kernels for complex dtypes in empty/lift_fresh/… (#5332) by @CLYtou
- fix(conf): add missing info for special_log_softmax (#5333) by @103yiran
- [KernelGen][Metax] Fix baddbmm: accuracy_fail (#5336) by @xzliu-opt
- fix(utils): add pure-triton j0 and log2 fallbacks for non-CUDA backends (#5348) by @tengqm
- fix(cambricon): tolerate triton forks lacking TRITON_MAX_TENSOR_NUMEL (#5351) by @tengqm
- Fix special_chebyshev_polynomial_w ut (#5352) by @0x45f
- fix: prevent run_tests.py deadlock when worker crashes (#5358) by @wuwentao
- fix flip op accuracy test failure on mthreads backend (#5374) by @lyujheng
- fix testcase iluvatar (#5375) by @jia-heng
- [KernelGen][Ascend] Fix specialerfc-import-error-on-ascend: import_error (#5380) by @xzliu-opt
- [KernelGen][Ascend] Fix any_dim: accuracy_fail (#5381) by @xzliu-opt
- Fix cat cpp wrapper kernel path, cat dtype mismatch, and rwkv_mm_sparsity precision (#5388) by @yysheng26
- fix: increase op test job timeout (#5391) by @wuwentao
- fix: skip kernel launch for zero-element tensors in empty() (#5410) by @lukatao
- Fix /test command parsing for space-separated params (#5425) by @tengqm
- Fix empty (#5438) by @0x45f
- [KernelGen][Mthreads] Fix conv2d_padding: accuracy_fail (#5445) by @xzliu-opt
- [Ascend] Fix input param of do_bench_npu (#5451) by @zhzhcookie
- fix(utils): add pure Triton lgamma fallback (#5464) by @huangyiqun
- fix(cambricon): stop index autotune from killing the process on untileable expand_shape (#5510) by @tengqm
- fix(utils): add pure-triton y0 fallback for non-CUDA backends (#5553) by @lukatao
- Fix empty tensor sum error (#5567) by @0x45f
- Fix ci import error (#5573) by @0x45f
- [metax] Fix ops registration and update heuristics/tune configs (#5576) by @yzw1128
- [metax] Fix upsample_linear1d autotune key mismatch (unblock metax CI) (#5591) by @yzw1128
- fix(i0): avoid unnecessary x.to() temporary tensor (#5592) by @kkkwb
- [KMCompiler][ASCEND] Fix tle import error for cholesky_solve op (#5616) by @YangLong114514
- [KMCompiler][Ascend] Fix bug (#5641) by @Willing2329
- [KMCompiler][Iluvatar] bug fix for operator nonzero_static (#5644) by @YangLong114514
- Fix M-Threads image and backend name in upload log (#5650) by @wuwentao
- [KMCompiler][Ascend] Fix replication pad2d backward bug (#5654) by @Willing2329
- [PPU][OP Test][Bug Fix] Fix upsample and quant operator test (#5655) by @tianyi-dialect
- [KMCompiler] fix linalg_vecdot: exclude fp64 benchmark for iluvatar and ascend platforms (#5658) by @rye985
- fix(flagtune): fall back when platform model package is unavailable (#5678) by @k6699k
- [KMComplier][NVIDIA][Thead][MetaX][iluvatar] Fix test_linalg_det hang with
--ref cpu(#5701) by @littlekid-zxt - [KMCompiler][iluvatar] Fix sparse_sampled_addmm compilation failure for K < 16 (tl.dot K >= 16 constraint) (#5705) by @littlekid-zxt
- fix(cambricon): use elementwise & | ~ instead of Python and/or/not on tensor masks (#5745) by @tengqm
- [Triton] Fix tl.where input param in randperm op (#5769) by @zhzhcookie
- fix dead ATen registrations and non-standard registration names (#5777) by @yysheng26
- fix _weight_norm: split underscore op tests into a dedicated mark (#5778) by @Yukun-Cui
- [Iluvatar] fix(pointwise): restore rank-1 vectorization on Triton 3.6 (#5783) by @awayzjj
- fix(cambricon): broadcast scalar max_virtual_index in paged load_from_kvcache (#5786) by @tengqm
- [kunlunxin] fix slice_scatter OOB src index on masked-out lanes (#5816) by @Cocytus486
- fixbug: exponential_ operator accuracy failure (#5830) by @lyujheng
- [KMCompiler][Nvidia]Fix igammac op register and its testcases related torch op name (#5831) by @lelongmask-cmp
- fix(utils): add pure-triton normcdfinv fallback for non-CUDA backends (#5864) by @lukatao
- [MTHREADS] Fix mthreads backend dispatch key (#5865) by @Oslomayor
- [NVIDIA] Fix Hopper MM split-K accumulation precision (#5876) by @Tiannn967
- [KernelGen][Nvidia] Fix index_put_impl (#5889) by @Yukun-Cui
- fix(benchmark): isolate MM operator measurements (#5902) by @Tiannn967
- fix:special_hermite_polynomial_h (#5930) by @douxetpur
- fix(ascend): fix multinomial NaN/OOM and cumsum device bug on Ascend backend (#5980) by @CLYtou
- fix(setup.sh): fix uv aarch64 unzip path error (#5999) by @103yiran
- fix(slice): fix missing default args and broken view semantics (#6004) by @kkkwb
- fix(broadcast_tensors): handle zero-size dims in broadcast shape computation (#6024) by @cccxxttt
- fix(hygon): KeyError: 'BLOCK_M' triggered by flash attention operations when running ViT encoder inference (#6040) by @xmhubj
- [KMCompiler][iluvatar] fix oom (#6070) by @YangLong114514
- fix(attention): prevent _attn_bwd runtime crashes on Hopper sm90 (#6075) by @cccxxttt
- fix(flash_attention): keep AABS from breaking tl.dot on decode backward (#6077) by @cccxxttt
- fix(index): cast pid_p to int64 to prevent int32 offset overflow (#6087) by @JosephNew
- Fix all and any zero div error (#6098) by @0x45f
- fix(nextafter): propagate NaN when target operand is NaN (#6100) by @cccxxttt
- fix(benchmark): write convert_weight_to_int4pack results as structured JSON (#6101) by @cccxxttt
- fix(benchmark): write int4pack_mm_with_scales_and_zeros results as structured JSON (#6118) by @cccxxttt
- [Fix] Respect FlagTune environment settings in MM tests and benchmarks (#6129) by @k6699k
- fix benchmark naming for irshift (#6131) by @freaking-man-in
- Fix expression syntax error in unittest.yaml concurrency group (#6137) by @103yiran
- fix: add missing DSA/init.py so non-editable installs include the subpackage (#6146) by @JosephNew
- Fix adaptive_avg_pool3d_backward dispatching to _adaptive_avg_pool3d_backward (#6147) by @kkkwb
- [KMCompiler][MetaX] Fix gru bug (#6155) by @Willing2329
- fix(.github): remove metax-maca3720 info (#6160) by @103yiran
- fix(sunrise): guard the vendor cumsum against empty input, add an asin fallback (#6165) by @tengqm
- fix(utils): add graph capture buffer reuse for pointwise_dynamic (#6171) by @CLYtou
- fix(ci): update kunlunxin/iluvatar flagtree ci docker image (#6192) by @wuwentao
- [Bug Fix] Enable SQLite WAL mode and busy_timeout to fix "database is locked" in autotune cache (#6203) by @modao1234
- fix(ci): stop sort_exports from deleting imports on double all (#6213) by @gavin0x01
- fix: gate Triton 3.6-only APIs (map_elementwise, triton.knobs) for the flagtree/Triton-3.5 line (#6236) by @JosephNew
- fix: add Triton 3.5 compatibility for flash_attention_backward (issue… (#6253) by @CLYtou
- [BugFix] Support boolean slice views (#6061) (#6350) by @0x45f
- Fix iluvatar import (#6228) (#6377) by @0x45f
- fix iluvatar randperm/sort (#5575) (#6380) by @0x45f
- fix(sqlite): tolerate WAL journal-mode switch race on multi-process cold start (#6388) by @JosephNew
- fix(autotune): reflect single table instead of whole db in build_sql_model_by_db (#6416) by @JosephNew
- [MetaX] Fix masked_fill broadcast OOB and illegal num_warps=16 (#6443) by @zhangjieBaai
- fix(enflame): chunk grid.x in index_put to respect the 65535 launch limit (#6538) by @JosephNew
- [Ascend] fix: improve Ascend cat/concat dispatch and implementation (#6547) by @modao1234
- fixbug: issue#6420 (#6566) by @lyujheng
- Fix variable reference for pytest operator check (
9479fcd1214d; commit) - Fix typo in pip install command in workflow (
b4d5d12adb13; commit) - Fix activation of virtual environment in workflow (
aff20564630e; commit) - Fix concurrency group expression in unittest.yaml (
3896f48357ab; commit)
Other
- [SiliconFlow] Operator Development adaptive_avg_pool2d (#2059) by @fpzh2011
- [Advanced Compiler] topk (#3329) by @AdvancedCompiler
- [SiliconFlow] Op: flash_attention_backward (#3386) by @LiangJiabaoY
- [Performance Optimization] Reduce test cases in QUICK_MODE (#3423) by @git-flyer
- Regist vllm ops in torch (#3497) by @0x45f
- scaled_dot_product_attention_backward support non-square tuning, non-causal, and unequal lengths of q and kv sequences (#3607) by @103yiran
- Prefer Triton over FlagTree on Iluvatar (#3666) by @tengqm
- Skip flash_mla_sparse_fwd operator (#3692) by @tengqm
- Be tolerant when running batch/daily tests (#3696) by @tengqm
- Initial version of container file for mthreads (#3697) by @tengqm
- Standardize Iluvatar backend debug log format (#3711) by @lukatao
- Standardize MetaX backend debug log format (#3712) by @lukatao
- Skip the fused_marlin_moe accuracy test due to assertion failures (#3734) by @tengqm
- Standardize Cambricon backend debug log format (#3737) by @lukatao
- Standardize Mthreads backend debug log format (#3738) by @lukatao
- Standardize Kunlunxin backend debug log format (#3739) by @lukatao
- Standardize Aipu backend debug log format (#3740) by @lukatao
- Standardize Hygon backend debug log format (#3742) by @lukatao
- Standardize Tsingmicro backend debug log format (#3743) by @lukatao
- Pin triton to 3.6.0 in setup script (#3758) by @tengqm
- Opname fix (#3780) by @tianxiao-baai
- update _sunrise backend (#3791) by @Dayrker
- Container file for the iluvatar backend (#3808) by @tengqm
- Downgrade torch from 2.12 to 2.11 for Nvidia (#3825) by @tengqm
- Re-enable tsingmicro backend test (#3827) by @tengqm
- Enable backend tests for enflame, spacemit and sunrise (#3830) by @tengqm
- Simplify pyproject.toml for extra dependency specification (#3843) by @tengqm
- Allowing for specifying all GPUs when testing (#3845) by @tengqm
- Try skip vendor.sh from the setup.sh script (#3848) by @tengqm
- Containerfile for MetaX 3.7.2.1 (#3849) by @tengqm
216 more changes
- Rename some containerfiles (#3850) by @tengqm
- Update pyproject.toml (#3851) by @tengqm
- Dockerfile for enflame (#3855) by @tengqm
- Containerfile for Hygon 26.04 (#3858) by @tengqm
- [Sunrise] Dev sunrise fix v4 (#3865) by @Dayrker
- Tsingmicro chip count (#3887) by @tengqm
- Update containerfile for mthreads backend (#3903) by @tengqm
- Update operator registry (#3925) by @tengqm
- No gpu op fix (#3935) by @tianxiao-baai
- Sort operators by name (#3938) by @tengqm
- Containerfile for Ascend backend (#3974) by @tengqm
- Clarify operator stages (#3979) by @tengqm
- Update package/environment setting for tsingmicro (#3982) by @tengqm
- Update triton version for sunrise backend (#3989) by @tengqm
- update tsingmicro container file to support flagtree (#4017) by @tengqm
- Simplify ascend container file (#4021) by @tengqm
- update containerfile for tsingmicro backend (#4035) by @tengqm
- Update version string for the master branch (#4038) by @tengqm
- Update CI setup for tsingmicro backend (#4044) by @tengqm
- Debug Ascend chip check (#4062) by @tengqm
- flaggems.ops to flaggems. (#4085) by @chongzhouyang
- Remove buggy plugin dragged in by Triton on kunlunxin (#4090) by @tengqm
- Amend the env.sh for environment setup on THead backend (#4107) by @tengqm
- update from enflame until 20260616 (#4111) by @chongzhouyang
- [SiliconFlow] Skip segment_reduce benchmark on MUSA (#4115) by @winfan1314
- [KMCompiler] Reimplement top_k_per_row_prefill and top_k_per_row_decode (#4132) by @liuxiao0909c
- Upgrade containerfile for nvidia 13.3 (#4135) by @tengqm
- Install vllm=0.21.0 on H20 for operator testing (#4136) by @tengqm
- Update containerfile for MetaX (#4140) by @tengqm
- Update container file for mthreads (#4149) by @tengqm
- Update containerfile for Iluvatar (#4150) by @tengqm
- Update container file for Ascend 8.5.0 (#4151) by @tengqm
- Update project setting for CANN 9.0.0 (#4152) by @tengqm
- Cleanse the operator registry again (#4162) by @tengqm
- Rename 'ilshift__' to 'ilshift' (#4163) by @tengqm
- Allow cann850 and cann9 to be specified for workflow (#4166) by @tengqm
- Speedup batch test run for non-existent test cases (#4171) by @tengqm
- Reorg CI workflows (#4172) by @tengqm
- Drop some useless tools that are no longer useful or maintained (#4173) by @tengqm
- Drop experimental operator 'abs' and 'abs_' (#4177) by @tengqm
- Drop experimental operator
addcdiv(#4180) by @tengqm - Drop experimental operator celu and celu_ (#4182) by @tengqm
- Drop experimental operator elu (#4184) by @tengqm
- Drop experimental operator 'exp2' (#4186) by @tengqm
- Drop experimental operator exp2_ (#4188) by @tengqm
- Drop the experimental operator glu (#4190) by @tengqm
- Drop experiemental operator 'logical_xor_' (#4192) by @tengqm
- Drop experimental operator maximum (#4194) by @tengqm
- Drop the experimental operator
mse_loss(#4196) by @tengqm - Update containerfile for tsingmicro (#4197) by @tengqm
- Drop experimental operator
eye(#4204) by @tengqm - Drop experimental operator
replication_pad3d(#4205) by @tengqm - Drop experimental operator
slice_backward(#4206) by @tengqm - Drop experimental operator
smooth_l1_loss(#4207) by @tengqm - Drop experimental operator
triu(#4208) by @tengqm - Drop experimental operator
leaky_relu(#4210) by @tengqm - Drop experimental operator 'mv' (#4212) by @tengqm
- Drop experimental operator 'slice_scatter' (#4213) by @tengqm
- Drop experimental operator 'deg2rad' (#4215) by @tengqm
- Drop experimental operator 'diag' (#4221) by @tengqm
- Drop experimental operator 'pixel_shuffle' (#4225) by @tengqm
- Drop experimental operator 'trace' (#4229) by @tengqm
- Remove skip annotation from trace benchmark (#4230) by @tengqm
- Drop experimental operator 'upsample_nearest1d' (#4233) by @tengqm
- Use correct vendor string for GPU check (#4235) by @tengqm
- Cache gh CLI install on self-hosted runners (#4241) by @tengqm
- Consolidate backend dispatch using matrix strategy (#4242) by @tengqm
- Change logging warning to info (#4246) by @0x45f
- Drop experimental operator
fix(#4249) by @tengqm - Drop experimental
masked_scatterimplementation (#4257) by @tengqm - Drop experimental operator
addcmul_(#4259) by @tengqm - Drop experimental operator
atanh_(#4260) by @tengqm - Drop experimental
masked_selectoperator (#4264) by @tengqm - Drop the experimental
reciprocaloperator (#4266) by @tengqm - [MACA] Support C++ wrappers with libtriton_jit (#4270) by @cjh9368
- Change to debug level (#4273) by @0x45f
- Setup script for Cambricon (#4307) by @tengqm
- Use vendor name in backend-tests job names (#4330) by @tengqm
- Change test count to time (#4338) by @0x45f
- Use internal mirror for gh CLI download (#4347) by @tengqm
- Persist corrected HOME to GITHUB_ENV in command.yaml (#4349) by @tengqm
- update vendor's debug logger (#4361) by @huangyiqun
- Make sure the preprocess step has source code checked out (#4373) by @tengqm
- Drop vendor script (#4379) by @tengqm
- Remove two operators from registry (#4380) by @tengqm
- Widen normal distribution test tolerance to reduce flakiness (#4386) by @tengqm
- Update configuration for enflame backend (#4406) by @tengqm
- Separate compiler selection from deps in backends.yaml (#4407) by @tengqm
- Drop the non-existent special.gammainc.out operator (#4408) by @tengqm
- Merge some operator code for deduplication (#4410) by @tengqm
- Post failure report when on-demand test fails to execute (#4413) by @tengqm
- Split nvidia backend into nvidia-cuda128 and nvidia-cuda133 (#4418) by @tengqm
- Update workflow files to use revised backend labels (#4421) by @tengqm
- [QC] Enable swap_ab on H20 for small-tile fp8 blockwise GEMM (#4434) by @monellz
- mul optimization from pointwise_dynamic to expert implementation (#4436) by @liuxin2560
- Update package config and env var for TsingMicro (#4449) by @tengqm
- [TSINGMICRO] update tsingmicro backend 0630 (#4450) by @tsingmicro-public-e
- [TSINGMICRO] skip failed tests and long time benchmarks (#4452) by @tsingmicro-public-e
- Use H20 backend for UT (#4462) by @tengqm
- Update code owners file (#4470) by @tengqm
- style: sort operator registrations alphabetically (#4477) by @lukatao
- Rename the artifact from on-demand workflow (#4479) by @tengqm
- Restore python-op CI job to jiuding machines (#4481) by @tengqm
- Disable flagtree for sunrise (requires GLIBC_2.39) (#4495) by @tengqm
- Simplify flag_gems version detection in run_tests.py (#4499) by @tengqm
- Simplify flag_gems version detection in run_tests.py (#4505) by @tengqm
- Remove extra LD_LIBRARY_PATH setting for metax (#4517) by @tengqm
- Remove the extra post-install step for Kunlunxin (#4518) by @tengqm
- run_tests: cap worker count to number of ops (#4529) by @tengqm
- [C++ Wrapper] Align softmax and cat with Python wrapper (#4542) by @yysheng26
- style: sort operator registrations alphabetically (#4559) by @lukatao
- Tune Triton-TLE topk paths (#4572) by @zheng1
- Update dependency and config for Sunrise (#4596) by @tengqm
- [TSINGMICRO] update tsingmicro backend to 0708 (#4603) by @tsingmicro-public-e
- Simplify NVIDIA backend config (#4606) by @tengqm
- [TSINGMICRO] align grouped_topk with vllm grouped_topk (#4610) by @tsingmicro-public-e
- [TSINGMICRO] align moe_align_block_size with vllm op (#4611) by @tsingmicro-public-e
- [TSINGMICRO] session.mege will introduce dirty data while cocurrently… (#4612) by @tsingmicro-public-e
- Skip
affine_grid_generatoraccuracy tests (#4614) by @tengqm - Remove pre_install from tsingmicro backends (#4617) by @tengqm
- chore: remove container template files — migrated to build-infra (#4620) by @tengqm
- chore: remove container/ — legacy files migrated to build-infra (#4624) by @tengqm
- chore: remove runtime_env from backends.yaml — migrated to build-infra (#4629) by @tengqm
- chore: remove tools/env.sh from CI workflows (#4641) by @tengqm
- [TSINGMICRO] update tsingmicro backend to 0709 (#4643) by @tsingmicro-public-e
- Increase timeout in run_cmd function (#4645) by @tianxiao-baai
- chore: remove requirements/ — superseded by backends.yaml (#4649) by @tengqm
- Refine masked fill test (#4665) by @0x45f
- Move optimized mul implementation to the general path (#4666) by @Tiannn967
- Move compiler (flagtree/triton) management exclusively to setup.sh (#4673) by @tengqm
- Debug ascend gpu_check failure; remove HOME override from command.yaml (#4677) by @tengqm
- Rename post_install to triton_post_install; remove post_uninstall (#4680) by @tengqm
- Remove debug logging from gpu_check_ascend.sh (#4681) by @tengqm
- Drop experimental negative operator (#4683) by @tengqm
- Always persist uv PATH to GITHUB_PATH (#4684) by @tengqm
- Promote deg2rad_ (in-place) from experimental to stable (#4685) by @tengqm
- Drop experimental _unsafe_view operator (#4723) by @tengqm
- Drop experimental amin operator (#4724) by @tengqm
- Drop experimental expand operator (#4725) by @tengqm
- Drop experimental permute_copy operator (#4728) by @tengqm
- Drop experimental relu operator (#4729) by @tengqm
- Unify vendor/backend naming across CI workflows (#4730) by @tengqm
- command.yaml: route h100 runner to nvidia-cuda128 backend (#4734) by @tengqm
- command.yaml: use COMPILER=triton for h100 runner (#4735) by @tengqm
- Cap setuptools below 77 to keep wheel Metadata-Version at 2.2 (#4738) by @tengqm
- Split C++ extension into a separate per-vendor wheel (#4745) by @tengqm
- Update env setting for Tsingmicro backend (#4796) by @tengqm
- Drop experimental operator sinc (#4810) by @zheng1
- Ship test and benchmark suites in the wheel (#4832) by @tengqm
- Minor revision for gate testing (#4878) by @tengqm
- [TSINGMICRO] update tsingmicro backend to 0720 (#4882) by @tsingmicro-public-e
- Cap setuptools-scm below 10 to fix
pip install .build failure (#4884) by @wuwentao - update code owner (#4941) by @huangyiqun
- [SiliconFlow] Operator fix adaptive_avg_pool2d (#4952) by @fpzh2011
- Relax numpy and packing version constraint (#5002) by @tengqm
- Make upload-artifact step optional in on-demand test workflow (#5007) by @tengqm
- [_sunrise] update sunrise ops (#5024) by @Dayrker
- Pick MXFP4 MoE block_m by padding cost, not an absolute M cutoff (#5027) by @zeroherolin
- [SiliconFlow] Restrict _scaled_mm FP8 tests to NVIDIA (#5062) by @winfan1314
- [TSINGMICRO] update tsingmicro backend to 0804 (#5181) by @tsingmicro-public-e
- [_sunrise] update sunrise ops (#5238) by @Dayrker
- Upgrade compiler sunrise (#5268) by @tengqm
- [KernelGen] Split xor/ixor operators and register ixor (#5283) by @xzliu-opt
- fused: int32 slot_mapping for enflame reshape_and_cache_flash (GCU300) (#5345) by @tengqm
- Append commit id to version for dev builds (#5346) by @tengqm
- [AddMM] Migrate correctness tests to public APIs (#5387) by @ZYX223
- Use Triton for metax CI (#5396) by @tengqm
- Restore flagtree backend for metax CI (#5405) by @tengqm
- Match CI backend to runner node for metax and ascend (#5408) by @tengqm
- Use space as parameter separator for command workflow (#5417) by @tengqm
- Rename fused_marlin_moe kernels by precision scheme and split benchmarks (#5437) by @zeroherolin
- Try fix Iluvatar CI job (#5441) by @tengqm
- [_sunrise] update commits of sunrise backend. [till 20260812] (#5478) by @Dayrker
- Update FlagTree and vendor docker images (#5628) by @wuwentao
- Update Huawei Ascent 910b docker image (#5681) by @wuwentao
- [KMCompiler] linalg vecdot: fix fp64 test logic (#5766) by @rye985
- Align operators.yaml (#5836) by @yysheng26
- [KMCompiler] Register out-of-place exponential and use aten out-of-place op as benchmark baseline (#5843) by @littlekid-zxt
- [KMCompiler]update hygon linalg_solve_triangular (#5943) by @YangLong114514
- [k] nn category ops fix (#5969) by @RobertLuobo
- [k] index and sorting category ops fix (#5971) by @RobertLuobo
- [k] linalg category ops fix (#5973) by @RobertLuobo
- [k]pecial function category ops fix (#5974) by @RobertLuobo
- [k]vision ops correctness fix (#5975) by @RobertLuobo
- [k] register new ops and update backend configs (#6009) by @RobertLuobo
- [KMCompiler] Operator linalg_solve_triangular, add skip for ascend baseline path. (#6209) by @YangLong114514
- Cherry-pick #4939 to 5.4.0-rc2: add missing init.py so fused/DSA and backend subpackages ship (#6391) by @shiptux
- Cherry-pick #5683 to 5.4.0-rc2: support SQLAlchemy 1.4 in the SQL persistent model (#6392) by @shiptux
- Cherry-pick #6400 to 5.4.0-rc2: Nexus upload caller (#6401) by @shiptux
- Update alpha version from 5.3 to 5.4 in operators.yaml (
3ae0fb8e16a3; commit) - Modify pip install command in setup.sh (
8895634ccaa1; commit) - Update command.yaml (
462843efd13a; commit) - Change runner label from 'hopper' to 'jiuding' (
98a513c1a86c; commit) - Change runner label from 'jiuding' to 'h20' (
03c03e161133; commit) - Clean up test-op-experimental.sh script (
a016074c5f2c; commit) - Integrate setup-gh action in command.yaml (
731570d38764; commit) - Update setup-gh action to actions4gh/setup-gh@v1 (
5f8e6f1e78ac; commit) - Remove github-token from setup-gh action (
4c24a33b905b; commit) - Update command.yaml (
91d31d44c576; commit) - Rename VENDOR variable to BACKEND in env.sh (
838ed92c3a39; commit) - Update regex for operator existence check in command.yaml (
5c1fcbc2f4bc; commit) - Update vendor to nvidia-cuda133 in unittest.yaml (
12ab37346a7f; commit) - Update unittest.yaml (
a1895ee80d10; commit) - Remove backend-test.yaml from unittest workflow (
86363bfdb6e3; commit) - Modify PATH for Python and CUDA in workflow (
a0f7e86efe4a; commit) - Modify PATH for build environment (
ad496ec782a0; commit) - Update pull request template to remove license section (
427fa4835620; commit) - Update black version in pre-commit config (
40baf00165c0; commit) - Update flake8 args to include E704 (
4b86060f3a6f; commit) - Modify flake8 args in pre-commit config (
50d5ad35e7ab; commit) - Update torch version to 2.9.1 in backends.yaml (
07f4c30c8d75; commit) - Update numpy dependency to allow versions greater than 2 (
146a1eb07ac1; commit) - Update upload-artifact action to version 7.0.0 (
5e8767a5593a; commit) - Revise numpy version requirement comment (
b4c271fd8e4d; commit) - Update CODEOWNERS to modify ownership paths (
8fcd87f88f59; commit) - Update backends.json (
34bd6d68928c; commit)
Contributors
Thanks to the contributors to the pull requests included in this release.
114 contributors
- @tengqm (262 PRs)
- @XDYuanzhuLee (132 PRs)
- @LoserCheems (53 PRs)
- @lukatao (38 PRs)
- @huangyiqun (36 PRs)
- @bwbwzzz (35 PRs)
- @0x45f (33 PRs)
- @NineAnnAnn (27 PRs)
- @Yukun-Cui (27 PRs)
- @CoyeCHEN (26 PRs)
- @chx7514 (26 PRs)
- @Dingxingdi (25 PRs)
- @ShawnsYing (25 PRs)
- @wuwentao (24 PRs)
- @xuanzhengdu-eng (24 PRs)
- @YangLong114514 (22 PRs)
- @Oslomayor (21 PRs)
- @103yiran (19 PRs)
- @winfan1314 (15 PRs)
- @Willing2329 (12 PRs)
- @gavin0x01 (12 PRs)
- @littlekid-zxt (11 PRs)
- @yzw1128 (11 PRs)
- @CLYtou (10 PRs)
- @Dayrker (10 PRs)
- @RobertLuobo (10 PRs)
- @YufanPeter (10 PRs)
- @lyujheng (10 PRs)
- @tsingmicro-public-e (10 PRs)
- @wuhansky (10 PRs)
- @awayzjj (9 PRs)
- @chongzhouyang (9 PRs)
- @dependabot (9 PRs)
- @JosephNew (8 PRs)
- @ZYX223 (8 PRs)
- @Caeruleann (7 PRs)
- @Tiannn967 (7 PRs)
- @chenkc2025 (7 PRs)
- @kkkwb (7 PRs)
- @xzliu-opt (7 PRs)
- @yysheng26 (7 PRs)
- @AdvancedCompiler (6 PRs)
- @cccxxttt (6 PRs)
- @douxetpur (6 PRs)
- @gao0624 (6 PRs)
- @Galaxy1458 (5 PRs)
- @cheersluvs (5 PRs)
- @ggbondbest (5 PRs)
- @zeroherolin (5 PRs)
- @zhzhcookie (5 PRs)
- @Jennie1004 (4 PRs)
- @July-h5kf3 (4 PRs)
- @LittleShun1214 (4 PRs)
- @Yangyang0906C (4 PRs)
- @fpzh2011 (4 PRs)
- @monellz (4 PRs)
- @shiptux (4 PRs)
- @stringl1l1l1l (4 PRs)
- @tianxiao-baai (4 PRs)
- @zheng1 (4 PRs)
- @Fattysand (3 PRs)
- @KingsleyLiu-NV (3 PRs)
- @WhatGhost (3 PRs)
- @elllzat (3 PRs)
- @k6699k (3 PRs)
- @liuxin2560 (3 PRs)
- @modao1234 (3 PRs)
- @rye985 (3 PRs)
- @sh1653487844 (3 PRs)
- @Alvin-YCHEN (2 PRs)
- @Salamanca001 (2 PRs)
- @ShawnXuan (2 PRs)
- @chaaa-a (2 PRs)
- @chen98038 (2 PRs)
- @fangzexian (2 PRs)
- @freaking-man-in (2 PRs)
- @git-flyer (2 PRs)
- @jia-heng (2 PRs)
- @kevinzs2048 (2 PRs)
- @lelongmask-cmp (2 PRs)
- @llaboon (2 PRs)
- @mygitljf (2 PRs)
- @tspyc072 (2 PRs)
- @xmhubj (2 PRs)
- @Ason93
- @CelysPr
- @CherryLemon
- @Cocytus486
- @DannyP0
- @HenryRenYz
- @Keriza
- @LiangJiabaoY
- @Onisen7
- @Salvationhope
- @Vincent-Xiao
- @YUANXIGUREN
- @bin913
- @chenmiao1919
- @cjh9368
- @dawnyang050-debug
- @huatuoli
- @liuxiao0909c
- @lzllx123
- @nianqi-tian
- @physics31415926
- @sgjzfzzf
- @shanyulu
- @tianyi-dialect
- @truong-v
- @w1120029931-bit
- @wangxshuai
- @wbavon
- @yilongma0110
- @zhangjieBaai