Repository navigation
b11453
The gather path attended over the selected latents with a plain matmul and softmax. It only ran with n_ubatch <= 16, and the flash attention backends now skip the masked rows through n_kv_max, so the scatter path covers every case. Drop the gather flag, the gathered attention branch and gather_mla_rows. set_input_kpool always maps padding to the n_kv sentinel, and the slot mask becomes sel_mask since only the scatter reads it.