Skip to content

Latest commit

 

History

37 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 

Repository files navigation

AI-knowledge

A common place for putting knowledge from different AIs.

Topics

Active plan

MiniMax-M3 MXFP8

  • Qualify the ATOM-style dense-attention path at the P-only SLA boundary. Kernel correctness and dispatch are already proven; the complete gather/dequant + BF16 FMHA + merge path is 2.12--2.16x faster than Unified Attention on production-like shapes. Run a clean, uninstrumented open-loop boundary A/B with identical stacks and confirm output quality, dispatch, TTFT, and the highest passing QPS. Do not repeat the fixed-shape UT. Dense attention covers only 3/60 layers and 6.712% of the current event-attributed GPU time, so the theoretical whole-model gain is about 3.7% before other overheads.

  • Raise the D-only SLA boundary above 2.20 QPS. First combine the already-qualified fused routed+shared-expert candidate with the complete deploy-safe D stack and remeasure the open-loop TPOT boundary. If the gain survives, optimize the Triton sparse-decode split-K + merge path. Current priorities and every measured candidate are tracked in the Decode optimization record.

About

A common place for putting knowledge from different AIs.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages