v0.55.0
Changes since v0.54.0:
Performance
- real-text H100 board — e2e 4/6 wins, decode 5/6, q35 decode flipped; q9 f16 mirrors per-model gated
- q35 shexp fused sigmoid-dot (+13.7% H100 decode), E4B head-last prime (111x prefill), matmul_pre empty-fallback guard
- g26 expert-down through the int8-MMA expert GEMM — H100 prefill 2901 -> 5042 tok/s (1.74x)
Fixes
- board prompt fox-repeat -> real text; recover uncommitted board evidence
- segment exec-update captures at kernel-class boundaries — q35 graph gate green
- route gemma4 through its dc/graph twins on the device-counter entry points
- decode-batch gate1 multi-seed recalibration; round 45 ledger
Documentation
- H100 board shows e2e tok/s per model — one number per arm, no pp/decode split
Boards + reproduction artifacts: https://huggingface.co/Avifenesh/memra-bench · full experiment log in research/tune-data/