Skip to content

NVMAI v3.0

Choose a tag to compare

@Pummelchen Pummelchen released this 11 Aug 01:18
· 992 commits to main since this release

Concise mode: per-quantization terse system prompts (server NVMAI_CONCISE_MODE, CLI --concise, app setting) cutting answer tokens 55-61% (measured 4/6/8-bit); model dirs renamed to qwen3.6_35B_A3B_{4,6,8}Bit; merged PR #1 (manifest family detection + HF redirect fix); fresh benchmark suite (Aug 2026). Full history: wiki Optimization-Journey.