v0.7.7
Cortiq 0.7.7
MiMo-V2.6-Flash
- Native CMF conversion and inference with q4tp MoE experts, mixed full/sliding attention, automatic resident/dynamic/hybrid GPU placement and GPU MTP verification.
- Text-only and full multimodal file sets share one backbone; MTP and image/video/audio towers are automatically discovered as companions.
- CLI and OpenAI-compatible server accept ordered image, silent-video and WAV inputs. Invalid media is rejected before SSE; text-only requests do not load media towers.
- Scoped GPU precision and worker-placement fixes preserve the reference numeric contracts. Overlapping MTP requests now use a counted row-exact scope, with overlap/nesting/unwind regression coverage. All six exact-weight GPU vision/video fixtures pass.
Verified measurements
On RTX PRO 6000 Blackwell 96 GB, the default 128-token core benchmark measured 32.20 / 41.40 / 40.85 tok/s (median 40.85). A 64000-MiB weight budget measured 44.48 tok/s median. Natural prompts vary and MTP is not always faster. CPU/full-GPU/24-GB PPL128 matches at 3.647; EN/RU/code plain/MTP sequences match all 128 token IDs. A 456-token prompt + 64-token CPU/GPU continuation also matches exactly with the CMF_SDOT=0 exact-float diagnostic; this is not a default-speed result.
Known limits
Budget tests used one 96-GB GPU, not physical cards at every capacity; temporary upload buffers can exceed the weight budget. Strict quantized-vision row cosine and precise video timestamp semantic checks remain open even though OCR, shape/chart and five ASR checks pass. No MP4, compressed audio, interleaved video/audio or media generation support is claimed. GPU MTP is greedy-only.
Install
cargo install cortiq-cli --version 0.7.7 --locked
cortiq --versionGPU support is enabled by default. Prebuilt CLI and FFI archives are attached with SHA-256 files.