Releases: Tylogi/MFQ
Release list
MFQ Studio 0.3.2 - DeepSeek V4.1 Prefill at 471 tok/s
MFQ Studio 0.3.2
This release focuses on DeepSeek V4.1 Flash raw-HF prefill performance on Apple M3 Ultra, while preserving the 0.3.1 token-generation and MTP paths.
Highlights
- Reaches up to 471.1 tok/s warmed prefill for DeepSeek V4.1 Flash raw-HF on a 512 GB Mac Studio M3 Ultra.
- Adds an exact FP32 Metal top-k path inspired by DeepSelect for DeepSeek V4.1 sparse attention.
- Accelerates large-M MoE prefill with the native sorted-MXFP4 path.
- Auto-selects a 5,440-token maximum prefill chunk for the resident M3 Ultra geometry, keeping routed expert rows below the MLX 32,768-row boundary.
- Fuses window-KV RMSNorm, partial RoPE, and activation fake quantization.
- Keeps the backbone and MoE experts resident in unified memory while offloading only Engram storage to the Mac's internal SSD.
- Refreshes the Tylogi AI Lab identity used in the repository documentation.
Measured DeepSeek V4.1 performance
Measured locally at batch size 1, 32,768-token context, temperature 0, and warmed execution on a 512 GB Apple M3 Ultra:
- Model load: 38.6 s.
- Prefill: 390.2 / 458.6 / 471.1 tok/s for representative 509 / 2,009 / 15,969-token deterministic repeated raw-completion prompts.
- Token generation with MTP off: 18.29 tok/s.
- High-acceptance MTP: 33.46 tok/s (+83.0%), with 132/132 drafted tokens accepted and output identical to non-MTP decoding.
- Steady-state process RSS after load: approximately 287 GiB.
MTP gains depend on workload predictability and remain disabled by default.
Validation
- 9 Metal component/source tests passed for the DeepSelect, MXFP4, chunk-autotuning, and fused KV preparation paths.
- Native AppleClang Release build completed on an Apple M3 Ultra.
- Metal runtime self-test, packaged CLI command checks, TypeScript build, Rust checks, Tauri release build, deep code-signature verification, and
hdiutil verifyall passed. - The packaged app and CLI report version 0.3.2; native executables are arm64.
Download
MFQ.Studio_0.3.2_aarch64.dmg: self-contained Apple Silicon macOS app.SHA256SUMS.txt: release checksum.
Build information
- Source tag:
v0.3.2 - Source commit:
fdfb12852a226b6219db0320f8afca3d75da591d - DMG size: 703,880,572 bytes
- DMG SHA-256:
cb2459650a74c5eeb3616fbb2884bbd10641f54eae062173ccacdc84b22599ee
This macOS build is ad-hoc signed and has not been Apple-notarized.
MFQ Studio 0.3.1 - Faster DeepSeek V4.1 and MTP
MFQ Studio 0.3.1
This release ships the latest DeepSeek V4.1 Flash raw-HF Metal prefill and MTP improvements for Apple Silicon, together with synchronized 0.3.1 version metadata across the CLI and desktop app.
Highlights
- Adds a direct circular sparse-attention path for DeepSeek V4.1 prefill and multi-row MTP verification.
- Avoids discarded intermediate vocabulary projections during chunked prefill.
- Uses variable-width DSpark proposals and skips unused diagnostic graphs during serving.
- Adds deterministic acceptance-based MTP depth control and a fast handoff to ordinary decoding after two low-yield cycles.
- Preserves DeepSeek V4/V4.1 attention, Engram, and DSpark prefix state in resident session snapshots so independent agent/session forks can reuse an exact prefix.
- Keeps the full backbone and MoE experts resident on a 512 GB M3 Ultra while offloading only Engram storage to the Mac's internal SSD.
Measured DeepSeek V4.1 performance
Measured locally at batch size 1, 4,096-token context, 2,048-token maximum prefill chunks, temperature 0, and warmed execution:
- Prefill: 293.8 tok/s at 511 tokens and 384.0 tok/s at 2,031 tokens.
- Repeated raw-completion prefill: up to 397.0 tok/s at 1,965 tokens.
- Token generation with MTP off: 18.29 tok/s.
- High-acceptance MTP: 33.46 tok/s, an 83.0% uplift, with 132/132 drafted tokens accepted and output identical to non-MTP decoding.
MTP gains depend on workload predictability. In a lower-acceptance code case (50% acceptance), MTP measured 17.52 tok/s versus 17.59 tok/s without MTP, then automatically handed off to ordinary decoding while preserving deterministic output. MTP remains disabled by default.
Other source updates
- Integrates NAQ imatrix calibration and unified NINTv2 GPTQ/GSQ workflows.
- Includes the latest packed CUDA SQ/NINT/NVQ/NEPQ/MXFP backward and small-M kernel improvements in the source release.
Validation
- 129 C++ runtime source-contract tests passed; 9 platform-conditional tests skipped.
- Native AppleClang Release build completed on an Apple M3 Ultra.
- Metal runtime self-test, packaged CLI command checks, TypeScript build, Rust checks, Tauri release build, deep code-signature verification, and
hdiutil verifyall passed. - The packaged app reports version 0.3.1; the bundled CLI reports
mfq 0.3.1; native executables are arm64.
Download
MFQ.Studio_0.3.1_aarch64.dmg: self-contained Apple Silicon macOS app.SHA256SUMS.txt: release checksum.
Build information
- Source tag:
v0.3.1 - Source commit:
de3a78fd23df6d160c95e9a02f89a072c522e81f - DMG size: 703,895,961 bytes
- DMG SHA-256:
18cc50dcf75d5a64bbe85371d72189c0da108f2f64f4a5f18fac4000a189d556
This macOS build is ad-hoc signed and has not been Apple-notarized.
MFQ Studio 0.3.0 - DeepSeek V4.1
MFQ Studio 0.3.0
This release adds optimized native DeepSeek V4.1 Flash raw-HF inference on Apple Silicon.
Highlights
- Loads the original DeepSeek V4.1 Flash Hugging Face checkpoint directly, without MFQ conversion.
- Adds an M3 Ultra full-resident path while keeping Engram tables offloaded to SSD.
- Accelerates Metal prefill, token generation, attention, MoE, hyper-connection, and scheduling paths.
- Automatically materializes deterministic tokenizer and Engram runtime assets.
- Preserves DeepSeek V4 Flash behavior with device-scoped V4.1 optimizations.
- Documents raw-HF benchmark results near the top of the README.
Measured performance
Measured on an M3 Ultra with 512 GB unified memory, the model on the internal SSD, 4K context, 512-token chunks, and MTP disabled:
- Warm load: 38.6 s
- Prefill: 285.6 tok/s at 511 tokens; 338.4 tok/s at 2031 tokens
- Stable token generation: 18.20 tok/s
- Runtime memory: approximately 286 GiB
The packaged app was also mounted from this DMG and used to register, load, and generate with the full DeepSeek V4.1 checkpoint. Its hot smoke runs reached 19.36 tok/s and 17.50 tok/s on a distinct prompt, with zero failed requests.
TODO
- Complete and optimize DeepSeek V4.1 MTP support; MTP remains disabled by default.
- Reduce cold startup and first-request overhead.
- Continue prefill and token-generation optimization.
Download
- MFQ.Studio_0.3.0_aarch64.dmg: Apple Silicon macOS app
- SHA256SUMS.txt: release checksum
Build information
- Source tag: v0.3.0
- Source commit: 2060191
- DMG SHA-256: af41c6a096c49c781d0ecf624008f7b512366f0cb4289f4fe5572056bcc152a7
This macOS build is ad-hoc signed and has not been Apple-notarized.
MFQ Studio 0.2.0
MFQ Studio 0.2.0
This release expands the MFQ Studio native runtime, multimodal model support, and high-throughput inference paths.
Highlights
- Adds native Flash-Next multimodal inference, DeepSeek V4 vision with DSpark MTP, and new Qwen4/Qwen3.5 runtime paths.
- Adds Qwen continuous batching with paged KV cache, cancellation-safe lifecycle handling, and batched greedy sampling.
- Adds tensor- and expert-parallel CUDA graph support, with improved multi-GPU peer copies and output reduction.
- Accelerates mixed-precision MoE and NINT/NVQ kernels across CUDA and Metal, including small-batch decode and prefill paths.
- Redesigns the Studio backend console, streams quantization output, and makes model status clearer.
- Refactors and internalizes the native text runtime while unifying model graph and CUDA runtime components.
Downloads
MFQ.Studio_0.2.0_x64-setup.exe— Windows x64 (CUDA)MFQ.Studio_0.2.0_aarch64.dmg— macOS Apple Silicon
Build information
- Commit:
477cc0666ccd52f2937799568172f0114c6a6553 - Tag:
v0.2.0 - Windows SHA-256:
d723a5cf71882867776563b98466f7474224887660717bc57f7bde213ea865b6 - macOS SHA-256:
2eb00bb2d8b0cc8e48deddc5b9ba9044487d2184dd97a62d2e5ffc261f66b4ea
The macOS build is ad-hoc signed and is not Apple-notarized. On first launch, macOS may require Control-clicking the app and choosing Open.
MFQ Studio 0.1.0 — Multimodal Reliability
MFQ Studio 0.1.0 — Multimodal Reliability
This Apple Silicon build focuses on reliable MiniCPM-o-4_5 multimodal conversations in MFQ Studio.
Highlights:
- Adds native recorded-audio input alongside text, image, and video input.
- Hardens image/audio tensor validation and transport for Metal and CUDA runtimes.
- Resumes interrupted Token2Wav component downloads instead of restarting them.
- Fixes runtime-metrics growth and reduces packaged-server log noise, with log rotation.
- Includes regression coverage for audio preprocessing, resumable downloads, telemetry retention, and packaged server options.
The packaged app was exercised with text, image, video, audio, and combined image-plus-audio prompts, as well as common session, role, sampling, edit/regenerate, and stop workflows.
The base DMG does not bundle Token2Wav. MFQ Studio offers the voice-output component when full-duplex mode is enabled, keeping the default download smaller.
Build information:
- Platform: macOS on Apple Silicon, Windows with CUDA
- Commit:
89156a15f75516e9043086c3ac63aa4a1380e6f8 - Tag:
studio-v0.1.0-multimodal-20260828 - SHA-256:
044e39f54030ae055c312aca2df382ba418f2e2a1ea6733038a2a1f08dfc9645
This preview build is ad-hoc signed and is not Apple-notarized. On first launch, macOS may require Control-clicking the app and choosing Open.
MFQ Studio v0.1.0 Preview
MFQ Studio v0.1.0 preview installer for Mac OS and Windows 10