Zyquo Local 1.0.0 — first public release 🎉
Run large language models 100 % locally on your Apple Silicon Mac. No API keys, no cloud, no data ever leaving your machine — powered by Apple's MLX.
✨ Highlights
- Private local inference — true token streaming with live tok/s, time-to-first-token and token-count stats under every response; cancellation that actually stops the GPU loop; multi-turn KV-cache reuse with automatic context-window management
- Browse & download models in-app — live Hugging Face search, a hand-curated Featured catalog of 30 verified models, and an industrial download manager (pause/resume across restarts via HTTP Range, retry on network drops, size verification, disk pre-check)
- RAM verdicts for your Mac — every model is stamped Fits / Tight / Too large before you download it
- Reasoning display — DeepSeek-R1 distills and Qwen3 thinking mode stream into a collapsible Thought process section
- Power tools — Quick Chat (global ⌥Space), two-model Compare mode, 56 prompt templates, personas, Markdown/PDF export, menu bar extra
- 57 supported architectures — Llama, Qwen 2/3/3.5/3.6, Mistral, Gemma 1–4, Phi-3/4, gpt-oss, GLM-4, SmolLM3, R1 distills, and more
- Verified — every capability validated end-to-end on real models (see docs/VERIFICATION.md): up to 222 tok/s on Llama-3.2-1B, memory verifiably released on unload
📦 Installation
- Download
ZyquoLocal.dmgbelow - Open it and drag Zyquo Local into Applications
- Launch — the app is Developer ID signed, notarized by Apple, and stapled (no Gatekeeper warnings)
- Pick a starter model sized for your Mac and start chatting — fully offline afterwards
🧰 Requirements
- Apple Silicon Mac (M1 or later) — Intel is not supported (MLX requirement)
- macOS 14.0+
- 8 GB RAM runs ≤4B models · 16 GB → 7–14B · 32 GB → 24–32B · 64 GB → 70B