Releases: 0xZKnw/mlxl3
Release list
MLXL3 Desktop v1.1.1 — One-click DFlash2 setup
MLXL3 Desktop v1.1.1 — one-click DFlash2 setup
For the Qwen3.6-35B-A3B EXL3 target, Generation → DFlash2 now installs
Inco AI's separate draft and applies the required greedy settings with one
switch. The app downloads only seven pinned draft/ files (~457 MiB), checks
their SHA-256 hashes and package structure, and resumes a partial download on
retry. An existing local draft can still be selected manually. The switch can
also cancel a download in progress.
DFlash2 remains experimental, opt-in and limited to this Qwen target. The DMG
contains the SwiftUI app, Rust/Metal engine and MLX runtime, not model or
draft weights. Speedups depend on the prompt and hardware; the first request
may compile Metal kernels. See the README
for setup, measurements and limitations.
Signing limitation: this build is ad-hoc signed, not Apple Developer ID
signed or notarized. macOS may require explicit approval in Privacy & Security.
Do not disable Gatekeeper globally; keep the previous app for rollback.
MLXL3 Desktop v1.1.0 — DFlash2 for Qwen3.6-35B-A3B
MLXL3 Desktop v1.1.0 — DFlash2 for Qwen3.6-35B-A3B
This release adds experimental, opt-in DFlash2 speculative decoding to
MLXL3 Desktop for Qwen3.6-35B-A3B EXL3. The native Rust/Metal engine proposes
tokens with Inco AI's draft and verifies them against the loaded target before
streaming them to the chat. The ordinary decoder remains the default.
In Generation → DFlash2, choose the folder containing the separate
draft/ weights and apply the greedy preset. DFlash2 currently requires greedy
sampling (temperature 0 or Top K 1, repetition penalty 1.0). Leave DFlash2 off
for other models or sampling settings; incompatible requests fail visibly rather
than silently changing the chosen sampling behavior.
The previously reported DFlash2 speedups were physical M5 benchmark results,
not promises for Desktop, other prompts or other Macs. The DMG includes the
SwiftUI app, Rust/Metal engine and MLX runtime, not the Qwen target or draft
weights. The first speculative request can be slower while Metal compiles its
kernels. See the README for setup and limitations.
Signing limitation: this build is ad-hoc signed, not Apple Developer ID
signed or notarized. macOS may require explicit approval in Privacy & Security.
Do not disable Gatekeeper globally; keep the previous app for rollback.
MLXL3 Desktop v1.0.3 — Ling 3.0 MCP compatibility
MLXL3 Desktop v1.0.3 — Ling 3.0 MCP compatibility
This is the release that makes Ling 3.0 Tiny's native tool-call format work
with MCP in MLXL3 Desktop. Ling could already generate text in v1.0.2, but its
<tool_call> output was not understood by the Rust bridge, so Exa and other
configured MCP tools were not called.
- The Rust bridge now parses Ling's
<arg_key>/<arg_value>arguments,
including JSON numbers, booleans and objects, while retaining the existing
Gemma, Qwen and LFM tool-call formats. - Malformed, duplicate or mixed prose/tool payloads remain non-executable.
- No model weights or quantization recipe changed in this release.
The DMG includes the native SwiftUI app, Rust/Metal engine and MLX runtime,
but not model weights. It requires Apple Silicon and macOS 26.2 or later.
Signing limitation: the app is ad-hoc signed, not Apple Developer ID signed
or notarized. macOS may require explicit approval in Privacy & Security. Do not
disable Gatekeeper globally; keep the previous app for rollback.
MLXL3 Desktop v1.0.2
MLXL3 Desktop v1.0.2
- Much faster native EXL3 checkpoint loading, especially for large MoE models.
- Batched TensorOps prefill for dense LFM2 models and Gemma 4.
- Keeps the previous serial path for LFM2 MoE and for Gemma contexts beyond its
validated sliding-attention batch window. - Includes the current Rust/Metal engine directly in the SwiftUI app and CLI.
On the tested M5, the fixed LFM2 dense workloads improved from 92 to 2,060 tok/s
and from 47 to 837 tok/s prefill, with identical greedy output hashes. Gemma 4
improved from 40 to 142 tok/s on the 82-token workload and from 36 to 127 tok/s
on 565 tokens, also with identical hashes. These measurements are specific to
the listed models, prompts and Mac; no decode improvement is claimed.
The DMG includes the native app, Rust engine and MLX runtime, not model
weights. It requires Apple Silicon and macOS 26.2 or later.
Signing limitation: the app is ad-hoc signed, not Apple Developer ID signed
or notarized. macOS may require explicit approval in Privacy & Security. Do not
disable Gatekeeper globally; keep the previous app for rollback.
MLXL3 Desktop v1.0.1
MLXL3 Desktop v1.0.1
- Updated the engine bundled with Desktop to the current published MLXL3 engine.
- Fixed the macOS 27 TensorOps compilation error that prevented Qwen from loading
(get_destination_cooperative_tensor). - Faster exact EXL3 MoE prefill on M5: larger row tiles and reusable weight-address
calculations, with the existing path retained for smaller workloads. - Includes the engine improvements published since v1.0.0 and a release smoke
test that now keeps the real model warmup enabled.
On the tested M5, paired 4k-token runs measured +12.4% prefill for Qwen3.6 35B A3B
and +34.9% for LFM2.5 8B A1B. A separate non-repetitive document measured +9.4%
and +25.6%, respectively. Text/cache parity checks passed. These are workload-
specific measurements, not universal speedups or a comparison against llama.cpp.
No decode or RAM improvement is claimed. See the
full performance report.
The DMG contains the native app, engine and dependencies, not model weights.
Requires Apple Silicon and macOS 26.2+. Physical validation is on M5;
other Mac generations and every supported architecture have not been retested.
Signing limitation: ad-hoc signed, not Apple Developer ID signed or notarized,
as in previous releases. macOS may require explicit approval under Privacy &
Security. Do not disable Gatekeeper globally. Keep the prior app for rollback.
MLXL3 Desktop v1.0.0
MLXL3 Desktop v1.0.0
Reliability and polish across the Metal engine, terminal chat and native Mac app.
- Protected chat saves, recovery/export/import, reliable cancellation and engine lifecycle.
- Safer MCP parsing, schema validation, bounded results and preserved tool context.
- Checkpoint/shard validation, cache recovery and model-memory cleanup.
- Smoother thinking/answer streaming, long-code layout, improved Markdown/math and per-chat drafts.
- Resumable variant-only HF downloads, folder relocation and HF sign-in in Settings.
- Build-aware updates, retained app backup and bundled dependency/source manifest.
The DMG includes the engine and its dependencies, not model weights.
Requires Apple Silicon, macOS 26.2+. CLI and Desktop share the same engine.
Distribution limitation: ad-hoc signed, not Apple Developer ID signed or
notarized, as in earlier releases. macOS may require explicit approval in
Privacy & Security. Do not disable Gatekeeper globally.
See validation, compatibility and remaining limits.
No universal model/device certification or new performance gain is claimed.
MLXL3 Desktop v0.4.6
MLXL3 Desktop v0.4.6
Download hotfix (same version)
- Download progress is delivered immediately instead of waiting for a full
16 KiB pipe buffer. Hugging Face Xet progress callbacks no longer fail on
disabled terminal progress bars. Existing partial downloads remain resumable. - The DMG and checksum were replaced with this corrected build. If you already
have v0.4.6, download the DMG again manually: the updater compares version
numbers and will not offer another v0.4.6.
Correct Metal memory accounting
- The menu-bar widget now reports the macOS physical memory footprint of
the engine and interface, usingproc_pid_rusage / ri_phys_footprintinstead
of RSS. GPU allocations are included: Qwen no longer appears to occupy only
a few hundred megabytes when its Metal buffers use roughly 12 GB. - Unavailable measurements display a dash instead of a misleading zero.
- The widget distinguishes model loading from the ready state.
- This is an accounting fix, not a reduction in memory use or a change to model
weights, inference kernels or generation quality.
Model library and refreshed interface
This release also ships the v0.4.5 work that had only been installed locally:
- Dark/ivory interface, serif headings, redesigned menu-bar panel and monochrome icons.
- Installed-model library: load, reveal in Finder, remove an entry or move a model
folder to Trash after confirmation. - Hugging Face EXL3 search, native Markdown model cards, branch/tag and variant
selection with download size. - Selective downloads pinned to a commit, progress, pause/resume and disk-space
checks. Shared tokenizer files are inherited; ambiguous weight layouts are
refused rather than mixed. - Large catalogue responses no longer block the CLI output pipe.
Installation
Use Settings → Check for updates, then restart when the download is ready;
or download the Apple Silicon DMG, open it and drag MLXL3 Desktop to Applications.
Requires an Apple Silicon Mac and macOS 26.2+. The engine/runtime are included;
model weights are downloaded separately. The app remains ad-hoc signed, not
Apple-notarized.
Validation includes the Python suite and packaged native checks, including a
64 MiB GPU-private Metal allocation that must increase the reported footprint.
MLXL3 Desktop v0.4.2
MLXL3 Desktop v0.4.2
Contexte et mémoire
- Compteur de tokens utilisé / limite dans la barre de chat, actualisé pendant
la génération. Il inclut le prompt, les résultats MCP et la génération ;
le brouillon non envoyé n'est pas inclus. - Limite de contexte réglable par modèle, avec préréglages et bouton
Enregistrer et recharger le modèle. Le réglage est conservé au redémarrage. - La limite est appliquée par le moteur : les prompts trop longs sont refusés
avant le calcul GPU, sans supprimer silencieusement l'historique. - Estimation modèle + contexte en Go, recalculée immédiatement pendant
la saisie. Elle utilise les caches réellement créés au chargement : dimensions,
précision, fenêtres glissantes et états récurrents.
L'estimation correspond à une conversation à contexte plein, pas au pic RAM
total. Les buffers de calcul, les autres conversations en cache et macOS
nécessitent une marge supplémentaire. Pour un cache inconnu ou sans warmup,
l'app indique que l'estimation est indisponible.
Français / English
Le choix de langue est disponible dans Réglages → Langue de l'app, à côté
des mises à jour. Il est conservé au redémarrage et ne traduit ni ne modifie
le contenu des conversations.
Installation
Télécharger le DMG ci-dessous et glisser MLXL3 Desktop dans Applications,
ou utiliser les réglages de mise à jour de l'app. Le moteur Python/MLX est inclus ;
les modèles sont téléchargés séparément.
- Mac Apple Silicon, macOS 26.2 ou plus récent.
- Application signée ad hoc, non notarisée : macOS peut demander une
autorisation à la première ouverture. - Exa reste préconfiguré, désactivé par défaut ; les choix MCP existants sont conservés.
Validation : 349 tests Python réussis, 2 tests ignorés faute de modèles
locaux ; contrôles Swift de chronologie, persistance, langue, contexte et mémoire ;
vérification visuelle FR/EN et du recalcul mémoire ; génération réelle sur Qwen.
Un fichier SHA-256 accompagne le DMG.
MLXL3 Desktop v0.4.1
MLXL3 Desktop v0.4.1
Un parcours MCP plus clair
- Les réflexions, appels outils et réponses s'affichent dans leur ordre réel.
La réflexion qui suit Exa apparaît désormais sous l'appel terminé. - Le traitement des résultats affiche un statut, le nombre de tokens nouveaux
et réutilisés, ainsi que le temps d'attente. - Cette chronologie est conservée à la réouverture des nouveaux échanges.
Les anciennes conversations restent lisibles, mais leur ordre détaillé
n'avait pas été enregistré et ne peut pas être reconstruit.
Prefill EXL3
Un nouveau kernel SIMD fusionne le gather, les scales et la transformation
Hadamard du prefill MoE SwiGLU. Sur Qwen3.6-35B-A3B EXL3 2.49 bpw, la dernière
paire contrôlée donne +5,7 % de débit prefill, avec une attente après outil
de 6,48 à 6,13 secondes pour 3 742 tokens évalués.
Les logits finaux et états du cache sont identiques dans la continuation de
contrôle. Aucun speculative decoding, changement de quantification ou résultat
MCP tronqué n'est introduit. Le chemin Gemma/GeGLU reste inchangé.
Limites : 20 essais sur un Mac M5 24 Gio sur batterie, avec variations de régime.
Ce n'est pas un gain universel ni un gain de decode. Le pic MLX augmente
d'environ 81 Mo sur cette charge. Les essais détaillés et variantes rejetées
sont documentés dans docs/tool-prefill-rd-2026-09-05.md.
Installation et mise à jour
Télécharger le DMG ci-dessous, puis glisser MLXL3 Desktop dans Applications.
Le moteur Python/MLX est inclus ; les poids des modèles sont téléchargés séparément.
La version est également détectable depuis les réglages de mise à jour de l'app.
- Mac Apple Silicon, macOS 26.2 ou plus récent.
- Application signée ad hoc, non notarisée : macOS peut demander une autorisation
à la première ouverture. - Exa reste préconfiguré ; le choix d'activation MCP reste mémorisé.
Validation : 347 tests Python réussis, 2 fixtures locales absentes ; contrôles
Swift de chronologie/persistance et préférences MCP ; vérification visuelle
du parcours après outil. Un fichier SHA256 accompagne le DMG.
MLXL3 Desktop v0.4.0
MLXL3 Desktop v0.4.0
Nouvelle interface
- Interface native repensée : palette noir et ivoire, navigation compacte et espace de lecture épuré.
- Messages utilisateur à droite, réponses du modèle à gauche.
- Historique recherchable et repliable, retour au dernier message et contrôles adaptés aux petites fenêtres.
- Réglages et gestionnaire de modèles harmonisés ; Markdown, code coloré, copie et LaTeX conservés.
Exa et interrupteur MCP
- Exa est préconfiguré dans le DMG : recherche web et lecture de pages, sans Node ni clé API pour l'offre gratuite d'Exa (limites applicables).
- Interrupteur MCP directement dans la barre de saisie, désactivé au premier lancement.
- Le choix est conservé entre les lancements et les conversations, jusqu'au prochain changement.
- Activation/désactivation sans recharger le modèle ; aucune connexion MCP tant que le mode reste désactivé.
- L'interrupteur est temporairement verrouillé pendant une génération ou une mise à jour de connexion.
- Connexion HTTPS native aux serveurs MCP Streamable HTTP (JSON/SSE), en plus des serveurs locaux stdio.
- Les recherches et URL envoyées à Exa quittent le Mac. L'inférence reste locale.
Installation
Mac Apple Silicon, macOS 26.2 ou ultérieur. Ouvrir le DMG et déplacer MLXL3 Desktop dans Applications, ou utiliser les mises à jour intégrées. Le moteur et l'interface sont inclus ; les poids des modèles se téléchargent séparément.
L'application est signée ad hoc, non notariée. Les instructions Gatekeeper habituelles du README restent applicables.
Vérifications
Tests ciblés du protocole MCP, du pont de génération et du registre ; autocontrôle natif de la préférence persistante ; essais visuels isolés et connexion Exa réelle. Pas de nouveau benchmark de débit ni de modification des kernels d'inférence dans cette version.