Releases: mannysinghx/LLMario
Release list
LLMario 0.2.0: now on Windows
LLMario now runs on Windows (preview), alongside macOS. LLMario is a free, open-source app for chatting with open models (Qwen, Gemma, Llama, gpt-oss, Mistral and more) on your own computer. Nothing you type leaves your machine.
Downloads
| File | ||
|---|---|---|
| macOS 13+ (Apple Silicon and Intel) | LLMario-0.2.0-macos-universal.dmg |
Open it and drag LLMario into Applications. Signed and notarized by Apple. |
| Windows 10/11 x64: app | LLMario-0.2.0-windows-x64-setup.exe |
Installs for your user account; no admin rights needed. |
| Windows 10/11 x64: command line | llmario-0.2.0-windows-x64.zip |
llmario.exe (CLI and local OpenAI-compatible API). Unzip anywhere. |
Each file has a .sha256 next to it. To check a download:
- macOS / Linux:
shasum -a 256 -c <file>.sha256 - Windows (PowerShell):
Get-FileHash <file>and compare with the.sha256file
Windows preview: SmartScreen warning
The Windows files are not code-signed yet. The first time you run the installer, Windows may show "Windows protected your PC". Click More info → Run anyway. Code signing is planned.
Before you start: install an engine
LLMario runs models with the open-source engine llama.cpp (and MLX-LM on Apple Silicon Macs). The app does not include an engine yet, so install one:
- Windows:
winget install ggml.llamacppin a terminal, then restart LLMario - macOS:
brew install llama.cpp, and optionallypip3 install mlx-lmon Apple Silicon (usually fastest)
The app's Models → Engines tab shows what it found.
What's new in 0.2.0
- Windows support for the desktop app and the
llmariocommand line.- Engines start without console windows and stop together with LLMario, even if it crashes.
- The Windows app includes the same features as on macOS: the model library, the memory check before loading, verified downloads, and drag-and-drop of your own
.gguffiles.
- Engine detection recognizes more llama.cpp builds when checking which model architectures an installed engine supports.
- macOS: no feature changes from 0.1.0.
Status
- Windows (preview): built and tested automatically on Windows (install and launch of the app, a real chat through llama.cpp, clean engine shutdown). Not yet used day-to-day on real Windows PCs, so please report issues. GPU acceleration uses NVIDIA GPUs; on AMD and Intel GPUs models run on the CPU for now.
- macOS: tested end to end on Apple Silicon (M4 Max) with both engines. The Intel part of the universal build is untested on real Intel hardware.
Website: https://llmario.com · Source: https://github.com/mannysinghx/LLMario
LLMario 0.1.0
First release of LLMario: a free, open-source Mac app for chatting with open models (Qwen, Gemma, Llama, gpt-oss, Mistral and more) on your own computer. Nothing you type leaves your Mac.
Download
LLMario-0.1.0-macos-universal.dmg: open it and drag LLMario into Applications. The app is signed with a Developer ID and notarized by Apple, so it opens without warnings.
To verify the download, compare its SHA-256 with the .sha256 file:
shasum -a 256 -c LLMario-0.1.0-macos-universal.dmg.sha256
SHA-256: 572860c9400d11eb432228563f78ed9b45d9f4a0a2dfa9605d3c93e1ef51e183
Before you start: install an engine
LLMario runs models with the open-source engines llama.cpp and MLX-LM. The app does not include them yet, so install at least one:
- llama.cpp (any Mac):
brew install llama.cpp - MLX-LM (Apple Silicon, usually fastest):
pip3 install mlx-lm
The app's Models → Engines tab shows which engines it found.
What's in 0.1.0
- Chat with streaming replies, Markdown and code blocks, a collapsible Thinking section for reasoning models, and a Stop button that actually cancels generation.
- Hardware-aware: picks the right engine for your Mac and checks that a model fits in memory before loading it.
- Model library: 37 open-model families and 69 verified downloads, with search, task filters and "fits this computer". Every file is checked against Hugging Face checksums and pinned to an exact version.
- Bring your own models: drag a
.gguffile or an MLX model folder onto the window. - Private by default: works offline once a model is downloaded, has no accounts and no telemetry, and never writes messages to logs.
- CLI and local API: build from source for the
llmariocommand and an OpenAI-compatible API on127.0.0.1(see the README).
Requirements and status
- macOS 13 or later.
- Tested end to end on Apple Silicon (M4 Max) with both engines.
- Intel Macs: included in the universal build (llama.cpp only). Not yet tested on real Intel hardware.
Website: https://llmario.com · Source: https://github.com/mannysinghx/LLMario · Report issues: https://github.com/mannysinghx/LLMario/issues