Skip to content

Releases: mannysinghx/LLMario

LLMario 0.2.0: now on Windows

Choose a tag to compare

@mannysinghx mannysinghx released this 29 Sep 23:11
96b8ad9

LLMario now runs on Windows (preview), alongside macOS. LLMario is a free, open-source app for chatting with open models (Qwen, Gemma, Llama, gpt-oss, Mistral and more) on your own computer. Nothing you type leaves your machine.

Downloads

File
macOS 13+ (Apple Silicon and Intel) LLMario-0.2.0-macos-universal.dmg Open it and drag LLMario into Applications. Signed and notarized by Apple.
Windows 10/11 x64: app LLMario-0.2.0-windows-x64-setup.exe Installs for your user account; no admin rights needed.
Windows 10/11 x64: command line llmario-0.2.0-windows-x64.zip llmario.exe (CLI and local OpenAI-compatible API). Unzip anywhere.

Each file has a .sha256 next to it. To check a download:

  • macOS / Linux: shasum -a 256 -c <file>.sha256
  • Windows (PowerShell): Get-FileHash <file> and compare with the .sha256 file

Windows preview: SmartScreen warning

The Windows files are not code-signed yet. The first time you run the installer, Windows may show "Windows protected your PC". Click More info → Run anyway. Code signing is planned.

Before you start: install an engine

LLMario runs models with the open-source engine llama.cpp (and MLX-LM on Apple Silicon Macs). The app does not include an engine yet, so install one:

  • Windows: winget install ggml.llamacpp in a terminal, then restart LLMario
  • macOS: brew install llama.cpp, and optionally pip3 install mlx-lm on Apple Silicon (usually fastest)

The app's Models → Engines tab shows what it found.

What's new in 0.2.0

  • Windows support for the desktop app and the llmario command line.
    • Engines start without console windows and stop together with LLMario, even if it crashes.
    • The Windows app includes the same features as on macOS: the model library, the memory check before loading, verified downloads, and drag-and-drop of your own .gguf files.
  • Engine detection recognizes more llama.cpp builds when checking which model architectures an installed engine supports.
  • macOS: no feature changes from 0.1.0.

Status

  • Windows (preview): built and tested automatically on Windows (install and launch of the app, a real chat through llama.cpp, clean engine shutdown). Not yet used day-to-day on real Windows PCs, so please report issues. GPU acceleration uses NVIDIA GPUs; on AMD and Intel GPUs models run on the CPU for now.
  • macOS: tested end to end on Apple Silicon (M4 Max) with both engines. The Intel part of the universal build is untested on real Intel hardware.

Website: https://llmario.com · Source: https://github.com/mannysinghx/LLMario

LLMario 0.1.0

Choose a tag to compare

@mannysinghx mannysinghx released this 29 Sep 18:50

First release of LLMario: a free, open-source Mac app for chatting with open models (Qwen, Gemma, Llama, gpt-oss, Mistral and more) on your own computer. Nothing you type leaves your Mac.

Download

LLMario-0.1.0-macos-universal.dmg: open it and drag LLMario into Applications. The app is signed with a Developer ID and notarized by Apple, so it opens without warnings.

To verify the download, compare its SHA-256 with the .sha256 file:

shasum -a 256 -c LLMario-0.1.0-macos-universal.dmg.sha256

SHA-256: 572860c9400d11eb432228563f78ed9b45d9f4a0a2dfa9605d3c93e1ef51e183

Before you start: install an engine

LLMario runs models with the open-source engines llama.cpp and MLX-LM. The app does not include them yet, so install at least one:

  • llama.cpp (any Mac): brew install llama.cpp
  • MLX-LM (Apple Silicon, usually fastest): pip3 install mlx-lm

The app's Models → Engines tab shows which engines it found.

What's in 0.1.0

  • Chat with streaming replies, Markdown and code blocks, a collapsible Thinking section for reasoning models, and a Stop button that actually cancels generation.
  • Hardware-aware: picks the right engine for your Mac and checks that a model fits in memory before loading it.
  • Model library: 37 open-model families and 69 verified downloads, with search, task filters and "fits this computer". Every file is checked against Hugging Face checksums and pinned to an exact version.
  • Bring your own models: drag a .gguf file or an MLX model folder onto the window.
  • Private by default: works offline once a model is downloaded, has no accounts and no telemetry, and never writes messages to logs.
  • CLI and local API: build from source for the llmario command and an OpenAI-compatible API on 127.0.0.1 (see the README).

Requirements and status

  • macOS 13 or later.
  • Tested end to end on Apple Silicon (M4 Max) with both engines.
  • Intel Macs: included in the universal build (llama.cpp only). Not yet tested on real Intel hardware.

Website: https://llmario.com · Source: https://github.com/mannysinghx/LLMario · Report issues: https://github.com/mannysinghx/LLMario/issues