Skip to content

LLMario 0.2.0: now on Windows

Latest

Choose a tag to compare

@mannysinghx mannysinghx released this 29 Sep 23:11
96b8ad9

LLMario now runs on Windows (preview), alongside macOS. LLMario is a free, open-source app for chatting with open models (Qwen, Gemma, Llama, gpt-oss, Mistral and more) on your own computer. Nothing you type leaves your machine.

Downloads

File
macOS 13+ (Apple Silicon and Intel) LLMario-0.2.0-macos-universal.dmg Open it and drag LLMario into Applications. Signed and notarized by Apple.
Windows 10/11 x64: app LLMario-0.2.0-windows-x64-setup.exe Installs for your user account; no admin rights needed.
Windows 10/11 x64: command line llmario-0.2.0-windows-x64.zip llmario.exe (CLI and local OpenAI-compatible API). Unzip anywhere.

Each file has a .sha256 next to it. To check a download:

  • macOS / Linux: shasum -a 256 -c <file>.sha256
  • Windows (PowerShell): Get-FileHash <file> and compare with the .sha256 file

Windows preview: SmartScreen warning

The Windows files are not code-signed yet. The first time you run the installer, Windows may show "Windows protected your PC". Click More info → Run anyway. Code signing is planned.

Before you start: install an engine

LLMario runs models with the open-source engine llama.cpp (and MLX-LM on Apple Silicon Macs). The app does not include an engine yet, so install one:

  • Windows: winget install ggml.llamacpp in a terminal, then restart LLMario
  • macOS: brew install llama.cpp, and optionally pip3 install mlx-lm on Apple Silicon (usually fastest)

The app's Models → Engines tab shows what it found.

What's new in 0.2.0

  • Windows support for the desktop app and the llmario command line.
    • Engines start without console windows and stop together with LLMario, even if it crashes.
    • The Windows app includes the same features as on macOS: the model library, the memory check before loading, verified downloads, and drag-and-drop of your own .gguf files.
  • Engine detection recognizes more llama.cpp builds when checking which model architectures an installed engine supports.
  • macOS: no feature changes from 0.1.0.

Status

  • Windows (preview): built and tested automatically on Windows (install and launch of the app, a real chat through llama.cpp, clean engine shutdown). Not yet used day-to-day on real Windows PCs, so please report issues. GPU acceleration uses NVIDIA GPUs; on AMD and Intel GPUs models run on the CPU for now.
  • macOS: tested end to end on Apple Silicon (M4 Max) with both engines. The Intel part of the universal build is untested on real Intel hardware.

Website: https://llmario.com · Source: https://github.com/mannysinghx/LLMario