Cadabra 2.0
Requirements: macOS 14.6 or later (Apple silicon)
Cadabra 2.0 is the first release of Cadabra, the successor to AIChat. It runs large language models locally on your Mac in a native chat window, with agentic tools and no web view. This release sums up the three 2.0 betas and adds the changes made since beta 3.
Cadabra is a new app, not a new version of AIChat. AIChat 1.2 stays in the repo and keeps working. Cadabra keeps its own settings and history under ~/Library/Application Support/Cadabra and never writes to AIChat's. On first launch it offers a one-time import of your AIChat 1.x chat history, and asks before taking anything.
Highlights
Native chat
- The chat is a native macOS view talking to the model over ACP (Agent Client Protocol). Streaming text, thinking, plans and tool calls all render in the transcript.
- The bundled llama.cpp WebUI is gone.
Three engines, chosen by the model
- MLX - an Apple silicon model folder (
config.jsonplus safetensors shards), run by the embedded mlx-agent. - GGUF - a single
.gguffile, run byllama-serverfrom llama.cpp. - Apple Foundation Models - on-device, nothing to download, on macOS 26 and later.
Picking a model is the whole configuration. GGUF and MLX models sit side by side in one list, and you can switch models in place from the chat toolbar.
Models
- Local Models browser with a RAM-fit advisory and a per-model benchmark.
- Download Models: search Hugging Face, filter by Any / MLX / GGUF, download and start.
- A first run with no models installed opens the download browser instead of an empty picker.
- Downloads are staged in a
.partfile and moved into place only when complete, so an interrupted download can no longer show up as a broken model. Interrupted downloads resume where they stopped, and Cancel still cleans up. - Delete downloaded models from inside the app.
- Run several models at once, one per window. GPU memory is budgeted across them, and a model offloads to the CPU rather than running out of memory.
Conversations
- Named, saved conversations you can rename, reveal in Finder, delete and continue.
- Resuming a long conversation can summarize its older half instead of replaying all of it. The summary is shown in the transcript, not hidden.
- You choose the summarizer: the model the chat is already running, Apple Intelligence (macOS 26 and later), or
auto. - The transcript records when a conversation was started or resumed and which model answered.
- Search across all conversations from the conversations window toolbar, or within one conversation with Cmd-F.
Agentic tools (MCP)
Tools are served to the model through MCP (Model Context Protocol), with a servers dialog and an inspector:
- Time
- Web search and fetch
- PDF, through the embedded
pdfutilwith no network access: inspection always, editing optional and approved per call, with results always written to new files - Local files and shell, sandboxed by
replay, with explicit read-only and read-write path lists rooted at the project folder you pick - An Allow Network switch that turns off the networked servers for the whole session
External agents
Cadabra can drive any local agent that speaks ACP, not just the bundled one. Agents are saved by name, and for each one you choose how much of the MCP tooling it gets: nothing, the read-only servers, or all of them.
Changes since 2.0 beta 3
- Search in a conversation gains a Regular Expression option in the find bar menu. A pattern that does not compile reports "Invalid expression" instead of "No matches".
- Search highlights now update right away as you type, press Cmd-G, or close the bar. Before, they did not appear until the transcript scrolled.
- App engine updated to OMC 5.3.
- llama.cpp updated to v0.4.1.
replaysandbox tool updated to 2.2.1.- Bundled Python packages updated, including mcp 1.30.0, uvicorn 0.53.0, anyio 4.15.1, PyJWT 2.14.0 and tzdata 2026.4.
- The embedded Python 3.14.7 is trimmed to what Cadabra uses, which makes the app smaller on disk (about 155 MB, down from 169 MB).
- The About box credits and the third-party license notices are complete.
Upgrading
From a 2.0 beta: replace the old Cadabra.app with the new one. Downloaded models, settings, and chat history are untouched.
From 2.0 beta 1 or earlier: a GGUF download interrupted under those versions may have left a truncated .gguf file under its real name. If a model fails to load, download it again from the model browser. Cadabra compares the file on disk with the size the server reports, so it downloads a short file again instead of treating it as already downloaded.
From AIChat 1.x: Cadabra installs alongside AIChat and offers to import your chat history on first launch.
Known limitations
Apple Foundation Models has a 4k-token context window. Enabled MCP servers use up a large part of it with tool definitions before the conversation starts, so tools are practical with Foundation Models only when very few servers are enabled.