Cadabra.app 2.0 beta 1
Pre-releaseCadabra requires macOS 14.6 and an Apple Silicon mac (arm64).
Cadabra is a new app, not a new version of AIChat. AIChat 1.2 stays in the repo and keeps
working, but the chat, the engine layer and the model management are all new. Cadabra keeps its own
settings and history under ~/Library/Application Support/Cadabra and never writes to AIChat's. On
first launch it offers a one-time import of your AIChat 1.x chat history, and asks before taking
anything.
Native chat, no web view. The chat is a native macOS view (ActionUI Chat) talking to the model
over ACP. The bundled llama.cpp WebUI is gone. Streaming text, thinking, plans and tool calls all
render in the transcript.
Two engines, chosen by the model:
- MLX - an Apple Silicon model directory (
config.jsonplus safetensors shards), run by the
embedded mlx-agent - GGUF - a single
.gguffile, run byllama-serverfrom
llama.cpp - Apple Foundation Models - on-device, nothing to download, on macOS 26 and later
Picking a model is the whole configuration: a file is treated as GGUF, a directory with shards as
MLX, and both sit side by side in one list. Models switch in place from the chat toolbar.
Models
- Local Models browser with a RAM-fit advisory and a per-model benchmark, presented at launch
- Hugging Face browser: search, filter by Any / MLX / GGUF, download and start
- delete downloaded models from inside the app
- run several models at once, one per window, with the GPU memory budget measured across them and
offloaded to CPU rather than pushed until something dies
Conversations
- named, persisted conversations you can rename, reveal, delete and continue
- resuming a long conversation can summarize its older half instead of replaying all of it, with
the summary shown in the transcript rather than hidden - if the digest missed something that
matters, you can supply it in the chat - you choose the summarizer: the model this chat is already running, Apple Intelligence
(macOS 26 and later), orautoto let the agent decide by measuring - the transcript records when a conversation was started or resumed and which model answered, so a
conversation continued with a different model still says so
Agentic tools (MCP), owned by mlx-agent over stdio, with a servers dialog and an inspector:
- Time
- Web Search and Fetch
- PDF, via the embedded
pdfutiland no network: inspection always, editing optional and
permission-gated per call, with outputs always written as new files rather than overwrites - Local files and shell, sandboxed via replay, with explicit
read-only and read-write path lists, rooted at the Project workspace you pick - a session-wide Allow Network switch that hard-gates the networked servers
External agents
Cadabra can drive any local ACP-speaking agent, not just the bundled one - saved by name, with a
choice of how much of the MCP tooling it is handed: nothing, the read-only servers, or all of
them.
Known limitations
- Apple Foundation Models runs in a 4k context window. Enabling MCP servers spends a large part of
it on tool definitions before the conversation starts, so tool use with Foundation Models is
practical only with very few servers enabled.
This is a beta: please report what breaks.