Skip to content

Releases: abra-code/Cadabra

Cadabra 2.0.1

Choose a tag to compare

@abra-code abra-code released this 19 Sep 08:49

Requirements: macOS 14.6 or later (Apple silicon)

Cadabra 2.0.1 is the bugfix release for macOS 27 compatibility.
It is updated with OMC 5.3.1 engine with ActionUI 0.8.3
Also bundles updated pdfutil 0.3 tool.

Cadabra 2.0

Choose a tag to compare

@abra-code abra-code released this 16 Sep 11:48

Requirements: macOS 14.6 or later (Apple silicon)

Cadabra 2.0 is the first release of Cadabra, the successor to AIChat. It runs large language models locally on your Mac in a native chat window, with agentic tools and no web view. This release sums up the three 2.0 betas and adds the changes made since beta 3.

Cadabra is a new app, not a new version of AIChat. AIChat 1.2 stays in the repo and keeps working. Cadabra keeps its own settings and history under ~/Library/Application Support/Cadabra and never writes to AIChat's. On first launch it offers a one-time import of your AIChat 1.x chat history, and asks before taking anything.

Highlights

Native chat

  • The chat is a native macOS view talking to the model over ACP (Agent Client Protocol). Streaming text, thinking, plans and tool calls all render in the transcript.
  • The bundled llama.cpp WebUI is gone.

Three engines, chosen by the model

  • MLX - an Apple silicon model folder (config.json plus safetensors shards), run by the embedded mlx-agent.
  • GGUF - a single .gguf file, run by llama-server from llama.cpp.
  • Apple Foundation Models - on-device, nothing to download, on macOS 26 and later.

Picking a model is the whole configuration. GGUF and MLX models sit side by side in one list, and you can switch models in place from the chat toolbar.

Models

  • Local Models browser with a RAM-fit advisory and a per-model benchmark.
  • Download Models: search Hugging Face, filter by Any / MLX / GGUF, download and start.
  • A first run with no models installed opens the download browser instead of an empty picker.
  • Downloads are staged in a .part file and moved into place only when complete, so an interrupted download can no longer show up as a broken model. Interrupted downloads resume where they stopped, and Cancel still cleans up.
  • Delete downloaded models from inside the app.
  • Run several models at once, one per window. GPU memory is budgeted across them, and a model offloads to the CPU rather than running out of memory.

Conversations

  • Named, saved conversations you can rename, reveal in Finder, delete and continue.
  • Resuming a long conversation can summarize its older half instead of replaying all of it. The summary is shown in the transcript, not hidden.
  • You choose the summarizer: the model the chat is already running, Apple Intelligence (macOS 26 and later), or auto.
  • The transcript records when a conversation was started or resumed and which model answered.
  • Search across all conversations from the conversations window toolbar, or within one conversation with Cmd-F.

Agentic tools (MCP)

Tools are served to the model through MCP (Model Context Protocol), with a servers dialog and an inspector:

  • Time
  • Web search and fetch
  • PDF, through the embedded pdfutil with no network access: inspection always, editing optional and approved per call, with results always written to new files
  • Local files and shell, sandboxed by replay, with explicit read-only and read-write path lists rooted at the project folder you pick
  • An Allow Network switch that turns off the networked servers for the whole session

External agents

Cadabra can drive any local agent that speaks ACP, not just the bundled one. Agents are saved by name, and for each one you choose how much of the MCP tooling it gets: nothing, the read-only servers, or all of them.

Changes since 2.0 beta 3

  • Search in a conversation gains a Regular Expression option in the find bar menu. A pattern that does not compile reports "Invalid expression" instead of "No matches".
  • Search highlights now update right away as you type, press Cmd-G, or close the bar. Before, they did not appear until the transcript scrolled.
  • App engine updated to OMC 5.3.
  • llama.cpp updated to v0.4.1.
  • replay sandbox tool updated to 2.2.1.
  • Bundled Python packages updated, including mcp 1.30.0, uvicorn 0.53.0, anyio 4.15.1, PyJWT 2.14.0 and tzdata 2026.4.
  • The embedded Python 3.14.7 is trimmed to what Cadabra uses, which makes the app smaller on disk (about 155 MB, down from 169 MB).
  • The About box credits and the third-party license notices are complete.

Upgrading

From a 2.0 beta: replace the old Cadabra.app with the new one. Downloaded models, settings, and chat history are untouched.

From 2.0 beta 1 or earlier: a GGUF download interrupted under those versions may have left a truncated .gguf file under its real name. If a model fails to load, download it again from the model browser. Cadabra compares the file on disk with the size the server reports, so it downloads a short file again instead of treating it as already downloaded.

From AIChat 1.x: Cadabra installs alongside AIChat and offers to import your chat history on first launch.

Known limitations

Apple Foundation Models has a 4k-token context window. Enabled MCP servers use up a large part of it with tool definitions before the conversation starts, so tools are practical with Foundation Models only when very few servers are enabled.

Cadabra.app 2.0 beta 3

Pre-release

Choose a tag to compare

@abra-code abra-code released this 05 Sep 00:55

Requires: macOS 14.6 or later (Apple silicon)

App engine upgraded to employ OMC 5.2.1, including Python 3.14.7
Latest llama.cpp v0.4.0 included.

Added search in all conversations (search field in conversations window toolbar) and search in one conversation - use cmd+F to open the bar for the chat view.

Cadabra.app 2.0 beta 2

Pre-release

Choose a tag to compare

@abra-code abra-code released this 30 Aug 09:26

Requires: macOS 14.6 or later (Apple silicon)

Summary: This beta is about the first ten minutes with the app: finding a model, downloading one, and never
being handed a broken one.

First run opens the model download browser

Launching Cadabra with no model installed used to open the Local Models picker on an empty table,
with nothing on screen saying where models come from. Cadabra now checks whether any model is
installed and opens the Hugging Face browser instead when the answer is none - downloading a model
being the only useful move at that point. Closing that browser hands back to the model picker, so a
first run no longer ends with the download finished and no window on screen.

Apple's on-device foundation model does not count as installed for this check: it needs no download
and is present on every eligible Mac, so counting it would retire the first-run branch on exactly
the machines it exists for.

Launch and the picker now read the same list of folders when they scan for models, so the two can
no longer disagree about what "no models installed" means.

"Download Models" instead of "Browse Hugging Face"

Hugging Face is not a household name, and someone new to running local models has no reason to know
what it is. The entry points now say Download Models - in the main window and in the Local
Models picker - while the browser window itself still names Hugging Face as the source it is
searching ("Download Models - Hugging Face").

Interrupted downloads can no longer masquerade as models

Downloads used to be written straight to the model's final name. Any death the app could not clean
up after - a force quit, a kernel panic, a power cut - left a truncated file sitting at exactly the
name the model picker scans for. The picker listed it, showed its size, and handed it to a loader
that could only fail.

Every transfer now lands in a <final>.part staging file and is moved into place only once it is
complete and the size the server declared. A staging name matches neither the *.gguf scan nor the
MLX config.json / *.safetensors shape, so an interrupted download is invisible to the picker by
construction.

What this means in practice:

  • Interrupted downloads resume. Whatever arrived is kept, and the next Download continues from
    there rather than starting over. If the server refuses to resume, Cadabra spends that one attempt
    and restarts the file from the top.
  • Quitting mid-download no longer throws away the bytes. Quit used to delete the partial file,
    because leaving it meant a corrupt model in the picker on the next launch. It is safe to keep now,
    so however many gigabytes had arrived survive the quit.
  • Cancel still cleans up. Pressing Cancel removes the staging file and its bookkeeping sidecar.
  • An already-downloaded file is now checked by size. A local file that is not the length the
    server declares is no longer reported as cached - it is downloaded again. Pressing Download is
    what repairs a file a previous build left truncated.
  • Incomplete MLX model folders are detected. An MLX download interrupted between shards leaves
    config.json beside only some of its weights, which used to look like a whole model. When the
    folder has a model.safetensors.index.json, Cadabra now requires every shard that index names to
    be present before it will offer the model. Folders with no index are unchanged - there is nothing
    to be incomplete against.

Upgrading from 2.0b1

Drop the new Cadabra.app in place of the old one. Downloaded models, settings, and chat history are
untouched.

One thing worth knowing: a GGUF download that was interrupted under 2.0b1 or earlier may have
left a truncated .gguf file under its real name, and the picker still lists it as a model. If one
fails to load, download it again from the model browser - Cadabra compares what is on disk against
the size the server declares now, so a short file is re-fetched rather than reported as already
downloaded. No manual deletion needed. Incomplete MLX folders are caught by the new shard check
regardless of which version created them, and simply stop appearing in the picker.

Cadabra.app 2.0 beta 1

Pre-release

Choose a tag to compare

@abra-code abra-code released this 27 Aug 22:25

Cadabra requires macOS 14.6 and an Apple Silicon mac (arm64).

Cadabra is a new app, not a new version of AIChat. AIChat 1.2 stays in the repo and keeps
working, but the chat, the engine layer and the model management are all new. Cadabra keeps its own
settings and history under ~/Library/Application Support/Cadabra and never writes to AIChat's. On
first launch it offers a one-time import of your AIChat 1.x chat history, and asks before taking
anything.

Native chat, no web view. The chat is a native macOS view (ActionUI Chat) talking to the model
over ACP. The bundled llama.cpp WebUI is gone. Streaming text, thinking, plans and tool calls all
render in the transcript.

Two engines, chosen by the model:

  • MLX - an Apple Silicon model directory (config.json plus safetensors shards), run by the
    embedded mlx-agent
  • GGUF - a single .gguf file, run by llama-server from
    llama.cpp
  • Apple Foundation Models - on-device, nothing to download, on macOS 26 and later

Picking a model is the whole configuration: a file is treated as GGUF, a directory with shards as
MLX, and both sit side by side in one list. Models switch in place from the chat toolbar.

Models

  • Local Models browser with a RAM-fit advisory and a per-model benchmark, presented at launch
  • Hugging Face browser: search, filter by Any / MLX / GGUF, download and start
  • delete downloaded models from inside the app
  • run several models at once, one per window, with the GPU memory budget measured across them and
    offloaded to CPU rather than pushed until something dies

Conversations

  • named, persisted conversations you can rename, reveal, delete and continue
  • resuming a long conversation can summarize its older half instead of replaying all of it, with
    the summary shown in the transcript rather than hidden - if the digest missed something that
    matters, you can supply it in the chat
  • you choose the summarizer: the model this chat is already running, Apple Intelligence
    (macOS 26 and later), or auto to let the agent decide by measuring
  • the transcript records when a conversation was started or resumed and which model answered, so a
    conversation continued with a different model still says so

Agentic tools (MCP), owned by mlx-agent over stdio, with a servers dialog and an inspector:

  • Time
  • Web Search and Fetch
  • PDF, via the embedded pdfutil and no network: inspection always, editing optional and
    permission-gated per call, with outputs always written as new files rather than overwrites
  • Local files and shell, sandboxed via replay, with explicit
    read-only and read-write path lists, rooted at the Project workspace you pick
  • a session-wide Allow Network switch that hard-gates the networked servers

External agents

Cadabra can drive any local ACP-speaking agent, not just the bundled one - saved by name, with a
choice of how much of the MCP tooling it is handed: nothing, the read-only servers, or all of
them.

Known limitations

  • Apple Foundation Models runs in a 4k context window. Enabling MCP servers spends a large part of
    it on tool definitions before the conversation starts, so tool use with Foundation Models is
    practical only with very few servers enabled.

This is a beta: please report what breaks.

AIChat.app version 1.2

Choose a tag to compare

@abra-code abra-code released this 08 Jun 08:53

AIChat now requires macOS 14.6 and an Apple Silicon mac (arm64)

What is new in version 1.2:

  • out-of-the-box agentic tooling via bundled MCP servers:
  • MCP Servers inspector window
  • configurable, agentic-tools and sandbox setup before model launch
  • LLM model choice help
  • updated llama.cpp to b9553

AIChat.app version 1.1

Choose a tag to compare

@abra-code abra-code released this 18 May 06:47
  • Added local model selector dialog, trying to find GGUF model files you might have downloaded with HuggingFace or LM Studio.
  • Added a HuggingFace model browser and downloader.
  • Improved calculations for context size. Added settings for reduced kv cache size, making even more context size available per RAM.
  • Added 10 min inactivity setting to llama-server to unload the model and free the memory - now the app can be left open in the background and will not wire the RAM permanently
  • update llama-server to b9204

AIChat now requires macOS 14.6. The release is only for Apple Silicon macs (arm64). Intel mac support dropped - you cannot get a reasonable inference speeds on these machines.

AIChat.app version 1.0.2

Choose a tag to compare

@abra-code abra-code released this 14 Nov 00:28

Misc fixes:

  • dark theme fix
  • attempt to set a better context size depending on system RAM
  • update llama.cpp to b7020
  • improve loading pages
  • persist last window size and position

AIChat.app version 1.0.1

Choose a tag to compare

@abra-code abra-code released this 03 Nov 08:59

Improve handing of slow llama-server launch with longer timeout, especially for the first launch of the app.
Track server launch progress in the WebView and give user more information if something goes wrong.
No executable binaries have been changed since v 1.0.

AIChat.app version 1.0

Choose a tag to compare

@abra-code abra-code released this 31 Oct 10:22

Initial release
AIChat.app is an app for running large language models in GGUF format locally.
Embedded llama-server is taken from this release:
https://github.com/ggml-org/llama.cpp/releases/tag/b6895

Download the archive with architecture appropriate for your computer.
Running some bigger models on outdated Intel Macs could be a painful experience. Apple Silicon Macs fare much better.