Releases: adatoo/quail
Releases · adatoo/quail
Release list
Quail 0.58.0
Added
- The comparison harness can time the engines (ADR D-063):
task bench:speedruns GuideLLM against each engine with the same prompts. It covers 512 → 256 tokens at 1, 2, 4 and 8 requests at once, and 4096 → 128 tokens.- Each level records time to first token, time between tokens, throughput, peak memory (GPU buffers included) and power.
- Every engine gets exactly the same number of requests. The prompts are exact lengths in the model's own tokenizer, and none repeats, so no engine gets a cache hit another doesn't. The Mac cools down before each level, and a level that ends hot is run again.
task bench:nativegives each lane's own ceiling on the same weights:llama-benchandllama-batched-benchfrom the llama.cpp release Quail bundles, andmlx_lm.benchmark.- A timed run won't start on a Mac that isn't quiet.
ALLOW_OTHERS=1runs it anyway as a dry run, and marks it as one.
- Rapid-MLX is loaded text-only, so Gemma 4 starts without mlx-vlm.
Quail 0.57.1
Fixed
- Gemma 4's tool calls work. Quail returned Gemma 4's calls as text, on both engines and on both the OpenAI and Anthropic routes, so agents couldn't use it. Now:
- Its calls come back as
tool_calls/tool_use, with the same arguments llama-server gives. - Its thinking (
<|channel>thought…) comes back as reasoning, not as part of the answer. This includes the empty thought block it writes after every tool result. - The tools it's offered reach it in the form it was trained on. The template renderer used to put a stray space after each
{in Gemma 4's tool declarations.
- Its calls come back as
Quail 0.57.0
Added
- A comparison harness in
bench/, for measuring Quail against llama-server, Ollama, oMLX and Rapid-MLX on the same Mac and the same weights with standard benchmarks (ADR D-063). This first part:task bench:setupfetches Ollama's official build and oMLX's release into scratch folders.task bench:doctorsays what's installed, what's missing, and how quiet the Mac is.task bench:smokestarts every engine with every model and records what each really does: whether a request's length can be held, whether thinking is off, whether tool calls parse on the OpenAI and Anthropic routes, and whether a repeated prompt hits a cache.- Nothing in
bench/ships in the app, and it never touches your own Quail, Ollama or oMLX setup.
Quail 0.56.1
Changed
- The plan records that
quail-ai.appandquail-ai.comare verified on the project's GitHub account, and the fresh-Mac checklist says Test… where it said Ping.
Quail 0.56.0
Added
- Quail links to its website, quail-ai.app:
- Quail Help in the menu opens the docs.
- "Learn more" beside the Server, General, Benchmark and model settings captions goes to the matching docs page.
- About has Website and Documentation links.
- The Homebrew cask's homepage is the website too.
Quail 0.55.2
Added
- A website: a landing page, and docs covering getting started, models, connecting tools, the network and API key, power, benchmarks, the
quailcommand, troubleshooting and privacy. It lives inwebsite/and is published with GitHub Pages;task site:serveshows it locally.
Quail 0.55.1
Changed
- The README describes Quail as it is now, calls it a beta heading for 1.0, and points at the notices file for licences.
Quail 0.55.0
Added
- Quail Help in the menu opens the documentation.
- About has links to the release notes, to report an issue, and to the source code. It also shows the MLX engine's versions, and Copy Details includes them.
- "Learn more" links go from the Server, General, Benchmark and model settings pages to the documentation. They appear once the website is published.
Changed
- Shorter explanations: the power, updates and benchmark captions are one line each now, with the details in the documentation.
Quail 0.54.0
Changed
-
Each model's settings are in plain sight. A sliders button on every row of the Models page, or a click on its "8K context" text, opens its settings:
- the model id to copy
- context size and KV cache, each option saying whether it fits on this Mac
- whether it loads when the server starts
- Benchmark… and Delete…
They used to hide in a small "ctx" menu and a star. Right-click a row for Model Settings… and Delete… too.
-
The Models page starts with what this Mac can run, "This Mac: Apple M4 Pro · 64 GB · comfortable up to ~55B", and Details goes to the About page. The Storage section is now "Storage and downloads".
-
Add model… in the menu, and Add Model… on the Server, Connect and Benchmark pages, open the Add Model sheet directly.
Quail 0.53.0
Added
- The Server page says what the server is doing and lets you act on it. It shows running, stopped or failed, and what's loaded. It has Start, Stop and Test…, and when the server failed, the reason and Show Logs. When a change needs a restart, a line says so with a Restart button; the menu shows the same line.
Changed
- The address you copy works. The Server page shows "On this Mac" and, when other devices can connect, "From other devices", with this Mac's network address. It no longer shows
http://0.0.0.0:8080. - Choose who can reach the server instead of typing an IP address: This Mac only, Local network, or Custom (any address, saved when you press Return). An address set before now shows as Custom, unchanged. Starting from the page asks once before the server listens on the network, as the menu does.
- The API key is hidden until you click Show. A new key needs a confirmation, since tools set up with the old one stop working.
- Changing the address, port, key or number of models loaded at once while the server runs now says to restart. The menu said so only for changed models.