Vocello 2.4.0
Released 2026-08-01. Vocello 2.4.0 is a voice-quality release: delivery presets were
rewritten and are now verified to steer the voice the way their names promise, Neutral
becomes a real steady default instead of an unconstrained one, and iPhone long-form
projects gain single-segment regeneration. Everything still runs locally on Apple
Silicon; no audio leaves the device unless you export it.
Upgrading from 2.3? Drag-install the new DMG as usual. Your installed models,
history, and saved voices carry over unchanged; no model re-download is needed for this
release.
Download v2.4.0 · All releases
Requirements
- macOS 26.0 or later
- Apple Silicon Mac
- 8 GB RAM minimum (16 GB+ recommended for Quality variants and long projects)
- ~6 GB free disk for the three recommended Speed (4-bit) model variants;
~15 GB if you also install Quality (8-bit) - Internet for one-time model downloads from Hugging Face
- iPhone: the public TestFlight beta is live (join link);
build 23 carries this release.
Headline themes
- Delivery presets that do what they say. The instruction text behind the Custom-mode
delivery presets was rewritten and every rewritten preset now clears a measured 0.85
adherence win-rate on the canonical hardware. The biggest wins are the presets that
previously barely moved the voice: dramatic and surprised now steer prosody reliably,
and happy, excited, angry, and whisper are stronger at the intensities that were weak.
Verification is autonomous: a new prosody-adherence harness measures pitch, rate,
pause, and energy movement against per-preset expectations across seeded takes, so
these claims are gate-checked rather than eyeballed. - Neutral is now a real preset. Selecting Neutral previously sent no delivery
instruction at all, which left the voice free to wander take to take. Neutral now
instructs a steady, even, slightly monotone delivery with measured pacing, so repeated
takes of the same text land closer together in tone. If you want the old unconstrained
behavior, clear the delivery selection entirely. - Regenerate a single long-form segment on iPhone. Long-form projects on iPhone now
let you regenerate just one segment: open the retained project's segments panel from
the setup chip, pick the segment, and confirm. The segment re-renders with a fresh
recorded seed, the joined output reassembles automatically, and History keeps the full
lineage searchable. The Mac app has had this since 2.3; the phone catches up. - Every take speaks the normalized script. Text normalization now happens once at the
engine entry, so prompt construction, language detection, quality checks, and the audio
itself all agree on the exact spoken text. Takes that were transformed record what
changed. This removes a class of subtle mismatches between what you typed and what the
quality gates thought was spoken. - Runtime refresh. The MLX runtime moved to the current upstream pair (mlx-swift
0.31.6, mlx-swift-lm 3.31.4) with a same-day A/B against the previous pins: warm
generation speed within the noise band, identical quality verdicts, and byte-equal
fixed-seed outputs on the clone path. Nothing changes in daily use; the engine now
rides the current dependency line.
Fixes
- iPhone: every interactive control now carries a stable accessibility identifier,
improving assistive-technology addressing and UI-automation coverage.
Under the hood
The delivery-quality claims above are backed by new standing QA machinery: a calibrated
per-preset prosody-expectation profile, a warn-first delivery gate composed into the
benchmark quality registry at canonical depth, an emotion-classifier advisory lane, and a
clone-fidelity lane (speaker identity plus reference-tone match). Benchmark history
records now carry the typed take-quality identity on UI records too (schema v3), and the
vendor-runtime contract re-anchored its evidence checks to a fresh canonical record
captured on the new runtime pins.
Evidence
Recorded per the standing release-candidate procedure (docs/reference/macos-release-qa.md):
- macOS smoke lane:
macos-xcui-smoke-20260801-175657-f956d835— PASS, all seven
journeys (navigation/readiness, Custom generation + History, cancellation, recording,
library surfaces, long-form project, line batch). - Engine regression bench: label
release-QA-2.4.0
(macos-engine-20260801-182014-e18ee022, schema v3) — clean-source run on the
release commit; audio QC clean on all 29 takes (one earlier unseeded attempt hit
the known low-rate dropout band and was correctly rejected by the fail-closed
engine QC; failed runs never publish), warm throughput in line with the canonical
record the public surfaces cite. - Preset adherence evidence: the R1 rewrite acceptance matrix (7 seeds × 18 paired
cells) and the Neutral-preset A/B are banked as delivery-calibration records; findings
indocs/reference/delivery-fidelity-report-2026-08-01.mdand
benchmarks/OPTIMIZATION.md§P–§Q. - Pin-bump A/B: 58-take matrix vs a same-day 0.30.6 control at the canonical seed —
warm RTF deltas +0.6…+2.0% (noise band), identical QC verdict distributions, byte-equal
clone durations;benchmarks/OPTIMIZATION.md§Q.
Install
- Download
Vocello-macos26.dmgfrom the v2.4.0 release. - Open the DMG and drag
Vocello.appto/Applications. - Open Vocello → Settings → Model downloads → install the recommended Speed packages
(~6 GB total). Upgrading installs keep their models; nothing re-downloads. - Generate from Custom Voice, Voice Design, or Voice Cloning.
The DMG is Developer ID signed and notarized with a stapled ticket. release-metadata.txt on
the release records build provenance.
TestFlight — What to Test (build 23)
Build 23 carries all of 2.4.0. Paste this section into the build's Test Details in App
Store Connect when distributing.
- Delivery presets. In Custom mode, generate the same short text with dramatic,
surprised, and whisper selected: each should sound clearly different from no preset,
and repeated takes should stay in character. - Neutral steadiness. Generate the same text three times with Neutral selected: the
takes should stay close in tone and pacing, with no random emotional swings. - Segment regeneration. Create a long-form project (past 900 characters), let it
finish, then open the segments panel from the setup chip and regenerate one segment:
it should re-render, the joined audio should update, and History should keep the
project lineage. - The usual pass. One short take in each mode (Custom, Design, Clone), and one
mid-take cancel: cancel should reset cleanly with nothing added to History.
Report anything odd through TestFlight feedback, including the moment it happened so it
can be matched against diagnostics.