Immutable
release. Only release title and notes can be modified.
What's new in v0.6.0
Added
- Computer use no longer fights you for the cursor (coexistence). When the
agent is driving the host desktop and you start using the machine, it now yields
the seat instead of stealing the pointer mid-action. Before any disturbing action
(click, drag, scroll, mouse-move, type, key, paste) it checks how long you have
been idle and, if you are active, briefly waits in-turn for a pause, then either
proceeds or holds the action back and tells you so./pausehands the machine
back on demand (the session stays alive and resumable — unlike/stop, which
tears it down) and/resumecontinues. Controlled by[fermix_core.computer_use] courtesy(yielddefault,off) andcourtesy_idle_ms. Idle detection is
macOS-only; where it is unavailable the arbiter proceeds rather than blocking. - Image generation can run on a ChatGPT/Codex subscription (keyless). A new
openai_codeximage backend drivesgpt-image-2over the Codex subscription
endpoint — billed to the subscription, noOPENAI_API_KEY— mirroring how the
Codex chat provider works. Opt-in and connection-gated (never a fallback); a
missing/unentitled token fails loud. Selectable in the CLI wizard and web setup. - One-tap directory-grant approval, and the grant loop is multi-user safe. The
/confirmprompt renders as a tap-to-copy code span, and Telegram/Discord get a
native Approve button that synthesizes the exact confirm through the unchanged
origin-bound, single-use, TTL path. Grants are now strictly operator-only (never
reachable through the guest command allowlist),/confirmpeek-validates before
consuming so a wrong-origin attempt can't burn the owner's token, and a non-owner
button tap is refused before consumption. (Telegram + Discord; Slack deferred.) - Owner-approval directory access, and the sandbox lives where you do. Standard
mode now auto-allows the launch/request working directory (under your real home)
plus the workspace and explicit grants, and open mode is your home minus the
protected credential/OS/state dirs — the mode roots key off your OS home, not
FERMIX_HOME, so "works where I am" is finally true. A new
request_directory_access(path, reason)tool prompts the owner for an
out-of-roots path and, on/confirm, persists the grant and auto-resumes the
original request.fermix sandbox explainannotates each root as (granted) vs
(mode). - Voice notes are heard everywhere, on a pluggable speech-to-text backend (M21).
Milestone 21 makes voice-note transcription real, configurable, and channel-wide
(see the transcription entries below). Also replaces the old Groq backend with a
native SpaceXAI STT backend, and relabels user-facing "xAI" → "SpaceXAI" (display
only — the:xaiatom,XAI_API_KEY,api.x.ai, andgrok-*model ids are
unchanged, so existing configs keep working). - FermixPet is now its own notarized macOS app, and the realtime wire is
versioned. The voice companion moved totezra-io/fermix-macos(SwiftPM source,
notarized universal2 DMG + Homebrew cask). The daemon's realtime socket gained a
versioned hello-first handshake (N/N-1 window) so the pet and daemon can ship
independently, andfermix doctor/voice statusnow flag the OpenAI-Platform-key
requirement (a Codex/OAuth login does not authorize the Realtime API). - Every spawned OS process is group-reaped (subprocess lifecycle). External
commands now run under a centralCommandHostin their own process group and are
killed on exit, crash, or daemon shutdown (a newkill_pgidNIF), and cron jobs
can delegate. Closes the orphaned-subprocess class across git tools, subagents,
jobs, sandbox env, plugin probes, and realtime. fermix statusandfermix doctornow warn when the running daemon's
version differs from the installed binary. A package-manager upgrade
(brew upgrade fermix) swaps the binary on disk while the launchd/systemd
service keeps running the old release until restarted, and nothing surfaced
that skew — right after a brew upgrade, doctor's upgrade check even reported
"on the latest version" while the daemon was stale.fermix statusnow
appends a warning line and doctor's daemon-socket check degrades to a
warning, both naming the two versions and thefermix restartfix. The
fermix upgrademanaged-install refusal also tells you to restart after
running the package-manager command, and the README, wiki, and Homebrew
caveats now document the restart-after-upgrade requirement.- Telegram voice notes are transcribed. Inbound Telegram voice notes, audio
files, audio-MIME documents, and round video notes now parse to a transcribable
audio attachment and are transcribed to text like the other channels — closing
the gap where Telegram (the primary channel) extracted only photos and answered
a voice note with "your message looks empty." Video notes ride their MP4
container straight to the hosted backend (no ffmpeg). - A voice note with a caption transcribes both. When an audio attachment
arrives with a caption, the caption is kept and the transcript is appended under
a[voice note transcript]delimiter, instead of the caption suppressing
transcription entirely. - Transcription is now a backend-pluggable capability.
[fermix_core.transcription] backendselects the speech-to-text provider —openai,xai, ordeepgram—
resolved through a fail-loud registry that lists the valid names on an unknown
choice (the on-devicelocalbackend is reserved for a later phase and says so).
Each backend has its own optional API-key slot:openai_api_keyand
xai_api_keyOVERRIDE the reused chat-provider key (or fall through to it if
unset), while Deepgram (nova-3, batch; no chat provider to reuse) requires its
deepgram_api_key. SpaceXAI's native/v1/sttis modelless (no model to pick)
and requires an API key — the Grok subscription OAuth token does not work for
STT, so paste one when your SpaceXAI provider is on OAuth. Every backend routes
its HTTP round-trip through the shared[:fermix, :provider, :call]telemetry
emitter (purpose: :transcription, no token cost), and a missing key fails loud
rather than degrading silently. - Transcription is now configurable through setup.
[fermix_core.transcription]
is a first-classconfig.tomlsection (backend,model, per-backend
openai_api_key/xai_api_key/deepgram_api_key,max_file_mb) with an unknown
key/backend/max_file_mbfailing config load loudly. The web-setup Transcription
card (moved to right after Channels) shows an API-key field for the selected
backend (all three ride secure-on-save to the OS keyring), plus a per-backend
model dropdown. Set it from the card, or withfermix setup --transcription-backend/--transcription-model/--transcription-api-key(the
generic key flag stores under the currently-selected backend's slot); switching
backend snaps the single sharedmodelto that backend's default so Deepgram
never inherits the OpenAI-shaped model (SpaceXAI is modelless and sends none).
fermix doctorgains atranscriptionrow that reports the active backend and
whether its credential resolves (offline — it never transcribes).
Changed
- GPT-5.6 (Sol, Terra, Luna) context window corrected to 272k (carried
forward from 0.5.8). The catalog listed the 5.6 generation at 372k; the
effective window is 272k on both the Codex (ChatGPT subscription) and OpenAI
direct-API paths. Auto-compaction thresholds, forced/compactbudgets, and the
fermix doctorcontext report key off the corrected value; the oldergpt-5.5/
gpt-5.4/gpt-5.4-miniwindows are unchanged. - Live-voice replies are quieter and shorter. The realtime seed no longer
licenses pre-announcing or narrating tool use ("act and lead with the result; a
slow action stays silent until you have the answer"), and the default length
tightened to one sentence, two at most — the "or multi-step" clause had made
narration the default rather than the exception (computer-use especially). - Built-in tool triggers lead with when to act, not just when not to. The
when_to_useforweb_search/web_fetch/subagents/memory_recallnow opens
with the affirmative trigger (the answer is current/changing, you already have
the URL, the request fans out, it may depend on stored facts) before its routing
exclusions — the documented lever for should-call rate on tool-conservative models. - Time-sensitive facts are grounded before answering from memory. A broad
"Verify, Don't Guess" principle now covers changeable external facts (rates,
figures, standings), not just local machine/repo state — verify with a tool
before answering, however sure the model feels. (Raised the capability-eval
web-research fire rate 32% → 68% with no query changes.) - The agent holds an evidence-backed answer under pushback. A new "Pushback
Gets Diligence, Not Deference" rule: re-check before conceding a challenge,
reconcile an apparent contradiction rather than capitulating, and change the
answer only when evidence changes it — confidence tracks evidence in both
directions, no caving and no digging in. - The default transcription model is now
gpt-4o-mini-transcribe(was the
legacywhisper-1) — better accuracy on the same OpenAI endpoint. Set
[fermix_core.transcription] modelto pin a different model.
Fixed
- Voice calls no longer freeze after a computer-use (or any slow) tool run.
The realtimeSessionServeris now a non-blocking coordinator — tool calls run
off the session loop on a supervised task, so mic audio, interrupts, and stop
keep flowing while a tool runs — and the connection handler owns its socket so
the pet always sees EOF on teardown. (Fixes a four-defect deadlock chain that
wedged the companion mid-call.) - Anthropic models now actually reason on time-sensitive turns. Opus 4.8 (and
4.6+/Sonnet 5/Fable/Mythos) runs without thinking unless the request carries
thinking: {type: "adaptive"}—reasoning_effortalone only calibrates token
spend — so the daemon now sends adaptive thinking on models that support it
(Haiku 4.5 stays gated off). The non-streaming clamp and buffered receive window
grew to fit thinking plus the visible answer. - The OpenAI direct-API (api-key) path works again. The Responses adapter was
sendingtemperature, which every model it serves (the gpt-5 reasoning family)
rejects with a 400 — breaking the whole direct-API catalog (the dev daemon never
saw it because it rides the Codex OAuth adapter). Temperature is dropped from the
Responses payloads; ChatCompletions is unaffected. web_searchno longer degrades to DuckDuckGo after an idle gap. All seven
search backends now route through the shared hardened Finch pool (15s idle cap +
one stale-socket retry) instead of per-request Req options that silently ran on
an infinite-idle pool — Cloudflare closed those keepalives and Finch handed out
the deadest one first, so the first searches after idle each burned a dead socket
(45 occurrences since June).fermix doctor --fullnow starts the pool before
its live probes.- A connect stall on a continuation call no longer kills the whole run. The
Codex adapter mislabeled a ~5s TCP/TLS connect stall as between-chunk stream
starvation and had no retry seam for continuations, so a mid-run blip was fatal
(and the failure-report delivery died on the same blip). Transportstageis now
a measured value, and continuations get bounded in-place retry for exactly the
proven zero-data timeout class (2 attempts, no tool replay, no provider switch) —
including scheduled runs. - Scheduled jobs honor their configured timeout.
Jobs.Runnerresolved
timeout precedence by key presence rather than value, so the Scheduler's
always-present-niltimeout_msshadowed each job's owntimeout_seconds/
inactivity_timeout_seconds— jobs set for longer were killed at the 30-minute
default and the inactivity watchdog never armed. Now keyed on the value. - Codex-style API errors surface their detail, and wire booleans are preserved.
Provider error composition falls back to a top-leveldetailfield (Codex
{"detail": …}errors were showing a generic message), and the introspection
wire no longer stringifies a baretrue/false/nil. - Failed voice-note transcription no longer drops silently. When
transcription isn't configured, the audio is over the size cap, or the provider
errors, the sender now receives a specific, actionable reply (not configured →
runfermix setup; too large → the size-cap limit; other failures →
transcription failed, try again) and no turn is scheduled — previously the
gateway logged the error and the sender heard nothing. A[fermix_core.transcription] max_file_mbcap (default 20, aligned with Telegram's bot limit) is enforced
before download when the size is declared, after download otherwise.
Internal
- Multi-OS CI gate + disposable eval/benchmark boxes (M22). The PR check job
is now a 3-leg matrix (linux-x64, linux-arm64, macOS-arm64) — 3 of the 4 shipped
Burrito targets had never run the suite in CI. The eval stack (including the
tiers too dangerous for a dev machine) runs unattended on disposable cloud boxes,
a destructive run additionally requiresFERMIX_EVAL_DISPOSABLE=1, and any failed
eval tier auto-files a deduped GitHub issue. Branch protection: the required check
splits into three per-leg names. - Eval boxes and the benchmark harness use a hosted Opik + an external judge.
Boxes talk to Comet-hosted Opik over its API (the in-box Docker stack is gone),
and the local harness calls the external judge API directly (make capability-autoseeds and tears down a throwaway capability daemon); new
chief-of-staff / epistemic-integrity suites added. - compux pinned to v0.5.0 (protocol 3) — the sidecar library behind computer
use, carrying the new idle-detection actions the coexistence arbiter uses; the
fermix pin verifies against the v0.5.0 signed-release checksum. - Docs:
self_knowledge, README, and wiki refreshed (reasoning-effort per-model
ceiling,schedule_jobtimeout args, bootstrap-template drift, FermixPet cask
migration, provider/channel/voice sections).
Verifying signatures
Each binary is signed with cosign using GitHub Actions OIDC keyless signing. Verify a download with:
cosign verify-blob \
--certificate fermix_<target>.pem \
--signature fermix_<target>.sig \
--certificate-identity-regexp "https://github.com/tezra-io/fermix/.github/workflows/release.yml@.*" \
--certificate-oidc-issuer https://token.actions.githubusercontent.com \
fermix_<target>The releases.json artifact lists every binary's URL, sha256, and signature URLs.