Skip to content

v4.11.0

Latest

Choose a tag to compare

@mudler mudler released this 02 Oct 21:56
· 9 commits to master since this release
58830f7

🎉 LocalAI 4.11.0 Release! 🚀




LocalAI 4.11.0 is out!

This release makes LocalAI more useful for audio understanding, resilient model serving, structured decisions, and day-to-day operation. Audio scenes can now combine transcription, diarization, sound detection, and remembered speaker names. Ordered failover chains keep a public model available across local and remote targets, while new Studio and operations pages expose these capabilities without requiring distributed mode.

The release also adds first-class decision models through /v1/systemone, signed OCI model galleries, and Kimodo text-to-animation. It includes 243 merged pull requests, 353 commits, 79 new gallery entries, focused fixes across APIs and backends, and broad backend-source updates.

Highlights:

  • 🎙️ Audio scenes and remembered speakers - combine speech-to-text, speaker diarization, sound-event detection, and speaker identification. Studio can diarize a recording, preview clean intervals, and register selected speakers by name.
  • 🔀 Model failover chains - place local and remote targets behind one model name, retry before response commitment, observe health and target switches, and pin a target from the API, MCP tools, or the React UI.
  • 🧭 Decision models - advertise the decisions use case and answer structured choice, score, and noul questions through /v1/systemone, with validation and dedicated gallery models.
  • 📦 Signed OCI galleries - distribute complete model galleries as OCI artifacts, verify them with Sigstore policies, and use digest-bound last-known-good caches.
  • 🕺 Kimodo text-to-animation - generate skeletal motion from text through POST /3d/animate, preview it in Studio, and export binary glTF animation.
  • 🖥️ Operate this machine - inspect resources and locally loaded models, view logs, and stop models on a single LocalAI host without enabling distributed mode.
  • 🛡️ Safer gallery and API behavior - path-confinement fixes, stricter credential checks, correct pre-stream errors, preserved streamed Responses items, and more reliable model installation.

Plus PDF attachment extraction in chat, deeper Hugging Face repository discovery, improved hardware detection, new Italian Piper voices, NeMo diarization and ASR models, and large model-gallery batches.

ui-diarization-speakers

📸 [ screenshot: diarize a recording and remember speakers by name ]

Studio turns anonymous speaker segments into reusable named speaker profiles.


📊 This release in numbers

Pull requests merged 243
Commits 353
Files changed 623 (+51,035 / -2,625)
Development window 15 days (2026-09-18 to 2026-10-02)
Human contributors 12, of whom 4 first-time
Gallery entries 1,847 to 1,926 (+79)

Where the work landed:

Area Change
core/ +23,883 / -807 across 350 files
backend/ +12,437 / -186 across 119 files
gallery/ +4,998 / -110
pkg/ +3,310 / -170 across 61 files
docs/ +2,142 / -1,115 across 50 files
swagger/ +2,254 / -100

📌 TL;DR

Area Summary
🎙️ Audio scenes parakeet-cpp can combine ASR, diarization, and sound detection in one model. POST /v1/audio/diarization can optionally include text and versioned speaker profiles. Transcription segments and words can carry speaker labels, and realtime sessions emit transcription-segment and sound-detection events.
🗣️ Remembered speakers Configure a compatible speaker_model to identify voices from the shared voice registry. Studio's Diarization page previews clean speaker intervals and registers a selected speaker only after an explicit Name and remember action. Speaker profiles are biometric data, remain opt-in, and are not proof of identity or consent.
🔀 Failover A model config can define an ordered failover.targets chain containing local or remote models. LocalAI retries eligible failures before committing a response, tracks health and recovery, exposes GET /api/failover plus an SSE event stream, adds admin pin and unpin controls, and reports the selected model in response headers. The localai-proxy backend connects chains to another LocalAI endpoint.
🧭 Decisions Models can explicitly declare known_usecases: [decisions] and serve structured requests through POST /v1/systemone. LocalAI validates request size, question count, state, IDs, options, levels, and noul criteria before backend invocation. The React UI exposes a Decisions capability chip and installation guidance.
📦 OCI galleries A gallery source can use oci://host/repository:tag or a digest-pinned reference. LocalAI resolves and verifies the digest, enforces layer and path limits, stages extraction atomically, and ties caches to the active verification policy. source_repository adds an exact Sigstore source constraint.
🕺 Kimodo POST /3d/animate accepts one UTF-8 prompt and generates a 60 to 150 frame skeletal animation at 30 FPS. Studio provides animation controls, preview, and local history. Output is a binary glTF .glb, not a humanoid mesh.
🖥️ Single-host operations Operate → This machine shows VRAM, RAM, CPU, models-disk usage, and running model processes. Operators can search and sort models, inspect logs, and stop a model. The existing Nodes workbench remains available when distributed mode is enabled.
🧠 Models The gallery reached 1,926 entries. Additions include large Qwen3.8 and community model batches, NeMo speech and diarization models, four Italian Piper voices, vllm-cpp structured-extraction models, and decision models including kev, Nimble, and CLM.

🚀 New Features & Major Enhancements

🎙️ Audio scenes: speech, speakers, and sounds

parakeet-cpp can now treat audio as a scene instead of a transcript alone. One model can run speech recognition, speaker diarization, and sound-event classification, then return aligned speaker and sound information for regular and realtime requests.

  • POST /v1/audio/diarization returns diarization segments and speaker summaries.
  • include_text=true combines diarization with ASR.
  • include_speaker_profiles=true explicitly requests versioned speaker profiles.
  • Transcription segments gain speaker_name; transcript words can carry a speaker label.
  • Realtime sessions emit conversation.item.input_audio_transcription.segment and conversation.item.sound_detection events.
  • Scene models can serve both transcription and sound_detection, with diarization enabled in the realtime pipeline.
  • New gallery entries cover Nemotron 3 diarization, diarization plus ASR, CED sound models, and realtime scene variants.

Speaker profiles are sensitive biometric data. Profile export is opt-in, requires voice-recognition permission, and does not itself register a person.

🔗 PRs: #12335, #12382

🗣️ Name and remember speakers

LocalAI can match diarized segments against voices registered in its shared voice registry. A compatible parakeet-cpp configuration uses speaker_model, speaker_threshold, and speaker_margin; known voices replace anonymous labels with names while unknown voices remain SPEAKER_NN.

Studio adds a Diarization page for enrollment from ordinary recordings. It finds clean intervals for each speaker, provides audio previews, and registers a profile only when the user selects Name and remember. Profiles and recordings are not persisted in browser storage and the relevant request bodies are excluded from API trace capture.

The registry is currently process-local and ephemeral. Names disappear after restart and are not synchronized across independent frontends. Encoder identity must match exactly between enrollment and recognition.

ui-diarization-speakers 📸 [ screenshot: speaker enrollment in Studio ]
Preview clean intervals, choose a speaker, and register the voice with an explicit action.

🔗 PR: #12414

🔀 Model failover chains and localai-proxy

One public model name can now represent an ordered chain of local and remote targets:

name: assistant-llm
failover:
  targets:
    - model: preferred-local
    - model: remote-localai
      warm: true

LocalAI retries the next target only when failure happens before response commitment. Transport errors, server errors, OOM, and rate limiting can trip a target; validation errors, ordinary client errors, and client cancellation do not. Recovery probes and a minimum fallback residence period prevent rapid target flapping.

The release adds:

  • GET /api/failover, GET /api/failover/{chain}, and GET /api/failover/events.
  • Admin pin and unpin operations, also available as list_failover_chains, pin_failover_target, and unpin_failover_target MCP tools.
  • X-LocalAI-Served-Model and X-LocalAI-Failover response headers.
  • localai.model.failover events for realtime sessions.
  • A Failover Chain model template, live health strip, pin controls, Installed Models badge, and Operate → Runtime → Failover page.
  • The localai-proxy OCI backend for routing LocalAI REST capabilities to another LocalAI endpoint.

ui-failover

📸 [ screenshot: live failover chains and target health ]

Inspect the active target, health state, and administrative pins from the runtime page.

🔗 PR: #12285

🧭 Structured decisions and zero-shot extraction

Decision models are now a first-class LocalAI capability. A model declares known_usecases: [decisions], appears with a Decisions chip, and serves the existing SystemOne contract through POST /v1/systemone. LocalAI does not infer this use case, which keeps decision models distinct from chat and token-classification models.

The request path now enforces a 64 KiB body limit, a maximum of 64 questions, and validation for state, IDs, option counts, levels, and noul criteria. vllm-cpp carries the structured result through its unified decision ABI, and hf_overrides can merge required top-level model configuration without modifying the downloaded snapshot.

Gallery entries add Laya, GLiNER2.5-Decide, Qwen3-VL, Tev1, kev, Nimble, and CLM decision models. Existing /permute and /separate endpoints remain for NER and token classification.

🔗 PRs: #12247, #12373, #12379, #12391, #12397

📦 Signed OCI model galleries

Model galleries can now be shipped as self-contained OCI artifacts. An artifact carries its index.yaml and relative model configuration files, so a registry can distribute a complete gallery without a separate web-hosted index.

LocalAI resolves tags to digests, verifies signatures against that digest, pulls the same digest, validates artifact type and layer metadata, enforces layer-count and size bounds, confines extracted paths, and promotes the gallery cache only after complete extraction and parsing. Verification policies can require an exact source_repository URL, and cache identity includes the active verification policy.

Strict integrity mode refuses OCI galleries without a verification block. A signature or policy failure does not fall back to an older cache; transient network failures can use content previously verified under the same policy.

🔗 PRs: #12167, #12138, #12235, #12238, #12239, #12243

🕺 Kimodo text-to-animation

LocalAI adds a native kimodocpp backend and a Studio workflow for skeletal motion generation. POST /3d/animate accepts exactly one UTF-8 prompt, plus frames, steps, seed, and text_guidance. It produces a binary glTF animation at 30 FPS and reports usage through metadata.usage.

The backend has CPU and Vulkan builds, multiple model and quantization gallery entries, tracing, and usage accounting. Studio provides generation controls, animation preview, and local history. The API generates a skeleton animation rather than a humanoid mesh, and frame count is constrained to 60 through 150.

🔗 PRs: #12095, #12162, #12184

🖥️ Operate this machine

Single-node installations now have an operations page for the local host. Operate → This machine shows resource gauges for VRAM, RAM, CPU, and the models disk, then lists each running model with its backend, resident memory, CPU share, uptime, and PID.

Operators can search and sort the table, view backend logs, and stop a model after confirmation. Operate overview → Running now shows the five heaviest models, and runtime navigation includes a running-model count. In distributed mode, the route continues to show the cluster Nodes workbench.

ui-this-machine

📸 [ screenshot: Operate → This machine ]

Inspect host capacity and control every locally loaded model from one page.

🔗 PR: #12189

🧰 Smaller features worth knowing about

  • Chat and Home extract text from PDF attachments before sending the request (#12374).
  • Hugging Face discovery lists repositories nested more than one directory deep (#12355).
  • Model capabilities report an alias target and fall back to the application default context size (#12183, #12216).
  • AMD APU detection includes GTT memory, and Intel GPU probing avoids startup hangs (#12094, #12206).
  • OCI gallery changes invalidate the React UI model-listing cache immediately (#12235).
  • The gallery adds four Italian community Piper voices (#12121).

🐛 Bug Fixes (recap)

  • fix(functions) - honor function_arguments_key when building tool grammar (#11677).
  • fix(responses) - wait for complete JSON tool calls and preserve streamed output items (#12001, #12048).
  • fix(api) - return correct saturation and no-node status codes, expose alias targets, and report per-slot context (#12113, #12183, #12204).
  • fix(gallery) - normalize OCI references, bind caches to verification policy, reduce repeated directory reads, and keep deletion inside the models directory (#12138, #12238, #12239, #12243, #12283, #12324, #12326).
  • fix(modelartifacts) - reuse committed sibling files when allow_patterns narrows an artifact (#11484).
  • fix(huggingface) - discover repositories nested beyond one directory level (#12355).
  • fix(backend) - re-probe media markers after cold vision-model loading (#12254).
  • fix(llama-cpp) - retain the default RAM cache, return pre-stream errors as errors, and let model parallel:1 override the environment (#12297, #12425, #12426).
  • fix(whisperx) - preserve the transcript when diarization fails and report the real error (#12427).
  • fix(distributed) - stage declared model files before loading (#12309).
  • fix(watchdog) - ignore stale backend evictions (#12333).
  • fix(quantization) - pin imported quantized models to the backend that produced them (#11879).
  • fix(auth) - require validated header credentials for the CSRF exemption (#12185).
  • fix(cosignverify) - locate bundles even when the index entry describes them incorrectly (#12165).
  • fix(react-ui) - extract text from PDF attachments and improve text-selection contrast (#12374, #12123).
  • fix(cloud-proxy) - surface Anthropic refusals instead of returning empty replies (#12424).
  • fix(ollama) - report on-disk size from /api/tags and /api/ps (#11989).
  • fix(xsysinfo) - include AMD GTT memory and avoid Intel GPU probe hangs (#12094, #12206).

🧠 Models

  • The gallery grew from 1,847 to 1,926 entries.
  • A consolidated gallery batch added 293 candidate entries before deduplication and cleanup (#12124).
  • A second batch added NeoHorse, Qwen3.8 Distill and Cyber variants, ByteShape, Flash Next GSQ-RCO, Occamy, Hy-MT2, Maple Preview, and more (#12221).
  • NeMo speech additions cover standalone diarization, ASR, and combined diarization plus ASR (#12265).
  • Four Italian community Piper voices are now available (#12121).
  • vllm-cpp gallery entries add Laya, CUA-S1 forms, GLiNER2.5, kev, Nimble, and CLM decision models (#12240, #12391, #12397).
  • Further additions include Hemmingway, MiMo Distill Qwen 9B, Sharp-Spark, Swift 1.5 GSQ-RCO, ThinkingCap, Agention, Qwopus Flash V2, Cyber-Tiel-Coder, and Cyber-Ornith (#12278, #12282, #12287, #12293, #12295, #12296, #12298, #12300, #12383).

👒 Dependencies

Backend sources and pinned repositories received regular updates:

  • ikawrakow/ik_llama.cpp: 15 updates.
  • 0xShug0/audio.cpp: 14 updates.
  • CrispStrobe/CrispASR: 14 updates.
  • ggml-org/llama.cpp: 12 updates.
  • leejet/stable-diffusion.cpp: 7 updates.
  • ServeurpersoCom/omnivoice.cpp: 7 updates.
  • ggml-org/whisper.cpp: 6 updates.
  • mudler/vllm.cpp: 6 updates.
  • PrismML-Eng/llama.cpp: 6 updates.
  • NVIDIA/NeMo-Speech.cpp: 5 updates.
  • mudler/parakeet.cpp: 4 updates.
  • TheTom/llama-cpp-turboquant: 3 updates.
  • LocalAGI and localai-org/ced.cpp: 2 updates each.
  • antirez/ds4, localai-org/voice-detect.cpp, PABannier/sam3.cpp, the documentation theme, nib, vllm-metal, and the vLLM CUDA wheel: 1 update each.

📖 Documentation

  • Explain mixed CPU and GPU inference (#12143).
  • Remove obsolete per-model documentation sections from the gallery docs (#12222).
  • Clarify that upstream API keys are optional for proxy targets (#12286).
  • Introduce decision models and the SystemOne API (#12428).
  • Fix dead documentation and community example links (#11546, #12320, #12392).
  • Correct gfx1151 environment-variable guidance (#12109).

🙌 New Contributors

Thank you to the first-time contributors in this release:

Full Changelog: v4.10.0...v4.11.0


What's Changed

Bug fixes 🐛

Exciting New Features 🎉

🧠 Models

📖 Documentation and examples

👒 Dependencies

Other Changes

New Contributors

Full Changelog: v4.10.0...v4.11.0