Highlights since v1.4.2
This is the cumulative summary of every release shipped between v1.4.2 (2026-02-26) and v1.5.8 (2026-04-26) — five months of work covering a major config overhaul, a brand-new auth system, the full translation / diarization / summarize pipelines, vLLM and MLX inference backends, ten new CLI commands, and the YouTube + ASS karaoke rewrites.
Heads up: v1.5.0 contained breaking changes to
CaptionConfigand CLI flags. If you are upgrading directly from v1.4.x, read the migration table below before bumping.
Breaking changes (v1.5.0)
CaptionConfig was flattened-into-nested. CLI dotted paths follow the same shape:
| Before (v1.4.x) | After (v1.5.x) |
|---|---|
CaptionConfig(split_sentence=True) |
CaptionConfig(input=CaptionInputConfig(split_sentence=True)) |
CaptionConfig(word_level=True) |
CaptionConfig(render=RenderConfig(word_level=True)) |
CaptionConfig(karaoke=KaraokeConfig(enabled=True)) |
CaptionConfig(ass=ASSConfig(karaoke_effect="sweep")) |
caption.write(path, word_level=…, karaoke_config=…) |
caption.write(path, render=…, format_config=…) |
caption.word_level=true (CLI) |
render.word_level=true (CLI) |
caption.karaoke.enabled=true (CLI) |
ass.karaoke_effect=sweep (CLI) |
AlignmentConfig(flush=N) |
AlignmentConfig(flush_interval=N) |
lai alignment youtube … |
lai youtube align … |
lai translate run … |
lai translate caption … |
New CLI commands (10+)
lai auth login | logout | trial | whoami # device-flow auth + 14-day trial
lai doctor # diagnostics with bundled selftest data
lai update # detects stale editable installs
lai config set KEY=VALUE # ~/.config/lattifai/config.toml
lai serve # local 4-tab playground (align / transcribe / convert / translate)
lai summarize caption # chapter-based summaries with metadata, SEO, quotes
lai youtube run | align # one-shot YouTube → aligned captions
lai translate caption | youtube # full multi-provider translation pipeline
lai diarize naming # LLM-based speaker name inferenceTop-level lai --version, --extra-index-url autoconfig via uv tool install, and a lai umbrella command replace the older lai-* scripts (the legacy entry points still work).
Authentication
- New device-flow login (
lai auth login) — opens browser, listens on a random local port, saves token to~/.config/lattifai/config.toml lai auth trial— no-signup quick start with auto-provisioned 14-day trial keyX-Device-Authheader auto-injection for device-bound API keyslai auth whoamishows account email + plan- v1.5.8 adds a one-shot pre-flight warning before backend-bound commands when the trial is expired or expiring within 2 days, so you see the issue before the server returns 401
Transcription
- vLLM / SGLang backend for Voxtral Realtime (WebSocket), Gemma-3n, Fun-ASR-Nano
- Qwen3-ASR local model support (vendored, decoupled from
transformers5.x conflict) - MLX backend (Apple Silicon):
MLXTranscriberwith dual-backend (mlx-audio+mlx-vlm); 30-second chunk cap for backend reliability - Google Gemma-4 E2B / E4B support; new Gemini 3 Pro / 2.5 Flash / 2.5 Flash-Lite
- Batch concurrent transcription for VAD chunks; dual API mode (transcriptions / chat); event detection; verbose_json
- Auto-inject ASR system prompt and temperature for general-purpose LLMs
- In-memory MP3 chunks, configurable
max_retries(default 2),system_prompt/audio_content_type/verbosefor vLLM chat mode
Translation pipeline
- Multi-provider (Gemini / OpenAI / OpenAI-compatible) end-to-end pipeline
- Three modes: quick (direct), normal (analyze → translate), refined (analyze → translate → review → polish)
- Glossary management, bilingual output with
translation_firstordering - Brand-name and no-meta-commentary rules in prompts
- Retry with exponential backoff and checkpoint/resume for long documents
Speaker diarization
- LLM-based speaker name inference using YouTube metadata and dialogue context
- Candidate extraction with voting, talk-format detection, post-LLM audience correction
lai diarize naming— dedicated command for speaker identification
YouTube
- 4+ transcript extraction strategies (Dwarkesh, Substack, podscripts.co, Rescript API)
- Chrome CDP and Safari fallbacks for SPA transcript pages
- SSL hijack detection, auto-detect caption language from video metadata
- Parent-channel metadata resolution + reordered speaker palette
Summarize (new in v1.5.5+)
- Chapter-based summaries with metadata, SEO fields, and notable quotes
honor_meta_chaptershard constraint — preserves user-supplied chapters verbatim- json-repair for robust LLM JSON parsing; retry on JSON failure with raw output logging
- v1.5.8:
SummarizationConfig.llmnow reads[summarization]inconfig.tomland falls back to theSUMMARIZATION_MODEL_NAMEenv var — point summarization at any OpenAI-compatible endpoint (SiliconFlow, vLLM, …) without code edits. Default stays Gemini.
Alignment
- Streaming flush-every-N-chunks for O(chunk) memory usage on long files
- Duplicate text-block detection with diff-aware detokenization (handles repeated subtitle segments across flush boundaries)
- Lyrics-aware: auto-suppresses duplicate output for lyrics input;
Goption to keep genuine duplicate blocks - Currency / percent / thousands aware tokenization
- RMS volume normalization before ONNX inference
start_margin/end_marginunified default of 0.10sAudioData.stats()+ CLI for amplitude distribution- Preserve input
Supervision.idwhensplit_sentence=False
ASS / Karaoke
- 12 karaoke color schemes (azure-gold, sakura-purple, mint-ocean, …)
- 10-color speaker palette with
ass.speaker_color=auto ASSConfigself-contained: font, colors, outline, shadow, positioning, karaoke effectkinetic_stylesurfaced in caption convert with auto-enable for word-level output
LLM module
- Unified
lattifai.llm:BaseLLM,GeminiLLM,OpenAICompatLLM - Shared
LLMConfigwith reasoning toggle and section-based auto-resolution - Gemini routed through OpenAI-compatible endpoint via shared client
Config system
config.tomlauto-resolution: CLI defaults sourced from~/.config/lattifai/config.toml- Structured
CaptionConfigwith sub-configs:CaptionInputConfig,CaptionOutputConfig,RenderConfig - Per-format configs:
ASSConfig,TTMLConfig,FCPXMLConfig,PremiereConfig,LRCConfig - Nested TOML sections for per-module LLM configs
- TOML round-trip preserves comments (tomlkit)
Notable fixes
- VTT speaker prefix preserved on consecutive same-speaker cues (was deduped, now matches user expectation that explicit speakers are emitted unconditionally)
- Auth/quota errors propagated correctly from alignment API instead of wrapping as
LatticeDecodingError - Multilingual tokenizer fixes; label-aware boundary merge; punctuation-inflated duplicate filter
- CJK support in
align_timestamps_from_ref - Numpy array wrapping for Qwen3-ASR; verbose_json fallback; correct Supervision timing
- Theme: dark orange
warnfor white-background terminal readability - 15 bugs from v1.5.0 third-party install test report (config flags, README examples, classifier mismatches, doctor false positives,
lai --version) - 6 additional bugs from v1.5.2 install test report
Dependencies
lattifai-core:>=0.7.5(was>=0.6.0at v1.4.2)lattifai-captions:>=0.4.11(was>=0.2.2at v1.4.2; new RenderConfig API at 0.4.0)lattifai-run:>=1.0.4lattifai-auth:>=0.2.1(new)k2py:0.4.0(upgraded from 0.2.4)- Python 3.10 – 3.14 supported (added 3.13 / 3.14)
- Replaced deprecated
cgi.FieldStoragewith stdlib-only multipart parser (Python 3.13 ready) - 23k lines of dead vendored
qwen_asrcode removed
Install
# CLI tool (recommended)
uv tool install "lattifai[all]" --extra-index-url https://lattifai.github.io/pypi/simple/
# or pip
pip install "lattifai[all]" --extra-index-url https://lattifai.github.io/pypi/simple/Full version history
For per-version detail see CHANGELOG.md:
- v1.5.8 — Trial-expiry pre-flight warning · VTT speaker prefix preservation · Summarization OpenAI-compat fallback
- v1.5.7 — Lyrics-aware dedup · currency tokenization ·
AudioData.stats() - v1.5.6 —
honor_meta_chapterssummarize constraint ·Supervision.idpreservation - v1.5.5 — Chapter-based summarize · brand-name translation rules · json-repair
- v1.5.4 — MLX backend (Apple Silicon)
- v1.5.2 —
X-Device-Authheader injection · 23k-lineqwen_asrcleanup - v1.5.1 — 15-bug fix pass from v1.5.0 install report
- v1.5.0 — Breaking: CaptionConfig overhaul · auth · translation · diarization · YouTube · summarize · CLI rewrite
Full diff: v1.4.2...v1.5.8