Skip to content

Releases: Dicklesworthstone/franken_tts

v0.1.10

Choose a tag to compare

@Dicklesworthstone Dicklesworthstone released this 30 Aug 18:53

Release 0.1.10 — 18 built-in voices across all platforms, Accelerate SGEMM kernels, microdecoder int4 AWQ/GPTQ pipeline, and iOS profiling & workbench tooling.

v0.1.9

Choose a tag to compare

@Dicklesworthstone Dicklesworthstone released this 25 Aug 09:48

Release 0.1.9

Oracle fixtures r1 — Qwen3-TTS 12Hz Base (native CUDA)

Choose a tag to compare

Native-CUDA ConformanceExact reference fixtures (bead frankentts-rf4).

Provenance

  • Source pin: QwenLM/Qwen3-TTS @ 022e286b98fbec7e1e916cb940cdf532cd9f488e
  • Weights pin: HF 5d83992436eae1d760afd27aff78a71d676296fc (hash-verified per docs/truth-pack/WEIGHTS.lfs.json)
  • Runtime (exact): qwen-tts 0.1.1, torch 2.7.1, torchaudio 2.7.1, transformers 4.57.3, accelerate 1.12.0, librosa 0.11.0, soundfile 0.13.1
  • Device: NVIDIA GeForce RTX 4090 (device_provenance inside provenance.json)
  • Corpus: docs/conformance/oracle_corpus.json (synthetic-tone-en; non-human deterministic reference — safe to publish)

Coverage: all four modes for the case — xvector/icl × non-streaming/streaming — with per-stage activation arrays (.npy) and hash-anchored manifests.

Anchors

  • fixture_manifest.json sha256: 57f6c273b397dbc13d27d74a636bef0263d48b90b58a7953ffba2e086273d7d1
  • provenance.json sha256: a146a81081be5145d626c4e5982bbf6d6fa29058a2dab621d210bb205be803f3
  • tarball sha256: 983d57a2b1dc8b1a77a5ec8334b2c8c0c4d86ebcfe6a7a1ea5259f0db120fe42

In-repo copies of the two manifests live under docs/conformance/fixtures-qwen3-tts-12hz-r1/.

v0.1.8 — model downloads survive GitHub throttling

Choose a tag to compare

@Dicklesworthstone Dicklesworthstone released this 13 Aug 02:55

The model download stops depending on GitHub's goodwill. Within hours of real traffic, GitHub's release-asset limiter started returning 503s for the chunked downloads both ftts pull and the browser playground perform — and each had a single source, so a throttled host meant a dead download.

ftts pull now tries ordered mirrors: the Hugging Face model repo first, the GitHub release second. Every asset carries its release name on both hosts, digest verification decides acceptance exactly as before, and only the last mirror's error surfaces when everything fails. The playground's /model proxy serves the same chain (HF, then R2, then GitHub), so any single host failing degrades the download to a slower one instead of a dead one.

Also new, experimental and off by default: ftts convert --embed-q8 stores the 622 MB cold text embedding as Q8 with one scale per 64-element group — a 1.02 GB artifact instead of 1.31 GB, with the worst-row SQNR floor at 35 dB. It ships as default only after the artifact-v2 listening and logit-parity gates pass; binaries at or below 0.1.7 refuse grouped artifacts loudly by name.

Full details in the CHANGELOG. Gate: 51 test suites and clippy -D warnings green at the tagged commit (standing model-gated ladder skips reported honestly by the skip audit).

🤖 Generated with Claude Code

v0.1.7 — the browser stops paying for the model twice

Choose a tag to compare

@Dicklesworthstone Dicklesworthstone released this 12 Aug 20:50

The browser stops paying for the model twice. The 2 GB model used to brush against the tab's memory ceiling because the codec was staged as a second safetensors copy on a wasm heap that can never shrink; it now streams into the engine tensor by tensor, the artifact's hot prefix is released before codec staging begins, only the decoder crosses the wire, and the page tears down its animated chrome while an engine is resident. The playground also proves what it plays: a browser-conformance harness compares real in-browser synthesis against a CLI-rendered golden of the same text and voice, sample by sample, with the first codec frames mirrored on both sides for token-level triage — and wasm now uses the native f32 reduction order, so the browser and the CLI run the same numerics route out of the box.

On the CLI side, the kernel team picks up the last quantized routes that were still single-core: FTTS_INT8=w8a16 fans out across the persistent worker team like the default route (3.6–5.0× per projection at the model's decode shapes in a kernel-level A/B, provisional pending a quiet-host run), batched prefill/verify calls route to the autotuned plan's measured batch-regime winner, and the armed quant path no longer allocates per call.

Full details in the CHANGELOG. Gate: 51 test suites and clippy -D warnings green at the tagged commit (standing model-gated ladder skips reported honestly by the skip audit).

🤖 Generated with Claude Code

v0.1.6 — voice cards: a picture that IS the voice

Choose a tag to compare

@Dicklesworthstone Dicklesworthstone released this 11 Aug 05:30

Voices become pictures. ftts card export renders any voice — a preset or your own .spk — as a shareable PNG whose green mosaic IS the full 1,024-float voiceprint: 144×144 cells at two bits each, QR-style finder patterns, interleaved Reed-Solomon error correction, and a lossless PNG chunk as a byte-exact fast path. The card survives screenshots and messaging-app recompression, imports back with ftts card import (PNG or JPEG, even quarter-turned), and ftts say --voice card.png speaks straight from the picture. The format is shared with the iOS app — cards made on a phone import here and vice versa, with the two encoders proven bit-identical and pinned by test.

Also in this release: the end-of-utterance tail trim now removes the noise burst it was built to remove (it previously kept the audible artifact and deleted the harmless silence after it — sample counts could not tell the difference; a content-asserting test now can), the iOS video exporter no longer deadlocks partway and renders several times faster, synthesis output on iOS runs through the same neural denoiser enrollment uses, and kernel worker threads on Apple platforms request an elevated QoS class so they stop idling on efficiency cores.

Full details in the CHANGELOG. Gate: 51/51 test suites and clippy -D warnings at the tagged commit.

v0.1.5 — the browser gets 9.4× faster

Choose a tag to compare

@Dicklesworthstone Dicklesworthstone released this 10 Aug 14:45

The browser engine runs 9.4× faster, Apple devices stop crashing, and enrollment cleans its own reference audio.

The browser at 0.31–0.43× real time

Measured in real Chromium against the real model, against 0.05× before. Per stage: codec 89.1 s → 6.8 s (13.2×), talker+microdecoder 7.6 s → 3.4 s (2.2×), total 97.3 s → 10.35 s.

Two levers, both bit-identical to the scalar reference rather than accuracy traded for speed:

  • A register-tiled, panel-packed f32 GEMM, ported from franken_numpy's fnp-linalg and adapted to this project's [n, k] weight layout so no transpose is materialized. WebAssembly has no BLAS, so every codec convolution and projection had been falling through to a dot product per output element. All six codec GEMM sites reach one function, so one kernel upgraded every convolution, ConvNeXt pointwise pair, and transformer projection at once.
  • The codec's dense route now dispatches across six kernel-team partitions. It had been 92% of frame time on a single thread while every worker sat parked.

Recorded as PERF-005 and labelled PROVISIONAL_LOCAL_WIN: single runs, unequal utterance lengths, no interleaved thermal pairs, and the incumbent is our own previous build rather than a pinned external one.

Apple devices survive

Measured on an iPhone 17 Pro Max: growing a shared wasm memory reclaims the tab past ~1 GB, while growing an unshared one to 2.75 GB is fine. Rust's allocator grows linear memory on every heap request, so a 2 GB model guarantees growth. The site now ships two builds and selects at runtime.

Also fixed: a message posted before a module worker finished evaluating was dispatched to no listener and lost, which bricked the engine in every browser — Chrome hung, Safari threw a null-property error, same cause.

Enrollment denoises itself

Every ftts enroll cleans the reference with a pure-Rust FastEnhancer-S port (207 K params) before computing the embedding, in the CLI and the browser alike. On a 15 dB-SNR reference it lowers the pause floor by 65.5 dB against classic spectral subtraction's 24.6 dB. --no-denoise opts out.

Also

  • install.ps1, a PowerShell one-liner for Windows, validated end to end on real hardware.
  • ftts say prints a human summary on a terminal instead of NDJSON; --robot forces the machine stream back.
  • Resident-daemon hardening: an unbounded pre-auth read, speaker vectors silently repaired instead of rejected, and a panic in one request killing the process holding the only warm model.

Full detail in CHANGELOG.md.

Install

brew install dicklesworthstone/tap/franken-tts
curl -fsSL https://raw.githubusercontent.com/Dicklesworthstone/franken_tts/main/install.sh | bash

Verify downloads against SHA256SUMS.txt.

v0.1.4 — faster than real time

Choose a tag to compare

@Dicklesworthstone Dicklesworthstone released this 09 Aug 19:40

The speed release: synthesis now runs faster than real time on Apple Silicon (1.4–1.6× typical on an M4 Pro under load), model load is 2.5× faster (~3.7 s warm), and first audio arrives ~450 ms after synthesis starts. Ships seven built-in preset voices (--voice matt|aria|ember|james|judy|leo|robert, matt is the out-of-box default), ftts pull quantized-artifact fetch, enroll --overwrite with backup, and opt-in enrollment denoise.

Highlights:

  • Persistent int8 worker team with bit-identical output partitioning; GQA attention on the same team
  • Codec decode overlaps generation (streamed output byte-identical to offline); ttfa_ms reported
  • Startup 2.5× faster: parallel tensor widening, artifact-native Q8 hydration, overlapped codec/tokenizer load
  • Codec dense projections use the reference's own BLAS form (faster and more oracle-faithful)
  • Seven built-in voices bundled in the binary; any enrolled or file voice outranks them
  • ftts enroll --overwrite (keeps a .spk.bak), --denoise, any-input-rate Lanczos resampling
  • Full changelog in CHANGELOG.md

Every optimization ships with a bit-identity or ledgered-equivalence proof; FTTS_INT8=0 restores the f32 reference route.

v0.1.3 — the optimized route goes library-wide

Choose a tag to compare

@Dicklesworthstone Dicklesworthstone released this 09 Aug 03:52

franken_tts v0.1.3 — the optimized route goes library-wide

A fast-follow to v0.1.2, all substance:

Changed

  • The optimized int8 route is now the library-wide default
    (ftts_kernels::route::optimized_default), not just a CLI environment
    default — library consumers get the same speed path as the binary.
    Conformance and oracle entry points pin the f32 reference route explicitly,
    so parity suites never measure the optimized numerics. FTTS_INT8=0 remains
    the master switch back to the bit-exact reference (DISC-003, amended).
  • Talker QKV and gate‖up projections fuse into single int8 dispatches.

Fixed

  • Enrollment from 44.1/48 kHz recordings — the default for phone and Mac
    voice memos. The system-decoder transcode now resamples to the speaker
    encoder's pinned 24 kHz mono instead of preserving the source rate and then
    refusing it. Found live with a real voice memo.
  • The codec int8 quantization memo keys on shape as well as pointer/length, so
    an allocator-reused address can never replay a memo entry under the wrong
    matrix geometry.

Quick start

brew install dicklesworthstone/tap/franken-tts   # or: curl -fsSL https://raw.githubusercontent.com/Dicklesworthstone/franken_tts/main/install.sh | bash
ftts pull                                        # one-time, ~2.0 GB, SHA-256 verified
ftts enroll voice_memo.m4a --default             # 44.1/48 kHz memos now work directly
ftts say "Hello from franken_tts" hello.m4a

Runtime code: MIT with OpenAI/Anthropic rider. Model weights: Apache-2.0 by
Qwen (attribution embedded and preserved). Full details in the CHANGELOG.

v0.1.2 — the optimized artifact ships, and it's the default

Choose a tag to compare

@Dicklesworthstone Dicklesworthstone released this 09 Aug 02:39

franken_tts v0.1.2 — the optimized artifact ships, and it's the default

Two things landed that change what a fresh install feels like:

ftts pull now fetches the quantized .fttsq (~2.0 GB total, was ~2.4 GB)

ftts convert works on the real checkpoint now — the converter packed one
container section per tensor (478) against the format's 64-section cap; sections
are now one per access class, with tensors located by offset inside them. The
resulting 1.3 GB artifact (talker/microdecoder hot projections int8
per-output-channel, everything else verbatim) is on the model release, and the
manifest embedded in this binary pulls it instead of the raw 1.7 GB main
checkpoint. Enrollment hydrates the speaker encoder from the artifact too —
bit-identical .spk output vs the raw checkpoint, proven on a real recording —
so the artifact-only bundle powers the whole workflow. Pulls are idempotent:
re-running verifies sizes and SHA-256 digests and downloads nothing.

The ftts binary defaults to the optimized int8 route

Measured 0.66–1.05× real time on a loaded M4 Pro (vs 6–7× slower for the f32
reference). The route: W8A8 int8 talker/microdecoder GEMMs on a six-way worker
team with fused QKV and gate‖up dispatches, int8 codec ConvNeXt projections
(spectral gate: 0.65 dB LSD, transparent), and vectorized SnakeBeta
(129.6 dB SNR, transparent). Honest boundary (DISC-003): sampled outputs are
different valid renditions, not the f32 waveform, and the listening-based
evaluation of the full sampled path is still open. FTTS_INT8=0 restores the
bit-exact f32 reference end to end.

Quick start

brew install dicklesworthstone/tap/franken-tts   # or: curl -fsSL https://raw.githubusercontent.com/Dicklesworthstone/franken_tts/main/install.sh | bash
ftts pull                                        # one-time, ~2.0 GB, SHA-256 verified
ftts enroll voice_memo.m4a --default             # clone a voice from any recording you may use
ftts say "Hello from franken_tts" hello.m4a

Runtime code: MIT with OpenAI/Anthropic rider. Model weights: Apache-2.0 by
Qwen (attribution embedded and preserved). Full details in the CHANGELOG.