Repository navigation
v0.5.0 — voice you can talk over
forge v0.5.0 — voice you can talk over
Voice mode is the headline of this release: streaming transcription, sentence-level TTS, barge-in, two
hands-free modes for people without headphones, and echo cancellation that was rebuilt against a real
microphone instead of a synthetic echo path.
Now on PyPI as handcraft-agent.
Install
# from PyPI
pip install "handcraft-agent[server,device]"
# or pin this tag from GitHub
pip install --upgrade "handcraft-agent[server,device] @ git+https://github.com/musokean/forge.git@v0.5.0"Added
- Voice mode Phase 2/3 (#11) — streaming transcription, sentence-level TTS, and barge-in (speak
while the agent is generating to interrupt it). --half-duplexand--ptt— for speakers plus a microphone, no headphones. Half duplex mutes
the microphone while the answer plays and reopens it once the speaker tail has died; PTT captures
only while you hold space, and pressing it stops playback.forge --voice --aec— hands-free echo cancellation: the microphone stays live while the answer
plays and the agent's own voice is subtracted, using the audio it is playing as the reference (a
pure-numpy block NLMS filter, no C extension).- Voice settings in the config file — a
voice:section inconfig/models.yaml, so plain
forge --voicecan be hands-free without retyping flags every session. - Residual echo suppression and online delay estimation — the two pieces that make echo
cancellation work on real hardware. A linear filter cannot model a laptop's speaker-to-microphone
path (the recording correlates with the played audio at 0.045 here), and the device delay is a
property of the hardware, not a constant: 460-520ms measured against the 26ms the driver reports.
Fixed
- Only the first half of an interrupted sentence was transcribed. Barge-in transcribed the audio
the moment it crossed the threshold; it now stops playback, waits for the user to finish, and
transcribes the stitched utterance. - The sentence being synthesized could still play after an interruption. Speaker rounds now carry
an epoch, and a worker whose epoch has passed retires. - The echo canceller's reference was anchored to the wrong moment — when the listener started,
100-300ms before playback began, and then advanced by sample count, so any late or dropped block
shifted it permanently. The filter could not converge at all on the real machine. - A safety net that quietly switched echo cancellation off — it bypassed any frame whose residual
was louder than its input, with no margin, which hit 54% of frames during convergence. config/models.yamlwas ignored for anyone who installed the package — the loader looked in
site-packagesafterpip install, so a config in the current directory was never read.
Verified
- 413 tests pass, 25 of them covering echo cancellation; CI runs the suite on Python 3.9, 3.11 and
3.13, plus the release-copy build check. - Published artifacts were checked before upload (no
config/models.yaml, no key-shaped strings in
either the wheel or the sdist) andtwine checkpassed on both. pip install "handcraft-agent[server,device]==0.5.0"from PyPI into a clean virtualenv reports
0.5.0 from both the package metadata andforge.server, withforge,forge-executorand
forge-device-simon PATH and the new echo-cancellation classes present.- With echo cancellation enabled on the test machine, a playback leaves a residual of 0.0084 against
the voice-activity threshold of 0.012 and the listener emits nospeech_end— it no longer
transcribes its own speech. Delay estimates stayed within 13ms across a playback. - On a clean synthetic echo path the filter converges to an ERLE of 55.7dB with nothing bypassed,
which is what says the reference timing is right.
A note on echo cancellation: if your speaker-to-microphone coupling is poor (microphone far from the
speaker, low volume), the linear stage will not cancel much, and --half-duplex or --ptt are the
reliable options. Full details, including how to interrupt in each mode, are in docs/voice.md.