Skip to content

v0.5.0 — voice you can talk over

Choose a tag to compare

@musokean musokean released this 03 Oct 15:02
· 36 commits to main since this release

forge v0.5.0 — voice you can talk over

Voice mode is the headline of this release: streaming transcription, sentence-level TTS, barge-in, two
hands-free modes for people without headphones, and echo cancellation that was rebuilt against a real
microphone instead of a synthetic echo path.

Now on PyPI as handcraft-agent.

Install

# from PyPI
pip install "handcraft-agent[server,device]"

# or pin this tag from GitHub
pip install --upgrade "handcraft-agent[server,device] @ git+https://github.com/musokean/forge.git@v0.5.0"

Added

  • Voice mode Phase 2/3 (#11) — streaming transcription, sentence-level TTS, and barge-in (speak
    while the agent is generating to interrupt it).
  • --half-duplex and --ptt — for speakers plus a microphone, no headphones. Half duplex mutes
    the microphone while the answer plays and reopens it once the speaker tail has died; PTT captures
    only while you hold space, and pressing it stops playback.
  • forge --voice --aec — hands-free echo cancellation: the microphone stays live while the answer
    plays and the agent's own voice is subtracted, using the audio it is playing as the reference (a
    pure-numpy block NLMS filter, no C extension).
  • Voice settings in the config file — a voice: section in config/models.yaml, so plain
    forge --voice can be hands-free without retyping flags every session.
  • Residual echo suppression and online delay estimation — the two pieces that make echo
    cancellation work on real hardware. A linear filter cannot model a laptop's speaker-to-microphone
    path (the recording correlates with the played audio at 0.045 here), and the device delay is a
    property of the hardware, not a constant: 460-520ms measured against the 26ms the driver reports.

Fixed

  • Only the first half of an interrupted sentence was transcribed. Barge-in transcribed the audio
    the moment it crossed the threshold; it now stops playback, waits for the user to finish, and
    transcribes the stitched utterance.
  • The sentence being synthesized could still play after an interruption. Speaker rounds now carry
    an epoch, and a worker whose epoch has passed retires.
  • The echo canceller's reference was anchored to the wrong moment — when the listener started,
    100-300ms before playback began, and then advanced by sample count, so any late or dropped block
    shifted it permanently. The filter could not converge at all on the real machine.
  • A safety net that quietly switched echo cancellation off — it bypassed any frame whose residual
    was louder than its input, with no margin, which hit 54% of frames during convergence.
  • config/models.yaml was ignored for anyone who installed the package — the loader looked in
    site-packages after pip install, so a config in the current directory was never read.

Verified

  • 413 tests pass, 25 of them covering echo cancellation; CI runs the suite on Python 3.9, 3.11 and
    3.13, plus the release-copy build check.
  • Published artifacts were checked before upload (no config/models.yaml, no key-shaped strings in
    either the wheel or the sdist) and twine check passed on both.
  • pip install "handcraft-agent[server,device]==0.5.0" from PyPI into a clean virtualenv reports
    0.5.0 from both the package metadata and forge.server, with forge, forge-executor and
    forge-device-sim on PATH and the new echo-cancellation classes present.
  • With echo cancellation enabled on the test machine, a playback leaves a residual of 0.0084 against
    the voice-activity threshold of 0.012 and the listener emits no speech_end — it no longer
    transcribes its own speech. Delay estimates stayed within 13ms across a playback.
  • On a clean synthetic echo path the filter converges to an ERLE of 55.7dB with nothing bypassed,
    which is what says the reference timing is right.

A note on echo cancellation: if your speaker-to-microphone coupling is poor (microphone far from the
speaker, low volume), the linear stage will not cancel much, and --half-duplex or --ptt are the
reliable options. Full details, including how to interrupt in each mode, are in docs/voice.md.