Releases: musokean/forge
Release list
v0.5.0 — voice you can talk over
forge v0.5.0 — voice you can talk over
Voice mode is the headline of this release: streaming transcription, sentence-level TTS, barge-in, two
hands-free modes for people without headphones, and echo cancellation that was rebuilt against a real
microphone instead of a synthetic echo path.
Now on PyPI as handcraft-agent.
Install
# from PyPI
pip install "handcraft-agent[server,device]"
# or pin this tag from GitHub
pip install --upgrade "handcraft-agent[server,device] @ git+https://github.com/musokean/forge.git@v0.5.0"Added
- Voice mode Phase 2/3 (#11) — streaming transcription, sentence-level TTS, and barge-in (speak
while the agent is generating to interrupt it). --half-duplexand--ptt— for speakers plus a microphone, no headphones. Half duplex mutes
the microphone while the answer plays and reopens it once the speaker tail has died; PTT captures
only while you hold space, and pressing it stops playback.forge --voice --aec— hands-free echo cancellation: the microphone stays live while the answer
plays and the agent's own voice is subtracted, using the audio it is playing as the reference (a
pure-numpy block NLMS filter, no C extension).- Voice settings in the config file — a
voice:section inconfig/models.yaml, so plain
forge --voicecan be hands-free without retyping flags every session. - Residual echo suppression and online delay estimation — the two pieces that make echo
cancellation work on real hardware. A linear filter cannot model a laptop's speaker-to-microphone
path (the recording correlates with the played audio at 0.045 here), and the device delay is a
property of the hardware, not a constant: 460-520ms measured against the 26ms the driver reports.
Fixed
- Only the first half of an interrupted sentence was transcribed. Barge-in transcribed the audio
the moment it crossed the threshold; it now stops playback, waits for the user to finish, and
transcribes the stitched utterance. - The sentence being synthesized could still play after an interruption. Speaker rounds now carry
an epoch, and a worker whose epoch has passed retires. - The echo canceller's reference was anchored to the wrong moment — when the listener started,
100-300ms before playback began, and then advanced by sample count, so any late or dropped block
shifted it permanently. The filter could not converge at all on the real machine. - A safety net that quietly switched echo cancellation off — it bypassed any frame whose residual
was louder than its input, with no margin, which hit 54% of frames during convergence. config/models.yamlwas ignored for anyone who installed the package — the loader looked in
site-packagesafterpip install, so a config in the current directory was never read.
Verified
- 413 tests pass, 25 of them covering echo cancellation; CI runs the suite on Python 3.9, 3.11 and
3.13, plus the release-copy build check. - Published artifacts were checked before upload (no
config/models.yaml, no key-shaped strings in
either the wheel or the sdist) andtwine checkpassed on both. pip install "handcraft-agent[server,device]==0.5.0"from PyPI into a clean virtualenv reports
0.5.0 from both the package metadata andforge.server, withforge,forge-executorand
forge-device-simon PATH and the new echo-cancellation classes present.- With echo cancellation enabled on the test machine, a playback leaves a residual of 0.0084 against
the voice-activity threshold of 0.012 and the listener emits nospeech_end— it no longer
transcribes its own speech. Delay estimates stayed within 13ms across a playback. - On a clean synthetic echo path the filter converges to an ERLE of 55.7dB with nothing bypassed,
which is what says the reference timing is right.
A note on echo cancellation: if your speaker-to-microphone coupling is poor (microphone far from the
speaker, low volume), the linear stage will not cancel much, and --half-duplex or --ptt are the
reliable options. Full details, including how to interrupt in each mode, are in docs/voice.md.
v0.4.1 — entry points actually ship
forge v0.4.1 — entry points actually ship
A packaging fix, no functional change. Found while verifying v0.4.0's published artifact: installing
from the tag put only main.py into site-packages, so a pip install could not run the two entry
points the docs tell you to run — you had to clone the repo.
Fixed
executor_agent.py(#17 client executor) anddevice_sim.py(#16 device-side simulator) are now
installed with the package- two console scripts join
forge:forge-executor --center http://<centre>:8080 --token <KEY> --id pc-01 --root D:/work— the
executor to run on a controlled PCforge-device-sim --transport socket --port 9130— the hardware simulator that speaks protocol v1
Verified
- v0.4.0's artifact (installed from the tag):
/healthzand/api/statusreport 0.4.0,/docsis
up, and all six/api/executor/*routes are present - v0.4.1's artifact (installed from the tag): version 0.4.1 in both places,
forge-executorand
forge-device-simon PATH, andforge executor --list-capsreporting the machine's real
capabilities
Upgrading
pip install --upgrade "handcraft-agent[server,device] @ git+https://github.com/musokean/forge.git@v0.4.1"v0.4.0 — hardware link and client executor
forge v0.4.0 — hardware link and client executor
All 18 modules are in. This release closes the last two: the hardware control plane (#16) and the client executor (#17) — an agent that drives other PCs. The offline suite is at 332 cases, CI covers 19 test files across Python 3.9 / 3.11 / 3.13.
Highlights
- Hardware control plane (#16, Phase 1) — one protocol, two transports:
- one message = one line of JSON with CRC and
seq/ack/state, protocol v1; the same frame goes over serial and MQTT, so swapping the transport does not touch the protocol pyserial's URL mechanism meansCOM5, asocket://bridge andloop://all run the same code path (the tests exercise the real one)- the control plane keeps the parts that bite in practice: device aliases that drift, commands crossing sites, inconsistent ACK semantics — plus a staged policy (stage, max level, max temperature, max runtime, cooldown), a command state machine with timeout, retry, rollback and an audit trail
device_sim.pyis the device-side stand-in that speaks the real protocol;hardware/esp32_beauty_device.inois a reference firmware skeleton (not compiled or flashed here — honest disclosure, seedocs/hardware.md)- try it:
/device·/device connect·/device mode serial socket://127.0.0.1:9132·/device audit 5
- one message = one line of JSON with CRC and
- Client executor (#17) — the agent drives other machines. A25's shape, the boring way:
- each controlled PC runs a light executor that only dials out (HTTP long poll, no inbound port, no firewall change, no new dependencies)
- capabilities are declared by the client and filtered twice: at the hub (allow-list, staged release
readonly/low_risk/approval/closed_loop, timeout, size caps) and again on the client (allow-list, path jail, size caps, andshellruns through that machine's own sandbox) - GUI sits behind an injectable driver. Without the optional
[executor]extra the client does not declare screenshot/input at all — commands fail withE_NO_GUIinstead of pretending to have worked - four-role Computer Use: planner → executor → (dispatch → raw evidence) → evaluator → supervisor. Four separate model calls with separate prompts, bindable to different models per role, and the evaluator only ever sees raw evidence — never the executor's own explanation
- try it:
forge --serveon the centre,python executor_agent.py --center <url> --token <KEY> --id pc-01 --root D:/workon the controlled PC, then/executor,/executor run pc-01 "whoami",/executor cua pc-01 "…"
- 20 tools, up from 14 — the six new ones are
executor_list/executor_run/executor_file/executor_screen/executor_input/cua_task
Fixes
- Dispatched REPL commands were also sent to the model.
/web,/serve,/logsand/devicewere missingcontinue, so the command ran and the same line went to the router as a task — a silent double execution that burned tokens.test_cli.pynow asserts at the source level that every_*_command()call in the dispatch chain is followed bycontinue - Upgraded installs could not switch modes. A config generated before a section existed made
/device mode …fail with "no device section". The config writers now append the missing section (device and sandbox) instead of giving up - The four-role loop reported finished tasks as failures. First real-machine run: 70s, six steps, ❌ — on a task that had actually been completed and verified. The executor simply never emitted
done. Added a completion gate (when the plan runs out, the evaluator judges the whole task from all the evidence); the executor is now told where it may write (the client's path jail travels in the device brief) and to returndoneas soon as the goal is met. Same task afterwards: 13.2s, 3 steps, ok pytest .no longer collects the release copies —build/andrelease/are excluded (norecursedirs), so a local full run works after a buildbuild/was accidentally committed once and has been removed and gitignored
Upgrading
pip install --upgrade "handcraft-agent[server,device] @ git+https://github.com/musokean/forge.git@v0.4.0"Optional extras: [server] HTTP API · [device] serial/MQTT hardware link · [executor] GUI capabilities for the client executor.
forge v0.3.0 — command sandbox and structured logs
forge v0.3.0 — command sandbox and structured logs
M5 is complete. This release closes the last two modules — the tool sandbox (#4) and full logging (#7) — and brings the offline test suite to 247 cases with CI now covering every module.
Highlights
- Command sandbox (#4) —
run_commandno longer executes straight on your machine:sandbox.mode: auto(default) runs inside Docker when it is available and otherwise falls back to hardened local execution;dockerrefuses to run at all without a daemon (use it in production);local/offfor explicit control- container runs are locked down:
--rm --network=none(no network by default), memory/CPU/PID caps, read-only rootfs with a writable/tmp, non-root user, working directory mounted read-only, and the container is force-removed on timeout - the local fallback still buys real protection: host-env allowlist, dangerous-command patterns (
rm -rf /,mkfs,ddto raw devices,shutdown…), timeout kill, output clipping - security fix: the old
run_commandhanded the wholeos.environto the child process — any command could readDEEPSEEK_API_KEY. The environment is now allowlisted - try it:
/sandbox(status) ·/sandbox mode docker(hot-reloaded) ·/sandbox test echo hi
- Structured logs (#7) — one JSON line per event in
data/logs/forge-YYYYMMDD.jsonl: daily files with size rotation, retention pruning, level filtering, per-run correlation ids (role / model / steps / tools / tokens / latency) and an HTTP request log- secret redaction: values under
key/token/secret/authorizationfields and anything shaped likesk-…/Bearer …/gho_…are written as*** - inspect with
/logs,/logs tail 20,/logs errors
- secret redaction: values under
- CI coverage —
test_device.pyandtest_voice.pyexisted but were never wired into CI; they are now, together with the new sandbox/logging suite
Fixes
- sandbox: a truthy default on
Sandbox.__init__(mode="auto")silently overrodesandbox.modefrom config, sodocker/offpolicies behaved asauto— caught by the new tests tools.run_command(): no longer leaks host environment variables to the commands it executes
Tests & CI
247 passed + 1 skipped on Python 3.9 / 3.11 / 3.13. 16 offline test files — no API keys, no network, no container runtime required. The single skip is the networked Edge-TTS case, by design.
Install
pip install "git+https://github.com/musokean/forge.git"
pip install "handcraft-agent[server] @ git+https://github.com/musokean/forge.git" # + HTTP APIforge # interactive chat — try /sandbox and /logs
forge --serve # HTTP API (Swagger at /docs)Notes
- The Docker isolation path is asserted parameter-by-parameter in the suite and needs no container runtime in CI; verifying it against a real container requires Docker installed on the host.
- Still open: hardware Phase 1 (real device link), voice Phase 2/3 (barge-in / streaming), client executor module.
forge v0.2.0 — voice mode, device layer, and an HTTP API
forge v0.2.0 — voice mode, device layer, and an HTTP API
A config-driven, multi-model AI agent with a readable codebase. Since v0.1.0 the project grew three new capability lines — and closed the M5 milestone.
Highlights
- Server mode (#14 · completes M5) —
forge --serveturns forge into an HTTP API:- multi-session persistence (SQLite): one Agent context per session, resume after a client disconnect, restore history after a restart
- API-key auth:
Authorization: Bearer <key>orX-API-Key; keys come from configserver.api_keysorFORGE_API_KEY(env:VARindirection supported). With no key configured it degrades to loopback-only — zero setup locally, a key required for exposure - rate limiting per caller (sliding window,
429+Retry-After) - write tools rejected by default (a service has no interactive approval channel); read-only tools work fully
- per-request access log,
/healthzliveness probe, Swagger UI at/docs - fastapi/uvicorn ship as an optional extra — the core keeps its zero-heavy-dependency promise and still imports without them
- Voice mode (Phase 1) — pluggable STT/TTS +
forge --voice: record → Whisper → Agent → Edge-TTS playback. The agent core is untouched; voice only swaps the I/O. - Device layer (hardware Phase 0) — a simulated triple-function beauty device (
fake_device.py) exposes power / level / temperature / current as agent tools, with writes behind the approval gate and built-in over-temperature protection. Proves the "hardware as tools" path without touching real hardware. - Engineering — every module now has offline tests; CI covers Python 3.9 / 3.11 / 3.13 with zero skips.
Fixes
- voice —
EdgeTTS.synthesizeno longer nestsasyncio.runwhen an event loop is already running (it dispatches to a worker thread instead) - eval — the export test no longer writes a report file into the working tree
- tests — the M1/M2 assertions now check mechanisms instead of one machine's config: a fallback role and a multi-model debate lineup are configuration, not code requirements
Install
pip install "git+https://github.com/musokean/forge.git" # core
pip install "handcraft-agent[server] @ git+https://github.com/musokean/forge.git" # + HTTP APIforge # interactive chat
forge "your question" # one-shot Q&A
forge --web # browser chat UI (zero-dependency HTTP server)
forge --voice # voice mode (needs the voice extras)
forge --serve # HTTP API on 127.0.0.1:8080 — interactive docs at /docsCI
Tests pass on Python 3.9 / 3.11 / 3.13 (193 tests) + release-copy build check + model smoke.
Still open
Docker sandbox and full logging (the remaining M5 tail), hardware Phase 1 (real device link over serial/MQTT), voice Phase 2/3 (barge-in, streaming, lower latency), and the client-executor module.
forge v0.1.0 — first public release
forge v0.1.0 — first public release
A config-driven, multi-model AI agent with a readable codebase.
Highlights
- ReAct core loop with rule-based pre-routing (greetings/simple Q&A answer in 0ms, no model call)
- Multi-agent orchestration: parallel task decomposition, debate mode, auto-routing
- Engineering backbone: read/write tool tiers, context truncation + rolling summaries, error retry + degradation, per-step tracing, write-operation approvals
- Structured output (JSON-schema enforced), local knowledge base (SQLite + FTS5), golden-set eval
- Web UI + Markdown export
- Config-driven: switch models/roles via config/models.yaml, no code changes
- Zero hard dependencies beyond openai + httpx
Install
pip install -e .
forge # interactive chat
forge "your question" # one-shot Q&ACI
Tests pass on Python 3.9 / 3.11 / 3.13 (168 tests) + release-copy build check + model smoke.