Skip to content

Releases: musokean/forge

v0.5.0 — voice you can talk over

Choose a tag to compare

@musokean musokean released this 03 Oct 15:02

forge v0.5.0 — voice you can talk over

Voice mode is the headline of this release: streaming transcription, sentence-level TTS, barge-in, two
hands-free modes for people without headphones, and echo cancellation that was rebuilt against a real
microphone instead of a synthetic echo path.

Now on PyPI as handcraft-agent.

Install

# from PyPI
pip install "handcraft-agent[server,device]"

# or pin this tag from GitHub
pip install --upgrade "handcraft-agent[server,device] @ git+https://github.com/musokean/forge.git@v0.5.0"

Added

  • Voice mode Phase 2/3 (#11) — streaming transcription, sentence-level TTS, and barge-in (speak
    while the agent is generating to interrupt it).
  • --half-duplex and --ptt — for speakers plus a microphone, no headphones. Half duplex mutes
    the microphone while the answer plays and reopens it once the speaker tail has died; PTT captures
    only while you hold space, and pressing it stops playback.
  • forge --voice --aec — hands-free echo cancellation: the microphone stays live while the answer
    plays and the agent's own voice is subtracted, using the audio it is playing as the reference (a
    pure-numpy block NLMS filter, no C extension).
  • Voice settings in the config file — a voice: section in config/models.yaml, so plain
    forge --voice can be hands-free without retyping flags every session.
  • Residual echo suppression and online delay estimation — the two pieces that make echo
    cancellation work on real hardware. A linear filter cannot model a laptop's speaker-to-microphone
    path (the recording correlates with the played audio at 0.045 here), and the device delay is a
    property of the hardware, not a constant: 460-520ms measured against the 26ms the driver reports.

Fixed

  • Only the first half of an interrupted sentence was transcribed. Barge-in transcribed the audio
    the moment it crossed the threshold; it now stops playback, waits for the user to finish, and
    transcribes the stitched utterance.
  • The sentence being synthesized could still play after an interruption. Speaker rounds now carry
    an epoch, and a worker whose epoch has passed retires.
  • The echo canceller's reference was anchored to the wrong moment — when the listener started,
    100-300ms before playback began, and then advanced by sample count, so any late or dropped block
    shifted it permanently. The filter could not converge at all on the real machine.
  • A safety net that quietly switched echo cancellation off — it bypassed any frame whose residual
    was louder than its input, with no margin, which hit 54% of frames during convergence.
  • config/models.yaml was ignored for anyone who installed the package — the loader looked in
    site-packages after pip install, so a config in the current directory was never read.

Verified

  • 413 tests pass, 25 of them covering echo cancellation; CI runs the suite on Python 3.9, 3.11 and
    3.13, plus the release-copy build check.
  • Published artifacts were checked before upload (no config/models.yaml, no key-shaped strings in
    either the wheel or the sdist) and twine check passed on both.
  • pip install "handcraft-agent[server,device]==0.5.0" from PyPI into a clean virtualenv reports
    0.5.0 from both the package metadata and forge.server, with forge, forge-executor and
    forge-device-sim on PATH and the new echo-cancellation classes present.
  • With echo cancellation enabled on the test machine, a playback leaves a residual of 0.0084 against
    the voice-activity threshold of 0.012 and the listener emits no speech_end — it no longer
    transcribes its own speech. Delay estimates stayed within 13ms across a playback.
  • On a clean synthetic echo path the filter converges to an ERLE of 55.7dB with nothing bypassed,
    which is what says the reference timing is right.

A note on echo cancellation: if your speaker-to-microphone coupling is poor (microphone far from the
speaker, low volume), the linear stage will not cancel much, and --half-duplex or --ptt are the
reliable options. Full details, including how to interrupt in each mode, are in docs/voice.md.

v0.4.1 — entry points actually ship

Choose a tag to compare

@musokean musokean released this 29 Sep 07:23

forge v0.4.1 — entry points actually ship

A packaging fix, no functional change. Found while verifying v0.4.0's published artifact: installing
from the tag put only main.py into site-packages, so a pip install could not run the two entry
points the docs tell you to run — you had to clone the repo.

Fixed

  • executor_agent.py (#17 client executor) and device_sim.py (#16 device-side simulator) are now
    installed with the package
  • two console scripts join forge:
    • forge-executor --center http://<centre>:8080 --token <KEY> --id pc-01 --root D:/work — the
      executor to run on a controlled PC
    • forge-device-sim --transport socket --port 9130 — the hardware simulator that speaks protocol v1

Verified

  • v0.4.0's artifact (installed from the tag): /healthz and /api/status report 0.4.0, /docs is
    up, and all six /api/executor/* routes are present
  • v0.4.1's artifact (installed from the tag): version 0.4.1 in both places, forge-executor and
    forge-device-sim on PATH, and forge executor --list-caps reporting the machine's real
    capabilities

Upgrading

pip install --upgrade "handcraft-agent[server,device] @ git+https://github.com/musokean/forge.git@v0.4.1"

v0.4.0 — hardware link and client executor

Choose a tag to compare

@musokean musokean released this 29 Sep 07:19

forge v0.4.0 — hardware link and client executor

All 18 modules are in. This release closes the last two: the hardware control plane (#16) and the client executor (#17) — an agent that drives other PCs. The offline suite is at 332 cases, CI covers 19 test files across Python 3.9 / 3.11 / 3.13.

Highlights

  • Hardware control plane (#16, Phase 1) — one protocol, two transports:
    • one message = one line of JSON with CRC and seq/ack/state, protocol v1; the same frame goes over serial and MQTT, so swapping the transport does not touch the protocol
    • pyserial's URL mechanism means COM5, a socket:// bridge and loop:// all run the same code path (the tests exercise the real one)
    • the control plane keeps the parts that bite in practice: device aliases that drift, commands crossing sites, inconsistent ACK semantics — plus a staged policy (stage, max level, max temperature, max runtime, cooldown), a command state machine with timeout, retry, rollback and an audit trail
    • device_sim.py is the device-side stand-in that speaks the real protocol; hardware/esp32_beauty_device.ino is a reference firmware skeleton (not compiled or flashed here — honest disclosure, see docs/hardware.md)
    • try it: /device · /device connect · /device mode serial socket://127.0.0.1:9132 · /device audit 5
  • Client executor (#17) — the agent drives other machines. A25's shape, the boring way:
    • each controlled PC runs a light executor that only dials out (HTTP long poll, no inbound port, no firewall change, no new dependencies)
    • capabilities are declared by the client and filtered twice: at the hub (allow-list, staged release readonly/low_risk/approval/closed_loop, timeout, size caps) and again on the client (allow-list, path jail, size caps, and shell runs through that machine's own sandbox)
    • GUI sits behind an injectable driver. Without the optional [executor] extra the client does not declare screenshot/input at all — commands fail with E_NO_GUI instead of pretending to have worked
    • four-role Computer Use: planner → executor → (dispatch → raw evidence) → evaluator → supervisor. Four separate model calls with separate prompts, bindable to different models per role, and the evaluator only ever sees raw evidence — never the executor's own explanation
    • try it: forge --serve on the centre, python executor_agent.py --center <url> --token <KEY> --id pc-01 --root D:/work on the controlled PC, then /executor, /executor run pc-01 "whoami", /executor cua pc-01 "…"
  • 20 tools, up from 14 — the six new ones are executor_list / executor_run / executor_file / executor_screen / executor_input / cua_task

Fixes

  • Dispatched REPL commands were also sent to the model. /web, /serve, /logs and /device were missing continue, so the command ran and the same line went to the router as a task — a silent double execution that burned tokens. test_cli.py now asserts at the source level that every _*_command() call in the dispatch chain is followed by continue
  • Upgraded installs could not switch modes. A config generated before a section existed made /device mode … fail with "no device section". The config writers now append the missing section (device and sandbox) instead of giving up
  • The four-role loop reported finished tasks as failures. First real-machine run: 70s, six steps, ❌ — on a task that had actually been completed and verified. The executor simply never emitted done. Added a completion gate (when the plan runs out, the evaluator judges the whole task from all the evidence); the executor is now told where it may write (the client's path jail travels in the device brief) and to return done as soon as the goal is met. Same task afterwards: 13.2s, 3 steps, ok
  • pytest . no longer collects the release copies — build/ and release/ are excluded (norecursedirs), so a local full run works after a build
  • build/ was accidentally committed once and has been removed and gitignored

Upgrading

pip install --upgrade "handcraft-agent[server,device] @ git+https://github.com/musokean/forge.git@v0.4.0"

Optional extras: [server] HTTP API · [device] serial/MQTT hardware link · [executor] GUI capabilities for the client executor.

forge v0.3.0 — command sandbox and structured logs

Choose a tag to compare

@musokean musokean released this 28 Sep 15:51

forge v0.3.0 — command sandbox and structured logs

M5 is complete. This release closes the last two modules — the tool sandbox (#4) and full logging (#7) — and brings the offline test suite to 247 cases with CI now covering every module.

Highlights

  • Command sandbox (#4) — run_command no longer executes straight on your machine:
    • sandbox.mode: auto (default) runs inside Docker when it is available and otherwise falls back to hardened local execution; docker refuses to run at all without a daemon (use it in production); local / off for explicit control
    • container runs are locked down: --rm --network=none (no network by default), memory/CPU/PID caps, read-only rootfs with a writable /tmp, non-root user, working directory mounted read-only, and the container is force-removed on timeout
    • the local fallback still buys real protection: host-env allowlist, dangerous-command patterns (rm -rf /, mkfs, dd to raw devices, shutdown…), timeout kill, output clipping
    • security fix: the old run_command handed the whole os.environ to the child process — any command could read DEEPSEEK_API_KEY. The environment is now allowlisted
    • try it: /sandbox (status) · /sandbox mode docker (hot-reloaded) · /sandbox test echo hi
  • Structured logs (#7) — one JSON line per event in data/logs/forge-YYYYMMDD.jsonl: daily files with size rotation, retention pruning, level filtering, per-run correlation ids (role / model / steps / tools / tokens / latency) and an HTTP request log
    • secret redaction: values under key / token / secret / authorization fields and anything shaped like sk-… / Bearer … / gho_… are written as ***
    • inspect with /logs, /logs tail 20, /logs errors
  • CI coverage — test_device.py and test_voice.py existed but were never wired into CI; they are now, together with the new sandbox/logging suite

Fixes

  • sandbox: a truthy default on Sandbox.__init__(mode="auto") silently overrode sandbox.mode from config, so docker / off policies behaved as auto — caught by the new tests
  • tools.run_command(): no longer leaks host environment variables to the commands it executes

Tests & CI

247 passed + 1 skipped on Python 3.9 / 3.11 / 3.13. 16 offline test files — no API keys, no network, no container runtime required. The single skip is the networked Edge-TTS case, by design.

Install

pip install "git+https://github.com/musokean/forge.git"
pip install "handcraft-agent[server] @ git+https://github.com/musokean/forge.git"   # + HTTP API
forge                # interactive chat — try /sandbox and /logs
forge --serve        # HTTP API (Swagger at /docs)

Notes

  • The Docker isolation path is asserted parameter-by-parameter in the suite and needs no container runtime in CI; verifying it against a real container requires Docker installed on the host.
  • Still open: hardware Phase 1 (real device link), voice Phase 2/3 (barge-in / streaming), client executor module.

forge v0.2.0 — voice mode, device layer, and an HTTP API

Choose a tag to compare

@musokean musokean released this 28 Sep 13:46

forge v0.2.0 — voice mode, device layer, and an HTTP API

A config-driven, multi-model AI agent with a readable codebase. Since v0.1.0 the project grew three new capability lines — and closed the M5 milestone.

Highlights

  • Server mode (#14 · completes M5) — forge --serve turns forge into an HTTP API:
    • multi-session persistence (SQLite): one Agent context per session, resume after a client disconnect, restore history after a restart
    • API-key auth: Authorization: Bearer <key> or X-API-Key; keys come from config server.api_keys or FORGE_API_KEY (env:VAR indirection supported). With no key configured it degrades to loopback-only — zero setup locally, a key required for exposure
    • rate limiting per caller (sliding window, 429 + Retry-After)
    • write tools rejected by default (a service has no interactive approval channel); read-only tools work fully
    • per-request access log, /healthz liveness probe, Swagger UI at /docs
    • fastapi/uvicorn ship as an optional extra — the core keeps its zero-heavy-dependency promise and still imports without them
  • Voice mode (Phase 1) — pluggable STT/TTS + forge --voice: record → Whisper → Agent → Edge-TTS playback. The agent core is untouched; voice only swaps the I/O.
  • Device layer (hardware Phase 0) — a simulated triple-function beauty device (fake_device.py) exposes power / level / temperature / current as agent tools, with writes behind the approval gate and built-in over-temperature protection. Proves the "hardware as tools" path without touching real hardware.
  • Engineering — every module now has offline tests; CI covers Python 3.9 / 3.11 / 3.13 with zero skips.

Fixes

  • voice — EdgeTTS.synthesize no longer nests asyncio.run when an event loop is already running (it dispatches to a worker thread instead)
  • eval — the export test no longer writes a report file into the working tree
  • tests — the M1/M2 assertions now check mechanisms instead of one machine's config: a fallback role and a multi-model debate lineup are configuration, not code requirements

Install

pip install "git+https://github.com/musokean/forge.git"                              # core
pip install "handcraft-agent[server] @ git+https://github.com/musokean/forge.git"    # + HTTP API
forge                  # interactive chat
forge "your question"  # one-shot Q&A
forge --web            # browser chat UI (zero-dependency HTTP server)
forge --voice          # voice mode (needs the voice extras)
forge --serve          # HTTP API on 127.0.0.1:8080 — interactive docs at /docs

CI

CI

Tests pass on Python 3.9 / 3.11 / 3.13 (193 tests) + release-copy build check + model smoke.

Still open

Docker sandbox and full logging (the remaining M5 tail), hardware Phase 1 (real device link over serial/MQTT), voice Phase 2/3 (barge-in, streaming, lower latency), and the client-executor module.

forge v0.1.0 — first public release

Choose a tag to compare

@musokean musokean released this 24 Aug 07:30

forge v0.1.0 — first public release

A config-driven, multi-model AI agent with a readable codebase.

Highlights

  • ReAct core loop with rule-based pre-routing (greetings/simple Q&A answer in 0ms, no model call)
  • Multi-agent orchestration: parallel task decomposition, debate mode, auto-routing
  • Engineering backbone: read/write tool tiers, context truncation + rolling summaries, error retry + degradation, per-step tracing, write-operation approvals
  • Structured output (JSON-schema enforced), local knowledge base (SQLite + FTS5), golden-set eval
  • Web UI + Markdown export
  • Config-driven: switch models/roles via config/models.yaml, no code changes
  • Zero hard dependencies beyond openai + httpx

Install

pip install -e .
forge                # interactive chat
forge "your question"  # one-shot Q&A

CI

CI
Tests pass on Python 3.9 / 3.11 / 3.13 (168 tests) + release-copy build check + model smoke.