Skip to content

forge v0.2.0 — voice mode, device layer, and an HTTP API

Choose a tag to compare

@musokean musokean released this 28 Sep 13:46
· 57 commits to main since this release

forge v0.2.0 — voice mode, device layer, and an HTTP API

A config-driven, multi-model AI agent with a readable codebase. Since v0.1.0 the project grew three new capability lines — and closed the M5 milestone.

Highlights

  • Server mode (#14 · completes M5) — forge --serve turns forge into an HTTP API:
    • multi-session persistence (SQLite): one Agent context per session, resume after a client disconnect, restore history after a restart
    • API-key auth: Authorization: Bearer <key> or X-API-Key; keys come from config server.api_keys or FORGE_API_KEY (env:VAR indirection supported). With no key configured it degrades to loopback-only — zero setup locally, a key required for exposure
    • rate limiting per caller (sliding window, 429 + Retry-After)
    • write tools rejected by default (a service has no interactive approval channel); read-only tools work fully
    • per-request access log, /healthz liveness probe, Swagger UI at /docs
    • fastapi/uvicorn ship as an optional extra — the core keeps its zero-heavy-dependency promise and still imports without them
  • Voice mode (Phase 1) — pluggable STT/TTS + forge --voice: record → Whisper → Agent → Edge-TTS playback. The agent core is untouched; voice only swaps the I/O.
  • Device layer (hardware Phase 0) — a simulated triple-function beauty device (fake_device.py) exposes power / level / temperature / current as agent tools, with writes behind the approval gate and built-in over-temperature protection. Proves the "hardware as tools" path without touching real hardware.
  • Engineering — every module now has offline tests; CI covers Python 3.9 / 3.11 / 3.13 with zero skips.

Fixes

  • voice — EdgeTTS.synthesize no longer nests asyncio.run when an event loop is already running (it dispatches to a worker thread instead)
  • eval — the export test no longer writes a report file into the working tree
  • tests — the M1/M2 assertions now check mechanisms instead of one machine's config: a fallback role and a multi-model debate lineup are configuration, not code requirements

Install

pip install "git+https://github.com/musokean/forge.git"                              # core
pip install "handcraft-agent[server] @ git+https://github.com/musokean/forge.git"    # + HTTP API
forge                  # interactive chat
forge "your question"  # one-shot Q&A
forge --web            # browser chat UI (zero-dependency HTTP server)
forge --voice          # voice mode (needs the voice extras)
forge --serve          # HTTP API on 127.0.0.1:8080 — interactive docs at /docs

CI

CI

Tests pass on Python 3.9 / 3.11 / 3.13 (193 tests) + release-copy build check + model smoke.

Still open

Docker sandbox and full logging (the remaining M5 tail), hardware Phase 1 (real device link over serial/MQTT), voice Phase 2/3 (barge-in, streaming, lower latency), and the client-executor module.