Local-First · Lightweight · Complete · Production-Ready general-purpose Agent framework, written in Go.
English | 简体中文
Most Agent frameworks are either bloated (screens of abstraction, hundreds of dependencies) or too thin (barely runs a demo). harness9 takes the middle path:
| Principle | Description |
|---|---|
| Local-First | All data lives on your machine (SQLite, tool_results, plans); tools run in a local Docker container. No cloud dependency, code never leaves your machine. |
| Lightweight | Minimal abstraction layers, straightforward code, very few direct dependencies. |
| Complete | Covers every core module an Agent needs to run. |
| Production-Ready | Error recovery, context management, timeout control, concurrent tool execution — production-grade, not a demo. |
# Install
curl -fsSL https://raw.githubusercontent.com/ZhangShenao/harness9/master/scripts/install.sh | bash
# Configure your API key
export OPENAI_API_KEY="sk-..."
# cd into your project and launch
cd /your/project && harness9
# See all CLI flags
harness9 --help
# Print the version
harness9 --versionFor full install options, Anthropic/OpenRouter/OrcaRouter configuration, AGENTS.md setup, and FAQ, see the Quick Start Guide.
Each feature below links to its full technical writeup on the documentation site.
- Full-screen TUI — Bubbletea-based, welcome/conversation dual-phase, streaming output, live tool spinners, Tab-completion.
- Shell execution (
!prefix) — run Bash commands straight from the input box; output is injected into the LLM context automatically. - Context Engineering — SQLite-backed session persistence, LLM-summarization compaction at an 80% threshold.
- Long-Term Memory — cross-session memory persisted to SQLite + FTS5, MEMORY.md materialized view injected into every prompt.
- Agent Skills — Progressive Disclosure: domain knowledge loaded on demand, keeping the system prompt lean.
- Human-in-the-Loop permissions — a rule engine auto-classifies risk; only genuinely risky actions pause for approval.
- Planning module — planning is a native capability: the LLM plans complex tasks via
plan_writeon its own, with checkpoint persistence, compaction immunity, and a stagnation detector. - File system capabilities — OffloadHook moves oversized tool output to disk; FilePlanWriter persists plans as markdown.
- Sub-Agent delegation — delegate well-scoped subtasks to isolated sub-agents with restricted tool sets.
- Observability — OpenTelemetry spans + metrics across the engine, LLM calls, and tool execution; ships with a Langfuse/Grafana/Jaeger-ready exporter.
- Test & Eval — deterministic
ScriptedProvider+ assertion framework + a 24-case golden dataset gating CI. - Sandbox — every tool call runs inside a locked-down Docker container by default, with automatic fallback to local execution.
- AutoDev (
/autodev) — a self-hosted development loop: clarify requirements → confirm a spec → delegate to a dev sub-agent that codes, tests, and opens the PR. - MCP integration — connect any Model Context Protocol server via
.mcp.json; tools appear transparently in the registry. - Web search & fetch —
web_search/web_fetchtools with SSRF hardening, no API key required. - Standard ReAct loop — one LLM call per turn with the full tool list; concurrent tool execution and self-healing (tool errors round-trip back to the LLM as observations, triggering automatic retries).
- Dual run modes — blocking
Runand streamingRunStream, sharing the same engine instance.
| Module | Description |
|---|---|
| TUI | Full-screen Bubbletea TUI: dual-phase, streaming output, spinner + precise timing, Tab completion, live token usage, shell mode. |
| Engine | Standard ReAct main loop, blocking + streaming, event stream (token updates, compaction, tool results, thinking deltas). |
| Hooks | Tool interceptors: HookRegistry (onion model) + OffloadHook + FilePlanWriter + DangerHook. |
| Permission | Human-in-the-loop: PermissionHook (JSON rules) + 5-option approval dialog + dynamic allowlist + hard-protected sensitive paths. |
| Sub-Agent | Task delegation: built-in general-purpose sub-agent, file-defined agents (.harness9/agents/*.md), foreground/background task tool, @agent direct invocation. |
| Planning | Native planning capability: PlanStore (session-level state machine), plan_write tool with anti-cheat validation, write-time checkpointing, compaction immunity, sub-agent isolation, auto-continue + stagnation detection. |
| Memory | Session persistence (SQLite WAL), SummarizationCompactor (default) + TokenBudgetCompactor (fallback). |
| LTM | Long-term memory store (SQLite + FTS5), MEMORY.md materialized view, extractor, Phase 3 seams (Provider/Embedder/Consolidator). |
| Context | System prompt assembly: base + AGENTS.md + skills index + planning/offload/sandbox/LTM sections. |
| Skills | Skill parsing, indexing, on-demand loading (use_skill tool). |
| Provider | Unified LLM interface, OpenAI/Anthropic adapters, real token usage extraction. |
| Schema | Shared core data types (Message, ToolCall, Usage, etc.). |
| Tools | Tool registry + built-ins (bash, read_file, write_file, edit_file, plan_write, memory_write/search, web_search/web_fetch). |
| Sandbox | Docker-level isolation: process sandboxing, per-agent containers, orphan reaping; on by default. |
| Observability | OpenTelemetry tracing/metrics across engine, LLM calls, and tools; noop by default. |
| Evals | Automated evaluation framework, golden dataset, CI quality gate. |
| MCP | Model Context Protocol client integration, transparent tool injection. |
| AutoDev | Self-hosted development loop (/autodev skill + dev sub-agent). |
| Env | Zero-dependency .env loader. |
| Framework | Origin | Difference from harness9 |
|---|---|---|
| DeepAgents | LangChain | Python, graph orchestration (LangGraph StateGraph); harness9 is an explicit Go ReAct loop with no graph engine dependency. |
| OpenHarness | HKUDS | Python, asyncio concurrency; harness9 uses goroutines natively. |
| OpenCode | Anomaly | TypeScript, delegates loop control to the Vercel AI SDK; harness9 owns its loop end to end. |
| OpenClaw | OpenClaw | TypeScript, multi-agent routing via the AI SDK; harness9 is a native Go single-agent ReAct loop. |
| HermesAgent | NousResearch | Python, ThreadPool concurrency, 3-tier context compaction; harness9 is lighter with goroutine concurrency. |
| Claude Agent SDK | Anthropic | Official SDK, Anthropic-only, black-box loop; harness9 supports multiple providers with a transparent, controllable loop. |
| OpenAI Agent SDK | OpenAI | Python, Handoffs multi-agent, depends on the OpenAI Compaction API; harness9 is native Go with its own compaction, no cloud API dependency. |
Since 2026-06-30 GitHub restricts stargazers API access to owners/collaborators, so third-party badges like star-history.com are broadly broken. The chart above is generated and committed daily as a static file by
.github/workflows/star-history.ymlusing the repo's ownGITHUB_TOKEN, so it doesn't depend on any third-party service's uptime.
harness9's Agent capability is evaluated on SWE-bench Lite. See the full benchmark methodology and run instructions for dataset setup, running cmd/swebench, and scoring results.
Full documentation, architecture write-ups, and technical blog posts live on the documentation site, including the AGENTS.md project conventions used to drive this repo's own Agent.


