Skip to content

Repository files navigation

agent-go

ago is a semantic edit protocol for Go: agents query the typechecked workspace and submit compiler-checked mutations instead of editing text.

Go as a language for agents. The protocol puts the Go toolchain between the model and the files: an edit that does not typecheck cannot reach disk. LSP-backed tools have the same compiler diagnostics available and can still write broken code, because checking is advisory there. Here invalid edits are unrepresentable, not merely checked.

The thesis under test: weak local models become effective repo-scale Go editors when raw file editing is replaced by semantic queries and validated mutations with typed rejections. The bench in this repo measures exactly that, on tasks mined from real commits in traefik, vault, and boundary.

Install

brew install guygrigsby/tap/ago
curl -fsSL https://ago.aeryx.ai/install.sh | sh
go install github.com/guygrigsby/agent-go/cmd/ago@latest

Windows: grab the zip from the releases page. Every channel ships the agent skill in the binary; ago skill install drops it into ~/.claude/skills for shell-capable coding agents. In an existing repo, ago init wires the agent files without touching the module: AGENTS.md (read natively by Codex, Copilot, Cursor, Windsurf, Zed, Claude Code, and the rest of the AGENTS.md ecosystem), MCP wiring, and a GEMINI.md pointer for the one holdout.

Or from a clone:

go build -o ago ./cmd/ago

Requires Go 1.26+. The daemon auto-spawns on first use, one per workspace, and exits after five minutes idle; there is nothing to start or configure.

How it works

One binary, three fronts over a per-workspace daemon (auto-spawned on first use, unix socket, idle exit):

$ ago --help
semantic edit protocol for Go: query the typechecked workspace, submit compiler-checked mutations

Usage:
  ago [command]

reads and setup:
  help        the versioned op catalog: args, examples, ceilings
  init        scaffold an agent-first module: MCP wiring, AGENTS.md
  inspect     kind, signature, position, doc for one symbol
  query       semantic questions by --kind (callers, implementations, ...)
  refs        every reference, tests included
  search      name fragment to exact symbol addresses
  status      load or refresh the snapshot: packages, files, errors
  view        declaration as annotated text with node handles

mutations (validated before anything touches disk):
  add-param   add a parameter, call sites updated with --default
  patch       ordered multi-op edit, atomic and generation-checked
  rename      rename a symbol, every reference proven to resolve
  set-body    replace a function body, typechecked first
  test        go test, scoped, structured pass/fail
  upsert      add or replace a whole declaration

lifecycle:
  daemon      run the workspace daemon in the foreground
  mcp         serve the MCP tools over stdio
  skill       install or print the embedded agent skill

Additional Commands:
  completion  Generate the autocompletion script for the specified shell
  help        Help about any command

Flags:
  -C, --dir string   workspace directory (default ".")
  -h, --help         help for ago

Use "ago [command] --help" for more information about a command.

A few invocations:

ago search -s MaxEntries
ago refs -p <pkg> -s <sym>
ago query -k implementations -p <pkg> -s <iface>
ago rename -p <pkg> -s DefaultLimiterMaxEntries --to DefaultLimiterMaxQuotas
echo 'return v << 1' | ago set-body -p <pkg> -s Double --body-file -
ago patch --body-file patch.json

Agents connect over MCP (ago mcp); ago init writes the wiring. For agents with a shell (Claude Code and kin), ago skill install teaches the CLI workflow in any Go repo with no per-repo setup; the skill ships embedded in the binary. patch is the full language: 14 statement ops, composable decl ops, table-driven test ops; rename/set-body/add-param/upsert are one-op sugar over it. Full catalog via ago help.

Every mutation validates before anything touches disk: the daemon re-typechecks only the affected packages against its in-memory graph (~200ms for a 29-reference rename on a 533-package repo) and rejects with the compiler's own diagnostics:

{
  "status": "rejected",
  "reason": "edit does not typecheck",
  "diagnostics": [
    {"pos": "config.go:114:29", "msg": "undefined: slices"},
    {"pos": "config.go:9:2", "msg": "\"reflect\" imported and not used"}
  ]
}

Rejections are data, not errors: the agent loop is query, mutate, and on rejection adjust and retry. Most rejections carry possible_repairs, complete paste-ready calls: the corrected mutation, the re-view that refreshes handles, the search that locates an undefined identifier. Resending a rejected call unchanged gets an escalated rejection instead of the same one forever. Rename also proves every rewritten reference still resolves to the renamed symbol, so shadowing capture is rejected even when the compiler is satisfied.

Bench

bench/ holds the raw-vs-semantic comparison: same local model, same harness (opencode), same mined tasks (rename and add-param kinds, from traefik, vault, boundary, and cobra); one mode gets shell and file editing, the other gets only the protocol.

The oracle is the bench's third arm: a scripted replay of each task's ground-truth commit through the protocol itself, no model in the loop. Every mined task carries the real change its source commit made; the oracle submits that change as protocol calls and must reach green. A task enters the roster only after the oracle certifies it, so "the model failed" is never confused with "the task was impossible". Oracle rejections are findings (an engine bug, a bad extraction, or a named protocol ceiling), its wall time is the task's time-to-green floor, and its accepted transcripts seed the fine-tuning corpus.

Scoring is goal predicate + typecheck + scoped tests under a time cap; the tests gate counts only where a pristine worktree passes the same tests. Every episode records its transcript, config, diff, score, token usage, failure kind, and the daemon's per-request log under bench/results/. Serving setups are pinned as named profiles in bench/profiles.json and embedded in each run's run.json.

# smoke: one certified task per kind, smallest repo, for first contact
AGO_BENCH_SUITE=smoke AGO_BENCH_PROFILE=<name> AGO_BENCH_SCRATCH=<clones dir> \
go test ./bench -bench Rename -benchtime 1x -timeout 0

# model round, k=3
AGO_BENCH_PROFILE=glm-flash AGO_BENCH_SCRATCH=<clones dir> \
go test ./bench -bench Rename -benchtime 3x -timeout 0

# oracle sweep: no model, episodes run in parallel
AGO_BENCH_MODES=oracle AGO_BENCH_SCRATCH=<clones dir> \
go test ./bench -run OracleSweep -parallel 20 -timeout 0

# mine candidate tasks from any clone; report across runs
go run ./cmd/bench mine -scratch <clones dir> <repo>
go run ./cmd/bench report bench/results/<run> ...

Development

go test ./...        # full suite; snapshot tests take ~40s
gofmt -l .           # must be empty

Rules of the road:

  • TDD: write the failing test first, watch it fail, then implement. Every op and repair in the tree landed that way.
  • Docs are drift-guarded by tests: the op catalog must match docs/specs/surface.md, help examples must be accepted against the test fixture, and README invocations must match the real CLI dispatch. Change behavior and the guard tells you which doc to touch.
  • The demo fixture (internal/snapshot/testdata/demo) is copied per test; never mutate it in place. demo/lib is frozen behind a recorded hash (TestDemoLibFixtureFrozen) because the view-handle tests depend on its exact layout. demo/sig is the place for new fixture shapes.
  • Race-sensitive changes (the parallel retypecheck) should run under go test -race ./internal/snapshot.

Docs

  • docs/specs/language.md: the full op catalog, decl through test ops
  • docs/specs/surface.md: shipped and foreseen calls, drift-guarded by test
  • docs/specs/protocol.md: protocol semantics and guarantees
  • docs/specs/bench.md: bench design
  • docs/specs/plan.md: status and build order
  • docs/optimizations/: per-model serving research and the cross-model build list
  • docs/tenets.md: the engineering principles this repo lives by
  • docs/model-strategies.md: brainstorm on coding with local models
  • docs/adr/: architecture decisions
  • idea.md: the original thesis

Work is tracked in beads: bd ready for the queue, bd list for everything; issues sync through .beads/issues.jsonl.

Experimental. Interfaces change without notice.

Inspired by zero, a programming language for agents.

About

ago: a semantic edit protocol for Go. Agents query the typechecked workspace and submit compiler-checked mutations instead of editing text.

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages