Repository navigation
Releases: knight22-21/DevAgent
Release list
v1.5.0
Changelog
All notable changes to this project will be documented in this file.
The format is based on Keep a Changelog,
and this project adheres to Semantic Versioning.
[1.5.0] - 2026-09-30
Added
REPL command expansion (Phases 32, 34, 36)
/fast— toggles fast mode (streaming vs. single-shot) on the fly without restarting the session/compact [focus]— compresses session history to a summary; optionalfocusstring biases what's kept/recap— prints a structured summary of the current session: files touched, tools called, decisions made/branch— shows the current git branch and lets you switch without leaving the REPL/theme <name>— switches the Rich colour theme live (dark, light, monokai, etc.)
Fan-out and background agents (Phases 37, 38)
/batch <glob>— fans out a sub-agent per matched file in parallel; results are collected and rendered as a table; each sub-agent runs in its own worktree whenisolation = "worktree"is set/background <task>— fires a task as a background daemon agent and returns immediately; status visible via/tasks
Notebook tools hardened (Phase 33)
- Suppressed the
nbformatversion warning that leaked into tool output on every call - Added 14 new notebook-tool tests covering cell execution, output capture, and round-trip save
emit_json output mode (Phase 34)
devagent do "<task>" --emit-json— prints a machine-readable JSON envelope ({event, text, tool_calls, …}) instead of Rich-formatted output; suitable for piping into jq or CI scripts
.mcp.json project MCP server config (Phase 35)
- Projects can drop a
.mcp.jsonat their root to declare MCP servers;MCPManagerloads and connects them automatically at session start - Supports
stdio,websocket, andssetransports in the same file devagent mcp listreads.mcp.jsonand shows connection status for each entry
WebSocket + SSE MCP transport (Phase 39)
connect_websocket(name, url)andconnect_sse(name, url, headers)async context managers indevagent/mcp/transports/connect_entry()dispatcher routestransport = "websocket"or"sse"entries from.mcp.jsonautomaticallyMCPManager._connect_remote_servers()best-effort connects all non-stdio servers at startup; unreachable servers are skipped with a warning- WebSocket import is lazy so older
mcpbuilds that lack the module still work
OAuth 2.0 PKCE auth for MCP servers (Phase 40)
- Full authorization-code + PKCE flow (
pkce_authorize): generates a code verifier, computes the SHA-256 challenge, opens the browser, starts a local redirect server, exchanges the code for tokens refresh_token_flow()— silently refreshes an expired access token- Platform keyring token cache (
devagent/mcp/auth/token_cache.py): tokens survive process restarts; 60-second expiry buffer prevents last-second failures; graceful no-op when keyring is unavailable OAuthConfigdataclass +MCPServerEntry.authfield — declare auth in.mcp.jsonunder anauth:key
Diff viewer before writes + --add-dir (Phase 41)
--diff-previewflag ondevagent run— before eachwrite_fileoredit_file, shows a syntax-highlighted unified diff (Rich + monokai) and prompts the user to accept or reject; the file is not touched on rejection--add-dir <path>flag (repeatable) onrunanddo— grants the agent read/write access to directories outsideproject_root; path-traversal guard still applies across all allowed rootsdiff_confirm_fncallback isNonein non-interactive /--baremode so batch workflows are unaffected
[1.4.0] - 2026-09-29
Added
Ollama Cloud auth (Phase 27)
devagent initnow accepts an Ollama Cloud API key andbase_url; auth is passed asAuthorization: Bearer <key>header via_ollama_client()helper_validate_ollama()skips the local HTTP check for the cloud endpoint and returns a fast confirmation- Cloud config fields (
base_url,api_key) survive config save/reload
Ollama Cloud model picker (Phase 28)
list_ollama_models()onLLMClient— fetches available models from the configured endpoint (local or cloud), returns[]on failuredevagent initCloud section shows a numbered list of available models; pick by number or type a name/modelwith no argument now shows a numbered list of available Ollama models; current model is marked with←/model <N>switches to the Nth model by index
REPL tab completion and session naming (Phase 29)
prompt_toolkitword completer for all/commands — press Tab to completedevagent run --name <name>— start a named session--resume <name>— resume by human-readable name (falls back to ID prefix)/rename <name>in REPL — rename the current session- Intro banner trimmed to one hint line; command list is discoverable via Tab
Auto-suggest and /sessions (Phase 30)
complete_while_typing=True— suggestions appear as you type without pressing Tab/sessionsREPL command — lists all sessions for the current project with name, short ID, last-updated time, and current-session markerSessionManager.list_by_project()— filters session store by project path
Fixed
/sessionscrash whenupdated_atwas a Unix float instead of a string- Doubled API key in config when
devagent initwas run twice
Changed
devagent chatis now fully functional — loads the selected.mdreport into a DevAgentSession REPL so users can ask questions interactively without re-runningdevagent analyze
[1.3.0] - 2026-09-25
Added
Agent definitions and worktree isolation (Phase 18–19)
AgentDef.isolation = "worktree"—devagent agent runcreates a fresh git worktree for an agent and cleans it up when done; prevents concurrent agents from clobbering each other's editsdevagent init-project --generate— LLM analysespyproject.toml,README.md, and directory layout to write a project-specificDEVAGENT.md; falls back to a static template if the LLM is offlineload_devagent_md()now walks ancestor directories (outermost first) so nested projects inherit parent-level instructions
Per-agent persistent memory (Phase 20)
AgentDef.memory = "project"wires.devagent/agent-memory/<name>/memory.md— the agent reads it at start and writes to it via theremember_persistenttool; survives between invocations
HTTP and prompt hook types (Phase 21)
- New hook types:
http(POST JSON to an external endpoint) andprompt(inject text into the next LLM call) - Lifecycle events
session_startandsession_endnow fire for all hook types, not justshell devagent hooks test <event> <tool>— dry-run any hook without starting a full session
Structured output flag (Phase 22)
devagent do "<task>" --json-schema '<schema>'— validatesFinalAnswerEvent.textagainst a JSON Schema; retries once with an augmented prompt on mismatch; exits with code 2 if the second attempt still fails- Requires
jsonschema>=4.0.0(added as a dependency)
/model mid-session switch (Phase 23)
/model <provider/model>insidedevagent runhot-swaps the active LLM for the rest of the session without restarting/model <model>(no slash) keeps the current provider and changes only the model name- Validates provider names; supported:
ollama,anthropic,openai,gemini,groq
autofix-pr watch mode (Phase 24)
devagent autofix-pr <pr-url>— polls a PR every N seconds (default 120); on each new failed CI run, fetches job logs and fires aDevAgentSessionto fix and push; on each new review comment, fires a session to address it--poll-intervaland--max-pollsflags for CI usage- Existing failures and comments at startup are seeded into
seensets so they are not re-processed
@agent-name REPL mention syntax (Phase 25)
@code-reviewer please check the auth changesinsidedevagent runspawns the named agent as a background daemon thread; progress is tracked via the task store- Respects
AgentDef.isolation = "worktree"for isolated execution - Use
/tasksto check status; agents are looked up from.devagent/agents/*.tomland the user config dir
Plugin bundle format (Phase 26)
PluginBundledataclass — third-party packages declare tools, skills, hooks, and MCP servers via thedevagent.pluginsentry-point group in theirpyproject.tomldevagent plugins list— discovers and renders all installed plugin bundles in a tabledevagent plugins install <package>— wrapspip installand reports success/failuredevagent tasks— list all background agent tasks in the current process (running, done, failed)
[1.2.0] - 2026-09-25
Added
- 24-task benchmark set covering Python (20), JavaScript (2), and Go (2) fixture projects
- JS fixture project (
js_project/) withadd-feature-001andtest-write-js-001tasks - Go fixture project (
go_project/) withadd-feature-go-001andtest-write-go-001tasks devagent bench leaderboardcommand — groups results by (model, provider), shows best and latest scoredevagent bench leaderboard --remote— fetches results frombench-resultsgit branchdevagent bench leaderboard --output <file>— writes markdown leaderboard to a fileLEADERBOARD.mdseeded with live benchmark scores (gpt-oss:20b: 21/24, llama3.2:3b: 9/20)- CI auto-update step in
bench-persist.yml: regeneratesLEADERBOARD.mdon main after every push devagent bench history --remoteflag for fetching result files from thebench-resultsbranch- Go binary validation in canary workflow (ensures Go fixtures compile before running)
- Native dry-run results persisted to
bench-resultsbranch alongside canary JS...
v1.3.0 — Worktree isolation, hooks, plugin bundles, autofix-pr, @mention
What's new in v1.3.0
Nine new capabilities shipped across Phases 18–26.
Worktree isolation for agents
AgentDef.isolation = "worktree" runs a named agent in a fresh git worktree so parallel agents never conflict. The worktree is created, used, and cleaned up automatically.
LLM-generated DEVAGENT.md
devagent init-project --generate uses the active LLM to analyse your codebase (pyproject.toml, README.md, directory layout) and write a project-specific DEVAGENT.md. Falls back to a static template when offline. The loader now also walks ancestor directories (outermost first) so nested projects inherit parent-level instructions.
Per-agent persistent memory
AgentDef.memory = "project" wires .devagent/agent-memory/<name>/memory.md — the agent reads it at session start and appends to it via the remember_persistent tool. Survives between invocations.
HTTP and prompt hook types
Two new hook types: http (POST JSON to an external endpoint) and prompt (inject text into the next LLM call). Lifecycle events session_start and session_end now fire for all hook types. New command: devagent hooks test <event> <tool>.
Structured output (--json-schema)
devagent do "extract the API endpoints from src/" --json-schema '{"type":"array","items":{"type":"string"}}'Validates FinalAnswerEvent.text against a JSON Schema. Retries once with an augmented prompt on mismatch. Exits code 2 if the second attempt still fails.
/model mid-session switch
> /model anthropic/claude-opus-4-8
Model switched to anthropic/claude-opus-4-8
Hot-swap the active LLM without restarting the session. Supports provider/model and model-only forms.
autofix-pr watch mode
devagent autofix-pr https://github.com/owner/repo/pull/42Polls a PR in a loop. On each new failed CI run it fetches job logs and fires an agent to fix and push. On each new review comment it fires an agent to address it. Supports --poll-interval and --max-polls.
@agent-name REPL mention syntax
> @code-reviewer please check the auth changes
Spawned @code-reviewer [a3f2d1c0]: please check the auth changes
Check progress with /tasks
Spawns a named agent definition as a background daemon thread. Use /tasks to monitor. Respects AgentDef.isolation = "worktree".
Plugin bundle format
Third-party packages can now extend DevAgent with tools, skills, hooks, and MCP servers by declaring a devagent.plugins entry point:
devagent plugins install devagent-docker
devagent plugins listFull changelog: CHANGELOG.md
v1.1.0 — Ollama Cloud + 20/20 benchmark
What's new in v1.1.0
Ollama Cloud support
DevAgent now works with Ollama Cloud out of the box. Set OLLAMA_HOST=https://ollama.com and OLLAMA_API_KEY in a .env file at your project root — the agent picks them up automatically at startup via python-dotenv. No code changes, no new flags, same Ollama API.
Available cloud models include gpt-oss:20b, nemotron-3-nano:30b, gemma4:31b, and 15+ others. See your account at ollama.com for the full list.
20/20 benchmark pass rate
With gpt-oss:20b on Ollama Cloud, DevAgent now passes all 20 tasks in the built-in benchmark — up from 9/20 (45%) with llama3.2:3b running locally. See BENCHMARKS.md for the full breakdown.
Benchmark results always saved
devagent bench native --live now always writes a timestamped JSON to benchmarks/results/ after each run. A rolling partial file is updated after each task so a crash mid-run does not lose completed results. The --output-json flag has been removed (it now happens automatically).
New docs
- INTEGRATIONS.md — setup guides for Claude Code, ChatGPT, Antigravity, GitHub Copilot, Cursor, Windsurf, VS Code, JetBrains, and Zed (moved from README)
- BENCHMARKS.md — per-task results, failure analysis, and how to run the benchmark against any model
Version 1.0.0
Added
- Web search —
web_searchtool queries Brave Search API or any self-hosted SearXNG instance. Configured viacfg.brave.api_keyorcfg.searchx.base_url. Provider can be overridden per call. - URL fetching —
fetch_urltool fetches any URL and returns plain text. HTML is stripped automatically using a stdlib parser. No extra dependencies required. - Image understanding —
read_imagetool loads PNG, JPEG, GIF, or WebP files and passes them to the LLM as native image content blocks (Anthropic and OpenAI). Gracefully degrades on providers without vision support. - Jupyter notebook tools —
notebook_read,notebook_edit, andnotebook_runlet the agent inspect, modify, and execute notebooks in place. Requiresnbformat; execution delegates tojupyter nbconvert. - Session todo list —
todo_writeandtodo_readtools give the agent a persistent task list scoped to the current session, backed by the existing SQLite store. - Effort levels —
--effort low|medium|high|xhigh|maxcontrolsmax_tokensand temperature per run. Changeable mid-session with/effort. - Anthropic extended thinking — activated at
xhighormaxeffort whenextended_thinking = truein config. Thinking blocks are surfaced in the terminal and instream-jsonoutput. - Bare mode —
--bareflag skips DEVAGENT.md injection, CodePrism overlay, and the permission gate. Intended for scripted or CI use. - stream-json output —
--output-format stream-jsonemits one JSON line per agent event to stdout, making it easy to pipe agent output into CI tooling or other processes. /effortand/thinkREPL commands — change effort level and toggle extended thinking live without restarting the session./undoand/diffREPL commands —/undorolls back the last file write;/diffshows a unified diff of all files changed this session.- Shell escape in REPL — prefix any line with
!to run a shell command inline without leaving the agent session. - Inline diff rendering — file edits and writes emit a unified diff; the terminal renders it with colour-coded
+/-lines immediately after each tool call. - Sub-agent spawning —
spawn_agenttool lets the agent delegate subtasks recursively to independent sub-agents. - Peer result sharing —
read_peer_resultslets orchestrated workers inspect outputs from completed sibling tasks before starting their own work. - Dependency context injection — workers that depend on other tasks automatically receive a summary of those tasks' outputs in their system prompt.
- DEVAGENT.md —
devagent init-projectscaffolds a project instruction file and.devagent/directory. The agent reads and injects it into every session's system prompt. - Project memory — named facts can be persisted across sessions in
.devagent/memory.jsonand are injected as context on each run. --allow-tools— shorthand flag to auto-allow specific tool names without writing full allow rules.
Changed
SearchXConfiggains abase_urlfield (defaulthttp://localhost:8888) so self-hosted SearXNG instances can be pointed at from TOML config.build_registrynow wires the active LLM provider and session ID through to vision and todo tools respectively.
Fixed
- Session ID and provider were not passed from
DevAgentSessionintobuild_registry, so vision tool image encoding and session-scoped todo storage were silently skipped. Now correctly wired.
Version 0.5.0
Added
- Permission gate with allow/deny rules for tool calls (--allow, --deny, --interactive-approval on devagent run)
- PermissionManager with glob pattern matching (e.g. --allow write_file:src/**)
- Agent loop pauses and prompts before executing unmatched tool calls
- POST /api/v1/sessions/{id}/approve now resolves live pending tool calls via call_id
- ApprovalNeededEvent emitted by the loop when approval is required
- Session-keyed permission registry so the REST API can reach the active manager
Changed
- devagent serve upgraded to FastAPI + uvicorn with full REST and WebSocket support
- WS /ws/v1/{session_id} streams live events from DB at 500ms polling
- GET /api/v1/sessions, /events, /memory, /orchestrate/graph endpoints added
- Session context auto-compression and multi-agent orchestration (devagent orchestrate)
Fixed
- Session store schema now initialised on app startup, fixing CI failures on fresh environments
Version 0.3.0
Added
- New
devagent watchcommand to monitor a GitHub repository for newly opened issues - Watcher storage, scheduler, and reporting flow for recurring repository health checks
- Cross-issue conflict detection for files touched by multiple open issues
- Watcher-specific terminal rendering for watched repos, health reports, and stored analyses
Changed
- GitHub client now supports listing repository issues for watcher checks
- Configuration now includes watcher defaults such as interval, labels, and cross-conflict behavior
- Added
apschedulerruntime support and enabled automatic asyncio handling for pytest
Fixed
- Added watcher-focused tests covering conflict detection, analysis building, and watcher storage
Version 0.2.0
Added
- Direct GitHub issue and pull request URL input for
devagent analyze --url - Interactive terminal chat sessions for exploring a generated gap analysis
- Compact chat-focused rendering and conversation history support
Changed
devagent analyzecan now jump straight into chat with--chat- Pull request URLs can be analyzed as specs
- Spec analysis server now uses FastMCP
Fixed
- Legacy
SPECSYNC_CONFIG_PATHconfig override still works - Test imports and assertions were aligned with the
devagentpackage naming