Repository navigation
Changelog
All notable changes to this project will be documented in this file.
The format is based on Keep a Changelog,
and this project adheres to Semantic Versioning.
[1.5.0] - 2026-09-30
Added
REPL command expansion (Phases 32, 34, 36)
/fast— toggles fast mode (streaming vs. single-shot) on the fly without restarting the session/compact [focus]— compresses session history to a summary; optionalfocusstring biases what's kept/recap— prints a structured summary of the current session: files touched, tools called, decisions made/branch— shows the current git branch and lets you switch without leaving the REPL/theme <name>— switches the Rich colour theme live (dark, light, monokai, etc.)
Fan-out and background agents (Phases 37, 38)
/batch <glob>— fans out a sub-agent per matched file in parallel; results are collected and rendered as a table; each sub-agent runs in its own worktree whenisolation = "worktree"is set/background <task>— fires a task as a background daemon agent and returns immediately; status visible via/tasks
Notebook tools hardened (Phase 33)
- Suppressed the
nbformatversion warning that leaked into tool output on every call - Added 14 new notebook-tool tests covering cell execution, output capture, and round-trip save
emit_json output mode (Phase 34)
devagent do "<task>" --emit-json— prints a machine-readable JSON envelope ({event, text, tool_calls, …}) instead of Rich-formatted output; suitable for piping into jq or CI scripts
.mcp.json project MCP server config (Phase 35)
- Projects can drop a
.mcp.jsonat their root to declare MCP servers;MCPManagerloads and connects them automatically at session start - Supports
stdio,websocket, andssetransports in the same file devagent mcp listreads.mcp.jsonand shows connection status for each entry
WebSocket + SSE MCP transport (Phase 39)
connect_websocket(name, url)andconnect_sse(name, url, headers)async context managers indevagent/mcp/transports/connect_entry()dispatcher routestransport = "websocket"or"sse"entries from.mcp.jsonautomaticallyMCPManager._connect_remote_servers()best-effort connects all non-stdio servers at startup; unreachable servers are skipped with a warning- WebSocket import is lazy so older
mcpbuilds that lack the module still work
OAuth 2.0 PKCE auth for MCP servers (Phase 40)
- Full authorization-code + PKCE flow (
pkce_authorize): generates a code verifier, computes the SHA-256 challenge, opens the browser, starts a local redirect server, exchanges the code for tokens refresh_token_flow()— silently refreshes an expired access token- Platform keyring token cache (
devagent/mcp/auth/token_cache.py): tokens survive process restarts; 60-second expiry buffer prevents last-second failures; graceful no-op when keyring is unavailable OAuthConfigdataclass +MCPServerEntry.authfield — declare auth in.mcp.jsonunder anauth:key
Diff viewer before writes + --add-dir (Phase 41)
--diff-previewflag ondevagent run— before eachwrite_fileoredit_file, shows a syntax-highlighted unified diff (Rich + monokai) and prompts the user to accept or reject; the file is not touched on rejection--add-dir <path>flag (repeatable) onrunanddo— grants the agent read/write access to directories outsideproject_root; path-traversal guard still applies across all allowed rootsdiff_confirm_fncallback isNonein non-interactive /--baremode so batch workflows are unaffected
[1.4.0] - 2026-09-29
Added
Ollama Cloud auth (Phase 27)
devagent initnow accepts an Ollama Cloud API key andbase_url; auth is passed asAuthorization: Bearer <key>header via_ollama_client()helper_validate_ollama()skips the local HTTP check for the cloud endpoint and returns a fast confirmation- Cloud config fields (
base_url,api_key) survive config save/reload
Ollama Cloud model picker (Phase 28)
list_ollama_models()onLLMClient— fetches available models from the configured endpoint (local or cloud), returns[]on failuredevagent initCloud section shows a numbered list of available models; pick by number or type a name/modelwith no argument now shows a numbered list of available Ollama models; current model is marked with←/model <N>switches to the Nth model by index
REPL tab completion and session naming (Phase 29)
prompt_toolkitword completer for all/commands — press Tab to completedevagent run --name <name>— start a named session--resume <name>— resume by human-readable name (falls back to ID prefix)/rename <name>in REPL — rename the current session- Intro banner trimmed to one hint line; command list is discoverable via Tab
Auto-suggest and /sessions (Phase 30)
complete_while_typing=True— suggestions appear as you type without pressing Tab/sessionsREPL command — lists all sessions for the current project with name, short ID, last-updated time, and current-session markerSessionManager.list_by_project()— filters session store by project path
Fixed
/sessionscrash whenupdated_atwas a Unix float instead of a string- Doubled API key in config when
devagent initwas run twice
Changed
devagent chatis now fully functional — loads the selected.mdreport into a DevAgentSession REPL so users can ask questions interactively without re-runningdevagent analyze
[1.3.0] - 2026-09-25
Added
Agent definitions and worktree isolation (Phase 18–19)
AgentDef.isolation = "worktree"—devagent agent runcreates a fresh git worktree for an agent and cleans it up when done; prevents concurrent agents from clobbering each other's editsdevagent init-project --generate— LLM analysespyproject.toml,README.md, and directory layout to write a project-specificDEVAGENT.md; falls back to a static template if the LLM is offlineload_devagent_md()now walks ancestor directories (outermost first) so nested projects inherit parent-level instructions
Per-agent persistent memory (Phase 20)
AgentDef.memory = "project"wires.devagent/agent-memory/<name>/memory.md— the agent reads it at start and writes to it via theremember_persistenttool; survives between invocations
HTTP and prompt hook types (Phase 21)
- New hook types:
http(POST JSON to an external endpoint) andprompt(inject text into the next LLM call) - Lifecycle events
session_startandsession_endnow fire for all hook types, not justshell devagent hooks test <event> <tool>— dry-run any hook without starting a full session
Structured output flag (Phase 22)
devagent do "<task>" --json-schema '<schema>'— validatesFinalAnswerEvent.textagainst a JSON Schema; retries once with an augmented prompt on mismatch; exits with code 2 if the second attempt still fails- Requires
jsonschema>=4.0.0(added as a dependency)
/model mid-session switch (Phase 23)
/model <provider/model>insidedevagent runhot-swaps the active LLM for the rest of the session without restarting/model <model>(no slash) keeps the current provider and changes only the model name- Validates provider names; supported:
ollama,anthropic,openai,gemini,groq
autofix-pr watch mode (Phase 24)
devagent autofix-pr <pr-url>— polls a PR every N seconds (default 120); on each new failed CI run, fetches job logs and fires aDevAgentSessionto fix and push; on each new review comment, fires a session to address it--poll-intervaland--max-pollsflags for CI usage- Existing failures and comments at startup are seeded into
seensets so they are not re-processed
@agent-name REPL mention syntax (Phase 25)
@code-reviewer please check the auth changesinsidedevagent runspawns the named agent as a background daemon thread; progress is tracked via the task store- Respects
AgentDef.isolation = "worktree"for isolated execution - Use
/tasksto check status; agents are looked up from.devagent/agents/*.tomland the user config dir
Plugin bundle format (Phase 26)
PluginBundledataclass — third-party packages declare tools, skills, hooks, and MCP servers via thedevagent.pluginsentry-point group in theirpyproject.tomldevagent plugins list— discovers and renders all installed plugin bundles in a tabledevagent plugins install <package>— wrapspip installand reports success/failuredevagent tasks— list all background agent tasks in the current process (running, done, failed)
[1.2.0] - 2026-09-25
Added
- 24-task benchmark set covering Python (20), JavaScript (2), and Go (2) fixture projects
- JS fixture project (
js_project/) withadd-feature-001andtest-write-js-001tasks - Go fixture project (
go_project/) withadd-feature-go-001andtest-write-go-001tasks devagent bench leaderboardcommand — groups results by (model, provider), shows best and latest scoredevagent bench leaderboard --remote— fetches results frombench-resultsgit branchdevagent bench leaderboard --output <file>— writes markdown leaderboard to a fileLEADERBOARD.mdseeded with live benchmark scores (gpt-oss:20b: 21/24, llama3.2:3b: 9/20)- CI auto-update step in
bench-persist.yml: regeneratesLEADERBOARD.mdon main after every push devagent bench history --remoteflag for fetching result files from thebench-resultsbranch- Go binary validation in canary workflow (ensures Go fixtures compile before running)
- Native dry-run results persisted to
bench-resultsbranch alongside canary JSON - Multi-task
-tflag (repeatable):devagent bench native -t bug-fix-001 -t refactor-002 - Model/provider metadata embedded in saved benchmark JSON files for leaderboard tracking
Fixed
code-review-001: raisedtimeout_secfrom 60 → 120 to prevent API call timeoutsrefactor-002: removed false task description premise ("multiply bug is already fixed"); agent now runs pytest first to discover and fix failures before adding type hintstest-write-002: added explicit import pattern, "run each function via shell to verify expected values before writing assertions", and raisedtimeout_secto 300- Multi-task
-tflag: switched fromstr | Nonetolist[str] | Noneso multiple-tvalues accumulate correctly - Unicode check marks in canary output replaced with ASCII for cross-platform terminal compatibility
Changed
- Benchmark result JSON format extended:
{"meta": {...}, "results": [...]}envelope (backward compatible with old plain-list format) BenchReport.save_jsonaccepts optionalmodelandproviderparametersBenchReport.render_leaderboardandgenerate_leaderboard_mdadded toreport.py
[1.1.0] - 2026-08-28
Added
INTEGRATIONS.md— setup guides for Claude Code, ChatGPT, Antigravity, GitHub Copilot, Cursor, Windsurf, VS Code, JetBrains, ZedBENCHMARKS.md— full benchmark documentation with per-task results table- Ollama Cloud provider support (
gpt-oss:20b,nemotron-3-nano:30b,gemma4:31b, 15+ hosted models) - First live 20-task benchmark run: gpt-oss:20b scored 20/20 (100%)
- B4 cost-to-correctness sweep infrastructure:
devagent bench sweepwith parameter grid devagent bench historycommand for cross-run trend comparisonexpected_files_touchedenforcement in benchmark runner — reports files missed or extra- Syntax checking for
.pyfiles written during benchmark runs (auto-surfaces errors to model) - Token-savings measurement in
bench_token_usage.pyusing real CodePrism graph queries - Real
call_counttracking for iteration counting in live benchmark runs - Phases 10–14: hooks infrastructure, permission modes, REPL UX improvements, agent definition files
Fixed
- Removed
--skip-legacyfrom CI — all legacy benchmark scripts pass reliably - Benchmark oracle fixes:
DEVAGENT_OUTPUT.txtauto-write, syntax check, task-id filter
Changed
- README overhauled to match the full current feature set (Phases 1–16 complete)
[0.3.0] - 2026-08-13
Added
- New
devagent watchcommand to monitor a GitHub repository for newly opened issues - Watcher storage, scheduler, and reporting flow for recurring repository health checks
- Cross-issue conflict detection to flag files touched by multiple open issues
- Watcher-specific terminal rendering for watched repos, health reports, and stored analyses
Changed
- GitHub client now supports listing repository issues for watcher checks
- Configuration now includes watcher defaults such as interval, labels, and cross-conflict behavior
- Added
apschedulerruntime support and enabled automatic asyncio handling for pytest
Fixed
- Added watcher-focused tests covering conflict detection, analysis building, and watcher storage
[0.2.0] - 2026-08-11
Added
- Direct GitHub issue and pull request URL input for
devagent analyzevia--url - Interactive terminal chat sessions for exploring a generated gap analysis
- Compact chat-focused report rendering and conversation history support
- Test coverage for GitHub URL parsing and chat prompt grounding
Changed
devagent analyzecan now open a chat session immediately with--chat- Pull request URLs are analyzed as specifications using PR title and description
- Spec analysis MCP server integration now uses
FastMCP
Fixed
- Configuration path resolution now supports the legacy
SPECSYNC_CONFIG_PATHoverride - Test imports and assertions were aligned with the
devagentpackage naming
[0.1.1] - 2026-08-10
Fixed
- Initial bug fixes and stability improvements
[0.1.0] - 2026-08-09
Added
- Initial release of DevAgent
- Automated gap analysis between specifications and codebase
- Local LLM support via Ollama
- Model Context Protocol (MCP) integration
- Semantic search with ChromaDB
- Rich terminal UI and Markdown report generation