Skip to content

Releases: knight22-21/DevAgent

v1.5.0

Choose a tag to compare

@knight22-21 knight22-21 released this 30 Sep 06:51

Changelog

All notable changes to this project will be documented in this file.

The format is based on Keep a Changelog,
and this project adheres to Semantic Versioning.

[1.5.0] - 2026-09-30

Added

REPL command expansion (Phases 32, 34, 36)

  • /fast — toggles fast mode (streaming vs. single-shot) on the fly without restarting the session
  • /compact [focus] — compresses session history to a summary; optional focus string biases what's kept
  • /recap — prints a structured summary of the current session: files touched, tools called, decisions made
  • /branch — shows the current git branch and lets you switch without leaving the REPL
  • /theme <name> — switches the Rich colour theme live (dark, light, monokai, etc.)

Fan-out and background agents (Phases 37, 38)

  • /batch <glob> — fans out a sub-agent per matched file in parallel; results are collected and rendered as a table; each sub-agent runs in its own worktree when isolation = "worktree" is set
  • /background <task> — fires a task as a background daemon agent and returns immediately; status visible via /tasks

Notebook tools hardened (Phase 33)

  • Suppressed the nbformat version warning that leaked into tool output on every call
  • Added 14 new notebook-tool tests covering cell execution, output capture, and round-trip save

emit_json output mode (Phase 34)

  • devagent do "<task>" --emit-json — prints a machine-readable JSON envelope ({event, text, tool_calls, …}) instead of Rich-formatted output; suitable for piping into jq or CI scripts

.mcp.json project MCP server config (Phase 35)

  • Projects can drop a .mcp.json at their root to declare MCP servers; MCPManager loads and connects them automatically at session start
  • Supports stdio, websocket, and sse transports in the same file
  • devagent mcp list reads .mcp.json and shows connection status for each entry

WebSocket + SSE MCP transport (Phase 39)

  • connect_websocket(name, url) and connect_sse(name, url, headers) async context managers in devagent/mcp/transports/
  • connect_entry() dispatcher routes transport = "websocket" or "sse" entries from .mcp.json automatically
  • MCPManager._connect_remote_servers() best-effort connects all non-stdio servers at startup; unreachable servers are skipped with a warning
  • WebSocket import is lazy so older mcp builds that lack the module still work

OAuth 2.0 PKCE auth for MCP servers (Phase 40)

  • Full authorization-code + PKCE flow (pkce_authorize): generates a code verifier, computes the SHA-256 challenge, opens the browser, starts a local redirect server, exchanges the code for tokens
  • refresh_token_flow() — silently refreshes an expired access token
  • Platform keyring token cache (devagent/mcp/auth/token_cache.py): tokens survive process restarts; 60-second expiry buffer prevents last-second failures; graceful no-op when keyring is unavailable
  • OAuthConfig dataclass + MCPServerEntry.auth field — declare auth in .mcp.json under an auth: key

Diff viewer before writes + --add-dir (Phase 41)

  • --diff-preview flag on devagent run — before each write_file or edit_file, shows a syntax-highlighted unified diff (Rich + monokai) and prompts the user to accept or reject; the file is not touched on rejection
  • --add-dir <path> flag (repeatable) on run and do — grants the agent read/write access to directories outside project_root; path-traversal guard still applies across all allowed roots
  • diff_confirm_fn callback is None in non-interactive / --bare mode so batch workflows are unaffected

[1.4.0] - 2026-09-29

Added

Ollama Cloud auth (Phase 27)

  • devagent init now accepts an Ollama Cloud API key and base_url; auth is passed as Authorization: Bearer <key> header via _ollama_client() helper
  • _validate_ollama() skips the local HTTP check for the cloud endpoint and returns a fast confirmation
  • Cloud config fields (base_url, api_key) survive config save/reload

Ollama Cloud model picker (Phase 28)

  • list_ollama_models() on LLMClient — fetches available models from the configured endpoint (local or cloud), returns [] on failure
  • devagent init Cloud section shows a numbered list of available models; pick by number or type a name
  • /model with no argument now shows a numbered list of available Ollama models; current model is marked with ←
  • /model <N> switches to the Nth model by index

REPL tab completion and session naming (Phase 29)

  • prompt_toolkit word completer for all / commands — press Tab to complete
  • devagent run --name <name> — start a named session
  • --resume <name> — resume by human-readable name (falls back to ID prefix)
  • /rename <name> in REPL — rename the current session
  • Intro banner trimmed to one hint line; command list is discoverable via Tab

Auto-suggest and /sessions (Phase 30)

  • complete_while_typing=True — suggestions appear as you type without pressing Tab
  • /sessions REPL command — lists all sessions for the current project with name, short ID, last-updated time, and current-session marker
  • SessionManager.list_by_project() — filters session store by project path

Fixed

  • /sessions crash when updated_at was a Unix float instead of a string
  • Doubled API key in config when devagent init was run twice

Changed

  • devagent chat is now fully functional — loads the selected .md report into a DevAgentSession REPL so users can ask questions interactively without re-running devagent analyze

[1.3.0] - 2026-09-25

Added

Agent definitions and worktree isolation (Phase 18–19)

  • AgentDef.isolation = "worktree" — devagent agent run creates a fresh git worktree for an agent and cleans it up when done; prevents concurrent agents from clobbering each other's edits
  • devagent init-project --generate — LLM analyses pyproject.toml, README.md, and directory layout to write a project-specific DEVAGENT.md; falls back to a static template if the LLM is offline
  • load_devagent_md() now walks ancestor directories (outermost first) so nested projects inherit parent-level instructions

Per-agent persistent memory (Phase 20)

  • AgentDef.memory = "project" wires .devagent/agent-memory/<name>/memory.md — the agent reads it at start and writes to it via the remember_persistent tool; survives between invocations

HTTP and prompt hook types (Phase 21)

  • New hook types: http (POST JSON to an external endpoint) and prompt (inject text into the next LLM call)
  • Lifecycle events session_start and session_end now fire for all hook types, not just shell
  • devagent hooks test <event> <tool> — dry-run any hook without starting a full session

Structured output flag (Phase 22)

  • devagent do "<task>" --json-schema '<schema>' — validates FinalAnswerEvent.text against a JSON Schema; retries once with an augmented prompt on mismatch; exits with code 2 if the second attempt still fails
  • Requires jsonschema>=4.0.0 (added as a dependency)

/model mid-session switch (Phase 23)

  • /model <provider/model> inside devagent run hot-swaps the active LLM for the rest of the session without restarting
  • /model <model> (no slash) keeps the current provider and changes only the model name
  • Validates provider names; supported: ollama, anthropic, openai, gemini, groq

autofix-pr watch mode (Phase 24)

  • devagent autofix-pr <pr-url> — polls a PR every N seconds (default 120); on each new failed CI run, fetches job logs and fires a DevAgentSession to fix and push; on each new review comment, fires a session to address it
  • --poll-interval and --max-polls flags for CI usage
  • Existing failures and comments at startup are seeded into seen sets so they are not re-processed

@agent-name REPL mention syntax (Phase 25)

  • @code-reviewer please check the auth changes inside devagent run spawns the named agent as a background daemon thread; progress is tracked via the task store
  • Respects AgentDef.isolation = "worktree" for isolated execution
  • Use /tasks to check status; agents are looked up from .devagent/agents/*.toml and the user config dir

Plugin bundle format (Phase 26)

  • PluginBundle dataclass — third-party packages declare tools, skills, hooks, and MCP servers via the devagent.plugins entry-point group in their pyproject.toml
  • devagent plugins list — discovers and renders all installed plugin bundles in a table
  • devagent plugins install <package> — wraps pip install and reports success/failure
  • devagent tasks — list all background agent tasks in the current process (running, done, failed)

[1.2.0] - 2026-09-25

Added

  • 24-task benchmark set covering Python (20), JavaScript (2), and Go (2) fixture projects
  • JS fixture project (js_project/) with add-feature-001 and test-write-js-001 tasks
  • Go fixture project (go_project/) with add-feature-go-001 and test-write-go-001 tasks
  • devagent bench leaderboard command — groups results by (model, provider), shows best and latest score
  • devagent bench leaderboard --remote — fetches results from bench-results git branch
  • devagent bench leaderboard --output <file> — writes markdown leaderboard to a file
  • LEADERBOARD.md seeded with live benchmark scores (gpt-oss:20b: 21/24, llama3.2:3b: 9/20)
  • CI auto-update step in bench-persist.yml: regenerates LEADERBOARD.md on main after every push
  • devagent bench history --remote flag for fetching result files from the bench-results branch
  • Go binary validation in canary workflow (ensures Go fixtures compile before running)
  • Native dry-run results persisted to bench-results branch alongside canary JS...
Read more

v1.3.0 — Worktree isolation, hooks, plugin bundles, autofix-pr, @mention

Choose a tag to compare

@knight22-21 knight22-21 released this 25 Sep 18:50

What's new in v1.3.0

Nine new capabilities shipped across Phases 18–26.

Worktree isolation for agents

AgentDef.isolation = "worktree" runs a named agent in a fresh git worktree so parallel agents never conflict. The worktree is created, used, and cleaned up automatically.

LLM-generated DEVAGENT.md

devagent init-project --generate uses the active LLM to analyse your codebase (pyproject.toml, README.md, directory layout) and write a project-specific DEVAGENT.md. Falls back to a static template when offline. The loader now also walks ancestor directories (outermost first) so nested projects inherit parent-level instructions.

Per-agent persistent memory

AgentDef.memory = "project" wires .devagent/agent-memory/<name>/memory.md — the agent reads it at session start and appends to it via the remember_persistent tool. Survives between invocations.

HTTP and prompt hook types

Two new hook types: http (POST JSON to an external endpoint) and prompt (inject text into the next LLM call). Lifecycle events session_start and session_end now fire for all hook types. New command: devagent hooks test <event> <tool>.

Structured output (--json-schema)

devagent do "extract the API endpoints from src/" --json-schema '{"type":"array","items":{"type":"string"}}'

Validates FinalAnswerEvent.text against a JSON Schema. Retries once with an augmented prompt on mismatch. Exits code 2 if the second attempt still fails.

/model mid-session switch

> /model anthropic/claude-opus-4-8
Model switched to anthropic/claude-opus-4-8

Hot-swap the active LLM without restarting the session. Supports provider/model and model-only forms.

autofix-pr watch mode

devagent autofix-pr https://github.com/owner/repo/pull/42

Polls a PR in a loop. On each new failed CI run it fetches job logs and fires an agent to fix and push. On each new review comment it fires an agent to address it. Supports --poll-interval and --max-polls.

@agent-name REPL mention syntax

> @code-reviewer please check the auth changes
Spawned @code-reviewer [a3f2d1c0]: please check the auth changes
Check progress with /tasks

Spawns a named agent definition as a background daemon thread. Use /tasks to monitor. Respects AgentDef.isolation = "worktree".

Plugin bundle format

Third-party packages can now extend DevAgent with tools, skills, hooks, and MCP servers by declaring a devagent.plugins entry point:

devagent plugins install devagent-docker
devagent plugins list

Full changelog: CHANGELOG.md

v1.1.0 — Ollama Cloud + 20/20 benchmark

Choose a tag to compare

@knight22-21 knight22-21 released this 07 Sep 14:03
0ce4fbf

What's new in v1.1.0

Ollama Cloud support

DevAgent now works with Ollama Cloud out of the box. Set OLLAMA_HOST=https://ollama.com and OLLAMA_API_KEY in a .env file at your project root — the agent picks them up automatically at startup via python-dotenv. No code changes, no new flags, same Ollama API.

Available cloud models include gpt-oss:20b, nemotron-3-nano:30b, gemma4:31b, and 15+ others. See your account at ollama.com for the full list.

20/20 benchmark pass rate

With gpt-oss:20b on Ollama Cloud, DevAgent now passes all 20 tasks in the built-in benchmark — up from 9/20 (45%) with llama3.2:3b running locally. See BENCHMARKS.md for the full breakdown.

Benchmark results always saved

devagent bench native --live now always writes a timestamped JSON to benchmarks/results/ after each run. A rolling partial file is updated after each task so a crash mid-run does not lose completed results. The --output-json flag has been removed (it now happens automatically).

New docs

  • INTEGRATIONS.md — setup guides for Claude Code, ChatGPT, Antigravity, GitHub Copilot, Cursor, Windsurf, VS Code, JetBrains, and Zed (moved from README)
  • BENCHMARKS.md — per-task results, failure analysis, and how to run the benchmark against any model

Version 1.0.0

Choose a tag to compare

@knight22-21 knight22-21 released this 04 Sep 15:34

Added

  • Web search — web_search tool queries Brave Search API or any self-hosted SearXNG instance. Configured via cfg.brave.api_key or cfg.searchx.base_url. Provider can be overridden per call.
  • URL fetching — fetch_url tool fetches any URL and returns plain text. HTML is stripped automatically using a stdlib parser. No extra dependencies required.
  • Image understanding — read_image tool loads PNG, JPEG, GIF, or WebP files and passes them to the LLM as native image content blocks (Anthropic and OpenAI). Gracefully degrades on providers without vision support.
  • Jupyter notebook tools — notebook_read, notebook_edit, and notebook_run let the agent inspect, modify, and execute notebooks in place. Requires nbformat; execution delegates to jupyter nbconvert.
  • Session todo list — todo_write and todo_read tools give the agent a persistent task list scoped to the current session, backed by the existing SQLite store.
  • Effort levels — --effort low|medium|high|xhigh|max controls max_tokens and temperature per run. Changeable mid-session with /effort.
  • Anthropic extended thinking — activated at xhigh or max effort when extended_thinking = true in config. Thinking blocks are surfaced in the terminal and in stream-json output.
  • Bare mode — --bare flag skips DEVAGENT.md injection, CodePrism overlay, and the permission gate. Intended for scripted or CI use.
  • stream-json output — --output-format stream-json emits one JSON line per agent event to stdout, making it easy to pipe agent output into CI tooling or other processes.
  • /effort and /think REPL commands — change effort level and toggle extended thinking live without restarting the session.
  • /undo and /diff REPL commands — /undo rolls back the last file write; /diff shows a unified diff of all files changed this session.
  • Shell escape in REPL — prefix any line with ! to run a shell command inline without leaving the agent session.
  • Inline diff rendering — file edits and writes emit a unified diff; the terminal renders it with colour-coded +/- lines immediately after each tool call.
  • Sub-agent spawning — spawn_agent tool lets the agent delegate subtasks recursively to independent sub-agents.
  • Peer result sharing — read_peer_results lets orchestrated workers inspect outputs from completed sibling tasks before starting their own work.
  • Dependency context injection — workers that depend on other tasks automatically receive a summary of those tasks' outputs in their system prompt.
  • DEVAGENT.md — devagent init-project scaffolds a project instruction file and .devagent/ directory. The agent reads and injects it into every session's system prompt.
  • Project memory — named facts can be persisted across sessions in .devagent/memory.json and are injected as context on each run.
  • --allow-tools — shorthand flag to auto-allow specific tool names without writing full allow rules.

Changed

  • SearchXConfig gains a base_url field (default http://localhost:8888) so self-hosted SearXNG instances can be pointed at from TOML config.
  • build_registry now wires the active LLM provider and session ID through to vision and todo tools respectively.

Fixed

  • Session ID and provider were not passed from DevAgentSession into build_registry, so vision tool image encoding and session-scoped todo storage were silently skipped. Now correctly wired.

Version 0.5.0

Choose a tag to compare

@knight22-21 knight22-21 released this 02 Sep 17:39

Added

  • Permission gate with allow/deny rules for tool calls (--allow, --deny, --interactive-approval on devagent run)
  • PermissionManager with glob pattern matching (e.g. --allow write_file:src/**)
  • Agent loop pauses and prompts before executing unmatched tool calls
  • POST /api/v1/sessions/{id}/approve now resolves live pending tool calls via call_id
  • ApprovalNeededEvent emitted by the loop when approval is required
  • Session-keyed permission registry so the REST API can reach the active manager

Changed

  • devagent serve upgraded to FastAPI + uvicorn with full REST and WebSocket support
  • WS /ws/v1/{session_id} streams live events from DB at 500ms polling
  • GET /api/v1/sessions, /events, /memory, /orchestrate/graph endpoints added
  • Session context auto-compression and multi-agent orchestration (devagent orchestrate)

Fixed

  • Session store schema now initialised on app startup, fixing CI failures on fresh environments

Version 0.3.0

Choose a tag to compare

@knight22-21 knight22-21 released this 13 Aug 17:53

Added

  • New devagent watch command to monitor a GitHub repository for newly opened issues
  • Watcher storage, scheduler, and reporting flow for recurring repository health checks
  • Cross-issue conflict detection for files touched by multiple open issues
  • Watcher-specific terminal rendering for watched repos, health reports, and stored analyses

Changed

  • GitHub client now supports listing repository issues for watcher checks
  • Configuration now includes watcher defaults such as interval, labels, and cross-conflict behavior
  • Added apscheduler runtime support and enabled automatic asyncio handling for pytest

Fixed

  • Added watcher-focused tests covering conflict detection, analysis building, and watcher storage

Version 0.2.0

Choose a tag to compare

@knight22-21 knight22-21 released this 11 Aug 13:58

Added

  • Direct GitHub issue and pull request URL input for devagent analyze --url
  • Interactive terminal chat sessions for exploring a generated gap analysis
  • Compact chat-focused rendering and conversation history support

Changed

  • devagent analyze can now jump straight into chat with --chat
  • Pull request URLs can be analyzed as specs
  • Spec analysis server now uses FastMCP

Fixed

  • Legacy SPECSYNC_CONFIG_PATH config override still works
  • Test imports and assertions were aligned with the devagent package naming