Repository navigation
v1.1.0 — Ollama Cloud + 20/20 benchmark
What's new in v1.1.0
Ollama Cloud support
DevAgent now works with Ollama Cloud out of the box. Set OLLAMA_HOST=https://ollama.com and OLLAMA_API_KEY in a .env file at your project root — the agent picks them up automatically at startup via python-dotenv. No code changes, no new flags, same Ollama API.
Available cloud models include gpt-oss:20b, nemotron-3-nano:30b, gemma4:31b, and 15+ others. See your account at ollama.com for the full list.
20/20 benchmark pass rate
With gpt-oss:20b on Ollama Cloud, DevAgent now passes all 20 tasks in the built-in benchmark — up from 9/20 (45%) with llama3.2:3b running locally. See BENCHMARKS.md for the full breakdown.
Benchmark results always saved
devagent bench native --live now always writes a timestamped JSON to benchmarks/results/ after each run. A rolling partial file is updated after each task so a crash mid-run does not lose completed results. The --output-json flag has been removed (it now happens automatically).
New docs
- INTEGRATIONS.md — setup guides for Claude Code, ChatGPT, Antigravity, GitHub Copilot, Cursor, Windsurf, VS Code, JetBrains, and Zed (moved from README)
- BENCHMARKS.md — per-task results, failure analysis, and how to run the benchmark against any model