Releases: haseeb-heaven/open-agent
Release list
OpenAgent v4.1.4 — Bundle version metadata fix
OpenAgent v4.1.4
This patch release fixes stale CLI version metadata in the published bundle.
Fixed
- Rebuilt core generated version metadata before bundling.
- Rebuilt the npm bundle from a clean output directory.
- Verified that the bundled executable and telemetry metadata report
4.1.4.
Installation
npm install -g @haseeb_heaven/open-agent@4.1.4OpenAgent v4.1.3 — Fast provider failover
OpenAgent v4.1.3
This release improves response latency for explicit model selections and makes
free-provider failover bounded and observable.
Performance
- Explicit Gemini model selections bypass the auto-model classifier on fresh
turns, while model changes during an active sequence retain re-routing. OPENAGENT_CLI_FAST_MODE=1caps compatible-provider completions at 1024
tokens, skips the optional next-speaker follow-up, bounds provider requests to
3 seconds, and limits fast failover to one alternate candidate.- Model registry and derived free catalog setup are cached across turns.
- Live model matrices support bounded concurrency and report wall-clock, p50,
p95, target misses, and per-model outcomes.
Reliability
- Fast-mode deadline failures enter the existing free-provider routing failure
path instead of waiting through the normal retry budget. - Normal mode retains existing retry, fallback, and model-quality behavior.
Validation
- Provider/client regression suite: 175 tests passed.
- Full core run: 8,108 tests passed; one existing workspace-policy test timed out
in the local environment. - A2A server suite: 140 tests passed.
- Real filesystem benchmarks were run across Google, Cerebras, Groq, OpenRouter,
and Hugging Face providers.
Installation
npm install -g @haseeb_heaven/open-agent@4.1.3v4.1.1 — Free-model picker, extensions market & post-release fixes
Ships the develop-branch feature set previously (mis-)tagged as an unreleased v4.1.0, renumbered to avoid colliding with main's already-released v4.1.0 (provider crash / test-isolation fixes — see that release).
Features (from the develop v4.1.0 feature set)
- feat(cli):
/free-modelsinteractive picker for free-tier models, plus/freeshortcut - feat(cli):
/extensions market— browse and install extensions from the registry
Bugfixes (this release)
- fix(tui):
/free-modelskey-entry dialog now rejects structurally-invalid input (leftover slash-command keystrokes, strings under 16 chars) before writing to.env, instead of silently persisting garbage as an API key - fix(providers):
resolveRegistryPathnow walks upward fromcwdforconfigs/models.tomlbefore falling back to a bundle-relative search, so running from a repo subdirectory no longer risks pinning a stalebundle/configs/models.tomlcopy for the process lifetime - fix(test): resolve a hanging/flaky kitty-protocol test suite in
gemini.test.tsx - fix(test): clear a stray force-exit timer on test teardown that was surfacing as unhandled
process.exitexceptions in unrelated test runs - fix(scripts): live-test harness (
run-live-tests.mjs) retries once on timeout before grading a scenario FAIL, and flags scenarios whose prompt references a missing media-dir subfolder
Full details in CHANGELOG.md.
v4.1.0 — Provider crash & test-isolation fixes
Rolls up recent fixes: provider crash on zero-API-key first-run auth, and elimination of a homedir-mock test-isolation bypass. Full packages/cli suite re-verified green (468 files, 6952 tests) before release.
- fix(tui): rename the interactive model picker command to
/models(with/modelretained as a compatibility alias); unavailable paid-provider models now offer their provider-specific.envkey setup directly in the dialog - fix(providers): preserve unique registry-key routing and endpoint overrides when aliases share a LiteLLM model id; free sessions now rotate through the fallback catalog after rate-limit or free-router failures
- fix(providers): stop crashing on first-run auth when zero API keys are configured
- fix(tests): eliminate a homedir-mock bypass in the extension manager and related suites that let tests leak real developer-machine state (
~/.gemini/~/.openagent) instead of the mocked temp home - test(providers): add fallback-chain coverage and validate OpenRouter's
openai/gpt-oss-20b:freewith both complete and streaming live requests - docs: update the branch-specific clone command; document
/models,--resume, and--yolo
Note: this release is from the main branch. A separate, unrelated v4.1.0 was shipped on develop (free-model picker + extensions marketplace) — see v4.1.1 for that lineage, renumbered to avoid a version collision.
Interpreter v3.6.0 — Stability Fixes & Live Testing Expansion
Highlights
Stability fixes
- Key-exhaustion no longer crashes the classic REPL; structured
AllKeysExhaustedErrorwithprovider/retry_after_etaattributes - Model router attempts one free-model fallback before surfacing key exhaustion
- Persistent banner/status line never wraps or truncates on narrow terminals
- Fixed Windows
rich.ConsoleOSError: Bad file descriptorcrash (forcedlegacy_windows=False) - Root-caused and fixed a test-suite stdout-fd corruption bug (
tests/interactive/helpers.py's mock interpreter) litellm.completion()calls now bounded by a 90s timeout to stop indefinite hangs- Fixed web-search DDGS import order, a chart prompt-injection contradiction, and live-flake soft-skip scoping
Live testing expansion
- 10 new create/analyze/summarize/convert/edit scenario cases covering zip, mp3, java, sqlite, docx, svg, webm (stdlib-only, no external codec/toolchain dependency in CI)
- Model-router multi-key rotation regression test
- Filled remaining all-modes e2e smoke gaps (yolo, gemini-style, and other previously-uncovered modes)
- Raised combined
libscoverage to 81% (≥80% CI gate)
New provider
- Added Cerebras (cloud.cerebras.ai) as a fully wired free/rate-limited model provider
TUI improvements
- No-args wizard answers now persist to
~/.code-interpreter/config.json;--configforces a re-run - "Configure advanced options?" → no now genuinely skips every subsequent advanced prompt
- Wizard cancellation (Ctrl+C/Esc) exits cleanly instead of an unhandled traceback
See CHANGELOG.md for the full list.
Interpreter v3.4.0 — MCP, Sessions, Data/Science, Sandbox & CI
Interpreter 3.4.0
Major feature wave on top of v3.3.0 agentic/free-LLM foundation.
Highlights
- MCP + autonomy: native FS/shell ToolRegistry,
--yolo,--mcp-server(#215) - Streaming + vision:
--stream/--no-stream,--image//image(#216) - Web search:
--search//search(DuckDuckGo / Tavily / Serper) (#217) - Codegen modes:
--mode generate/--mode projectwithout execution (#212) - Structured output:
--output-format json|markdown|plain(#219) - Persistent sessions:
--session//session(#218) - Identity & onboarding: free/local positioning + first-run UX (#220)
- Local-first:
--attach//file,--ollama,--local(#221) - Data analysis: EDA, charts, SQL, exports (#222)
- Science: notebooks, themes, ML helpers, PDF reports (#223)
- CI/coverage: matrix CI, coverage gate, Codecov (#224)
- Sandbox hardening: subprocess/Docker backends,
--timeout,--safety, audit, secret scan (#225) - Interactive tests: slash/REPL/session/live-exec coverage (#226)
Quick start
python interpreter.py --list-free
python interpreter.py --local --attach data.csv "summarize this"
python interpreter.py --cli --yes --output-format json -m local-model -f task.txt
python interpreter.py --sandbox docker --timeout 60 "analyze sales.csv"Full history: CHANGELOG.md
Interpreter v3.3.0 — Agentic Free LLMs & Resilience
Interpreter 3.3.0
Foundation release for agentic free-LLM UX and production resilience.
Highlights
- Gemini-CLI-style
--gemini-styleReAct REPL with free/cheap catalog (--free,--list-free,/free) - Multi-key rotation, rate limiter, circuit breaker
- Metrics CLI (
/key-status,/reload-keys,/metrics) - Non-interactive
--yes/INTERPRETER_YESfor CI - Multi-agent
--agentand ReAct--agenticloops
See v3.4.0 for MCP, streaming, sessions, data/science, sandbox levels, and CI coverage that landed after this tag.
Full history: CHANGELOG.md
v3.2.3
Interpreter 3.2.3 Latest
@haseeb-heaven haseeb-heaven released this Jul 11, 2026
3.2.3
🔥 Release highlights:
- Resolved Command Injection vulnerability in file opening by replacing
subprocess.callwithos.startfileon Windows. - Mitigated Path Traversal vulnerability in
UtilityManager.get_full_file_pathwith strict boundary checks. - Enhanced stability with timeouts on all external HTTP requests to prevent application hanging.
- Boosted performance by pre-compiling all regex patterns in
ExecutionSafetyManager. - Improved Terminal UI accessibility, ensuring prompt choices fallbacks are explicitly visible in non-TTY environments.
- Fixed Ollama "NoneType" Error, allowing robust extraction of direct string responses and dictionary outputs for local models like Mistral.
- Fixed Ollama API Key Error, intentionally bypassing
HUGGINGFACE_API_KEYrequirements for local and Ollama models. - Updated legacy model configurations to point to the modern 2026 stable aliases (e.g.
gpt-4.1,claude-sonnet-4-6). - Expanded unit test coverage to 263 tests, directly validating security fixes, UX enhancements, and API Key robustness logic.
📜 Changelog:
- v3.2.3 - Fixed Windows command injection, resolved path traversal in utility manager, implemented HTTP timeouts, optimized SafetyManager by pre-compiling regexes, improved non-TTY fallback prompts, fixed Ollama/local model API Key extraction and output parsing, updated legacy model configurations, and expanded test suite with full coverage.
- v3.2.2 - Added sandbox mode (default ON) with /sandbox and /unsafe toggles, replaced --unsafe with --sandbox / --no-sandbox, improved subprocess security delegation, increased SAFE timeout to 300s, fixed watchdog timer issues, strengthened safe-mode protection, added process-group kill on timeout, improved Python detection using AST parsing, fixed multiple security vulnerabilities (P0/P1/P2).
📦 Assets:
- interpreter.zip
- Source code (zip)
- Source code (tar.gz)
v3.2.2 Sandbox Security and Code Interpreter Architecture
Full Changelog: v3.2.2...v3.3.0
VERSION v3.3.0 (Summary of Key Changes)
This update mainly improves security, stability, and how code runs in a controlled environment.
MAIN FEATURE: SANDBOX MODE
- A new sandbox system is now the default way code runs
- You can control it using:
- /sandbox → safe execution
- /unsafe → less restricted execution
- The old --unsafe flag is replaced with --sandbox / --no-sandbox
- Sandbox now has stronger protection against harmful commands and file access
SECURITY IMPROVEMENTS
- Better blocking of dangerous system-level commands
- Fixed ways users could bypass safety (file writes, absolute paths, symlinks)
- Added stricter checks for file operations like .write()
- Improved detection of unsafe patterns across Python and other scripts
- Safer subprocess handling and execution delegation
EXECUTION & PERFORMANCE
- Increased timeout limit in safe mode to 300 seconds for long-running code
- Improved handling of stuck or long processes using process group kill
- Better detection of Python code using AST parsing
- Fixed watchdog timer issues in sandbox
STABILITY & BUG FIXES
- Fixed syntax errors and test failures
- Cleaned up execution logic formatting
- Fixed temporary file execution issues
- Resolved multiple high-priority (P0/P1/P2) security bugs
BUILD & TOOLING
- Improved build_release.sh script with:
- better error handling
- cleaner structure
- more reliable release process
Overall, this release focuses on making the interpreter safer, more reliable, and better at handling real-world code execution, especially with the new sandbox-first approach.
v3.2.1
What's Changed
- Enhance model catalog, execution safety, and TUI features by @haseeb-heaven in #24
- Update model configs to newer LLM releases and fix routing issues by @haseeb-heaven in #23
Full Changelog: v3.1.0...v3.2.1