Skip to content

Releases: siddharth23P/otto_agent

Otto 0.1.4

Choose a tag to compare

@siddharth23P siddharth23P released this 17 Sep 09:30
8542bb6

What's new

  • Vision fallback: image questions walk a chain (gemini-3.8-flash → 3.7-flash → 3.6-flash → Claude Haiku 4.5 → GPT-5 mini) instead of failing when one provider is out of credit. Router.chain() and describe_with_fallback() are new.
  • Documents: make_document writes PDF, Word, Excel and PowerPoint from Markdown (or from a file) into documents/. openpyxl, python-docx and reportlab are now installed everywhere; python-pptx is new.
  • Phone as a last resort: the phone decision prefers otto's own tools, judges attached files by name only, and when no model can decide, uses the phone only for an explicit device request.
  • Phone action dictionary: a bridge that offers actions() / run_action() gets phone_find (search the phone's premapped actions) and phone_action (run one by name). The prompt stays the same size however many actions the phone has. Actions the person finishes hand the phone over; message text and recipients stay out of the log.
  • invalid is a phone error code.

Otto 0.1.3

Choose a tag to compare

@siddharth23P siddharth23P released this 17 Sep 06:32
32f691a

Input-token (prompt) caching with all four providers (#15).

Added

  • Anthropic: automatic prompt caching (a top-level cache_control, 5-minute TTL). An agent loop re-reads its transcript at a tenth of the input price.
  • OpenAI: a prompt_cache_key per model, so repeated prefixes hit one cache. Sent only to OpenAI itself.
  • Gemini: implicit caching, as before; cached tokens are counted.
  • Inception: Mercury's cache hits (prompt_tokens_details.cached_tokens) are now read and priced.
  • OTTO_PROMPT_CACHE=0 turns caching off. A route can override or drop its own setting through model_kwargs.

Fixed

  • An Anthropic cache write reported by its TTL (ephemeral_5m_input_tokens) was priced as plain input; it is now priced as a write.
  • The usage snapshot now carries cache_write_tokens.

Otto 0.1.2

Choose a tag to compare

@siddharth23P siddharth23P released this 16 Sep 12:17

Otto as a library, and the phone as a place to work (#12, #13, #14).

Added

  • Embedding API (agent/embed.py, API_VERSION = 1): configure(home, environ=...) with a keystore mode that never writes a key file, Runtime, sessions whose run streams JSON events, and answer / cancel from another thread. OTTO_HOME / OTTO_ENV_FILE move otto's state.
  • Phone tools (agent/phone/): a device-agnostic backend protocol, a bounded screen digest that never shows a password field, and tools of which only phone_install and phone_commit change anything. Used by the Otto Android app.
  • A money guard that judges the whole page (rules v3): secure, payment, checkout and cart pages are scored from the page's structure, not from single words, with a page corpus shared with the app. Explicit pay controls are never tapped, and PIN/OTP/CVV/password fields are never typed into.
  • otto serve: the same runtime behind a token-protected WebSocket.

Changed

  • fastembed and langgraph-cli are optional extras; a missing vendor SDK disables that vendor, not the router.
  • numpy>=1.26.2, so otto installs on Android CPython 3.13.
  • A connection error names its root cause, and a failed provider check logs its whole exception chain.
  • A provider that answers with a credits or quota 429 is cooled down for an hour.

Otto 0.1.1

Choose a tag to compare

@siddharth23P siddharth23P released this 13 Sep 17:25

Fixes found while recording the demo and running CI, on top of 0.1.0.

Fixed

  • Langfuse exporter flooding the REPL when its host is down (#9). With LANGFUSE_BASE_URL pointing at a collector that is not running, the SDK retried every span batch with a connection error on stderr. The host is now probed once per process with a one-second connect; unreachable means one warning and tracing off for the process. No keys means silence with no warning.
  • A traceback on quitting the TUI mid-turn (#10). ValueError: <Token ...> was created in a different Context from the stream generator being finalised on another thread. Every per-run binder restores by value when its token cannot be reset, and the streaming entry points finish quietly on that one error.
  • Setup screen crashing on a slow machine. on_mount queried widgets inside TabbedContent panes before they were mounted. The fill now waits a refresh for them.
  • Three setup-screen tests asserted on UI state after a single pause; they wait on the condition now.

Docs and packaging

  • Every folder has its own README; the root README is a map plus the measured properties, with the one-minute demo inline and seven screenshots taken mid-run.
  • GitHub Pages site at https://siddharth23p.github.io/otto_agent/ , contributing guide, code of conduct, security policy, issue and PR templates.
  • Published to PyPI as otto-cli-agent through trusted publishing.

Install

pip install --upgrade otto-cli-agent

1,554 tests.

Otto 0.1.0

Choose a tag to compare

@siddharth23P siddharth23P released this 13 Sep 14:59

First release of Otto: a terminal AI agent that works on a codebase, a container, a browser or a desktop, uses what it built, and judges its own work against criteria it wrote before it started.

What is in it

  • One agent loop with four modes over four vendors (Inception, OpenAI, Anthropic, Gemini) and any OpenAI-compatible endpoint, with routing that learns from outcomes, per-model cooldowns and a provider circuit breaker.
  • A rubric-first evaluator: criteria written from the task before any attempt, shared by the loop and the judge.
  • Once-only holds before an irreversible action, before finishing on unrun code, and before finishing without using what was built.
  • 18 tools: shell and Python, files, code_map, a browser, exercise (walk through a page, a served app, a CLI, an API, a terminal program or a device app and report each step as the machine saw it), a desktop, web search, workspace RAG and memory recall.
  • Tiered memory with type-aware compaction and two-stage semantic recall (96% on LoCoMo at ~2,600 tokens a query), a lesson bank, and sessions that survive the process.
  • A Textual TUI with a setup screen, a directory browser, live progress, a token and dollar ledger, themes; a REPL with the same pipeline.
  • Six benchmark harnesses (golden set, SWE-bench Verified, Claw-Eval, LoCoMo, compaction, Humanity's Last Exam) with the rules that keep a number honest.
  • 1,534 tests that need no key and no network, on macOS, Linux and Windows.

Install

pip install otto-cli-agent
otto doctor
otto tui

Also on PyPI, or from source with uv sync. INCEPTION_API_KEY is required; the other vendors are optional. See the README for setup, commands and environment variables.