Skip to content

Otto 0.1.3

Choose a tag to compare

@siddharth23P siddharth23P released this 17 Sep 06:32
32f691a

Input-token (prompt) caching with all four providers (#15).

Added

  • Anthropic: automatic prompt caching (a top-level cache_control, 5-minute TTL). An agent loop re-reads its transcript at a tenth of the input price.
  • OpenAI: a prompt_cache_key per model, so repeated prefixes hit one cache. Sent only to OpenAI itself.
  • Gemini: implicit caching, as before; cached tokens are counted.
  • Inception: Mercury's cache hits (prompt_tokens_details.cached_tokens) are now read and priced.
  • OTTO_PROMPT_CACHE=0 turns caching off. A route can override or drop its own setting through model_kwargs.

Fixed

  • An Anthropic cache write reported by its TTL (ephemeral_5m_input_tokens) was priced as plain input; it is now priced as a write.
  • The usage snapshot now carries cache_write_tokens.