Repository navigation
Otto 0.1.3
Input-token (prompt) caching with all four providers (#15).
Added
- Anthropic: automatic prompt caching (a top-level
cache_control, 5-minute TTL). An agent loop re-reads its transcript at a tenth of the input price. - OpenAI: a
prompt_cache_keyper model, so repeated prefixes hit one cache. Sent only to OpenAI itself. - Gemini: implicit caching, as before; cached tokens are counted.
- Inception: Mercury's cache hits (
prompt_tokens_details.cached_tokens) are now read and priced. OTTO_PROMPT_CACHE=0turns caching off. A route can override or drop its own setting throughmodel_kwargs.
Fixed
- An Anthropic cache write reported by its TTL (
ephemeral_5m_input_tokens) was priced as plain input; it is now priced as a write. - The usage snapshot now carries
cache_write_tokens.