Skip to content

v0.20.0: Prompt Caching

Choose a tag to compare

@timofriedlberlin timofriedlberlin released this 21 Mar 00:45
· 272 commits to main since this release

Prompt Caching

Add opt-in provider-side prompt caching for Anthropic, Bedrock (Claude), and OpenAI via a unified CacheHint API. Non-breaking — all new fields are optional and zero-value behaviour is identical to before.

New API

// Top-level automatic caching (easy path)
events, err := provider.CreateStream(ctx, llm.StreamOptions{
    Model:     "anthropic/claude-sonnet-4-6",
    Messages:  messages,
    CacheHint: &llm.CacheHint{Enabled: true},
})

// Per-message explicit breakpoints (advanced)
&llm.SystemMsg{
    Content:   largeSystemPrompt,
    CacheHint: &llm.CacheHint{Enabled: true},
}

// Inspect cache usage
event.Usage.CachedTokens     // tokens read from cache
event.Usage.CacheWriteTokens // tokens written to cache

Provider support

Provider Mode
Anthropic (direct + OAuth) Explicit breakpoints via cache_control
Bedrock (Claude models) Explicit breakpoints via cachePoint blocks
OpenAI Always automatic; CacheHint{TTL: "1h"} enables extended retention

Other changes

  • Cost calculation: Usage.Cost now correctly prices cache reads (0.1× input) and writes (1.25× input) separately for Anthropic and Bedrock
  • llmcli: cache token counts appear automatically in -v verbose output when non-zero
  • Tests: 26 new unit tests + 7 cache integration tests; live tests migrated out of provider packages — provider packages are now unit-tests only; TestProviders table now covers both OpenAI Chat Completions and Responses API paths separately

Breaking changes

None.