v0.20.0: Prompt Caching
·
272 commits
to main
since this release
Prompt Caching
Add opt-in provider-side prompt caching for Anthropic, Bedrock (Claude), and OpenAI via a unified CacheHint API. Non-breaking — all new fields are optional and zero-value behaviour is identical to before.
New API
// Top-level automatic caching (easy path)
events, err := provider.CreateStream(ctx, llm.StreamOptions{
Model: "anthropic/claude-sonnet-4-6",
Messages: messages,
CacheHint: &llm.CacheHint{Enabled: true},
})
// Per-message explicit breakpoints (advanced)
&llm.SystemMsg{
Content: largeSystemPrompt,
CacheHint: &llm.CacheHint{Enabled: true},
}
// Inspect cache usage
event.Usage.CachedTokens // tokens read from cache
event.Usage.CacheWriteTokens // tokens written to cacheProvider support
| Provider | Mode |
|---|---|
| Anthropic (direct + OAuth) | Explicit breakpoints via cache_control |
| Bedrock (Claude models) | Explicit breakpoints via cachePoint blocks |
| OpenAI | Always automatic; CacheHint{TTL: "1h"} enables extended retention |
Other changes
- Cost calculation:
Usage.Costnow correctly prices cache reads (0.1× input) and writes (1.25× input) separately for Anthropic and Bedrock - llmcli: cache token counts appear automatically in
-vverbose output when non-zero - Tests: 26 new unit tests + 7 cache integration tests; live tests migrated out of provider packages — provider packages are now unit-tests only;
TestProviderstable now covers both OpenAI Chat Completions and Responses API paths separately
Breaking changes
None.