Skip to content

[Feature]: Managed Prompt Caching For Anthropic Models #5885

Description

@yigitkonur

The Feature

The new prompt caching feature in Anthropic models offers a great opportunity for cost savings to many people. However, managing this individually on the client side is very challenging. It would be awesome if the LiteLLM proxy offered a managed cache feature, where you could add a single parameter to the YAML configuration to automatically cache all previous messages for repetitive requests to the model. That would be super helpful.

Motivation, pitch

This technique can reduce costs by 3-4 times for users of conversational AI products, but unfortunately, no LLM proxy has this feature. I wish you would add it.

Twitter / LinkedIn details

@yigitkonur

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions