The Feature
The new prompt caching feature in Anthropic models offers a great opportunity for cost savings to many people. However, managing this individually on the client side is very challenging. It would be awesome if the LiteLLM proxy offered a managed cache feature, where you could add a single parameter to the YAML configuration to automatically cache all previous messages for repetitive requests to the model. That would be super helpful.
Motivation, pitch
This technique can reduce costs by 3-4 times for users of conversational AI products, but unfortunately, no LLM proxy has this feature. I wish you would add it.
Twitter / LinkedIn details
@yigitkonur
The Feature
The new prompt caching feature in Anthropic models offers a great opportunity for cost savings to many people. However, managing this individually on the client side is very challenging. It would be awesome if the LiteLLM proxy offered a managed cache feature, where you could add a single parameter to the YAML configuration to automatically cache all previous messages for repetitive requests to the model. That would be super helpful.
Motivation, pitch
This technique can reduce costs by 3-4 times for users of conversational AI products, but unfortunately, no LLM proxy has this feature. I wish you would add it.
Twitter / LinkedIn details
@yigitkonur