0.2.4
v0.2.4 — safer semantic cache in prod.
Multimodal: content as a list of blocks; media — SHA-256 in the key, without full base64 in the query.
Isolation by model: one prompt does not give a hit between gpt-4o and gpt-4o-mini.
Message type, multimodal in WarmupEntry.
After upgrading, if necessary, reset the old Redis/Qdrant entries (legacy entries do not have a model field).
Migration Note
After upgrading, legacy Redis/Qdrant entries written by ≤ v0.2.3 have no model field in their payloads . Because _entry_from_payload reads payload.get("model") (defaults to None) and _model_matches(None, requested_model) returns False when a model is requested, those old entries will never hit — they just accumulate as dead weight. Flush them manually:
Redis:
bash
Delete all entries under the namespace (default: llm_cache_router)
redis-cli --scan --pattern "llm_cache_router:entry:*" | xargs redis-cli DEL
redis-cli DEL llm_cache_router:entries
Qdrant:
bash
Drop and recreate the collection
curl -X DELETE http://localhost:6333/collections/llm_cache_router
Or call the built-in cache clear:
python
await router._cache.clear()