Skip to content

0.2.4

Choose a tag to compare

@svalench svalench released this 24 May 10:47
· 6 commits to main since this release

v0.2.4 — safer semantic cache in prod.

Multimodal: content as a list of blocks; media — SHA-256 in the key, without full base64 in the query.
Isolation by model: one prompt does not give a hit between gpt-4o and gpt-4o-mini.
Message type, multimodal in WarmupEntry.
After upgrading, if necessary, reset the old Redis/Qdrant entries (legacy entries do not have a model field).
Migration Note
After upgrading, legacy Redis/Qdrant entries written by ≤ v0.2.3 have no model field in their payloads . Because _entry_from_payload reads payload.get("model") (defaults to None) and _model_matches(None, requested_model) returns False when a model is requested, those old entries will never hit — they just accumulate as dead weight. Flush them manually:

Redis:

bash

Delete all entries under the namespace (default: llm_cache_router)

redis-cli --scan --pattern "llm_cache_router:entry:*" | xargs redis-cli DEL
redis-cli DEL llm_cache_router:entries
Qdrant:

bash

Drop and recreate the collection

curl -X DELETE http://localhost:6333/collections/llm_cache_router
Or call the built-in cache clear:

python
await router._cache.clear()