Skip to content

v0.3.0

Choose a tag to compare

@svalench svalench released this 04 Sep 08:06
· 3 commits to main since this release

Release v0.3.0

Date: 2026-09-04

Summary

Any OpenAI-compatible endpoint as a provider, cache invalidation and key versioning shipped ahead of v1.0, plus safer routing defaults and community infrastructure.

Highlights

  • openai_compatible provider — one parameterized base_url covers OpenRouter, vLLM, llama.cpp server, LiteLLM proxy and any self-hosted inference speaking the OpenAI Chat Completions protocol. api_key is optional for keyless local servers.
  • Cache invalidation — LLMRouter.invalidate_cache(model=...) (and invalidate(model=...) on every backend) drops cached entries per model or entirely, returning the number of removed entries. Available on memory, Redis and Qdrant backends.
  • Key versioning — new CacheConfig.key_version: bump it when deploying a new system prompt and stale entries become invisible immediately, aging out via TTL — no flush needed.
  • exact_match mode — new CacheConfig.exact_match=True returns cache hits only for byte-identical queries (semantic search disabled) for correctness-sensitive workloads.
  • Safer CHEAPEST_FIRST — models with unknown pricing are now treated as infinitely expensive instead of free, so self-hosted models no longer win routing by default.
  • OpenAI provider fixes — no more Authorization: Bearer None header when api_key is unset; base_url trailing slashes are normalized.
  • Community infrastructure — CONTRIBUTING.md, bug report / feature request issue templates, PR template, GitHub Discussions enabled.
  • Repo hygiene — committed llm_cache_router.egg-info/ build artifact and .cursor/ scratchpad removed; uv.lock excluded from sdist.
  • README — new «Cache Invalidation, Versioning & Exact Match» section with threshold false-positive guidance, OpenAI-compatible provider docs, roadmap reordered (OpenTelemetry before Django helpers).

Upgrade notes

  • CHEAPEST_FIRST behavior change: if you relied on an unpriced model (e.g. self-hosted via openai_compatible) being selected as cheapest, pin its pricing via PricingManager(pricing_override={...}) — unknown pricing now loses to any known-priced option.
  • provider_used label: OpenAI-family responses now use the configured provider name (previously hardcoded to "openai"). Same value for standard setups.
  • Test suite grew from 40 to 82 tests; no public API removals.

Install

pip install llm-cache-router==0.3.0