Releases: sheepdestroyer/LLM-Routing
Release list
v0.1.75: On-Demand Dev Lifecycle, HAProxy Standby Backup & Quadlet Hardening
Release v0.1.75: On-Demand Dev Lifecycle, HAProxy Standby Backup & Quadlet Hardening
Highlights
• On-Demand Dev Stack Lifecycle (#699): Transition dev-router-pod to disabled and down by default. Dev Quadlet systemd units strip default.target so the dev pod will never auto-start on host boot or user login, while retaining WantedBy=llm-routing-dev-pod.service so all containers start in lockstep when invoked manually for qualification or qualification testing.
• Graceful Teardown & Port Cleanup (#699): Added --stop and --down options to start-stack.sh for graceful pod teardown, quadlet re-rendering, and zombie port cleanup (DEV_ENV_FILE=.env.dev ./start-stack.sh --stop).
• HAProxy Standby Backup Routing: Configured server dev_* <dev-port> check backup on production backends (llm-routing-backend, litellm-backend, langfuse-backend) for automated zero-traffic failover if production is offline and dev is activated.
• Uptime Kuma Noise Elimination: Paused all 9 dev LLM-Routing monitors (active = 0) in Uptime Kuma so they display ⏸️ PAUSED and prevent false alert noise while dev is offline.
• Version Bump (#700): Bumped pyproject.toml version to 0.1.75.
• Governance & Documentation: Updated Rules E1 and E7 in AGENTS.md, updated SERVICES.md, doc/manuals/LLM-Routing.md, and server wiki knowledge base.
• Quality Gates & Coverage: Maintained 100.00% statement and branch coverage gate across 888 unit tests, with clean ruff and mypy validations.
v0.1.74: Dependabot Updates, Annotation Auth & Manifest Sync
Release v0.1.74: Dependabot Updates, Annotation Auth & Manifest Sync
Highlights
- LiteLLM Gateway Upgrade (#679): Bump
ghcr.io/berriai/litellmfromv1.100.0tov1.100.1. - Langfuse Web & Worker Upgrade (#678): Bump
langfuse/langfuseandlangfuse/langfuse-workerfrom4.30.0to4.33.0. - Manifest Synchronization (#698): Synchronize
docker-compose.ymlservice definitions in lockstep withpod.yamland bumppyproject.tomlto0.1.74. - Annotation Save Authentication & Error Resilience (#691, #696): Enforce client authentication (
_authenticate_client_request) on@app.post("/dashboard/save-annotations"). Improve visualizer token extraction with decoupledtry...catchblocks for sandboxed iframe environments, plusres.okstatus checking. - HTTP Client Singleton Identity Verification (#692, #695): Add unit test assertions in
tests/test_models_proxy.pyverifying pointer identity (assert client1 is client2), proper reset instantiation, and connection limit inheritance. - Code Health & Dead Import Pruning (#693, #694): Remove unused
import urllib.errorinscripts/benchmark_classifier.py. - Documentation (#697): Document audio endpoint proxy (
/v1/audio/*), visualizer annotation authentication, and update router API table. - Quality Gates & Coverage: Maintained 100.00% statement and branch coverage gate across all 887 unit tests.
v0.1.73: Security Hardening, Performance Optimizations & Multi-Review Remediations
Release v0.1.73: Security Hardening, Performance Optimizations & Multi-Review Remediations
Highlights
- Proxy Client Authentication & Path Security (#680, #687, #690): Enforce fail-closed Bearer client authentication on
/v1/memoryand/v1/audio(ROUTER_API_KEY,LITELLM_MASTER_KEY,GATEWAY_KEY,MEMORY_API_KEY). Eliminate double-encoded path traversal attacks (%252e%252e%252f) via iterative decoding, preserve exact root URLs, and redirect stdio MCP diagnostics inmemory_mcp.pytosys.stderr. - Parallel Startup Model Registration & Atomic Swap (#681, #688, #690): Concurrently register OpenRouter and Ollama model rosters at startup with
asyncio.Semaphore(10), defensive dictionary lookups, and atomic in-memory roster swap into_registered_free_models. - Visualizer UX Accessibility (#682, #686, #690): Migrate Clear Annotation action to semantic
<button type="button">, WCAG AA contrast ratio compliance, and focus-visible indicators. - Chat Streaming Normalization & Tests (#683, #685, #690): Parse standard message payloads and streaming SSE
deltachunks while preserving token boundary whitespaces and indentation. - High-Performance Annotations Cache (#684, #689, #690): Read raw async bytes with
orjsondeserialization, freshness invalidation trackingst_ino,st_mtime_ns, andst_size, with immediate cache eviction on decode failure to prevent persistent cache poisoning. - Quality Gates & Test Coverage: 100.00% statement and branch coverage gate maintained across 878 unit tests and 319 integration tests, clean ruff and mypy static analysis.
v0.1.72: Bump LiteLLM to v1.100.0 & Langfuse to v4.30.0
Release v0.1.72: Bump LiteLLM to v1.100.0 & Langfuse to v4.30.0
Highlights
- LiteLLM Gateway Upgrade (#676): Bump
ghcr.io/berriai/litellmfromv1.99.1tov1.100.0. - Langfuse Web & Worker Upgrade (#675): Bump
langfuse/langfuseandlangfuse/langfuse-workerfrom4.28.1to4.30.0. - Manifest Synchronization (#677): Synchronize
docker-compose.ymlservice definitions forlitellm-gateway,langfuse-web, andlangfuse-workerin lockstep withpod.yaml.
v0.1.71: Local Qwen Vision & Model Aliases, Quadlet Health Probe Timeouts, Default Parallel Pytest
Release v0.1.71: Local Qwen Vision & Model Aliases, Quadlet Health Probe Timeouts, Default Parallel Pytest
Highlights
-
Model Registry & Alias Harmonization (#668):
- Register
local-qwen-vlandlocallama-qwen-vlvision models across LiteLLM, router backends, and runtime model sync with 65,536 token context windows and vision metadata. - Fully harmonize and register all host presets:
local-qwen,local-qwen-hass,local-qwen-routing,local-nomic-embed,whisper-1,llm-routing-agy,llm-routing-agy-sse, andagy-sse. - Prune active aliases from
DEPRECATED_MODEL_NAMESinrouter/model_sync.pyso background sync cycles do not inadvertently purge active models. - Update
LANGFUSE_MANAGED_MODELSand context limits inrouter/main.py. - Fix Jinja2 formatting in
dashboard.htmlwhenbest_free_modelorcontext_lengthis undefined.
- Register
-
Quadlet Health Probe Resilience (#669, #670, #671, #672, #673):
- Add explicit
timeout=3and bumpHealthTimeout=10sto prevent indefinite socket hangs during heavy load on LiteLLM and router container probes. - Replace
node -ehealth check inlangfuse-workerwith a native HTTP wget probe.
- Add explicit
-
Developer Experience & Validation Speed (#668):
- Configure
pytest-xdistwith-n autoas default inpytest.ini, reducing test suite execution from ~35s down to ~11s while strictly enforcing 100% statement and branch test coverage.
- Configure
v0.1.70: Prevent Llama-Server Autoload Loop During Metrics Polling
Release v0.1.70: Prevent Llama-Server Autoload Loop During Metrics Polling
- Llama-Server Metrics Query Hardening: In
router/main.py:get_llamacpp_metrics(), avoid querying/slots?model=<model>when no model is currently loaded in memory. - Prevent Autoload & Timeout Thrashing: Previously, when all models were unloaded, the router fell back to querying
/slotsfor the first listed model, triggeringensure_model_readyin llama.cpp router mode. Because this cold load takes 6-8s while the metrics poll timed out after 3s, it resulted in an aborted load and an endless retry loop every 5s. - Test Coverage: Updated unit tests in
router/tests/test_get_llamacpp_metrics.pymaintaining 100% test coverage gate.
v0.1.69
Release v0.1.69: LiteLLM Client Auth Log Reclassification & Negative Key Caching
- LiteLLM Client Auth Reclassification: Added
ClientAuthLogFilterandMinLevelFilterinlitellm/entrypoint.pyto reclassify client 401/404 authentication and virtual key lookup failures fromERRORtoWARNING. - Traceback Suppression: Strips redundant multi-line Python tracebacks from client auth errors so they produce clean, single-line
[WARNING]events routed tostdout. Real server and upstream LLM provider errors retain full[ERROR]severity and tracebacks. - Virtual Key Negative Caching: Added bounded negative cache in
router/main.py(_INVALID_VIRTUAL_KEY_CACHE, 60s TTL, max 5000 entries with FIFO eviction) for 400/404/blocked keys to prevent repeated hammering of/key/info. - Master Key Rejection Telemetry: Differentiates 401/403 responses from LiteLLM
/key/info, explicitly logging master key rejections as[ERROR]without polluting the client negative key cache.
v0.1.68
Release v0.1.68: LiteLLM Single-Line Logging, PostgreSQL Checkpoint Filtering & Langfuse 4.28.1
- LiteLLM Error Logging & Traces: Added custom
SingleLineFormatterandsys.excepthookinlitellm/entrypoint.pythat strips ANSI color codes, collapses multiline stack traces into a single line delimited by|, stamps bracketed severity tags ([INFO],[WARNING],[ERROR],[CRITICAL]), and preserves LiteLLM correlation context metadata. - Langfuse Media Upload Suppression: Automatically monkey-patches
MediaManager.process_media_in_eventto no-op whenLANGFUSE_MEDIA_UPLOAD_ENABLED=false(or unset), preventing spurious 500 error logs when media blob storage is disabled. - Static Model Tier Deployments: Configured safety-net fallback local endpoints for
agent-reasoning-core,agent-complex-core,agent-medium-core, andagent-simple-coreinlitellm/config.yaml. - PostgreSQL Checkpoint Log Suppression: Added
-c log_checkpoints=off -c log_min_messages=warningto container arguments/exec and applied runtimeALTER SYSTEMpersistence instart-stack.shto eliminate routine 5-minute checkpoint logs from systemd logs. - Dependencies & Manifest Parity: Upgraded Langfuse and Langfuse Worker to
4.28.1inpod.yamland synchronizeddocker-compose.yml.
v0.1.67
Release v0.1.67: WUD Tag Filters and Transforms for ClickHouse & MinIO
- ClickHouse: Added
wud.tag.include=^[0-9]+[.][0-9]+[.][0-9]+[.][0-9]+-distroless$andwud.tag.transform=^([0-9]+)[.]([0-9]+)[.]([0-9]+)[.]([0-9]+)-distroless$ => $1.$2.$3-$4to prevent WUD from falsely flagging-alpinetags as updates over-distroless. - MinIO: Replaced escaped regex classes with
[0-9]and[.]to avoid Quadlet generatorunsupported escape charlabel drops, and addedwud.tag.transform=^RELEASE[.]([0-9]{4})-([0-9]{2})-([0-9]{2})T([0-9]{2})-([0-9]{2})-([0-9]{2})Z$ => $1.$2$3.$4$5$6to ensure CalVer release dates are sorted in chronological order instead of coercing to equal2025.0.0SemVer. - LiteLLM & Router: Refined
wud.tag.include=^v[0-9]+[.][0-9]+[.][0-9]+$to avoid generator escape drops and prevent un-prefixed tags from being reported as updates. - Tests: Updated Quadlet template assertions in
tests/test_quadlet_templates.py.
v0.1.66
Release v0.1.66: 100% Pytest Statement & Branch Coverage
- Achieved 100.00% statement and branch coverage across all modules in router/ and scripts/host_agy_daemon.py (4,013 statements, 1,428 branches, 0 missed).
- Added 6 comprehensive test suites bringing the unit test suite to 710 passing tests.
- Enforced strict
--cov-fail-under=100gate in pyproject.toml and GitHub Actions CI. - All static checks (ruff, mypy) passing with zero errors.