Skip to content

Releases: sheepdestroyer/LLM-Routing

v0.1.75: On-Demand Dev Lifecycle, HAProxy Standby Backup & Quadlet Hardening

Choose a tag to compare

@sheepdestroyer sheepdestroyer released this 12 Sep 16:32
166701c

Release v0.1.75: On-Demand Dev Lifecycle, HAProxy Standby Backup & Quadlet Hardening

Highlights

• On-Demand Dev Stack Lifecycle (#699): Transition dev-router-pod to disabled and down by default. Dev Quadlet systemd units strip default.target so the dev pod will never auto-start on host boot or user login, while retaining WantedBy=llm-routing-dev-pod.service so all containers start in lockstep when invoked manually for qualification or qualification testing.
• Graceful Teardown & Port Cleanup (#699): Added --stop and --down options to start-stack.sh for graceful pod teardown, quadlet re-rendering, and zombie port cleanup (DEV_ENV_FILE=.env.dev ./start-stack.sh --stop).
• HAProxy Standby Backup Routing: Configured server dev_* <dev-port> check backup on production backends (llm-routing-backend, litellm-backend, langfuse-backend) for automated zero-traffic failover if production is offline and dev is activated.
• Uptime Kuma Noise Elimination: Paused all 9 dev LLM-Routing monitors (active = 0) in Uptime Kuma so they display ⏸️ PAUSED and prevent false alert noise while dev is offline.
• Version Bump (#700): Bumped pyproject.toml version to 0.1.75.
• Governance & Documentation: Updated Rules E1 and E7 in AGENTS.md, updated SERVICES.md, doc/manuals/LLM-Routing.md, and server wiki knowledge base.
• Quality Gates & Coverage: Maintained 100.00% statement and branch coverage gate across 888 unit tests, with clean ruff and mypy validations.

v0.1.74: Dependabot Updates, Annotation Auth & Manifest Sync

Choose a tag to compare

@sheepdestroyer sheepdestroyer released this 12 Sep 10:29
c57f9a2

Release v0.1.74: Dependabot Updates, Annotation Auth & Manifest Sync

Highlights

  • LiteLLM Gateway Upgrade (#679): Bump ghcr.io/berriai/litellm from v1.100.0 to v1.100.1.
  • Langfuse Web & Worker Upgrade (#678): Bump langfuse/langfuse and langfuse/langfuse-worker from 4.30.0 to 4.33.0.
  • Manifest Synchronization (#698): Synchronize docker-compose.yml service definitions in lockstep with pod.yaml and bump pyproject.toml to 0.1.74.
  • Annotation Save Authentication & Error Resilience (#691, #696): Enforce client authentication (_authenticate_client_request) on @app.post("/dashboard/save-annotations"). Improve visualizer token extraction with decoupled try...catch blocks for sandboxed iframe environments, plus res.ok status checking.
  • HTTP Client Singleton Identity Verification (#692, #695): Add unit test assertions in tests/test_models_proxy.py verifying pointer identity (assert client1 is client2), proper reset instantiation, and connection limit inheritance.
  • Code Health & Dead Import Pruning (#693, #694): Remove unused import urllib.error in scripts/benchmark_classifier.py.
  • Documentation (#697): Document audio endpoint proxy (/v1/audio/*), visualizer annotation authentication, and update router API table.
  • Quality Gates & Coverage: Maintained 100.00% statement and branch coverage gate across all 887 unit tests.

v0.1.73: Security Hardening, Performance Optimizations & Multi-Review Remediations

Choose a tag to compare

@sheepdestroyer sheepdestroyer released this 10 Sep 19:28
61d3d91

Release v0.1.73: Security Hardening, Performance Optimizations & Multi-Review Remediations

Highlights

  • Proxy Client Authentication & Path Security (#680, #687, #690): Enforce fail-closed Bearer client authentication on /v1/memory and /v1/audio (ROUTER_API_KEY, LITELLM_MASTER_KEY, GATEWAY_KEY, MEMORY_API_KEY). Eliminate double-encoded path traversal attacks (%252e%252e%252f) via iterative decoding, preserve exact root URLs, and redirect stdio MCP diagnostics in memory_mcp.py to sys.stderr.
  • Parallel Startup Model Registration & Atomic Swap (#681, #688, #690): Concurrently register OpenRouter and Ollama model rosters at startup with asyncio.Semaphore(10), defensive dictionary lookups, and atomic in-memory roster swap into _registered_free_models.
  • Visualizer UX Accessibility (#682, #686, #690): Migrate Clear Annotation action to semantic <button type="button">, WCAG AA contrast ratio compliance, and focus-visible indicators.
  • Chat Streaming Normalization & Tests (#683, #685, #690): Parse standard message payloads and streaming SSE delta chunks while preserving token boundary whitespaces and indentation.
  • High-Performance Annotations Cache (#684, #689, #690): Read raw async bytes with orjson deserialization, freshness invalidation tracking st_ino, st_mtime_ns, and st_size, with immediate cache eviction on decode failure to prevent persistent cache poisoning.
  • Quality Gates & Test Coverage: 100.00% statement and branch coverage gate maintained across 878 unit tests and 319 integration tests, clean ruff and mypy static analysis.

v0.1.72: Bump LiteLLM to v1.100.0 & Langfuse to v4.30.0

Choose a tag to compare

@sheepdestroyer sheepdestroyer released this 07 Sep 15:29
d21f8d8

Release v0.1.72: Bump LiteLLM to v1.100.0 & Langfuse to v4.30.0

Highlights

  • LiteLLM Gateway Upgrade (#676): Bump ghcr.io/berriai/litellm from v1.99.1 to v1.100.0.
  • Langfuse Web & Worker Upgrade (#675): Bump langfuse/langfuse and langfuse/langfuse-worker from 4.28.1 to 4.30.0.
  • Manifest Synchronization (#677): Synchronize docker-compose.yml service definitions for litellm-gateway, langfuse-web, and langfuse-worker in lockstep with pod.yaml.

v0.1.71: Local Qwen Vision & Model Aliases, Quadlet Health Probe Timeouts, Default Parallel Pytest

Choose a tag to compare

@sheepdestroyer sheepdestroyer released this 05 Sep 09:40
6863458

Release v0.1.71: Local Qwen Vision & Model Aliases, Quadlet Health Probe Timeouts, Default Parallel Pytest

Highlights

  • Model Registry & Alias Harmonization (#668):

    • Register local-qwen-vl and locallama-qwen-vl vision models across LiteLLM, router backends, and runtime model sync with 65,536 token context windows and vision metadata.
    • Fully harmonize and register all host presets: local-qwen, local-qwen-hass, local-qwen-routing, local-nomic-embed, whisper-1, llm-routing-agy, llm-routing-agy-sse, and agy-sse.
    • Prune active aliases from DEPRECATED_MODEL_NAMES in router/model_sync.py so background sync cycles do not inadvertently purge active models.
    • Update LANGFUSE_MANAGED_MODELS and context limits in router/main.py.
    • Fix Jinja2 formatting in dashboard.html when best_free_model or context_length is undefined.
  • Quadlet Health Probe Resilience (#669, #670, #671, #672, #673):

    • Add explicit timeout=3 and bump HealthTimeout=10s to prevent indefinite socket hangs during heavy load on LiteLLM and router container probes.
    • Replace node -e health check in langfuse-worker with a native HTTP wget probe.
  • Developer Experience & Validation Speed (#668):

    • Configure pytest-xdist with -n auto as default in pytest.ini, reducing test suite execution from ~35s down to ~11s while strictly enforcing 100% statement and branch test coverage.

v0.1.70: Prevent Llama-Server Autoload Loop During Metrics Polling

Choose a tag to compare

@sheepdestroyer sheepdestroyer released this 05 Sep 01:06
f5e8e70

Release v0.1.70: Prevent Llama-Server Autoload Loop During Metrics Polling

  • Llama-Server Metrics Query Hardening: In router/main.py:get_llamacpp_metrics(), avoid querying /slots?model=<model> when no model is currently loaded in memory.
  • Prevent Autoload & Timeout Thrashing: Previously, when all models were unloaded, the router fell back to querying /slots for the first listed model, triggering ensure_model_ready in llama.cpp router mode. Because this cold load takes 6-8s while the metrics poll timed out after 3s, it resulted in an aborted load and an endless retry loop every 5s.
  • Test Coverage: Updated unit tests in router/tests/test_get_llamacpp_metrics.py maintaining 100% test coverage gate.

v0.1.69

Choose a tag to compare

@sheepdestroyer sheepdestroyer released this 05 Sep 00:21
707bdd9

Release v0.1.69: LiteLLM Client Auth Log Reclassification & Negative Key Caching

  • LiteLLM Client Auth Reclassification: Added ClientAuthLogFilter and MinLevelFilter in litellm/entrypoint.py to reclassify client 401/404 authentication and virtual key lookup failures from ERROR to WARNING.
  • Traceback Suppression: Strips redundant multi-line Python tracebacks from client auth errors so they produce clean, single-line [WARNING] events routed to stdout. Real server and upstream LLM provider errors retain full [ERROR] severity and tracebacks.
  • Virtual Key Negative Caching: Added bounded negative cache in router/main.py (_INVALID_VIRTUAL_KEY_CACHE, 60s TTL, max 5000 entries with FIFO eviction) for 400/404/blocked keys to prevent repeated hammering of /key/info.
  • Master Key Rejection Telemetry: Differentiates 401/403 responses from LiteLLM /key/info, explicitly logging master key rejections as [ERROR] without polluting the client negative key cache.

v0.1.68

Choose a tag to compare

@sheepdestroyer sheepdestroyer released this 04 Sep 23:43
ec538ca

Release v0.1.68: LiteLLM Single-Line Logging, PostgreSQL Checkpoint Filtering & Langfuse 4.28.1

  • LiteLLM Error Logging & Traces: Added custom SingleLineFormatter and sys.excepthook in litellm/entrypoint.py that strips ANSI color codes, collapses multiline stack traces into a single line delimited by |, stamps bracketed severity tags ([INFO], [WARNING], [ERROR], [CRITICAL]), and preserves LiteLLM correlation context metadata.
  • Langfuse Media Upload Suppression: Automatically monkey-patches MediaManager.process_media_in_event to no-op when LANGFUSE_MEDIA_UPLOAD_ENABLED=false (or unset), preventing spurious 500 error logs when media blob storage is disabled.
  • Static Model Tier Deployments: Configured safety-net fallback local endpoints for agent-reasoning-core, agent-complex-core, agent-medium-core, and agent-simple-core in litellm/config.yaml.
  • PostgreSQL Checkpoint Log Suppression: Added -c log_checkpoints=off -c log_min_messages=warning to container arguments/exec and applied runtime ALTER SYSTEM persistence in start-stack.sh to eliminate routine 5-minute checkpoint logs from systemd logs.
  • Dependencies & Manifest Parity: Upgraded Langfuse and Langfuse Worker to 4.28.1 in pod.yaml and synchronized docker-compose.yml.

v0.1.67

Choose a tag to compare

@sheepdestroyer sheepdestroyer released this 03 Sep 23:37
a225c6b

Release v0.1.67: WUD Tag Filters and Transforms for ClickHouse & MinIO

  • ClickHouse: Added wud.tag.include=^[0-9]+[.][0-9]+[.][0-9]+[.][0-9]+-distroless$ and wud.tag.transform=^([0-9]+)[.]([0-9]+)[.]([0-9]+)[.]([0-9]+)-distroless$ => $1.$2.$3-$4 to prevent WUD from falsely flagging -alpine tags as updates over -distroless.
  • MinIO: Replaced escaped regex classes with [0-9] and [.] to avoid Quadlet generator unsupported escape char label drops, and added wud.tag.transform=^RELEASE[.]([0-9]{4})-([0-9]{2})-([0-9]{2})T([0-9]{2})-([0-9]{2})-([0-9]{2})Z$ => $1.$2$3.$4$5$6 to ensure CalVer release dates are sorted in chronological order instead of coercing to equal 2025.0.0 SemVer.
  • LiteLLM & Router: Refined wud.tag.include=^v[0-9]+[.][0-9]+[.][0-9]+$ to avoid generator escape drops and prevent un-prefixed tags from being reported as updates.
  • Tests: Updated Quadlet template assertions in tests/test_quadlet_templates.py.

v0.1.66

Choose a tag to compare

@sheepdestroyer sheepdestroyer released this 03 Sep 23:00
5e40145

Release v0.1.66: 100% Pytest Statement & Branch Coverage

  • Achieved 100.00% statement and branch coverage across all modules in router/ and scripts/host_agy_daemon.py (4,013 statements, 1,428 branches, 0 missed).
  • Added 6 comprehensive test suites bringing the unit test suite to 710 passing tests.
  • Enforced strict --cov-fail-under=100 gate in pyproject.toml and GitHub Actions CI.
  • All static checks (ruff, mypy) passing with zero errors.