You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
This commit was created on GitHub.com and signed with GitHub’s verified signature.
Release v0.1.70: Prevent Llama-Server Autoload Loop During Metrics Polling
Llama-Server Metrics Query Hardening: In router/main.py:get_llamacpp_metrics(), avoid querying /slots?model=<model> when no model is currently loaded in memory.
Prevent Autoload & Timeout Thrashing: Previously, when all models were unloaded, the router fell back to querying /slots for the first listed model, triggering ensure_model_ready in llama.cpp router mode. Because this cold load takes 6-8s while the metrics poll timed out after 3s, it resulted in an aborted load and an endless retry loop every 5s.
Test Coverage: Updated unit tests in router/tests/test_get_llamacpp_metrics.py maintaining 100% test coverage gate.