Skip to content

v0.1.70: Prevent Llama-Server Autoload Loop During Metrics Polling

Choose a tag to compare

@sheepdestroyer sheepdestroyer released this 05 Sep 01:06
· 28 commits to master since this release
f5e8e70

Release v0.1.70: Prevent Llama-Server Autoload Loop During Metrics Polling

  • Llama-Server Metrics Query Hardening: In router/main.py:get_llamacpp_metrics(), avoid querying /slots?model=<model> when no model is currently loaded in memory.
  • Prevent Autoload & Timeout Thrashing: Previously, when all models were unloaded, the router fell back to querying /slots for the first listed model, triggering ensure_model_ready in llama.cpp router mode. Because this cold load takes 6-8s while the metrics poll timed out after 3s, it resulted in an aborted load and an endless retry loop every 5s.
  • Test Coverage: Updated unit tests in router/tests/test_get_llamacpp_metrics.py maintaining 100% test coverage gate.