v0.1.71: Local Qwen Vision & Model Aliases, Quadlet Health Probe Timeouts, Default Parallel Pytest
Release v0.1.71: Local Qwen Vision & Model Aliases, Quadlet Health Probe Timeouts, Default Parallel Pytest
Highlights
-
Model Registry & Alias Harmonization (#668):
- Register
local-qwen-vlandlocallama-qwen-vlvision models across LiteLLM, router backends, and runtime model sync with 65,536 token context windows and vision metadata. - Fully harmonize and register all host presets:
local-qwen,local-qwen-hass,local-qwen-routing,local-nomic-embed,whisper-1,llm-routing-agy,llm-routing-agy-sse, andagy-sse. - Prune active aliases from
DEPRECATED_MODEL_NAMESinrouter/model_sync.pyso background sync cycles do not inadvertently purge active models. - Update
LANGFUSE_MANAGED_MODELSand context limits inrouter/main.py. - Fix Jinja2 formatting in
dashboard.htmlwhenbest_free_modelorcontext_lengthis undefined.
- Register
-
Quadlet Health Probe Resilience (#669, #670, #671, #672, #673):
- Add explicit
timeout=3and bumpHealthTimeout=10sto prevent indefinite socket hangs during heavy load on LiteLLM and router container probes. - Replace
node -ehealth check inlangfuse-workerwith a native HTTP wget probe.
- Add explicit
-
Developer Experience & Validation Speed (#668):
- Configure
pytest-xdistwith-n autoas default inpytest.ini, reducing test suite execution from ~35s down to ~11s while strictly enforcing 100% statement and branch test coverage.
- Configure