Skip to content

v0.1.71: Local Qwen Vision & Model Aliases, Quadlet Health Probe Timeouts, Default Parallel Pytest

Choose a tag to compare

@sheepdestroyer sheepdestroyer released this 05 Sep 09:40
· 25 commits to master since this release
6863458

Release v0.1.71: Local Qwen Vision & Model Aliases, Quadlet Health Probe Timeouts, Default Parallel Pytest

Highlights

  • Model Registry & Alias Harmonization (#668):

    • Register local-qwen-vl and locallama-qwen-vl vision models across LiteLLM, router backends, and runtime model sync with 65,536 token context windows and vision metadata.
    • Fully harmonize and register all host presets: local-qwen, local-qwen-hass, local-qwen-routing, local-nomic-embed, whisper-1, llm-routing-agy, llm-routing-agy-sse, and agy-sse.
    • Prune active aliases from DEPRECATED_MODEL_NAMES in router/model_sync.py so background sync cycles do not inadvertently purge active models.
    • Update LANGFUSE_MANAGED_MODELS and context limits in router/main.py.
    • Fix Jinja2 formatting in dashboard.html when best_free_model or context_length is undefined.
  • Quadlet Health Probe Resilience (#669, #670, #671, #672, #673):

    • Add explicit timeout=3 and bump HealthTimeout=10s to prevent indefinite socket hangs during heavy load on LiteLLM and router container probes.
    • Replace node -e health check in langfuse-worker with a native HTTP wget probe.
  • Developer Experience & Validation Speed (#668):

    • Configure pytest-xdist with -n auto as default in pytest.ini, reducing test suite execution from ~35s down to ~11s while strictly enforcing 100% statement and branch test coverage.