Skip to content

v0.1.2

Choose a tag to compare

@loopyd loopyd released this 02 May 02:03
· 72 commits to main since this release

Highlights

  • Default ./run.sh build targets TheTom/llama-cpp-turboquant@feature/turboquant-kv-cache.
  • ./run.sh warmup [model...] can load models early through llama-swap's upstream health route.
  • Hugging Face -hf assets are documented as lazy downloads, with warmup and preload guidance.
  • Rerank support is documented as a first-class route with qmd-rerank and /v1/rerank examples.
  • config.yml.example is synced with the active QMD aliases: qmd-generate, qmd-embed, and qmd-rerank.
  • README wording was cleaned up to keep the docs concise and approachable.

Included Commits

  • 5a6368e Simplify README wording
  • fc3f841 Document QMD rerank config and API
  • 5ecc23e Add warmup flow and turboquant defaults