Skip to content

v0.3.13 — unified repo env vars, concurrency limits, CUDA unified memory

Choose a tag to compare

@loopyd loopyd released this 03 Jul 23:52
· 49 commits to main since this release

Refactoring release: unified repo source env vars, llama-swap concurrency limits, and CUDA unified memory for larger models.

Changed

  • Unified repo source env vars: Removed mode-specific turboquant_llama_cpp_repo, turboquant_llama_cpp_ref, spiritbuun_llama_cpp_repo, spiritbuun_llama_cpp_ref, mtp_llama_cpp_repo, and mtp_llama_cpp_ref from pyproject.toml, Settings dataclass, env var resolution, and Dockerfile build args. All llama.cpp variants now share one LLAMA_CPP_REPO/LLAMA_CPP_REF pair set via LLAMACPP_LLAMA_CPP_REPO/LLAMACPP_LLAMA_CPP_REF env vars or pyproject.toml defaults.
  • llama-swap concurrency limits: Added concurrencyLimit: 4 to every model in all five runtime configs (basic, turboquant, mtp, spiritbuun, lucebox). llama-swap now enforces at most 4 parallel requests per upstream model via its semaphore-based throttling.
  • CUDA unified memory: Set ENV GGML_CUDA_ENABLE_UNIFIED_MEMORY=1 in the Dockerfile runtime stage, allowing oversubscribed VRAM for larger models on RTX 5090 GPUs.

Fixed

  • Fixed ruff I001 (unsorted imports) in config.py, helpers.py, and runtime.py.
  • Fixed ruff F401 (unused imports) in config.py (ProgressReporter, shutil_which).
  • Fixed ruff F811 (duplicate dataclass import) in config.py.
  • Fixed pyright str | object → int() type errors in config.py load_settings.
  • Fixed pyright Popen[bytes] / Popen[str] mismatch in runtime.py _stop_proc.
  • Fixed pyright signal handler type error in runtime.py.

Validation

  • ruff check easyllama/ passes with zero errors.
  • pyright easyllama/ passes for all user code; 11 pre-existing docker SDK type-stub warnings remain.