You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Refactoring release: unified repo source env vars, llama-swap concurrency limits, and CUDA unified memory for larger models.
Changed
Unified repo source env vars: Removed mode-specific turboquant_llama_cpp_repo, turboquant_llama_cpp_ref, spiritbuun_llama_cpp_repo, spiritbuun_llama_cpp_ref, mtp_llama_cpp_repo, and mtp_llama_cpp_ref from pyproject.toml, Settings dataclass, env var resolution, and Dockerfile build args. All llama.cpp variants now share one LLAMA_CPP_REPO/LLAMA_CPP_REF pair set via LLAMACPP_LLAMA_CPP_REPO/LLAMACPP_LLAMA_CPP_REF env vars or pyproject.toml defaults.
llama-swap concurrency limits: Added concurrencyLimit: 4 to every model in all five runtime configs (basic, turboquant, mtp, spiritbuun, lucebox). llama-swap now enforces at most 4 parallel requests per upstream model via its semaphore-based throttling.
CUDA unified memory: Set ENV GGML_CUDA_ENABLE_UNIFIED_MEMORY=1 in the Dockerfile runtime stage, allowing oversubscribed VRAM for larger models on RTX 5090 GPUs.
Fixed
Fixed ruff I001 (unsorted imports) in config.py, helpers.py, and runtime.py.
Fixed ruff F401 (unused imports) in config.py (ProgressReporter, shutil_which).
Fixed ruff F811 (duplicate dataclass import) in config.py.
Fixed pyright str | object → int() type errors in config.pyload_settings.
Fixed pyright Popen[bytes] / Popen[str] mismatch in runtime.py_stop_proc.
Fixed pyright signal handler type error in runtime.py.
Validation
ruff check easyllama/ passes with zero errors.
pyright easyllama/ passes for all user code; 11 pre-existing docker SDK type-stub warnings remain.