The stack deployment of vLLM
What's Changed
- feat(helm): support externally-managed PVC via pvcExistingClaimName (#938) by @Anai-Guo in #988
- [Bugfix][Helm] skip lmcache dashboard when cache server is disabled by @bitnik in #987
- [BugFux] Fix help text formatting for Sentry sample rates in argument parser t… by @xiaoajie738 in #946
- [CI/Build] Add Gatekeeper aggregate status check for auto-merge by @ruizhang0101 in #1015
- [Router][Bugfix] Handle all router types correctly in proxy_multipart_request by @shernshiou in #1021
- [Bugfix] Align vllm version in requirements-test.txt with pyproject.toml by @arijitroy003 in #1009
- feat(operator): support shmSize on VLLMRuntime for tensor parallelism (#899) by @Anai-Guo in #994
- [CI/Build] Add CODEOWNERS for automated PR review assignment by @arijitroy003 in #1011
- Feat(helm): Modify the existing serviceAccount for use by router and lora-controller by @ghkdwlgns612 in #1005
- [CI/Build] Move test scripts and assets from .github/ to tests/ (part of #531) by @ighutake-debug in #1024
- [Doc] Fix references to non-existent tutorial values files by @latent-9 in #1026
- Bump Kubernetes client to 36.0.3 by @arnaudgeiser in #1014
- [CI/Build] Add checkout to Helm functionality jobs by @lfsun02 in #1027
- [Router][Bugfix] Return 400 for malformed request payloads by @lfsun02 in #1017
- [CI/Build] Clean up timed-out K8s E2E runs and fail fast on backend exits by @lfsun02 in #1031
- [Feat] Migrate inference kgateway to agentgateway by @danehans in #1033
- [CI/Build] Increase Secure-Minimal-Example validation timeout to 10 minutes by @ruizhang0101 in #1036
- [Router][Feat] loadaware routing: KV-cache-aware placement weighted by live load by @bazakeliad in #1035
- [Bugfix][Router] Remove ghost endpoint on evicted-worker MODIFIED event by @jiawen7777 in #1012
- [Feat] Add priority routing by @punitvara in #996
- [CI/Build] Scope Shaoting-Feng's code ownership to router, observability, and benchmarks by @Shaoting-Feng in #1046
- [CI/Build] Request one GPU per replica in the two-pods Helm functionality test by @Shaoting-Feng in #1047
- [Doc][Bugfix] Pin KubeRay pipeline-parallel tutorial to a vLLM image that ships the ray CLI by @rishabhsinha17 in #1056
- [Bugfix][Router] Fix audio multipart routing and response handling by @lfsun02 in #1038
- [Router][Bugfix] Select longest LMCache prefix match by @dsxyy in #1055
- [Router][Bugfix] Cache KV-aware tokenizers per model by @dsxyy in #1077
- [Bugfix][Router] Default transcription language to None (auto-detect) by @WaelRabah in #1070
- [Bugfix][Router] Soften fastapi pin to allow patched starlette by @dundysm in #1088
- [Bugfix][Router] Handle container.command=None in the service-name sleep-mode check by @chrikrah in #1096
- [Bugfix][Router] Fix leaked in-flight counters in RequestStatsMonitor by @lfsun02 in #1072
- feat(helm): optional runtime AI inventory via k8s-aibom (default off) by @glenmessenger in #1099
New Contributors
- @xiaoajie738 made their first contribution in #946
- @arijitroy003 made their first contribution in #1009
- @ighutake-debug made their first contribution in #1024
- @latent-9 made their first contribution in #1026
- @arnaudgeiser made their first contribution in #1014
- @lfsun02 made their first contribution in #1027
- @danehans made their first contribution in #1033
- @bazakeliad made their first contribution in #1035
- @jiawen7777 made their first contribution in #1012
- @punitvara made their first contribution in #996
- @rishabhsinha17 made their first contribution in #1056
- @dsxyy made their first contribution in #1055
- @dundysm made their first contribution in #1088
- @chrikrah made their first contribution in #1096
- @glenmessenger made their first contribution in #1099
Full Changelog: vllm-stack-0.1.12...vllm-stack-0.1.13