Skip to content

vllm-stack-0.1.13

Latest

Choose a tag to compare

@github-actions github-actions released this 29 Sep 23:48
e8cb495

The stack deployment of vLLM

What's Changed

  • feat(helm): support externally-managed PVC via pvcExistingClaimName (#938) by @Anai-Guo in #988
  • [Bugfix][Helm] skip lmcache dashboard when cache server is disabled by @bitnik in #987
  • [BugFux] Fix help text formatting for Sentry sample rates in argument parser t… by @xiaoajie738 in #946
  • [CI/Build] Add Gatekeeper aggregate status check for auto-merge by @ruizhang0101 in #1015
  • [Router][Bugfix] Handle all router types correctly in proxy_multipart_request by @shernshiou in #1021
  • [Bugfix] Align vllm version in requirements-test.txt with pyproject.toml by @arijitroy003 in #1009
  • feat(operator): support shmSize on VLLMRuntime for tensor parallelism (#899) by @Anai-Guo in #994
  • [CI/Build] Add CODEOWNERS for automated PR review assignment by @arijitroy003 in #1011
  • Feat(helm): Modify the existing serviceAccount for use by router and lora-controller by @ghkdwlgns612 in #1005
  • [CI/Build] Move test scripts and assets from .github/ to tests/ (part of #531) by @ighutake-debug in #1024
  • [Doc] Fix references to non-existent tutorial values files by @latent-9 in #1026
  • Bump Kubernetes client to 36.0.3 by @arnaudgeiser in #1014
  • [CI/Build] Add checkout to Helm functionality jobs by @lfsun02 in #1027
  • [Router][Bugfix] Return 400 for malformed request payloads by @lfsun02 in #1017
  • [CI/Build] Clean up timed-out K8s E2E runs and fail fast on backend exits by @lfsun02 in #1031
  • [Feat] Migrate inference kgateway to agentgateway by @danehans in #1033
  • [CI/Build] Increase Secure-Minimal-Example validation timeout to 10 minutes by @ruizhang0101 in #1036
  • [Router][Feat] loadaware routing: KV-cache-aware placement weighted by live load by @bazakeliad in #1035
  • [Bugfix][Router] Remove ghost endpoint on evicted-worker MODIFIED event by @jiawen7777 in #1012
  • [Feat] Add priority routing by @punitvara in #996
  • [CI/Build] Scope Shaoting-Feng's code ownership to router, observability, and benchmarks by @Shaoting-Feng in #1046
  • [CI/Build] Request one GPU per replica in the two-pods Helm functionality test by @Shaoting-Feng in #1047
  • [Doc][Bugfix] Pin KubeRay pipeline-parallel tutorial to a vLLM image that ships the ray CLI by @rishabhsinha17 in #1056
  • [Bugfix][Router] Fix audio multipart routing and response handling by @lfsun02 in #1038
  • [Router][Bugfix] Select longest LMCache prefix match by @dsxyy in #1055
  • [Router][Bugfix] Cache KV-aware tokenizers per model by @dsxyy in #1077
  • [Bugfix][Router] Default transcription language to None (auto-detect) by @WaelRabah in #1070
  • [Bugfix][Router] Soften fastapi pin to allow patched starlette by @dundysm in #1088
  • [Bugfix][Router] Handle container.command=None in the service-name sleep-mode check by @chrikrah in #1096
  • [Bugfix][Router] Fix leaked in-flight counters in RequestStatsMonitor by @lfsun02 in #1072
  • feat(helm): optional runtime AI inventory via k8s-aibom (default off) by @glenmessenger in #1099

New Contributors

Full Changelog: vllm-stack-0.1.12...vllm-stack-0.1.13