Popular repositories Loading
-
vllm-v100-2080ti-recipes
vllm-v100-2080ti-recipes PublicWorking recipes and measured benchmarks for serving large LLMs on pre-Ampere NVIDIA GPUs — Tesla V100 (sm_70) and RTX 2080 Ti (sm_75). vLLM forks, llama.cpp tensor parallelism, AWQ/MoE gotchas.
Something went wrong, please refresh the page to try again.
If the problem persists, check the GitHub status page or contact support.
If the problem persists, check the GitHub status page or contact support.