You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
AlpinDale edited this page Sep 3, 2024
·
5 revisions
Aphrodite Engine
Caution
This wiki is deprecated. Please refer to the official documentation website instead!
Aphrodite Engine is designed for serving LLMs at scale, based on vLLM. It supports the majority of HuggingFace models, including Llama, Mistral, and Mixtral.
Aphrodite also supports multiple weight quantization methods for not-at-scale (and at-scale!) use-cases. Please see this page for details.