Releases: aws-neuron/vllm-omni-neuron
Releases · aws-neuron/vllm-omni-neuron
Release list
v0.24.0.0.1.0 (Beta)
vLLM Omni Neuron v0.24.0.0.1.0 (Beta)
Highlights
This first release introduces the vLLM Omni Neuron plugin, an out-of-tree Neuron backend for vLLM Omni that runs diffusion and multimodal generation models on AWS Trainium. It reuses vLLM Omni's upstream serving components like request APIs, scheduling, and output processing, and adds Neuron-specific execution, caching, and sharding on top. See the vLLM Omni Neuron documentation.
- Supported models: Wan2.2-T2V-A14B (text-to-video) and Wan2.2-I2V-A14B (image-to-video) on Trn2, Trn3
- Offline generation via the Omni.generate API and online serving via the OpenAI-compatible /v1/videos endpoint
- Neuron-optimized NKI kernels for the diffusion hot paths (fused QKV projection, const-max and ring attention, fused adaptive LayerNorm with FP8 quantization, and MLP), included as reference implementations you can adapt for your own model.
- Parallelism support includes tensor (TP), context (CP), Megatron sequence (SP), classifier-free guidance (CFG), and VAE patch parallelism
- Quantization support includes BF16 and FP8 (fp8_row_mx) DiT quantization
- Cache-DiT diffusion caching support for faster Wan2.2-T2V generation
- VAE spatial tiling and temporal chunking for high-resolution generation
- torch.compile with compilation caching artifacts support
- Accuracy validated with VBench and VBench-I2V
Other Details
- The vLLM Omni Neuron plugin uses the versioning scheme . (e.g., 0.24.0.0.1.0 = vLLM Omni 0.24.0, plugin 0.1.0)
- Support for vLLM Omni 0.24.0 and Neuron SDK 2.32.0
- Instances supported: Trn2 and Trn3
Known Issues & Limitations
- Initial support is validated on specific configurations (480P and 720P). See the model recipes for the exact supported configurations per model and resolution.
- This initial release optimizes the single-request (batch size 1) flow for performance.
- FP8 is supported for Wan2.2-T2V on Trn3 only; image-to-video runs in BF16.
- CFG parallelism (cfg_parallel_size=2) is not validated for I2V at 720P; see the I2V model recipe for the recommended 720P configuration.