llava
Here are 191 public repositories matching this topic...
[NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.
-
Updated
Aug 12, 2024 - Python
SGLang is a fast serving framework for large language models and vision language models.
-
Updated
Jun 7, 2025 - Python
SUPIR aims at developing Practical Algorithms for Photo-Realistic Image Restoration In the Wild. Our new online demo is also released at suppixel.ai.
-
Updated
May 12, 2025 - Python
An efficient, flexible and full-featured toolkit for fine-tuning LLM (InternLM2, Llama3, Phi3, Qwen, Mistral, ...)
-
Updated
May 29, 2025 - Python
中文nlp解决方案(大模型、数据、模型、训练、推理)
-
Updated
Jun 2, 2025 - Jupyter Notebook
Build multimodal language agents for fast prototype and production
-
Updated
Mar 19, 2025 - Python
Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
-
Updated
Jun 6, 2025 - Python
[ACL 2024 🔥] Video-ChatGPT is a video conversation model capable of generating meaningful conversation about videos. It combines the capabilities of LLMs with a pretrained visual encoder adapted for spatiotemporal video representation. We also introduce a rigorous 'Quantitative Evaluation Benchmarking' for video-based conversational models.
-
Updated
Mar 29, 2025 - Python
MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.
-
Updated
Jun 4, 2025 - Python
Pocket-Sized Multimodal AI for content understanding and generation across multilingual texts, images, and 🔜 video, up to 5x faster than OpenAI CLIP and LLaVA 🖼️ & 🖋️
-
Updated
Jan 3, 2025 - Python
Tag manager and captioner for image datasets
-
Updated
May 21, 2025 - Python
Famous Vision Language Models and Their Architectures
-
Updated
Feb 24, 2025 - Markdown
🔥🔥 LLaVA++: Extending LLaVA with Phi-3 and LLaMA-3 (LLaVA LLaMA-3, LLaVA Phi-3)
-
Updated
Jul 10, 2024 - Python
A Framework of Small-scale Large Multimodal Models
-
Updated
Apr 26, 2025 - Python
Paddle Multimodal Integration and eXploration, supporting mainstream multi-modal tasks, including end-to-end large-scale multi-modal pretrain models and diffusion model toolbox. Equipped with high performance and flexibility.
-
Updated
Jun 5, 2025 - Python
Improve this page
Add a description, image, and links to the llava topic page so that developers can more easily learn about it.
Add this topic to your repo
To associate your repository with the llava topic, visit your repo's landing page and select "manage topics."