large-multimodal-models

Here are 21 public repositories matching this topic...

zjysteven / lmms-finetune

A unified codebase for finetuning (full, lora) large multimodal models, supporting llava-1.5, qwen-vl, llava-interleave, llava-next-video, etc.

finetuning multimodal vision-language foundation-models instruction-tuning large-language-model llava visual-instruction-tuning multimodal-large-language-models large-multimodal-models qwen-vl llava-next

Updated Jul 20, 2024
Python

thunlp / LEGENT

Star

Open Platform for Embodied Agents

physics-engine robot-simulator language-grounding embodied-ai large-multimodal-models

Updated Jul 20, 2024
Python

OpenAdaptAI / OpenAdapt

Sponsor

Star

AI-First Process Automation with Large ([Language (LLMs) / Action (LAMs) / Multimodal (LMMs)] / Visual Language (VLMs)) Models

python transformers openai agents process-mining ai-agents process-automation huggingface huggingface-transformers ultralytics gpt-4 large-language-models anthropic segment-anything ai-agents-framework large-multimodal-models gpt4-vision google-gemini large-action-model

Updated Jul 21, 2024
Python

Psycoy / MixEval

Star

The official evaluation suite and dynamic data release for MixEval.

benchmark evaluation benchmarking-suite evaluation-framework benchmarking-framework foundation-models large-language-models large-language-model llm-inference llm-evaluation large-multimodal-models llm-evaluation-framework benchmark-mixture mixeval

Updated Jul 17, 2024
Python

TinyLLaVA / TinyLLaVA_Factory

Star

A Framework of Small-scale Large Multimodal Models

nlp transformers llama vision-language llava large-multimodal-models tinyllama

Updated Jul 17, 2024
Python

shijian2001 / VQAPromptBench

Star

A Benchmark for VQA prompt sensitivity

benchmark evaluation large-multimodal-models

Updated Jul 17, 2024
Python

MileBench / MileBench

Star

This repo contains evaluation code for the paper "MileBench: Benchmarking MLLMs in Long Context"

benchmark machine-learning natural-language-processing deep-neural-networks computer-vision deep-learning evaluation multimodality visual-question-answering multimodal foundation-models large-language-models llm llms long-context-transformers multimodal-large-language-models large-multimodal-models long-context-modeling

Updated Jul 11, 2024
Python

ShareGPT4Omni / ShareGPT4Video

Star

An official implementation of ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

gpt sora text-to-video large-language-models chatgpt large-vision-language-models large-multimodal-models gpt-4v large-video-language-models

Updated Jul 8, 2024
Python

ParadoxZW / LLaVA-UHD-Better

Star

A bug-free and improved implementation of LLaVA-UHD, based on the code from the official repo

multimodal large-language-models llava large-multimodal-models

Updated Jul 6, 2024
Python

AIFEG / BenchLMM

Star

[ECCV 2024] BenchLMM: Benchmarking Cross-style Visual Capability of Large Multimodal Models

benchmark cv dataset large-language-models large-multimodal-models

Updated Jul 3, 2024
Python

ShareGPT4Omni / ShareGPT4V

Star

[ECCV 2024] ShareGPT4V: Improving Large Multi-modal Models with Better Captions

gpt language-model large-language-models chatgpt instruction-tuning vision-language-model large-vision-language-models gpt4v large-multimodal-models gpt-4v eccv2024

Updated Jul 1, 2024
Python

MMMU-Benchmark / MMMU

Star

This repo contains evaluation code for the paper "MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI"

machine-learning natural-language-processing deep-neural-networks computer-vision deep-learning evaluation question-answering stem multimodality multimodal-learning visual-question-answering multimodal multimodal-deep-learning foundation-models large-language-models llm llms large-multimodal-models

Updated Jul 1, 2024
Python

eric-ai-lab / ProbMed

Star

"Worse than Random? An Embarrassingly Simple Probing Evaluation of Large Multimodal Models in Medical VQA"

evaluation vision-and-language medical-vqa medical-diagnosis llms large-multimodal-models

Updated Jun 24, 2024
Python

bzluan / TextCoT

Star

The official repo for “TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding”.

chain-of-thought large-multimodal-models

Updated Jun 20, 2024
Python

rohit901 / VANE-Bench

Star

Contains code and documentation for our VANE-Bench paper.

benchmark-datasets multimodal-deep-learning video-anomaly-detection large-language-models multimodal-large-language-models large-multimodal-models

Updated Jun 18, 2024
Python

shikiw / OPERA

Star

[CVPR 2024 Highlight] OPERA: Alleviating Hallucination in Multi-Modal Large Language Models via Over-Trust Penalty and Retrospection-Allocation

chatbot llama multimodal gpt-4 chatgpt vision-language-model vision-language-learning large-multimodal-models

Updated Jun 16, 2024
Python

xiaoachen98 / Open-LLaVA-NeXT

Star

An open-source implementation of LLaVA-NeXT.

chatbot llama multimodal multi-modality gpt-4 visual-language-learning chatgpt vision-language-model llava large-multimodal-models llama3 gpt4o llava-next

Updated Jun 12, 2024
Python

VisualWebBench / VisualWebBench

Star

Evaluation framework for paper "VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?"

machine-learning natural-language-processing computer-vision deep-learning evaluation question-answering visual-question-answering multimodal multimodal-deep-learning foundation-models large-language-models llm llms mllm multimodal-large-language-models large-multimodal-models

Updated May 31, 2024
Python

MMStar-Benchmark / MMStar

Star

This repo contains evaluation code for the paper "Are We on the Right Way for Evaluating Large Vision-Language Models"

evaluation multimodality multimodal-learning visual-question-answering multimodal large-language-models llm llms large-vision-language-model large-vision-language-models large-multimodal-models lvlms lvlm

Updated Apr 17, 2024
Python

sshh12 / multi_token

Star

Embed arbitrary modalities (images, audio, documents, etc) into large language models.

multimodal multi-modality large-language-models llm vision-language-model llava large-context large-multimodal-models

Updated Mar 27, 2024
Python

Improve this page

Add a description, image, and links to the large-multimodal-models topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the large-multimodal-models topic, visit your repo's landing page and select "manage topics."

Learn more

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

large-multimodal-models

Here are 21 public repositories matching this topic...

zjysteven / lmms-finetune

thunlp / LEGENT

OpenAdaptAI / OpenAdapt

Psycoy / MixEval

TinyLLaVA / TinyLLaVA_Factory

shijian2001 / VQAPromptBench

MileBench / MileBench

ShareGPT4Omni / ShareGPT4Video

ParadoxZW / LLaVA-UHD-Better

AIFEG / BenchLMM

ShareGPT4Omni / ShareGPT4V

MMMU-Benchmark / MMMU

eric-ai-lab / ProbMed

bzluan / TextCoT

rohit901 / VANE-Bench

shikiw / OPERA

xiaoachen98 / Open-LLaVA-NeXT

VisualWebBench / VisualWebBench

MMStar-Benchmark / MMStar

sshh12 / multi_token

Improve this page

Add this topic to your repo