model: add Janus Pro for image understanding #16906

ravenouse · 2025-10-31T23:24:02Z

This pull request introduces support for the Janus‑Pro 1B and Janus‑Pro 7B models within the llama.cpp framework.

The focus of this update is on image understanding (i.e., visual-input → textual or conceptual output).
Image generation is not covered by this PR.

Usage & Current Progress

Convert models to GGUF files:

# Convert the base Janus-Pro 1B model
python convert_hf_to_gguf.py deepseek-community/Janus-Pro-1B \
    --remote \
    --outfile janus-pro-1b-f16.gguf \
    --outtype f16

# Convert the mmproj component
python convert_hf_to_gguf.py deepseek-community/Janus-Pro-1B \
    --remote \
    --outfile mmproj-janus-pro-1b-f16.gguf \
    --outtype f16 \
    --mmproj

The converted GGUF files can be accessed here: https://huggingface.co/Ericwang/Janus-Pro-1B-GGUF

Run the model:

# Build the project:
cmake -B build
cmake --build build --target llama-mtmd-cli

./build/bin/llama-mtmd-cli \
    -m janus-pro-1b-f16.gguf \
    --mmproj mmproj-janus-pro-1b-f16.gguf \
    --chat-template deepseek

References

Janus-Pro 1B model card:
https://huggingface.co/deepseek-community/Janus-Pro-1B

Janus-Pro 7B model card:
https://huggingface.co/deepseek-community/Janus-Pro-7B

Configurations:
https://huggingface.co/deepseek-community/Janus-Pro-1B/blob/main/config.json
https://huggingface.co/deepseek-community/Janus-Pro-7B/blob/main/config.json

HF Implementation:
https://github.com/huggingface/transformers/tree/main/src/transformers/models/janus

ravenouse · 2025-10-31T23:37:40Z

Tested with this image: https://www.pixelstalk.net/wp-content/uploads/2016/04/Golden-retriever-dogs-high-definition-wallpapers.jpg

gguf-py/gguf/tensor_mapping.py

convert_hf_to_gguf.py

tools/mtmd/clip.cpp

Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com>

Co-authored-by: Xuan-Son Nguyen <son@huggingface.co>

ravenouse · 2025-11-02T07:34:56Z

Hi @CISC and @ngxson ,

Thank you for the thorough review and valuable feedback.
I've addressed all the comments. I also re-ran the conversion and inference workflows, and both are working as expected.

Ready for another look when you have a moment. Thanks a lot!

tools/mtmd/clip.cpp

Co-authored-by: Xuan-Son Nguyen <thichthat@gmail.com>

Co-authored-by: Xuan-Son Nguyen <son@huggingface.co>

ravenouse · 2025-11-02T18:31:41Z

Thanks again for the review!

Just updated the code and tested it again.

I've updated the code and retested it with the following image:
https://1.bp.blogspot.com/-tLB0HRLcOp4/Tj4Pvhsq6vI/AAAAAAAAAG8/h6ahy6g4GJI/s1600/Llama_lying_down.jpg

PS: The forced push was made to correct a formatting typo in the commit message.

ngxson

Looking good! Merging once the CI passes

CISC · 2025-11-02T19:22:22Z

Fix all the whitespace errors though. :)
https://github.com/ggml-org/llama.cpp/actions/runs/19016845485/job/54305935425?pr=16906

tools/mtmd/clip.cpp

Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com>

ravenouse · 2025-11-03T16:02:53Z

Hi @CISC and @ngxson ,

Thank you so much for the quick review and support to finalize this PR. Truly appreciated!

* origin/master: (169 commits) opencl: support imrope (ggml-org#16914) fix: Viewing multiple PDF attachments (ggml-org#16974) model-conversion : pass config to from_pretrained (ggml-org#16963) server : add props.model_alias (ggml-org#16943) ggml: CUDA: add head size 72 for flash-attn (ggml-org#16962) mtmd: add --image-min/max-tokens (ggml-org#16921) mtmd: pad mask for qwen2.5vl (ggml-org#16954) ggml : LoongArch fixes (ggml-org#16958) sync: minja (glm 4.6 & minmax m2 templates) (ggml-org#16949) SYCL: optimized repeat_back kernel (3× fewer asm instructions, 2× faster)Feature/sycl repeat back opt (ggml-org#16869) feat(webui): improve LaTeX rendering with currency detection (ggml-org#16508) test-backend-ops : fix segfault in moe-expert-reduce test in support mode and coverage (ggml-org#16936) ci : disable failing riscv cross build (ggml-org#16952) model: add Janus Pro for image understanding (ggml-org#16906) clip : use FA (ggml-org#16837) server : support unified cache across slots (ggml-org#16736) common : move gpt-oss reasoning processing to init params (ggml-org#16937) docs: remove llama_sampler_accept reference in sampling sample usage (ggml-org#16920) CUDA: add FLOOR, CEIL, ROUND, TRUNC unary ops (ggml-org#16917) devops: fix failing s390x docker build (ggml-org#16918) ...

Add support for Janus Pro

5471f50

ravenouse requested review from CISC and ngxson as code owners October 31, 2025 23:24

github-actions bot added examples python python script changes labels Oct 31, 2025

DajanaV mentioned this pull request Nov 1, 2025

UPSTREAM PR #16906: model: add Janus Pro for image understanding auroralabs-loci/llama.cpp#32

Closed

CISC reviewed Nov 1, 2025

View reviewed changes

gguf-py/gguf/tensor_mapping.py Outdated Show resolved Hide resolved

gguf-py/gguf/tensor_mapping.py Outdated Show resolved Hide resolved

convert_hf_to_gguf.py Outdated Show resolved Hide resolved

ngxson reviewed Nov 1, 2025

View reviewed changes

tools/mtmd/clip.cpp Outdated Show resolved Hide resolved

tools/mtmd/clip.cpp Outdated Show resolved Hide resolved

tools/mtmd/clip.cpp Outdated Show resolved Hide resolved

tools/mtmd/clip.cpp Outdated Show resolved Hide resolved

ravenouse and others added 5 commits November 1, 2025 09:05

Update gguf-py/gguf/tensor_mapping.py

01bd163

Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com>

Update gguf-py/gguf/tensor_mapping.py

d6069df

Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com>

Address reviewer suggestions

d92205e

Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com>

Add JANUS_PRO constant

e260b0e

Update clip model handling

5794785

Co-authored-by: Xuan-Son Nguyen <son@huggingface.co>

ngxson reviewed Nov 2, 2025

View reviewed changes

tools/mtmd/clip.cpp Outdated Show resolved Hide resolved

tools/mtmd/clip.cpp Outdated Show resolved Hide resolved

tools/mtmd/clip.cpp Outdated Show resolved Hide resolved

ravenouse and others added 2 commits November 2, 2025 10:05

Update tools/mtmd/clip.cpp

9601dc8

Co-authored-by: Xuan-Son Nguyen <thichthat@gmail.com>

Refactor JANUS_PRO handling in clip.cpp

38ff44f

Co-authored-by: Xuan-Son Nguyen <son@huggingface.co>

ravenouse force-pushed the januspro branch from 4ba66ee to 38ff44f Compare November 2, 2025 18:24

Merge branch 'master' into januspro

5b35faa

ngxson approved these changes Nov 2, 2025

View reviewed changes

CISC approved these changes Nov 2, 2025

View reviewed changes

CISC requested changes Nov 2, 2025

View reviewed changes

tools/mtmd/clip.cpp Show resolved Hide resolved

CISC reviewed Nov 2, 2025

View reviewed changes

tools/mtmd/clip.cpp Outdated Show resolved Hide resolved

ngxson and others added 2 commits November 2, 2025 21:14

Update tools/mtmd/clip.cpp

c06440f

Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com>

em whitespace

63f7cf3

ngxson requested a review from CISC November 2, 2025 20:18

CISC approved these changes Nov 2, 2025

View reviewed changes

ngxson merged commit 6b9a524 into ggml-org:master Nov 2, 2025
75 of 79 checks passed

model: add Janus Pro for image understanding #16906

model: add Janus Pro for image understanding #16906

Conversation

ravenouse commented Oct 31, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Usage & Current Progress

References

Uh oh!

ravenouse commented Oct 31, 2025

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

ravenouse commented Nov 2, 2025

Uh oh!

Uh oh!

Uh oh!

Uh oh!

ravenouse commented Nov 2, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Uh oh!

ngxson left a comment

Choose a reason for hiding this comment

Uh oh!

CISC commented Nov 2, 2025

Uh oh!

Uh oh!

Uh oh!

Uh oh!

ravenouse commented Nov 3, 2025

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

3 participants

ravenouse commented Oct 31, 2025 •

edited

Loading

ravenouse commented Nov 2, 2025 •

edited

Loading