A curated list of checkpoints, quants, prompt engines, LoRAs, and tooling for Qwen-Image 2.1 — Alibaba's 7B unified text-to-image and image-editing model.
Note
Qwen-Image 2.1 was released 2026-09-20 with day-0 support in Diffusers, ComfyUI, vLLM-Omni, SGLang, and LightX2V. It unifies generation and editing in one 7B DiT, adds native RGBA output, and accepts up to 10 reference images. Licensed under the Qwen Research License.
Table of Contents
Pick your entry point based on the runtime you already have.
| I want to… | Use | Why |
|---|---|---|
| Run the reference model in Diffusers | Qwen | Official BF16 diffusers repo, QwenImage21Pipeline |
| Use it in ComfyUI | Comfy-Org | Pre-split folder layout, Day-0 native nodes |
| Fit it in 8–12 GB VRAM | INT4ConvRot-ComfyUI | Complete ComfyUI pack incl. int4 ConvRot DiT + TE |
| Run locally with llama.cpp / GGUF | Unsloth GGUF | Widest quant spread, Q2_K → Q8_0 |
| Generate in 4 steps | Viggle Turbo | Official turbo distill + LoRA variants |
| On a Mac (Apple Silicon) | MLX-4bit | MLX 4-bit, native unified memory |
| Improve prompt quality | PE-T2I | Official prompt rewriter + aspect-ratio picker |
Official resources
- Qwen-Image-2.1 model card — official weights, license, aspect-ratio table
- Qwen-Image-2.1 GitHub repo — inference code, news, supported frameworks
- Qwen blog: Qwen-Image-2.1 — release notes with architecture and benchmark figures
- ComfyUI native workflow example — official ComfyUI tutorial
- ModelScope mirror — official CN mirror
Model facts worth knowing
- 7B parameters in the visual generation component, 32 single-stream DiT layers.
- Mixed-granularity attention (token-level causal for text, chunk-level for image) with prefix KV cache reuse for multi-reference editing.
- Native RGBA output — the prompt decides whether the result has an alpha channel.
- Editing accepts up to 10 reference images, plus circles, painted annotations, or separate masks for local edits.
- Qwen3-VL-8B is the text encoder, so the text encoder is the largest single download at ~17.5 GB BF16.
- Default sample setting is 40 steps without classifier-free guidance; CFG is available for prompt adherence.
◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆
The reference weights. Start here before touching any community conversion.
· · · · · · · · · · · · · ·
| Name | Precision | Layout | DiT | Text encoder | VAE | Links |
|---|---|---|---|---|---|---|
| Qwen-Image-2.1 | diffusers | 14.23 GB | 17.53 GB (Qwen3-VL-8B) | 1.35 GB |
The reference release, and the only repo you need for a standard Diffusers setup. The DiT and text encoder ship as numbered safetensors shards with an index, so pull the repo rather than a single file.
· · · · · · · · · · · · · ·
Comfy-Org — the Day-0 ComfyUI repackage. Files land directly in models/ with no renaming.
| Name | Precision | Size | Links |
|---|---|---|---|
| Image Model | 14.23 GB | ||
| Image Model | 7.26 GB | ||
| Text Encoder | 17.53 GB | ||
| Text Encoder | 9.35 GB | ||
| Text Encoder | 6.31 GB | ||
| Prompt Engine T2I | 9.47 GB | ||
| Prompt Engine I2I | 9.47 GB | ||
| VAE | 0.68 GB |
Tip
ConvRot files are ComfyUI's native rotated-channel integer format. Use a recent ComfyUI build and load them with the standard diffusion-model and text-encoder loaders — no custom nodes required.
Destination folders: Image Model → models/diffusion_models/, Text Encoder and both Prompt Engine rows → models/text_encoders/, VAE → models/vae/.
◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆
Qwen-Image 2.1 conditions on Qwen3-VL-8B, which at BF16 is the largest file in the stack. It also ships dedicated prompt-rewriting models — separate text models that take a short request in any language and return a detailed English prompt plus a recommended aspect ratio. Neither is part of the diffusion path; both are optional, and you can run the base model without either.
· · · · · · · · · · · · · ·
The official pair is fine-tuned Qwen3.5-VL 9B. The Pocket builds are fine-tuned from Qwen3.5 base models instead — roughly 5× smaller, text-to-image only.
| Name | Task | Params | Precision | Size | Links |
|---|---|---|---|---|---|
| PE-T2I | 9B | 18.82 GB | |||
| PE-I2I | 9B | 18.82 GB | |||
| PE-T2I Pocket 2B | 2B | 3.76 GB | |||
| PE-T2I Pocket 0.8B | 0.8B | 1.50 GB |
The official rewriters ship a system_prompt.txt and work with AutoModelForCausalLM + AutoTokenizer. Output is JSON after a reasoning block, so split on the think-tag before parsing. The T2I rewriter also returns a recommended wh_ratio you can map straight to a size.
Each official repo splits into four model-0000N.safetensors shards plus a model.safetensors.index.json, so download the repo — there is no single-file build to link to. The 0.8B Pocket build also ships a Q8_0 GGUF (0.81 GB) for llama.cpp.
· · · · · · · · · · · · · ·
Warning
These repos have the safety refusal direction abliterated. The model no longer refuses prompt content, which means it will happily rewrite anything you ask. Same Qwen Research License as the base weights — the license does not grant you additional rights.
Applied to both halves of the text stack: the text encoder (the conditioning signal itself no longer refuses) and the prompt rewriter.
Start with the TE GGUF build if you want one download: Q4_K_M (5.03 GB), fp8 (9.34 GB), bf16 (17.53 GB), plus a 1.16 GB mmproj projector, all in one repo.
· · · · · · · · · · · · · ·
The INT8 ConvRot pack covers both PE-T2I and PE-I2I in one download. The heretic int8 ConvRot pair does too, and unlike the tensorwise build it keeps the MTP head — 1,395 tensors against 829, one file per task at 9.96 GB.
◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆
Community conversions of the base DiT. Sizes are total weight bytes per repo. Distilled turbo GGUFs live under Turbo instead.
· · · · · · · · · · · · · ·
Transformer-only weights for llama.cpp, sorted from the highest quant down. Q4_K_M is the usual quality/size balance point. Unsloth is the primary source — it carries the widest ladder, so it wins every quant it ships. Where two repos offer a quant Unsloth does not, both are linked in the same cell.
· · · · · · · · · · · · · ·
Every FP4 / NVFP4 / MXFP4 / INT8 / INT4 / W4A4 conversion in one place.
· · · · · · · · · · · · · ·
· · · · · · · · · · · · · ·
SVDQ-based 4-bit for Nunchaku, which targets low-VRAM systems and 4090-class cards.
| Name | Precision | Size | Links | Notes |
|---|---|---|---|---|
| nunchaku | 8.75 GB | Two builds: best_quality_fp4 and svdq-fp4_r32. |
||
| nunchaku lite int4 | 4.15 GB | Half the size of the FP4 build. |
◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆
Few-step students of the base model, distilled with Distribution Matching Distillation. They trade fidelity for speed: useful for iteration and batch work, weaker than 40 steps on multi-reference composition and identity-preserving edits.
Current Viggle release: v0.2.1 (2026-09-24), 6 steps. It supersedes v0.2 (5-step name, sampled at 6) and v0.1 (4 steps). Sample with sigmas=[1.0, 0.9375, 0.875, 0.75, 0.5, 0.25], CFG 1.0, empty negative prompt. Official source: Viggle/Qwen-Image-2.1-viggle-turbo.
Three independent distillation lines are listed below, and their checkpoints are not interchangeable — pick one family. Viggle is the reference DMD line above. Alibaba PAI ships an official 4-step line via Parallel Decoding Distillation (PDD) in VideoX-Fun. Pruna is a third-party DMD line with 8-step and 5-step adapters.
The Pruna adapters are strict about sampling. Each has its own sigma schedule and they are not interchangeable, so load exactly one: 8-step wants sigmas=[1.0, 14/15, 6/7, 10/13, 2/3, 6/11, 0.4, 2/9], 5-step wants sigmas=[1.0, 0.94, 6/7, 2/3, 0.4]. Both need shift=1.0 with dynamic shifting off (otherwise the sigmas get shifted twice), no CFG, no negative prompt, and LoRA strength 1.0. Trained at 1K with up to 3 reference images; 2K runs but sits outside training coverage. Requires a pinned diffusers commit.
· · · · · · · · · · · · · ·
DiT-only quants for ComfyUI-GGUF, from the v0.1 full fine-tune (4-step). Both repos keep FP32 attention norms and BF16 patch embeddings, so do not stack a Viggle LoRA on top. Repos: realrebelai (wider ladder) and Abiray (adds Q4_K_S). Where both ship a quant, both are linked; sizes differ because the builds differ.
| Quant | Size | Download |
|---|---|---|
| 7.69 / 7.59 GB | ||
| 6.91 / 5.88 GB | ||
| 5.96 / 5.01 GB | ||
| 5.56 / 4.19 GB | ||
| 4.06 GB | ||
| 4.19 / 3.19 GB | ||
| 3.77 GB |
· · · · · · · · · · · · · ·
Repo links: Viggle · alibaba-pai · Pruna · chfm (v0.2 snapshot) · t8star · xingewh · RunningHubAI · addlabsviral · cgb
Two ComfyUI conversions of the v0.2.1 adapters exist and are not interchangeable. xingewh re-keys the weights for ComfyUI and loads with the stock Load LoRA node. t8star fuses gate/up for ComfyUI's merged MLP, which is why its files are larger; load those with the built-in LoraLoaderBypassModelOnly at strength 1.0 on a BF16 base, because ordinary merging loaders lose adapter updates to rounding.
| Name | Steps | Precision | Size | Links | Notes |
|---|---|---|---|---|---|
| v0.2.1 LoRA r256 | 6 | 1.36 / 1.36 / 1.76 GB | Start here. Sharpest and most faithful to the 40-step base; runs the demo Space. Load on the base transformer at runtime — do not merge. | ||
| v0.2.1 LoRA r128 | 6 | 0.68 / 0.68 / 0.88 GB | Same adapter cut to rank 128; what the shipped ComfyUI workflows use. | ||
| v0.2.1 adapter (peft) | 6 | 2.72 GB | The v0.2.1 LoRA in PEFT key format, F32 as trained. | ||
| v0.2 LoRA r256 | 5 / 6 | 1.36 GB | Step 600 of the same run. The 5step in the name is the launch schedule; sample at 6 like v0.2.1. |
||
| v0.2 LoRA r128 | 5 / 6 | 0.68 GB | Rank-128 cut of v0.2. | ||
| v0.1 full fine-tune | 4 | 14.23 GB | Merged transformer, no LoRA needed. This is what the GGUF quants above were built from. | ||
| v0.1 LoRA r64 | 4 | 0.34 GB | Superseded — diversity collapsed to 0.75× base and the output was visibly softer. Kept for reproducibility. | ||
| Fun-Acc 4Step (PDD) | 4 | 0.35 GB | Official Alibaba PAI line, rank 64, via Parallel Decoding Distillation. Independent of Viggle — 4 NFE for both T2I and instruction editing. | ||
| Pruna 8Step | 8 | 0.34 GB | Third-party DMD line, rank 64 / alpha 128. Better of the two and the card's default, but v0.1 is marked work in progress. | ||
| Pruna 5Step | 5 | 0.34 GB | Same line, faster and visibly weaker. One adapter at a time — see the schedule note above. | ||
| UltraFast 8Step | 8 | 0.17 GB | Editing-focused rather than T2I, trained on MagicBrush. Rank 32, but loaded by the repo's own inject_dit_lora script, not a stock PEFT or ComfyUI loader. |
||
| Turbo BF16 diffusers | 4 | 32.44 GB | Full pipeline (TE + DiT + VAE), ready to load with QwenImage21Pipeline. |
||
| Turbo FP4 diffusers | 4 | 11.41 GB | Same v0.1 pipeline with an FP4 DiT; the TE is the larger half at 6.73 GB. | ||
| Turbo ONNX (browser) | 4 | ~17.2 GB | r64 LoRA merged into the denoiser, then Q4 MatMulNBits. WebGPU in-browser; needs the FreeGen pipeline and a desktop adapter. Experimental. | ||
| UltraFast Q4_K | 8 | 4.05 GB | DiT-only Q4_K with the UltraFast adapter already merged — do not stack the adapter on top. Still needs the text/vision encoders and the matching VAE. Quality and speed unbenchmarked. | ||
| Pruna 8Step SDNQ | 8 | 11.51 GB | The Pruna 8-step adapter merged into the base and re-quantized with SDNQ — UINT4 dynamic, Hadamard rotation at group size 256. Full diffusers pipeline (DiT 4.10 + TE 6.74 + VAE 0.68), so the text encoder is quantized too. Needs SDNQ 0.2.0+, Triton and diffusers from git. Follows the Pruna sigma schedule above. |
◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆
Style, control, and fix adapters. All target the base DiT unless noted. Grouped by type, then by size.
Caution
Repos marked
◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆
Non-CUDA runtimes and specialized accelerator backends.
· · · · · · · · · · · · · ·
| Name | Format | Precision | Size | Links | Notes |
|---|---|---|---|---|---|
| MLX 8bit | MLX (mflux) | 24.04 GB | For mflux. 13 TE shards. | ||
| Coreml | CoreML .mlpackage |
14.74 GB | 4 transformer blocks (~3.5 GB each) + embed + VAE decoder. | ||
| QIPACK base | QIPACK1 .qipack |
14.23 GB | For the native C++/Metal qwen-image-cplus runtime. 40-step base, defaults to TaylorSeer caching. | ||
| QIPACK distilled | QIPACK1 .qipack |
14.23 GB | Same runtime, 4-step. The Viggle v0.1 full fine-tune, not its LoRA and not v0.2.1. | ||
| MLX 4bit | MLX | 11.59 GB | Smallest viable Apple build. | ||
| Uncensored MLX | MLX (4/6/8-bit) | 7.56 GB | The uncensored line's Apple build, from the GGUF repo. DiT only — 7.56 GB at 8-bit, 5.78 at 6-bit, 4.00 at 4-bit, and the repo ships no MLX text encoder or config, so pair it with one of the two repos above. |
The repo is self-contained: alongside the two packs it now ships the Qwen3-VL-8B text encoder (4 shards, 17.53 GB), the VAE (1.35 GB) and the processor files, so the support download is no longer needed. Needs macOS 14+. 1024×1024 works; 2048×2048 does not yet.
· · · · · · · · · · · · · ·
Alibaba MNN runtime for on-device inference. The full repos are large — the MNN build bundles the text encoder, DiT, and VAE together.
| Name | Precision | Size | Links | Notes |
|---|---|---|---|---|
| MNN fp16 | 30.72 GB | Highest fidelity, and by far the largest. | ||
| MNN int8 | 21.42 GB | |||
| MNN int4 | 14.39 GB | |||
| MNN | 10.62 GB | llm.mnn.weight 4.73 + dit.mnn.weight 4.47 + VAE 0.51 GB. |
· · · · · · · · · · · · · ·
p150 is the standout entry here. A full port to a single Tenstorrent Blackhole p150a via tt-nn, with all three sub-models resident on-chip. ~20.5 s for a 40-step 1024² generation (513 ms/step sustained), and it beats an RTX 5090 with CPU offload by 1.4–1.6× end to end. Includes editing support for 1–4 condition images. No weights are redistributed — it pulls the official ones.
ROCm gfx1151 targets AMD Strix Halo / RX 9070-class iGPUs. Code only; no weights in the repo.
Intel Arc (SYCL). Frosty40 is a serving package for the sd.cpp SYCL backend, not a fine-tune — "turbo" refers to the recipe, not to step distillation. Despite the name the weights are Qwen's own, quantized from a BF16 master.
| Name | Precision | Size | Links | Notes |
|---|---|---|---|---|
| SYCL quality | 7.89 GB | Attention qkv/o kept at F16; 73 BF16 + 128 F16 + 96 Q5_0. | ||
| SYCL lean | 5.07 GB | All-exception Q5_0, 73 BF16 + 224 Q5_0. |
Both ship with a pinned Qwen3-VL-8B-Instruct Q8_0 text encoder (8.71 GB) and the BF16 VAE (0.68 GB). Two non-obvious requirements: --vae-tiling is mandatory at 1024² or the SYCL VAE overflows int32, and one sd-cli per GPU — a second concurrent run wedges the xe driver and needs a root-only reset. The 3.84× speedup comes from --eager-load --cache-mode easycache, not from fewer steps.
FlagOS ships eight builds for Chinese NPUs, same 28-file layout in each:
· · · · · · · · · · · · · ·
Same weights, re-laid-out for a ComfyUI-side loader. Each ships a YAML manifest next to the shards, so no diffusers config is needed.
· · · · · · · · · · · · · ·
| Name | Precision | Size | Links | Notes |
|---|---|---|---|---|
| hdr vae test | 0.68 GB | Untested in the wild; treat as an experiment, not a drop-in replacement. |
◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆
Gotchas that will save you an afternoon
QwenImage21Pipelinerequires diffusersmain— release0.40.0does not have it. Install from git.- The CFG kwarg is
true_cfg_scale, notguidance_scale, and it needs anegative_promptto engage. - Width and height must be multiples of 32, max edge 2048.
- With CPU offload, the generator should live on
'cpu'. - The default sample is 40 steps with no CFG. Adding CFG changes the look; don't assume more steps help.
- Chinese prompts work natively — the PE models are what rewrite them into English.
◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆
This list is a set of Markdown files, not one. README.md is generated and
should never be edited directly — your change will be overwritten on the next
build. Each section lives in its own file under blocks/:
| To change… | Edit this file |
|---|---|
| Title, badges, table of contents | blocks/00_header.md |
| Base model or official ComfyUI files | blocks/10_official_checkpoints.md |
| Prompt engines, rewriters, text encoders | blocks/20_text_encoders_prompt_engines.md |
| GGUF, 4/8-bit, FP8, Nunchaku quants | blocks/30_quantizations.md |
| Turbo and step-distilled models | blocks/40_turbo.md |
| LoRAs and adapters | blocks/50_loras.md |
| Apple, MNN, AMD, VAE ports | blocks/60_platform_ports.md |
| Tools, notebooks, training scripts | blocks/70_tools.md |
| Badge and link definitions | blocks/99_footer.md |
The file prefix controls the order — the build concatenates blocks/*.md in
filename order, so 10_ always lands before 20_. A new section just needs a
new number.
Awesome Qwen-Image 2.1
