Skip to content

Latest commit

 

History

24 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 

Repository files navigation

Awesome Qwen-Image 2.1

A curated list of checkpoints, quants, prompt engines, LoRAs, and tooling for Qwen-Image 2.1 — Alibaba's 7B unified text-to-image and image-editing model.

awesome-qwen-image

Hugging Face License Last Commit PRs Welcome

Note

Qwen-Image 2.1 was released 2026-09-20 with day-0 support in Diffusers, ComfyUI, vLLM-Omni, SGLang, and LightX2V. It unifies generation and editing in one 7B DiT, adds native RGBA output, and accepts up to 10 reference images. Licensed under the Qwen Research License.

Table of Contents

⌬ Quick start

Pick your entry point based on the runtime you already have.

I want to… Use Why
Run the reference model in Diffusers Qwen Official BF16 diffusers repo, QwenImage21Pipeline
Use it in ComfyUI Comfy-Org Pre-split folder layout, Day-0 native nodes
Fit it in 8–12 GB VRAM INT4ConvRot-ComfyUI Complete ComfyUI pack incl. int4 ConvRot DiT + TE
Run locally with llama.cpp / GGUF Unsloth GGUF Widest quant spread, Q2_K → Q8_0
Generate in 4 steps Viggle Turbo Official turbo distill + LoRA variants
On a Mac (Apple Silicon) MLX-4bit MLX 4-bit, native unified memory
Improve prompt quality PE-T2I Official prompt rewriter + aspect-ratio picker

Official resources

Model facts worth knowing

  • 7B parameters in the visual generation component, 32 single-stream DiT layers.
  • Mixed-granularity attention (token-level causal for text, chunk-level for image) with prefix KV cache reuse for multi-reference editing.
  • Native RGBA output — the prompt decides whether the result has an alpha channel.
  • Editing accepts up to 10 reference images, plus circles, painted annotations, or separate masks for local edits.
  • Qwen3-VL-8B is the text encoder, so the text encoder is the largest single download at ~17.5 GB BF16.
  • Default sample setting is 40 steps without classifier-free guidance; CFG is available for prompt adherence.

◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆

▓ Official checkpoints

The reference weights. Start here before touching any community conversion.

· · · · · · · · · · · · · ·

▣ Base model

Name Precision Layout DiT Text encoder VAE Links
Qwen-Image-2.1 bf16 diffusers 14.23 GB 17.53 GB (Qwen3-VL-8B) 1.35 GB

The reference release, and the only repo you need for a standard Diffusers setup. The DiT and text encoder ship as numbered safetensors shards with an index, so pull the repo rather than a single file.

· · · · · · · · · · · · · ·

▣ ComfyUI official

Comfy-Org — the Day-0 ComfyUI repackage. Files land directly in models/ with no renaming.

Name Precision Size Links
Image Model bf16 14.23 GB
Image Model int8 7.26 GB
Text Encoder bf16 17.53 GB
Text Encoder int8 9.35 GB
Text Encoder w4a8 6.31 GB
Prompt Engine T2I int8 9.47 GB
Prompt Engine I2I int8 9.47 GB
VAE bf16 0.68 GB

Tip

ConvRot files are ComfyUI's native rotated-channel integer format. Use a recent ComfyUI build and load them with the standard diffusion-model and text-encoder loaders — no custom nodes required.

Destination folders: Image Model → models/diffusion_models/, Text Encoder and both Prompt Engine rows → models/text_encoders/, VAE → models/vae/.

◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆

⚲ Text encoders & prompt engines

Qwen-Image 2.1 conditions on Qwen3-VL-8B, which at BF16 is the largest file in the stack. It also ships dedicated prompt-rewriting models — separate text models that take a short request in any language and return a detailed English prompt plus a recommended aspect ratio. Neither is part of the diffusion path; both are optional, and you can run the base model without either.

· · · · · · · · · · · · · ·

▣ Prompt rewriters

The official pair is fine-tuned Qwen3.5-VL 9B. The Pocket builds are fine-tuned from Qwen3.5 base models instead — roughly 5× smaller, text-to-image only.

Name Task Params Precision Size Links
PE-T2I text → image 9B bf16 18.82 GB
PE-I2I image → image 9B bf16 18.82 GB
PE-T2I Pocket 2B text → image 2B bf16 3.76 GB
PE-T2I Pocket 0.8B text → image 0.8B Q8_0 1.50 GB

The official rewriters ship a system_prompt.txt and work with AutoModelForCausalLM + AutoTokenizer. Output is JSON after a reasoning block, so split on the think-tag before parsing. The T2I rewriter also returns a recommended wh_ratio you can map straight to a size.

Each official repo splits into four model-0000N.safetensors shards plus a model.safetensors.index.json, so download the repo — there is no single-file build to link to. The 0.8B Pocket build also ships a Q8_0 GGUF (0.81 GB) for llama.cpp.

· · · · · · · · · · · · · ·

▣ Heretic & abliterated

Warning

These repos have the safety refusal direction abliterated. The model no longer refuses prompt content, which means it will happily rewrite anything you ask. Same Qwen Research License as the base weights — the license does not grant you additional rights.

Applied to both halves of the text stack: the text encoder (the conditioning signal itself no longer refuses) and the prompt rewriter.

Type Name Task Precision Size Links
TE Heretic TE GGUF Q4_K_M fp8 bf16 33.07 GB
TE Heretic TE NVFP4 nvfp4 6.31 GB
TE Heretic TE int8 ConvRot int8 9.35 GB
TE Heretic TE W4A8 w4a8 6.31 GB
TE Heretic TE BF16 bf16 17.53 GB
PE PE-T2I Heretic GGUF text → image Q4_K_M 5.89 GB
PE PE-I2I Heretic GGUF image → image Q4_K_M 6.81 GB
PE ComfyUI PE bundle both Q4_K_M 18.07 GB
PE PE-I2I Abliterated image → image bf16 18.82 GB
PE PE-T2I Heretic text → image bf16 18.82 GB
PE PE-I2I Heretic image → image bf16 18.82 GB
PE PE-T2I Heretic NVFP4 text → image nvfp4 11.20 GB
PE PE-I2I Heretic NVFP4 image → image nvfp4 11.20 GB

Start with the TE GGUF build if you want one download: Q4_K_M (5.03 GB), fp8 (9.34 GB), bf16 (17.53 GB), plus a 1.16 GB mmproj projector, all in one repo.

· · · · · · · · · · · · · ·

▣ Quantized & ported

Name Format Precision Size Links
PE ComfyUI pack BF16 + int8 ConvRot bf16 int8 64.68 GB
PE-T2I MLX MLX (4/8/16-bit) int4 int8 bf16 35.20 GB
PE-I2I MLX MLX (4/8/16-bit) int4 int8 bf16 35.20 GB
Prompt Enhancement INT8 int8 ConvRot int8 24.69 GB
Heretic T2I int8 tensorwise int8 tensorwise ConvRot int8 9.99 GB
Heretic PE int8 ConvRot int8 ConvRot (T2I + I2I) int8 19.91 GB

The INT8 ConvRot pack covers both PE-T2I and PE-I2I in one download. The heretic int8 ConvRot pair does too, and unlike the tensorwise build it keeps the MTP head — 1,395 tensors against 829, one file per task at 9.96 GB.

◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆

◈ Quantizations

Community conversions of the base DiT. Sizes are total weight bytes per repo. Distilled turbo GGUFs live under Turbo instead.

· · · · · · · · · · · · · ·

▣ GGUF

Transformer-only weights for llama.cpp, sorted from the highest quant down. Q4_K_M is the usual quality/size balance point. Unsloth is the primary source — it carries the widest ladder, so it wins every quant it ships. Where two repos offer a quant Unsloth does not, both are linked in the same cell.

Quant Size Download
F16 14.23 GB
BF16 14.23 GB
Q8_0 7.64 GB
Q6_K_XL 6.72 GB
Q6_K 6.27 GB
Q5_K_M 5.39 GB
Q5_K_S 4.50 GB
Q4 · HVQ3 (sd.cpp) 5.96 GB
NVFP4 4.05 GB ┊
Q4_K_M 4.20 GB
Q4_0 4.15 GB ┊
Q4_K_S 3.91 GB
Q3_K_XL 3.61 GB
Q3_K_M 3.17 GB
Q3_K_S 2.72 GB
Q2_K 2.47 GB
Text encoder NVFP4 6.30 GB
Text encoder Q4_K_M 5.03 GB
mmproj projector Q8_0 0.75 GB
VAE BF16 0.68 GB
VAE F16 0.68 GB

· · · · · · · · · · · · · ·

▣ 4-bit & 8-bit

Every FP4 / NVFP4 / MXFP4 / INT8 / INT4 / W4A4 conversion in one place.

Name Precision Size Links Notes
INT4ConvRot-ComfyUI bf16 int8 int4 w4a8 65.91 GB Best all-in-one ComfyUI pack. DiT and TE at three precisions plus VAE, in correct folder layout.
Darkstar ModelOpt W4A4 NVFP4 nvfp4 23.66 GB NVIDIA ModelOpt-derived, full repo.
INT8 int8 17.96 GB Full diffusers repo.
DiT NVFP4 ComfyUI nvfp4 13.82 GB Three ComfyUI cuts: nvfp4, nvfp4_T2, nvfp4_T3.
FP4 fp4 11.74 GB DiT 6.46 + TE 4.94 + VAE 0.34.
INT4 int4 11.08 GB Full diffusers repo.
bnb 4bit int4 11.41 GB Standard bitsandbytes NF4.
MXFP4 (Paiton) mxfp4 9.31 GB Paiton backend, 57 shards.
MXFP4 Paiton RDNA4 mxfp4 9.31 GB RDNA4-specific kernel variant, identical layout.
Uncensored MXFP4 Paiton mxfp4 9.31 GB ⚠️ Uncensored, derived from the abenzerps GGUF.
W4A4 NVFP4 nvfp4 4.88 GB DiT only; near-identical to the INT4 build.
W4A4 INT4 int4 4.66 GB DiT only.

· · · · · · · · · · · · · ·

▣ FP8 & bf16

Name Precision Size Links Notes
Darkstar ModelOpt FP8 fp8 26.33 GB NVIDIA ModelOpt-derived, full repo.
FP8 fp8 17.96 GB Closest thing to a drop-in smaller BF16.
Uncensored BF16 SafeTensor bf16 14.23 GB ⚠️ ┊ Single 14.23 GB file — the full DiT in one piece, mirrored by both repos.
DF11 ComfyUI bf16 9.72 GB qwen_image_2.1_bf16-DF11.safetensors.

· · · · · · · · · · · · · ·

▣ Nunchaku (SVDQ)

SVDQ-based 4-bit for Nunchaku, which targets low-VRAM systems and 4090-class cards.

Name Precision Size Links Notes
nunchaku fp4 8.75 GB Two builds: best_quality_fp4 and svdq-fp4_r32.
nunchaku lite int4 int4 4.15 GB Half the size of the FP4 build.

◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆

▸ Turbo & step distillation

Few-step students of the base model, distilled with Distribution Matching Distillation. They trade fidelity for speed: useful for iteration and batch work, weaker than 40 steps on multi-reference composition and identity-preserving edits.

Current Viggle release: v0.2.1 (2026-09-24), 6 steps. It supersedes v0.2 (5-step name, sampled at 6) and v0.1 (4 steps). Sample with sigmas=[1.0, 0.9375, 0.875, 0.75, 0.5, 0.25], CFG 1.0, empty negative prompt. Official source: Viggle/Qwen-Image-2.1-viggle-turbo.

Three independent distillation lines are listed below, and their checkpoints are not interchangeable — pick one family. Viggle is the reference DMD line above. Alibaba PAI ships an official 4-step line via Parallel Decoding Distillation (PDD) in VideoX-Fun. Pruna is a third-party DMD line with 8-step and 5-step adapters.

The Pruna adapters are strict about sampling. Each has its own sigma schedule and they are not interchangeable, so load exactly one: 8-step wants sigmas=[1.0, 14/15, 6/7, 10/13, 2/3, 6/11, 0.4, 2/9], 5-step wants sigmas=[1.0, 0.94, 6/7, 2/3, 0.4]. Both need shift=1.0 with dynamic shifting off (otherwise the sigmas get shifted twice), no CFG, no negative prompt, and LoRA strength 1.0. Trained at 1K with up to 3 reference images; 2K runs but sits outside training coverage. Requires a pinned diffusers commit.

· · · · · · · · · · · · · ·

▣ Turbo GGUF

DiT-only quants for ComfyUI-GGUF, from the v0.1 full fine-tune (4-step). Both repos keep FP32 attention norms and BF16 patch embeddings, so do not stack a Viggle LoRA on top. Repos: realrebelai (wider ladder) and Abiray (adds Q4_K_S). Where both ship a quant, both are linked; sizes differ because the builds differ.

Quant Size Download
Q8_0 7.69 / 7.59 GB ┊
Q6_K 6.91 / 5.88 GB ┊
Q5_K_M 5.96 / 5.01 GB ┊
Q4_K_M 5.56 / 4.19 GB ┊
Q4_K_S 4.06 GB
Q3_K_M 4.19 / 3.19 GB ┊
Q2_K 3.77 GB

· · · · · · · · · · · · · ·

▣ Official & converted

Repo links: Viggle · alibaba-pai · Pruna · chfm (v0.2 snapshot) · t8star · xingewh · RunningHubAI · addlabsviral · cgb

Two ComfyUI conversions of the v0.2.1 adapters exist and are not interchangeable. xingewh re-keys the weights for ComfyUI and loads with the stock Load LoRA node. t8star fuses gate/up for ComfyUI's merged MLP, which is why its files are larger; load those with the built-in LoraLoaderBypassModelOnly at strength 1.0 on a BF16 base, because ordinary merging loaders lose adapter updates to rounding.

Name Steps Precision Size Links Notes
v0.2.1 LoRA r256 6 bf16 1.36 / 1.36 / 1.76 GB ┊ ┊ Start here. Sharpest and most faithful to the 40-step base; runs the demo Space. Load on the base transformer at runtime — do not merge.
v0.2.1 LoRA r128 6 bf16 0.68 / 0.68 / 0.88 GB ┊ ┊ Same adapter cut to rank 128; what the shipped ComfyUI workflows use.
v0.2.1 adapter (peft) 6 fp32 2.72 GB The v0.2.1 LoRA in PEFT key format, F32 as trained.
v0.2 LoRA r256 5 / 6 bf16 1.36 GB ┊ ┊ Step 600 of the same run. The 5step in the name is the launch schedule; sample at 6 like v0.2.1.
v0.2 LoRA r128 5 / 6 bf16 0.68 GB Rank-128 cut of v0.2.
v0.1 full fine-tune 4 bf16 14.23 GB Merged transformer, no LoRA needed. This is what the GGUF quants above were built from.
v0.1 LoRA r64 4 bf16 0.34 GB ┊ ┊ Superseded — diversity collapsed to 0.75× base and the output was visibly softer. Kept for reproducibility.
Fun-Acc 4Step (PDD) 4 bf16 0.35 GB Official Alibaba PAI line, rank 64, via Parallel Decoding Distillation. Independent of Viggle — 4 NFE for both T2I and instruction editing.
Pruna 8Step 8 bf16 0.34 GB ┊ Third-party DMD line, rank 64 / alpha 128. Better of the two and the card's default, but v0.1 is marked work in progress.
Pruna 5Step 5 bf16 0.34 GB ┊ Same line, faster and visibly weaker. One adapter at a time — see the schedule note above.
UltraFast 8Step 8 bf16 0.17 GB Editing-focused rather than T2I, trained on MagicBrush. Rank 32, but loaded by the repo's own inject_dit_lora script, not a stock PEFT or ComfyUI loader.
Turbo BF16 diffusers 4 bf16 32.44 GB Full pipeline (TE + DiT + VAE), ready to load with QwenImage21Pipeline.
Turbo FP4 diffusers 4 fp4 11.41 GB Same v0.1 pipeline with an FP4 DiT; the TE is the larger half at 6.73 GB.
Turbo ONNX (browser) 4 int4 ~17.2 GB r64 LoRA merged into the denoiser, then Q4 MatMulNBits. WebGPU in-browser; needs the FreeGen pipeline and a desktop adapter. Experimental.
UltraFast Q4_K 8 Q4_K 4.05 GB DiT-only Q4_K with the UltraFast adapter already merged — do not stack the adapter on top. Still needs the text/vision encoders and the matching VAE. Quality and speed unbenchmarked.
Pruna 8Step SDNQ 8 int4 11.51 GB The Pruna 8-step adapter merged into the base and re-quantized with SDNQ — UINT4 dynamic, Hadamard rotation at group size 256. Full diffusers pipeline (DiT 4.10 + TE 6.74 + VAE 0.68), so the text encoder is quantized too. Needs SDNQ 0.2.0+, Triton and diffusers from git. Follows the Pruna sigma schedule above.

◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆

◉ LoRA & adapters

Style, control, and fix adapters. All target the base DiT unless noted. Grouped by type, then by size.

Name Type Precision Size Links Notes
ControlNet-Union Control bf16 7.55 GB Official Alibaba PAI / VideoX-Fun. One checkpoint for 8 conditions (Canny, Depth, Grayscale, HED, Lineart, MLSD, Pose, Scribble) plus inpainting. Control branch only, 16 injection points, loaded strict=False.
Object Mover Bbox Preview Control bf16 0.50 GB Bbox object moving, 6 checkpoints.
Object Remover Bbox Preview Control bf16 0.50 GB Bbox object removal, full-quality variant.
Object Remover Bbox turbo Control bf16 0.42 GB 4-step-compatible variant.
BFS Head Swap v1 Control bf16 0.32 GB Head replacement, not a whole-face blend: identity, hair, eye colour and nose come from image 2 while gaze direction, head rotation and expression stay with image 1. Rank 64, 5,000 steps, MIT. Image order is load-bearing — swapping the two inputs swaps who is retargeted. Trigger with the head_swap: prefix. The repo's other 17 adapters target Qwen Image Edit 2509/2511, Flux 2 Klein, Krea 2 and LTX-2, not this base.
Orbit Alpha Control bf16 0.17 GB Official ML-Intern-lab. One RGBA image in, the same object from a new viewpoint out. Rank 32, 2,000 steps at 768 px. Use the _gate_up_split file — the other checkpoint in the repo does not load correctly. <orbit> grammar, 40 steps, no CFG.
Outpaint v2 Control bf16 0.16 GB Pad the picture with flat #808080, hand the padded canvas to the model as the reference; the adapter fills the gray and keeps the original pixel-registered. One side, a corner, or all four. Rank 32, ComfyUI keys, 2,000 steps. 25 steps, CFG 1, resolution 0, target 1–2 MP. Do not pin the known area with a latent noise mask — on this model it draws a visible rectangle at the seam.
Outpaint v1 Control bf16 0.16 GB Same adapter one step earlier, trained at ≤1 MP over more extreme zoom-outs. The better of the two on very large extensions; the two are within noise on ordinary crops. The repo also keeps the step-500 and step-1250 intermediates.
Fix Fix bf16 0.11 GB The most-liked community LoRA.
De-AI LoRA pack Style bf16 2.45 GB 8-file "remove the AI look" pack (CN filenames).
Natural Exposure LoRA Style bf16 0.42 GB Exposure correction, 5 checkpoints.
De-AI + lighting v5 Style bf16 0.17 GB Two revisions of the same de-AI + lighting adapter; v5 is the newer. The repo also has two lighting-only adapters, 0.24 and 0.09 GB. CN filenames.
Sts2 Cards Drawer Style fp16 0.10 GB deckbuilder_cardart_style_lora_v1_fp16.
Normal2RGB Utility bf16 0.25 GB Normal map → RGB render, 3 checkpoints.
Hips / buttocks NSFW bf16 0.47 GB ⚠️ Body-shape slider, ported from a Qwen-2.5-09 derivative. CN filename.
Breasts and hips NSFW bf16 0.17 GB ⚠️ Combined body-shape adapter from the same pack. CN filename.
RadianceChrome Voluptuous NSFW bf16 0.17 GB ⚠️ Character-style LoRA.
NSFW LoRA NSFW bf16 0.16 GB ⚠️ Civitai original by TheseAlpacas, mirrored unmodified. Rank 32, 192 targets. The author asks for 25+ steps, er_sde and the beta scheduler on an INT8 ConvRot base — it is not a few-step adapter.
NSFW Image Edit NSFW bf16 0.08 GB ⚠️ Editing LoRA, uploaded 2026-09-24.
Breasts Slider V1 NSFW bf16 0.003 GB ⚠️ Slider control, 3 MB.

Caution

Repos marked ⚠️ are uncensored, abliterated, or NSFW. They are listed for completeness because they are widely used — the uncensored GGUF in particular is the most-downloaded repo in this ecosystem. They carry the same Qwen Research License as the base model; a research license is not a license to do whatever you want, and you remain responsible for how you use them. Sizes are the LoRA file only; none of these merge a base model.

◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆

⬡ Platform ports

Non-CUDA runtimes and specialized accelerator backends.

· · · · · · · · · · · · · ·

▣ Apple Silicon

Name Format Precision Size Links Notes
MLX 8bit MLX (mflux) int8 24.04 GB For mflux. 13 TE shards.
Coreml CoreML .mlpackage bf16 14.74 GB 4 transformer blocks (~3.5 GB each) + embed + VAE decoder.
QIPACK base QIPACK1 .qipack fp16 14.23 GB For the native C++/Metal qwen-image-cplus runtime. 40-step base, defaults to TaylorSeer caching.
QIPACK distilled QIPACK1 .qipack fp16 14.23 GB Same runtime, 4-step. The Viggle v0.1 full fine-tune, not its LoRA and not v0.2.1.
MLX 4bit MLX int4 11.59 GB Smallest viable Apple build.
Uncensored MLX MLX (4/6/8-bit) int4 int8 7.56 GB ⚠️ The uncensored line's Apple build, from the GGUF repo. DiT only — 7.56 GB at 8-bit, 5.78 at 6-bit, 4.00 at 4-bit, and the repo ships no MLX text encoder or config, so pair it with one of the two repos above.

The repo is self-contained: alongside the two packs it now ships the Qwen3-VL-8B text encoder (4 shards, 17.53 GB), the VAE (1.35 GB) and the processor files, so the support download is no longer needed. Needs macOS 14+. 1024×1024 works; 2048×2048 does not yet.

· · · · · · · · · · · · · ·

▣ Mobile & edge (MNN)

Alibaba MNN runtime for on-device inference. The full repos are large — the MNN build bundles the text encoder, DiT, and VAE together.

Name Precision Size Links Notes
MNN fp16 fp16 30.72 GB Highest fidelity, and by far the largest.
MNN int8 int8 21.42 GB
MNN int4 int4 14.39 GB
MNN int4 10.62 GB llm.mnn.weight 4.73 + dit.mnn.weight 4.47 + VAE 0.51 GB.

· · · · · · · · · · · · · ·

▣ AMD & domestic accelerators

p150 is the standout entry here. A full port to a single Tenstorrent Blackhole p150a via tt-nn, with all three sub-models resident on-chip. ~20.5 s for a 40-step 1024² generation (513 ms/step sustained), and it beats an RTX 5090 with CPU offload by 1.4–1.6× end to end. Includes editing support for 1–4 condition images. No weights are redistributed — it pulls the official ones.

ROCm gfx1151 targets AMD Strix Halo / RX 9070-class iGPUs. Code only; no weights in the repo.

Intel Arc (SYCL). Frosty40 is a serving package for the sd.cpp SYCL backend, not a fine-tune — "turbo" refers to the recipe, not to step distillation. Despite the name the weights are Qwen's own, quantized from a BF16 master.

Name Precision Size Links Notes
SYCL quality q5_0 7.89 GB Attention qkv/o kept at F16; 73 BF16 + 128 F16 + 96 Q5_0.
SYCL lean q5_0 5.07 GB All-exception Q5_0, 73 BF16 + 224 Q5_0.

Both ship with a pinned Qwen3-VL-8B-Instruct Q8_0 text encoder (8.71 GB) and the BF16 VAE (0.68 GB). Two non-obvious requirements: --vae-tiling is mandatory at 1024² or the SYCL VAE overflows int32, and one sd-cli per GPU — a second concurrent run wedges the xe driver and needs a root-only reset. The 3.84× speedup comes from --eager-load --cache-mode easycache, not from fewer steps.

FlagOS ships eight builds for Chinese NPUs, same 28-file layout in each:

Name Target Precision Size Links
BF16 nvidia NVIDIA (FlagOS path) bf16 33.13 GB
BF16 hygon Hygon DCU bf16 33.13 GB
BF16 ascend Ascend NPU bf16 33.13 GB
BF16 metax MetaX CGC bf16 33.13 GB
BF16 enflame Enflame GCU bf16 33.13 GB
BF16 zhenwu Zhenwu MUSA bf16 33.13 GB
BF16 mthreads Moore Threads MUSA bf16 33.13 GB
W8A8 arm ARM w8a8 29.38 GB

· · · · · · · · · · · · · ·

▣ ComfyUI package formats

Same weights, re-laid-out for a ComfyUI-side loader. Each ships a YAML manifest next to the shards, so no diffusers config is needed.

Name Format Precision Size Links Notes
libwaifu bf16 libwaifu (yaml + 8 shards) bf16 30.03 GB Full-precision layout for the libwaifu loader.
libwaifu fp8 libwaifu (yaml + 4 shards) fp8 15.98 GB Same layout, weight_format: fp8. Half the download.

· · · · · · · · · · · · · ·

▣ Experimental VAE

Name Precision Size Links Notes
hdr vae test fp16 0.68 GB Untested in the wild; treat as an experiment, not a drop-in replacement.

◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆

⚙ Tools & notebooks

Name Type Links Notes
training assistant v1 Training A working pytorch_lora_weights.safetensors + training_details.json you can inspect to learn a real SimpleTuner config.
qwen_image2.1_colab Notebook 7-cell notebook with form UI and enable_model_cpu_offload(). Works on L4 22 GB; T4 16 GB will OOM by design.
qwen_image2.1_molab Notebook Script version of the above, tuned for Marimo and Blackwell.
Qwen-Image-2.1-Skills Agent skill Turns a short scene into a structured prompt for believable casual phone photography. 22 example images.
qwen-image-2.1-p150 Port Tenstorrent Blackhole p150a. tt-model pull --with-weights then tt-model serve.

Gotchas that will save you an afternoon

  1. QwenImage21Pipeline requires diffusers main — release 0.40.0 does not have it. Install from git.
  2. The CFG kwarg is true_cfg_scale, not guidance_scale, and it needs a negative_prompt to engage.
  3. Width and height must be multiples of 32, max edge 2048.
  4. With CPU offload, the generator should live on 'cpu'.
  5. The default sample is 40 steps with no CFG. Adding CFG changes the look; don't assume more steps help.
  6. Chinese prompts work natively — the PE models are what rewrite them into English.

◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆◇◆

✎ Contributing

This list is a set of Markdown files, not one. README.md is generated and should never be edited directly — your change will be overwritten on the next build. Each section lives in its own file under blocks/:

To change… Edit this file
Title, badges, table of contents blocks/00_header.md
Base model or official ComfyUI files blocks/10_official_checkpoints.md
Prompt engines, rewriters, text encoders blocks/20_text_encoders_prompt_engines.md
GGUF, 4/8-bit, FP8, Nunchaku quants blocks/30_quantizations.md
Turbo and step-distilled models blocks/40_turbo.md
LoRAs and adapters blocks/50_loras.md
Apple, MNN, AMD, VAE ports blocks/60_platform_ports.md
Tools, notebooks, training scripts blocks/70_tools.md
Badge and link definitions blocks/99_footer.md

The file prefix controls the order — the build concatenates blocks/*.md in filename order, so 10_ always lands before 20_. A new section just needs a new number.

Awesome Qwen-Image 2.1

About

Qwen-Image 2.1. Checkpoints, quants, prompt engines, LoRAs, and tooling

Topics

Resources

Stars

107 stars

Watchers

1 watching

Forks

Contributors