Repository: localaiLicense: apache-2.0

We are thrilled to release Qwen-Image, an image generation foundation model in the Qwen series that achieves significant advances in complex text rendering and precise image editing. Experiments show strong general capabilities in both image generation and editing, with exceptional performance in text rendering, especially for Chinese.
Links
Tags

Qwen-Image-Edit is a model for image editing, which is based on Qwen-Image.
Links
Tags

Qwen-Image-Edit is a model for image editing, which is based on Qwen-Image.
Links
Tags
Repository: localaiLicense: qwen-research

Qwen-Image 2.1 is the Qwen image generation foundation model, with strong prompt adherence and accurate text rendering (English and Chinese). It uses Qwen3-VL-8B as the text encoder and its own VAE, and supports both text-to-image and image editing: pass reference images to edit them. This is the Q4_K (4-bit) quantization (~4.2GB diffusion model) by leejet for stable-diffusion.cpp. The bundle also pulls the Qwen3-VL-8B-Instruct text encoder with its vision projector (used for image editing) and the Qwen-Image 2.1 VAE. Use image dimensions divisible by 32.
Links
Tags
Repository: localaiLicense: qwen-research

Qwen-Image 2.1 is the Qwen image generation foundation model, with strong prompt adherence and accurate text rendering (English and Chinese). It uses Qwen3-VL-8B as the text encoder and its own VAE, and supports both text-to-image and image editing: pass reference images to edit them. This is the Q8_0 (8-bit) quantization (~7.7GB diffusion model) by leejet for stable-diffusion.cpp. The bundle also pulls the Qwen3-VL-8B-Instruct text encoder with its vision projector (used for image editing) and the Qwen-Image 2.1 VAE. Use image dimensions divisible by 32.
Links
Tags
Repository: localaiLicense: other

🤖 ModelScope | 🤗 HuggingFace | 📑 Blog | 🖥️ Demo | 🫨 Discord | 💬 WeChat ## Introduction We are excited to open-source **Qwen-Image-2.1**, a unified text-to-image generation and image editing model in the Qwen family. With just **7B parameters in its visual generation component** (32 Single-Stream DiT layers), Qwen-Image-2.1 balances generation quality, inference efficiency, and versatility. Four key improvements define this release: ...
Links
Tags
Repository: localaiLicense: apache-2.0
> 🍱 **All-GGUF Qwen-Image-2.1:** text encoder GGUF + DiT GGUF (Q8_0 / Q6_K / Q4_K_M, loads in stock ComfyUI-GGUF) + PE-T2I GGUF. > [!IMPORTANT] > **The 4 `model-0000x-of-00004` shards are the HF `transformers` checkpoint (bf16) — they do not load in ComfyUI.** > ComfyUI text encoders use a different key layout (the `model.language_model.` prefix is dropped > when repacking), so pointing ComfyUI at these shards will not work. > > For ComfyUI use one of these instead — all keep the vision tower, which 2.1 needs for editing. > **Mac / non-CUDA:** use the single file `qwen3vl_8b_bf16_heretic.safetensors` in *this* repo (full bf16, 17.5 GB — also in the GGUF repo) — > put it in `models/text_encoders/` and load it with the stock `CLIPLoader`, type `qwen_image`, then `TextEncodeQwenImage21`. > It is the only format here that needs no CUDA-specific kernels. > > | Repo | File | Loader | > |---|---|---| > | **this repo** | `qwen3vl_8b_bf16_heretic.safetensors` (bf16 single file, any device) | `CLIPLoader`, type `qwen_image` | > | `…-NVFP4` | `qwen3vl_8b_nvfp4_heretic.safetensors` | `CLIPLoader`, type `qwen_image` | > | `…-W4A8` | `qwen3vl_8b_w4a8_heretic.safetensors` | `CLIPL ...
Links
Tags