English

Model ReleasesQwen-Image-2.1

Qwen-Image-2.1 Released as Open-Source Unified Image Generation and Editing Model

The Qwen team has open-sourced Qwen-Image-2.1, a unified model designed for text-to-image generation and image editing. The visual generation component utilizes 7 billion parameters with a 32 single-stream DiT (Diffusion Transformer) layer architecture, balancing generation quality with inference efficiency and versatility.

Key capabilities of the new model include:

  • Text-to-Image Generation: Supports native 2K resolution and various aspect ratios.
  • Image Editing: Supports single-image editing and multi-image composition, allowing up to 10 reference images for multi-subject tasks.
  • Transparent Generation: Natively supports RGBA output for generating images with transparency.

The release also includes fine-tuned Qwen3.5-VL 9B checkpoints designed for prompt rewriting. These models take vague editing instructions and input images to produce precise, actionable prompts optimized for downstream image editing tasks.

For optimal results, the developers recommend using the provided prompt rewriting models to expand short prompts into detailed descriptions. The model is available under the Qwen Research License Agreement and supports various inference frameworks, including Diffusers, vLLM-Omni, and SGLang-Diffusion.

Sources

  1. Qwen/Qwen-Image-2.1-PE-I2I (HF: Qwen, 2026-09-20)
  2. GitHub repo
1 more sourcesHide sources
  1. Qwen/Qwen-Image-2.1-PE-T2I (HF: Qwen, 2026-09-20)