Qwen-Image-2.1: Compact, Efficient, and Unified Image Creation
Qwen-Image-2.1 open-sources a 7B parameter model unifying text-to-image generation and editing with native transparency support. The release introduces mixed-granularity attention to reduce memory usage during multi-image editing tasks. Developers can access weights via Hugging Face and ModelScope, with benchmarks comparing it against other open and closed-source image models.

Qwen released an open-source image model named Qwen-Image-2.1 that combines text-to-image generation with editing features. The system includes a visual generation component with 7B parameters and 32 Single-Stream DiT layers. The architecture supports native transparency for creating and modifying images with transparent backgrounds. Users can access the model weights through Hugging Face and ModelScope platforms. The release aims to balance high image quality with lower computational costs for developers. It introduces a mixed-granularity attention mechanism that reduces memory usage during complex tasks. The model accepts up to 10 reference images to create coherent compositions from multiple assets. It also supports local editing via circles, painted annotations, or separate masks to control specific regions. The provided text does not specify the exact latency or speed improvements resulting from the new attention architecture. It lacks concrete numerical benchmark scores comparing the model to other open or closed-source systems. The description does not detail the specific licensing terms for commercial use of the generated content. The availability of the release on other hosting platforms beyond the two mentioned remains unspecified.