Qwen-Image-2.1 packs image generation and editing into a 7B model

Qwen-Image-2.1: Compact, efficient, and unified image creation

Qwen-Image-2.1 is an open-source image model that unifies text-to-image generation and editing in a single 7B visual generation component. It natively generates and edits transparent images, accepts up to 10 reference images for composition, supports local edits via circles, painted annotations, or masks, and preserves portrait identity and product details. A mixed-granularity attention architecture with KV cache reuse speeds up multi-image editing while improving typography and portrait realism.

With a compact 7B visual generation component, Qwen-Image-2.1 brings image generation, transparency, and a wide range of editing capabilities together in one model.
  1. vunderba

    So thoughts

    Positives

    • It's a heck of a lot smaller than Qwen-Image 1 (20b parameters) at only 7b, making it one of the smaller open-weight models available (Z-Image Turbo is one of the few that is smaller at 6b) when compared to Ideogram, Krea2, Flux2, etc.

    • It supports native transparency (Qwen's team, as far as I know, is the only one attempting to tackle this). Even though it's relatively trivial to set up background removal postprocessors, it's also neat to see it natively supported.

    • It's fast using QwenImage2.1 convrot, a 1MP image took around ~5 seconds on an RTX4090.

    Negatives

    • The license (assuming you respect it) is far more restrictive. The original Qwen Image 1 was released under the standard Apache license; this one explicitly forbids commercial usage without obtaining a separate license. On the other hand, a lot of us didn't expect the Qwen team to ever release "weights-available" ever again.

    Qwen-Image 1.0, released about a year ago, only scored 4/15 on my GenAI Showdown Benchmarks. Since that time, they've been upstaged by Krea 2 (6/15) and Ideogram4 (8/15). I'll post the new results once I have some more time to run them.

    https://genai-showdown.specr.net

  2. jjcm

    I run a prompt-to-ui design site that uses image models for the design process[1]. The text rendering especially makes this model deeply interesting to me, despite the license. Here are some tests using my harness comparing the outputs of gpt-image-2 and qwen 2.1:

    https://html.non.io/qwen-comparison/

    The text rendering definitely is much, much better than anything else on the open weights market right now. Small text fidelity is quite good. It seems like the text encoder however gets a little bit overloaded with larger prompts - note the presence of hex codes in the design output, those were inputs from the expanded prompt.

    I'll be trying a post-training run on this for web design, it has some serious potential.

    [1] diffui.ai

  3. jfoster

    A lot of the previous Qwen models seem to have used Apache licenses, among others:

    https://en.wikipedia.org/wiki/Qwen#List_of_models

    Unfortunately, it looks like this model is using a much more restrictive license:

    https://github.com/QwenLM/Qwen-Image-2.1/blob/main/LICENSE

  4. fishfasell

    The capabilities of local LLM text-to-image is honestly pretty damn impressive. IMO, I think local image generation is currently ahead of local code generation. I can get an image in seconds locally with the quality being way higher than what I'd expect from a local model. However with coding it's much slower and much less impressive. I'm sure there's a reason for this and I'm not an AI expert so I'll let the smarter folks tell me why, but that's just been my observation thus far.

  5. mdp2021

    How do you use this model locally, similarly to using `llama-server -m <model>`?

    (I mean: outside direct or substantial use of Python, and running the Neural Network in the most efficient way.)

  6. trentor

    They finally fixed their VAE. It really held back their models over the last 2 years.

    EDIT: It still produces artifacts it's better but unusable for production work. In midvalues you will see a slight dot pattern.

  7. trains39472

    A 7B diffusion model can now render CJK text better than Microsoft Windows.

  8. gunalx

    Its happy to see a new open image model from qwen. But the license is a let down. And it dosent even beat their closed qwen3 image wich is already a bit old.

More from this day

2026-09-20