1
 
 

cross-posted from: https://lemdro.id/post/36196733

Just finished reading the report on Qwen-Image-2.0 that dropped the other day. This looks like the efficiency breakthrough we've been waiting for.

The "Headline" Stats:

  • Model Size: 7B parameters.
  • Previous Gen: The old Qwen-Image-2512 was a heavy 20B model.
  • Architecture: Unified "Omni" model (handles both generation and editing in the same weights).
  • Resolution: Native 2K (2048x2048).

The 20B to 7B Optimization: This is the most important part for us. The previous 20B model was a pain to run locally without 24GB VRAM. Crushing that performance down to a 7B model means this should theoretically run on:

  • 12GB Cards (3060/4070): Comfortably at FP16 or Q8.
  • 8GB Cards: Likely possible with aggressive quantization (Q4/Q5) once the community gets hold of it.

Beating "Nano Banana" (Gemini 2.5 Flash Image): The technical report explicitly calls out their performance on blind leaderboards (ELO score). They are claiming Qwen-Image-2.0 achieves a higher ELO rating than Gemini 2.5 Flash Image (aka. Nano Banana) in blind human preference testing.

  • Why this matters: Nano Banana is currently regarded as the SOTA for instruction following and complex prompt adherence. If a 7B local model is actually beating it in ELO, that is insane efficiency.

The "Catch": Weights are not open yet. It is currently available via their API and Demo (Qwen Chat). However, Qwen has an excellent track record (Apache 2.0 releases for almost everything eventually). Given that they released the 20B weights previously, it is highly likely we see the 7B weights in a matter of weeks.

TL;DR: They optimized the 20B heavy-hitter down to a consumer-viable 7B, it claims to beat Google's best efficiency model in ELO, and now we wait for the HF upload to see if the quantization holds up.

2
3
4
submitted 2 years ago by to c/stablediffusion@lemmy.ml
5
6
 
 

Without paywall: https://archive.ph/QD9v1

7
 
 

Excerpt from the relevant “ComfyUI dev” Matrix room:

matt3o
and what is it then?

comfyanonymous
"safety training"

matt3o
why does it trigger on certain keywords and it's like it's scrambling the image?

comfyanonymous
the 2B wasn't the one I had been working on so I don't really know the specifics

matt3o
I was even able to trick it by sending certain negatives

comfyanonymous
I was working on a T5 only 4B model which would ironically had been safer without breaking everything
because T5 doesn't know any image data so it was only able to generate images in the distribution of the filtered training data

comfyanonymous
but they canned my 4B and I wasn't really following the 2B that closely

[…]

comfyanonymous
yeah they did something with the weights
the model arch of the 2B was never changed at all

BVH
weights directly?
oh boy, abliteration, the worst kind

comfyanonymous
also they apparently messed up the pretraining on the 2B so it was never supposed to actually be released

[…]

comfyanonymous
yeah the 2B apparently was a bit of a failed experiment by the researchers that left
but there was a strong push by the top of the company to release to 2B instead of the 4B and 8B

Additional excerpt (after the Reddit post) from Stable Diffusion Discord “#sd3”:

comfy
Yes I resigned over 2 weeks ago and Friday was my last day at stability

8
9
submitted 2 years ago by to c/stablediffusion@lemmy.ml
10
11
 
 

Basic ComfyUI workflows without 10k custom nodes or impossible to follow workflows.

12
13
14
15
 
 

Without paywall: https://archive.ph/8QkSl

16
17
18
19
20
21
 
 

I created a custom SDXL Lora using my dataset. I created the dataset using a previous generative art tool I build to visualize factorio blueprints: https://github.com/piebro/factorio-blueprint-visualizer. I like the lora to create interesting patterns.

22
23
24
25
Happy Halloween (programming.dev)
 
 

Using ParchArtXL civitai.com/models/141471/parchartxl

view more: next ›