‹ BackHN Continuity

Thread

Qwen Image 2.1

740 points · 199 comments · jmillikin

  1. vunderba · · focus · HN ↗
    So thoughts

    Positives

    • It's a heck of a lot smaller than Qwen-Image 1 (20b parameters) at only 7b, making it one of the smaller open-weight models available (Z-Image Turbo is one of the few that is smaller at 6b) when compared to Ideogram, Krea2, Flux2, etc.

    • It supports native transparency (Qwen's team, as far as I know, is the only one attempting to tackle this). Even though it's relatively trivial to set up background removal postprocessors, it's also neat to see it natively supported.

    • It's fast using QwenImage2.1 convrot, a 1MP image took around ~5 seconds on an RTX4090.

    Negatives

    • The license (assuming you respect it) is far more restrictive. The original Qwen Image 1 was released under the standard Apache license; this one explicitly forbids commercial usage without obtaining a separate license. On the other hand, a lot of us didn't expect the Qwen team to ever release "weights-available" ever again.

    Qwen-Image 1.0, released about a year ago, only scored 4/15 on my GenAI Showdown Benchmarks. Since that time, they've been upstaged by Krea 2 (6/15) and Ideogram4 (8/15). I'll post the new results once I have some more time to run them.

    <a href="https:&#x2F;&#x2F;genai-showdown.specr.net" rel="nofollow">https:&#x2F;&#x2F;genai-showdown.specr.net

    1. Epitaque · · focus · HN ↗
      That benchmark might have some issues. You prompted the models to generate an image of striking a ring against a crucible. Then you (presumably, manually?) scored the images that depicted an anvil higher than the ones striking something resembling a crucible.
      1. vunderba · · focus · HN ↗
        That’s a good catch. Yes, all scoring is done through manual review since relying on a VL model for these kinds of meta-metrics is a sort of loose equivalent of gödel&#x27;s second incompleteness theorem.

        I’ll have to think about this one. When I crafted the prompt, I wasn’t really thinking about the differences between a crucible and an anvil. It was more the visual of an archangel smelting halos for newly arrived heavenly beings.

        1. speerer · · focus · HN ↗
          I&#x27;m not sure why one would even strike metal against a crucible! It&#x27;s a container for liquid metal. One of the outputs shows it being smashed by the manoeuvre, which is probably the most realistic outcome of all of them.

          Sorry, I&#x27;m not trying to nitpick. I&#x27;m just joining in because I&#x27;m interested in how the models dealt with the request.

          1. vunderba · · focus · HN ↗
            Well this is HN - original home of the &quot;ummm actually...&quot; - so I appreciate when people pick all the nits. :)

            Even though I prompted for a crucible in the prompt, I think the fact that the prompt also contained terms like “blacksmith” and “hammer,” caused it to lean towards anvils over crucibles in some of the pictures (which as you brought up makes more sense anyway).

            1. seemaze · · focus · HN ↗
              pedants unite!
            2. apothegm · · focus · HN ↗
              Perhaps for the particular image you liked. But choosing something “typical” or average over what it was instructed to do is a massively common failure mode for AI. One that makes the difference between a useful model and one that makes you want to throw your laptop out a window.
              1. vunderba · · focus · HN ↗
                I agree. In fact the entire reason I initially built GenAI Showdown was because many of the comparative tests on places like Image Arena Leaderboard [1] are not designed to challenge models on prompt adherence. Even when they are, a considerable number of amateur judges tend to prioritize aesthetics over adherence or instruction-following.

                I&#x27;ll likely be redoing that particular bench with added minimum passing criteria of an anvil.

                [1] - <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;spaces&#x2F;ArtificialAnalysis&#x2F;Text-to-Image-Leaderboard" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;spaces&#x2F;ArtificialAnalysis&#x2F;Text-to-Ima...

                1. RugnirViking · · focus · HN ↗
                  ive been taking a look there. I think for some of the image to image ones you really ought to use real images as the base prompt, there are a lot of weird things happening where the base image has issues and then its hard to say if the model should be correcting the flaws or not. Things like &quot;childrens drawing app where each crayon is clearly sized to be tapped on&quot; but the base image has the rainbow cut off partway across with only half the red and black crayons visible.

                  Also &quot;vintage&quot; photography being corrected, where the original is clearly ai generated with unrealistic sharp focus everywhere

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.