‹ BackHN Continuity

Thread

Qwen Image 2.1

740 points · 199 comments · jmillikin

  1. vunderba · · focus · HN ↗
    So thoughts

    Positives

    • It's a heck of a lot smaller than Qwen-Image 1 (20b parameters) at only 7b, making it one of the smaller open-weight models available (Z-Image Turbo is one of the few that is smaller at 6b) when compared to Ideogram, Krea2, Flux2, etc.

    • It supports native transparency (Qwen's team, as far as I know, is the only one attempting to tackle this). Even though it's relatively trivial to set up background removal postprocessors, it's also neat to see it natively supported.

    • It's fast using QwenImage2.1 convrot, a 1MP image took around ~5 seconds on an RTX4090.

    Negatives

    • The license (assuming you respect it) is far more restrictive. The original Qwen Image 1 was released under the standard Apache license; this one explicitly forbids commercial usage without obtaining a separate license. On the other hand, a lot of us didn't expect the Qwen team to ever release "weights-available" ever again.

    Qwen-Image 1.0, released about a year ago, only scored 4/15 on my GenAI Showdown Benchmarks. Since that time, they've been upstaged by Krea 2 (6/15) and Ideogram4 (8/15). I'll post the new results once I have some more time to run them.

    <a href="https:&#x2F;&#x2F;genai-showdown.specr.net" rel="nofollow">https:&#x2F;&#x2F;genai-showdown.specr.net

    1. vunderba · · focus · HN ↗
      Well, the results are in, at least for text-to-image (the editing bench will come later).

      Qwen-Image 2.1 is definitely a pretty big leap over the last open-weight version, Qwen-Image 1.0, released back in August of last year and managed to score 7 out of 15 as opposed to its predecessor which scored 4 out of 15.

      Even though it&#x27;s significantly smaller, 7b vs 20b, it&#x27;s multimodal (so you don&#x27;t need a separate image-to-image model like you did with Qwen-Edit), more coherent, and significantly faster even when outputting at higher 2K resolutions. However, in my testing, I found that I had to play with dialing up the CFG depending on the complexity of the prompt.

      I&#x27;ve also added a progress dropdown under Model Performance so you can see how cloud vs. local models have been trending since 2024. Spoiler: June of this year released some of the biggest bangers (Krea 2, Ideogram 4, and the kind of slept-on Boogu-Image 0.1).

      Downsides:

      - It was clearly trained on at least some level of synthetic training data, and it shows in some of the subpar outputs in terms of fidelity. Some of this you might be able to iron out with a refiner model downstream or a custom LoRA but time will tell.

      - They&#x27;ve moved away from the permissive Apache license. Commercial usage is only allowed by request.

      Comparisons:

      <a href="https:&#x2F;&#x2F;genai-showdown.specr.net" rel="nofollow">https:&#x2F;&#x2F;genai-showdown.specr.net

      If you just want to compare local models only:

      <a href="http:&#x2F;&#x2F;genai-showdown.specr.net&#x2F;?models=local" rel="nofollow">http:&#x2F;&#x2F;genai-showdown.specr.net&#x2F;?models=local

      1. SV_BubbleTime · · focus · HN ↗
        &gt; and the kind of slept-on Boogu-Image 0.1

        Not slept on at all. It was absolute trash, and I’m super curious why people pretend otherwise. There isn’t a single thing that model did better than any temporal peer.

      2. chr15m · · focus · HN ↗
        What&#x27;s to stop somebody using the output at scale on local hardware to distill their own model and then making that available open weights?
        1. gspr · · focus · HN ↗
          This is the fundamental problem with where the AI race is heading, IMHO. Broadly, there are two possible legal interpretations (to my layman&#x27;s mind):

          * A model is derived work of its training data. This seems sane to me. Open, but copyrighted, works (like FOSS) remain protected from abuse. There&#x27;s some legal moat around AI models. But on the other hand it seems unlikely that there&#x27;s enough liberally licensed (or public domain) training data to go around. The little guy&#x27;s status quo remains, the frontier labs&#x27; work slows down massively.

          * A model is not derived work of its training data. This seems to me insane, but a lot of the world seems to hold this view (including the frontier labs). Stuff like FOSS or indie art is under huge threat of copyrightwashing. But on the other hand, there&#x27;s also zero legal moat around the models. The little guy is eviscerated, but so are the frontier labs.

          Neither interpretation seems, to me, to be capable of sustaining the last couple of years&#x27; developments. But what do I know.

          1. skykooler · · focus · HN ↗
            If a model is derived work of its training data, surely all the existing frontier models that have been training on copyrighted work have a big legal issue, and therefore so does anything produced with them?
            1. gspr · · focus · HN ↗
              Yes. Of course. They&#x27;d have to be rebuilt with acceptably licensed training data. And since there might not be a enough of it, the model owners are screwed.

              My point is that they&#x27;re also screwed in the opposite scenario, because they rely on the same legal protection (against deriving works) as the works they trained on.

              That&#x27;s why I don&#x27;t understand how any of this can be sustained.

          2. dale_glass · · focus · HN ↗
            &gt; A model is not derived work of its training data. This seems to me insane, but a lot of the world seems to hold this view

            Why insane? Models don&#x27;t take the content as-is, they take measurements. I don&#x27;t owe you royalties just because I used your photo to get the proportions and coloring of a duck right. Go watch artist streams, you&#x27;ll often see people to go Google Images for references. I&#x27;ve never seen that result in credit or payments.

            The alternative is that we hand out lots of money to a few large companies specializing in content archives, and there&#x27;s really no benefit to anyone else anyway. On the long term I would expect a few fat cats to get fatter, the small guy to get nothing, and AI still work but get there slowly. I don&#x27;t see the point or the benefit.

            1. gspr · · focus · HN ↗
              &gt; Models don&#x27;t take the content as-is, they take measurements.

              At some point, enough measurements constitute a copy. If I redistribute the average value of all the pixels in your photo, I&#x27;m obviously not in violation of your copyright. If I measure and redistribute 90% of its DCT coefficients (i.e. make a slightly compressed JPEG), I am.

              The interesting stuff happens between those extremes. We cannot just take as a given that all LLMs always are on the safe side. It is not at all obvious.

            2. jml78 · · focus · HN ↗
              Yeah but here is the thing. Say you train a Lora for a model. You then merge those weights into the open weight model.

              Now prompt it for an original image, it will pretty much be able to reproduce that exact image.

              You can say it is just measurements but at some point, it can just reproduce with high enough accuracy to just be seen as a copy

              Given that corporations buy up any valuable IP, I personally think the answer is to abolish copyright because right now it is really only protecting the rich and corporations . Individuals have the illusion of protection but if Disney steals your shit, good luck with the pain and suffering you experience trying to win a court case against them

              1. boplicity · · focus · HN ↗
                Abolishing copyright would make it much harder for smaller players to protect their interests. Their original work would be gobbled up by the big players, and made easily available under the umbrella of a large corporations pre-existing market share. Not a fun situation at all to be in. Especially for the little guys.
            3. selicos · · focus · HN ↗
              &gt; I don&#x27;t owe you royalties just because I used your photo to get the proportions and coloring of a duck right.

              No but a tribute or citation would be nice, especially if the (software) license requires it.

              1. popalchemist · · focus · HN ↗
                Such things are beyond what copyright protects. This line of reasoning would only work if there were a radical reimagining of copyright itself.
          3. Kostchei · · focus · HN ↗
            Lemme fix that for you in the style of Accelerando.... This is the fundamentally great thing about where the AI spend is heading, looks like there is no moat and will be a democratization of cognition simply by dint of the way the tech works. Which is awesome. I shed 1x tear for the private investors and large existing monopolists who poured money into this, only because they might not make the same mistake with the next great technology, but given their greed, that&#x27;s unlikely. They also might get government backing to mitigate their loss, which would be horrible and inflationary, but given the state of open-weights, that seems unlikely outside of the USA.
      3. meherabhossain · · focus · HN ↗

        [dead]

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.