‹ BackHN Continuity

Thread

Qwen Image 2.1

740 points · 199 comments · jmillikin

  1. jjcm · · focus · HN ↗
    I run a prompt-to-ui design site that uses image models for the design process[1]. The text rendering especially makes this model deeply interesting to me, despite the license. Here are some tests using my harness comparing the outputs of gpt-image-2 and qwen 2.1:

    <a href="https:&#x2F;&#x2F;html.non.io&#x2F;qwen-comparison&#x2F;" rel="nofollow">https:&#x2F;&#x2F;html.non.io&#x2F;qwen-comparison&#x2F;

    The text rendering definitely is much, much better than anything else on the open weights market right now. Small text fidelity is quite good. It seems like the text encoder however gets a little bit overloaded with larger prompts - note the presence of hex codes in the design output, those were inputs from the expanded prompt.

    I&#x27;ll be trying a post-training run on this for web design, it has some serious potential.

    [1] diffui.ai

    1. cloudking · · focus · HN ↗
      Those simple prompts produce nearly the exact same layout in the 2 different models?
      1. jjcm · · focus · HN ↗
        My harness expands the prompt into a json representation that specifies layout much more rigorously, which is why you see such that amount of alignment between the two.

        That internal json backing helps significantly when you want to maintain consistent design system components&#x2F;patterns across multiple pages. The aligned layout is it working as intended.

        1. j-bos · · focus · HN ↗
          Private harness?
          1. jjcm · · focus · HN ↗
            It is, yes. This is for diffui.ai, which for now is closed source.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.