‹ BackHN Continuity

Thread

FLUX 3 Image

437 points · 96 comments · minimaxir

  1. mdp2021 · · focus · HN ↗
    What are the best benchmarks to compare these kind (image, video etc.) of generative models - difficult to compare?

    That would be to compare e.g. Qwen Image 3.0 with FLUX 3, with Midjourney etc.

    I had seen some attempts - but I do not know well how they try to approach objectivity.

    1. vunderba · · focus · HN ↗
      I wouldn’t say that mine, the GenAI Showdown, is necessarily any more or less objective than any others but I definitely do a lot of manual curation.

      Prompt results are graded based a weighted calculation which includes: adherence to the prompt, image fidelity, and steerability.

      My comparison benchmark also tends to favor prompt adherence, which a lot of others don’t. Most of ones that I've seen tend towards rather simplistic prompts (e.g. "neon-lit city facing a robotic uprising, with high-tech battles, in anime style"), whereas the prompts I've created try to test high specificity.

      I’ve been running them all the way back to SDXL.

      You can compare specific models using the "View All Models" so if you want to see the progression of open-weight models, or model X vs model Y, you can do so.

      Just a heads up - I haven't added Flux 3 as I'm waiting until BFL drops the open-weights version.

      Generative Comparisons:

      <a href="https:&#x2F;&#x2F;genai-showdown.specr.net" rel="nofollow">https:&#x2F;&#x2F;genai-showdown.specr.net

      Editing Comparisons:

      <a href="https:&#x2F;&#x2F;genai-showdown.specr.net&#x2F;image-editing" rel="nofollow">https:&#x2F;&#x2F;genai-showdown.specr.net&#x2F;image-editing

      1. tangotaylor · · focus · HN ↗
        This is great. Looking forward to your Flux 3 results.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.