‹ BackHN Continuity

Thread

FLUX 3 Image

437 points · 96 comments · minimaxir

  1. vunderba · · focus · HN ↗
    One of the things they seem to be emphasizing here is the UX around being able to place specific elements where you want them in an image. If the positions of the components in the overall composition are very important, this seems to make that a lot easier and kind of reminds me of InvokeAI.

    Ideogram V4, an open-weight model released back in June can also do this [1], but you have to use a relatively cumbersome JSON structure to describe all the different bounding boxes. So it’s definitely a bit of a hassle.

    I'll probably be waiting until it goes open-weight (hopefully soon) like they did with Flux.2 / Klein.

    [1] - <a href="https:&#x2F;&#x2F;docs.ideogram.ai&#x2F;using-ideogram&#x2F;getting-started&#x2F;prompting-guide&#x2F;4.-json-prompting-ideogram-4.0" rel="nofollow">https:&#x2F;&#x2F;docs.ideogram.ai&#x2F;using-ideogram&#x2F;getting-started&#x2F;prom...

    1. kranke155 · · focus · HN ↗
      You just get an LLM to do the bounding box stuff or use the ComfyUI node that provides a GUI for bounding box generation
      1. vunderba · · focus · HN ↗
        Yeah, given how much better Ideogram v4 outputs are when you use the proper structured JSON (background, elements, etc.) I think most users probably have stuck some kind of Qwen&#x2F;Gemma-based LLM between their raw prompt and the CLIP encoder.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.