‹ BackHN Continuity

Thread

FLUX 3 Image

437 points · 96 comments · minimaxir

  1. vunderba · · focus · HN ↗
    One of the things they seem to be emphasizing here is the UX around being able to place specific elements where you want them in an image. If the positions of the components in the overall composition are very important, this seems to make that a lot easier and kind of reminds me of InvokeAI.

    Ideogram V4, an open-weight model released back in June can also do this [1], but you have to use a relatively cumbersome JSON structure to describe all the different bounding boxes. So it’s definitely a bit of a hassle.

    I'll probably be waiting until it goes open-weight (hopefully soon) like they did with Flux.2 / Klein.

    [1] - <a href="https:&#x2F;&#x2F;docs.ideogram.ai&#x2F;using-ideogram&#x2F;getting-started&#x2F;prompting-guide&#x2F;4.-json-prompting-ideogram-4.0" rel="nofollow">https:&#x2F;&#x2F;docs.ideogram.ai&#x2F;using-ideogram&#x2F;getting-started&#x2F;prom...

    1. kranke155 · · focus · HN ↗
      You just get an LLM to do the bounding box stuff or use the ComfyUI node that provides a GUI for bounding box generation
      1. CuriouslyC · · focus · HN ↗
        I always hated Comfy&#x27;s node based UI, but agents make it tolerable. Now I just have them set up a workflow and I go in and tweak it manually if the results aren&#x27;t where I want them. I even have agents cherry doing multiple runs and cherry picking the best outputs, models have gotten good enough that it&#x27;s a real time saver, assuming you have references they can and a rubric to check against.
        1. hdjrudni · · focus · HN ↗
          How do you get agents to set up a workflow? You just get them to modify the JSON directly and then import it, or do you have a tighter integration (e.g. in the UI)?
          1. CuriouslyC · · focus · HN ↗
            The agents can interact with Comfy via API pretty well, which afaik ends up being directly with JSON.
        2. swiftcoder · · focus · HN ↗
          &gt; I always hated Comfy&#x27;s node based UI

          It&#x27;s one of the most uniquely hostile user experiences I&#x27;ve ever had the (dis)pleasure of working with

          1. bavell · · focus · HN ↗
            Awhile back I built a project which reads comfy&#x27;s API and builds a typescript sdk from it. Much nicer to work with and still benefits from node caching and the comfy ecosystem. Perhaps I should clean it up and open source it, though I haven&#x27;t looked around and there may already be other projects doing this out there.
            1. user43928 · · focus · HN ↗
              Using AI for the ComfyUI API or workflows worked pretty well for me even early this year.

              However, today I don&#x27;t see a reason to use ComfyUI at all.

              For Qwen Image 2.1, I had Opus 5.5 create a backend outside of ComfyUI and it was able to make generation take 20% less time with some optimizations.

              The optimizations it implemented were caching the text computation in Qwen Image 2.1 rather than including it in every step, fusing projections into a larger matrix multiplication, and decoding the VAE in horizontal bands or something like that.

              If there was anything interesting in ComfyUI nodes, I imagine I could just have the AI adopt the relevant code instead of dealing with ComfyUI or custom nodes.

          2. orbital-decay · · focus · HN ↗
            I remember precisely how ComfyUI started, and the motivation was something along the lines of &quot;look at the cool guys working in Houdini and Resolve, they must know something&quot; (I wish I was kidding). Of course nobody thought that DAGs are just poorly suited for the task, even when plugin authors started trying to slap loops on top of them shortly after that.
        3. jarjoura · · focus · HN ↗
          It&#x27;s definitely not for me.

          From where I&#x27;m sitting, it&#x27;s just turning python functions into boxes and instead of write the function yourself, you drag from the output of one box to the input of another. For 2 or 3 boxes, this is cool, but I opened up a professional workflow and was taken into a view with 100s of boxes and wires all over the place. Uhh, ok?

          For myself, I&#x27;d rather just create my own python environment, write some quick pytorch or mlx calls, wire up some cli to it and share that in a GitHub.

          1. sorenjan · · focus · HN ↗
            I recently found a project[0] that takes a ComfyUI workflow and turns it into a simple UI. I have no affiliation with it and haven&#x27;t tested it myself, but it looks like it might be handy once you have a finished workflow you want to use.

            [0] <a href="https:&#x2F;&#x2F;github.com&#x2F;saintbrodie&#x2F;Orange" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;saintbrodie&#x2F;Orange

            1. sheepscreek · · focus · HN ↗
              ComfyUI has a feature called “App” which is basically this. You pick the inputs and outputs in a workflow and then they become a kind of a subgraph.

              Disclaimer: never used it for actual generation so I don’t know if it’s doing anything special other than being a subgraph with a different name.

        4. rf15 · · focus · HN ↗
          As someone with general experience with some generative AI tools I started exploring Comfy last week and it was the easiest to pick up by far. This is not meant in a troll way, it just reflects how you do it programmatically the most, which I&#x27;ve seen the most before. The other UIs are always opinionated layouts on top of the actual logic where you constantly have to look up&#x2F;sift through menus for what you want to do. In Comfy, you just use the searchbar and get the right node.

          That being said: defining composition made me immediately think that someone probably made a gui like this with easy to move bounding boxes, and I&#x27;m happy they did.

      2. vunderba · · focus · HN ↗
        Yeah, given how much better Ideogram v4 outputs are when you use the proper structured JSON (background, elements, etc.) I think most users probably have stuck some kind of Qwen&#x2F;Gemma-based LLM between their raw prompt and the CLIP encoder.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.