‹ BackHN Continuity

Thread

FLUX 3 Image

437 points · 96 comments · minimaxir

  1. armcat · · focus · HN ↗
    Does anyone know if it can be used to generate accurate frame-by-frame sprite sequences? I found that no image model can do this well (with sufficient fidelity) - neither with one shot (full spritesheet), nor single frame conditioning. It would be great if an imagegen model could do this. What I do now (I use my own tool <a href="https:&#x2F;&#x2F;github.com&#x2F;acatovic&#x2F;ai-game-studio" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;acatovic&#x2F;ai-game-studio) is basically generate a reference image, then condition on that image to generate a very short video, then extract and prune frames. Then I get indie-level sprite fidelity about 90% of the time.
    1. popalchemist · · focus · HN ↗
      The task you&#x27;re describing is a video model task, not an image model task. It&#x27;s inherently temporal.

      Generate a sprite in an image editor, then use a video model to make the loop you want; then turn the resulting video back into individual sprite images.

      1. armcat · · focus · HN ↗
        Sure and that&#x27;s what I do, but a video can be seen as a causal generation on discreet sequence of images, each image conditioned on the one before it. It can also be seen as a series of image editing tasks. It would be cool to get this working in imagegen because of the amount of control you would get. Right now with video generation you can at best specify start and end frame and hope for the best.
        1. popalchemist · · focus · HN ↗
          Image edit models can probably do a grid, but the temporal accuracy &#x2F; coherence will never match what a video model, which is really a world model, can do.
          1. armcat · · focus · HN ↗
            Regarding your world model statement. This is completely FALSE. Learning the visual statistics of a physical world is NOT the same thing as learning its causal dynamics. The difference is observational likelihood versus intervention-dependent dynamics. There have been great studies disproving video models as world models, like this ICML paper: <a href="https:&#x2F;&#x2F;proceedings.mlr.press&#x2F;v267&#x2F;kang25g.html" rel="nofollow">https:&#x2F;&#x2F;proceedings.mlr.press&#x2F;v267&#x2F;kang25g.html. Unfortunately lot of people treat them as world models, mostly because of their ability to reproduce increasingly convincing physical behaviour without ever discovering the underlying physical laws. This is due to many things that I could write an essay about, but better conditioning, latent space represtnation, scaling etc, all make them look awesome.

            I can still get absolutely insane results with MiniMax H3 - insane in the sense that it would not make sense at all and would make your head spin.

            1. popalchemist · · focus · HN ↗
              They are proto world models (lots written about this - flux being an example of a video model whose weights also power world-action-engines used in robots) in that they attempt to model causality in time, the thing that is required for what OP is asking for and which image models will never do because it is out of domain.
              1. bobcatsmith · · focus · HN ↗

                [dead]

                1. popalchemist · · focus · HN ↗
                  <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;html&#x2F;2603.28489v1" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;html&#x2F;2603.28489v1

                  Video Generation Models as World Models: Efficient Paradigms, Architectures and Algorithms

                  <a href="https:&#x2F;&#x2F;bfl.ai&#x2F;models&#x2F;flux-3-action" rel="nofollow">https:&#x2F;&#x2F;bfl.ai&#x2F;models&#x2F;flux-3-action

                  The idea to use predictive models of future observations for decision making has a long history, including early demonstrations of robot control which combined action-conditioned video prediction with model-predictive control.

                  Clearly I&#x27;m the uneducated one here eh

            2. yorwba · · focus · HN ↗
              &gt; There have been great studies disproving video models as world models, like this ICML paper: <a href="https:&#x2F;&#x2F;proceedings.mlr.press&#x2F;v267&#x2F;kang25g.html" rel="nofollow">https:&#x2F;&#x2F;proceedings.mlr.press&#x2F;v267&#x2F;kang25g.html.

              That paper sets up a task where generating a correct video requires correctly modeling physical laws. From the failure to always generate the correct video, they infer that the model has failed to correctly model the physical laws. The whole premise of the experiment is that learning visual statistics is equivalent to learning causal dynamics, such that failure at one implies failure at the other.

              The main difference in applications is that the bar for entertainment is lower, so that even a very bad world model may be acceptable.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.