Does anyone know if it can be used to generate accurate frame-by-frame sprite sequences? I found that no image model can do this well (with sufficient fidelity) - neither with one shot (full spritesheet), nor single frame conditioning. It would be great if an imagegen model could do this. What I do now (I use my own tool <a href="https://github.com/acatovic/ai-game-studio" rel="nofollow">https://github.com/acatovic/ai-game-studio) is basically generate a reference image, then condition on that image to generate a very short video, then extract and prune frames. Then I get indie-level sprite fidelity about 90% of the time.
The task you're describing is a video model task, not an image model task. It's inherently temporal.
Generate a sprite in an image editor, then use a video model to make the loop you want; then turn the resulting video back into individual sprite images.
Sure and that's what I do, but a video can be seen as a causal generation on discreet sequence of images, each image conditioned on the one before it. It can also be seen as a series of image editing tasks. It would be cool to get this working in imagegen because of the amount of control you would get. Right now with video generation you can at best specify start and end frame and hope for the best.
Image edit models can probably do a grid, but the temporal accuracy / coherence will never match what a video model, which is really a world model, can do.
Regarding your world model statement. This is completely FALSE. Learning the visual statistics of a physical world is NOT the same thing as learning its causal dynamics. The difference is observational likelihood versus intervention-dependent dynamics. There have been great studies disproving video models as world models, like this ICML paper: <a href="https://proceedings.mlr.press/v267/kang25g.html" rel="nofollow">https://proceedings.mlr.press/v267/kang25g.html. Unfortunately lot of people treat them as world models, mostly because of their ability to reproduce increasingly convincing physical behaviour without ever discovering the underlying physical laws. This is due to many things that I could write an essay about, but better conditioning, latent space represtnation, scaling etc, all make them look awesome.
I can still get absolutely insane results with MiniMax H3 - insane in the sense that it would not make sense at all and would make your head spin.
They are proto world models (lots written about this - flux being an example of a video model whose weights also power world-action-engines used in robots) in that they attempt to model causality in time, the thing that is required for what OP is asking for and which image models will never do because it is out of domain.
The idea to use predictive models of future observations for decision making has a long history, including early demonstrations of robot control which combined action-conditioned video prediction with model-predictive control.
> There have been great studies disproving video models as world models, like this ICML paper: <a href="https://proceedings.mlr.press/v267/kang25g.html" rel="nofollow">https://proceedings.mlr.press/v267/kang25g.html.
That paper sets up a task where generating a correct video requires correctly modeling physical laws. From the failure to always generate the correct video, they infer that the model has failed to correctly model the physical laws. The whole premise of the experiment is that learning visual statistics is equivalent to learning causal dynamics, such that failure at one implies failure at the other.
The main difference in applications is that the bar for entertainment is lower, so that even a very bad world model may be acceptable.
An (apparently, as I haven't tried it) ready-to-go ComfyUI graph for the above. In theory you should be able to drop these into ComfyUI and have it work:
I've stumbled upon a reliablish pipeline you create a reference sheet and a single image pose then trellis v2 for the body and unirig for tigging, then you generate the poses you need in a sheet and send the pose sheet plus the reference images, and use these as 2d anymation cycles and you compose on top of the scene with lanes, this focus all attention of the model to fidelity instead of background integration
I’m looking for the same thing, on one side making sprites should be easier because of the lower complexity of pixel art; on the other hand making something with a specific style or with sprite frame-by-frame coherence seems harder.
Imagine online procedural MMO with old gen final fantasy / chrono trigger styles :)
armcat · · focus · HN ↗
popalchemist · · focus · HN ↗
Generate a sprite in an image editor, then use a video model to make the loop you want; then turn the resulting video back into individual sprite images.
armcat · · focus · HN ↗
popalchemist · · focus · HN ↗
armcat · · focus · HN ↗
I can still get absolutely insane results with MiniMax H3 - insane in the sense that it would not make sense at all and would make your head spin.
popalchemist · · focus · HN ↗
bobcatsmith · · focus · HN ↗
[dead]
popalchemist · · focus · HN ↗
Video Generation Models as World Models: Efficient Paradigms, Architectures and Algorithms
<a href="https://bfl.ai/models/flux-3-action" rel="nofollow">https://bfl.ai/models/flux-3-action
The idea to use predictive models of future observations for decision making has a long history, including early demonstrations of robot control which combined action-conditioned video prediction with model-predictive control.
Clearly I'm the uneducated one here eh
yorwba · · focus · HN ↗
That paper sets up a task where generating a correct video requires correctly modeling physical laws. From the failure to always generate the correct video, they infer that the model has failed to correctly model the physical laws. The whole premise of the experiment is that learning visual statistics is equivalent to learning causal dynamics, such that failure at one implies failure at the other.
The main difference in applications is that the bar for entertainment is lower, so that even a very bad world model may be acceptable.
nicerice · · focus · HN ↗
<a href="https://media.discordapp.net/attachments/1401891025970008154/1552060356220686336/2511_vs_21.mp4?ex=6ac16b58&is=6ac019d8&hm=bb259d2b5e1ce5169721568ddff15ba227685d639660b21c1a111b3afdcf0794" rel="nofollow">https://media.discordapp.net/attachments/1401891025970008154...
I don't know what went into making it, but their twitter is @araminta_k if you're curious
justinclift · · focus · HN ↗
* <a href="https://alvdansen.github.io/animating-on-twos/" rel="nofollow">https://alvdansen.github.io/animating-on-twos/
* <a href="https://github.com/alvdansen/animating-on-twos" rel="nofollow">https://github.com/alvdansen/animating-on-twos
* <a href="https://huggingface.co/alvdansen/h3-keyframe-animation" rel="nofollow">https://huggingface.co/alvdansen/h3-keyframe-animation
## Quick Start
An (apparently, as I haven't tried it) ready-to-go ComfyUI graph for the above. In theory you should be able to drop these into ComfyUI and have it work:
<a href="https://huggingface.co/alvdansen/h3-keyframe-animation#quick-start" rel="nofollow">https://huggingface.co/alvdansen/h3-keyframe-animation#quick...
[deleted] · · focus · HN ↗
[deleted]
bobcatsmith · · focus · HN ↗
avereveard · · focus · HN ↗
Lucasoato · · focus · HN ↗
Imagine online procedural MMO with old gen final fantasy / chrono trigger styles :)