‹ BackHN Continuity

Thread

Training Text-to-Image Models 3.6× Faster

56 points · 10 comments · schopra909

  1. schopra909 · · focus · HN ↗
    Author here, feel free to drop questions below. Will try to answer to best of my ability!
    1. E-Reverance · · focus · HN ↗
      Slightly off-topic but thoughts on this <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;pdf&#x2F;2410.08159" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;pdf&#x2F;2410.08159
      1. schopra909 · · focus · HN ↗
        Just skimmed the paper (haven&#x27;t seen this before). Might need to read more carefully, but at first glance I don&#x27;t understand their intuition why the &quot;Markovian property limits the model’s ability to fully utilize the generation trajectory&quot;.

        When it comes to multi-resolution training (e.g. matryoshka training), there are precedents that don&#x27;t require this AR formulation.

        They cite MAR (<a href="https:&#x2F;&#x2F;arxiv.org&#x2F;pdf&#x2F;2406.11838" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;pdf&#x2F;2406.11838), which I think is a much clearer articulation of &quot;autoregressive diffusion&quot;. MAR uses an autoregressive base (like an LLM) and staples on a small MLP on-top that&#x27;s trained as a diffusion head. That makes more sense to me, since you can leverage the &quot;knowledge prior&quot; from an LLM and have it generate images. That&#x27;s probably how nano-bannana and GPT-Image broadly work.

        Zooming out, it&#x27;s not clear to me from any of the work WHY autoregressive diffusion on it&#x27;s own is better than regular diffusion. Most autoregressive diffusion models use speculative decoding, because pure autoregressive diffusion is too slow to run at inference time.

        The one ~magical~ thing about auto-regressive diffusion (IMO) has nothing to do with autoregressive vs. fully-bidirectional diffusion. But simply, the face that you can do it on-top of a LLM. That means the LLM can look at it&#x27;s generation, assess it&#x27;s quality, think on what&#x27;s broken, and then call itself to edit the image and fix it. With methods like RLVR, that means you can essentially guarantee the correctness of your image (along the axes you&#x27;ve RLVR-ed).

        1. E-Reverance · · focus · HN ↗
          In my opinion its a great idea because you get reuse hidden states which allows you to do something akin to thinking&#x2F;[online learning]. If you only evolve use input space outputs you&#x27;re dealing with more decoding pressure.

          Also related to what I was saying see these:

          <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;pdf&#x2F;2609.16372v1" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;pdf&#x2F;2609.16372v1

          <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;pdf&#x2F;2609.01449v1" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;pdf&#x2F;2609.01449v1 (this one is quite mindblowing cause they noise they the actual input space every time but it still works)

          <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;pdf&#x2F;2609.11801" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;pdf&#x2F;2609.11801

          1. schopra909 · · focus · HN ↗
            Thanks for the references! I’ll have to take a deeper look, haven’t read these before.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.