‹ BackHN Continuity

Thread

Qwen Image 2.1

740 points · 199 comments · jmillikin

  1. fishfasell · · focus · HN ↗
    The capabilities of local LLM text-to-image is honestly pretty damn impressive. IMO, I think local image generation is currently ahead of local code generation. I can get an image in seconds locally with the quality being way higher than what I'd expect from a local model. However with coding it's much slower and much less impressive. I'm sure there's a reason for this and I'm not an AI expert so I'll let the smarter folks tell me why, but that's just been my observation thus far.
    1. tarcon · · focus · HN ↗
      I think there was a lot more brainpower invested in the media generation side of things. The noise-based diffusion technique is further developed. It had a discovery of applying a physics-based understanding of Brownian motion to guide it. Image generation has comparatively simple training process - this is an image with dog, and without dog (contrastive learning).

      Might be worth to watch the diffusion based LLMs.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.