‹ BackHN Continuity

Thread

Qwen Image 2.1

740 points · 199 comments · jmillikin

  1. mdp2021 · · focus · HN ↗
    How do you use this model locally, similarly to using `llama-server -m <model>`?

    (I mean: outside direct or substantial use of Python, and running the Neural Network in the most efficient way.)

    1. peri-cl · · focus · HN ↗
      I've set it up on my local machine just now, as my first local image diffuser. I can confirm it's very easy.

      I tried stable-diffusion.cpp, following its compile guide here[0], and its Qwen Image-2.1 specific instructions here[1]. It works out of the box. I made a test pelican[2]. It took 3 minutes on a CPU.

      [0] <a href="https:&#x2F;&#x2F;github.com&#x2F;leejet&#x2F;stable-diffusion.cpp&#x2F;blob&#x2F;master&#x2F;docs&#x2F;build.md" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;leejet&#x2F;stable-diffusion.cpp&#x2F;blob&#x2F;master&#x2F;d...

      [1] <a href="https:&#x2F;&#x2F;github.com&#x2F;leejet&#x2F;stable-diffusion.cpp&#x2F;blob&#x2F;master&#x2F;docs&#x2F;qwen_image_2.1.md" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;leejet&#x2F;stable-diffusion.cpp&#x2F;blob&#x2F;master&#x2F;d...

      [2] <a href="https:&#x2F;&#x2F;i.ibb.co&#x2F;yMknC2K&#x2F;output.png" rel="nofollow">https:&#x2F;&#x2F;i.ibb.co&#x2F;yMknC2K&#x2F;output.png

      1. mdp2021 · · focus · HN ↗
        Thank you! Can you please check how much RAM does it consume (and require)?
        1. peri-cl · · focus · HN ↗
          This is what the runtime reports, at Q8:

              total params memory size = 15645.19MB (VRAM 15645.19MB, RAM 0.00MB):
              text_encoders 7669.77MB(VRAM),
              diffusion_model 7331.05MB(VRAM),
              vae 644.38MB(VRAM),
              controlnet 0.00MB(N&#x2F;A),
              extensions 0.00MB(N&#x2F;A)
          1. mdp2021 · · focus · HN ↗
            That suggests that 16GB RAM will not be enough.

            In fact, like it appears in the reports above, it is &quot;7b&quot; as in

            &gt; 7B parameters in its visual generation component

            It seems they calibrated the size to fill a 16GB VRAM near the limit. RAM requirements will vary.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.