‹ BackHN Continuity

Thread

Qwen Image 2.1

740 points · 199 comments · jmillikin

  1. mdp2021 · · focus · HN ↗
    How do you use this model locally, similarly to using `llama-server -m <model>`?

    (I mean: outside direct or substantial use of Python, and running the Neural Network in the most efficient way.)

    1. embedding-shape · · focus · HN ↗
      Probably ComfyUI is one of the easiest way to get started with local image/video models. Or perhaps vLLM, if they have support for it already, would be something like `vllm serve <model> --omni --port 9080`
    2. utopiah · · focus · HN ↗
      why not just as you suggested i.e. <a href="https:&#x2F;&#x2F;qwen.readthedocs.io&#x2F;en&#x2F;latest&#x2F;run_locally&#x2F;llama.cpp.html#llama-server" rel="nofollow">https:&#x2F;&#x2F;qwen.readthedocs.io&#x2F;en&#x2F;latest&#x2F;run_locally&#x2F;llama.cpp.... then get the result either via a UI or wget&#x2F;curl it back?
      1. mdp2021 · · focus · HN ↗
        I am not sure that llama.cpp also supports image generation models.
        1. utopiah · · focus · HN ↗
          it&#x27;s multimodal, see <a href="https:&#x2F;&#x2F;github.com&#x2F;ggml-org&#x2F;llama.cpp&#x2F;blob&#x2F;master&#x2F;docs&#x2F;multimodal.md" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;ggml-org&#x2F;llama.cpp&#x2F;blob&#x2F;master&#x2F;docs&#x2F;multi...
          1. exe34 · · focus · HN ↗
            Multimodal doesn&#x27;t guarantee input and output.

            &gt; Currently, we support image, audio and video input.

            1. utopiah · · focus · HN ↗
              Seems I&#x27;m missing something. Does this model support other inputs?

              Image outputs are supported, videos I&#x27;m not sure but I don&#x27;t think that&#x27;s an output, just a preview of the equirectangular example, so, same question here, what does this model outputs that isn&#x27;t supported?

              1. [deleted] · · focus · HN ↗

                [deleted]

              2. exe34 · · focus · HN ↗
                It&#x27;s a diffusion model, completely different from autoregressive attention models.
                1. utopiah · · focus · HN ↗
                  like <a href="https:&#x2F;&#x2F;github.com&#x2F;ggml-org&#x2F;llama.cpp&#x2F;tree&#x2F;master&#x2F;examples&#x2F;diffusion" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;ggml-org&#x2F;llama.cpp&#x2F;tree&#x2F;master&#x2F;examples&#x2F;d... ?

                  At the risk of stating the obvious llama.cpp isn&#x27;t just about LLaMa as <a href="https:&#x2F;&#x2F;github.com&#x2F;ggml-org&#x2F;llama.cpp&#x2F;blob&#x2F;master&#x2F;src&#x2F;llama-arch.h#L13" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;ggml-org&#x2F;llama.cpp&#x2F;blob&#x2F;master&#x2F;src&#x2F;llama-... someone else pointed out.

                  1. exe34 · · focus · HN ↗
                    Aha I was wrong. Thanks for sharing that!
                    1. utopiah · · focus · HN ↗
                      no worries, I was wrong too, it is multimodal but only for inputs apparently, so for now there seem to only be text as output but no image as output
              3. mdp2021 · · focus · HN ↗
                Not all architectures are supported by llama.cpp . The GGUF format encodes the NN in a standardized way, but then you need code that can use that NN structure.

                I understand that llama.cpp could only output text, last time I checked (I do not know how to find a good source for that though).

                See <a href="https:&#x2F;&#x2F;github.com&#x2F;ggml-org&#x2F;llama.cpp&#x2F;blob&#x2F;master&#x2F;src&#x2F;llama-arch.h" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;ggml-org&#x2F;llama.cpp&#x2F;blob&#x2F;master&#x2F;src&#x2F;llama-... , the

                  enum llm_arch {
                
                ...
                1. utopiah · · focus · HN ↗
                  I haven&#x27;t used llama.cpp for image generation either but I recall an issue about it. Unfortunately I can&#x27;t pinpoint it now and there is the older closed issue <a href="https:&#x2F;&#x2F;github.com&#x2F;ggml-org&#x2F;llama.cpp&#x2F;issues&#x2F;4408" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;ggml-org&#x2F;llama.cpp&#x2F;issues&#x2F;4408 so unless mtmd supports also multimodal outputs out of the box safe to assume output is still limited to text.
    3. fp64 · · focus · HN ↗
      on the linked GitHub page they list support Diffusers, ComfyUI, vLLM-Omni, SGLang, and LightX2V with links to each
      1. mdp2021 · · focus · HN ↗
        &gt; Diffusers, ComfyUI, vLLM-Omni, SGLang, and LightX2V

        I think that&#x27;s all Python (not a direct executable).

        You could just do (see the &quot;Quick Start&quot;) four `pip install` and have a dozen lines script to generate the image. But `llama.cpp` and similar do not require e.g. installing Torch (or PyTorch) - you can use `llama.cpp` on a non-specialized machine.

        1. wgd · · focus · HN ↗
          &quot;just&quot;

          I don&#x27;t think I have ever once run &quot;pip install transformers&quot; and had it work without three rounds of fiddling

          1. mdp2021 · · focus · HN ↗
            Or trying to install the whole of CUDA on machines that do not even have a GPU (not Nvidia, not anything past the embedded)...

            Yep, that&#x27;s (also) what I meant ;)

            Lean, efficient... Also sensible and trouble-less.

            1. iugtmkbdfil834 · · focus · HN ↗
              The &#x27;not nvidia&#x27; is no longer a deal breaker by itself.
    4. Iolaum · · focus · HN ↗
      There is difussion.cpp which is intended for those types of models. I set up krea-2-turbo with the help of ChatGPT 2 months ago, if you have a capable computer that&#x27;s what I would suggest once it becomes supported.
    5. rwmj · · focus · HN ↗
      Additional question is what kind of local hardware would be required for this? 7B parameters sounds very light weight, but I&#x27;m not sure. (Edit: The download is 33 GB).

      Edit x2: As usual I&#x27;m in a twisty maze of pip packages that don&#x27;t work together, with obscure errors about missing modules, even though I followed the instructions on the page to the letter. I really wish people didn&#x27;t use Python for this stuff. A simple C&#x2F;C++ program would be so much better.

      1. peri-cl · · focus · HN ↗
        It&#x27;s about 16 GiB at Q8 quants (combining both the image and language parts). (Meaning, community quantized models from HuggingFace).

        I think it will technically run on anything that has enough memory. I just tried it on a standard laptop (dual-channel DDR5), and it took about 3 minutes for a 512x512. If you&#x27;d want to run it at interactive speeds, you would want a GPU (one which fits this in VRAM).

        &gt; &quot;I really wish people didn&#x27;t use Python for this stuff. A simple C&#x2F;C++ program would be so much better.&quot;

        You mean besides stable-diffusion.cpp ?

        1. rwmj · · focus · HN ↗
          &gt; You mean besides stable-diffusion.cpp ?

          Yes, thanks, I didn&#x27;t know about that. Will try it.

    6. [deleted] · · focus · HN ↗

      [deleted]

    7. nkhgfugjk · · focus · HN ↗
      I am using sd.cpp, which is the cousin of llama.cpp: <a href="https:&#x2F;&#x2F;github.com&#x2F;leejet&#x2F;stable-diffusion.cpp" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;leejet&#x2F;stable-diffusion.cpp

      it already has day-0 qwen image 2.1 support!

    8. peri-cl · · focus · HN ↗
      I&#x27;ve set it up on my local machine just now, as my first local image diffuser. I can confirm it&#x27;s very easy.

      I tried stable-diffusion.cpp, following its compile guide here[0], and its Qwen Image-2.1 specific instructions here[1]. It works out of the box. I made a test pelican[2]. It took 3 minutes on a CPU.

      [0] <a href="https:&#x2F;&#x2F;github.com&#x2F;leejet&#x2F;stable-diffusion.cpp&#x2F;blob&#x2F;master&#x2F;docs&#x2F;build.md" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;leejet&#x2F;stable-diffusion.cpp&#x2F;blob&#x2F;master&#x2F;d...

      [1] <a href="https:&#x2F;&#x2F;github.com&#x2F;leejet&#x2F;stable-diffusion.cpp&#x2F;blob&#x2F;master&#x2F;docs&#x2F;qwen_image_2.1.md" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;leejet&#x2F;stable-diffusion.cpp&#x2F;blob&#x2F;master&#x2F;d...

      [2] <a href="https:&#x2F;&#x2F;i.ibb.co&#x2F;yMknC2K&#x2F;output.png" rel="nofollow">https:&#x2F;&#x2F;i.ibb.co&#x2F;yMknC2K&#x2F;output.png

      1. mdp2021 · · focus · HN ↗
        Thank you! Can you please check how much RAM does it consume (and require)?
        1. peri-cl · · focus · HN ↗
          This is what the runtime reports, at Q8:

              total params memory size = 15645.19MB (VRAM 15645.19MB, RAM 0.00MB):
              text_encoders 7669.77MB(VRAM),
              diffusion_model 7331.05MB(VRAM),
              vae 644.38MB(VRAM),
              controlnet 0.00MB(N&#x2F;A),
              extensions 0.00MB(N&#x2F;A)
          1. mdp2021 · · focus · HN ↗
            That suggests that 16GB RAM will not be enough.

            In fact, like it appears in the reports above, it is &quot;7b&quot; as in

            &gt; 7B parameters in its visual generation component

            It seems they calibrated the size to fill a 16GB VRAM near the limit. RAM requirements will vary.

    9. leumon · · focus · HN ↗
      Unsloth Desktop is the easiest way imo. There are already gguf quants of this model, or simply wait until the official one comes out.
    10. finnjohnsen2 · · focus · HN ↗
      I use opencode + &lt;a decent saas llm&gt; to set up all this new ai generation stuff. GLM-5.3 is my current gun. Safely inside podman containers too because I dont trust this fast moving python eco system at all. Never do I want this running on my main OS.

      I have FLUX.2 klein and dev, Ideogram, LaDA-Image and SenseNova locally. Works great. Ive never touched a file.

      The days of making container yamls myself is over. I read them but I dont edit anymore.

    11. iugtmkbdfil834 · · focus · HN ↗
      I am on AI max 395, comfyUI+qwen models is all you technically need. With today&#x27;s release, I just built a quick and dirty html that allow simpler prompt use and edits ( via headless comfyui ).. its not bad for a day&#x27;s work, but a little too unpolished to publish. I would say, try comfyUI first ( complex, but it worked OOTB ).
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.