‹ BackHN Continuity

Thread

Qwen Image 2.1

740 points · 199 comments · jmillikin

  1. mdp2021 · · focus · HN ↗
    How do you use this model locally, similarly to using `llama-server -m <model>`?

    (I mean: outside direct or substantial use of Python, and running the Neural Network in the most efficient way.)

    1. utopiah · · focus · HN ↗
      why not just as you suggested i.e. <a href="https:&#x2F;&#x2F;qwen.readthedocs.io&#x2F;en&#x2F;latest&#x2F;run_locally&#x2F;llama.cpp.html#llama-server" rel="nofollow">https:&#x2F;&#x2F;qwen.readthedocs.io&#x2F;en&#x2F;latest&#x2F;run_locally&#x2F;llama.cpp.... then get the result either via a UI or wget&#x2F;curl it back?
      1. mdp2021 · · focus · HN ↗
        I am not sure that llama.cpp also supports image generation models.
        1. utopiah · · focus · HN ↗
          it&#x27;s multimodal, see <a href="https:&#x2F;&#x2F;github.com&#x2F;ggml-org&#x2F;llama.cpp&#x2F;blob&#x2F;master&#x2F;docs&#x2F;multimodal.md" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;ggml-org&#x2F;llama.cpp&#x2F;blob&#x2F;master&#x2F;docs&#x2F;multi...
          1. exe34 · · focus · HN ↗
            Multimodal doesn&#x27;t guarantee input and output.

            &gt; Currently, we support image, audio and video input.

            1. utopiah · · focus · HN ↗
              Seems I&#x27;m missing something. Does this model support other inputs?

              Image outputs are supported, videos I&#x27;m not sure but I don&#x27;t think that&#x27;s an output, just a preview of the equirectangular example, so, same question here, what does this model outputs that isn&#x27;t supported?

              1. exe34 · · focus · HN ↗
                It&#x27;s a diffusion model, completely different from autoregressive attention models.
                1. utopiah · · focus · HN ↗
                  like <a href="https:&#x2F;&#x2F;github.com&#x2F;ggml-org&#x2F;llama.cpp&#x2F;tree&#x2F;master&#x2F;examples&#x2F;diffusion" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;ggml-org&#x2F;llama.cpp&#x2F;tree&#x2F;master&#x2F;examples&#x2F;d... ?

                  At the risk of stating the obvious llama.cpp isn&#x27;t just about LLaMa as <a href="https:&#x2F;&#x2F;github.com&#x2F;ggml-org&#x2F;llama.cpp&#x2F;blob&#x2F;master&#x2F;src&#x2F;llama-arch.h#L13" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;ggml-org&#x2F;llama.cpp&#x2F;blob&#x2F;master&#x2F;src&#x2F;llama-... someone else pointed out.

                  1. exe34 · · focus · HN ↗
                    Aha I was wrong. Thanks for sharing that!
                    1. utopiah · · focus · HN ↗
                      no worries, I was wrong too, it is multimodal but only for inputs apparently, so for now there seem to only be text as output but no image as output
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.