Probably ComfyUI is one of the easiest way to get started with local image/video models. Or perhaps vLLM, if they have support for it already, would be something like `vllm serve <model> --omni --port 9080`
why not just as you suggested i.e. <a href="https://qwen.readthedocs.io/en/latest/run_locally/llama.cpp.html#llama-server" rel="nofollow">https://qwen.readthedocs.io/en/latest/run_locally/llama.cpp.... then get the result either via a UI or wget/curl it back?
Seems I'm missing something. Does this model support other inputs?
Image outputs are supported, videos I'm not sure but I don't think that's an output, just a preview of the equirectangular example, so, same question here, what does this model outputs that isn't supported?
like <a href="https://github.com/ggml-org/llama.cpp/tree/master/examples/diffusion" rel="nofollow">https://github.com/ggml-org/llama.cpp/tree/master/examples/d... ?
At the risk of stating the obvious llama.cpp isn't just about LLaMa as <a href="https://github.com/ggml-org/llama.cpp/blob/master/src/llama-arch.h#L13" rel="nofollow">https://github.com/ggml-org/llama.cpp/blob/master/src/llama-... someone else pointed out.
Not all architectures are supported by llama.cpp . The GGUF format encodes the NN in a standardized way, but then you need code that can use that NN structure.
I understand that llama.cpp could only output text, last time I checked (I do not know how to find a good source for that though).
See <a href="https://github.com/ggml-org/llama.cpp/blob/master/src/llama-arch.h" rel="nofollow">https://github.com/ggml-org/llama.cpp/blob/master/src/llama-... , the
I haven't used llama.cpp for image generation either but I recall an issue about it. Unfortunately I can't pinpoint it now and there is the older closed issue <a href="https://github.com/ggml-org/llama.cpp/issues/4408" rel="nofollow">https://github.com/ggml-org/llama.cpp/issues/4408 so unless mtmd supports also multimodal outputs out of the box safe to assume output is still limited to text.
> Diffusers, ComfyUI, vLLM-Omni, SGLang, and LightX2V
I think that's all Python (not a direct executable).
You could just do (see the "Quick Start") four `pip install` and have a dozen lines script to generate the image. But `llama.cpp` and similar do not require e.g. installing Torch (or PyTorch) - you can use `llama.cpp` on a non-specialized machine.
There is difussion.cpp which is intended for those types of models. I set up krea-2-turbo with the help of ChatGPT 2 months ago, if you have a capable computer that's what I would suggest once it becomes supported.
Additional question is what kind of local hardware would be required for this? 7B parameters sounds very light weight, but I'm not sure. (Edit: The download is 33 GB).
Edit x2: As usual I'm in a twisty maze of pip packages that don't work together, with obscure errors about missing modules, even though I followed the instructions on the page to the letter. I really wish people didn't use Python for this stuff. A simple C/C++ program would be so much better.
It's about 16 GiB at Q8 quants (combining both the image and language parts). (Meaning, community quantized models from HuggingFace).
I think it will technically run on anything that has enough memory. I just tried it on a standard laptop (dual-channel DDR5), and it took about 3 minutes for a 512x512. If you'd want to run it at interactive speeds, you would want a GPU (one which fits this in VRAM).
> "I really wish people didn't use Python for this stuff. A simple C/C++ program would be so much better."
I am using sd.cpp, which is the cousin of llama.cpp: <a href="https://github.com/leejet/stable-diffusion.cpp" rel="nofollow">https://github.com/leejet/stable-diffusion.cpp
I've set it up on my local machine just now, as my first local image diffuser. I can confirm it's very easy.
I tried stable-diffusion.cpp, following its compile guide here[0], and its Qwen Image-2.1 specific instructions here[1]. It works out of the box. I made a test pelican[2]. It took 3 minutes on a CPU.
I use opencode + <a decent saas llm> to set up all this new ai generation stuff. GLM-5.3 is my current gun. Safely inside podman containers too because I dont trust this fast moving python eco system at all. Never do I want this running on my main OS.
I have FLUX.2 klein and dev, Ideogram, LaDA-Image and SenseNova locally. Works great. Ive never touched a file.
The days of making container yamls myself is over. I read them but I dont edit anymore.
I am on AI max 395, comfyUI+qwen models is all you technically need. With today's release, I just built a quick and dirty html that allow simpler prompt use and edits ( via headless comfyui ).. its not bad for a day's work, but a little too unpolished to publish. I would say, try comfyUI first ( complex, but it worked OOTB ).
mdp2021 · · focus · HN ↗
(I mean: outside direct or substantial use of Python, and running the Neural Network in the most efficient way.)
embedding-shape · · focus · HN ↗
utopiah · · focus · HN ↗
mdp2021 · · focus · HN ↗
utopiah · · focus · HN ↗
exe34 · · focus · HN ↗
> Currently, we support image, audio and video input.
utopiah · · focus · HN ↗
Image outputs are supported, videos I'm not sure but I don't think that's an output, just a preview of the equirectangular example, so, same question here, what does this model outputs that isn't supported?
[deleted] · · focus · HN ↗
[deleted]
exe34 · · focus · HN ↗
utopiah · · focus · HN ↗
At the risk of stating the obvious llama.cpp isn't just about LLaMa as <a href="https://github.com/ggml-org/llama.cpp/blob/master/src/llama-arch.h#L13" rel="nofollow">https://github.com/ggml-org/llama.cpp/blob/master/src/llama-... someone else pointed out.
exe34 · · focus · HN ↗
utopiah · · focus · HN ↗
mdp2021 · · focus · HN ↗
I understand that llama.cpp could only output text, last time I checked (I do not know how to find a good source for that though).
See <a href="https://github.com/ggml-org/llama.cpp/blob/master/src/llama-arch.h" rel="nofollow">https://github.com/ggml-org/llama.cpp/blob/master/src/llama-... , the
...utopiah · · focus · HN ↗
fp64 · · focus · HN ↗
mdp2021 · · focus · HN ↗
I think that's all Python (not a direct executable).
You could just do (see the "Quick Start") four `pip install` and have a dozen lines script to generate the image. But `llama.cpp` and similar do not require e.g. installing Torch (or PyTorch) - you can use `llama.cpp` on a non-specialized machine.
wgd · · focus · HN ↗
I don't think I have ever once run "pip install transformers" and had it work without three rounds of fiddling
mdp2021 · · focus · HN ↗
Yep, that's (also) what I meant ;)
Lean, efficient... Also sensible and trouble-less.
iugtmkbdfil834 · · focus · HN ↗
Iolaum · · focus · HN ↗
rwmj · · focus · HN ↗
Edit x2: As usual I'm in a twisty maze of pip packages that don't work together, with obscure errors about missing modules, even though I followed the instructions on the page to the letter. I really wish people didn't use Python for this stuff. A simple C/C++ program would be so much better.
peri-cl · · focus · HN ↗
I think it will technically run on anything that has enough memory. I just tried it on a standard laptop (dual-channel DDR5), and it took about 3 minutes for a 512x512. If you'd want to run it at interactive speeds, you would want a GPU (one which fits this in VRAM).
> "I really wish people didn't use Python for this stuff. A simple C/C++ program would be so much better."
You mean besides stable-diffusion.cpp ?
rwmj · · focus · HN ↗
Yes, thanks, I didn't know about that. Will try it.
[deleted] · · focus · HN ↗
[deleted]
nkhgfugjk · · focus · HN ↗
it already has day-0 qwen image 2.1 support!
peri-cl · · focus · HN ↗
I tried stable-diffusion.cpp, following its compile guide here[0], and its Qwen Image-2.1 specific instructions here[1]. It works out of the box. I made a test pelican[2]. It took 3 minutes on a CPU.
[0] <a href="https://github.com/leejet/stable-diffusion.cpp/blob/master/docs/build.md" rel="nofollow">https://github.com/leejet/stable-diffusion.cpp/blob/master/d...
[1] <a href="https://github.com/leejet/stable-diffusion.cpp/blob/master/docs/qwen_image_2.1.md" rel="nofollow">https://github.com/leejet/stable-diffusion.cpp/blob/master/d...
[2] <a href="https://i.ibb.co/yMknC2K/output.png" rel="nofollow">https://i.ibb.co/yMknC2K/output.png
mdp2021 · · focus · HN ↗
peri-cl · · focus · HN ↗
mdp2021 · · focus · HN ↗
In fact, like it appears in the reports above, it is "7b" as in
> 7B parameters in its visual generation component
It seems they calibrated the size to fill a 16GB VRAM near the limit. RAM requirements will vary.
leumon · · focus · HN ↗
finnjohnsen2 · · focus · HN ↗
I have FLUX.2 klein and dev, Ideogram, LaDA-Image and SenseNova locally. Works great. Ive never touched a file.
The days of making container yamls myself is over. I read them but I dont edit anymore.
iugtmkbdfil834 · · focus · HN ↗