Nothing comparable but inspired from DwarfStar I wrote a little inference engine for Intel Xe-LP (no XMX) 32GB laptops. The only model supported right now is a quantized Gemma-4, but I don't exclude in the future to support other MoE of similar size. Too bad we have no Qwen 3.8 35B-A3B yet.
I'm also looking into expanding the protocol and the engine to support various steering techniques.
Yeah I see the value but I built Xenolith to target smaller models.
I heard antirez saying that he designed DwarfStar also to be forked and tuned to everyone's specific needs. Do you have a specific machine/spec in mind?
There were ROCM commits to ds4 in late summer, so that's probably related. We include a ROCm build with ds4go (links elsewhere in this comment section), but it is absolutely untested by us whereas the Mac and DGX Spark are very tested.
Just tried this on my Intel Ultra 7 255H, I also only have an iGPU. This does ~22tps! Love this.
I just had to do a little patch to support my iGPU device that is a bit newer than Intel Xe-LP, maybe I'll do a PR.
On a side note the other day I was experimenting with Sonnet 5.5. I gave it the llama cpp repo and told it to extract in a single file inference for a single model + backend (qwen3.5 4b mtp + sycl) and (after a long time) it actually worked! It produced a ~1400 lines file with no deps. I need to check the quality of inference yet but I think this is still a great achievement.
I'm pretty sure 2027 will be a very interesting year for local models and inference.
Indeed! I let GLM5.3 write a patch as well, I will test it and submit to you as PR if it works. I've learned in the process that the 255H ships with DeepLink, which helps balancing the work between CPU and GPU (sounds like an OpenCL derivative?).
It worked, the change I had to do was also minimal for the compilation of the .cu unit. You can compile it for the whole iGPU family, not only a specific model.
I've submitted it as PR, as well as an initial implementation for an OpenAI like API
simoiacos · · focus · HN ↗
I'm also looking into expanding the protocol and the engine to support various steering techniques.
<a href="https://github.com/simoneiacomino/xenolith" rel="nofollow">https://github.com/simoneiacomino/xenolith
ilaksh · · focus · HN ↗
Maybe Intel and AMD should help them with that.
simoiacos · · focus · HN ↗
I heard antirez saying that he designed DwarfStar also to be forked and tuned to everyone's specific needs. Do you have a specific machine/spec in mind?
ilaksh · · focus · HN ↗
ABS · · focus · HN ↗
neomantra · · focus · HN ↗
aziis98 · · focus · HN ↗
I just had to do a little patch to support my iGPU device that is a bit newer than Intel Xe-LP, maybe I'll do a PR.
On a side note the other day I was experimenting with Sonnet 5.5. I gave it the llama cpp repo and told it to extract in a single file inference for a single model + backend (qwen3.5 4b mtp + sycl) and (after a long time) it actually worked! It produced a ~1400 lines file with no deps. I need to check the quality of inference yet but I think this is still a great achievement.
I'm pretty sure 2027 will be a very interesting year for local models and inference.
simoiacos · · focus · HN ↗
If your GPU supports XMX we could also explore using it to improve the prefill kernel, but I don't have the hardware to test it myself.
madduci · · focus · HN ↗
simoiacos · · focus · HN ↗
As I said, I don't have the hardware to test it myself, so let me know how it goes!
madduci · · focus · HN ↗
madduci · · focus · HN ↗
I've submitted it as PR, as well as an initial implementation for an OpenAI like API
jacquesm · · focus · HN ↗
simoiacos · · focus · HN ↗