Nothing comparable but inspired from DwarfStar I wrote a little inference engine for Intel Xe-LP (no XMX) 32GB laptops. The only model supported right now is a quantized Gemma-4, but I don't exclude in the future to support other MoE of similar size. Too bad we have no Qwen 3.8 35B-A3B yet.
I'm also looking into expanding the protocol and the engine to support various steering techniques.
Just tried this on my Intel Ultra 7 255H, I also only have an iGPU. This does ~22tps! Love this.
I just had to do a little patch to support my iGPU device that is a bit newer than Intel Xe-LP, maybe I'll do a PR.
On a side note the other day I was experimenting with Sonnet 5.5. I gave it the llama cpp repo and told it to extract in a single file inference for a single model + backend (qwen3.5 4b mtp + sycl) and (after a long time) it actually worked! It produced a ~1400 lines file with no deps. I need to check the quality of inference yet but I think this is still a great achievement.
I'm pretty sure 2027 will be a very interesting year for local models and inference.
Indeed! I let GLM5.3 write a patch as well, I will test it and submit to you as PR if it works. I've learned in the process that the 255H ships with DeepLink, which helps balancing the work between CPU and GPU (sounds like an OpenCL derivative?).
It worked, the change I had to do was also minimal for the compilation of the .cu unit. You can compile it for the whole iGPU family, not only a specific model.
I've submitted it as PR, as well as an initial implementation for an OpenAI like API
simoiacos · · focus · HN ↗
I'm also looking into expanding the protocol and the engine to support various steering techniques.
<a href="https://github.com/simoneiacomino/xenolith" rel="nofollow">https://github.com/simoneiacomino/xenolith
aziis98 · · focus · HN ↗
I just had to do a little patch to support my iGPU device that is a bit newer than Intel Xe-LP, maybe I'll do a PR.
On a side note the other day I was experimenting with Sonnet 5.5. I gave it the llama cpp repo and told it to extract in a single file inference for a single model + backend (qwen3.5 4b mtp + sycl) and (after a long time) it actually worked! It produced a ~1400 lines file with no deps. I need to check the quality of inference yet but I think this is still a great achievement.
I'm pretty sure 2027 will be a very interesting year for local models and inference.
madduci · · focus · HN ↗
simoiacos · · focus · HN ↗
As I said, I don't have the hardware to test it myself, so let me know how it goes!
madduci · · focus · HN ↗
madduci · · focus · HN ↗
I've submitted it as PR, as well as an initial implementation for an OpenAI like API