‹ BackHN Continuity

Thread

From the creator of Redis; run LLM locally with ds4

359 points · 103 comments · fibo

  1. simoiacos · · focus · HN ↗
    Nothing comparable but inspired from DwarfStar I wrote a little inference engine for Intel Xe-LP (no XMX) 32GB laptops. The only model supported right now is a quantized Gemma-4, but I don't exclude in the future to support other MoE of similar size. Too bad we have no Qwen 3.8 35B-A3B yet.

    I'm also looking into expanding the protocol and the engine to support various steering techniques.

    <a href="https:&#x2F;&#x2F;github.com&#x2F;simoneiacomino&#x2F;xenolith" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;simoneiacomino&#x2F;xenolith

    1. aziis98 · · focus · HN ↗
      Just tried this on my Intel Ultra 7 255H, I also only have an iGPU. This does ~22tps! Love this.

      I just had to do a little patch to support my iGPU device that is a bit newer than Intel Xe-LP, maybe I&#x27;ll do a PR.

      On a side note the other day I was experimenting with Sonnet 5.5. I gave it the llama cpp repo and told it to extract in a single file inference for a single model + backend (qwen3.5 4b mtp + sycl) and (after a long time) it actually worked! It produced a ~1400 lines file with no deps. I need to check the quality of inference yet but I think this is still a great achievement.

      I&#x27;m pretty sure 2027 will be a very interesting year for local models and inference.

      1. madduci · · focus · HN ↗
        Can you share it? I have the same hardware, would be interested in trying it
        1. simoiacos · · focus · HN ↗
          I just added experimental support for Xe-LPG and Xe-LPG+ in main so you should be able to try it without a patch now.

          As I said, I don&#x27;t have the hardware to test it myself, so let me know how it goes!

          1. madduci · · focus · HN ↗
            It worked, the change I had to do was also minimal for the compilation of the .cu unit. You can compile it for the whole iGPU family, not only a specific model.

            I&#x27;ve submitted it as PR, as well as an initial implementation for an OpenAI like API

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.