‹ BackHN Continuity

Thread

From the creator of Redis; run LLM locally with ds4

359 points · 103 comments · fibo

  1. simoiacos · · focus · HN ↗
    Nothing comparable but inspired from DwarfStar I wrote a little inference engine for Intel Xe-LP (no XMX) 32GB laptops. The only model supported right now is a quantized Gemma-4, but I don't exclude in the future to support other MoE of similar size. Too bad we have no Qwen 3.8 35B-A3B yet.

    I'm also looking into expanding the protocol and the engine to support various steering techniques.

    <a href="https:&#x2F;&#x2F;github.com&#x2F;simoneiacomino&#x2F;xenolith" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;simoneiacomino&#x2F;xenolith

    1. ilaksh · · focus · HN ↗
      I wish someone would add Intel support to ds4. And also improve AMD support.

      Maybe Intel and AMD should help them with that.

      1. simoiacos · · focus · HN ↗
        Yeah I see the value but I built Xenolith to target smaller models.

        I heard antirez saying that he designed DwarfStar also to be forked and tuned to everyone&#x27;s specific needs. Do you have a specific machine&#x2F;spec in mind?

        1. ilaksh · · focus · HN ↗
          The recent Intel GPU&#x2F;AI cards. Really the same type of models as ds4
      2. ABS · · focus · HN ↗
        AMD sent antirez a Strix Halo back in June for this purpose
        1. neomantra · · focus · HN ↗
          There were ROCM commits to ds4 in late summer, so that&#x27;s probably related. We include a ROCm build with ds4go (links elsewhere in this comment section), but it is absolutely untested by us whereas the Mac and DGX Spark are very tested.
    2. aziis98 · · focus · HN ↗
      Just tried this on my Intel Ultra 7 255H, I also only have an iGPU. This does ~22tps! Love this.

      I just had to do a little patch to support my iGPU device that is a bit newer than Intel Xe-LP, maybe I&#x27;ll do a PR.

      On a side note the other day I was experimenting with Sonnet 5.5. I gave it the llama cpp repo and told it to extract in a single file inference for a single model + backend (qwen3.5 4b mtp + sycl) and (after a long time) it actually worked! It produced a ~1400 lines file with no deps. I need to check the quality of inference yet but I think this is still a great achievement.

      I&#x27;m pretty sure 2027 will be a very interesting year for local models and inference.

      1. simoiacos · · focus · HN ↗
        Please open a PR! I was too conservative with the supported devices.

        If your GPU supports XMX we could also explore using it to improve the prefill kernel, but I don&#x27;t have the hardware to test it myself.

      2. madduci · · focus · HN ↗
        Can you share it? I have the same hardware, would be interested in trying it
        1. simoiacos · · focus · HN ↗
          I just added experimental support for Xe-LPG and Xe-LPG+ in main so you should be able to try it without a patch now.

          As I said, I don&#x27;t have the hardware to test it myself, so let me know how it goes!

          1. madduci · · focus · HN ↗
            Indeed! I let GLM5.3 write a patch as well, I will test it and submit to you as PR if it works. I&#x27;ve learned in the process that the 255H ships with DeepLink, which helps balancing the work between CPU and GPU (sounds like an OpenCL derivative?).
          2. madduci · · focus · HN ↗
            It worked, the change I had to do was also minimal for the compilation of the .cu unit. You can compile it for the whole iGPU family, not only a specific model.

            I&#x27;ve submitted it as PR, as well as an initial implementation for an OpenAI like API

    3. jacquesm · · focus · HN ↗
      I can&#x27;t run it because I don&#x27;t have that hardware but that looks pretty neat, kudos!
      1. simoiacos · · focus · HN ↗
        Thank you, appreciate it!
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.