‹ BackHN Continuity

Thread

From the creator of Redis; run LLM locally with ds4

359 points · 103 comments · fibo

  1. cuttothechase · · focus · HN ↗
    Wondering how well this does with tool calling. Any one has any numbers or videos or anything using this?

    From the github repo it seems like you really don't need a big Mac with huge amounts of RAM but SSD is sufficient.

    If this is anywhere near 50 TPS, that would be a game changer in the personal LLM space!

    1. neomantra · · focus · HN ↗
      I just put up videos of some of the ds4go I mentioned. Most of it is long and boring, but it's there.

      <a href="https:&#x2F;&#x2F;gist.github.com&#x2F;neomantra&#x2F;d49df05d6b137b9e6844186499715756" rel="nofollow">https:&#x2F;&#x2F;gist.github.com&#x2F;neomantra&#x2F;d49df05d6b137b9e6844186499...

      I started playing with local LLM+MCP in April 2025... I had to beg Qwen to look at the tool list and try anything.

      These ds4 models, will happily call tools and all those harnesses I&#x27;ve made are composed of custom tools.

      Once the HuggingFace+OpenAI showed how powerful notes are, I added a scratchpad tool to ds4go to improve self-improvement.

      While you can do 64G&#x2F;96G with the Qwen3.8 model, realistically you need 128G. Also, despite tons of playing with local models, the cloud-hosted models on bigger iron are smarter and faster. I don&#x27;t truly code with my local models and don&#x27;t recommend this path right now to replace something like Opus&#x2F;Astra or full-brain DeepSeek4.

      The &quot;frontier-ness&quot; of ds4 is great though! It has vast knowledge and thinking capability. Look at that steering video especially. I&#x27;m now exploring using ds4 for high-level thinking to create prompts for denser coding models.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.