‹ BackHN Continuity

Thread

From the creator of Redis; run LLM locally with ds4

359 points · 103 comments · fibo

  1. HoldOnAMinute · · focus · HN ↗
    How is this different from other LLM runners?
    1. csmlab_notes · · focus · HN ↗

      [dead]

    2. pydry · · focus · HN ↗
      My instinctive reaction from the readme is that it isnt. It's apparently a vibe coded knock off of llama.CPP.
      1. jminnl · · focus · HN ↗
        The llama.cpp guys don't get nearly enough credit for their work. Though the quality of the codebase is dropping over time, it is still quite high compared to most of the alternatives, and it is still one of the most stable ways to run a large variety of models.

        Definitely worth looking at if you have only a single 5090 is ninfer, and various hardware specific forks (3090, 4090).

    3. ilaksh · · focus · HN ↗
      Emphasis on performance and usable coding/agentic ability for consumer AI hardware. Does not attempt to handle all models or hardware at once but rather focuses on optimizing the best options for that category of hardware.
    4. simonw · · focus · HN ↗
      It's more likely to work. Most LLM runners are meant to work with any model, which means there are all kinds of ways you might misconfigure them in a way that causes function tooling not to work, or performance to be less than you would like.

      DwarfStar's selling point is that it only supports a small set of carefully chosen models, but it supports them really well.

      1. locknitpicker · · focus · HN ↗
        I'm sorry, it's hard for me to understand what point you were trying to make. So existing LLM runners are designed to support all models, and they run all models, but they might be misconfigured? And DS4 is better because it's unable to run all models?
        1. simonw · · focus · HN ↗
          It has better defaults.
          1. rrgok · · focus · HN ↗
            So wouldn't be easier just to provide a repository of better defaults for each model? Like lsp-config for neovim?
            1. jminnl · · focus · HN ↗
              The problem is that with diverse hardware such a set of defaults is much harder to make. You get this matrix of possibilities: gpu, VRAM and memory configurations and then the model axis. This leads to way too many options. The better alternative would be to have the runner self-benchmark what the best settings are given that it already has access to that one particular configuration.

              With llama.cpp once you have the model + the runner on the same box you have from 1 ... 40+ configurations of GPUs (depending on how many gpus you have and how many sub-classes of GPUs) for basic options that will load the model. Then you can start multiplying by different batch sizes (1024, 4096, 8192), CPU thread counts (4, 8, 16), tensor splits (this can get really hairy), P2P enabled/disabled, various caching options, speculative decoding options and so on.

              The effect is that you can easily spend a day or more benchmarking. On first run of a new model the software should figure this out by itself.

              llama-bench is next to useless for this purpose.

          2. aflinik · · focus · HN ↗
            Does that imply that with enough effort spent on fine-tuning the configuration, other LLM runners can achieve similar results?
            1. simonw · · focus · HN ↗
              Yes.
    5. ttoinou · · focus · HN ↗
      Lots of small details are taken care of so it runs smoothly. For example ds4-agent is append only, never rewriting history of messages, keeping KV cache prefix reusable. Huge benefit
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.