It's more likely to work. Most LLM runners are meant to work with any model, which means there are all kinds of ways you might misconfigure them in a way that causes function tooling not to work, or performance to be less than you would like.
DwarfStar's selling point is that it only supports a small set of carefully chosen models, but it supports them really well.
I'm sorry, it's hard for me to understand what point you were trying to make. So existing LLM runners are designed to support all models, and they run all models, but they might be misconfigured? And DS4 is better because it's unable to run all models?
The problem is that with diverse hardware such a set of defaults is much harder to make. You get this matrix of possibilities: gpu, VRAM and memory configurations and then the model axis. This leads to way too many options. The better alternative would be to have the runner self-benchmark what the best settings are given that it already has access to that one particular configuration.
With llama.cpp once you have the model + the runner on the same box you have from 1 ... 40+ configurations of GPUs (depending on how many gpus you have and how many sub-classes of GPUs) for basic options that will load the model. Then you can start multiplying by different batch sizes (1024, 4096, 8192), CPU thread counts (4, 8, 16), tensor splits (this can get really hairy), P2P enabled/disabled, various caching options, speculative decoding options and so on.
The effect is that you can easily spend a day or more benchmarking. On first run of a new model the software should figure this out by itself.
HoldOnAMinute · · focus · HN ↗
simonw · · focus · HN ↗
DwarfStar's selling point is that it only supports a small set of carefully chosen models, but it supports them really well.
locknitpicker · · focus · HN ↗
simonw · · focus · HN ↗
rrgok · · focus · HN ↗
jminnl · · focus · HN ↗
With llama.cpp once you have the model + the runner on the same box you have from 1 ... 40+ configurations of GPUs (depending on how many gpus you have and how many sub-classes of GPUs) for basic options that will load the model. Then you can start multiplying by different batch sizes (1024, 4096, 8192), CPU thread counts (4, 8, 16), tensor splits (this can get really hairy), P2P enabled/disabled, various caching options, speculative decoding options and so on.
The effect is that you can easily spend a day or more benchmarking. On first run of a new model the software should figure this out by itself.
llama-bench is next to useless for this purpose.
aflinik · · focus · HN ↗
simonw · · focus · HN ↗