I’m a little confusedds4 is referring to “dwarfstar” “4” and references DeepSeek V4 most of the timebut its model agnostic-ishand benchmarks compared to what? what do these large MoE models typically get in tokens per second?I’m garnering this is just an easier way to load large models per expert on consumer hardware? as opposed to the hackier solutions?I’m intruiged. Note that the blogpost says 64gb Macs are good minimums while the github says 96gb is a minimum
I could be wrong but my read from following this on Twitter was that it started on Deepseek 4 (flash). It probably started out focused solely on that model with no guarantee the techniques would transfer to anything else.
yieldcrv · · focus · HN ↗
ds4 is referring to “dwarfstar” “4” and references DeepSeek V4 most of the time
but its model agnostic-ish
and benchmarks compared to what? what do these large MoE models typically get in tokens per second?
I’m garnering this is just an easier way to load large models per expert on consumer hardware? as opposed to the hackier solutions?
I’m intruiged. Note that the blogpost says 64gb Macs are good minimums while the github says 96gb is a minimum
futhey · · focus · HN ↗