All these new models are such tease for us folks with 128GB of shared memory. Buying another unit now to expand to 256GB is a mortgage payment but it’s getting tempting…
You can definitely offload n-gram embeddings to storage; they're very sparsely used (only a few KB fetched per token) so this is quite effective. Loading to DRAM only becomes necessary if they are a bottleneck to overall performance (which might happen if you're doing very wide batches and everything else uses super fast VRAM/HBM).
I was looking at the qwen-next-flash, and the weights would fill my OEM Spark on their own, before the n-gram. I'm unclear if offloading to disk can work here, is that what you are implying is possible?!
I have a quirky vLLM on k8s on 2x OEM sparks with about 9 models available to me. I'm not keen to run nightly vLLM, too many issues with it in the past
I have a watchful eye on the diffusion ~ Jev/Kev PR
For what it's worth, Flash Next outperforms every other model that is available to us on the GB10 in all of my testing; though if you have two sparks then the TP=2 version is even better and easier (I don't think you'll need the nightly for that at all, just use the recipe)
Nah, I’m streaming ngrams off NVMe on my Spark-alike right now. Works surprisingly well (except for when I accidentally bottlenecked it through my NAS)
GB10 boxes have way more compute than they have memory bandwidth, which nicely fits medium sized MoE models with speculative execution (MTP, DSpark/DFlash, etc)
Qwen 3.8 Flash Next (what I'm running basically entirely now) sees 30 / 35.0 / 45 tk/s for prose, analysis and code respectively for actual use (not short context benchmarking) with Pi. Thinking blocks are ~35tk/s or so.
The GB10 having so much compute is great for prefill too, 2000-3000/s for 14k to 64k token prompts (cold cache too) in the quick benchmark I did. 3500tk/s for warm cache which is nice :)
When I accidentally streamed my ngrams over the 2.5Gb/s network, it cut all the throughput down in half basically. Especially notable for the time-to-first-token, which is what clued me in that I'd messed up somehow!
For Qwen 3.8 27B, I got it up to a consistent 20tk-25tk/s but 27B thinks so much that it was honestly too painful: Flash Next is as smart, as useful, but much faster for real agentic dev usage IMO
Laguna S 2.1 saw similar numbers to Flash Next if I remember right, but their latest updates means it doesn't quite fit a GB10 128GB anymore at full context which is a shame.
Note: these are all NVFP4 quants (usually a dynamic one where some tensor layers are left at full precision though)
I personally stopped caring as much about the tok/s as the agents are largely in the background, and so have also moved preference from MoE to dense
I want to see about fine-tuning these models a bit on the GB10 to tame that over thinking and some other behaviors (like using tools I don't use)
qwen 3.8 seems to have been trained with some `rkt` that messes with tool outputs to "save tokens"
syntaxing · · focus · HN ↗
brcmthrowaway · · focus · HN ↗
stymaar · · focus · HN ↗
[1] <a href="https://sebastianraschka.com/llm-architecture-gallery/per-layer-embeddings/" rel="nofollow">https://sebastianraschka.com/llm-architecture-gallery/per-la...
[2]: See DS 4.1-Flash and Qwen-3.8-Next.
verdverm · · focus · HN ↗
stymaar · · focus · HN ↗
verdverm · · focus · HN ↗
zozbot234 · · focus · HN ↗
verdverm · · focus · HN ↗
girvo · · focus · HN ↗
It’s an NVFP4 quant, but it fits, and is surprisingly capable.
verdverm · · focus · HN ↗
(or is it somewhere else)
verdverm · · focus · HN ↗
<a href="https://github.com/spark-arena/eugr-recipes" rel="nofollow">https://github.com/spark-arena/eugr-recipes
girvo · · focus · HN ↗
This one!
verdverm · · focus · HN ↗
I have a watchful eye on the diffusion ~ Jev/Kev PR
<a href="https://github.com/vllm-project/vllm/pull/57250" rel="nofollow">https://github.com/vllm-project/vllm/pull/57250
girvo · · focus · HN ↗
I'm so tempted to buy a second one...
verdverm · · focus · HN ↗
I'm running embedding, reranking, and policy tuned models too. Flash Next is not a substitute for those
girvo · · focus · HN ↗
verdverm · · focus · HN ↗
jkingsman · · focus · HN ↗
verdverm · · focus · HN ↗
girvo · · focus · HN ↗
Qwen 3.8 Flash Next (what I'm running basically entirely now) sees 30 / 35.0 / 45 tk/s for prose, analysis and code respectively for actual use (not short context benchmarking) with Pi. Thinking blocks are ~35tk/s or so.
The GB10 having so much compute is great for prefill too, 2000-3000/s for 14k to 64k token prompts (cold cache too) in the quick benchmark I did. 3500tk/s for warm cache which is nice :)
When I accidentally streamed my ngrams over the 2.5Gb/s network, it cut all the throughput down in half basically. Especially notable for the time-to-first-token, which is what clued me in that I'd messed up somehow!
For Qwen 3.8 27B, I got it up to a consistent 20tk-25tk/s but 27B thinks so much that it was honestly too painful: Flash Next is as smart, as useful, but much faster for real agentic dev usage IMO
Laguna S 2.1 saw similar numbers to Flash Next if I remember right, but their latest updates means it doesn't quite fit a GB10 128GB anymore at full context which is a shame.
Note: these are all NVFP4 quants (usually a dynamic one where some tensor layers are left at full precision though)
verdverm · · focus · HN ↗
I want to see about fine-tuning these models a bit on the GB10 to tame that over thinking and some other behaviors (like using tools I don't use)
qwen 3.8 seems to have been trained with some `rkt` that messes with tool outputs to "save tokens"
petu · · focus · HN ↗