All these new models are such tease for us folks with 128GB of shared memory. Buying another unit now to expand to 256GB is a mortgage payment but it’s getting tempting…
You can definitely offload n-gram embeddings to storage; they're very sparsely used (only a few KB fetched per token) so this is quite effective. Loading to DRAM only becomes necessary if they are a bottleneck to overall performance (which might happen if you're doing very wide batches and everything else uses super fast VRAM/HBM).
I was looking at the qwen-next-flash, and the weights would fill my OEM Spark on their own, before the n-gram. I'm unclear if offloading to disk can work here, is that what you are implying is possible?!
syntaxing · · focus · HN ↗
brcmthrowaway · · focus · HN ↗
stymaar · · focus · HN ↗
[1] <a href="https://sebastianraschka.com/llm-architecture-gallery/per-layer-embeddings/" rel="nofollow">https://sebastianraschka.com/llm-architecture-gallery/per-la...
[2]: See DS 4.1-Flash and Qwen-3.8-Next.
verdverm · · focus · HN ↗
zozbot234 · · focus · HN ↗
verdverm · · focus · HN ↗
girvo · · focus · HN ↗
It’s an NVFP4 quant, but it fits, and is surprisingly capable.
verdverm · · focus · HN ↗
(or is it somewhere else)
verdverm · · focus · HN ↗
<a href="https://github.com/spark-arena/eugr-recipes" rel="nofollow">https://github.com/spark-arena/eugr-recipes