‹ BackHN Continuity

Thread

MiMo v2.6

1130 points · 483 comments · volf_

  1. syntaxing · · focus · HN ↗
    All these new models are such tease for us folks with 128GB of shared memory. Buying another unit now to expand to 256GB is a mortgage payment but it’s getting tempting…
    1. brcmthrowaway · · focus · HN ↗
      Is there a gamechanger around the corner to reduce DRAM requirements?
      1. stymaar · · focus · HN ↗
        n-gram per-layer embeddings[1][2] might be it.

        [1] <a href="https:&#x2F;&#x2F;sebastianraschka.com&#x2F;llm-architecture-gallery&#x2F;per-layer-embeddings&#x2F;" rel="nofollow">https:&#x2F;&#x2F;sebastianraschka.com&#x2F;llm-architecture-gallery&#x2F;per-la...

        [2]: See DS 4.1-Flash and Qwen-3.8-Next.

        1. verdverm · · focus · HN ↗
          this is to offload VRAM to DRAM (for GP comment), and makes no difference for URAM
          1. stymaar · · focus · HN ↗
            Am I missing a joke? WTF is URAM?
            1. verdverm · · focus · HN ↗
              unified memory, not sure if anyone uses URAM, I human hallucinated it
          2. zozbot234 · · focus · HN ↗
            You can definitely offload n-gram embeddings to storage; they&#x27;re very sparsely used (only a few KB fetched per token) so this is quite effective. Loading to DRAM only becomes necessary if they are a bottleneck to overall performance (which might happen if you&#x27;re doing very wide batches and everything else uses super fast VRAM&#x2F;HBM).
            1. verdverm · · focus · HN ↗
              I was looking at the qwen-next-flash, and the weights would fill my OEM Spark on their own, before the n-gram. I&#x27;m unclear if offloading to disk can work here, is that what you are implying is possible?!
              1. girvo · · focus · HN ↗
                Check out eugr’s TP=1 sparkrun recipe :)

                It’s an NVFP4 quant, but it fits, and is surprisingly capable.

                1. verdverm · · focus · HN ↗
                  do you have a HF link? HF search is not uncovering it for me

                  (or is it somewhere else)

                  1. verdverm · · focus · HN ↗
                    looks like this is likely it

                    <a href="https:&#x2F;&#x2F;github.com&#x2F;spark-arena&#x2F;eugr-recipes" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;spark-arena&#x2F;eugr-recipes

                  2. girvo · · focus · HN ↗
                    <a href="https:&#x2F;&#x2F;github.com&#x2F;spark-arena&#x2F;eugr-recipes&#x2F;blob&#x2F;main&#x2F;recipes&#x2F;qwen3.8-flash-next-nvfp4-solo.yaml" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;spark-arena&#x2F;eugr-recipes&#x2F;blob&#x2F;main&#x2F;recipe...

                    This one!

                    1. verdverm · · focus · HN ↗
                      I have a quirky vLLM on k8s on 2x OEM sparks with about 9 models available to me. I&#x27;m not keen to run nightly vLLM, too many issues with it in the past

                      I have a watchful eye on the diffusion ~ Jev&#x2F;Kev PR

                      <a href="https:&#x2F;&#x2F;github.com&#x2F;vllm-project&#x2F;vllm&#x2F;pull&#x2F;57250" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;vllm-project&#x2F;vllm&#x2F;pull&#x2F;57250

                      1. girvo · · focus · HN ↗
                        For what it&#x27;s worth, Flash Next outperforms every other model that is available to us on the GB10 in all of my testing; though if you have two sparks then the TP=2 version is even better and easier (I don&#x27;t think you&#x27;ll need the nightly for that at all, just use the recipe)

                        I&#x27;m so tempted to buy a second one...

                        1. verdverm · · focus · HN ↗
                          prices have gone up quite a bit...

                          I&#x27;m running embedding, reranking, and policy tuned models too. Flash Next is not a substitute for those

          3. girvo · · focus · HN ↗
            Nah, I’m streaming ngrams off NVMe on my Spark-alike right now. Works surprisingly well (except for when I accidentally bottlenecked it through my NAS)
            1. verdverm · · focus · HN ↗
              interesting, peer comment seems to indicate this is a possibility as well, will have to take a deeper look
            2. jkingsman · · focus · HN ↗
              What kind of throughput do you see on what models?
              1. verdverm · · focus · HN ↗
                check out the spark arena website, its the raison d&#x27;etre
              2. girvo · · focus · HN ↗
                GB10 boxes have way more compute than they have memory bandwidth, which nicely fits medium sized MoE models with speculative execution (MTP, DSpark&#x2F;DFlash, etc)

                Qwen 3.8 Flash Next (what I&#x27;m running basically entirely now) sees 30 &#x2F; 35.0 &#x2F; 45 tk&#x2F;s for prose, analysis and code respectively for actual use (not short context benchmarking) with Pi. Thinking blocks are ~35tk&#x2F;s or so.

                The GB10 having so much compute is great for prefill too, 2000-3000&#x2F;s for 14k to 64k token prompts (cold cache too) in the quick benchmark I did. 3500tk&#x2F;s for warm cache which is nice :)

                When I accidentally streamed my ngrams over the 2.5Gb&#x2F;s network, it cut all the throughput down in half basically. Especially notable for the time-to-first-token, which is what clued me in that I&#x27;d messed up somehow!

                For Qwen 3.8 27B, I got it up to a consistent 20tk-25tk&#x2F;s but 27B thinks so much that it was honestly too painful: Flash Next is as smart, as useful, but much faster for real agentic dev usage IMO

                Laguna S 2.1 saw similar numbers to Flash Next if I remember right, but their latest updates means it doesn&#x27;t quite fit a GB10 128GB anymore at full context which is a shame.

                Note: these are all NVFP4 quants (usually a dynamic one where some tensor layers are left at full precision though)

                1. verdverm · · focus · HN ↗
                  I personally stopped caring as much about the tok&#x2F;s as the agents are largely in the background, and so have also moved preference from MoE to dense

                  I want to see about fine-tuning these models a bit on the GB10 to tame that over thinking and some other behaviors (like using tools I don&#x27;t use)

                  qwen 3.8 seems to have been trained with some `rkt` that messes with tool outputs to &quot;save tokens&quot;

          4. petu · · focus · HN ↗
            n-grams can be kept on SSD, no need to hold them in any kind of RAM (at least w&#x2F;o batching)
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.