‹ BackHN Continuity

Thread

MiMo v2.6

1130 points · 483 comments · volf_

  1. syntaxing · · focus · HN ↗
    All these new models are such tease for us folks with 128GB of shared memory. Buying another unit now to expand to 256GB is a mortgage payment but it’s getting tempting…
    1. brcmthrowaway · · focus · HN ↗
      Is there a gamechanger around the corner to reduce DRAM requirements?
      1. stymaar · · focus · HN ↗
        n-gram per-layer embeddings[1][2] might be it.

        [1] <a href="https:&#x2F;&#x2F;sebastianraschka.com&#x2F;llm-architecture-gallery&#x2F;per-layer-embeddings&#x2F;" rel="nofollow">https:&#x2F;&#x2F;sebastianraschka.com&#x2F;llm-architecture-gallery&#x2F;per-la...

        [2]: See DS 4.1-Flash and Qwen-3.8-Next.

        1. verdverm · · focus · HN ↗
          this is to offload VRAM to DRAM (for GP comment), and makes no difference for URAM
          1. girvo · · focus · HN ↗
            Nah, I’m streaming ngrams off NVMe on my Spark-alike right now. Works surprisingly well (except for when I accidentally bottlenecked it through my NAS)
            1. jkingsman · · focus · HN ↗
              What kind of throughput do you see on what models?
              1. verdverm · · focus · HN ↗
                check out the spark arena website, its the raison d&#x27;etre
              2. girvo · · focus · HN ↗
                GB10 boxes have way more compute than they have memory bandwidth, which nicely fits medium sized MoE models with speculative execution (MTP, DSpark&#x2F;DFlash, etc)

                Qwen 3.8 Flash Next (what I&#x27;m running basically entirely now) sees 30 &#x2F; 35.0 &#x2F; 45 tk&#x2F;s for prose, analysis and code respectively for actual use (not short context benchmarking) with Pi. Thinking blocks are ~35tk&#x2F;s or so.

                The GB10 having so much compute is great for prefill too, 2000-3000&#x2F;s for 14k to 64k token prompts (cold cache too) in the quick benchmark I did. 3500tk&#x2F;s for warm cache which is nice :)

                When I accidentally streamed my ngrams over the 2.5Gb&#x2F;s network, it cut all the throughput down in half basically. Especially notable for the time-to-first-token, which is what clued me in that I&#x27;d messed up somehow!

                For Qwen 3.8 27B, I got it up to a consistent 20tk-25tk&#x2F;s but 27B thinks so much that it was honestly too painful: Flash Next is as smart, as useful, but much faster for real agentic dev usage IMO

                Laguna S 2.1 saw similar numbers to Flash Next if I remember right, but their latest updates means it doesn&#x27;t quite fit a GB10 128GB anymore at full context which is a shame.

                Note: these are all NVFP4 quants (usually a dynamic one where some tensor layers are left at full precision though)

                1. verdverm · · focus · HN ↗
                  I personally stopped caring as much about the tok&#x2F;s as the agents are largely in the background, and so have also moved preference from MoE to dense

                  I want to see about fine-tuning these models a bit on the GB10 to tame that over thinking and some other behaviors (like using tools I don&#x27;t use)

                  qwen 3.8 seems to have been trained with some `rkt` that messes with tool outputs to &quot;save tokens&quot;

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.