‹ BackHN Continuity

Thread

MiMo v2.6

1130 points · 483 comments · volf_

  1. syntaxing · · focus · HN ↗
    All these new models are such tease for us folks with 128GB of shared memory. Buying another unit now to expand to 256GB is a mortgage payment but it’s getting tempting…
    1. brcmthrowaway · · focus · HN ↗
      Is there a gamechanger around the corner to reduce DRAM requirements?
      1. stymaar · · focus · HN ↗
        n-gram per-layer embeddings[1][2] might be it.

        [1] <a href="https:&#x2F;&#x2F;sebastianraschka.com&#x2F;llm-architecture-gallery&#x2F;per-layer-embeddings&#x2F;" rel="nofollow">https:&#x2F;&#x2F;sebastianraschka.com&#x2F;llm-architecture-gallery&#x2F;per-la...

        [2]: See DS 4.1-Flash and Qwen-3.8-Next.

        1. verdverm · · focus · HN ↗
          this is to offload VRAM to DRAM (for GP comment), and makes no difference for URAM
          1. zozbot234 · · focus · HN ↗
            You can definitely offload n-gram embeddings to storage; they&#x27;re very sparsely used (only a few KB fetched per token) so this is quite effective. Loading to DRAM only becomes necessary if they are a bottleneck to overall performance (which might happen if you&#x27;re doing very wide batches and everything else uses super fast VRAM&#x2F;HBM).
            1. verdverm · · focus · HN ↗
              I was looking at the qwen-next-flash, and the weights would fill my OEM Spark on their own, before the n-gram. I&#x27;m unclear if offloading to disk can work here, is that what you are implying is possible?!
              1. girvo · · focus · HN ↗
                Check out eugr’s TP=1 sparkrun recipe :)

                It’s an NVFP4 quant, but it fits, and is surprisingly capable.

                1. verdverm · · focus · HN ↗
                  do you have a HF link? HF search is not uncovering it for me

                  (or is it somewhere else)

                  1. girvo · · focus · HN ↗
                    <a href="https:&#x2F;&#x2F;github.com&#x2F;spark-arena&#x2F;eugr-recipes&#x2F;blob&#x2F;main&#x2F;recipes&#x2F;qwen3.8-flash-next-nvfp4-solo.yaml" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;spark-arena&#x2F;eugr-recipes&#x2F;blob&#x2F;main&#x2F;recipe...

                    This one!

                    1. verdverm · · focus · HN ↗
                      I have a quirky vLLM on k8s on 2x OEM sparks with about 9 models available to me. I&#x27;m not keen to run nightly vLLM, too many issues with it in the past

                      I have a watchful eye on the diffusion ~ Jev&#x2F;Kev PR

                      <a href="https:&#x2F;&#x2F;github.com&#x2F;vllm-project&#x2F;vllm&#x2F;pull&#x2F;57250" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;vllm-project&#x2F;vllm&#x2F;pull&#x2F;57250

                      1. girvo · · focus · HN ↗
                        For what it&#x27;s worth, Flash Next outperforms every other model that is available to us on the GB10 in all of my testing; though if you have two sparks then the TP=2 version is even better and easier (I don&#x27;t think you&#x27;ll need the nightly for that at all, just use the recipe)

                        I&#x27;m so tempted to buy a second one...

                        1. verdverm · · focus · HN ↗
                          prices have gone up quite a bit...

                          I&#x27;m running embedding, reranking, and policy tuned models too. Flash Next is not a substitute for those

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.