‹ BackHN Continuity

Thread

MiMo v2.6

1130 points · 483 comments · volf_

  1. syntaxing · · focus · HN ↗
    All these new models are such tease for us folks with 128GB of shared memory. Buying another unit now to expand to 256GB is a mortgage payment but it’s getting tempting…
    1. brcmthrowaway · · focus · HN ↗
      Is there a gamechanger around the corner to reduce DRAM requirements?
      1. stymaar · · focus · HN ↗
        n-gram per-layer embeddings[1][2] might be it.

        [1] <a href="https:&#x2F;&#x2F;sebastianraschka.com&#x2F;llm-architecture-gallery&#x2F;per-layer-embeddings&#x2F;" rel="nofollow">https:&#x2F;&#x2F;sebastianraschka.com&#x2F;llm-architecture-gallery&#x2F;per-la...

        [2]: See DS 4.1-Flash and Qwen-3.8-Next.

        1. verdverm · · focus · HN ↗
          this is to offload VRAM to DRAM (for GP comment), and makes no difference for URAM
          1. girvo · · focus · HN ↗
            Nah, I’m streaming ngrams off NVMe on my Spark-alike right now. Works surprisingly well (except for when I accidentally bottlenecked it through my NAS)
            1. verdverm · · focus · HN ↗
              interesting, peer comment seems to indicate this is a possibility as well, will have to take a deeper look
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.