‹ BackHN Continuity

Thread

ESP32S3 cluster running 1.58-bit (BitNet) Language model

150 points · 31 comments · nkko

  1. ladyanita22 · · focus · HN ↗
    This is something I've been fantasizing about for long.

    Let's say we took Rust, a language that makes parallelization easier than others (as it helps you avoid some common footguns). How difficult would it be to have a massively parallel computer system made out of many tiny, simple microcontroller-like chips? Let's say we picked many little Risc-V's. Surely this would be an interesting experiment (though I'm not sure whether it'd make economic sense or not...)

    1. sigmoid10 · · focus · HN ↗
      It would certainly not make any economic sense, and I guess that's also why noone is seriously looking into stuff like volunteer/enthusiast clusters of home computers to do inference in the same way that e.g. LHC@home works. The main bottleneck for LLMs is still memory bandwidth. Any memory bus not directly soldered on your GPU is terribly slow. That's why one big GPU with twice the VRAM will always perform significantly better than two GPUs with half the VRAM each. And it's also not like you can just solder more memory onto a chip. At modern speeds, the speed of light is a hard limit. For current GDDR7, signals may only travel like 10mm per cycle.

      If you spread such a system out over dozens or hundreds of tiny chips, you'll be wasting most of its resources and lose hard to anyone who built a single chip setup.

    2. ur-whale · · focus · HN ↗
      > How difficult would it be to have a massively parallel computer system made out of many tiny, simple microcontroller-like chips?

      It's scaling the communication that becomes hard.

      In this project they daisy-chain SPI. I don't believe that would scale very far.

    3. akavel · · focus · HN ↗
      See GreenArrays' 144-core Forth chips by Chuck Moore.
    4. alfanick · · focus · HN ↗
      Look at Xmos, founded by a transputers-dad. Small uCs, that can be connected together into a massive cluster, while being first-citizen of their xC language.
    5. tyingq · · focus · HN ↗
      Not exactly what you're describing, and not shipping yet, but a cluster in a box. 8 cores/node, 8 nodes.

      <a href="https:&#x2F;&#x2F;milkv.io&#x2F;cluster-08" rel="nofollow">https:&#x2F;&#x2F;milkv.io&#x2F;cluster-08

    6. godojo · · focus · HN ↗
      Man I miss Slashdot&#x27;s Beowolf culsters
    7. [deleted] · · focus · HN ↗

      [deleted]

    8. giancarlostoro · · focus · HN ↗
      Mojo might be what you want, especially with &quot;MAX&quot; which is their AI modeling framework, where in other languages you need NVidia&#x27;s libraries, or AMDs, etc the Max libraries just let you talk directly to the GPU &#x2F; CPU &#x2F; ASIC with Mojo. I think Mojo is very underrated in this space right now, but assuming they don&#x27;t mess it up, it could be a major contender in AI. In theory, if someone releases a board, and Modular (company that &#x27;owns&#x27; Mojo) adds it to Max, you&#x27;re basically in the green to experiment as much as you want.

      <a href="https:&#x2F;&#x2F;max.modular.com&#x2F;" rel="nofollow">https:&#x2F;&#x2F;max.modular.com&#x2F;

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.