‹ BackHN Continuity

Thread

MicroLLM Lab – Try 7 tiny LLM's in the browser

283 points · 113 comments · logicallee

  1. bhouston · · focus · HN ↗
    I built something like this just last month, using a few of the same models, but I used ThreeJS&#x27;s Three-Shading-Language abstraction to do it: <a href="https:&#x2F;&#x2F;three-llm.ben3d.ca&#x2F;?model=qwen3.5-0.8b" rel="nofollow">https:&#x2F;&#x2F;three-llm.ben3d.ca&#x2F;?model=qwen3.5-0.8b
    1. logicallee · · focus · HN ↗
      Awesome! I tried it and after loading (which took a while as it is a large model) got 12 tokens&#x2F;second and very coherent output. Great demonstration.
    2. sourweasel · · focus · HN ↗
      I would recommend adding the MiniCPM5-1B model. Surprisingly coherent for a 1B model and performs well. On my pixel 9 I get 33 tok&#x2F;s on CPU. On GPU I get about 26 tok&#x2F;s but prefill jumps to nearly 500.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.