Show HN: Maple-Preview – Ternary 20B MoE running at 120 tok/s on a iPhone
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Show HN: Maple-Preview – Ternary 20B MoE running at 120 tok/s on a iPhone
Unofficial Hacker News client; not affiliated with Y Combinator.
beautiful_apple · · focus · HN ↗
I didn't notice the version difference when first reading the article! So this is a heads up to people like me.
ricardobeat · · focus · HN ↗
walrus01 · · focus · HN ↗
Really basic stuff. But then again, the entire thing was running in <6GB of RAM.
But before anyone says 1-bit bonsai 27B beats anything, please actually run it and ask it some questions about topics you already know the answer to.
While Qwen 3.6 35B A3B in Q8 with full context capability (llama-server in no-mmap mode with 262k context will eat 47GB, so not comparable in size either) knows a great deal. The 35B-A3B can even translate multiple pages of English into Farsi and its Farsi output is not far off the quality of what Google Translate does.
I haven't tested something as badly quantized as 35B A3B Q2 which is somewhere around 12GB on disk. <a href="https://huggingface.co/unsloth/Qwen3.6-35B-A3B-GGUF" rel="nofollow">https://huggingface.co/unsloth/Qwen3.6-35B-A3B-GGUF
cmrx64 · · focus · HN ↗
nxtfari · · focus · HN ↗
swiftcoder · · focus · HN ↗
It's also the aspect of LLMs that degrades fastest with quantisation. You can't reasonably expect accurate knowledge of everything in the world in a few gigabytes.
Which means that small models need to be conditioned to rely more heavily on tools to fetch accurate information, and ideally not try and generate facts purely based on their (extremely lossy) internal knowledge
seba_dos1 · · focus · HN ↗