Show HN: Maple-Preview – Ternary 20B MoE running at 120 tok/s on a iPhone
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Show HN: Maple-Preview – Ternary 20B MoE running at 120 tok/s on a iPhone
Unofficial Hacker News client; not affiliated with Y Combinator.
walrus01 · · focus · HN ↗
"please write 250 words on the etymology and history of the word schlong"
<a href="https://pastes.io/uhshFgn4" rel="nofollow">https://pastes.io/uhshFgn4
The actual origin of the word is from middle high German and Yiddish-speaking Ashkenazi Jewish communities.
For comparison qwen 3.6 35B A3B does perfect on this and will give a solid description of the word's real origins and how it has made it into casual profanity/vulgarity as used in US English, and even mentions specific stand-up comedians and famous public figures of specific ethnic/religious origin in the US NE who introduced it into wider use.
Ask it for something that's not a narrow niche scientific or technical field, but something that would be less common to make it into a 20B size model, and see just how it does.
chat test link: <a href="https://chat.deepgrove.ai/">https://chat.deepgrove.ai/
brainless · · focus · HN ↗
walrus01 · · focus · HN ↗
It's also something I've seen has great results with esoteric individual pieces of knowledge that works fine in a Q6 or Q8 quantized LLM but breaks down in a bad way at worse quantization.
sznio · · focus · HN ↗
20b parameters * 1.5 bits per parameter is just 30 billion bits, about 3.75gb
a full 20b fp16 is about 40GB.
I find it weird how a smaller model still produces decent text, except it bullshits all the way.