‹ BackHN Continuity

Thread

Show HN: Maple-Preview – Ternary 20B MoE running at 120 tok/s on a iPhone

169 points · 52 comments · edwardbzhang

  1. kamranjon · · focus · HN ↗
    “Current approaches to low precision primarily focus on converting models trained in full precision to lower bitwidths. We view this as fundamentally the wrong approach…”

    Very excited to see how it performs, I’ve been a bit skeptical of the efficacy of converting existing models - really cool to see one trained from scratch in the ternary format.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.