‹ BackHN Continuity

Thread

Show HN: Maple-Preview – Ternary 20B MoE running at 120 tok/s on a iPhone

169 points · 52 comments · edwardbzhang

  1. walrus01 · · focus · HN ↗
    I wish that "small" LLMs would stop being confidently very incorrect. Admittedly this is a bit of an intentionally esoteric test, but the confident way in which it presents a totally incorrect answer is a bit concerning.

    "please write 250 words on the etymology and history of the word schlong"

    <a href="https:&#x2F;&#x2F;pastes.io&#x2F;uhshFgn4" rel="nofollow">https:&#x2F;&#x2F;pastes.io&#x2F;uhshFgn4

    The actual origin of the word is from middle high German and Yiddish-speaking Ashkenazi Jewish communities.

    For comparison qwen 3.6 35B A3B does perfect on this and will give a solid description of the word&#x27;s real origins and how it has made it into casual profanity&#x2F;vulgarity as used in US English, and even mentions specific stand-up comedians and famous public figures of specific ethnic&#x2F;religious origin in the US NE who introduced it into wider use.

    Ask it for something that&#x27;s not a narrow niche scientific or technical field, but something that would be less common to make it into a 20B size model, and see just how it does.

    chat test link: <a href="https:&#x2F;&#x2F;chat.deepgrove.ai&#x2F;">https:&#x2F;&#x2F;chat.deepgrove.ai&#x2F;

    1. sajithdilshan · · focus · HN ↗
      I think for smaller models, they need to be more defensive on unknown information and frontier model level tool calling capabilities.

      LLMs are kind of a compact knowledge box of its training data and it&#x27;s understandable it would not have information about every topic and in that case just do a web search or a proper tool invocation to get the data and then synthesize.

      1. miohtama · · focus · HN ↗
        You need a small LLM that can reason very well and use tools like web search very well.

        No one is going to compress human knowledge into few bits.

        1. sajithdilshan · · focus · HN ↗
          Not at the moment, but who knows what types of models and storage types we would have in 100 years.
          1. jasonjmcghee · · focus · HN ↗
            It&#x27;s far more likely basic consumer devices will advance such that much larger models efficiently run than finding novel ways to compress all of human experience to fit on today&#x27;s mobile hardware.

            Information can only be compressed so much

            1. wolttam · · focus · HN ↗
              We’re already seeing incredible knowledge compression out of the models the GP mentioned like Qwen 35B-A3B, which feels well within the realm of “runs on a phone” in the next handful of years.

              And by then we’ll probably have been further surprised by just how much information and capacity for reasoning can be crammed into a few gigs of weights. Models just keep getting better for a given size, it’ll be interesting to see where the limit of that is.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.