‹ BackHN Continuity

Thread

Ember-1

589 points · 249 comments · gmays

  1. nico · · focus · HN ↗
    > The problem: thinking models think too much

    This is partly the appeal of Jev et al; having a quick model for simple tasks, that doesn’t require that much thinking

    It’s amazing all the workflows that models like that can unlock. And yes, classifiers and other ML models have been around for a while for these types of tasks, but Jev has made it easy and cheap to play and experiment. This in turn, is incentivizing people to try them for a bunch of stuff, unlocking creativity and producing a lot of new cool (and eventually potentially very useful) applications

    1. elcomet · · focus · HN ↗
      Why not using a cheap LLM with thinking completely disabled ? I don't think it will be much more expensive than jev.
      1. ssivark · · focus · HN ↗
        LLM inference has two very different regimes of work: prefill & decode. You can think of the former roughly as processing a pre-specified prompt, and the latter as sequential processing (auto-regressive token generation) eg. "chain of thought". The latter is very important for LLMs and cannot be ignored; it deeply influences infra design, even necessitates copious amounts of high-bandwidth memory. Jev-like models can ignore the latter and therefore optimize much better for the former, consequently operating at both better cost and latency.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.