‹ BackHN Continuity

Thread

Tokens too cheap to meter

354 points · 227 comments · teoruiz

  1. breadislove · · focus · HN ↗
    One super important thing missing: Speculative decoding. Things like Dflash(2), Dspark etc. help to do one forwards pass and get 6-7 tokens out of it. (For completeness, the embeddings from the forward pass are passed into a diffusion model which predicts the next tokens, and the model just verifies it (very cheap operation)). So we can produce way more tokens for roughly a similar amount of compute.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.