‹ BackHN Continuity

Thread

Faster prompt lookup drafting in llama.cpp

89 points · 12 comments · pptadversary

  1. S0y · · focus · HN ↗
    How much does this speedup inference for the end user in terms of tk/s ?
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.