‹ BackHN Continuity

Thread

Contrastive Language Models

176 points · 59 comments · erichocean

  1. sdan · · focus · HN ↗
    I tried running this on a H100 and got 190ms compared to Jev's 170ms. Maybe I set it up wrong?
    1. kevmo314 · · focus · HN ↗
      It seems like caching has a huge contribution towards the low latency numbers they report. With a single request and no cache it appears to be quite slow, roughly in line with what you observed.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.