‹ BackHN Continuity

Thread

How OpenAI Used Its Own LLMs to Design Its Jalapeño Chip

206 points · 139 comments · maxall4

  1. globnomulous · · focus · HN ↗
    > Jalapeño can reduce end-to-end latency (the time between prompt to last token) by up to 3.6 times

    I'm never sure what on earth this kind of impressionistic math is supposed to tell me. Is the comparison between 4.6 and 1.0? 3.6 and 1.0? Clearly the comparison isn't supposed to be 1.0 and -2.6, even though that's what the words literally mean. I can't be the only person who finds this infuriating and distracting. These numbers shouldn't be impressionistic. They should be precise. That this is an article on spectrum.ieee.org makes the imprecision all the stranger. I'd expect their readershipt to care, for instance, about what's even being measured. Is this the geometric mean of something? The arithmetic mean? And what latency has improved?

    1. perching_aix · · focus · HN ↗
      It's... written right there? Like what?

      Suppose you send in your marvelous prompt and hit Enter.

      Machine churns for 18 seconds, types out a "reply", then yields back control.

      18 / 3.6 = 5

      So now the machine will only churn for 5 seconds before yielding back control.

      This is confusing how exactly?

      Why would an "up to" figure be a mean, or a geometric mean? It's clearly a max, that's why it's called "up to"...

      Am I missing something?

      1. sebzim4500 · · focus · HN ↗
        You're missing the part of your brain that wants to be pedantic more than it wants to understand someone.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.