‹ BackHN Continuity

Thread

How accurately calibrated is Jev?

53 points · 22 comments · dblack12705

  1. cannedbread · · focus · HN ↗
    IMO by asking Jev underspecified questions like this, you&#x27;re essentially using it as a random number generator (similar to the dice example). On actual NLP problems (including ones with uncertainty under human review) it does appear to be well calibrated: <a href="https:&#x2F;&#x2F;leonardgrazian.com&#x2F;blog&#x2F;jev-calibration&#x2F;" rel="nofollow">https:&#x2F;&#x2F;leonardgrazian.com&#x2F;blog&#x2F;jev-calibration&#x2F;
    1. unholiness · · focus · HN ↗
      In the dice example[0] and in this one, the expected output is a distribution e.g. &quot;1: 16.7%, 2: 16.7%...&quot;, not a random generation. I agree that in implementation, the architecture is not designed to accurately calculate or incorporate any known uncertainties like these, but the task itself is completely reasonable and arguably trivial.

      IMO the lesson here is: even in trivial cases, Jev&#x27;s outputs are just ~reasonableness scores which do not correspond to actual probabilities. They should not be treated as actual probabilities without careful calibration and plenty of meta-uncertainty about how well that calibration extrapolates.

      The problem is, most of the value proposition of Jev is that it gives you the probabilities without doing that, which it doesn&#x27;t.

      [0] <a href="https:&#x2F;&#x2F;kantahayashiai.github.io&#x2F;posts&#x2F;jev-does-not-play-dice&#x2F;" rel="nofollow">https:&#x2F;&#x2F;kantahayashiai.github.io&#x2F;posts&#x2F;jev-does-not-play-dic...

      1. cannedbread · · focus · HN ↗
        I think you hit on my core criticism of this and similar analyses: LLMs are a tool, and using them to draw from a distribution is misuse of the tool. It&#x27;s not what it&#x27;s designed&#x2F;tuned for. Asking follow ups like &quot;is Jev calibrated when I mis-use it?&quot; is asking the wrong question. The right question would be, &quot;is Jev calibrated for expected use cases?&quot;. And based on some initial exploration, I do think Jev is calibrated for common natural language questions
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.