‹ BackHN Continuity

Thread

Jev Can't Be Calibrated

65 points · 61 comments · alexmolas

  1. kantahayashi · · focus · HN ↗
    I tested Jev with a fair die 400 times without telling it the die result. The true probability of face 1 is 1/6, but Jev always chose face 1 and the probability it returned was about 83%. I also tested with a fair coin 200 times and got 0.92 probability.

    I did several tests and I think Jev is good at problems with a correct answer but weak at problems about actual probabilities whose answers can't be known at all.

    Write-up: &quot;Jev Does Not Play Dice&quot; <a href="https:&#x2F;&#x2F;kantahayashiai.github.io&#x2F;posts&#x2F;jev-does-not-play-dice&#x2F;" rel="nofollow">https:&#x2F;&#x2F;kantahayashiai.github.io&#x2F;posts&#x2F;jev-does-not-play-dic...

    1. seizethecheese · · focus · HN ↗
      Maybe I’m confused here, but it’s perfectly reasonable to just guess the same dice roll every time right?
      1. kantahayashi · · focus · HN ↗
        Yes. There&#x27;s no problem with choosing the same face every time. The problem is the probability it attached to the choice. Jev gave face 1 an 83% probability while the true probability is 1&#x2F;6.
        1. seizethecheese · · focus · HN ↗
          Okay, I see, you&#x27;re expecting Jev to properly give 1&#x2F;6 probability for each option. This is different from my intuition of how LLMs work, where their probabilities don&#x27;t really work like this (I would expect LLM to also do something like 0.83 for 1).
          1. [deleted] · · focus · HN ↗

            [deleted]

          2. kantahayashi · · focus · HN ↗
            That&#x27;s right. It&#x27;s normal behavior of LLMs. But what matters is TypeSafe argues it&#x27;s different exactly on this point. The selling point of Jev is &quot;calibrated probabilities&quot;, so I checked it on probability problems.
          3. [deleted] · · focus · HN ↗

            [deleted]

          4. maayank · · focus · HN ↗
            Jev and LLMs give other promises. Jev&#x27;s RLCD training aims to make its probabilities calibrated such that given many cases where it assigns label Y about X% probability, Y should be the correct label about X% of the time.
          5. NeutralCrane · · focus · HN ↗
            One of the entire value props for Jev is that is exactly how it is supposed to work. That is one of the big claimed advantages over regular LLMs.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.