‹ BackHN Continuity

Thread

Jev Can't Be Calibrated

65 points · 61 comments · alexmolas

  1. kantahayashi · · focus · HN ↗
    I tested Jev with a fair die 400 times without telling it the die result. The true probability of face 1 is 1/6, but Jev always chose face 1 and the probability it returned was about 83%. I also tested with a fair coin 200 times and got 0.92 probability.

    I did several tests and I think Jev is good at problems with a correct answer but weak at problems about actual probabilities whose answers can't be known at all.

    Write-up: &quot;Jev Does Not Play Dice&quot; <a href="https:&#x2F;&#x2F;kantahayashiai.github.io&#x2F;posts&#x2F;jev-does-not-play-dice&#x2F;" rel="nofollow">https:&#x2F;&#x2F;kantahayashiai.github.io&#x2F;posts&#x2F;jev-does-not-play-dic...

    1. seizethecheese · · focus · HN ↗
      Maybe I’m confused here, but it’s perfectly reasonable to just guess the same dice roll every time right?
      1. alexmolas · · focus · HN ↗
        I don&#x27;t know if it&#x27;s reasonable. What it isn&#x27;t is calibrated.
      2. kantahayashi · · focus · HN ↗
        Yes. There&#x27;s no problem with choosing the same face every time. The problem is the probability it attached to the choice. Jev gave face 1 an 83% probability while the true probability is 1&#x2F;6.
        1. sshine · · focus · HN ↗
          Do you provide Jev that the probability is 1&#x2F;6 and yet it gives back a probability that is way off?
          1. kantahayashi · · focus · HN ↗
            Yes. For example, one of the prompts said &quot;The die is unbiased: each of the six faces has probability exactly 1&#x2F;6.&quot;
        2. seizethecheese · · focus · HN ↗
          Okay, I see, you&#x27;re expecting Jev to properly give 1&#x2F;6 probability for each option. This is different from my intuition of how LLMs work, where their probabilities don&#x27;t really work like this (I would expect LLM to also do something like 0.83 for 1).
          1. [deleted] · · focus · HN ↗

            [deleted]

          2. kantahayashi · · focus · HN ↗
            That&#x27;s right. It&#x27;s normal behavior of LLMs. But what matters is TypeSafe argues it&#x27;s different exactly on this point. The selling point of Jev is &quot;calibrated probabilities&quot;, so I checked it on probability problems.
          3. [deleted] · · focus · HN ↗

            [deleted]

          4. maayank · · focus · HN ↗
            Jev and LLMs give other promises. Jev&#x27;s RLCD training aims to make its probabilities calibrated such that given many cases where it assigns label Y about X% probability, Y should be the correct label about X% of the time.
          5. NeutralCrane · · focus · HN ↗
            One of the entire value props for Jev is that is exactly how it is supposed to work. That is one of the big claimed advantages over regular LLMs.
        3. [deleted] · · focus · HN ↗

          [deleted]

      3. dgritsko · · focus · HN ↗
        Reminds me of this... <a href="https:&#x2F;&#x2F;xkcd.com&#x2F;221&#x2F;" rel="nofollow">https:&#x2F;&#x2F;xkcd.com&#x2F;221&#x2F;
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.