I tested Jev with a fair die 400 times without telling it the die result. The true probability of face 1 is 1/6, but Jev always chose face 1 and the probability it returned was about 83%. I also tested with a fair coin 200 times and got 0.92 probability.
I did several tests and I think Jev is good at problems with a correct answer but weak at problems about actual probabilities whose answers can't be known at all.
Write-up: "Jev Does Not Play Dice"
<a href="https://kantahayashiai.github.io/posts/jev-does-not-play-dice/" rel="nofollow">https://kantahayashiai.github.io/posts/jev-does-not-play-dic...
Yes. There's no problem with choosing the same face every time. The problem is the probability it attached to the choice. Jev gave face 1 an 83% probability while the true probability is 1/6.
Okay, I see, you're expecting Jev to properly give 1/6 probability for each option. This is different from my intuition of how LLMs work, where their probabilities don't really work like this (I would expect LLM to also do something like 0.83 for 1).
Jev and LLMs give other promises. Jev's RLCD training aims to make its probabilities calibrated such that given many cases where it assigns label Y about X% probability, Y should be the correct label about X% of the time.
kantahayashi · · focus · HN ↗
I did several tests and I think Jev is good at problems with a correct answer but weak at problems about actual probabilities whose answers can't be known at all.
Write-up: "Jev Does Not Play Dice" <a href="https://kantahayashiai.github.io/posts/jev-does-not-play-dice/" rel="nofollow">https://kantahayashiai.github.io/posts/jev-does-not-play-dic...
seizethecheese · · focus · HN ↗
kantahayashi · · focus · HN ↗
seizethecheese · · focus · HN ↗
maayank · · focus · HN ↗