I tested Jev with a fair die 400 times without telling it the die result. The true probability of face 1 is 1/6, but Jev always chose face 1 and the probability it returned was about 83%. I also tested with a fair coin 200 times and got 0.92 probability.
I did several tests and I think Jev is good at problems with a correct answer but weak at problems about actual probabilities whose answers can't be known at all.
Write-up: "Jev Does Not Play Dice"
<a href="https://kantahayashiai.github.io/posts/jev-does-not-play-dice/" rel="nofollow">https://kantahayashiai.github.io/posts/jev-does-not-play-dic...
But "problems about actual probabilities whose answers can't be known at all" are exactly the problems where calibration is important. Since calibration is one of the big claims about Jev I'd expect it to perform well in these problems.
I agree. I think it's odd behavior too. Jev should be good at actual probability problems given the phrase "calibrated probabilities" TypeSafe uses for Jev. Maybe the reason is the data used in their training method (RLCD). If all the data consists of problems with a correct answer, I think this kind of odd behavior could happen.
kantahayashi · · focus · HN ↗
I did several tests and I think Jev is good at problems with a correct answer but weak at problems about actual probabilities whose answers can't be known at all.
Write-up: "Jev Does Not Play Dice" <a href="https://kantahayashiai.github.io/posts/jev-does-not-play-dice/" rel="nofollow">https://kantahayashiai.github.io/posts/jev-does-not-play-dic...
alexmolas · · focus · HN ↗
kantahayashi · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]