It seems to me that Jev cannot be usefully calibrated out of the box, in a fairly strict sense. Suppose I have a pull request and the state is the title, description, etc. The question is “Will it be merged?” (it doesn’t really matter whether it’s a choice or a “noul” [0]).
Now consider that two different projects may have radically different criteria for accepting a PR. And the two projects may have different probabilities for acceptance of a random PR from the distribution of PRs they get (i.e. the overall fraction of PRs that are accepted). So what could Jev possibly return that is “calibrated” for both? It doesn’t even have an “I have no idea” option because the output schema cannot distinguish between “I am confident that there is a 50% probability that the answer is yet conditioned on the state” and “there is no useful information contained in the state that I can extract and therefore you should assume that you posterior distribution is the same as your prior”.
For fun, I gave Jev some irrelevant state and asked it various questions for which the state was useless (I picked sporting outcomes), and it was 0-for-3 at giving yes/no probabilities that were particularly close to the obviously correct no-information answers or close to 0.5 in cases where the prior was far from 0.5.
This is a silly test, but I’ve personally encountered genuine production situations where the best classifier available (or at least the best one available at any cost remotely close to what it was worth) was, drumroll please, a constant. But it was a calibrated constant: we measured it! And there is no way to feed this sort of information to Jev. (Yes, I tried it. Even literally stating the distribution in the state does not work well, although it does appear to have some effect on the outputs.)
[0] Is “noul” even a word? I know what a binary classifier is…
amluto · · focus · HN ↗
Now consider that two different projects may have radically different criteria for accepting a PR. And the two projects may have different probabilities for acceptance of a random PR from the distribution of PRs they get (i.e. the overall fraction of PRs that are accepted). So what could Jev possibly return that is “calibrated” for both? It doesn’t even have an “I have no idea” option because the output schema cannot distinguish between “I am confident that there is a 50% probability that the answer is yet conditioned on the state” and “there is no useful information contained in the state that I can extract and therefore you should assume that you posterior distribution is the same as your prior”.
For fun, I gave Jev some irrelevant state and asked it various questions for which the state was useless (I picked sporting outcomes), and it was 0-for-3 at giving yes/no probabilities that were particularly close to the obviously correct no-information answers or close to 0.5 in cases where the prior was far from 0.5.
This is a silly test, but I’ve personally encountered genuine production situations where the best classifier available (or at least the best one available at any cost remotely close to what it was worth) was, drumroll please, a constant. But it was a calibrated constant: we measured it! And there is no way to feed this sort of information to Jev. (Yes, I tried it. Even literally stating the distribution in the state does not work well, although it does appear to have some effect on the outputs.)
[0] Is “noul” even a word? I know what a binary classifier is…