‹ BackHN Continuity

Thread

Decision models like Jev don't beat LLM-as-a-judge or traditional classifiers

114 points · 43 comments · tomncooper

  1. NeumannGod · · focus · HN ↗
    This is a false comparison. What Jev does is fundamentally different from what LLMs are doing as a judge.

    For a quick primer, Jev is able to provide confidence scores on its classification, i.e. it is able to calibrate how well it is able to predict. Being able to predict in a distribution is different from being able to calibrate confidence of the predictions which should happen from the question or domain distribution from which the decisions are predicted - being able to do that is tough and is not same as using LLMs logit probabilities which are predictions in the vocab space. Though both are loosely correlated and might converge as LLMs keep getting better, the former is a much stronger decision-making signal than the latter. Jev not beating LLM-as-a-judge might be due to various other reasons such as world knowledge etc, but Jev as a concept will always provide more reliable decisions / outputs than LLM-as-a-judge giving a scalar score.

    1. fifilura · · focus · HN ↗
      I don't know about Jev, but in my experience, 90% of the time, the confidence score such machine outputs (e.g. simplest case a kalman filter) is bogus because the model is wrong. Or not even wrong, but just not perfect.

      Mathematicians build this, and they love the beauty of it so much that they loose touch of reality.

      1. bunderbunder · · focus · HN ↗
        Usably good calibration is hard. It's hard with logistic regression, it's even harder with linear support vector machines, and it makes me question my life choices with non-linear models.

        It's not necessarily because the model is wrong. I've had trouble getting useful calibration out of models with an F1 of 0.9. The fundamental problem is twofold. First, it turns out that [0.0, 1.0] is a much larger set than {0, 1}. Second, the kinds of use cases where you care about calibration tend to be fussy and demanding.

        As a fun anecdote, I got a chance to ask the person who invented the method scikit-learn uses for calibrated SVMs if he had any advice, and his answer was basically, "good luck."

        All that said, sometimes you don't actually need calibration; you just need decent ranking. "Items with a score of 0.9 should be more likely to be positive than ones with a score of 0.5," is an easier requirement than "90% of items with a score of 0.9 should be positive." But I've had trouble getting that out of highly non-linear neural models, too, because they oftentimes produce results where the relationship between score and probability of being in the positive class is not even remotely monotonic.

        But I haven't poked at Jev like this either, so I do have to allow that maybe they've found the secret sauce.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.