‹ BackHN Continuity

Thread

Decision models like Jev don't beat LLM-as-a-judge or traditional classifiers

129 points · 46 comments · tomncooper

  1. NeumannGod · · focus · HN ↗
    This is a false comparison. What Jev does is fundamentally different from what LLMs are doing as a judge.

    For a quick primer, Jev is able to provide confidence scores on its classification, i.e. it is able to calibrate how well it is able to predict. Being able to predict in a distribution is different from being able to calibrate confidence of the predictions which should happen from the question or domain distribution from which the decisions are predicted - being able to do that is tough and is not same as using LLMs logit probabilities which are predictions in the vocab space. Though both are loosely correlated and might converge as LLMs keep getting better, the former is a much stronger decision-making signal than the latter. Jev not beating LLM-as-a-judge might be due to various other reasons such as world knowledge etc, but Jev as a concept will always provide more reliable decisions / outputs than LLM-as-a-judge giving a scalar score.

    1. ViscountPenguin · · focus · HN ↗
      Fine-tuning language models for calibrated probability predictions has been a thing since the Bert era though, that's really nothing special.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.