‹ BackHN Continuity

Thread

Decision models like Jev don't beat LLM-as-a-judge or traditional classifiers

96 points · 34 comments · tomncooper

  1. NeumannGod · · focus · HN ↗
    This is a false comparison. What Jev does is fundamentally different from what LLMs are doing as a judge.

    For a quick primer, Jev is able to provide confidence scores on its classification, i.e. it is able to calibrate how well it is able to predict. Being able to predict in a distribution is different from being able to calibrate confidence of the predictions which should happen from the question or domain distribution from which the decisions are predicted - being able to do that is tough and is not same as using LLMs logit probabilities which are predictions in the vocab space. Though both are loosely correlated and might converge as LLMs keep getting better, the former is a much stronger decision-making signal than the latter. Jev not beating LLM-as-a-judge might be due to various other reasons such as world knowledge etc, but Jev as a concept will always provide more reliable decisions / outputs than LLM-as-a-judge giving a scalar score.

    1. kwinkunks · · focus · HN ↗
      I agree that there are differences between Jev and LLM-as-a-judge (e.g. Almeida&#x27;s assertion that LLM probabilities have been irrevocably biased by RLHF), but I am not sure about your description of &#x27;confidence&#x27;. Perhaps I misinterpret you or the docs, but I understand it as simply being computed from the probability distribution: <a href="https:&#x2F;&#x2F;docs.typesafe.ai&#x2F;confidence#how-confidence-is-calculated" rel="nofollow">https:&#x2F;&#x2F;docs.typesafe.ai&#x2F;confidence#how-confidence-is-calcul...
    2. [deleted] · · focus · HN ↗

      [deleted]

    3. yieldcrv · · focus · HN ↗
      for the uninitiated:

      LLM’s are not able to give confidence scores, they make them up.

      Your AI driven app is making that up. Your product manager and executive team’s demand for confidence in the UI is a totally fictional cosmetic telling them nothing. Your company sold bullshit confidence to your clients.

      I’ve done this for many organizations that “formed a new team to work with the CTO on their AI strategy”, and the trappings are the same

      You can have an LLM tell you how much of a schema it was able to get information about. And derive a “confidence” or level of compliance from the completeness of the schema

      But this is layers upon layers of cruft that a classification model wouldn’t need

      1. andy99 · · focus · HN ↗
        Softmax over logits doesn’t give calibrated probabilities either as a rule. I don’t want to comment specifically on Jev but as a rule it’s very hard to get good calibration because it’s somewhat in tension with minimizing training loss for neural networks, e.g. Guo et al (2017) <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;1706.04599" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;1706.04599

        While I know there are ways to improve calibration, I’d personally want to see a lot of evidence the probabilities were actually more meaningful before trusting them. I agree of course that asking an LLM to provide a confidence estimate is meaningless.

    4. ViscountPenguin · · focus · HN ↗
      Fine-tuning language models for calibrated probability predictions has been a thing since the Bert era though, that&#x27;s really nothing special.
    5. fifilura · · focus · HN ↗
      I don&#x27;t know about Jev, but in my experience, 90% of the time, the confidence score such machine outputs (e.g. simplest case a kalman filter) is bogus because the model is wrong. Or not even wrong, but just not perfect.

      Mathematicians build this, and they love the beauty of it so much that they loose touch of reality.

    6. jvanderbot · · focus · HN ↗
      Partially so.

      The benefits of Jev, and the reason I&#x27;m excited about them:

      * Pre-trained&#x2F;tuned to provide only structured output

      * Small, fast, essentially trivial to locally run - this alone makes them an actual candidate for real-world planning systems

      * Their confidence scores are useful, even if those scores are not correct&#x2F;true probabilities

      For me and my work, they look like nearly turn-key, tiny, fast, locally-hostable models that can manage state transitions in a deeply autonomous system.

      LLM-as judge requires, well, what we know as a full LLM. A fully trained classifier model requires fully training - something you cannot &#x2F; won&#x27;t do for a embedded autonomous planning system (by contradiction - if this worked, we&#x27;d have used it everywhere already!). Totally unfair comparison in my mind.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.