‹ BackHN Continuity

Thread

Decision models like Jev don't beat LLM-as-a-judge or traditional classifiers

138 points · 49 comments · tomncooper

  1. NeumannGod · · focus · HN ↗
    This is a false comparison. What Jev does is fundamentally different from what LLMs are doing as a judge.

    For a quick primer, Jev is able to provide confidence scores on its classification, i.e. it is able to calibrate how well it is able to predict. Being able to predict in a distribution is different from being able to calibrate confidence of the predictions which should happen from the question or domain distribution from which the decisions are predicted - being able to do that is tough and is not same as using LLMs logit probabilities which are predictions in the vocab space. Though both are loosely correlated and might converge as LLMs keep getting better, the former is a much stronger decision-making signal than the latter. Jev not beating LLM-as-a-judge might be due to various other reasons such as world knowledge etc, but Jev as a concept will always provide more reliable decisions / outputs than LLM-as-a-judge giving a scalar score.

    1. jvanderbot · · focus · HN ↗
      Partially so.

      The benefits of Jev, and the reason I'm excited about them:

      * Pre-trained/tuned to provide only structured output

      * Small, fast, essentially trivial to locally run - this alone makes them an actual candidate for real-world planning systems

      * Their confidence scores are useful, even if those scores are not correct/true probabilities

      For me and my work, they look like nearly turn-key, tiny, fast, locally-hostable models that can manage state transitions in a deeply autonomous system.

      LLM-as judge requires, well, what we know as a full LLM. A fully trained classifier model requires fully training - something you cannot / won't do for a embedded autonomous planning system (by contradiction - if this worked, we'd have used it everywhere already!). Totally unfair comparison in my mind.

      1. monocasa · · focus · HN ↗
        With llama.cpp, you can give it an ebnf grammar and constrain the output to any formal grammar you wish.

        And how is the confidence score any different than the softmaxed token probabilities you get out of running every llm?

        1. jvanderbot · · focus · HN ↗
          > And how is the confidence score any different than the softmaxed token probabilities you get out of running every llm?

          You're asking me how a structured output which assigns probabilities to classifications of the input text is different than the next-token weights which are calculated while generating that output?

          One is internal (token prob), one is output generated by that internal mechanism e.g.,

              "{ next_state: { "evade": 0.84, "land": 0.10, "search": 0.01"...}}"
        2. ekidd · · focus · HN ↗
          > With llama.cpp, you can give it an ebnf grammar and constrain the output to any formal grammar you wish.

          With llama-server, you can use basically any GGUF model as a Jev-like classifier. See Pi.dev's codemode and the accompanying llama backend to the "classify" function. The advantages of Jev-like inference are that:

          - You generate 1 token per question in parallel, instead of sequentially generating a bunch of JSON punctuation. So you get much better latency and utilization.

          - You can use a custom "sampler" that sees the raw logprobs for these tokens, including all the tokens that might have been generated. So you can generate some "probability" or "confidence" numbers based on how likely the model was to have generated a different answer.

          So really, no difference? It's just a nice API with some computational efficiencies.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.