‹ BackHN Continuity

Thread

Decision models like Jev don't beat LLM-as-a-judge or traditional classifiers

114 points · 43 comments · tomncooper

  1. segmondy · · focus · HN ↗
    duh, this is not news. (general, fast and cheap) before decision models, you could pick only 2.

    LLM as judges - generalized, but too slow. If you had to make millions of classifications a day, this will be the wrong approach. you won't/shouldn't use LLM to classify spam/no spam. hot dog/or something.

    traditional classifiers, very specific 1 trick pony, super fast and cheap once built. If you need to make tons and tons of classifications, this would be the approach. but if you wanted a classifier right now for a novel problem, you need an expert to curate data, train and deploy.

    decision models/jev - are generic, you can throw them at most generic classification problems, and they are good enough. it's a fine balance between general, fast and cheap. you get all 3

    1. lostmsu · · focus · HN ↗
      Is Jev any faster, cheaper, or more precise than medium LLMs like Luna 6?
      1. IanCal · · focus · HN ↗
        Luna 6 is 10c per million input tokens and charges 5x that for output. Not sure if it’s still true but it used to be the case that structured outputs took time to process and cache which is relevant if the structure changes. It doesn’t give a percentage you can use for thresholds, and I’d want to know if Jen treats the questions as independent (they aren’t with luna, order of questions will change the result).

        Jev is 4.2c/m tokens in and free out.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.