‹ BackHN Continuity

Thread

Decision models like Jev don't beat LLM-as-a-judge or traditional classifiers

96 points · 34 comments · tomncooper

  1. segmondy · · focus · HN ↗
    duh, this is not news. (general, fast and cheap) before decision models, you could pick only 2.

    LLM as judges - generalized, but too slow. If you had to make millions of classifications a day, this will be the wrong approach. you won't/shouldn't use LLM to classify spam/no spam. hot dog/or something.

    traditional classifiers, very specific 1 trick pony, super fast and cheap once built. If you need to make tons and tons of classifications, this would be the approach. but if you wanted a classifier right now for a novel problem, you need an expert to curate data, train and deploy.

    decision models/jev - are generic, you can throw them at most generic classification problems, and they are good enough. it's a fine balance between general, fast and cheap. you get all 3

    1. lostmsu · · focus · HN ↗
      Is Jev any faster, cheaper, or more precise than medium LLMs like Luna 6?
      1. ShinTakuya · · focus · HN ↗
        Yes. Not as fast/cheap as Typesafe claims, but lots of benchmarks suggest around 3-4 times cheaper, and 7-14 times faster.

        - <a href="https:&#x2F;&#x2F;www.ml6.eu&#x2F;en&#x2F;blog&#x2F;jev-vs-gpt-6-luna-vs-bert-text-classification" rel="nofollow">https:&#x2F;&#x2F;www.ml6.eu&#x2F;en&#x2F;blog&#x2F;jev-vs-gpt-6-luna-vs-bert-text-cl... - <a href="https:&#x2F;&#x2F;tessl.io&#x2F;blog&#x2F;jev-is-136x-faster-and-27x-cheaper-than-gpt-luna-6-for-tessl-verifiers-try-it-yourself" rel="nofollow">https:&#x2F;&#x2F;tessl.io&#x2F;blog&#x2F;jev-is-136x-faster-and-27x-cheaper-tha... - <a href="https:&#x2F;&#x2F;x.com&#x2F;fazxes&#x2F;status&#x2F;2100300097695232164" rel="nofollow">https:&#x2F;&#x2F;x.com&#x2F;fazxes&#x2F;status&#x2F;2100300097695232164 (this last one is Luna 5.6 but that isn&#x27;t too different from 6 besides accuracy and cost)

        1. CharlieDigital · · focus · HN ↗
          If you can do it offline, batching solves this. We used a test dataset from CFPB and at n=20, it was 1.6x faster and 1.2x more expensive with no statistically meaningful accuracy dropoff. Did not tune for batch size, but it&#x27;s possible that we could get to 30 and see better perf.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.