‹ BackHN Continuity

Thread

Decision models like Jev don't beat LLM-as-a-judge or traditional classifiers

96 points · 34 comments · tomncooper

  1. segmondy · · focus · HN ↗
    duh, this is not news. (general, fast and cheap) before decision models, you could pick only 2.

    LLM as judges - generalized, but too slow. If you had to make millions of classifications a day, this will be the wrong approach. you won't/shouldn't use LLM to classify spam/no spam. hot dog/or something.

    traditional classifiers, very specific 1 trick pony, super fast and cheap once built. If you need to make tons and tons of classifications, this would be the approach. but if you wanted a classifier right now for a novel problem, you need an expert to curate data, train and deploy.

    decision models/jev - are generic, you can throw them at most generic classification problems, and they are good enough. it's a fine balance between general, fast and cheap. you get all 3

    1. lostmsu · · focus · HN ↗
      Is Jev any faster, cheaper, or more precise than medium LLMs like Luna 6?
      1. pokeapallascat · · focus · HN ↗
        not really
        1. nicce · · focus · HN ↗
          I don't think Luna is fast enough by any means
      2. IanCal · · focus · HN ↗
        Luna 6 is 10c per million input tokens and charges 5x that for output. Not sure if it’s still true but it used to be the case that structured outputs took time to process and cache which is relevant if the structure changes. It doesn’t give a percentage you can use for thresholds, and I’d want to know if Jen treats the questions as independent (they aren’t with luna, order of questions will change the result).

        Jev is 4.2c/m tokens in and free out.

      3. ShinTakuya · · focus · HN ↗
        Yes. Not as fast/cheap as Typesafe claims, but lots of benchmarks suggest around 3-4 times cheaper, and 7-14 times faster.

        - <a href="https:&#x2F;&#x2F;www.ml6.eu&#x2F;en&#x2F;blog&#x2F;jev-vs-gpt-6-luna-vs-bert-text-classification" rel="nofollow">https:&#x2F;&#x2F;www.ml6.eu&#x2F;en&#x2F;blog&#x2F;jev-vs-gpt-6-luna-vs-bert-text-cl... - <a href="https:&#x2F;&#x2F;tessl.io&#x2F;blog&#x2F;jev-is-136x-faster-and-27x-cheaper-than-gpt-luna-6-for-tessl-verifiers-try-it-yourself" rel="nofollow">https:&#x2F;&#x2F;tessl.io&#x2F;blog&#x2F;jev-is-136x-faster-and-27x-cheaper-tha... - <a href="https:&#x2F;&#x2F;x.com&#x2F;fazxes&#x2F;status&#x2F;2100300097695232164" rel="nofollow">https:&#x2F;&#x2F;x.com&#x2F;fazxes&#x2F;status&#x2F;2100300097695232164 (this last one is Luna 5.6 but that isn&#x27;t too different from 6 besides accuracy and cost)

        1. [deleted] · · focus · HN ↗

          [deleted]

        2. CharlieDigital · · focus · HN ↗
          If you can do it offline, batching solves this. We used a test dataset from CFPB and at n=20, it was 1.6x faster and 1.2x more expensive with no statistically meaningful accuracy dropoff. Did not tune for batch size, but it&#x27;s possible that we could get to 30 and see better perf.
    2. [deleted] · · focus · HN ↗

      [deleted]

    3. xfalcox · · focus · HN ↗
      Doesn&#x27;t the article covers the speed part by showing that Qwen 3.6 35A3B has lower latency and same accuracy?
      1. jvanderbot · · focus · HN ↗
        Laya, a local Jev alternative, is a 400Million parameter model. It can play doom.

        Find me a another class of 0.4 B model that can handle structured output decision problems with the same latency and accuracy as Qwen 3.6.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.