‹ BackHN Continuity

Thread

Decision models like Jev don't beat LLM-as-a-judge or traditional classifiers

96 points · 34 comments · tomncooper

  1. AnthusAI · · focus · HN ↗
    That was a pretty simple task they gave it, and sure you can use BERT with sequence classification for simple classification tasks.

    In our benchmarks, Jev did a LOT better at multi-step reasoning tasks than any open decision model we have tested so far, and it was also better than GLiDE which was specifically designed for that kind of task. And also better than Luna. On accuracy and also confidence calibration but also time and cost.

    <a href="https:&#x2F;&#x2F;hard-decisions.anth.us&#x2F;models&#x2F;" rel="nofollow">https:&#x2F;&#x2F;hard-decisions.anth.us&#x2F;models&#x2F;

    1. meander_water · · focus · HN ↗
      Just looking through your results, seems like gpt-6 luna was run with reasoning:off for a lot (all?) results. Seems like an unfair comparison.
      1. zurfer · · focus · HN ↗
        Not if you care about latency and cost. Reasoning is slow and expensive.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.