‹ BackHN Continuity

Thread

Decision models like Jev don't beat LLM-as-a-judge or traditional classifiers

96 points · 34 comments · tomncooper

  1. aidiveyt · · focus · HN ↗
    The block/allow framing is the part I'd push on. I had a batch where automated checks passed all 99 outputs and reading each one by hand found 8 broken. The scorer only catches the failure modes its rubric already names, and a two-option guardrail bench inherits that ceiling whichever model sits behind it.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.