‹ BackHN Continuity

Thread

Decision models like Jev don't beat LLM-as-a-judge or traditional classifiers

96 points · 34 comments · tomncooper

  1. deepsquirrelnet · · focus · HN ↗
    BART is quite an old model for this kind of test, and probably not a good very good choice for much these days. I'm working on replicating their benchmark on my own NLI model that targets zero-shot guardrail applications. I don't think it'll beat much larger models, but should give a better baseline for what a crossencoder can do.

    <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;dleemiller&#x2F;crossingguard-nli-l" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;dleemiller&#x2F;crossingguard-nli-l

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.