‹ BackHN Continuity

Thread

Show HN: JevBench, a reproducible benchmark for typed decision models

154 points · 39 comments · florianstandhar

  1. pushpendraw · · focus · HN ↗
    the slop detector giving 86% confidence on a keysmash is the real finding here, not the leaderboard score. confident and wrong is worse than an LLM that just hedges.
    1. TheWayWithin · · focus · HN ↗

      [dead]

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.