‹ BackHN Continuity

Thread

Show HN: JevBench, a reproducible benchmark for typed decision models

154 points · 39 comments · florianstandhar

  1. faangguyindia · · focus · HN ↗
    Can you please add our model to benchmark? <a href="https:&#x2F;&#x2F;gambler-relay-us-west1.leo-fish.ts.net&#x2F;demo" rel="nofollow">https:&#x2F;&#x2F;gambler-relay-us-west1.leo-fish.ts.net&#x2F;demo
    1. florianstandhar · · focus · HN ↗
      Happy to — please open an issue at <a href="https:&#x2F;&#x2F;github.com&#x2F;fstandhartinger&#x2F;jevbench&#x2F;issues" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;fstandhartinger&#x2F;jevbench&#x2F;issues with an endpoint or runnable code, the model and licence, and whether it was trained on the public items. Every entrant runs through the same harness, including the sealed set. I definitely prefer open models I can run locally, because sending our private held-out set of tasks to an external API produces some headaches on ur side and basically means we have to rotate the test set a lot, to avoid contaimination. We can call API hosted models though, we&#x27;ll just flag the leaderboard entries appropriately.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.