‹ BackHN Continuity

Thread

Show HN: JevBench, a reproducible benchmark for typed decision models

154 points · 39 comments · florianstandhar

  1. ks2048 · · focus · HN ↗
    I was trying to figure out what exactly the tests here are. I guess I found some of the questions (here: <a href="https:&#x2F;&#x2F;github.com&#x2F;fstandhartinger&#x2F;jevbench&#x2F;blob&#x2F;main&#x2F;datasets&#x2F;public&#x2F;easy.jsonl" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;fstandhartinger&#x2F;jevbench&#x2F;blob&#x2F;main&#x2F;datase...)

    e.g.,

      &quot;instructions&quot;: &quot;Which intent does the user&#x27;s message express?&quot;,
      &quot;labels&quot;:[&quot;set_alarm&quot;, &quot;play_music&quot;, &quot;weather&quot;, &quot;send_message&quot;, &quot;turn_off_lights&quot;],
      &quot;state&quot;: &quot;Play some Taylor Swift.&quot;,
      &quot;expected&quot;: &quot;play_music&quot;
    1. florianstandhar · · focus · HN ↗

      [dead]

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.