‹ BackHN Continuity

Thread

Show HN: JevBench, a reproducible benchmark for typed decision models

154 points · 39 comments · florianstandhar

  1. swyx · · focus · HN ↗
    jev ceo on why he eschewed benchmarking: <a href="https:&#x2F;&#x2F;www.latent.space&#x2F;i&#x2F;216783460&#x2F;privacy-benchmarking-and-trusting-intelligence" rel="nofollow">https:&#x2F;&#x2F;www.latent.space&#x2F;i&#x2F;216783460&#x2F;privacy-benchmarking-an...
    1. jldugger · · focus · HN ↗
      Interesting; was curious how this didn&#x27;t fall into trouble with ToS. Apparently the &quot;no benchmarks&quot; clause was intended for &quot;limited preview&quot; audiences and didn&#x27;t get removed at launch on accident.
      1. Maxious · · focus · HN ↗
        But also no big deal because &quot;I’m extremely anti-public benchmarks.&quot;
    2. arbot360 · · focus · HN ↗
      Many SaaS vendors forbid benchmarking, I find it crazy that such anti-competitive terms are standard across the industry but they are. Generally the goal of such terms is to &quot;control the narrative&quot; around the product, regardless of the truth of performance being better or worse than competitors.
    3. tomrod · · focus · HN ↗
      I mean, that&#x27;s a great reason to ignore JEV entirely.

      &quot;Trust, but verify&quot; isn&#x27;t just a catchy cliche. It&#x27;s the only way to operate where models and code are fast to market.

    4. meander_water · · focus · HN ↗
      They released some examples of what their workflow evals are like. I&#x27;m sure you could reverse engineer a benchmark from that

      <a href="https:&#x2F;&#x2F;evals.typesafe.ai&#x2F;" rel="nofollow">https:&#x2F;&#x2F;evals.typesafe.ai&#x2F;

    5. florianstandhar · · focus · HN ↗

      [dead]

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.