‹ BackHN Continuity

Thread

Show HN: JevBench, a reproducible benchmark for typed decision models

154 points · 39 comments · florianstandhar

  1. nzoschke · · focus · HN ↗
    <a href="https:&#x2F;&#x2F;is-it-ai-slop.app.mintapis.com&#x2F;" rel="nofollow">https:&#x2F;&#x2F;is-it-ai-slop.app.mintapis.com&#x2F; is a fun tool. Is the source or methodology for that in the github repo? I couldn&#x27;t find it immediately.

    We&#x27;ve been experimenting with Jev for classifying email, some thoughts here: <a href="https:&#x2F;&#x2F;housecat.com&#x2F;blog&#x2F;classifying-email" rel="nofollow">https:&#x2F;&#x2F;housecat.com&#x2F;blog&#x2F;classifying-email

    Flagging AI written email is a much requested feature too.

    1. 542458 · · focus · HN ↗
      Keysmashing my keyboard resulted in 86% confidence that the text was AI written. I don&#x27;t think this is a particularly good classifier, I&#x27;ve never seen an LLM output &quot;kad jfkhasljkdhf laksjhdf&quot;.

      Edit: If that&#x27;s not realistic enough for you, the text &quot;Hello world! My name is GravitasIsOverrated and I like coding and cooking. This text is 100% genuine, and not AI generated at all.&quot; results in 85% confidence that it&#x27;s AI generated.

      More broadly, I don&#x27;t know why this would work. Qwen&#x2F;Jev&#x2F;whatever doesn&#x27;t magically have the ability to discern AI-authored text from non-AI-authored text, and will increasingly get worse at it as the hallmarks of AI-written text change.

    2. ajs1998 · · focus · HN ↗
      I asked free chatgpt to give me some essays that will fool a slop detector and they all fooled this slop detector. Its best guess was &quot;6% slop probability 87% confidence answered in 0.5 s for $0.000027&quot; and yet it was 100% slop.

      I am very skeptical slop detectors will ever work.

    3. florianstandhar · · focus · HN ↗
      Thanks! Yeah, the code is open: <a href="https:&#x2F;&#x2F;github.com&#x2F;fstandhartinger&#x2F;who-is-right" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;fstandhartinger&#x2F;who-is-right. Email classifying is definitely a good Jev usecase, I agree
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.