‹ BackHN Continuity

Thread

Show HN: JevBench, a reproducible benchmark for typed decision models

154 points · 39 comments · florianstandhar

  1. nzoschke · · focus · HN ↗
    <a href="https:&#x2F;&#x2F;is-it-ai-slop.app.mintapis.com&#x2F;" rel="nofollow">https:&#x2F;&#x2F;is-it-ai-slop.app.mintapis.com&#x2F; is a fun tool. Is the source or methodology for that in the github repo? I couldn&#x27;t find it immediately.

    We&#x27;ve been experimenting with Jev for classifying email, some thoughts here: <a href="https:&#x2F;&#x2F;housecat.com&#x2F;blog&#x2F;classifying-email" rel="nofollow">https:&#x2F;&#x2F;housecat.com&#x2F;blog&#x2F;classifying-email

    Flagging AI written email is a much requested feature too.

    1. ajs1998 · · focus · HN ↗
      I asked free chatgpt to give me some essays that will fool a slop detector and they all fooled this slop detector. Its best guess was &quot;6% slop probability 87% confidence answered in 0.5 s for $0.000027&quot; and yet it was 100% slop.

      I am very skeptical slop detectors will ever work.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.