Show HN: JevBench, a reproducible benchmark for typed decision models
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Show HN: JevBench, a reproducible benchmark for typed decision models
Unofficial Hacker News client; not affiliated with Y Combinator.
nzoschke · · focus · HN ↗
We've been experimenting with Jev for classifying email, some thoughts here: <a href="https://housecat.com/blog/classifying-email" rel="nofollow">https://housecat.com/blog/classifying-email
Flagging AI written email is a much requested feature too.
542458 · · focus · HN ↗
Edit: If that's not realistic enough for you, the text "Hello world! My name is GravitasIsOverrated and I like coding and cooking. This text is 100% genuine, and not AI generated at all." results in 85% confidence that it's AI generated.
More broadly, I don't know why this would work. Qwen/Jev/whatever doesn't magically have the ability to discern AI-authored text from non-AI-authored text, and will increasingly get worse at it as the hallmarks of AI-written text change.
ajs1998 · · focus · HN ↗
I am very skeptical slop detectors will ever work.
florianstandhar · · focus · HN ↗