I think you can't trust Pangram in a high stakes situation, but it is absolutely better than random noise at detecting AI-generated text. Which isn't surprising. If the distribution of probabilities can yield blatant Claudisms, it's not surprising it would also have more subtle deviations.
(Addendum: As I recall, LLM-generated outputs roughly follow Zipf's law, but the distribution still tends to have some subtle distinctions vs human text; pretty interesting, but I don't know where I heard this, so nothing to cite. Sorry.)
Pangram has an extremely low false positive rate. Even on adversarial examples.
One trade-off is even some obviously LLM text won't get detected by them, but they work really hard to ensure false positives are rare since a false accusation is much worse for society than someone getting away with LLM meatpuppetry.
claiir · · focus · HN ↗
Damn even Nvidia is putting out fully Claude-written articles.
keybrd-intrrpt · · focus · HN ↗
Why "even Nvidia"?
They are fully behind using AI for basically everything.
What's next? "Damn, even McDonald's is putting out unhealthy food"
jchw · · focus · HN ↗
keybrd-intrrpt · · focus · HN ↗
They are just better at hiding it or configuring Claude.
I have several skills that reformat text to remove AI-speak tells.
jchw · · focus · HN ↗
WD-42 · · focus · HN ↗
I put the "humanized" output through Pangram and it still comes out as 100% AI generated.
breezybottom · · focus · HN ↗
jchw · · focus · HN ↗
(Addendum: As I recall, LLM-generated outputs roughly follow Zipf's law, but the distribution still tends to have some subtle distinctions vs human text; pretty interesting, but I don't know where I heard this, so nothing to cite. Sorry.)
meowface · · focus · HN ↗
One trade-off is even some obviously LLM text won't get detected by them, but they work really hard to ensure false positives are rare since a false accusation is much worse for society than someone getting away with LLM meatpuppetry.