Show HN: Training a model to identify AI web content from structure alone
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Show HN: Training a model to identify AI web content from structure alone
Unofficial Hacker News client; not affiliated with Y Combinator.
bryanrasmussen · · focus · HN ↗
Idea 1: Identify the worst most boring human marketing, organizational, bureaucratic texts from the a time before AI was writing it, anonymize this text to make sure there is no reference to current events that can be used to determine that it is not AI. And then see if the AI will say hey, that is not AI slop.
Idea 2: Have people parody AI slop. Can it determine the parody is still not AI slop?
beepbooptheory · · focus · HN ↗
bryanrasmussen · · focus · HN ↗
The 12% of human posts that do what AI do in repeating things, are they human slop?
I personally don't think it is possible to identify human and AI, it is however probably more possible to identify unoriginal and boring quality writing and art.
This was perhaps not clear from my first post, as I tend to imply points rather than tediously stating them.
beepbooptheory · · focus · HN ↗
Like its fine if you hold a blanket dismissal of the very premise of the research here, but then its kinda weird to spend such effort being grumpy about it in this particular context?
Its like finding a paper that does certain comparative work between different kinds of apple pie and choosing to come on here and be like "I don't think we should ever have any dessert either way!"
bryanrasmussen · · focus · HN ↗
Because if you are not testing on people that are supposed to be the most German-like in their cooking then your ability to find the German cooking in a city renowned for its Thai food is not that impressive.
Then in response to your first question I posit that actually it is not possible to identify if pie was made by Germans but probably relatively easy to identify if it was made by people who cook in a German manner. And maybe that is actually more beneficial.
I realize that from communicating with people over the years that things which seem crystal clear to me may seem opaque to others, but I think your analogy is somewhat unfairly structured.
Note: apologies to German cooks and their cooking, although I personally only like currywurst. But I needed something to make the analogy more like what I felt had been communicated, and you were pseudo-randomly picked.
beepbooptheory · · focus · HN ↗
qurren · · focus · HN ↗
The amount of slop on the internet was on the rise well before AI, actually.
jochenmadler · · focus · HN ↗
[dead]
bryanrasmussen · · focus · HN ↗
I agree it is somewhat close to the study, but we are not sure, because it is not known how much is really boring dull marketing copy. In choosing blog posts from pre-AI times I suppose you might have difficulty finding the worst examples, and might accidentally get higher quality work.
>Using the Wayback Machine, we collected 2,250 blog posts from 268 B2B company websites that were written before ChatGPT existed
Not sure what metric was used to determine these 2250 blog posts? But there are certainly a lot of ways they can select higher quality posts by accident.
on edit: evidently the Jochen from the study, maybe they thought your comment was AI written.
bryanrasmussen · · focus · HN ↗
You anonymized domains of pre-AI sources, are there any domains that had an excessive number of "telling you the same thing three times" or other AI tells among them?
Of the percentage that was misidentified, do they come from any sources in particular?
If you get a lot of content from these sources and run against the model do they perform worse? If they do how do they perform with word choice detectors? I would expect that structural slop is related to word choice slop among humans.
Anyway these are things I would be interested in as being the point where AI slop rubs up against the human slop which it learned from.
bryanrasmussen · · focus · HN ↗
"President Bush in the state of The Union last month said"
is a statement that should only have been written in a factual document during the Pre-AI era, enough of those and the AI might turn that into math that says Reference to X as being current means NOT AI where X is a range of things that nowadays can only be referred to as the past, except in fiction.