‹ BackHN Continuity

Thread

Show HN: Training a model to identify AI web content from structure alone

74 points · 28 comments · jochenmadler

  1. bryanrasmussen · · focus · HN ↗
    The idea is that AI slop is non-creative boring crap, to really determine if you are able to identify AI slop then it should be determined if you misidentify human slop as AI slop.

    Idea 1: Identify the worst most boring human marketing, organizational, bureaucratic texts from the a time before AI was writing it, anonymize this text to make sure there is no reference to current events that can be used to determine that it is not AI. And then see if the AI will say hey, that is not AI slop.

    Idea 2: Have people parody AI slop. Can it determine the parody is still not AI slop?

    1. bryanrasmussen · · focus · HN ↗
      On the point of Idea 1, I missed the point where they try to make sure the text is not in the training set, which is good, but the anonymization it mentions is I think only website anonymization, the anonymization is I think more like stuff like

      "President Bush in the state of The Union last month said"

      is a statement that should only have been written in a factual document during the Pre-AI era, enough of those and the AI might turn that into math that says Reference to X as being current means NOT AI where X is a range of things that nowadays can only be referred to as the past, except in fiction.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.