‹ BackHN Continuity

Thread

Show HN: Training a model to identify AI web content from structure alone

74 points · 28 comments · jochenmadler

  1. evantbyrne · · focus · HN ↗
    Immediately saw false positives on content written before ChatGPT. Not difficult to see how the methodology is wrong when it considers restating the thesis in the conclusion to be signal.
    1. jochenmadler · · focus · HN ↗
      Can you elaborate on the false positives?

      Re methodology: Restating the thesis is one signal of many. 77% of the AI posts close that way, but so do 12% of the human ones. The classifier in the paper never decides on one feature, but their combination.

      1. evantbyrne · · focus · HN ↗
        Human writing was flagged as AI generated by the tool with dubious explanations given. Other human written posts were marked down substantially. I mentioned one feature, which seems to be double counted to mark down posts a whopping 20%, but none of the features listed on the results page gave me confidence in the classifier's ability to distinguish between undergraduate essays and slop.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.