‹ BackHN Continuity

Thread

The Implications of Linguistic Illegibility for LLM Security

79 points · 29 comments · tomjakubowski

  1. bloppe · · focus · HN ↗
    I thought this was about all the illegible jargon
    1. TZubiri · · focus · HN ↗
      >" However, various strands of evidence indicate that a"

      Strands of evidence? My best guess would be that:

      0- this is ai generated slop

      1- it's using that watermarking technique

      2- it's obviously detectable and degrades quality

      3- it's amplified when inferencing on its own content and generates slop

      1. EagnaIonat · · focus · HN ↗
        I would have more contention with "Dynamic taint", but the paper doesn't appear to be AI slop at all.

        As I understand the paper they are saying the reasoning/thinking you see is actually a translation of what is actually going on, and stuff can be lost in the translation. Similar to what was observed in j-space.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.