‹ BackHN Continuity

Thread

The Implications of Linguistic Illegibility for LLM Security

79 points · 29 comments · tomjakubowski

  1. bloppe · · focus · HN ↗
    I thought this was about all the illegible jargon
    1. TZubiri · · focus · HN ↗
      >" However, various strands of evidence indicate that a"

      Strands of evidence? My best guess would be that:

      0- this is ai generated slop

      1- it's using that watermarking technique

      2- it's obviously detectable and degrades quality

      3- it's amplified when inferencing on its own content and generates slop

      1. akk0 · · focus · HN ↗
        Maybe you shouldn't be so quick to jump to conclusions, as "strands of evidence" is not a rare turn of phrase and long predates LLMs

        <a href="https:&#x2F;&#x2F;ludwig.guru&#x2F;s&#x2F;strand+of+evidence" rel="nofollow">https:&#x2F;&#x2F;ludwig.guru&#x2F;s&#x2F;strand+of+evidence

        1. a_t48 · · focus · HN ↗
          With all (a lot?) of this language, it&#x27;s not that it&#x27;s using new turns of phrases, it&#x27;s that it will use the same rare but valid phrases over and over. Everybody has their own speaking patterns. But imagine if one person&#x27;s idiosyncrasies were everywhere. That&#x27;s what&#x27;s happened here.
          1. harimau777 · · focus · HN ↗
            That doesn&#x27;t seem that uncommon to me. A lot of people have phrases that they commonly overuse. Sometimes quite dramatically. I wouldn&#x27;t be suprised if that&#x27;s even more the case when communicating complex topics since, at least personally, once I find an approach that lets me communicate some difficult part of an argument I tend to reuse it.
      2. tomjakubowski · · focus · HN ↗
        if Mickens has fallen and resorted to publishing slop then there is no hope for the rest of us
      3. EagnaIonat · · focus · HN ↗
        I would have more contention with &quot;Dynamic taint&quot;, but the paper doesn&#x27;t appear to be AI slop at all.

        As I understand the paper they are saying the reasoning&#x2F;thinking you see is actually a translation of what is actually going on, and stuff can be lost in the translation. Similar to what was observed in j-space.

      4. WithinReason · · focus · HN ↗
        Watermarking doesn&#x27;t degrade output (this is a provable fact) and it is not detectable by a human reading the text.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.