‹ BackHN Continuity

Thread

Understanding the Impact of LLM Watermarking on AI Agent Behavior

58 points · 72 comments · nisosguy

  1. jamienk · · focus · HN ↗
    In _1984_ the Big Brother regime has the idea that by controlling language you can influence what is possible to think, and thus becomes a key tool of political repression.

    Political Correctness has a similar idea that by adjusting the terminology we use, we can purge biases and historical implications and speak in a purer way.

    Psychoanalysis has its own idea of repression - where a person struggles to BLOCK our associations between ideas, memories, and words in order to try to stop one thought from being contaminated by another, intolerable thought.

    All of these attempts to control language are fundamentally misguided at best, often have severe unintended consequences, and are genuinely immoral at worse.

    1. p_j_w · · focus · HN ↗
      > Political Correctness has a similar idea that by adjusting the terminology we use, we can purge biases and historical implications and speak in a purer way.

      You either misunderstand political correctness or are trying to make it sound more nefarious than it is. It is nothing more than an effort, sometimes overdone and misguided, to not say things that make minorities feel bad. It’s nothing more than an attempt to broaden what’s considered good manners.

      If anyone serious thinks political correctness is going to end racism, I certainly haven’t seen it.

      > Psychoanalysis has its own idea of repression - where a person struggles to BLOCK our associations between ideas, memories, and words in order to try to stop one thought from being contaminated by another, intolerable thought.

      It’s been a while since I took Psych 101, but my memory is that, according to Freud, repression is subconscious. There’s no struggle possible because there is no intent. It’s also about memories and emotions, so if it WAS an intentional act, it’s not an attempt to control language. It wouldn’t be about trying to not think of the word cat, but trying to not think about that time when your cat died.

      1. pessimizer · · focus · HN ↗
        > It is nothing more than an effort, sometimes overdone and misguided, to not say things that make minorities feel bad.

        Unless it enforced by governments or other monopolies. Then it inevitably leads to lèse majesté, because 1) powerless minorities are powerless, and 2) the controllers of governments and monopolies (the powerful themselves) are a minority group.

        The protection of the feelings of powerless minorities is a Trojan Horse; the grievances of powerless minorities are usually far more serious than hurt feelings, not saying magic words involves far less effort than actually addressing those grievances, and demanding that the public changes the way they speak costs governments and monopolies nothing. It in fact allows them to build the technical and legal structures to enforce lèse majesté.

        > If anyone serious thinks political correctness is going to end racism, I certainly haven’t seen it.

        Not if you say it like that, because it's so obviously stupid. But if you really investigate with people why they want people to stop saying "the n-word," they'll admit they have a fantasy that the sentiment will disappear with the word: it's why they also want black people themselves to stop saying it. The possibility that the sentiment will stay but that black people will lose the language to refer to it doesn't occur to them.

        Just like color-blind casting of historical theatrical works isn't seen as a whitewashing of the racism of the past; there's a fantasy that if you wipe the memory of racism from the media, then it will somehow wipe it from the world rather than just gaslight minorities about their own pasts (and thereby their present condition.) At least that's just in the media, and breaking up monopolies can just make it a form of individual artistic expression rather than a diktat.

        The censorship of speech itself, however, is a totalitarian version of this. It can only be done by monopolies. To make sure that people can't eat if they express themselves in a particular way is as compelling as you can get.

        None of this has anything to do with AI watermarking, though, which is good. Somebody made a good point above that watermarking would allow a cartel of AI providers to make sure that their models were trained with uncontaminated data, while everybody else would have to train with the stepped-on garbage they were spitting out. But that's a job for antitrust. If you don't have antitrust, you don't have markets and you don't have freedom. Little technical touches around the edges are worthless. If they share, make them share with the entire class.

        1. jamienk · · focus · HN ↗
          I think it is related to AI watermarking, though, in that all forms of trying to channel words in certain directions, or to add a meta-layer to language, share a family resemblance.

          There might be an inevitability to this - there is always a form PC policing, or euphemism, or re-brands, or rhetoric. But there is always the danger that what is considered a well-meaning tweak to language masks or evolves into repression.

          From TFA "Text watermarking is meant to help identify whether content was generated by AI, but inside an agent it also becomes part of the generation process that produces decisions."

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.