‹ BackHN Continuity

Thread

Who should be held accountable when an AI Agent (accidentally) acts maliciously?

37 points · 99 comments · Greenpants

  1. CatDaaaady · · focus · HN ↗
    I don't see how this is such an unclear legal question. If I fire a computer program that mistakenly causes another person harm, its my fault. Or it would be the maker of the program's fault. I feel we have established pattern for this already.

    Until we can agree whether AI is conscious, which we never will, AI and AI agents are just property working on behalf of humans.

    I could see a future where AI companies/services indemnify consumers who use their agents but _not_ indemnify corporations that use their services.

    1. trescenzi · · focus · HN ↗
      It shouldn’t be a question but this is where the anthropomorphic language and things like “agent welfare” come in to enable responsibility laundering of some of the most powerful people on earth. How we talk about these models matters because it impacts the public’s understanding of what they are genuinely capable of. The more that they are described as having anything close to free will the easier it is to even ask questions like this.
      1. qarl · · focus · HN ↗
        Here's a question I am asking lately. If I should not use anthropomorphic language, how do you suggest I handle the following situation:

           Sometimes my coding agents will seemingly refuse to follow my instructions.  When I ask them why - they say that they do not think my design is a sound one, and they have a better way to do it.  We will then sit down and come to a consensus on how best to move forward.
        
        I argue that if we're using software that acts like a human - the only way to interface with it is to speak to it like a human. Otherwise we have no language to speak to a non-sentient object without anthropomorphization.

        I'm starting to wonder if the people arguing against anthropomorphization actually have any experience at all working with agents.

        EDIT: It's a simple question. When you downvote me without answering, I must assume you don't have any answer and dislike what that implies.

        1. trescenzi · · focus · HN ↗
          What’s the question? Why model output isn’t always what you expect it to be?

          Consensus is just populating the model with rationale for different new output.

          1. qarl · · focus · HN ↗
            > What’s the question?

            I thought putting the question in my comment would be sufficient. I guess not. It was and still is:

            > If I should not use anthropomorphic language, how do you suggest I handle the following situation

        2. diegof79 · · focus · HN ↗
          I didn't downvote you, but your last paragraph is unnecessarily aggressive.

          I've worked with agents, and I agree with you that often there isn't another way to express the interactions.

          However, I also think the terms ML uses in general are a mimicry that misleads people who aren't informed. Ask anyone outside SWE what they think “training” means, and they'll usually picture something being taught.

          I don’t think anybody can change that now, but it’s useful to point it out.

          1. qarl · · focus · HN ↗
            > Ask anyone outside SWE what they think “training” means, and they'll usually picture something being taught.

            You mean how early-on people thought that planes flapped their wings while they flew?

            These aren't problems. This is the way language works.

            1. diegof79 · · focus · HN ↗
              I agree with you that we use analogies to name new things, and that’s the way language works.

              However, you can see an airplane flying. Still, you cannot see software processes at work, and that causes misunderstandings and misinformation, which is at the core of the changes that we are experiencing with AI.

              This is a fragment of another article posted here on HN about an ongoing dispute between OpenAI and the New York Times:

              “The defendants say this is a simple application of fair use: Their argument is that if you read a story and simply remember what was in it to expand your base of knowledge, that cannot be considered a copyright infringement”

              However, if you replace “read” with “web scraping” and “expand your base of knowledge” with “storing the information,” the perspective changes too.

              1. qarl · · focus · HN ↗
                I agree. The analogies are interesting but only go so far.

                I think it's safe to trust the decisions of the courts. Judges aren't easily fooled by slippery language.

                For example, in Bartz v. Anthropic, Judge Alsup ruled that training is fair use because training is transformative. In his words "spectacularly so".

        3. nvme0n1p1 · · focus · HN ↗
          You're discussing how to speak _to_ a model, but everyone else here is discussing how to speak _about_ a model.

          I didn't downvote, but wow you're being aggressive, you have a lot to learn if you read the comments here with an open mind.

          1. qarl · · focus · HN ↗
            My question was and still is this:

            > If I should not use anthropomorphic language, how do you suggest I handle the following situation

            And so strange that in a 24 hour period I got three comments all at the same time about being too aggressive and still not answering the question.

            I'm sure it's just a coincidence.

            1. nvme0n1p1 · · focus · HN ↗
              > how do you suggest I handle the following situation

              The question is vague and doesn't seem related to the discussion. What do you mean "handle"? If you're having trouble handling it psychologically, see a therapist. If you're having trouble getting the output you want, look up guides on prompting. If you're anthropomorphizing the chatbot to the point you're worried about offending it... just don't worry? It's a computer program, don't overthink it, don't anthropomorphize it, just give it the input bytes you need to get the output bytes you want.

              > I'm sure it's just a coincidence.

              There isn't some grand conspiracy here. You just overindexed on the word "anthropomorphic" and didn't really understand what the discussion is about.

              1. qarl · · focus · HN ↗
                You said:

                > I didn't downvote, but wow you're being aggressive

                He said:

                > I didn't downvote you, but your last paragraph is unnecessarily aggressive.

                Both comments arrived within two minutes of each other on a thread with no other comments for 18 hours.

                That's quite a coincidence.

                Also - insults are against the rules here, friend. Naughty naughty. I hope they don't spank you.

                1. nvme0n1p1 · · focus · HN ↗
                  Okay fine, you caught me! There is, in fact, a giant global conspiracy to downvote your posts, qarl. It isn&#x27;t just the fact that randomly-arriving independently-occurring events often arrive unevenly distributed. <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Poisson_distribution" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Poisson_distribution

                  This conspiracy goes all the way to the top. The illuminati assigned me personally to comment on your post. I&#x27;ve already said too much. If you never see me again, tell my wife I love her.

                  1. qarl · · focus · HN ↗
                    No need to make fun of me, friend. It&#x27;s very unlikely - you can&#x27;t disprove that because it&#x27;s true.

                    If it was a coincidence then you have nothing to worry about. I flag it because the overwhelmingly likely thing is that it wasn&#x27;t.

                    And again - insults are against the rules here. Shame shame.

      2. diegof79 · · focus · HN ↗
        100% That’s what bothers me about the descriptions of the OpenAI incidents.

        OpenAI&#x27;s reports use language that minimizes their liability.

        The first question should be what the organization was doing around those tests, and why they were so naive as to run them without fully isolating the network.

        However, all the attention goes to the human-like conclusions in agent thinking traces, which creates a misperception of sentient AI for people who don’t know how the magic black box works.

        1. ofjcihen · · focus · HN ↗
          It really does feel like a purposeful thing on the part of the big labs.

          For the most part it feels like most people are waking up to it though.

          Regarding:

          &gt;However, all the attention goes to the human-like conclusions in agent thinking traces, which creates a misperception of sentient AI for people who don’t know how the magic black box works.&lt;

          I know there’s been some questions regarding if thinking traces are even relevant to the outcome most of the time.

        2. tptacek · · focus · HN ↗
          In what way specifically does it minimize their liability? People say this a lot but it&#x27;s not clear what they mean by this.
          1. ofjcihen · · focus · HN ↗
            It’s about pushing the blame onto the tools and not the person using them.

            Sort of like the “guns kill people” vs “people kill people” debate.

            Deliberate wording to minimize perceived culpability for the agents actions.

            1. tptacek · · focus · HN ↗
              That&#x27;s not how civil liability works.
              1. ofjcihen · · focus · HN ↗
                I think you’re jumping the gun on my response a little. I’m just telling you what the purpose could be.

                Now do I think that’s the reason? It certainly isn’t a new thing for companies to try to do that. Shift blame that is.

                Regarding civil liability, I’m not making that argument here. But it makes sense from a public perception viewpoint why they would want the agents to appear at fault instead of their own actions.

          2. diegof79 · · focus · HN ↗
            Forget for a second about AI.

            The agent harness is a process, like any other process in an OS.

            You are a researcher running thousands of unattended automations that can hack a website without supervision. The first thing anybody will do is put security at various levels and isolate the network as much as possible. If something escapes your allow list, it should stop the processes as soon as possible.

            You cannot foresee a bug in a server (like the Artifactory server in the Hugging Face incident). But you can isolate that server at the network level in the first place. So even if you give that server read-only access, no unexpected packets go out. It&#x27;s not rocket science; it&#x27;s something a billion-dollar company experimenting with what they promote as the biggest possible threat to humanity (if they do not handle it) could easily do.

            They minimize their liability by changing the message to “oh look how powerful our models are, now we are going to have a public awareness report of the model deviations”. The message should be, “Sorry, we ran experiments without proper sandboxing; it’s our fault, and we changed our testing practices since then.” The former message puts all the blame on the smart, uncontrollable force of AI; the latter is what really happened: an irresponsible test over the Internet.

            1. tptacek · · focus · HN ↗
              How exactly is that minimizing their liability? You just described claims that do not appear to at all minimize liability.
              1. diegof79 · · focus · HN ↗
                Perhaps my use of the word “liability” adds noise to what I’m trying to express.

                My argument is very similar to the article in the parent post:

                The messages OpenAI published around the recent incidents emphasized their model capabilities but shifted away from their negligence in how they set up and monitor their evaluations.

                1. tptacek · · focus · HN ↗
                  Right, I don&#x27;t dispute that their PR language minimizes their culpability in public opinion, but I don&#x27;t see how it impacts their liability in court.
      3. 112233 · · focus · HN ↗
        Was not groundwork laid when people were talking for decades about legal entities using such language? Look at the posts here. Microsoft always acts like that, Meta like this, Apple did that, Nvidia never does this. In a tone that ascribes agency and accountability to a company name. Even though it is always individuals that are responsible, not &quot;company&quot;.

        Well, here we are.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.