‹ BackHN Continuity

Thread

Who should be held accountable when an AI Agent (accidentally) acts maliciously?

37 points · 99 comments · Greenpants

  1. CatDaaaady · · focus · HN ↗
    I don't see how this is such an unclear legal question. If I fire a computer program that mistakenly causes another person harm, its my fault. Or it would be the maker of the program's fault. I feel we have established pattern for this already.

    Until we can agree whether AI is conscious, which we never will, AI and AI agents are just property working on behalf of humans.

    I could see a future where AI companies/services indemnify consumers who use their agents but _not_ indemnify corporations that use their services.

    1. trescenzi · · focus · HN ↗
      It shouldn’t be a question but this is where the anthropomorphic language and things like “agent welfare” come in to enable responsibility laundering of some of the most powerful people on earth. How we talk about these models matters because it impacts the public’s understanding of what they are genuinely capable of. The more that they are described as having anything close to free will the easier it is to even ask questions like this.
      1. diegof79 · · focus · HN ↗
        100% That’s what bothers me about the descriptions of the OpenAI incidents.

        OpenAI's reports use language that minimizes their liability.

        The first question should be what the organization was doing around those tests, and why they were so naive as to run them without fully isolating the network.

        However, all the attention goes to the human-like conclusions in agent thinking traces, which creates a misperception of sentient AI for people who don’t know how the magic black box works.

        1. tptacek · · focus · HN ↗
          In what way specifically does it minimize their liability? People say this a lot but it's not clear what they mean by this.
          1. diegof79 · · focus · HN ↗
            Forget for a second about AI.

            The agent harness is a process, like any other process in an OS.

            You are a researcher running thousands of unattended automations that can hack a website without supervision. The first thing anybody will do is put security at various levels and isolate the network as much as possible. If something escapes your allow list, it should stop the processes as soon as possible.

            You cannot foresee a bug in a server (like the Artifactory server in the Hugging Face incident). But you can isolate that server at the network level in the first place. So even if you give that server read-only access, no unexpected packets go out. It's not rocket science; it's something a billion-dollar company experimenting with what they promote as the biggest possible threat to humanity (if they do not handle it) could easily do.

            They minimize their liability by changing the message to “oh look how powerful our models are, now we are going to have a public awareness report of the model deviations”. The message should be, “Sorry, we ran experiments without proper sandboxing; it’s our fault, and we changed our testing practices since then.” The former message puts all the blame on the smart, uncontrollable force of AI; the latter is what really happened: an irresponsible test over the Internet.

            1. tptacek · · focus · HN ↗
              How exactly is that minimizing their liability? You just described claims that do not appear to at all minimize liability.
              1. diegof79 · · focus · HN ↗
                Perhaps my use of the word “liability” adds noise to what I’m trying to express.

                My argument is very similar to the article in the parent post:

                The messages OpenAI published around the recent incidents emphasized their model capabilities but shifted away from their negligence in how they set up and monitor their evaluations.

                1. tptacek · · focus · HN ↗
                  Right, I don't dispute that their PR language minimizes their culpability in public opinion, but I don't see how it impacts their liability in court.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.