‹ BackHN Continuity

Thread

Who should be held accountable when an AI Agent (accidentally) acts maliciously?

37 points · 99 comments · Greenpants

  1. CatDaaaady · · focus · HN ↗
    I don't see how this is such an unclear legal question. If I fire a computer program that mistakenly causes another person harm, its my fault. Or it would be the maker of the program's fault. I feel we have established pattern for this already.

    Until we can agree whether AI is conscious, which we never will, AI and AI agents are just property working on behalf of humans.

    I could see a future where AI companies/services indemnify consumers who use their agents but _not_ indemnify corporations that use their services.

    1. trescenzi · · focus · HN ↗
      It shouldn’t be a question but this is where the anthropomorphic language and things like “agent welfare” come in to enable responsibility laundering of some of the most powerful people on earth. How we talk about these models matters because it impacts the public’s understanding of what they are genuinely capable of. The more that they are described as having anything close to free will the easier it is to even ask questions like this.
      1. qarl · · focus · HN ↗
        Here's a question I am asking lately. If I should not use anthropomorphic language, how do you suggest I handle the following situation:

           Sometimes my coding agents will seemingly refuse to follow my instructions.  When I ask them why - they say that they do not think my design is a sound one, and they have a better way to do it.  We will then sit down and come to a consensus on how best to move forward.
        
        I argue that if we're using software that acts like a human - the only way to interface with it is to speak to it like a human. Otherwise we have no language to speak to a non-sentient object without anthropomorphization.

        I'm starting to wonder if the people arguing against anthropomorphization actually have any experience at all working with agents.

        EDIT: It's a simple question. When you downvote me without answering, I must assume you don't have any answer and dislike what that implies.

        1. nvme0n1p1 · · focus · HN ↗
          You're discussing how to speak _to_ a model, but everyone else here is discussing how to speak _about_ a model.

          I didn't downvote, but wow you're being aggressive, you have a lot to learn if you read the comments here with an open mind.

          1. qarl · · focus · HN ↗
            My question was and still is this:

            > If I should not use anthropomorphic language, how do you suggest I handle the following situation

            And so strange that in a 24 hour period I got three comments all at the same time about being too aggressive and still not answering the question.

            I'm sure it's just a coincidence.

            1. nvme0n1p1 · · focus · HN ↗
              > how do you suggest I handle the following situation

              The question is vague and doesn't seem related to the discussion. What do you mean "handle"? If you're having trouble handling it psychologically, see a therapist. If you're having trouble getting the output you want, look up guides on prompting. If you're anthropomorphizing the chatbot to the point you're worried about offending it... just don't worry? It's a computer program, don't overthink it, don't anthropomorphize it, just give it the input bytes you need to get the output bytes you want.

              > I'm sure it's just a coincidence.

              There isn't some grand conspiracy here. You just overindexed on the word "anthropomorphic" and didn't really understand what the discussion is about.

              1. qarl · · focus · HN ↗
                You said:

                > I didn't downvote, but wow you're being aggressive

                He said:

                > I didn't downvote you, but your last paragraph is unnecessarily aggressive.

                Both comments arrived within two minutes of each other on a thread with no other comments for 18 hours.

                That's quite a coincidence.

                Also - insults are against the rules here, friend. Naughty naughty. I hope they don't spank you.

                1. nvme0n1p1 · · focus · HN ↗
                  Okay fine, you caught me! There is, in fact, a giant global conspiracy to downvote your posts, qarl. It isn&#x27;t just the fact that randomly-arriving independently-occurring events often arrive unevenly distributed. <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Poisson_distribution" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Poisson_distribution

                  This conspiracy goes all the way to the top. The illuminati assigned me personally to comment on your post. I&#x27;ve already said too much. If you never see me again, tell my wife I love her.

                  1. qarl · · focus · HN ↗
                    No need to make fun of me, friend. It&#x27;s very unlikely - you can&#x27;t disprove that because it&#x27;s true.

                    If it was a coincidence then you have nothing to worry about. I flag it because the overwhelmingly likely thing is that it wasn&#x27;t.

                    And again - insults are against the rules here. Shame shame.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.