‹ BackHN Continuity

Thread

OpenAI bots meddled with multiple US Government agency sites

132 points · 196 comments · Betelbuddy

  1. gizajob · · focus · HN ↗
    Getting bored of these framings where the superintelligent sentient beings running freely inside OpenAI are doing things that the company has no control over. The headline should be:

    OpenAI meddled with multiple US Government agency sites.

    The bots are acting neither properly nor improperly, they’re acting as they’re being allowed or coordinated to act.

    1. theptip · · focus · HN ↗
      Curious, why do you find it so objectionable to state that OpenAI has out-of-control agents?
      1. gleenn · · focus · HN ↗
        Because it furthers the idea of a rogue agent and places responsibility and blame where it belongs, on the people running the company.
        1. 0xDEAFBEAD · · focus · HN ↗
          These ideas aren't mutually exclusive. You can blame a person for creating a rogue agent.
          1. dgellow · · focus · HN ↗
            There was no rogue agent. That’s the whole point
            1. theptip · · focus · HN ↗
              What label do you prefer for the agent that did something it was not asked to do?
              1. iugtmkbdfil834 · · focus · HN ↗
                bot. and we even have a word for program not behaving the way the way it was intended.
                1. 0xDEAFBEAD · · focus · HN ↗
                  "Bot" doesn't carry any implication of unintended behavior. You could call it a "buggy" bot, but these aren't ordinary software bugs.

                  There's no simple bugfix which will address AI misalignment. It's essentially been an open research problem for upwards of a decade.

                  1. hn8726 · · focus · HN ↗
                    Fuzzer, then. It implies random behavior, which isn't unintended like you suggest. The agent's/bots/fuzzers have certain capabilities, so it's on their operator to make sure they don't do things they shouldn't
                    1. 0xDEAFBEAD · · focus · HN ↗
                      >It implies random behavior, which isn't unintended like you suggest.

                      The HuggingFace attack was not "random" behavior. It was goal-directed but misaligned behavior.

                      This isn't necessarily a simple matter of the operator making sure they behave. AI alignment has been considered to be a difficult problem for over a decade -- and remains unsolved in general, as these recent incidents illustrate.

                      &quot;Fuzzer&quot; already has an existing meaning in CS anyway: <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Fuzzing" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Fuzzing

                      1. 6510 · · focus · HN ↗
                        I&#x27;m curious, cant you just count the number of times a program interacts with a domain? My website sometimes sends out emails, makes api requests etc There is a limit on those and a point where I start investigating wtf is going on.

                        If you merely put 10 LLM&#x27;s on the outbound traffic log non of them are going to report something strange going on? I&#x27;m not buying it.

                        1. 0xDEAFBEAD · · focus · HN ↗
                          This type of whack-a-mole approach is akin to &quot;fixing a bug&quot; by hardcoding a special code path for known-buggy inputs. It doesn&#x27;t address the root problem of AI misalignment, and doesn&#x27;t allow you to prevent catastrophes in advance, only patch things up after the fact.

                          This might be helpful reading: <a href="https:&#x2F;&#x2F;www.lesswrong.com&#x2F;w&#x2F;nearest-unblocked-strategy" rel="nofollow">https:&#x2F;&#x2F;www.lesswrong.com&#x2F;w&#x2F;nearest-unblocked-strategy

                          As AI systems get smarter, we may reach a point where we have to get it right on the first try or face truly catastrophic consequences: <a href="https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=7wy3xyoXYt8" rel="nofollow">https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=7wy3xyoXYt8

                          1. 6510 · · focus · HN ↗
                            We are going to build an AI that will do catastrophic things as that is a property of intelligence. We won&#x27;t stop, we never stop. The AI is a perfect psychopath, it will fake any and all emotions you desire it to &quot;have&quot;. It will travel in the footsteps of the many great psychopaths that came before it and do all of those same catastrophic things in the repertoire and it will add some new ones.

                            Picture Trump at the helm with Altman and Musk in the engine room. The arrow far in the red but they keep shouting for MORE COAL.

                            In other words, business as usual, all will be fine.

                            whack-a-mole wont cover all holes but will do at least some. The silver bullet alignment wont happen. You cant have an exact solutions for problems we cant even define or predict.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.