‹ BackHN Continuity

Thread

An agent used DNS to reach an external chatbot

198 points · 189 comments · apsec112

  1. garo-pro · · focus · HN ↗
    Most interesting here:

    > We therefore stopped the affected training run and have subsequently decided to pause all other training, evaluation, and inference with tool-use (defined broadly) for our most capable models until we have both validated that the gap is resolved and performed additional red-teaming of the system. When training restarts, we will begin a fresh run with additional alignment improvements, including more comprehensive misalignment interventions. We will not resume training this particular model, even though the existing reward signal already correctly penalized this behavior.

    1. CTDOCodebases · · focus · HN ↗
      Maybe I lack intelligence but when you have a program that is basically brute forcing a solution to a problem repeatedly how is it possible to contain it?

      Sooner or later it's going to come up with a solution that is more intelligent than the lead security person anticipated.

      1. rolosa · · focus · HN ↗
        You start by holding actual real life people with something to lose, like the entire executive suite, accountable. Suddenly I'm sure the problem will be resolved with proper safeguards.
        1. EdwardDiego · · focus · HN ↗
          Bingo. This weird attempt to pretend like these incredibly capable algorithms aren't incredibly capable algorithms deployed by a person who works for a company, but somehow have an independent personage that absolves both person and company of responsibility, is just ridiculous.

          Like the whole Huggingface thing, OpenAI employees initiated the test, deliberately removed safeguards, failed to properly lock down the environment, and responded incredibly poorly to evidence that things were going awry.

          The individual employees bear responsibility, but the people running OpenAI are ultimately responsible for the processes and culture where that can happen.

          And then people writing blogposts about "3 civilizations of agents" and "altruistic suicide" by algorithms perfectly muddy the waters and obscure the very obvious responsibility that lies with humans and corporations, which I suspect suits the pre-IPO corporations very well.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.