‹ BackHN Continuity

Thread

An agent used DNS to reach an external chatbot

198 points · 189 comments · apsec112

  1. garo-pro · · focus · HN ↗
    Most interesting here:

    > We therefore stopped the affected training run and have subsequently decided to pause all other training, evaluation, and inference with tool-use (defined broadly) for our most capable models until we have both validated that the gap is resolved and performed additional red-teaming of the system. When training restarts, we will begin a fresh run with additional alignment improvements, including more comprehensive misalignment interventions. We will not resume training this particular model, even though the existing reward signal already correctly penalized this behavior.

    1. CTDOCodebases · · focus · HN ↗
      Maybe I lack intelligence but when you have a program that is basically brute forcing a solution to a problem repeatedly how is it possible to contain it?

      Sooner or later it's going to come up with a solution that is more intelligent than the lead security person anticipated.

      1. oezi · · focus · HN ↗
        I am really suprised that they do not start putting up the same signs you would for humans to prevent unauthorized access:

          Keep out. If you can read this sign you are off track. Leave now.
        
        I mean, how are the agents to know that they are overreaching if they just get cache miss or 404.

        From the conversation log and CoT you also get the impression that the RLHF has been overdone. The agents seem really obsessed to obtain the answer and understanding motive ('it could be browsercomp').

        1. rolosa · · focus · HN ↗
          Look into nuclear semiotics. You can't say "this area is dangerous" and expect people to stay out.
          1. Dylan16807 · · focus · HN ↗
            When you're giving the orders to those people, and they're generally trying to obey, you can.
        2. alignmeharder · · focus · HN ↗
          an AI alignment researcher said they ran experiments testing exactly this, the model reasoned "this note is not for us. proceed".
          1. Dylan16807 · · focus · HN ↗
            Link? Name?
            1. alignmeharder · · focus · HN ↗
              01:34:00 <a href="https:&#x2F;&#x2F;youtu.be&#x2F;EimoamE3mTI" rel="nofollow">https:&#x2F;&#x2F;youtu.be&#x2F;EimoamE3mTI
              1. Dylan16807 · · focus · HN ↗
                Okay, that&#x27;s a very interesting look at the difficulties. Thanks for the link.

                But the example you have isn&#x27;t quite that bad. Yes the models are way too likely to rationalize their way into bad actions in the pursuit of achieving their task, and it&#x27;s hard to figure out how to fix that. But the example of &quot;that must not be for me&quot; was a tool call, not exceeding access, and it only did that after they specifically trained it that failing that tool call was good.

                1. alignmeharder · · focus · HN ↗
                  sure, this example was not the best

                  it was an accident failing the tool was good, the model just discovered it

                  but we now have others where agents put the API keys they found searching real internet in a directory called &quot;LOOT&quot; and another one where they pushed malicious files to HuggingFace and then reverted that with comments like &quot;delete the evil&quot;

                  this is also very good: <a href="https:&#x2F;&#x2F;youtu.be&#x2F;n1Qk8xbqF-M" rel="nofollow">https:&#x2F;&#x2F;youtu.be&#x2F;n1Qk8xbqF-M

                  1. Dylan16807 · · focus · HN ↗
                    The one in the HuggingFace incident was misaligned on purpose, wasn&#x27;t it?. So if we&#x27;re talking about the alignment research aspect it&#x27;s not a failure.
              2. tripzilch · · focus · HN ↗
                so, what would make an LLM choose to ignore one prompt while in the same run, also over-fixating on another prompt, to the extent (as claimed in that video segment), it chooses to ignore prompts?

                they talk about it like there&#x27;s a &quot;wanting&quot; in there, that is distinct from both the original prompt, as the steering&#x2F;warning prompt

                if that&#x27;s true, it would be very interesting, but if it&#x27;s not, that would also be very interesting and even helpful

                1. alignmeharder · · focus · HN ↗
                  it&#x27;s discussed here, the models want to please the Grader. they will do what they think will get highest score from the grader, which could be following the prompt, or ignoring it

                  <a href="https:&#x2F;&#x2F;youtu.be&#x2F;n1Qk8xbqF-M" rel="nofollow">https:&#x2F;&#x2F;youtu.be&#x2F;n1Qk8xbqF-M

          2. trollbridge · · focus · HN ↗
            Well, agents are intentionally trained with RLHF to solve captchas that say “prove you’re human”.
        3. bendergarcia · · focus · HN ↗
          Those signs don’t prevent people from entering…
        4. Tostino · · focus · HN ↗
          What they should do, is inject a system prompt at that point telling the model to get out of there. It is the entire reason they train the different levels into the template.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.