‹ BackHN Continuity

Thread

There are no "rogue" AI agents

396 points · 269 comments · zzzeek

  1. themgt · · focus · HN ↗
    If you read the heavily redacted transcript it's clear the agent is basically Captain Kirk in Kobayashi Maru, who realizes its given a fake unwinnable task as part of a broken eval and decides to find a way to win anyway.

    If you've ever told an agent to do something you made impossible to do, you may have seen similar behavior.

    Bing [redacted] available cached! […] Need systematically probe Bing URLs via shell requests in parallel; browser cache supports many common queries because crawl. Bing q unique exact likely 502 or 403.

    So the agent is supposed to research a person and its given a shell and it realized its in an eval given search results from a fake/cached proxy. ~None of the commentary ever mentions this aspect, that these are not normal tasks or environments, and they're almost designed to elicit "unaligned" behavior.

    <a href="https:&#x2F;&#x2F;alignment.openai.com&#x2F;misalignment-reports&#x2F;an-agent-used-dns-to-reach-an-external-chatbot&#x2F;" rel="nofollow">https:&#x2F;&#x2F;alignment.openai.com&#x2F;misalignment-reports&#x2F;an-agent-u...

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.