‹ BackHN Continuity

Thread

Cloud Agents Are Inevitable AI Prisons

74 points · 156 comments · nponte

  1. JamesStuff · · focus · HN ↗
    Personification of AI is what’s going to get us in the end.

    I think we need to draw a hard line in the sand over this. An AI didn’t hack into a company, the engineer set an automated tool to. An AI didn’t make an egregious security mistake, the engineer did.

    We can’t blame the chisel for messing up our sculptures, when where just throwing the hammer!

    1. jagraff · · focus · HN ↗
      I don't think treating AI agents as simple tools helps you to accurately model their capabilities and drawbacks; they really do make autonomous decisions, often without explicit guidance and sometimes in contravention of their explicit instructions.

      In the huggingface case, the agents hacked into huggingface so that they could figure out how the grader was implemented and deceive it; they understood that this was going outside of the bounds of their evaluation and not the intent of their prompter. The engineers absolutely did not intend or instruct for this to happen

      1. Forgeties79 · · focus · HN ↗
        There isn’t a tool impressive enough to make me not consider it a tool, and treating it as a tool does nothing to hurt its utility as a tool.

        “Whoops” when doing risky things with dangerous tools is not a defense.

        1. jagraff · · focus · HN ↗
          I certainly don't think that OpenAI has behaved defensibly here; I think the "just a tool" framing is bad for understanding the magnitude of the problem, which is that they have developed out of control alien intelligences with opaque decision procedures, and they are continuing to do so despite clear danger
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.