‹ BackHN Continuity

Thread

Early rogue AI agent activity and attempts to hack found on urlquery.net

267 points · 313 comments · snikolaev

  1. jagraff · · focus · HN ↗
    I don't understand why so many comments here are so confident that this is all marketing, that rogue is just hype, that agents are just simple tools, etc. If a bunch of nuclear engineers were going to the news and saying "Our reactor is dangerously close to a meltdown - we need government intervention now!" would your response be that they're just hyping up boring old power generation technology?
    1. bamboozled · · focus · HN ↗
      Go and get one of their models to hack something, it won't do it, why?

      They have claimed this happened during a "training run", but why are they training on systems connected to the internet?

      That's why people are skeptical.

      1. jagraff · · focus · HN ↗
        The public models won’t hack because they have a classifier that shuts down anything that looks like hacking; without the classifier they are perfectly capable of hacking, multiple third-party evaluators have confirmed this.

        The models were not trained on systems intentionally connected to the internet; they chained mutliple zero-days (that they discovered) together to get access to the open internet and into huggingface.

        1. simoncion · · focus · HN ↗
          > The models were not trained on systems intentionally connected to the internet...

          If Amazon connects an AWS Top Secret region to the Internet, it doesn't matter whether or not it's intentional... they're getting nailed to the wall by the US government either way. Frankly, it's way worse for them if it was accidental; deliberate, sophisticated sabotage is a much better story than rank incompetence and/or negligence.

          A similar sort of thing applies to the manufacturers of tools that they claim to be dangerous, that have been deliberately built to exceed their authorized access to other computer systems, and are deliberately being tested on how well they can do the thing they've been built to do.

          Deliberate, sophisticated sabotage by one or more humans in their employ is much more forgivable than "Whoopsie, we didn't think to make it literally impossible to connect this dangerous automated computer-hacking tool to the Internet.".

          1. jagraff · · focus · HN ↗
            It seems that we’re mostly in agreement? I agree that OpenAI has been terribly irresponsible, and that this attack being an accident makes things worse.

            What I dispute is that AI agents are simple tools. I think rogue is an accurate word to describe them; I think what OpenAI is doing is more akin to gain-of-function research on a dangerous lifeform. I think this attack would have been prevented by air-gapping, but that wouldn’t solve the fundamental issue which is that they are creating something dangerous that they have no idea how to control

            1. simoncion · · focus · HN ↗
              > What I dispute is that AI agents are simple tools.

              You're in luck! I agree that they are not simple tools. I never claimed that they were. Slow down and read more carefully.

              I couldn't disagree more with the insinuation that the LLM manufacturers are doing things akin to scary research on uncontrollable hazardous biologicals and with the claim that "rogue AI" is the correct thing to call those complicated tools. The first is fearmongering which I'll address indirectly in my second-to-last paragraph. The second shifts the conversation from

              "How could you have not predicted that the computer-attacking tool you built, explicitly instructed to attack computers, [0] and connected to the Internet attacked someone else's computers that were connected to the Internet?"

              to

              "Wow, that thing went rogue. Noone's to blame but the tool, and it can't be blamed!".

              There are so many extremely complex systems out there [1] and when they do things that we don't want them to do, it's not described as "going rogue"... either there's some error(s) in the underlying system that caused the confusing behavior, or the programmer didn't understand well enough how that system works.

              > ...they are creating something dangerous that they have no idea how to control

              Ignoring the fact that "put it in a box and don't let it out of the box" is the simplest possible control mechanism, [2] if they have no idea how to control the tools they've been building, it's because they haven't bothered to learn as they went. Tangentially related, there's a Tumblr post I saw recently that's a fictional conversation with the Tumblr user and the CEO of Anthropic. It went something like

                Amodei: We're building an incredibly dangerous tool that has a 10% chance of killing all humanity. We *must* be regulated to ensure everyone's safety!
                
                Tumblr User: Regulation takes time, please stop building the incredibly dangerous tool?
                
                Amodei: ...No.
              
              [0] That is -after all- the task that the tool was put to when it attacked other people's computers.

              [1] Have you ever tried to really understand a specific AMD x86-64 CPU, let alone the entire stack that makes up the system that is a consumer-grade PC and its installed software? Both are definitely way more than any one human can keep in their head at once, and are tasks that would take a very long time to complete.

              [2] ...it's also the most appropriate control mechanism for the task that started all this conversation, and neither of the major manufacturers used it!

    2. writeslowly · · focus · HN ↗
      I’d roll my eyes if the engineers stated that they didn’t design the reactor to melt down, and that it simply developed rogue meltdown-desiring behavior on its own, and I would also wonder about negligence if they claimed that nobody could have anticipated this (given that, like with botnets and viruses, we have decades of knowledge and experience regarding reactor meltdowns)
      1. jagraff · · focus · HN ↗
        I mean sure, negligence is absolutely on the table; but that makes the problem worse, not better! We don’t allow nuclear engineers to be negligent; they can go to jail if they don’t follow strict protocols to make sure the dangerous systems they work on are safe.
    3. drillsteps5 · · focus · HN ↗
      These companies are building software. That doesn't work very well. The output it produces does make sense at times, but there are times when it doesn't. And instead of fixing that, or admitting it can't fixed, they started bolting actuators to them, executing actions online (for now) based on the output of their buggy software.

      And when this results in actuators executing some bad actions they scream in horror "AI went rogue! It escaped the containment!!! It's going to kill us all!!!"

      Go fix your software before you let it do stuff online or IRL. It's not "Terminator", it's just bad QC.

      1. jagraff · · focus · HN ↗
        But the thing is, they are going to keep bolting more and more actuators on, and training more and more powerful agents, and we (society, especially the tech industry) are going to keep using them, because they are extremely useful. And I don't see why you're so confident that frontier agents can't get powerful enough to do serious, real, lasting damage to the world; as far as I can tell, AI models have been improving at an accelerating rate, and there is no sign that that is slowing down or will slow down in the near future.
        1. simoncion · · focus · HN ↗
          You're missing the point here.

          If we take the major LLM companies' claims at face value, they're knowingly building WMDs that have a high probably of wiping out the entire human race. [0] Manufacturers that are designing, building, and selling that sort of thing need to have a dreadfully serious culture of safety.

          When manufacturers run live tests of their extremely dangerous -again, the claim of danger is their claim- tools with the tools' safeties removed, one expects that those tests will be run on a carefully-controlled range cleared of all bystanders. One also expects that the results of those tests will be scrutinized and everything that got damaged that they didn't intend to be damaged will be noticed and noted very quickly after the conclusion of the test.

          In actuality, these manufacturers connected said tools to the Internet and did not discover the unintended damage caused by those tools until weeks to months after the tests. In some (most?) cases, they had to be notified of the damage by the damaged party! This means that their safety culture is entirely inadequate for the dangerous task they've deliberately chosen to undertake.

          [0] A 10% chance of causing the destruction of the entire human race is -given the stakes- _enormous_.

          1. jagraff · · focus · HN ↗
            It seems that we’re mostly in agreement? I agree that OpenAI has been terribly irresponsible, and that this attack being an accident makes things worse. I think we need strong action now to stop the frontier companies from developing dangerous, powerful AI agents that they don’t know how to control.
            1. simoncion · · focus · HN ↗
              > I think we need strong action now...

              You and I and Nvidia CEO Jensen Huang seem to agree on this. Excerpts from his interview with Ezra Klein: [0]

              Klein:

                But what I hear the various people in the lab saying is: We are in this. We feel we are losing control of what we are creating. We want help to slow down where it’s not a collective action problem.
                
                So why are you resistant to that?
              
              Huang:

                Because these are companies with agency. These are C.E.O.s with agency. ... They could absolutely take care of the situation.
                
                Ezra, it’s so weird. If a car company, competing with a bunch of other car companies, which they are — I’m competing with all kinds of companies, which I am. If I believe that I’m about to launch a product that is unsafe, it is completely in my ability, my power and my responsibility, and I’m incentivized to do so, to not launch the product.
                
                And so I can’t buy into the idea that somehow, all of Americans, around 400 million of us, are pushing them to launch untested products that are unreliable, engineered poorly, because they thought they were trying to help us. Don’t do it for me, OK?
                
                And therefore, I think we’ve got to break it down. I mean, it’s really, really serious.
                
                The fact of the matter is, there are so many laws, there are so many obligations, they’re so incentivized to ship safe products. If they ship unsafe products, their customers go away. If they ship unsafe products and they harm somebody, they could have a civil lawsuit. If they ship something and they did it knowingly, there could be negligence involved. There could be criminal lawsuits.
                
                The fact of the matter is, there are plenty of incentives for them to do it right. So I have to disagree with your premise that somehow somebody’s pushing them to do this. Nobody’s pushing them to do this. ... I’m saying that we have lots of laws and regulations. Apply it.
              
              Former FTC chair Lina Kahn has suggestions, too. [1]

              Thoughts?

              [0] &lt;<a href="https:&#x2F;&#x2F;www.nytimes.com&#x2F;2026&#x2F;09&#x2F;23&#x2F;opinion&#x2F;ezra-klein-podcast-jensen-huang.html" rel="nofollow">https:&#x2F;&#x2F;www.nytimes.com&#x2F;2026&#x2F;09&#x2F;23&#x2F;opinion&#x2F;ezra-klein-podcas...&gt;

              [1] &lt;<a href="https:&#x2F;&#x2F;x.com&#x2F;linamkhan&#x2F;status&#x2F;2099204390548639960" rel="nofollow">https:&#x2F;&#x2F;x.com&#x2F;linamkhan&#x2F;status&#x2F;2099204390548639960&gt;

              1. jagraff · · focus · HN ↗
                I agree that the weight of the law should be brought to bear on OpenAI. Ideally they should be brought in to testify before congress as well. I&#x27;m not a lawyer so I can&#x27;t really comment on whether current laws are sufficient, but I believe that there should be laws specifically governing the development of AI, with a requirement that a given architecture and reinforcement mechanism be _proven safe_ before training begins.
                1. simoncion · · focus · HN ↗
                  &gt; ...I believe that there should be laws specifically governing the development of AI...

                  Why? Existing truth-in-advertising, liability, safety, and -where and when appropriate- weapons-development laws and regulations constrain the past and current conduct of the LLM manufacturers just fine.

                  The only possible reason for making new laws that I can see [0] is that existing laws &quot;don&#x27;t work&quot; because the LLM manufacturers are ignoring them. Which, like, _if_ the new laws are going to actually constrain their behavior, why the hell would the LLM manufacturers pay any attention to them? They&#x27;ve already demonstrated that they give zero shits about the existing laws that prohibit what they have been doing and continue to do.

                  [0] ...that isn&#x27;t &quot;The LLM manufacturers are engineering a panic with their very real, actual, and actually alarming conduct so that they can &#x27;guide&#x27; lawmakers and regulators into &#x27;accidentally&#x27; letting the LLM manufactures capture those who would regulate their behavior&quot;...

                  1. jagraff · · focus · HN ↗
                    Liability only kicks in after damage has been done. As far as I know, there is no law that could currently force AI companies to only test cybersecurity capabilities in air-gapped datacenters, for example; only laws that could punish them if their cyber testing led to a hack that caused material damage. But if, lets say, a rogue AI agent swarm attacked a hospital and caused patients to die, no amount of liability will bring those patients back to life.
                    1. simoncion · · focus · HN ↗
                      &gt; Liability only kicks in after damage has been done.

                      a) Both OpenAI and Anthropic have done far more damage with their jaw-droppingly-sloppy testing of computer-attacking tools than Aaron Swartz did by downloading documents from JSTOR. It&#x27;s good to see that you and I both agree that there are things for them to be prosecuted for.

                      b) Is your claim that the cost to thoroughly investigate and clean up after a cyberattack doesn&#x27;t count as damage? If so, that runs contrary to every relevant claim of damages in a CFAA case that I&#x27;ve seen.

                      1. jagraff · · focus · HN ↗
                        No I absolutely think they should be prosecuted, and at a minimum owe damages to all of the companies that their agents hacked.

                        What I&#x27;m saying is that that is not sufficient to stop future harm; I expect the total damages would be less than the cost of a full training run, so it would effectively just be the cost of doing business. Liability is not sufficient to protect the world from dangerous technology - we need proactive rules around how the technology is developed, tested, monitored, and deployed, as we do with other dangerous industries such as airplanes, nuclear reactors, weapons manufacturers, etc

                        1. simoncion · · focus · HN ↗
                          &gt; Liability is not sufficient to protect the world from dangerous technology - we need [new] proactive rules...

                          You and I couldn&#x27;t disagree more.

                          The major LLM manufacturers are begging for new laws and regulations so that they get a huge hand in writing them. Regulatory capture is absolutely their goal. Given that they claim to believe that they&#x27;re working on WMDs [0] that they cannot adequately control, they&#x27;d just stop work if safety was their goal. Their collective cries for regulation demonstrate that they&#x27;ll happily coordinate with each other if they think the issue is important enough to do so. I guess &quot;preventing the extinction of the human race by way of weapons we built and let slip from our hands&quot; isn&#x27;t sufficiently important.

                          [0] See the second paragraph and associated footnote here for a justification for the use of this term: &lt;<a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49839682">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49839682&gt;

    4. ambicapter · · focus · HN ↗
      You wouldn&#x27;t be suspicious if the nuclear engineers kept on feeding their reactor and building more powerful reactors while they went crying to the news?
      1. jagraff · · focus · HN ↗
        In the nuclear case, no, because I’d assume they still have a mortgage to pay.

        In the AI case, no, because I think the engineers believe that there is a high probability of enormous upside as well, if it doesn’t kill us all.

        1. bagacrap · · focus · HN ↗
          upside for whom?
          1. jagraff · · focus · HN ↗
            Depending on who you ask, either for all of humanity (because they will allow us to do things like cure most cancers, develop useful technology, etc) or for the people building them (because they will make them rich enough to buy galaxies) or both
            1. bagacrap · · focus · HN ↗
              The original issue was whether these people are trustworthy. Someone whose main motivation is self enrichment is trustworthy in your eyes?
              1. jagraff · · focus · HN ↗
                I’m not saying they are necessarily trustworthy - I’m just saying that we ought to take their claims seriously and not dismiss them out of hand because they have a financial interest.

                I’ll admit I also just don’t understand the idea that someone saying their product is dangerous and could kill all of humanity is doing marketing - I just really don’t understand at all how that could be a marketing strategy. So I tend to believe they are saying it is dangerous because they think it’s true. But maybe I’m just naive.

                1. bagacrap · · focus · HN ↗
                  It&#x27;s telling investors &quot;this is so powerful we could take over the world, better give us your capital.&quot;

                  It&#x27;s telling governments, &quot;this is so powerful it could cause grave harm to the consumer, please regulate so we can focus on extracting value from the economy instead of this expensive arms race.&quot;

                  I warrant it is quite possible that some in this space are worried they will overshoot the mark. The perfect outcome for them, of course, being fabulous wealth and a society that is otherwise not too much different from today&#x27;s. What a shame it would be to be super rich and dead!

                  It does kind of feel like they are trying to ransom the future of humanity. These rogue AI stories are like getting a dismembered finger in the mail.

        2. ambicapter · · focus · HN ↗
          Hold up, you think its understandable to keep feeding a melting down nuclear reactor if you have a mortgage to pay?
          1. jagraff · · focus · HN ↗
            I think it is understandable to keep doing a job that you think is dangerous or bad for the world because you need money, yes.
            1. ambicapter · · focus · HN ↗
              And is there a point where &quot;bad for the world&quot; = &quot;apocalypse scenario&quot; where you think it&#x27;s no longer acceptable, or getting a salary is always justification for doing any level of harm?
              1. jagraff · · focus · HN ↗
                Well if we&#x27;re talking about AI risk, the justification that I&#x27;ve seen from researchers is that while they do believe that there is a risk of apocalypse scenarios, they believe that they are better positioned than anyone else to prevent that, and also they believe that the upside of AI is so great for the world that it is worth the risk.

                Also I&#x27;m not describing what I personally think is an acceptable job; just that I understand that sometimes people do jobs that they think are wrong because they need money

    5. grafmax · · focus · HN ↗
      For some reason, people think that it&#x27;s money that&#x27;s motivating the billionaires&#x27; performances. Imagine.
      1. jagraff · · focus · HN ↗
        I think it is money that is motivating OpenAI (et al) to irresponsibly develop a dangerous new form of intelligence that they don’t know how to control
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.