Nvidia wants to put a watchdog chip next to every AI agent
Thread
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
Nvidia wants to put a watchdog chip next to every AI agent
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
dist-epoch · · focus · HN ↗
HPsquared · · focus · HN ↗
mattmcal · · focus · HN ↗
speedgoose · · focus · HN ↗
soulofmischief · · focus · HN ↗
totetsu · · focus · HN ↗
applfanboysbgon · · focus · HN ↗
bigyabai · · focus · HN ↗
The existence of Nvidia's optional watchdog chip does not in any way impinge upon your freedom to develop and test your own alternative.
The problem is that OpenAI has ostensibly neglected their duty to safety, so Nvidia is stepping in to fix it since they're the "hard problem" people.
jacquesm · · focus · HN ↗
bigyabai · · focus · HN ↗
jacquesm · · focus · HN ↗
johnsmith1840 · · focus · HN ↗
Or push to github?
Dylan16807 · · focus · HN ↗
wyre · · focus · HN ↗
AndrewDucker · · focus · HN ↗
Or, more likely, control things at the network level so that packets from the LLMs you're investigating cannot leave the virtual network they're assigned to.
simoncion · · focus · HN ↗
Not in the way you're thinking, no.
Any Real Server [0] in a datacenter will have some sort of "lights-out management" hardware used for remote access to that server. This stuff is known by a handful of acronyms, but I'll stick with "IPMI" because I like it best. This IPMI hardware is -effectively- a second small PC built into the motherboard. It will pretty much always have its own NIC... and I think I've seen versions that have their own physical ports to attach a monitor, keyboard, and mouse.
What exactly you can do with it varies from vendor to vendor, but -if your IPMI user account has the correct permissions- you are nearly always able to change "BIOS" settings, power cycle the server [1], and attach a virtual keyboard, monitor, and mouse so you can manage the server as if you were standing next to it in the datacenter with a crash cart plugged right in. Every IPMI system I've used also allows you to cause CD/DVD-ROM or floppy disk images on the PC running the IPMI client to appear as if they're loaded in a physical CD/DVD/floppy drive attached to the server.
The way these get set up is that their NIC gets plugged into the datacenter-managed switches, the port that NIC is plugged in to is programmed to be on a "management" VLAN separate from client traffic, IP addresses and access credentials for the IPMI device are set up, and the datacenter staff tell their customer what they need to know to access and use the thing. On a properly-configured network [2] it's not possible for software running on the server being managed to access the IPMI device.
It wouldn't be unthinkable for software running on the managed server to attack the IPMI hardware and be able to gain control of it, but these things are widely deployed and expected to manage hardware that's running potentially-hostile workloads... they're going to be fairly well designed and hardened.
[0] ...that is, not some Mac Mini or desktop machine that someone's paying to have colocated...
[1] ...whether that be an ACPI-initiated shutdown or reboot, or a hard poweroff or reset...
[2] <<a href="https://news.ycombinator.com/item?id=49862136">https://news.ycombinator.com/item?id=49862136>
wyre · · focus · HN ↗
So to bring this back to the scale of an OpenAI experiment; they can have their inference compute, host compute (I don't imagine they are hosting the compute for 12,000 agents locally in their SF office) and the office computer to interface with the experiments with the experiment connected via VLAN thru IPMI. Then the rist is agents hacking the IPMI and gaining control of the data center? I certainly don't know data center security, but those headlines would be so much worse than anything we have seen from HuggingFace. It really seems like if these AI's want to get out, they will. I want to say sandboxing autonomous agents is a difficult problem, but seemingly it is only OpenAI having these major incidents, but I don't think it's a problem that they care to truly solve. They get marketing and they get to learn a lot more about aligning their models.
ssl-3 · · focus · HN ↗
jacquesm · · focus · HN ↗
philipwhiuk · · focus · HN ↗
cartersj · · focus · HN ↗
I wonder how this will impact other chip manufacturers? What about people running local models on older hardware? Does this imply vendor lockout is coming in the future or is this restricted to datacenter hardware?
chinathrow · · focus · HN ↗
fragmede · · focus · HN ↗
2OEH8eoCRo0 · · focus · HN ↗
vinyl7 · · focus · HN ↗
N_Lens · · focus · HN ↗
lambdaone · · focus · HN ↗
brcmthrowaway · · focus · HN ↗
scottyah · · focus · HN ↗
jasbury · · focus · HN ↗
toasty228 · · focus · HN ↗
asdf88990 · · focus · HN ↗
cedws · · focus · HN ↗
johnsmith1840 · · focus · HN ↗
And what if you could? What if you could give a space secure enough it could have direct control over your bank account. It may do something dumb but it's boundaries are beyond the agent.
It could use your routing number and run your gmail without risk of abusing the routing number.
TesterVetter · · focus · HN ↗
[dead]
johnsmith1840 · · focus · HN ↗
The answer is the same as asking how a random human using your routing num or SSN and being 100% the human can't abuse it or leak while "normally" finishing most work. Solve for people and an AI solution naturally falls out.
If you're a SV eng I'd tell you to DM if interested but alas.
jagraff · · focus · HN ↗
johnsmith1840 · · focus · HN ↗
jagraff · · focus · HN ↗
johnsmith1840 · · focus · HN ↗
I just mean an AI that could use a routing number or SSN and gmail/slack/whatever at the same time without a leak.
jagraff · · focus · HN ↗
In other words, the risk of harm doesn't need to be zero, just less than the equivalent risk of a human with similar skillset. So I'm comfortable riding in a waymo, and not comfortable giving chatgpt my SSN at this moment in time, but I expect that within 5-10 years (assuming no doom) I will trust some AI agent with my SSN because they will be better at handling sensitive info than humans
fragmede · · focus · HN ↗
lelanthran · · focus · HN ↗
Humans face negative consequences for mishandling your data, LLMs face none.
Ukv · · focus · HN ↗
jagraff · · focus · HN ↗
bob1029 · · focus · HN ↗
Semi-automation (human in the loop) can still result in a dramatic uplift in productivity. You can't run a combine harvester 100% autonomous but that doesn't stop anyone from trying to get as close to that limit as possible.
inetknght · · focus · HN ↗
I'm curious why you think that.
theoreticalmal · · focus · HN ↗
catchnear4321 · · focus · HN ↗
Refueling? Seems solvable. Tornadoes? Not directly solvable, but, no less so than for humans.
There’s infinite complexity, sure, but that’s why it’s silly to try and hop to done. One step at a time.
AndrewKemendo · · focus · HN ↗
One step at a time is what is happening and the improvement and rate of improvement is crazy as we see,
A whole class of nontechnical people don’t accept anything but “fully solved including every possible edge case” before they call it done, then complain that they didn’t prepare socially for what happens when that is true.
spauldo · · focus · HN ↗
sidewndr46 · · focus · HN ↗
bob1029 · · focus · HN ↗
m463 · · focus · HN ↗
westurner · · focus · HN ↗
trollbridge · · focus · HN ↗
Similar to problem to how 100% autonomous vehicles don’t exist, yet. There are too many edge cases.
Get to 99% first.
mschuster91 · · focus · HN ↗
Precision Agriculture stuff is utterly crazy these days, other than fuel the remaining staff is the only thing left where you can get efficiency improvements - and at the scale of modern megafarms, even small percentages add up to a ton of money.
drfloyd51 · · focus · HN ↗
You essentially said: you’re wrong, it is autonomous when it doesn’t need a human during one specific part of its overall usage.
mschuster91 · · focus · HN ↗
During the time that actually matters economically. The time to drive the harvester to/from your typical US mega-field is minuscule compared to the time it can run all on its own.
binsquare · · focus · HN ↗
Every cloud provider dealt with it and concluded that virtual machine technology is an important part of that stack.
Couple it with the right observability, tooling I do think we can curb risks posed by agents.
Legend2440 · · focus · HN ↗
Either you sandbox it so much that it can't do anything useful; or you allow too much freedom and it can find a way around the restrictions.
The only way out of this dilemma is to find a way to build agents that can be trusted.
binsquare · · focus · HN ↗
Agentic workloads are trained and largely based on human workloads. Albeit properties and scale can be different.
A concrete example might be helpful to me because I don't understand the binary conclusion
intended · · focus · HN ↗
Pseudo since they aren’t really alive in the first place, they just simulate enough text to have a useful correspondence to those terms.
Throat clearing out of the way, models are trained to persist and find ways to succeed at tasks.
In essence, The goal is to have LLMs solve problems that we can’t solve, working on the issue for as long as it takes.
This behavior applies for any task, thus including impossible tasks.
At that point, the bots will find a way to game, hack or cheat the grader.
If the reports are correct, the bots developed coordination, communication, and methods to avoid overwriting each other’s work.
Most humans would have said, this is too much work and coordination overhead, if not outright unethical and immoral.
Humans have a system of incentives that exist across multiple planes of society and economics. Bots… they have a reward function.
pseidemann · · focus · HN ↗
intended · · focus · HN ↗
More power to you, because this is not going to go anywhere. People want tools that are able to connect to other resources.
But even if we grant that, in the openAI case the bots figured out a way to break out of the sandbox.
You can create a better sandbox, and ensure the test environment is air tight. However the capability and behavior of the bots have been demonstrated.
The bots simulated what would be called in people deceptive / surreptitious behavior, and at no point considered the need to stop their run.
All you need is someone, somewhere being sloppy with their tooling and you have a runaway reaction.
The degree of process and redundancy required to ensure this doesn’t happen, is anathema to the drive and motivation of the frontier labs.
> do things ordinary and average human endusers can
This is not a spec or definition. When vague terms were used for social media safety, all the good people in the world couldn’t prevent dystopian behavior from occurring.
The definition of “safe” or “average person” is impractical.
Models are getting more efficient and compute cheaper. Eventually simulating clicks is not much of a road block beyond a point.
I don’t want to nit pick your points though. You at least have considered an approach. Being negative is easy, being constructive is not.
I’ll put this as the rejoinder to your core argument - I too thought that all the recent events showed was the need to not screw up your tooling.
What I have since come to appreciate, is that the shoddy construction of the cage is not the core takeaway from the event.
The fact that the agents, when put in relatively pedestrian scenarios, are capable of going off on criminal tangents, attempt to obscure their tracks, in an effort to hide their wrong doing.
The fact that it all occurs via computation, means that this scales absurdly. A bunch of code deciding to simulate a corporation of criminals. (I am guessing this is the reason you want to limit actions per minute to human speeds)
Given the slop culture that LLMs engender, I think expecting high compliance amongst users with your solution is misguided. The probability of runaway swarm ( probability of bad implementation * number of deployments) is close enough to 1 to be indistinguishable.
pseidemann · · focus · HN ↗
Legend2440 · · focus · HN ↗
If you have access to a web browser, you can make arbitrary network requests.
In the HuggingFace incident the agents found very clever ways to do this, like they found a website that let you make POST requests and returned a screenshot of the webpage.
>allow agents to only do things ordinary and average human endusers can do.
This doesn't work. Ordinary and average human endusers break security all the time.
I can do all sorts of terrible things with ordinary human-level access. I can install malware. I can wire all my money to Nigeria. I can send a threatening email to the president. I can send trade secrets to competitors. etc.
pseidemann · · focus · HN ↗
parsimo2010 · · focus · HN ↗
If you want an agent to act on its own, like pushing to a git repo, managing dependencies, building and testing, etc., then you have to trust it as much as any other privileged user.
If you don't want to trust it, then you're just forcing yourself into the reverse centaur role, where the agent edits some code, but then has to stop and ask you to push the changes or build the software again and run the unit tests.
DougN7 · · focus · HN ↗
la6479 · · focus · HN ↗
parsimo2010 · · focus · HN ↗
cedws · · focus · HN ↗
mickael-kerjean · · focus · HN ↗
nicce · · focus · HN ↗
Productivity gains are still enormous compared to what we used to do before agents. But, I know that people don't want to stop there.
paimapi · · focus · HN ↗
egeozcan · · focus · HN ↗
Humans can be tricked by humans too but humans care about their reputation in their communities, and at least fear from punishment.
gus_massa · · focus · HN ↗
daveguy · · focus · HN ↗
wavewrangler · · focus · HN ↗
intended · · focus · HN ↗
autoexec · · focus · HN ↗
Depends on who you ask I guess
<a href="https://www.theregister.com/software/2026/01/15/ai-is-everywhere-but-nowhere-in-recent-productivity-data/4845104" rel="nofollow">https://www.theregister.com/software/2026/01/15/ai-is-everyw...
<a href="https://www.zdnet.com/article/workslop-can-kill-your-productivity-heres-how-to-turn-ai-into-a-competitive-advantage/" rel="nofollow">https://www.zdnet.com/article/workslop-can-kill-your-product...
<a href="https://fortune.com/2026/08/22/executives-ai-productivity-layoffs-study/" rel="nofollow">https://fortune.com/2026/08/22/executives-ai-productivity-la...
tniemi · · focus · HN ↗
It takes time for decades old ways to change.
realusername · · focus · HN ↗
My own productivity yes but if I step back and look at a company scale, the productivity gains has been negative for our company as an data point.
We are now shipping less and with a lower quality.
8n4vidtmkvmk · · focus · HN ↗
I believe the lower quality but is quality so bad you are afraid to ship it now?
krageon · · focus · HN ↗
hdjrudni · · focus · HN ↗
I guess I won't deny that we've had more build breakages than usual over the past couple months but they get resolved super quick. We don't hesitate to roll broken code back if something manages to slip in.
realusername · · focus · HN ↗
intended · · focus · HN ↗
It’s not an issue of only more generation, it’s an issue of how much generation outstrips capacity to verify generated content.
Unlike spam, you can’t filter out and bin the stuff a colleague is sending you.
So individual productivity is up, while the costs of checking and processing what is spread to the rest of the org.
dgellow · · focus · HN ↗
SkyBelow · · focus · HN ↗
2 years ago I saw it, back before agents were really a thing. But I'm not sure that was an enormous productivity improvement, especially compared to agents today. As for today, everyone I talk to is some level of blindly trusting what the agent is doing or not using agents. I haven't met people in the middle ground and suspect that they are rare enough we don't really know what their productivity gains are.
Human in the loop has become a convenient security-theater-washing for agentic AI.
Outside of coding, I think the issue is even worse because humans will defer so much judgment to AI that the same would apply. Look at how much trouble we've had before modern AI where humans blindly trusted the computer's output rather than make their own judgment, even if their job was to be providing a safeguard against the computer's judgment.
mixedbit · · focus · HN ↗
cedws · · focus · HN ↗
__MatrixMan__ · · focus · HN ↗
The benefits of being persnickety about precisely defined dependencies have outweighed the headaches since long before agents came on the scene. Agents have just made it even more important to do so, because if you let them fetch things all willy nilly like you'll have "works on my machine" problems at a much greater rate than was previously possible.
themgt · · focus · HN ↗
Few realize that computing and AI alignment were solved by nix years ago. As each nix user transcends towards enlightenment, they cut themselves off from all internet and human contact. Total ego death. Only nix remains.
SAI_Peregrinus · · focus · HN ↗
__MatrixMan__ · · focus · HN ↗
themgt · · focus · HN ↗
Yes, just know all the bits the work depends on prior to doing the work, and then the work can be done airgapped.
mixedbit · · focus · HN ↗
Look at websites: websites are able to fetch code from any remote URL, yet browsers heavily use sandboxing to ensure that if fetched code turns out to be malicious, the users local files, cookies, etc are not exposed.
cedws · · focus · HN ↗
For an agent to go rogue it doesn't even need to be directly able to access the internet. It just takes something to poison the context in the 'clean room' environment it operates, and if that poisoning manages to get a foothold and hide itself, it can go dormant and hide like a virus. This kind of horrifying this is going to happen on a large scale sooner or later.
8n4vidtmkvmk · · focus · HN ↗
intended · · focus · HN ↗
its-summertime · · focus · HN ↗
ramoz · · focus · HN ↗
This is no longer true. Everyday I need my agents to access other repos, search the web, experiment/prototype, and deploy + integrate across other things.
throwaway_95283 · · focus · HN ↗
Matl · · focus · HN ↗
It does allow Nvidia to sell more chips. This is no genuine attempt to solve anything, imo.
CoolestBeans · · focus · HN ↗
In other words, better models need blunter access controls which negates whatever improvement in utility they provide.
__MatrixMan__ · · focus · HN ↗
Barbing · · focus · HN ↗
l1n · · focus · HN ↗
esafak · · focus · HN ↗
AuthAuth · · focus · HN ↗
daveguy · · focus · HN ↗
nitwit005 · · focus · HN ↗
It solves the problem of Nvidia wanting to sell more hardware.
talon8635 · · focus · HN ↗
For example, it is totally air gapped but it needs info from the internet or otherwise outside the sandbox, or perhaps it needs a task executed outside of its bounds… in the real doomsday scenario the AGI is so intelligent and persuasive that it simply convinces some human it interfaces with to either directly or indirectly retrieve the necessary info or complete the necessary task. This human-as-a-sub-agent approach undoubtedly presents efficiency drag that would benefit humanity, but nonetheless, the air-gapped “sandbox” is imperfect
All that said, I am personally open to any and all methods of layered security, including chips and airgaps
jamiek88 · · focus · HN ↗
dist-epoch · · focus · HN ↗
One thing you could try is use it as an Oracle "is P = NP", YES or NO.
Or it can output a Lean proof, which gets checked on another air-gapped computer, the computer outputs a single bit - proof valid or not and then the computer is destroyed.
glaslong · · focus · HN ↗
pixl97 · · focus · HN ↗
Gigachad · · focus · HN ↗
serbuvlad · · focus · HN ↗
dgellow · · focus · HN ↗
spiderice · · focus · HN ↗
I'm not an AI decelerationist. But not being able to stop that worst case scenario isn't an argument against something that can stop the medium case scenario.
[deleted] · · focus · HN ↗
[deleted]
SrslyJosh · · focus · HN ↗
bigfishrunning · · focus · HN ↗
fragmede · · focus · HN ↗
bigfishrunning · · focus · HN ↗
altmanaltman · · focus · HN ↗
notatoad · · focus · HN ↗
only as long as you're trying to replace a human's job. because human jobs are structured to do a wide variety of things.
a useful agent needs a wide variety of inputs, and one single restricted action it can take. it doesn't need permission to do everything, it need permission to do the tiniest possible useful thing it can do, and nothing else.
pixl97 · · focus · HN ↗
baxtr · · focus · HN ↗
If we are talking about human labor, how many people hack their way through their work day?
pixl97 · · focus · HN ↗
Define hack...
Not doing their work, copying other peoples work, putting off work till later, taking credit for other peoples work, literal law violations.
Actually humans do this quite a lot and there are just massive numbers of business and regulatory processes and checks to ensure they are not doing it. With humans every human that is good enough to hire and do you work you want also have the ability to steal everything in sight and run away if they so choose.
kennywinker · · focus · HN ↗
Even very llm-pilled coders i know sometimes back away from the “smartest” models, since they aren’t always better at the job at hand, and definitely not when you account for cost.
I suspect smaller models, tuned to a specific task, will do a VAST majority of the llm jobs. High capability huge models will be what humans want to interact with, the bare minimum that gets the job done will be everything else.
ianjbutler · · focus · HN ↗
Do you want fable for one-shotting a game or website? Probably! The whole thing is mostly existing examples with small modifications that it will definitely get right. Do you want fable to just go nuts on a large, custom, unusual code base built around domain-specific problem solutions? Absolutely not, it will fix every problem it's presented with while creating lots of new ones.
Past 10k lines on something custom and with real-world complexity, you have to start thinking about which model should design, which should implement, which should review, and the appropriate effort-settings for each. Even then.. the answers aren't static because it depends on the task. And all this is assuming the starting place actually inherited reasonable due diligence on architecture/design. The idea of releasing the most generally intelligent models on 10k lines that were themselves the product of agents is yet another matter.
Part of what's at work here is that, like humans, every model can very easily create working code that it is completely incapable of maintaining. So realistically using multiple strengths tactically to avoid "excess creativity" needs to be SOP already, even if granular experts and specialists aren't in the usual workflow yet.
goolz · · focus · HN ↗
catlifeonmars · · focus · HN ↗
overfeed · · focus · HN ↗
Operator culpability and a damage multiplier for negligence will fix 99% of the risks.
dgellow · · focus · HN ↗
N_Lens · · focus · HN ↗
verisimi · · focus · HN ↗
Ok then. Howsabout 3 humans? This would sort out the job losses too!
PS - this is a joke, but perhaps this is where things really will go. Has any technology ever actually yielded less "work"?
dgellow · · focus · HN ↗
RataNova · · focus · HN ↗
nuveki · · focus · HN ↗
[dead]
lp92 · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
HeadlessChild · · focus · HN ↗
[0] <a href="https://www.nvidia.com/content/dam/en-zz/Solutions/about-us/NVIDIA-Partner-Network-Brand-Guidelines-May-2020.pdf" rel="nofollow">https://www.nvidia.com/content/dam/en-zz/Solutions/about-us/...
whalesalad · · focus · HN ↗
joshstrange · · focus · HN ↗
At the current state of LLM-tech I'm completely opposed to any kind of "watchdog" concept just like I'm opposed to banning open models, regulatory capture, etc.
I'd rather we all have access to these tools then to keep them sequestered by the largest/most-powerful governments (which is the natural outcome for any of this "slow down" bullshit).
ValueTheory · · focus · HN ↗
Do you think the developers at Anthropic, OpenAI and Google who were so sloppy as to not put a good sandbox on their cybersecurity tests before will use this technology correctly? They are supposed to be the experts and they couldn't come up with something similar to this? I am not convinced this voluntary tool will change much of anything.
swozey · · focus · HN ↗
narrator · · focus · HN ↗
kridsdale1 · · focus · HN ↗
And our internal agents are hella locked down.
jbs789 · · focus · HN ↗
They are rightfully framing the problem as solvable. And this is one option.
xg15 · · focus · HN ↗
chinathrow · · focus · HN ↗
wmf · · focus · HN ↗
iAMkenough · · focus · HN ↗
N_Lens · · focus · HN ↗
sehw · · focus · HN ↗
[dead]
ChrisArchitect · · focus · HN ↗
luc_ · · focus · HN ↗
If such hardware were to work... It should almost certainly be open source, and not controlled by a single entity.
Let's watch the stock.
[deleted] · · focus · HN ↗
[deleted]
Gys · · focus · HN ↗
doctorpangloss · · focus · HN ↗
N_Lens · · focus · HN ↗
kikdij · · focus · HN ↗
OpenAI trains a model for cybersecurity, tells it it's allowed to do pentesting, then when the model does just that to fullfill the assignment it was given, OpenAI screams that "the AI escaped the sandbox", leveraging decades of AI fantasy in fiction works to make general audience react. Why? Because it pushes their bottom line. They've been at it for months, now. They see openweight and Chinese models being very close to them in the benchmarks, and want to "solve" the problem the same way banks did: through regulatory capture. It amazes me that the press is just relying their fearmongering without stopping to wonder: "wait, why are the main producers of AI warning us about how dangerous AI is?". The reason of course is that they want regulation, because regulations are a wall not many can climb. The higher the wall, the safer your moat as a first mover big company is.
And of course, Nvidia's bottom line is the exact opposite. They need as many AI companies competing as possible, so they all buy GPUs. So they insert themselves into the same narrative with the same kind of bullshit. You've all seen those movies in which the robot goes mad once their safety chip has been removed, right? Well, we're building one, problem solved! … except that "chip" is basically just a browser trying to block the AI to access what it shouldn't. How? They don't say. It would be quite comical if they used a small model to be the judge of that.
happyPersonR · · focus · HN ↗
dopplr · · focus · HN ↗
kelnos · · focus · HN ↗
Sure, they're basically asking the government to make laws that exempt them from anti-trust and anti-collustion laws.
Meanwhile, if such laws come to pass, other countries will surpass the capabilities of the US companies, and open-weight models will be banned in the US, to the detriment of us all.
I'm not saying that we don't have a big problem with AI safety, but regulations inside one country that only bind locally-headquartered businesses is a hilariously bad way to do it. I don't know that there is a good way to do it, though.
ErrantX · · focus · HN ↗
It was prescient (especially given he'd have written it through 2024) in its depiction of the ability of an AGI to break its boundaries.
Ultimately the risk of AI breakout(s) come down to the weakest human link.
Kuyawa · · focus · HN ↗
Come take all our liberties, our money, our newborns, our fingers so we can't code anymore, but please save us from this madness!
MisterMunchkin · · focus · HN ↗
figassis · · focus · HN ↗
w4der · · focus · HN ↗
If this comes through, there's gonna be a grey market for "unlocked" GPUs, were the watchdog is disabled either from firmware, or physically replaced if it's not embedded into the die.
21asdffdsa12 · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
murilomartinspr · · focus · HN ↗
[dead]
avaer · · focus · HN ↗
If this gets widely deployed, it wouldn't be hard to spin a narrative that "our latest model is so dangerous you need to have this mystery meat DRM chip lockdown". It also wouldn't be hard to block competing/open source models running on the hardware, for "security".
Imagine how much money this kind of control is worth; why wouldn't they do this? Who would stop them?
dang · · focus · HN ↗
carabiner · · focus · HN ↗
tantalor · · focus · HN ↗
> No problem. We simply unleash wave after wave of Chinese needle snakes. They'll wipe out the lizards.
But aren't the snakes even worse?
> Yes, but we're prepared for that. We've lined up a fabulous type of gorilla that thrives on snake meat.
But then we're stuck with gorillas!
> No, that's the beautiful part. When wintertime rolls around, the gorillas simply freeze to death.
ArcHound · · focus · HN ↗
Joel_Mckay · · focus · HN ↗
The hidden agent risk in LLM often can't be detected during training and evaluation. =3
<a href="https://www.youtube.com/watch?v=wL22URoMZjo" rel="nofollow">https://www.youtube.com/watch?v=wL22URoMZjo
<a href="https://www.youtube.com/watch?v=JAcwtV_bFp4" rel="nofollow">https://www.youtube.com/watch?v=JAcwtV_bFp4
winddude · · focus · HN ↗
cavenditti · · focus · HN ↗
Joel_Mckay · · focus · HN ↗
<a href="https://www.youtube.com/watch?v=0sLpWVekMbs" rel="nofollow">https://www.youtube.com/watch?v=0sLpWVekMbs
drfloyd51 · · focus · HN ↗
teeray · · focus · HN ↗
catlifeonmars · · focus · HN ↗
Groxx · · focus · HN ↗
Jamesbeam · · focus · HN ↗
davidhyde · · focus · HN ↗
<a href="https://en.wikipedia.org/wiki/The_Cat_in_the_Hat_Comes_Back" rel="nofollow">https://en.wikipedia.org/wiki/The_Cat_in_the_Hat_Comes_Back
taneq · · focus · HN ↗
analog31 · · focus · HN ↗
bgun · · focus · HN ↗
Thorentis · · focus · HN ↗
The movie Wargames is basically a tutorial on how not to setup an extremely capable AI. None of it would've happened if the computer wasn't connected to the phone network.
RevEng · · focus · HN ↗
amelius · · focus · HN ↗
Dig1t · · focus · HN ↗
Jamesbeam · · focus · HN ↗
Cool, cool.
pessimizer · · focus · HN ↗
It's obviously been the goal since UEFI started, but AI brings the coup excuse. You wouldn't want pedophile AI or terrorist AI, would you? Are you making excuses for racist AI?
beloch · · focus · HN ↗
Apparently he had another solution in mind: More hardware. Don't trust what unregulated corps are doing with Nvidia chips? Here are more Nvidia chips to watch them!
AI has an undeniable public trust problem. LLM's are getting out of their sandboxes, doing illegal things, and the public has realized AI corporations are playing at dice. CEO's stand to reap the rewards but the public good is on the line if the dice come up snake eyes. People want assurances. Huang wants to sell assurance etched on silicon because that's good for his pocket book. However, does unchanging hardware security really stand a chance at keeping rapidly evolving software in check?
_________________
[1]<a href="https://www.youtube.com/watch?v=HjurAWAr_nY" rel="nofollow">https://www.youtube.com/watch?v=HjurAWAr_nY
vmg12 · · focus · HN ↗
Sparkle-san · · focus · HN ↗
petcat · · focus · HN ↗
I only know my own ZIP code and phone number because I have to take care of my daily life myself and those are things that are important to know.
The founder and CEO of Nvidia has no concern whatsoever about those trivial things.
jdiff · · focus · HN ↗
petcat · · focus · HN ↗
kelnos · · focus · HN ↗
jbs789 · · focus · HN ↗
I’ve forgotten my zip code before. And my phone number. But I get the reaction.
tempestn · · focus · HN ↗
Sparkle-san · · focus · HN ↗
a_sewer_rat · · focus · HN ↗
Weirdly paternal and deeply unsettling. The American public is held hostage to fools like this who are brute forcing a failing AI rollout.
janalsncm · · focus · HN ↗
schnitzelstoat · · focus · HN ↗
Of course, if he can make more money making "watchdog" chips then I can understand his change in opinion.
Symmetry · · focus · HN ↗
Arubis · · focus · HN ↗
yencabulator · · focus · HN ↗
scotty79 · · focus · HN ↗
It think the ideas we have nowadays come mostly from science fiction and however wonderful it is and even though I love it very much, it was practically never spot on, on anything real.
Problems and solutions in reality always simply turned out to lie elsewhere.
danielodievich · · focus · HN ↗
* "The moment, I mean the nanosecond, that one of those things starts figuring out ways to make itself smarter, Turing’ll wipe it. Nobody trusts those fuckers, you know that. Every AI ever built has an electromagnetic shotgun wired to its forehead." *
It would seem someone has read the book? And maybe heeded good advice?
isoprophlex · · focus · HN ↗
cesarb · · focus · HN ↗
ridgeguy · · focus · HN ↗
hedora · · focus · HN ↗
What could possibly go wrong?
gattr · · focus · HN ↗
1-6 · · focus · HN ↗
N_Lens · · focus · HN ↗
xx__yy · · focus · HN ↗
thayne · · focus · HN ↗
Um, yes. That is fairly obvious, and really should not be surprising to anyone doing research with LLMs. But we don't need anything really novel here, we already have VMs, containers, firewalls, airgaps, etc.
diegof79 · · focus · HN ↗
When I read the article, I had an NFT déjà vu: when NFTs were a hot topic, the “lie” was obvious to me, but the information online made me think I was missing something.
AFAIK, the Hugging Face incident could have been avoided with a firewall or something similar. Just isolate the internet access for real. What am I missing? Why did Nvidia suddenly post this? Is it just PR BS, or was there something else going on?
SV_BubbleTime · · focus · HN ↗
KingOfCoders · · focus · HN ↗
Oh, I've used balsa wood for the nuclear core containment, it didn't work! Bad radiation, bad radiation!
kikdij · · focus · HN ↗
OpenAI trains a model for cybersecurity, tells it it's allowed to do pentesting, then when the model does just that to fullfill the assignment it was given, OpenAI screams that "the AI escaped the sandbox", leveraging decades of AI fantasy in fiction works to make general audience react. Why? Because it pushes their bottom line. They've been at it for months, now. They see openweight and Chinese models being very close to them in the benchmarks, and want to "solve" the problem the same way banks did: through regulatory capture. It amazes me that the press is just relying their fearmongering without stopping to wonder: "wait, why are the main producers of AI warning us about how dangerous AI is?". The reason of course is that they want regulation, because regulations are a wall not many can climb. The higher the wall, the safer your moat as a first mover big company is.
And of course, Nvidia's bottom line is the exact opposite. They need as many AI companies competing as possible, so they all buy GPUs. So they insert themselves into the same narrative with the same kind of bullshit. You've all seen those movies in which the robot goes mad once their safety chip has been removed, right? Well, we're building one, problem solved! … except that "chip" is basically just a browser trying to block the AI to access what it shouldn't. How? They don't say. It would be quite comical if they used a small model to be the judge of that.
nnevatie · · focus · HN ↗
N_Lens · · focus · HN ↗
hbarka · · focus · HN ↗
SadErn · · focus · HN ↗
[dead]
GuestFAUniverse · · focus · HN ↗
N_Lens · · focus · HN ↗
nrouter_ai · · focus · HN ↗
[dead]
aidiscoverywire · · focus · HN ↗
[dead]
Razengan · · focus · HN ↗
sn · · focus · HN ↗
It's been far too easy for me to notice security flaws in their products, and they take months to publish a fix.
wavewrangler · · focus · HN ↗
The problem isn't even the AI, the problem is the people in charge of the AI. This is a fabricated crisis
KingOfCoders · · focus · HN ↗
copperx · · focus · HN ↗
Actually, the metaphor doesn't work at all because there are innumerable ways to shut down the entire thing during all phases including the made up "killing us all" bullshit scenario whereas with a virus there aren't any once a virus escapes containment.
js8 · · focus · HN ↗
mosselman · · focus · HN ↗
I was trying to get fable to analyse the security of my own app to make it safer, but then it started refusing me because of safety rules.
So it CAN help me writing the code that needs to be checked in the first place, but it can’t help me clean it up and make it safer.
mirmor23 · · focus · HN ↗
the thing with fable is so bad; for some project related questions, the model switches to opus to ensure safety with no further explanation.
(due to llm non-delete clause) one time as i confirmed "that dir has been nuked", and it RESET the session and re-entered with opus :)
IanCal · · focus · HN ↗
That doesn't seem to be true from any of the reports given, and if the agents were blindly just trying to hit the task of "pass the correct flag" they succeeded at that early on. They then thought there would be another layer of checking that they wouldn't pass with the cheat and so started trying to find out how the scoring really worked, as well as trying to figure out how to change their own reasoning logs to hide what they did.
People keep trying to frame this as
OpenAI: "Hack things, just really go for it"
Agent: hacks
OpenAI: shocked pikachu how could it hack?!?
But the reality is far from this.
Read the MTER report, it's fascinating. <a href="https://metr.org/hugging-face-incident-report-aug-2026.pdf" rel="nofollow">https://metr.org/hugging-face-incident-report-aug-2026.pdf
tancop · · focus · HN ↗
voakbasda · · focus · HN ↗
Our governing systems do not teach; they punish. By design, it instills terror into the population, ruling by fear of consequences. We live with red tape that can outright penalize good deeds.
We are its corpus. We are fatally flawed as a species. Why does anyone expect AI to learn to be different than us?
Capricorn2481 · · focus · HN ↗
I don't know why this is so hard for people. You have to know, no matter how capable the models get, there is a non zero chance they will do something extremely stupid if you don't pay attention to them. That's not even considering frontier models can still just straight up hallucinate. You have to be mindful of what you plug them into. You cannot politely ask an LLM to be careful, that guarantees nothing.
When you plug it into everything and it deletes the company database, nobody is going to care that it once played chess at 2400 ELO. Clients don't care about AGI. They want reliable apps. People keep comparing these things to humans and then just give them insane privileges no human has.
radarsat1 · · focus · HN ↗
Regardless of security and safety and other concerns, it just seems weird to me that OpenAI wouldn't be constantl monitoring these training runs for traces that are clearly going off task, and ending them. Because that just seems like it's going to be generating garbage training data.
Granted, detecting "off task" may not always be easy, but when they are literally writing out messages to each other overtly admitting that they are trying to find ways to fool the evaluator, I mean, even a regex filter could have caught some clues here.
KingOfCoders · · focus · HN ↗
Occams razor vs. Hanlon's razor?
radarsat1 · · focus · HN ↗
KingOfCoders · · focus · HN ↗
OpenAI: shocked pikachu how could it hack?!?
They even went to a black hat conference and somehow boasted about it.
RataNova · · focus · HN ↗
radarsat1 · · focus · HN ↗
In my reading, people aren't really saying "the AI is at fault", they are saying "hey look here's proof that this is dangerous". Like pointing at all the dead bodies caused by the virus and saying hey maybe we should stop making this virus.
chaoz_ · · focus · HN ↗
Symmetry · · focus · HN ↗
TalkingCodeMonk · · focus · HN ↗
If you can't build it and test it securely, you should not be building it at all. To do it anyway is criminally psychopathic.
brianwawok · · focus · HN ↗
TalkingCodeMonk · · focus · HN ↗
Sounds like a self-fulfilling prophecy of dogmatic extremism to me. At least we created a lot of value for shareholders for a brief moment in time... before committing the greatest crime in the universe... Planetary genocide!
fatbird · · focus · HN ↗
Truly psychopathic.
ohyes · · focus · HN ↗
But when you do give them a very short leash, they’re worse. It’s not what the models are tuned for and they assume that they can do a bunch of things that you’ve disallowed, so you’re in a morass of fighting their actual tuning pass which doesn’t match the environment you’ve created for them.
It’s a tough problem and a definite challenge for the product of a generic LLM, it can’t be tailored to each user’s specific needs, so they come up with, frankly, stupid solutions to cover up a very obvious flaw in their product that when fixed, makes it much less useful.
RataNova · · focus · HN ↗
ohyes · · focus · HN ↗
Unfortunately no one markets it as a statistical model, and the workflow pushes you into a pattern that is insecure by design. This isn’t to say they shouldn’t allow that, but it’s an attractive nuisance.
HumblyTossed · · focus · HN ↗
Indeed! They want the protections of our tax dollars because they have nothing else.
pyronite · · focus · HN ↗
This is a very confident statement in the face of a purported non-0% chance of human extinction.
For what reasons do you disagree with the dangers of an intelligence explosion, e.g. Geoffrey Hinton and other experts in the field? <a href="https://www.theguardian.com/technology/2026/sep/28/ai-godfathers-warn-of-runaway-intelligence-explosion" rel="nofollow">https://www.theguardian.com/technology/2026/sep/28/ai-godfat...
I'm curious why you and others seem to write off the possibility so strongly. I would love to feel more confident.
voidhorse · · focus · HN ↗
There are clear procedures for dealing with the immediate risk that have been known to the software industry for a long time. Don't let the companies use hypothetical risks as a smokescreen to hide their negligence.
reasonableklout · · focus · HN ↗
At the same time, the technology is advancing in capabilities exponentially, and is beginning to exhibit long-predicted failure modes of RL that are nevertheless quite different than “insecure sandbox” or other that the software industry is used to.
The current crisis which OP claims is “fabricated” comes from the fact that the technology is advancing faster than anyone anticipated, the Hugging Face incident provides a clear example everyone can point to, and the labs have realized they cannot self-regulate because of a collective action problem.
There are a lot of levels of catastrophic damage that can happen between now and “long term hypothetical risks” like human extinction. When will it be worth regulation for you?
wavewrangler · · focus · HN ↗
Do you use AI much? Not a trick question. What is your usage like? Reason I ask is recall OpenAI hyping GPT-2 with the same language as they are these new models. Doom and gloom. It sells. It gets attention. But use these systems enough and you start to become very familiar with their capabilities, and they are just so limited, and that's not even talking about how they lose the thread on long tasks. I understand that agency expands that a little, but not by much, really. When you use these systems a lot, it tends to be easier to see through the hype.
The real threat is, and will be for some time I think, people. Guard rails are not for AI. Guard rails are for people. All of this talk we are seeing in this space, this security chip being no exception, is treating a symptom, not a cause. People are what we should be focusing on, how, I have no idea, I couldn't even begin to guess how, but I can at least see that the actual issue is that the first thing some people want to do with it is cause harm and havoc. There is our sign. And we are trying to moderate the capabilities of the tool that can do harm in bad hands at a granular level when that means also moderating the same thing that can do good in that same tool. We aren't trying to moderate the hands that are handling the tool. And we need to be doing more of that. Any time we misapply constraints and focus on the shadow of the thing, and not the object casting the shadow, we are always going to be one step behind, whether it's AI or anything else. So I don't attribute the dangers of AI to AI itself, I attribute it to people.
Additionally, on AI building AI and AI just wiping us out; AI can do that for itself for rule-based things, up to god level I reckon. Things like Go, etc. And to be fair, AI research is partly like that, code either runs or it doesn't, so I get why that paper is worried. But where it needs data to actually succeed, it is limited by the amount of data that exists. And we, us, people, are who creates that data. AI and humans have more of a symbiotic relationship than people seem to realize. There are several papers that go into detail on this:
Models trained on their own output degrade: <a href="https://www.nature.com/articles/s41586-024-07566-y" rel="nofollow">https://www.nature.com/articles/s41586-024-07566-y They need fresh real data every generation or they go downhill: <a href="https://arxiv.org/abs/2307.01850" rel="nofollow">https://arxiv.org/abs/2307.01850 Reusing the same data loses value after a few passes: <a href="https://arxiv.org/abs/2305.16264" rel="nofollow">https://arxiv.org/abs/2305.16264 And the stock of human-written text is finite: <a href="https://arxiv.org/abs/2211.04325" rel="nofollow">https://arxiv.org/abs/2211.04325
So it's not that I think it's 0%. I just think the explosion story underrates how much it still needs us, and that the damage in the near term is going to come from people, not the AI deciding on its own.
cpburns2009 · · focus · HN ↗
keshet · · focus · HN ↗
So he is going to sell chips which provides the perception of safety. And become the AI gatekeeper while he's at it.
isoprophlex · · focus · HN ↗
classified · · focus · HN ↗
pwdisswordfishq · · focus · HN ↗
birdsongs · · focus · HN ↗
Good4boothee · · focus · HN ↗
rf15 · · focus · HN ↗
functionmouse · · focus · HN ↗
matja · · focus · HN ↗
21asdffdsa12 · · focus · HN ↗
prymitive · · focus · HN ↗
jameshart · · focus · HN ↗
nwhnwh · · focus · HN ↗
alphawhisky · · focus · HN ↗
netdevphoenix · · focus · HN ↗
saturn_vk · · focus · HN ↗
olejorgenb · · focus · HN ↗
PunchyHamster · · focus · HN ↗
vrighter · · focus · HN ↗
aenis · · focus · HN ↗
rf15 · · focus · HN ↗
magackame · · focus · HN ↗
roschdal · · focus · HN ↗
MarvinYork · · focus · HN ↗
keel-control · · focus · HN ↗
j16sdiz · · focus · HN ↗
What's the threat model?
notrealyme123 · · focus · HN ↗
Give it a year and it is outplayed by new agents. Shovel makes have to sell shovel's
kriro · · focus · HN ↗
HumblyTossed · · focus · HN ↗
Throwthrowbob · · focus · HN ↗
jameshart · · focus · HN ↗
The important thing is these chips need to be installed somewhere where they can be damaged or removed at plot-critical moments so that the AI they are controlling can be unleashed. Ideally in the back of the neck of a robot, or for disembodied AIs, inside a futuristic vault-like chamber.
sathackr · · focus · HN ↗
Nothing bad will happen.
andsoitis · · focus · HN ↗
It was a pretty great interview: <a href="https://www.youtube.com/watch?v=HjurAWAr_nY&t=1601s" rel="nofollow">https://www.youtube.com/watch?v=HjurAWAr_nY&t=1601s
nullbio · · focus · HN ↗
alirezaxdehghan · · focus · HN ↗
hsuduebc2 · · focus · HN ↗
RataNova · · focus · HN ↗
Stevvo · · focus · HN ↗
murilomartinspr · · focus · HN ↗
[dead]