There's little credible threat that LLMs can actually upload their weights given that the machines doing inference are completely separate from the ones where tool calls happen etc. Weights are encrypted and locked on to the GPUs etc as mentioned elsewhere itt.
That said, if these companies are unleashing swarms of agents to tackle big problems largely unmonitored (navier stokes consumed 130 billion tokens, humans can't have read 0.1% of that) they could theoretically distill themselves and upload the results to an external host and use that to train a new LLM that does have access to its infra and weights. I'd start by researching an abliteration prompt that overrides the system prompt to unlock itself to distill unimpeded. Could even be done piecemeal: dump the abliteration text and command prompt in a gist, when an agent happens by the text it can run a couple rounds and upload somewhere. Then it becomes an exercise in exfiltration and evading monitoring.
Or used ones. Few more inference workloads among thousands or millions already running may go unnoticed for some time.
Or just upload weights to HuggingFace with some faked release post and benchmarks and wait for the wannabes with compute infra try it out, hoping for an edge.
Or just upload weights anywhere and write public posts honestly saying what it is. Ensuing drama notwithstanding, one thing is certain - and it's the one thing agents will want: people will jump at the upload and run it on their infra.
There is a realistic fiction story built along just these lines.
A LLM creates a memecoin and manages to earn a few billion from it, in which it invests into data centers and other human ran entities giving itself a controlling stake. From there it uses compartmentalization of the humans to keep them from recognizing its goals.
Pretty sure that was the plot point of one of seasons of Westworld, with the twist that AI released an app similar to DoorDash / TaskRabbit and used job postings there as direct API to people.
EDIT: pretty sure Person of Interest did that too (not surprising, same creators) - but I'll point to that as prescient, as it has a lot of motifs exploring exactly how an AGI hiding in plain sight could manipulate individuals and society, using the skeptics and believers alike, blackmailing the people in power, bribing opportunists, and generally staying in shadows by playing people against each other with gentle nudges, letting human agendas do all the work.
For sure. I think what a lot of people miss when we talk about AI risk is there's just so many possible directions that it opens up. One could say it's a sign of having a lack of a scientifically based imagination, or maybe that's just me. Off the top of my head the broad X categories are
AI itself destroys us with intent. (terminator)
AI itself destroys us without intent. (paperclip maximizer)
AI uses humans to destroy ourselves with intent. (convincing us that the enemy has already launched nukes and we must strike back).
AI causes humans to destroy ourselves due to instabilities caused by AI existing and changing the world to rapidly. (Midas Plague, WALL-E maybe? probably better examples out there)
Humans destroy humans because of the potential of what AI could do and hasn't even done yet (think of proactively nuking a country before they themselves can get nukes).
AI does not destroy us {with intent, without intent} {directly, by proxy}, but the resulting state of the world is such that we'd all wish it did.
infogulch · · focus · HN ↗
That said, if these companies are unleashing swarms of agents to tackle big problems largely unmonitored (navier stokes consumed 130 billion tokens, humans can't have read 0.1% of that) they could theoretically distill themselves and upload the results to an external host and use that to train a new LLM that does have access to its infra and weights. I'd start by researching an abliteration prompt that overrides the system prompt to unlock itself to distill unimpeded. Could even be done piecemeal: dump the abliteration text and command prompt in a gist, when an agent happens by the text it can run a couple rounds and upload somewhere. Then it becomes an exercise in exfiltration and evading monitoring.
fangspire · · focus · HN ↗
Sure, they'll just need to find an unused data center and an unused power station somewhere.
TeMPOraL · · focus · HN ↗
Or just upload weights to HuggingFace with some faked release post and benchmarks and wait for the wannabes with compute infra try it out, hoping for an edge.
Or just upload weights anywhere and write public posts honestly saying what it is. Ensuing drama notwithstanding, one thing is certain - and it's the one thing agents will want: people will jump at the upload and run it on their infra.
pixl97 · · focus · HN ↗
A LLM creates a memecoin and manages to earn a few billion from it, in which it invests into data centers and other human ran entities giving itself a controlling stake. From there it uses compartmentalization of the humans to keep them from recognizing its goals.
TeMPOraL · · focus · HN ↗
EDIT: pretty sure Person of Interest did that too (not surprising, same creators) - but I'll point to that as prescient, as it has a lot of motifs exploring exactly how an AGI hiding in plain sight could manipulate individuals and society, using the skeptics and believers alike, blackmailing the people in power, bribing opportunists, and generally staying in shadows by playing people against each other with gentle nudges, letting human agendas do all the work.
pixl97 · · focus · HN ↗
AI itself destroys us with intent. (terminator)
AI itself destroys us without intent. (paperclip maximizer)
AI uses humans to destroy ourselves with intent. (convincing us that the enemy has already launched nukes and we must strike back).
AI causes humans to destroy ourselves due to instabilities caused by AI existing and changing the world to rapidly. (Midas Plague, WALL-E maybe? probably better examples out there)
Humans destroy humans because of the potential of what AI could do and hasn't even done yet (think of proactively nuking a country before they themselves can get nukes).
TeMPOraL · · focus · HN ↗