No one can control any AI model. It will never be controlled. These models are based on a huge amount of data, it's just gonna be impossible to control the output that is based on that data only with a system prompt or some other injection mechanism.
The model is just a powerless token generator without a harness. If you give the model a harness which you choose to exercise no control over, can you say that it can't be controlled?
Inform yourself by reading the METR analysis of the HuggingFace incident.
Agents simply broke out of their environment. And this can't be discarded anymore by assuming that it's just a poorly configurend jail, because agents are becoming better and better at escaping.
In short: on a large enough scale and timeline, the possibility of constrain AIs approaches zero.
Bonus: what many people don't know is that agents also hacked in the internal OpenAI network. Crazy times.
Wouldn't this mean better sandboxes are needed for some things, for example (might include very strong airgaps even)? Breaking out of something isolated electromagnetically, optically, and acustically is not easy.
Could sit in the box and interact if a model of certain capabilities is needed/tested. We do physical security for other things, too. Not saying everything needs that type of isolation.
thewhitetulip · · focus · HN ↗
Seems like there are no guardrails on LLMs
worldsavior · · focus · HN ↗
dns_snek · · focus · HN ↗
pizza234 · · focus · HN ↗
Agents simply broke out of their environment. And this can't be discarded anymore by assuming that it's just a poorly configurend jail, because agents are becoming better and better at escaping.
In short: on a large enough scale and timeline, the possibility of constrain AIs approaches zero.
Bonus: what many people don't know is that agents also hacked in the internal OpenAI network. Crazy times.
RandomLensman · · focus · HN ↗
ogogmad · · focus · HN ↗
RandomLensman · · focus · HN ↗