The challenge here is in trying to decide whether LLMs are intelligent or have a mind. Famously, the criteria for intelligence seem to slip with each advancement in technology. But going back to Turing, his test was actually more carefully phrased than we remember: he said that when machines could pass the test, the question of whether they are intelligent would become moot. That seems to be what we're actually seeing: if people can't tell the difference, it kind of won't matter whether they're "truly intelligent" or not.
I'm a bit more prosaic. I think if we engineered ways for LLMs to begin conversations, rather than just respond, we'd be more open to the concept of their intelligence. Without perceived "will" to do things, they operate as a next-gen search engine or encyclopedia.
We are way past that point, any harness can trivially make LLMs start conversations or pursue goals. An encyclopaedia wouldn't have hacked Huggingface on its own.
Ah, are you saying that because most people don’t interact with agents, they aren’t aware that LLMs can have initiative and pursue goals?
I think the line is blurring though, mainstream chat interfaces are adding more and more “agentic” features.
ChatGPT will happily execute code in a sandbox, search the web and design downloadable PDFs purely through the standard OpenAI chat interface. They can also send you emails or do tasks on a repeated schedule.
It's more like that people's typical experience of LLMs doesn't go beyond human-initiated conversations or conversations triggered on cron or some obvious event handler coded in deterministic/"legacy"/"boring" way. Most of us, I believe, also try and steer agents away from messaging other people when such possibility exists.
It would be interesting if we didn't - if it became common that AI, in the middle of some task, starts chatting with people to e.g. gather more context. The perception of those "third parties" may suddenly become different - an agent striking conversation first, obviously pursuing some agenda of its own that it's not completely sharing, and communicating on its own schedule that's clearly not just a hook firing on timer or pattern-match, and not random, but visibly causally related to things happening at work in broader context.
sethev · · focus · HN ↗
tomrod · · focus · HN ↗
joefourier · · focus · HN ↗
tomrod · · focus · HN ↗
The hacking agents being tested have goals beforehand, from the frontier lab or from a superior agent, that they execute immediately.
But the perceived experience most people have is a chatbot, which is the encyclopedia form.
joefourier · · focus · HN ↗
I think the line is blurring though, mainstream chat interfaces are adding more and more “agentic” features.
ChatGPT will happily execute code in a sandbox, search the web and design downloadable PDFs purely through the standard OpenAI chat interface. They can also send you emails or do tasks on a repeated schedule.
TeMPOraL · · focus · HN ↗
It would be interesting if we didn't - if it became common that AI, in the middle of some task, starts chatting with people to e.g. gather more context. The perception of those "third parties" may suddenly become different - an agent striking conversation first, obviously pursuing some agenda of its own that it's not completely sharing, and communicating on its own schedule that's clearly not just a hook firing on timer or pattern-match, and not random, but visibly causally related to things happening at work in broader context.
tomrod · · focus · HN ↗