I don't care if it's "intelligent", I don't care if it "has a mind". I don't care if it is "really reasoning", I don't care if it "understands". I don't care if it is "sentient" or "conscious".
None of this matters for the practical outcome.
You'd think that this has been understood over the last 4 years, but apparently it keeps circling back to this.
[Edit: I see that it was written back in 2023. Then (2023) should be added in the submission title]
If it generates functional output that works, then it works. And it works. It's not a psychic's con when it outputs Lean-verified proofs. It isn't a con when it can find and exploit zero-days.
The OP is still in the "denial" phase. Most I see are already in "anger" (a blurry fury against everything AI-shaped, from vague reasons piling on all "bad stuff" political reasons they already hated before) or "bargaining" (mathematicians scrambling to come up with a new definition of their job and retcon that it was always the main part anyway). A few are already in "depression" and feel like spectators on the Titanic, and the tiniest sliver is at "acceptance" with some kind of well-informed plan for their future.
My biggest issue isn't being too agreeable (ie the psychic con), it's being confidently wrong, including outright hallucinations.
If you ask a common question to an LLM with unusual qualifiers, it tends to ignore the qualifiers and give you the typical answer. I saw a demonstration of this with the whole "the surgeon is my mother" "puzzle" that people use to expose implicit gender bias (ie where they assume the surgeon is a man). Ask variations of this and it'll keep going back to the standard form.
Another one I saw was multiplying large numbers. The starting and ending digits tended to be correct but the middle digits were wrong. Why? Because it's really not doing multiplication at all. It's looking for statistical answers. It's unlikely to have met the exact pair of very large numbers you're multiplying before.
Now pundits will argue that all of these are solvable problems and individually they are. But my suspicion is that there will be a neverending stream of such edge cases and it'll be impossible to trust an LLM's output unless you are knowledgeable enough to fact check it yourself.
Now if your example of identifying zero days, this comes up with what I can only describe as "light positives", meaning it's technically a bug but essentially impossible to exploit. IIRC this came up with the demonstration where someone pointed Fable at some BSD code. I'm not sure if there have been any true false positives and obviously false negatives are impossible to know.
I guess my point is that I think LLMs are way more limited than a lot of people think.
bonoboTP · · focus · HN ↗
None of this matters for the practical outcome.
You'd think that this has been understood over the last 4 years, but apparently it keeps circling back to this.
[Edit: I see that it was written back in 2023. Then (2023) should be added in the submission title]
If it generates functional output that works, then it works. And it works. It's not a psychic's con when it outputs Lean-verified proofs. It isn't a con when it can find and exploit zero-days.
The OP is still in the "denial" phase. Most I see are already in "anger" (a blurry fury against everything AI-shaped, from vague reasons piling on all "bad stuff" political reasons they already hated before) or "bargaining" (mathematicians scrambling to come up with a new definition of their job and retcon that it was always the main part anyway). A few are already in "depression" and feel like spectators on the Titanic, and the tiniest sliver is at "acceptance" with some kind of well-informed plan for their future.
jmyeet · · focus · HN ↗
If you ask a common question to an LLM with unusual qualifiers, it tends to ignore the qualifiers and give you the typical answer. I saw a demonstration of this with the whole "the surgeon is my mother" "puzzle" that people use to expose implicit gender bias (ie where they assume the surgeon is a man). Ask variations of this and it'll keep going back to the standard form.
Another one I saw was multiplying large numbers. The starting and ending digits tended to be correct but the middle digits were wrong. Why? Because it's really not doing multiplication at all. It's looking for statistical answers. It's unlikely to have met the exact pair of very large numbers you're multiplying before.
Now pundits will argue that all of these are solvable problems and individually they are. But my suspicion is that there will be a neverending stream of such edge cases and it'll be impossible to trust an LLM's output unless you are knowledgeable enough to fact check it yourself.
Now if your example of identifying zero days, this comes up with what I can only describe as "light positives", meaning it's technically a bug but essentially impossible to exploit. IIRC this came up with the demonstration where someone pointed Fable at some BSD code. I'm not sure if there have been any true false positives and obviously false negatives are impossible to know.
I guess my point is that I think LLMs are way more limited than a lot of people think.