Thinking fast and slow in AI: The role of metacognition (2021)
Thread
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
Thinking fast and slow in AI: The role of metacognition (2021)
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
jannyfer · · focus · HN ↗
(In case people miss that before discussion)
tomhow · · focus · HN ↗
bbor · · focus · HN ↗
I do wonder if any of these folks ever got a chance to try this at one of the big labs, tho...
BOOSTERHIDROGEN · · focus · HN ↗
crorella · · focus · HN ↗
mdk2578 · · focus · HN ↗
[dead]
alansaber · · focus · HN ↗
red75prime · · focus · HN ↗
I guess my box labelled 'subconsciousness' is trying to say that low-level mechanisms that give rise to the observed cognitive phenomena might have nothing to do with neat boxes.
bananaflag · · focus · HN ↗
<a href="https://dl.acm.org/doi/pdf/10.1145/1045339.1045340" rel="nofollow">https://dl.acm.org/doi/pdf/10.1145/1045339.1045340
arethuza · · focus · HN ↗
bananaflag · · focus · HN ↗
arethuza · · focus · HN ↗
CipherText · · focus · HN ↗
[dead]
Phemist · · focus · HN ↗
sinuhe69 · · focus · HN ↗
cjg · · focus · HN ↗
aidiscoverywire · · focus · HN ↗
[dead]
simianwords · · focus · HN ↗
bpshaver · · focus · HN ↗
lelanthran · · focus · HN ↗
Tell me you didn't read Daniel Khaneman's book without telling me you didn't read Daniel Khaneman's book.
simianwords · · focus · HN ↗
lelanthran · · focus · HN ↗
No, it hasn't. Maybe you have a different definition of System 1 and System 2. I last read the book well over a decade ago (2011, maybe? 2012?), but System 1 and System 2 are different systems. IOW, System 2 is not a more computational version of System 1.
The argument you made implies that System 2 is just a more capable System 1, which is not what the book (nor this paper, AIUI) proposes.
In computery terms, System 1 runs in O(1) time, System 2 runs in O(log n) (or maybe just O(n)) time.
This means that any System 1 will run the input once through the heuristics, using the same computational power and taking the same time whether the input is 100 tokens or 1 million tokens, for quick but perhaps wrong decision (not "answer"). We don't have LLMs that do that. We have System 2 - run in O(log n) time and produce an answer.
System 1 is completely bereft of thought.
simianwords · · focus · HN ↗
No, system 2 is the emergent capability to reason and increase the space of places to find the answer. Forget the paper's proposal, and look at the problem it is trying to solve. Ability to give quick answers, ability to give thought out answers, and the ability to know when to choose what. Adaptive reasoning does all three.
> This means that any System 1 will run the input once through the heuristics, using the same computational power and taking the same time whether the input is 100 tokens or 1 million tokens, for quick but perhaps wrong decision (not "answer").
No, I don't think we humans use o(1) to for understanding 1000 tokens or 2 tokens. I simply don't think that's the case. There's a new model called "Jev" and it is literally named System 1 (from the book) and even it is billed per input token.
hoppp · · focus · HN ↗
System 1 can do multiple passes and all that, the speed is related to how aware are you of the computation occuring and system 2 must be consciously managed.
pixl97 · · focus · HN ↗
They also share a lot of overlap in brain structures and they interplay while executing. A thought can start out Sys1 and quickly migrate to Sys2 as the pattern fails to match. Or a System 2 chain of thinking can be made from a bunch of smaller system 1 actions. Heck, in the middle of a system 2 thought you can plunge into system 1 system/actions. It's more of a who matches the pattern up with reality the fastest and acts on it.
> We don't have LLMs that do that. We have System 2
Eh. LLMs are system 1 thinkers by default. "Quick" response with no reflection is where LLMs started. It's later we added Chain of Thought and reflection layers and all kinds of other things like harnesses and agent training to make them act like system 2 thinkers. Of course we have other technologies being tested on LLMs these days like Dynamic Sparse Attention that likely match more of your thoughts on what system 1 thinking is too. Where the network doesn't have to parse the full context of the prompt and instead pattern matches with a much smaller percentage of the input giving responses back in ms versus seconds.
creativeSlumber · · focus · HN ↗
I don't think there's any problem category that is strictly a quick heuristic decision or something that you will think through thoroughly always. I think it's more about how much time you have. If you don't have time you'll make a quick heuristic based decision. If you have more time you will think through it more. Imagine you're driving and suddenly like a branch falls on the road right in front of you. And you need to avoid it. Your brain will quickly use a heuristic based approach to avoid the branch. But on the other hand, if this branch was already there on the road and you saw it from far away, you will probably take a lot more time to think through and figure out which path you need to take.
globnomulous · · focus · HN ↗
* The person who posted it likely wasn't posting it as an out-of-date paper but as an interesting idea. Your comment ignores the idea and focuses on the out-of-dateness.
* You say "this has been solved" without defining what "this" is.
* The solution to the problem -- different effort levels -- seems to indicate that you misunderstand the idea that the paper is proposing. If I understand their proposal, it's that the system itself decides how to reason based on the nature of the problem it faces, given the model's world model.
* The title is an allusion to a book by Daniel Kahneman. The brisk dismissal without acknowledging the idea or the history doesn't leave a good impression, even if I'm mistaken and you're right.
In short, Hacker News readers tend to reward depth and detail (the FAQ specifically encourages thoughtful contributions and explicitly discourages dismissal). Your comment doesn't provide them, and it appears to make a mistake that further undermines its value as a contribution to discussion.
simianwords · · focus · HN ↗
do you even know how adaptive reasoning works?
globnomulous · · focus · HN ↗
simianwords · · focus · HN ↗
<a href="https://openai.com/index/gpt-5-1/" rel="nofollow">https://openai.com/index/gpt-5-1/
It says literally the thing you wanted from system 2. Its almost exactly that.
This is what you said btw:
"it's that the system itself decides how to reason based on the nature of the problem it faces"
globnomulous · · focus · HN ↗
I don't know whether that's why anybody downvoted you, but I think that's probably part of the reason. On the other hand, the Hacker News guidelines ask us to respond to the strongest interpretation of a comment or post. In that spirit, just as you could have said something to the effect of "Yep, AI researchers learned from behavioral economics and research into human decision making and here's where OpenAI mentions it," I could have said just "the relevance may not be that it's a new idea but rather that it's a neat idea, to the person posting it at least. Maybe someone who knows AI systems more deeply wouldn't find it interesting. I do though!" That's a bit of hypocrisy on my part, I think.
Anyhow, for the record, yep, I knew of adaptive reasoning, but, no, I didn't know that it matched what this paper proposes so closely. Thanks for the correction!
zfoong · · focus · HN ↗
creativeSlumber · · focus · HN ↗
I know a large organization who's built their AI framework completely around this concept, and I feel that it's not really meaningful concept with the capabilities of current models.
Retric · · focus · HN ↗
That seems to fit the fast vs slow model of human thought reasonably well.
usernametaken29 · · focus · HN ↗
That’s still several orders of magnitudes too slow to fit fast vs slow. Think of 30ms vs 3-4 seconds to get an idea of what we’re talking about here
Retric · · focus · HN ↗
usernametaken29 · · focus · HN ↗
Retric · · focus · HN ↗
hoppp · · focus · HN ↗
So to talk about system 2 in AI we need to talk about consciousness. As long as AI is not officially conscious there is no System 2 thinking implemented
Retric · · focus · HN ↗
The process of internal refinement without external action however fits.
hoppp · · focus · HN ↗
The systems apply for humans and describe conscious and subconscious processing.
Retric · · focus · HN ↗
Using conscious vs subconscious processing is not however a meaningful definition. Subconscious processing isn’t necessarily fast.
hoppp · · focus · HN ↗
But for example the amygdala is super fast at tagging dangerious situations with emotions.
And humans socialize subconsciously, so all the small talk conversations are all subconsciously generated.
The definition from the book was:
System 1 is subconscious, fast, automatic and relies on heuristics
System 2 is conscious, slow, deliberate and effortful
Retric · · focus · HN ↗
hoppp · · focus · HN ↗
I am not buying that danger classification is slow because it's very clear that it's not. Animals must react fast or die.
Thats just natural selection.
Retric · · focus · HN ↗
Emotions are extremely complex and hormones are inherently slow to diffuse through the brain. Pair bonding between a mother and child needs to happen, it doesn’t need to happen in 50ms.
hoppp · · focus · HN ↗
If a tiger attacks you, you better react fast, but one strictly human example is noticing and reacting to angry faces.
An angry person comes at you with a bloody hammer you react before you understand what's even happening
Retric · · focus · HN ↗
I can only assume you have some sort of philosophical objection here, but this is information processing utilizing biological systems that are inherently messy here.
Neurons are reacting to a range of chemical signals not just a single neurotransmitter because they don’t map to what CS style neural networks. Understanding the brain requires understanding the nuances not simplified models.
hoppp · · focus · HN ↗
The key word is modulation.
Thats why I would not add the endocrine system as a system 3 or whatever.
DNA copying is also an information processing system but it's not considered in cognitive psychology but transcription factors will greatly effect behavior also.
hoppp · · focus · HN ↗
But it humans that processing is more nuanced and the olfactory system is less advanced.
pixl97 · · focus · HN ↗
When are you measuring?
Systems 1 thinking is closer to precomputed tables in some ways. That is by evolution or massive amounts of training your neural network has a narrow fast path it can execute with as little compute at execution as needed.
Retric · · focus · HN ↗
Slower paths means loops here.
usernametaken29 · · focus · HN ↗
Retric · · focus · HN ↗
We’re talking about biological processes which have physical limitations on how fast a neuron can fire as well as a signaling system that based on how rapidly neurons fire. Loops anywhere in the system enforce a minimum 5ms latency. And a typical neuron is firing closer to 0.1 to 2 times a second.
schainks · · focus · HN ↗
olgava · · focus · HN ↗
[dead]
anotha_one · · focus · HN ↗
[dead]
Shorel · · focus · HN ↗
virgilp · · focus · HN ↗
Do we? I just learned from a speaker[1] that we literally need words to recognize emotions. People who have a poor vocabulary have lower emotional intelligence because without being able to attach a word to an emotion, the brain is unable to recognize & process it.
[1] Dude seemed to be knowledgeable about the subject. He's a specialized trainer, should be educated in this exact field. So hopefully I'm not lying to anyone here :)
globalise83 · · focus · HN ↗
Shorel · · focus · HN ↗
Someone expresses an emotion but doesn't know how to react to it, their inner group all have an opinion about it, and the consensus is selected as the "appropiate" reaction to it. The individuals who react this way will claim this consensus is the same as emotional intelligence.
Just as there are also people who react in one way, and completely disregard any external opinion about it. They simply have firm opinions and don't need the consensus.
I will not comment on who can belong to each group, that's an exercise for the reader.
Shorel · · focus · HN ↗
ghm2199 · · focus · HN ↗
TeMPOraL · · focus · HN ↗
Except, we do.
emp17344 · · focus · HN ↗
TeMPOraL · · focus · HN ↗
This choice doesn't really constrain what the organization can do, either. Pop-sci books have plenty of wiggle room in interpretation, and afford a lot of "you're holding it wrong" dismissals of criticism, that with a bit of clever copywriting, the organization can do absolutely anything and still claim it's embodying the framework/theory of the book.
glial · · focus · HN ↗
jhrmnn · · focus · HN ↗
hoppp · · focus · HN ↗
So the entire debate is fubar.
hoppp · · focus · HN ↗
glial · · focus · HN ↗
'slow' means making one or several action-dependent forecasts, evaluating the expected value of the outcomes, and making a decision based on that.
Neither map exactly to the situation with LLMs, but very roughly, the first is analogous to trained classifiers and the second to reasoning models.
The analogy breaks down, since each instance of token being produced is an example of a policy execution (system 1), and reasoning is just stringing lots of these together. But there are those who argued, before LLMs, that system 2 is just "policy composition" anyway...
creativeSlumber · · focus · HN ↗
glial · · focus · HN ↗
theptip · · focus · HN ↗
Of course the shapes of what an AI can do in fast vs slow are quite different.
vist_orn · · focus · HN ↗
readthenotes1 · · focus · HN ↗
I shouldn't be surprised that it shows up in a screed on AI
gchamonlive · · focus · HN ↗
Obscurity4340 · · focus · HN ↗
Like for text, what would that involve? How do you compress a string or multi-line string without losing information and hopefully structure (paragraphs, would it be like replacing periods and the following space with just sticking the starting capitalized letter of the following word to the previous sentence's last letter and when it decompresses theres some kind of note that converts that back into the. First letter of the next sentence
mpalmer · · focus · HN ↗
gchamonlive · · focus · HN ↗
nkmnz · · focus · HN ↗
selicos · · focus · HN ↗
doctoboggan · · focus · HN ↗
gchamonlive · · focus · HN ↗
fizzbuzzbarbazz · · focus · HN ↗
``` th #1 - numbers are just comments ng #2, this has a space at the end information #3 compress #4 letter #5 this has a space at the start and #6 this has a space at the start and end sentence #7 this has a space at the start
How can 1ings be 4ed without losi23 or structure?
Like for text, what would 1at involve? How do you 4 a stri2or multi-line stri2wi1out losi236hopefully structure (paragraphs, would it be like replaci2periods61e followi2space wi1 just sticki21e starti2capitalized5 of 1e followi2word to 1e previous7's last56when it de4es 1eres some kind of note 1at converts 1at back into 1e. First5of the next7 ```
I'm on a phone, so I may have mistakes here, but I'm pretty sure that's shorter than your original text, in bytes, by about (9+18+20+21+12+12+16=108), minus the dictionary size of 51 -- so, 57 bytes shorter, but still containing your full text. With predistributed compressor binaries and a lot of analysis, you can even predistribute a global dictionary for common sequences, and simply specify "xyz0", x, y, and z being 24-bit numbers, or whatever bit size can index into your full reference dictionary, and 0 meaning end-of-file-dictonary. then, assuming the byte sequences in your text above are common enough to be in the 24-bit indexed dictionary, that initial dictionary could be just 22 bytes (21 and a terminator) -- so, 86 bytes shorter than the original, but still containing your original message unaltered. ..assuming i didn't make mistakes in my hand-compression.
sourdecor · · focus · HN ↗
acuozzo · · focus · HN ↗
<a href="https://en.wikipedia.org/wiki/Hutter_Prize" rel="nofollow">https://en.wikipedia.org/wiki/Hutter_Prize
BigDogAU2026 · · focus · HN ↗
[dead]
arbirk · · focus · HN ↗
ot4t · · focus · HN ↗
[dead]
hoppp · · focus · HN ↗
Meta-cognition makes sense in a dynamic and updatable and modular system, for example I can monitor thoughts coming from my amygdala with my prefrontal cortex and then adjust how I process these thoughts.
In LLMs it makes zero sense, even if you feed the output of one model into another, there is no way they can update the heuristics behind how those were computed.
tug2024 · · focus · HN ↗
[dead]
locitra · · focus · HN ↗
[dead]