‹ BackHN Continuity

Thread

Thinking fast and slow in AI: The role of metacognition (2021)

177 points · 84 comments · teleforce

  1. creativeSlumber · · focus · HN ↗
    How relevant is this fast/slow thinking thing with regards to current frontier models?

    I know a large organization who's built their AI framework completely around this concept, and I feel that it's not really meaningful concept with the capabilities of current models.

    1. Retric · · focus · HN ↗
      You can ask a model for output directly and stop, or you can recursively ask it to keep refining the output.

      That seems to fit the fast vs slow model of human thought reasonably well.

      1. usernametaken29 · · focus · HN ↗
        > You can ask a model for output directly and stop

        That’s still several orders of magnitudes too slow to fit fast vs slow. Think of 30ms vs 3-4 seconds to get an idea of what we’re talking about here

        1. Retric · · focus · HN ↗
          That’s a function of the amount of processing power involved not the underlying architecture of decision making.
          1. usernametaken29 · · focus · HN ↗
            In terms of making an LLM faster but not in terms of meta-cognition. System 1 thinking as defined by Kahneman doesn’t have 100000x more compute than System 2, it is actually the opposite. That completely contradicts your claim
            1. Retric · · focus · HN ↗
              The ratio between a single pass and multiple passes is unchanged when you throw more processing power at both.

              In a human 30ms vs 3-4 seconds is a 1:100 ratio. Single vs multiple passes with an LLM varies but a 1:100 ratio isn’t unrealistic. So with enough compute and the right workload single vs multiple pass LLM could sit in that exact same 30ms vs 3-4 second timeframe.

              1. hoppp · · focus · HN ↗
                Multiple passes doesn't make it system 2. The defining characteristic of system 2 is consciousness which is expensive and slows down the system.

                So to talk about system 2 in AI we need to talk about consciousness. As long as AI is not officially conscious there is no System 2 thinking implemented

                1. Retric · · focus · HN ↗
                  Consciousness is ill defined hogwash when used in such descriptions.

                  The process of internal refinement without external action however fits.

                  1. hoppp · · focus · HN ↗
                    Using system 1 and 2 for AI is ill defined hogwash indeed.

                    The systems apply for humans and describe conscious and subconscious processing. Without that the entire reference to thinking fast and slow is bullshit.

                    1. Retric · · focus · HN ↗
                      System 1 and 2 is a description of physical process in people.

                      Using conscious vs subconscious processing is not however a meaningful definition. Subconscious processing isn’t necessarily fast. Visual processing and sensory integration can be quite slow without any conscious input.

                      1. hoppp · · focus · HN ↗
                        Not all subconscious processing is fast, for sure.

                        But for example the amygdala is super fast at tagging dangerious situations with emotions.

                        And humans socialize subconsciously, so all the small talk conversations are all subconsciously generated.

                        The definition from the book was:

                        System 1 is subconscious, fast, automatic and relies on heuristics

                        System 2 is conscious, slow, deliberate and effortful

                        1. Retric · · focus · HN ↗
                          The amygdala isn’t particularly fast here. It’s got significant connections to higher order centers in the brain and it’s own memory because classifying dangers let alone more complex emotional states can be quite complicated.

                          If he’s defining System 3+ then that kind of specificity could still work, but from what I understand he doesn’t thus making subconscious a poor fit for the systems described.

                          1. hoppp · · focus · HN ↗
                            Reaction to danger must be fast, otherwise the animal dies.

                            Saying that danger classification is slow is very clearly false because animals must react fast to danger or die.

                            You bringing up something bogus like system 3 makes me think this conversation is pointless and Im just chatting with a system 1 that does not have good shortcuts to have a good conversation about this so it's making new up contexts.

                            1. Retric · · focus · HN ↗
                              > something bogus like system 3

                              These systems exist, if the author you’re referring to is unaware or unwilling to discuss major issues with their classification then that’s a hit to their credibility.

                              Danger isn’t some universal thing, people are adapted to notice fact moving objects heading towards them quickly very fast. A very dangerous snake isn’t in that category of danger and can take a significant chunk of time to notice after you’ve stopped looking in that direction. Noticing social danger is extremely critical and glacial by any of the metrics previously discussed.

                              Emotions cover a huge range of complex interactions that depend on many different parts of the brain but hormones are inherently slow to diffuse through the brain. Pair bonding between a mother and child needs to happen, it doesn’t need to happen in 50ms.

                              1. hoppp · · focus · HN ↗
                                Yeah, pair bonding doesn't happen instantly but thats also not about processing and reacting to signals, it also involves the endocrine system as you say, so it's not an information processing system in the same sense.

                                If a tiger attacks you, you better react fast, but one strictly human example is noticing and reacting to angry faces.

                                An angry person comes at you with a bloody hammer you react before you understand what's even happening

                                1. Retric · · focus · HN ↗
                                  > thats also not about processing and reacting to signals > not an information processing system in the same sense

                                  I can only assume you have some sort of philosophical objection here, but this is information processing.

                                  Neurons are reacting to a range of chemical signals not just a single neurotransmitter because they don’t map to what CS style neural networks. Understanding the brain requires understanding the nuances not simplified models.

                                  1. hoppp · · focus · HN ↗
                                    I think it modulates information processing, the endocrine system does deliver information but it's internally recieved.

                                    The key word is modulation.

                                    Thats why I would not add the endocrine system as a system 3 or whatever.

                                    DNA copying is also an information processing system but it's not considered in cognitive psychology but transcription factors will greatly effect behavior also.

                                  2. hoppp · · focus · HN ↗
                                    But in the case of rats I would agree, their brains have more advanced olfactory systems and they communicate directly via pheromones to each other using the endocrine system.

                                    But it humans that processing is more nuanced and the olfactory system is less advanced.

            2. pixl97 · · focus · HN ↗
              >System 1 thinking as defined by Kahneman doesn’t have 100000x more compute than System 2, it is actually the opposite.

              When are you measuring?

              Systems 1 thinking is closer to precomputed tables in some ways. That is by evolution or massive amounts of training your neural network has a narrow fast path it can execute with as little compute at execution as needed.

              1. Retric · · focus · HN ↗
                Slower paths means loops here for humans where the output of a neuron gets feed back into itself. The fastest path = a feed forward neural network without loops.

                LLM’s operate strictly feed forward neural networks.

                1. usernametaken29 · · focus · HN ↗
                  Your assumption is wrong. Unlike artificial neurons, real neurons process massively and concurrently, let’s say 10-20 thousand inputs at once in less than 1 millisecond. In Kahnemans theory system 2 always involves the anterior cingulate cortex and arises when there are multiple conflicting streams of information. That’s the defining characteristic. Both system 1 and system two process in the same way, including looping back to previous areas. So really, no, stopping an LLM early vs letting it run has nothing to do with Kahnemans theory. In general LLMs are so far removed from what we consider natural cognition that it’s very hard to apply neuroscience concepts to LLMs simply because they share no real world similarity. At best, you’re simulating something in a very inefficient way.
                  1. Retric · · focus · HN ↗
                    We’re talking about biological processes which have physical limitations on how frequently a neuron can fire as well as a signaling system that based on how rapidly neurons fire. Loops anywhere in the system enforce a minimum 5ms latency and more realistically several times that. A typical neuron is firing closer to 0.1 to 2 times a second. Add to that the physical transmission speed and rapid response doesn’t allow for loops.

                    There’s a reason reflexes for the feet are handled by the spine that’s got nothing to do with total processing power.

      2. schainks · · focus · HN ↗
        Not quite. The better analogy for you is the "thinking" setting on your model.
    2. olgava · · focus · HN ↗

      [dead]

    3. anotha_one · · focus · HN ↗

      [dead]

    4. Shorel · · focus · HN ↗
      That's because an LLM thinks in terms of language, while we think in a different way, then convert the ideas to language. It can be said that language is a tool for the serialization (writing) and deserialization (reading) of human ideas. It is also an incredible useful and powerful tool by itself. This last sentence has been proved true by LLMs themselves. However, since it is working on the serialized version of ideas, I agree with you in that's not the optimal way to think and something not serialized (maybe world models) can be invented that's better for thinking. All this in no way diminishes the usefulness of language and of automated language generation.
      1. virgilp · · focus · HN ↗
        > That's because an LLM thinks in terms of language, while we think in a different way, then convert the ideas to language

        Do we? I just learned from a speaker[1] that we literally need words to recognize emotions. People who have a poor vocabulary have lower emotional intelligence because without being able to attach a word to an emotion, the brain is unable to recognize & process it.

        [1] Dude seemed to be knowledgeable about the subject. He's a specialized trainer, should be educated in this exact field. So hopefully I'm not lying to anyone here :)

        1. globalise83 · · focus · HN ↗
          Sounds uncannily similar to the pseudofacts you hear a lot in Neurolinguistic Programming training courses for sales reps. Would ask for a scientific publication reference on that one.
          1. Shorel · · focus · HN ↗
            It seems to me that idea is rooted in social consensus.

            Someone expresses an emotion but doesn't know how to react to it, their inner group all have an opinion about it, and the consensus is selected as the "appropiate" reaction to it. The individuals who react this way will claim this consensus is the same as emotional intelligence.

            Just as there are also people who react in one way, and completely disregard any external opinion about it. They simply have firm opinions and don't need the consensus.

            I will not comment on who can belong to each group, that's an exercise for the reader.

        2. Shorel · · focus · HN ↗
          I can't agree about the "unable to recognize and process it", simply because that idea is totally contrary to my own experience. I have in fact many memories which have emotions in them, without words or other external elements. However, seeing that language serialization seems to enable a vastly extended memory (entire sagas remembered as songs), it is understandable that something is gained by serialization of emotional experiences, just as something more immediate is lost.
    5. ghm2199 · · focus · HN ↗
      Structurally speaking we learn nothing like AI, we don't use vast amounts of information to pick up completely new skills. We also make decisions by using prior knowledge and emotions.The latter part is important, Thinking fast and slow cannot operate in a world of AIs as they stand today unless we are willing to grant them rights — because you have to teach them to make decisions based on all kinds of emotions — which is tricky at best.
      1. TeMPOraL · · focus · HN ↗
        > we don't use vast amounts of information to pick up completely new skills.

        Except, we do.

    6. emp17344 · · focus · HN ↗
      Seems like a terrible idea in the first place to build an entire organization around a single pop-sci book, but that’s just me.
      1. TeMPOraL · · focus · HN ↗
        It's indeed a terrible idea - as in, it's great. You get benefits of cross-marketing: you ride on a popularity of a well-known book, and as you also drive more sales of it, even if you don't have a deal and don't benefit from that directly, you strengthen the loop and solidify your brand.

        This choice doesn't really constrain what the organization can do, either. Pop-sci books have plenty of wiggle room in interpretation, and afford a lot of "you're holding it wrong" dismissals of criticism, that with a bit of clever copywriting, the organization can do absolutely anything and still claim it's embodying the framework/theory of the book.

      2. glial · · focus · HN ↗
        System 1/2 is the pop version, but the fast/slow distinction is prevalent in both RL and computational cognitive science, sometimes going under different names: procedural/deliberative, model-free/model-based, automatic/controlled, associative/rule-based, autonomous/algorithmic, etc.
    7. jhrmnn · · focus · HN ↗
      No clue what’s the consensus on this but my internal mental model is absolutely that LLM AI is pure fast mode, no slow mode. The “reasoning” loops are an attempt to mimic the slow mode but ultimately it doesn’t really work. I’m curious about the recent maths advances though, they seem to possibly challenge this.
      1. hoppp · · focus · HN ↗
        Its just hype talk. Slow mode is conscious in humans, fast is subconscious, so we need to discuss AI consciousness to talk about system 2 thinking.

        So the entire debate is fubar.

    8. hoppp · · focus · HN ↗
      Not relevant at all. Its a way to hype.
    9. glial · · focus · HN ↗
      'fast' means executing a policy, that is, a state-action mapping. A trained RL model does this.

      'slow' means making one or several action-dependent forecasts, evaluating the expected value of the outcomes, and making a decision based on that.

      Neither map exactly to the situation with LLMs, but very roughly, the first is analogous to trained classifiers and the second to reasoning models.

      The analogy breaks down, since each instance of token being produced is an example of a policy execution (system 1), and reasoning is just stringing lots of these together. But there are those who argued, before LLMs, that system 2 is just "policy composition" anyway...

      1. creativeSlumber · · focus · HN ↗
        Wouldn't you need a classifier to even decide if it is system 1 or 2? How capable does this classier need to be?
        1. glial · · focus · HN ↗
          I think of System 1 as a hash map. If you have a map, and see a new state/key whose action/value is not defined in the map, you have to go with System 2.
    10. theptip · · focus · HN ↗
      Very relevant. Modern models use CoT to do “slow thinking” and this enables them to achieve much greater performance. You can also turn off thinking and answer directly which is quite similar to “fast thinking”, good at approximate maths, not capable of algorithms, etc.

      Of course the shapes of what an AI can do in fast vs slow are quite different.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.