Dont be fooled. Reasoning does not happen in the prediction of the next token, it happens virtually in the text that is created.
The next token prediction is just "the hardware" following the underlying rules. Like the basic set of rules.. in a sense similar to how the "game of life" does not really contain gliders. Gliders are just a self stabilised system that arrises from the simple rules.
When i look at the things said by LLMs i clearly see stuff that i would call reasoning.
Its reasoning does however have certain "bugs" that a humans reasoning would never have.
My guess is that much of those bugs appear cause the reasoning that a LLM does is not itself built on a self stabilised system. In humans you get coupling between levels of self stability which acts as a constraint, stabilising the system even more. The next level being predictive coding modelling the world. There is no next level in a LLM; they are trained as refiners in teacher forcing mode, a paradigm where self stabilisation is not a driving factor.
reliablereason · · focus · HN ↗
The next token prediction is just "the hardware" following the underlying rules. Like the basic set of rules.. in a sense similar to how the "game of life" does not really contain gliders. Gliders are just a self stabilised system that arrises from the simple rules.
lstodd · · focus · HN ↗
My position is that it is not possible.
qarl · · focus · HN ↗
Seems so to me.
So explain then what you mean by "not possible".
lstodd · · focus · HN ↗
qarl · · focus · HN ↗
So what about reasoning is not computable in your mind?
reliablereason · · focus · HN ↗
Its reasoning does however have certain "bugs" that a humans reasoning would never have.
My guess is that much of those bugs appear cause the reasoning that a LLM does is not itself built on a self stabilised system. In humans you get coupling between levels of self stability which acts as a constraint, stabilising the system even more. The next level being predictive coding modelling the world. There is no next level in a LLM; they are trained as refiners in teacher forcing mode, a paradigm where self stabilisation is not a driving factor.