I don't know what you would call it, but reasoning/thinking/whatever it is, is a way for LLMs to tighten their sampling space, while allowing for wide sampling to still happen. This is akin to people brainstorming ideas.
I know that sounds confusing, let me break down how I think about this.
1. LLMs don't pick the token that ends up being used. This is by design, if the LLM gives a wide choice, it can better adapt to real world scenarios. i.e. generalize.
2. Without reasoning, this means that the LLM either locks in on whatever the sampler picked. Or decides mid-sentence/response to correct itself. This is what used to happen before reasoning, still happens if you turn reasoning off.
3. With reasoning, the LLM can make as many mistakes as it wants and explore its sampling space. Then use its vast pattern matching capabilities to decide which parts of the reasoning make sense and which were idiot ideas.
4. Enabling reasoning makes it so LLMs are much more confident on the final response, and the logits should theoretically all be near 99% on a single token for every token, i.e. much closer to greedy decoding. It analyzed all the possible options and figured out the best outcome, so a stray sample doesn't cause the answer to go awry.
This is why reasoning traces are filled with "but wait". I don't know if those were added in organically or artificially in the RL training, but regardless they're a good way to let the LLM keep generating other options and explore it's sampling space to the fullest.
Note: I haven't tested any of this and it's just my theory, but I'm sure if you really wanna know you can have claude run some smoke tests :)
nodja · · focus · HN ↗
I know that sounds confusing, let me break down how I think about this.
1. LLMs don't pick the token that ends up being used. This is by design, if the LLM gives a wide choice, it can better adapt to real world scenarios. i.e. generalize.
2. Without reasoning, this means that the LLM either locks in on whatever the sampler picked. Or decides mid-sentence/response to correct itself. This is what used to happen before reasoning, still happens if you turn reasoning off.
3. With reasoning, the LLM can make as many mistakes as it wants and explore its sampling space. Then use its vast pattern matching capabilities to decide which parts of the reasoning make sense and which were idiot ideas.
4. Enabling reasoning makes it so LLMs are much more confident on the final response, and the logits should theoretically all be near 99% on a single token for every token, i.e. much closer to greedy decoding. It analyzed all the possible options and figured out the best outcome, so a stray sample doesn't cause the answer to go awry.
This is why reasoning traces are filled with "but wait". I don't know if those were added in organically or artificially in the RL training, but regardless they're a good way to let the LLM keep generating other options and explore it's sampling space to the fullest.
Note: I haven't tested any of this and it's just my theory, but I'm sure if you really wanna know you can have claude run some smoke tests :)