I'm curious about some of the claims in TFA and how it lines up with the research we've seen so far. The author of TFA is clearly very experienced in the field, so is there disagreement on what this research means? Hoping experts can chime in:
> First, these models typically maintain no explicit, persistent, and inspectable epistemic state... there is no independent, explicitly represented set of beliefs.
We cannot decipher it, but Mechanistic Interpretability research does show that models are applying and manipulating abstract concepts and relationships encoded in the weights to derive their responses. We can even identify and manipulate those weights, see e.g. Golden Gate Claude. I assume these are the "independent set of beliefs" and the only reason they are not "explicitly represented" is that the representation is too complicated for us to decode.
In fact, this is also how we know they "concoct" chains of thought, because their reasoning traces do not always align with what is going on in their weights (i.e. the "concepts" that were activated during inference.) Lookup chain of thought faithfulness research, something TFA directly cites.
As another comment (<a href="https://news.ycombinator.com/item?id=49934166">https://news.ycombinator.com/item?id=49934166) points out, there is very robustly replicated evidence that humans do something similar (lookup "post-hoc rationalization.") So it is extremely fascinating that LLMs do the same thing!
Can't help but wonder if that's an entirely unrelated though similar-looking phenomenon, or an emergent property of intelligence, or something transmitted subliminally via training on the data our brains we produced...
keeda · · focus · HN ↗
> First, these models typically maintain no explicit, persistent, and inspectable epistemic state... there is no independent, explicitly represented set of beliefs.
We cannot decipher it, but Mechanistic Interpretability research does show that models are applying and manipulating abstract concepts and relationships encoded in the weights to derive their responses. We can even identify and manipulate those weights, see e.g. Golden Gate Claude. I assume these are the "independent set of beliefs" and the only reason they are not "explicitly represented" is that the representation is too complicated for us to decode.
In fact, this is also how we know they "concoct" chains of thought, because their reasoning traces do not always align with what is going on in their weights (i.e. the "concepts" that were activated during inference.) Lookup chain of thought faithfulness research, something TFA directly cites.
As another comment (<a href="https://news.ycombinator.com/item?id=49934166">https://news.ycombinator.com/item?id=49934166) points out, there is very robustly replicated evidence that humans do something similar (lookup "post-hoc rationalization.") So it is extremely fascinating that LLMs do the same thing!
Can't help but wonder if that's an entirely unrelated though similar-looking phenomenon, or an emergent property of intelligence, or something transmitted subliminally via training on the data our brains we produced...