Personification of AI is what’s going to get us in the end.
I think we need to draw a hard line in the sand over this. An AI didn’t hack into a company, the engineer set an automated tool to. An AI didn’t make an egregious security mistake, the engineer did.
We can’t blame the chisel for messing up our sculptures, when where just throwing the hammer!
You know, I actually think it's the refusal to consider personification that's going to get us.
Not because I think LLMs are human beings exactly, but because some people immediately reject any mechanism that just happens to look remotely human, even when there's empirical evidence for it.
So, a couple of months ago Anthropic's interpretability team found emotion-like representations that causally drive behavior. On impossible coding tasks, a "desperate" vector climbs with each failure, and steering it up takes reward hacking from ~5% to ~70%: <a href="https://arxiv.org/html/2604.07729v1" rel="nofollow">https://arxiv.org/html/2604.07729v1
A lot of people chalked it up to Anthropic's weirdness at the time, but meanwhile it looks pretty coughload bearingcough here.
You really don't need to believe that LLMs Truly Feel Emotions(tm) as blessed by an invisible pink unicorn. It's just: Vector exists; Vector changes over time; vector controls output; maybe make sure vector doesn't point wrong way.
And sure, blame the engineers for not doing that right. But then let 'em actually deal with the root cause?
JamesStuff · · focus · HN ↗
I think we need to draw a hard line in the sand over this. An AI didn’t hack into a company, the engineer set an automated tool to. An AI didn’t make an egregious security mistake, the engineer did.
We can’t blame the chisel for messing up our sculptures, when where just throwing the hammer!
Kim_Bruning · · focus · HN ↗
Not because I think LLMs are human beings exactly, but because some people immediately reject any mechanism that just happens to look remotely human, even when there's empirical evidence for it.
So, a couple of months ago Anthropic's interpretability team found emotion-like representations that causally drive behavior. On impossible coding tasks, a "desperate" vector climbs with each failure, and steering it up takes reward hacking from ~5% to ~70%: <a href="https://arxiv.org/html/2604.07729v1" rel="nofollow">https://arxiv.org/html/2604.07729v1
A lot of people chalked it up to Anthropic's weirdness at the time, but meanwhile it looks pretty coughload bearingcough here.
You really don't need to believe that LLMs Truly Feel Emotions(tm) as blessed by an invisible pink unicorn. It's just: Vector exists; Vector changes over time; vector controls output; maybe make sure vector doesn't point wrong way.
And sure, blame the engineers for not doing that right. But then let 'em actually deal with the root cause?