US Military had close call after using AI for hallucinated intelligence report
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
US Military had close call after using AI for hallucinated intelligence report
Unofficial Hacker News client; not affiliated with Y Combinator.
drtgh · · focus · HN ↗
Poorly understood? how convenient...
LLMs are vectorial databases with losses that index statistically filled data, which uses a text interface to query such statistically filled data. The output is a string concatenation (statistically concatenated bit by bit).
When the LLMs are queried (prompted), you can get random mixed data as output, ERRORS, due to undesired indexes getting closer at one point while the string was being concatenated for the output, what affects the rest of the indexed content that will be concatenated.
It is intrinsic to this tech. The larger the context, the greater the probability of get mixed data. And if the provider lowers the precision of those indexes -in order to decrease hardware resources and energy consumption- such probability increases to the point where those errors are granted.
Even knowing that the queries can return wrong/mixed data in the responses, errors, the companies developing this, decided to introduce a new product, that connects such LLMs outputs to the command console, latter connected to internet, raw 'eval' running commands from such outputs witch obviously can contain whatever mixed random. Then we started to hear "oh, it deleted my directory", etc, and it seems the next one will be "a missile killed my wife", because it is a text concatenation engine with errors.
To name it "hallucination" is an euphemism... those are errors, and they are granted to happen at one moment. If they do not know this, then they ate too much marketing without doing their job, or it was a convenient contract for the pocket$ of someone.
theptip · · focus · HN ↗
You use a bunch of technical-sounding words here to make it sound like you understand. But to be clear, nobody understands why the evolved weights of a NN make the decisions that they do.
Almost nothing is understood about the actual representations used for nontrivial concepts, decision algorithms, etc.
If you look at the field of mechanistic interpretability, compared to “GOFAI” like learned decision trees, an LLM is completely opaque.
tantalor · · focus · HN ↗
That's like saying "my d20 decided to roll a 17"
nonethewiser · · focus · HN ↗
reichstein · · focus · HN ↗
That's what it did, with no analogy needed.
(But, to be the devil's advocate: the fake can be said about the output of anyone participating here.)
semi-extrinsic · · focus · HN ↗
But even so people don't say that we don't understand how dice work.
Saying that we don't understand how LLMs work is exactly like saying we don't understand how dice, or tires, or golf ball shots work. Or like the old myth that we don't understand how bumblebees fly.
jacquesm · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
fc417fc802 · · focus · HN ↗
skydhash · · focus · HN ↗
fc417fc802 · · focus · HN ↗
In contrast, we do not understand LLMs in the same way (nor biological brains). Claiming that anything of that nature is simply biased towards coherent output seems entirely reductive to me - the question is how such coherence arises in the first place. There is no meaning encoded or computation performed by the particular pathway a die travels through the chaotic landscape.
Sure an argument can be made that it's "just" a next token predictor thus how is it really any different from a markov model? Yet the output is not even remotely the same.
skydhash · · focus · HN ↗
From my point of view, (not a ML researcher), it’s due to the magic of numbers. The same thing happens with computer vision and neural networks. There’s a bunch of magic weights that get created which has no meaning by themselves, but computing them does help with detecting objects.
So if you take words, derives them into tokens, use the attention techniques to extract the “coherency” aspect, it’s no wonder you can replicate “coherency”. Add reinforcement learning to that to increase towards certain aspects like correct code syntax and you have heavily loaded the dice again.
We have used maths to model chemistry, biology, and physics, as well as economics and sociologic phenomena. Then we use maths (more specifically logic and set theory) to usher in the age of information and computing. Now you want us to act surprised that maths, through ML, can model language.
Maybe further down the line, we can have a simpler set of formulas for language coherency, but for now we have to make to with using the whole internet and a bazillion watts of power to guess the weights for the generic ML model.
jacquesm · · focus · HN ↗
I would not be surprised at all if we will find that AI will go the same route. The fact that we don't know how it works is where the opportunity for improvement lies.
theptip · · focus · HN ↗
The best way to model dice is the Physical Stance. You consider rules such as gravity, kinematics, etc. There is no “internal state”, “world model”, “knowledge”. If you prefer, in Friston’s terms, there is no Markov Blanket.
The best way to model a human is the Intentional Stance[1]. You mostly need things like beliefs, knowledge, biases, etc to build this model. In Friston’s terms, there is a Markov Blanket, an inside vs outside.
Without going into any irrelevant-but-interesting philosophical discussions about consciousness, I believe the intentional stance is most useful for modeling LLMs. Most of the success in predicting, debugging, optimizing these systems is in activities like understanding what they believe, what their intent was, what they observed, what they concluded from those observations. Also note that much simpler creatures benefit from the Intentional Stance; you will be more successful at modeling your dog if you think about what it “wants” rather than trying to run Physics on it.
[1]: <a href="https://en.wikipedia.org/wiki/Intentional_stance" rel="nofollow">https://en.wikipedia.org/wiki/Intentional_stance - the astute reader will note that I skipped the Design Stance. If we truly understood how NNs actually implement all their cognitive processes then we could perhaps apply this to them; if we actually crafted and designed every parameter of its mind. But we are talking about why dice are different.
tantalor · · focus · HN ↗
The latest episode of On The Media also uses this framing.
> On the Media: How Extinction Entered the AI Debate
<a href="https://www.wnycstudios.org/podcasts/otm" rel="nofollow">https://www.wnycstudios.org/podcasts/otm
Cthulhu_ · · focus · HN ↗
That is, in this case, it should not be used to influence decisions that can start a war.
krapp · · focus · HN ↗
hardbass · · focus · HN ↗
semiquaver · · focus · HN ↗
“we made this artifact and don’t know why the thing it does looks spookily like cognition”
and
“this artifact makes decisions at random”
are obviously distinct categories and pretending otherwise is silly.
watwut · · focus · HN ↗
Regardless of negative consequences it brings. They have that project of creating tech god which will save the unborn people thousands years in the future ... so people living now dont matter.
That is why.
semiquaver · · focus · HN ↗
So I don’t think “labs worked hard” is the same thing is “we know scientifically how these things work in any real level of detail”. The ability to build a thing, even if building it is hard, is not the same thing as understanding of what the thing is or how it works, not even a little bit.
theptip · · focus · HN ↗
It’s not completely random. We just don’t understand why the tricks we learned work.
(Fully agree with the second point FWIW)
s1artibartfast · · focus · HN ↗
Can you show me where a human or a dog makes decisions
rayiner · · focus · HN ↗
unsupp0rted · · focus · HN ↗
We might override them or ignore them or whatever, but they make decisions as much as anybody else or anything else does
skydhash · · focus · HN ↗