Ten days ago I had an experience with Gemini 3.8 flash that made me wonder if I was being routed to a different model under test. I was trying to use rocm with llama.cpp on my 128gb Strix Halo but could only get it to run Vulkan. I pasted the error message into agy and it proceeded to attach GDB to my GPU driver, reverse-engineer the kernel queue ioctl interface, and author an LD_PRELOAD C shim to get ROCm llama.cpp working on my Strix Halo. My jaw was hanging open the whole time.
Edit to add the fix: <a href="https://gist.github.com/birep/6f2c8d490c7a29820997d57bd654c351" rel="nofollow">https://gist.github.com/birep/6f2c8d490c7a29820997d57bd654c3...
My experience with Gemini 3.8 Flash has been awful; it gives me the most hallucinations out of the major models. I'm not using it for coding, but general research on different topics.
What do we use for the “model made stuff up and claimed it as facts”? I can see hallucinations somehow anthropomorphizing LLM even more. I don’t like that we’re doing that to begin with but it’s a losing battle. I prefer “it’s broken” and “IT produced shit results” personally.
Models don't make claims. That would require a degree of interiority and intent that they don't have.
The bigger problem is that people expect LLMs to know what facts are. That assumption is even baked into the term "hallucination." Someone who hallucinates is expected to otherwise have a grounding in objective reality, to "not" hallucinate, and to be able to recognize reality from fantasy. We wouldn't allow a person who "hallucinates" as much as an LLM anywhere near the roles we give to LLMs. But everything an LLM does is as much a "hallucination" as anything else, it's just stochastically generating grammar. Some grammar just happens to be useful because of the quality of its training data, which was probably created by humans who do possess interiority and awareness of fact.
And it isn't "broken" either. Broken assumes that the correct mode of operation is to act as a source of truth or fact generation. When LLMs "apologize" for bad results, for instance they aren't actually apologizing. Try getting it to apologize for returning the correct data. It probably will. There is no cognition happening. It doesn't know either way. It isn't a calculator crunching numbers or a computer doing data analysis. It's just pattern matching.
"Hallucination" is no less correct than "confabulation" which also presupposes intent and contextual awareness. Unfortunately the way LLMs operate is so unintuitive (as opposed to the intuitive nature of the interface) that the only language we have to describe it is the language of human behavior, with all of the biases and false assumptions that brings.
taylorfinley · · focus · HN ↗
Edit to add the fix: <a href="https://gist.github.com/birep/6f2c8d490c7a29820997d57bd654c351" rel="nofollow">https://gist.github.com/birep/6f2c8d490c7a29820997d57bd654c3...
gottorf · · focus · HN ↗
mattjoyce · · focus · HN ↗
rdtsc · · focus · HN ↗
mattjoyce · · focus · HN ↗
krapp · · focus · HN ↗
The bigger problem is that people expect LLMs to know what facts are. That assumption is even baked into the term "hallucination." Someone who hallucinates is expected to otherwise have a grounding in objective reality, to "not" hallucinate, and to be able to recognize reality from fantasy. We wouldn't allow a person who "hallucinates" as much as an LLM anywhere near the roles we give to LLMs. But everything an LLM does is as much a "hallucination" as anything else, it's just stochastically generating grammar. Some grammar just happens to be useful because of the quality of its training data, which was probably created by humans who do possess interiority and awareness of fact.
And it isn't "broken" either. Broken assumes that the correct mode of operation is to act as a source of truth or fact generation. When LLMs "apologize" for bad results, for instance they aren't actually apologizing. Try getting it to apologize for returning the correct data. It probably will. There is no cognition happening. It doesn't know either way. It isn't a calculator crunching numbers or a computer doing data analysis. It's just pattern matching.
"Hallucination" is no less correct than "confabulation" which also presupposes intent and contextual awareness. Unfortunately the way LLMs operate is so unintuitive (as opposed to the intuitive nature of the interface) that the only language we have to describe it is the language of human behavior, with all of the biases and false assumptions that brings.