Ten days ago I had an experience with Gemini 3.8 flash that made me wonder if I was being routed to a different model under test. I was trying to use rocm with llama.cpp on my 128gb Strix Halo but could only get it to run Vulkan. I pasted the error message into agy and it proceeded to attach GDB to my GPU driver, reverse-engineer the kernel queue ioctl interface, and author an LD_PRELOAD C shim to get ROCm llama.cpp working on my Strix Halo. My jaw was hanging open the whole time.
Edit to add the fix: <a href="https://gist.github.com/birep/6f2c8d490c7a29820997d57bd654c351" rel="nofollow">https://gist.github.com/birep/6f2c8d490c7a29820997d57bd654c3...
My experience with Gemini 3.8 Flash has been awful; it gives me the most hallucinations out of the major models. I'm not using it for coding, but general research on different topics.
While that's annoying, the other frontier models easily overcome this with appropriate tool usage. I do a lot of research with frontier models and they're very good about identifying where their parametric knowledge is insufficient and searching for the correct knowledge on the internet. 3.8 Flash is HORRIFIC. The majority of the time it doesn't use any tools and infers things from its parametric knowledge. Things which should clearly have implied tool calls. Historical statistics, legal precedent, economic data, etc. I think it's incredibly clear that it has been tuned for speed and not accuracy.
Of course, it's called "flash," and that implies its purpose. I have little use for speed and a LOT of use for accuracy, so I'm hopeful 4.0 is much better. I saw a benchmark earlier today showing that it is much less prone to hallucinations. Let's see.
taylorfinley · · focus · HN ↗
Edit to add the fix: <a href="https://gist.github.com/birep/6f2c8d490c7a29820997d57bd654c351" rel="nofollow">https://gist.github.com/birep/6f2c8d490c7a29820997d57bd654c3...
gottorf · · focus · HN ↗
WarmWash · · focus · HN ↗
I'm assuming that Argon has at least a June 2026 date, but man, the 3 series models were a mess with newer information.
Gareth321 · · focus · HN ↗
Of course, it's called "flash," and that implies its purpose. I have little use for speed and a LOT of use for accuracy, so I'm hopeful 4.0 is much better. I saw a benchmark earlier today showing that it is much less prone to hallucinations. Let's see.