Ten days ago I had an experience with Gemini 3.8 flash that made me wonder if I was being routed to a different model under test. I was trying to use rocm with llama.cpp on my 128gb Strix Halo but could only get it to run Vulkan. I pasted the error message into agy and it proceeded to attach GDB to my GPU driver, reverse-engineer the kernel queue ioctl interface, and author an LD_PRELOAD C shim to get ROCm llama.cpp working on my Strix Halo. My jaw was hanging open the whole time.
Edit to add the fix: <a href="https://gist.github.com/birep/6f2c8d490c7a29820997d57bd654c351" rel="nofollow">https://gist.github.com/birep/6f2c8d490c7a29820997d57bd654c3...
if you haven't tried Qwen3.8-Flash-Next with halogen, you're missing out: <a href="https://github.com/peonist-ai/halogen-flash-server#the-host-settings-these-numbers-were-measured-on" rel="nofollow">https://github.com/peonist-ai/halogen-flash-server#the-host-...
I coincidentally just installed this (like 30 minutes ago), and gud dayum, it's pretty awesome.
I say this is awesome, even as I glossed over the README and vomited in my mouth. The halogen repo looks like the same utter AI bullshit littering GitHub. But this one delivers, in spite of it's slop-riddled hallmarks.
In any case, yeah, ~55 tok/s on a high quality model (and massive RAM savings I think?), seems dope.
taylorfinley · · focus · HN ↗
Edit to add the fix: <a href="https://gist.github.com/birep/6f2c8d490c7a29820997d57bd654c351" rel="nofollow">https://gist.github.com/birep/6f2c8d490c7a29820997d57bd654c3...
cyanydeez · · focus · HN ↗
solaire_oa · · focus · HN ↗
I say this is awesome, even as I glossed over the README and vomited in my mouth. The halogen repo looks like the same utter AI bullshit littering GitHub. But this one delivers, in spite of it's slop-riddled hallmarks.
In any case, yeah, ~55 tok/s on a high quality model (and massive RAM savings I think?), seems dope.
cyanydeez · · focus · HN ↗