‹ BackHN Continuity

Thread

Gemini 4 Argon

1699 points · 1187 comments · bradleyg223

  1. taylorfinley · · focus · HN ↗
    Ten days ago I had an experience with Gemini 3.8 flash that made me wonder if I was being routed to a different model under test. I was trying to use rocm with llama.cpp on my 128gb Strix Halo but could only get it to run Vulkan. I pasted the error message into agy and it proceeded to attach GDB to my GPU driver, reverse-engineer the kernel queue ioctl interface, and author an LD_PRELOAD C shim to get ROCm llama.cpp working on my Strix Halo. My jaw was hanging open the whole time.

    Edit to add the fix: <a href="https:&#x2F;&#x2F;gist.github.com&#x2F;birep&#x2F;6f2c8d490c7a29820997d57bd654c351" rel="nofollow">https:&#x2F;&#x2F;gist.github.com&#x2F;birep&#x2F;6f2c8d490c7a29820997d57bd654c3...

    1. cyanydeez · · focus · HN ↗
      if you haven&#x27;t tried Qwen3.8-Flash-Next with halogen, you&#x27;re missing out: <a href="https:&#x2F;&#x2F;github.com&#x2F;peonist-ai&#x2F;halogen-flash-server#the-host-settings-these-numbers-were-measured-on" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;peonist-ai&#x2F;halogen-flash-server#the-host-...
      1. solaire_oa · · focus · HN ↗
        I coincidentally just installed this (like 30 minutes ago), and gud dayum, it&#x27;s pretty awesome.

        I say this is awesome, even as I glossed over the README and vomited in my mouth. The halogen repo looks like the same utter AI bullshit littering GitHub. But this one delivers, in spite of it&#x27;s slop-riddled hallmarks.

        In any case, yeah, ~55 tok&#x2F;s on a high quality model (and massive RAM savings I think?), seems dope.

        1. cyanydeez · · focus · HN ↗
          Yeah, unfortunately, it does deliver for its platform.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.