‹ BackHN Continuity

Thread

Gemini 4 Argon

1699 points · 1187 comments · bradleyg223

  1. taylorfinley · · focus · HN ↗
    Ten days ago I had an experience with Gemini 3.8 flash that made me wonder if I was being routed to a different model under test. I was trying to use rocm with llama.cpp on my 128gb Strix Halo but could only get it to run Vulkan. I pasted the error message into agy and it proceeded to attach GDB to my GPU driver, reverse-engineer the kernel queue ioctl interface, and author an LD_PRELOAD C shim to get ROCm llama.cpp working on my Strix Halo. My jaw was hanging open the whole time.

    Edit to add the fix: <a href="https:&#x2F;&#x2F;gist.github.com&#x2F;birep&#x2F;6f2c8d490c7a29820997d57bd654c351" rel="nofollow">https:&#x2F;&#x2F;gist.github.com&#x2F;birep&#x2F;6f2c8d490c7a29820997d57bd654c3...

    1. gottorf · · focus · HN ↗
      My experience with Gemini 3.8 Flash has been awful; it gives me the most hallucinations out of the major models. I&#x27;m not using it for coding, but general research on different topics.
      1. WarmWash · · focus · HN ↗
        The achilles heel of 3.8 flash is it&#x27;s january 2025 knowledge cutoff date. Yes, almost 2 years ago.

        I&#x27;m assuming that Argon has at least a June 2026 date, but man, the 3 series models were a mess with newer information.

        1. Gareth321 · · focus · HN ↗
          While that&#x27;s annoying, the other frontier models easily overcome this with appropriate tool usage. I do a lot of research with frontier models and they&#x27;re very good about identifying where their parametric knowledge is insufficient and searching for the correct knowledge on the internet. 3.8 Flash is HORRIFIC. The majority of the time it doesn&#x27;t use any tools and infers things from its parametric knowledge. Things which should clearly have implied tool calls. Historical statistics, legal precedent, economic data, etc. I think it&#x27;s incredibly clear that it has been tuned for speed and not accuracy.

          Of course, it&#x27;s called &quot;flash,&quot; and that implies its purpose. I have little use for speed and a LOT of use for accuracy, so I&#x27;m hopeful 4.0 is much better. I saw a benchmark earlier today showing that it is much less prone to hallucinations. Let&#x27;s see.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.