‹ BackHN Continuity

Thread

Gemini 4 Argon

1699 points · 1187 comments · bradleyg223

  1. taylorfinley · · focus · HN ↗
    Ten days ago I had an experience with Gemini 3.8 flash that made me wonder if I was being routed to a different model under test. I was trying to use rocm with llama.cpp on my 128gb Strix Halo but could only get it to run Vulkan. I pasted the error message into agy and it proceeded to attach GDB to my GPU driver, reverse-engineer the kernel queue ioctl interface, and author an LD_PRELOAD C shim to get ROCm llama.cpp working on my Strix Halo. My jaw was hanging open the whole time.

    Edit to add the fix: <a href="https:&#x2F;&#x2F;gist.github.com&#x2F;birep&#x2F;6f2c8d490c7a29820997d57bd654c351" rel="nofollow">https:&#x2F;&#x2F;gist.github.com&#x2F;birep&#x2F;6f2c8d490c7a29820997d57bd654c3...

    1. gottorf · · focus · HN ↗
      My experience with Gemini 3.8 Flash has been awful; it gives me the most hallucinations out of the major models. I&#x27;m not using it for coding, but general research on different topics.
      1. staticman2 · · focus · HN ↗
        The web version of Gemini is awful at search but I don&#x27;t think that&#x27;s the models fault.
      2. MILP · · focus · HN ↗
        I&#x27;m also not using it for coding but I&#x27;ve found Flash 3.8 to generate much better HTML output than Sonnet or Opus.
        1. robobo96 · · focus · HN ↗
          Only html or also css? Opus seems a bit more creative than most other models i&#x27;ve seen.
      3. mattjoyce · · focus · HN ↗
        Hallucination seems a very dated term.
        1. nkozyra · · focus · HN ↗
          Why? It&#x27;s the same concept and root cause it was when we first started using it.
        2. xdavidliu · · focus · HN ↗
          there are many dated expressions, including

          - AI is just a tool, like excel; it does what the human operating it tells it to

          - next token prediction cannot be true understanding

          - models can have no desires and goals, don&#x27;t anthropomorphize it

          However, &quot;hallucination&quot; is very much not one of them

        3. gottorf · · focus · HN ↗
          Hallucination is accurate for what I&#x27;m seeing -- e.g. it&#x27;s making up information about the 2nd gen Toyota Tundra that has no basis in reality. When challenged, it corrects itself.
          1. alluro2 · · focus · HN ↗
            My colleague wanted to diagnose a specific error code on his car himself, and Gemini told him that it&#x27;s simple to do with an OBD2 dongle - he asked it about the details thoroughly, to confirm, and bought the dongle.

            It didn&#x27;t work. Gemini: &quot;Oh yeah, that obviously cannot work, it&#x27;s not possible to do it through OBD2&quot; (paraphrasing)

            It was quite funny to me, but a bit less so to my colleague.

            1. Gareth321 · · focus · HN ↗
              I&#x27;ve had many similar experiences. It&#x27;s confidently incorrect to a shocking degree. Worse than ChatGPT from two years ago.
          2. mattjoyce · · focus · HN ↗
            Its always been a bad term. If we wanted an accurate term then it&#x27;s &#x27;confabulation&#x27;, but &#x27;muddled&#x27; or just &#x27;wrong&#x27; are also good.
        4. rdtsc · · focus · HN ↗
          What do we use for the “model made stuff up and claimed it as facts”? I can see hallucinations somehow anthropomorphizing LLM even more. I don’t like that we’re doing that to begin with but it’s a losing battle. I prefer “it’s broken” and “IT produced shit results” personally.
          1. mattjoyce · · focus · HN ↗
            I agree with you, except I really don&#x27;t hear that term much. People just say it wrong or confused. Good riddance, it was always a bad term.
          2. krapp · · focus · HN ↗
            Models don&#x27;t make claims. That would require a degree of interiority and intent that they don&#x27;t have.

            The bigger problem is that people expect LLMs to know what facts are. That assumption is even baked into the term &quot;hallucination.&quot; Someone who hallucinates is expected to otherwise have a grounding in objective reality, to &quot;not&quot; hallucinate, and to be able to recognize reality from fantasy. We wouldn&#x27;t allow a person who &quot;hallucinates&quot; as much as an LLM anywhere near the roles we give to LLMs. But everything an LLM does is as much a &quot;hallucination&quot; as anything else, it&#x27;s just stochastically generating grammar. Some grammar just happens to be useful because of the quality of its training data, which was probably created by humans who do possess interiority and awareness of fact.

            And it isn&#x27;t &quot;broken&quot; either. Broken assumes that the correct mode of operation is to act as a source of truth or fact generation. When LLMs &quot;apologize&quot; for bad results, for instance they aren&#x27;t actually apologizing. Try getting it to apologize for returning the correct data. It probably will. There is no cognition happening. It doesn&#x27;t know either way. It isn&#x27;t a calculator crunching numbers or a computer doing data analysis. It&#x27;s just pattern matching.

            &quot;Hallucination&quot; is no less correct than &quot;confabulation&quot; which also presupposes intent and contextual awareness. Unfortunately the way LLMs operate is so unintuitive (as opposed to the intuitive nature of the interface) that the only language we have to describe it is the language of human behavior, with all of the biases and false assumptions that brings.

        5. UpsideDownRide · · focus · HN ↗
          They still happen.
      4. WarmWash · · focus · HN ↗
        The achilles heel of 3.8 flash is it&#x27;s january 2025 knowledge cutoff date. Yes, almost 2 years ago.

        I&#x27;m assuming that Argon has at least a June 2026 date, but man, the 3 series models were a mess with newer information.

        1. blinding-streak · · focus · HN ↗
          Incorrect (to some degree)

          &gt; The knowledge cutoff date for Gemini 3.8 Flash is March 2026

          <a href="https:&#x2F;&#x2F;deepmind.google&#x2F;models&#x2F;model-cards&#x2F;gemini-3-8-flash&#x2F;" rel="nofollow">https:&#x2F;&#x2F;deepmind.google&#x2F;models&#x2F;model-cards&#x2F;gemini-3-8-flash&#x2F;

          1. WarmWash · · focus · HN ↗
            &gt;The knowledge cutoff date for Gemini 3.8 Flash is March 2026 – users can expect updated information for some domains while in others they may experience the model’s knowledge is limited to January 2025 (in line with the Gemini 3 Model Family). For more information about known limitations, see the Gemini 3.7 Flash

            The &quot;some domains&quot; are very narrow. They likely just RL&#x27;ed popular queries.

        2. Gareth321 · · focus · HN ↗
          While that&#x27;s annoying, the other frontier models easily overcome this with appropriate tool usage. I do a lot of research with frontier models and they&#x27;re very good about identifying where their parametric knowledge is insufficient and searching for the correct knowledge on the internet. 3.8 Flash is HORRIFIC. The majority of the time it doesn&#x27;t use any tools and infers things from its parametric knowledge. Things which should clearly have implied tool calls. Historical statistics, legal precedent, economic data, etc. I think it&#x27;s incredibly clear that it has been tuned for speed and not accuracy.

          Of course, it&#x27;s called &quot;flash,&quot; and that implies its purpose. I have little use for speed and a LOT of use for accuracy, so I&#x27;m hopeful 4.0 is much better. I saw a benchmark earlier today showing that it is much less prone to hallucinations. Let&#x27;s see.

      5. esafak · · focus · HN ↗
        I would expect a flash model, with its reduced size, to suffer on tail tasks. That is the trade-off you make.
      6. Gareth321 · · focus · HN ↗
        I agree. It&#x27;s much worse than the cheap Chinese models. They appear to heavily bias parametric knowledge and discourage tool use. That&#x27;s fine for things like &quot;how do I perform CPR?&quot; but worse than useless for any kind of research. [There is one benchmark showing far lower rates of hallucination, so let&#x27;s see how accurate this is.](<a href="https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;singularity&#x2F;comments&#x2F;1wuj72j&#x2F;gemini_4_argon_solved_hallucinations&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;singularity&#x2F;comments&#x2F;1wuj72j&#x2F;gemini...)
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.