‹ BackHN Continuity

Thread

Gemini 3.8 Live and 3.8 Live Extended Thinking

493 points · 328 comments · leumon

  1. galkk · · focus · HN ↗
    I don’t understand good experiences people are having with Gemini. It’s the only model that sometimes loses/forgets context in literally next message. Plus feeding unasked product links to responses.
    1. Gareth321 · · focus · HN ↗
      I strongly agree. I suspect it's people who have not yet used the paid models from OpenAI and Anthropic. Gemini is comparable to free models from other providers, but not in the same universe as paid models.

      This is frustrating because when I discuss AI with laypeople they think it's still incapable of counting the number of Rs in "strawberry." They believe it to be essentially useless and incapable of basic tasks. Which, to be fair, is the case with the free models.

      1. TacticalCoder · · focus · HN ↗
        > I suspect it's people who have not yet used the paid models from OpenAI and Anthropic. Gemini is comparable to free models from other providers, but not in the same universe as paid models.

        I totally disagree. I pay for both ChatGPT and Anthropic (haven't tried the chinese models yet) and yet Gemini is my go-to model (and I pay for it too through Google Workspace subscriptions for several domain names tied to Google/GMail) for anything that is not coding.

        I find Gemini better/quicker/more polished for basically every single subject out there that is not "write me lines of code".

        1. Gareth321 · · focus · HN ↗
          To be fair, it has been about six months since I tried a paid Google model. I will give it another go to compare. Hallucinations were the main issue back then but perhaps it has come a long way.
          1. mathgeek · · focus · HN ↗
            Much respect for being able to admit that you haven’t used a model in a while after saying folks probably haven’t used the models you use.
            1. Gareth321 · · focus · HN ↗
              Okay I just tested 3.8 Flash (High) on some real world problems I have used Opus (High) and Sol (high) to solve. The outcome here is terrible.

              One of the problems is local zoning laws regarding an expansion of my house. Comparison of annex vs extension, boundaries, precedent, costs, etc.

              3.8 Flash didn't check most of the required zoning laws. It relied on parametric knowledge, which is outdated and inaccurate. It checked zero precedents. It did made a very cursory check of the boundary area, but didn't validate it, so it missed a lot of important nuance and exceptions to the boundary. Its cost estimates were wildly inaccurate. Ostensibly because it was inferring an average based on historical pricing data rather than gathering current info.

              I could go on, but if I had to judge this attempt I would give it a 3/10. It's very fast, but wildly inaccurate. It's clear that the model is designed for speed over accuracy.

              But don&#x27;t take my word for it. [Most benchmarks show it to be significantly below frontier models like Astra.](<a href="https:&#x2F;&#x2F;llm-stats.com&#x2F;models&#x2F;compare&#x2F;gemini-3.8-flash-vs-gpt-6-astra">https:&#x2F;&#x2F;llm-stats.com&#x2F;models&#x2F;compare&#x2F;gemini-3.8-flash-vs-gpt...)

              This has been a useful exercise. It&#x27;s important to understand the developments taking place. I am disappointed to see that Google has made very little progress in six months relative to the frontier labs.

              1. staticman2 · · focus · HN ↗
                Gemini&#x27;s search harness in the Google app is (ironically) bad so it makes the model look bad.

                If you really want to compare apples to apples you need to test Gemini models against other models using the same third party search harness.

                Otherwise you are largely measuring how much computation the model provider is allocating to a search harness.

                1. Gareth321 · · focus · HN ↗
                  I used [AI Studio.](<a href="https:&#x2F;&#x2F;aistudio.google.com&#x2F;" rel="nofollow">https:&#x2F;&#x2F;aistudio.google.com&#x2F;) It&#x27;s possible the same issues are present there, but AI Studio is intended for serious work. I don&#x27;t see why they would intentionally hobble the model&#x27;s capabilities in AI Studio.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.