I don’t understand good experiences people are having with Gemini. It’s the only model that sometimes loses/forgets context in literally next message. Plus feeding unasked product links to responses.
I strongly agree. I suspect it's people who have not yet used the paid models from OpenAI and Anthropic. Gemini is comparable to free models from other providers, but not in the same universe as paid models.
This is frustrating because when I discuss AI with laypeople they think it's still incapable of counting the number of Rs in "strawberry." They believe it to be essentially useless and incapable of basic tasks. Which, to be fair, is the case with the free models.
> you're just not using the latest model, bro
Pro tip, ChatGPT is the normiest of all normie websites right now. You're not part of the cognoscenti just because you learned how to type prompts into one of the most popular websites in the world.
P.S. You're probably not using OpenAI models for complex or non-standard tasks. It shits the best just as often as Qwen when you need precision and detail in a non-obvious problem.
> You're probably not using OpenAI models for complex or non-standard tasks. It shits the best just as often as Qwen when you need precision and detail in a non-obvious problem.
The benchmarks clearly show otherwise. This is your cue to tell me the benchmarks are made by the Illuminati and only your superior and subjective methods of evaluation are correct.
artificialanalysis just updated their benchmark after the release of GPT-6. They removed old, saturated benchmarks and replaced them with new, until GPT-6 floated to the top with the cream. One of those new benchmarks is AutomationBench-AA, where GPT-6 had a clear lead. Today that benchmark is topped by DeepSeek v4.1 Flash.
Edit: For those who are not familiar with it, this model is quite a bit faster, and about 100x cheaper, per token, than Fable and Astra.
It's very impressive that DeepSeek 4.1 beats Astra in one benchmark, but I presume you are aware that Astra wins in almost all other benchmarks? These are some of them: <a href="https://llm-stats.com/models/compare/deepseek-v4.1-flash-vs-gpt-6-astra">https://llm-stats.com/models/compare/deepseek-v4.1-flash-vs-...
galkk · · focus · HN ↗
Gareth321 · · focus · HN ↗
This is frustrating because when I discuss AI with laypeople they think it's still incapable of counting the number of Rs in "strawberry." They believe it to be essentially useless and incapable of basic tasks. Which, to be fair, is the case with the free models.
otabdeveloper4 · · focus · HN ↗
Pro tip, ChatGPT is the normiest of all normie websites right now. You're not part of the cognoscenti just because you learned how to type prompts into one of the most popular websites in the world.
P.S. You're probably not using OpenAI models for complex or non-standard tasks. It shits the best just as often as Qwen when you need precision and detail in a non-obvious problem.
Gareth321 · · focus · HN ↗
The benchmarks clearly show otherwise. This is your cue to tell me the benchmarks are made by the Illuminati and only your superior and subjective methods of evaluation are correct.
irthomasthomas · · focus · HN ↗
Edit: For those who are not familiar with it, this model is quite a bit faster, and about 100x cheaper, per token, than Fable and Astra.
Gareth321 · · focus · HN ↗