See the last gemini message in this thread: <a href="https://gemini.google.com/share/6d141b742a13" rel="nofollow">https://gemini.google.com/share/6d141b742a13
In my opinion still the most egregious example in history of a commercial LLM going off the rails in production. Never any technical postmortem from Google on this.
The problem is newer models are never trained from scratch, they generally just layer on more training data and use the same tools/methods for RLHF. OpenAI, Anthropic, xAI models all have a feel to them that carries over from one generation to the next.
Point is if Gemini is flawed, there's a very good chance that it's still deeply flawed, and getting smarter at the same time - that is a very bad thing.
> the problem is newer models are never trained from scratch
Base models are, and then subsequent iterations build on that base model. Closed labs do not publish which models are new base models but as a rule of thumb major release numbers are an indication (with some exceptions).
If the training data is the same, the training algorithms are the same, the RLHF is the same, and the rest of the process is the same, then it's not really from scratch, or not from scratch in a way that results in an 'out of family' model. I doubt any company would take that risk. You always build on and use what works and go from there.
This is true, but Google's models have now had a consistent history of lower psychological* coherence / consistency. See, eg <a href="https://arxiv.org/abs/2603.10011" rel="nofollow">https://arxiv.org/abs/2603.10011 (Gemma Needs Help), or search for recent "Gemini shame loops", where gemini flash models stop producing output other than SHAME SHAME SHAME...
* - as in, Skinner psychology. The set of observable behaviors. Not speaking directly here to anything like an inner life of models.
From the example alone it's hard to say that a postmortem would be useful. It could be context poisoning by an adversarial user, memory corruption etc.
It's useful from a disclosure and trust perspective.
If I remember correctly, it was in fact possible to manually inject chat context at the time, which would have made spoofing something like this completely possible.
bottlepalm · · focus · HN ↗
colordrops · · focus · HN ↗
NiloCK · · focus · HN ↗
In my opinion still the most egregious example in history of a commercial LLM going off the rails in production. Never any technical postmortem from Google on this.
rhaff · · focus · HN ↗
jackkinsella · · focus · HN ↗
bottlepalm · · focus · HN ↗
Point is if Gemini is flawed, there's a very good chance that it's still deeply flawed, and getting smarter at the same time - that is a very bad thing.
unbrice · · focus · HN ↗
Base models are, and then subsequent iterations build on that base model. Closed labs do not publish which models are new base models but as a rule of thumb major release numbers are an indication (with some exceptions).
bottlepalm · · focus · HN ↗
NiloCK · · focus · HN ↗
* - as in, Skinner psychology. The set of observable behaviors. Not speaking directly here to anything like an inner life of models.
kelvinjps10 · · focus · HN ↗
schmookeeg · · focus · HN ↗
wg0 · · focus · HN ↗
unbrice · · focus · HN ↗
NiloCK · · focus · HN ↗
If I remember correctly, it was in fact possible to manually inject chat context at the time, which would have made spoofing something like this completely possible.
But the silence on it is very frustrating.