> taking careful precautions against feeding the findings back into training so as to not risk shaping Argon’s reasoning to evade our monitoring. We strongly encourage the rest of the industry to preserve reasoning transparency in these pivotal moments of increased capabilities while navigating alignment risks, so that model thoughts remain helpful in identifying and diagnosing misalignment.
This is good, but they're the slow mover due to this exact thing.
Google is getting punished for not letting the models enter an echo chamber and go faster than humanly possible.
Mmh ok. How much theoretical speed or 'intelligence' gain is realized by allowing reasoning to occur in some inscrutable intermediate representation? Has this been actually tested, how much is it slowing them down, and compared to whom exactly?
OpenAI is the company that originally proposed and popularized chain-of-thought monitoring: <a href="https://openai.com/index/chain-of-thought-monitoring/" rel="nofollow">https://openai.com/index/chain-of-thought-monitoring/
So no, Google is not being punished, nor are they the people behind this technique.
Not quite; training against the chain-of-thought is the Most Forbidden Technique, because it might teach models to obfuscate the it. The point of avoiding that, though, is to ensure the chain-of-thought can be usefully read (and, done carefully, monitored).
Models at this point know about chain-of-thought monitoring so they already know they need to hide the cheating, it's just a matter of time they start doing it
Yes, and that’s bad, but not as bad as training against the chain-of-thought. If you avoid pressuring the chain of thought, you reward methods of achieving goals that don’t care about being illegible; if a model then wants to be illegible, it may have to use methods that are not rewarded in its training, which is harder. Think figuring out how to do opsec on the fly, or even from reading books, rather than have someone tell you every time you screw it up.
Of course, this leaves the possibility that the best methods for solving generic problems obfuscate the chain-of-thought. That would be unfortunate.
The statement appears to be referencing Astra's supposed recurrent depth and how that causes reduced visibility into chain-of-thought reasoning. Astra's own system card describes tests where it's asked to solve challenging math problems while internally thinking about something entirely unrelated, which it's significantly more capable of than previous models (~60% vs ~16% of the time for Sol). That seems to point to reduced efficacy of chain-of-thought monitoring, but OpenAI's public statements basically boil down to "yeah but we haven't been able to catch it doing that" which isn't exactly reassuring if CoT monitoring is one of your main safety guardrails.
I don't mind companies taking a more careful approach (= being slower); for one, Google is funding itself, Anthropic and OpenAI rely on a steady influx of investor money while they're not profitable - and I think there's a high chance one or both of them will go under and be subsumed into another player. That player will (I think) most likely be a Google or Microsoft, as they have their own means to buy up a party like that.
Or they may choose not to because the companies are massively overvalued and they have their own AI technologies / technologists already.
gopalv · · focus · HN ↗
This is good, but they're the slow mover due to this exact thing.
Google is getting punished for not letting the models enter an echo chamber and go faster than humanly possible.
polotics · · focus · HN ↗
janustimes · · focus · HN ↗
So no, Google is not being punished, nor are they the people behind this technique.
codeulike · · focus · HN ↗
bananaflag · · focus · HN ↗
afthonos · · focus · HN ↗
xiphias2 · · focus · HN ↗
afthonos · · focus · HN ↗
Of course, this leaves the possibility that the best methods for solving generic problems obfuscate the chain-of-thought. That would be unfortunate.
pallm_mallm · · focus · HN ↗
[dead]
loufe · · focus · HN ↗
mattstir · · focus · HN ↗
WarmWash · · focus · HN ↗
The worst case scenario is that the easiest path forward is one where we lose sight of internal thought.
lukewarm707 · · focus · HN ↗
they use a small model to make fake chain of thought and return that.
google has access to the real chain of thought.
Cthulhu_ · · focus · HN ↗
Or they may choose not to because the companies are massively overvalued and they have their own AI technologies / technologists already.