‹ BackHN Continuity

Thread

Gemini 4 Argon

1699 points · 1187 comments · bradleyg223

  1. gopalv · · focus · HN ↗
    > taking careful precautions against feeding the findings back into training so as to not risk shaping Argon’s reasoning to evade our monitoring. We strongly encourage the rest of the industry to preserve reasoning transparency in these pivotal moments of increased capabilities while navigating alignment risks, so that model thoughts remain helpful in identifying and diagnosing misalignment.

    This is good, but they're the slow mover due to this exact thing.

    Google is getting punished for not letting the models enter an echo chamber and go faster than humanly possible.

    1. janustimes · · focus · HN ↗
      OpenAI is the company that originally proposed and popularized chain-of-thought monitoring: <a href="https:&#x2F;&#x2F;openai.com&#x2F;index&#x2F;chain-of-thought-monitoring&#x2F;" rel="nofollow">https:&#x2F;&#x2F;openai.com&#x2F;index&#x2F;chain-of-thought-monitoring&#x2F;

      So no, Google is not being punished, nor are they the people behind this technique.

      1. mattstir · · focus · HN ↗
        The statement appears to be referencing Astra&#x27;s supposed recurrent depth and how that causes reduced visibility into chain-of-thought reasoning. Astra&#x27;s own system card describes tests where it&#x27;s asked to solve challenging math problems while internally thinking about something entirely unrelated, which it&#x27;s significantly more capable of than previous models (~60% vs ~16% of the time for Sol). That seems to point to reduced efficacy of chain-of-thought monitoring, but OpenAI&#x27;s public statements basically boil down to &quot;yeah but we haven&#x27;t been able to catch it doing that&quot; which isn&#x27;t exactly reassuring if CoT monitoring is one of your main safety guardrails.
        1. WarmWash · · focus · HN ↗
          Unfortunately, this reduced insight seems to also give Astra incredible &quot;intelligence-per-token&quot;.

          The worst case scenario is that the easiest path forward is one where we lose sight of internal thought.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.