‹ BackHN Continuity

Thread

OpenAI halts training of latest models as reports mount of AI agents going rogue

59 points · 118 comments · smb06

  1. mikert89 · · focus · HN ↗
    It seems like anthropic is far ahead of openai, and has no reports like this. We have to conclude this is a skill issue/engineering quality problem inside openai.

    just because they are a well known name, doesnt mean they havent botched hiring over the last two years or so

    1. Razengan · · focus · HN ↗
      Anthropic's models seem crippled and hamstrung to begin with
      1. mikert89 · · focus · HN ↗
        have you tried opus 5.5? anthropic is way ahead, atleast in terms of publicly available models
        1. Razengan · · focus · HN ↗
          Someone else&#x27;s experience with Opus 5.5: <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49821657">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49821657

          &gt; I had it try to prepare a code review for me. Not only did it refuse, it refused to even tell me what the prompt (written by another Claude!) was. Why?

          &gt; When I had another model read the session (all of the &quot;stupider&quot; models handled it just fine) it explained that it had the word &quot;reasoning&quot; in it

          &gt; That&#x27;s the entirety of Anthropic&#x27;s billions of dollars of research: any prompt with the word &quot;reasoning&quot; is trying to hack Claude to figure out how it reasons!

          &gt; A model like that should never have gotten out of QA, let alone been released.

          1. verdverm · · focus · HN ↗
            we have GLM flash catching Claude errors in our PR review system, costs a few pennies

            I&#x27;ve seen the same pattern regardless of open v closed, don&#x27;t have the same family that wrote the code also review the code

            diversity has this way of making things better across everything humans do

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.