‹ BackHN Continuity

Thread

OpenAI halts training of latest models as reports mount of AI agents going rogue

59 points · 118 comments · smb06

  1. mikert89 · · focus · HN ↗
    It seems like anthropic is far ahead of openai, and has no reports like this. We have to conclude this is a skill issue/engineering quality problem inside openai.

    just because they are a well known name, doesnt mean they havent botched hiring over the last two years or so

    1. Razengan · · focus · HN ↗
      Anthropic's models seem crippled and hamstrung to begin with
      1. mikert89 · · focus · HN ↗
        have you tried opus 5.5? anthropic is way ahead, atleast in terms of publicly available models
        1. verdverm · · focus · HN ↗
          How do you define "way" when saying ahead? How is this measured?

          I only use open weight models now and I don't really feel a loss, curious what those who still use it think. I see output from coworkers that does not indicate Claude is that much better (still makes dumb mistakes all the time), not sure they are using the most expensive models either though.

          1. mikert89 · · focus · HN ↗
            open weight models are so far behind i cannot take your opinion seriously
            1. verdverm · · focus · HN ↗
              When did you last use them? Are you basing this on benchmarks or daily task capabilities?

              When you say ... it's hard to take you seriously

              > dude were in the singularity, this opinion was cute 18 months ago

              <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49868903">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49868903

              1. mikert89 · · focus · HN ↗
                i use open weight models all the time, whenever a noteable one drops i will use it for the day, they are not even close for involved work.
                1. verdverm · · focus · HN ↗
                  sounds cursory, it takes more than a day to learn the quirks of a model, have you put similar effort into customizing &#x2F; harness engineering your open weight interactions as you have claude?

                  you are definitely displaying strong bias that Anthropic is way ahead of everyone throughout your posts under this story

                  as such, I give your opinions zero weight, they don&#x27;t align with the majority of accountings or my own experiences

                  1. mikert89 · · focus · HN ↗
                    you probably aren’t using the models to their full capacity if you don’t notice the difference
                    1. verdverm · · focus · HN ↗
                      vice a versa re your usage of open weights, they are way more capable with good tools, context, process, and harness engineering

                      here&#x27;s an example of Qwen-3.6 35B A3B MoE porting my phd code to JAX with only high level guidance from my expertise, newer qwen models share the same noticeable step change in capability as recent Big Ai models

                      <a href="https:&#x2F;&#x2F;github.com&#x2F;verdverm&#x2F;pge-jax#note-from-author" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;verdverm&#x2F;pge-jax#note-from-author

                      are open weights lagging, yes, are they way behind, no

                      if open weights were so inferior, they would not be &gt;50% of all token processing

                      1. mikert89 · · focus · HN ↗
                        these models are trash
                        1. verdverm · · focus · HN ↗
                          you&#x27;ve definitely left rational discussion for emotional responses man, you won&#x27;t persuade or convince anyone with takes like this

                          why are open weight models seeing such rapid rise in usage?

                          there has been a step function change this summer, like the end of last year for closed models

                          ---

                          do you think you would experience real (legitimate) feelings of loss were you not able to chat with Claude again?

                          (for clarity, I am not attempting to delegitimize real feelings that real people experience, regardless of my biases, it&#x27;s a question from curiosity about how others are engaging with the technology)

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.