‹ BackHN Continuity

Thread

Pacing the Frontier is not the actual goal for AI labs

83 points · 94 comments · brlewis

  1. nonethewiser · · focus · HN ↗
    We already have a window into the future.

    - Anthropic told everyone Mythos was dangerous because it's proficiency with biologics and cyber security

    - Anthropic didn't release Mythos like everything else. They released a neutered fable. They didn't get rid of Mythos

    - Anthropic opens a lab in SF

    There was always a quesiton of "will the labs stop releasing their models and start building around them instead?" Yes - they already have. Anthropic is a biologic and cyber security company, in addition to intelligience.

    Personally I wonder if they've been holding back a lot. Opus 5.5 was a good release after a little stagnation. Open AI releases good models and everyone says Anthropic sucks and -- Oh would you look at that - a better model finally and all of a sudden.

    1. bpodgursky · · focus · HN ↗
      Yes, both labs already have monstrous internal teacher models they don't sell for inference, this is generally acknowledged. They cut releases for the public just to keep revenues growing, it's not their actual frontier.
      1. sowhat1 · · focus · HN ↗
        Ok. Given this hypothesis, why is the software they release generally considered crappy by competitor standards, benchmarks, and open source standards?

        Claude Code is an awful codebase, has leaked its own source code multiple times, and scores the worst on number of tokens burned vs pass rate percentages.

        Is anyone even using their Figma competitor?

        1. monocasa · · focus · HN ↗
          Probably because bad code that you create initially without thinking that it's a core piece of your stack becomes depended on for its crappy behavior, and then you can't change much without breaking workflows.

          I seem to recall Fred Brooks talking about that experience with OS/360 JCL (maybe just straight up in The Mythical Man Month?).

          1. CoolestBeans · · focus · HN ↗
            If agents really are superpowerful at programming tasks why not just have it rewrite the tool that the majority of your customers use and have it recreate the bugs? I mean presumably its the primary force behind the current version so what's the major cost there?
            1. XenophileJKO · · focus · HN ↗
              I imagine it comes down to economics.. there isn't much upside to fixing the last 20% of issues that the dumber faster models are missing.

              The cost to serve, latency profile ,and internal demand for a maximal intelligence model would probably keep it pointed at harder and more valuable problems most of the time.

              1. _aavaa_ · · focus · HN ↗
                Like front-running Millennium prize solutions?
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.