‹ BackHN Continuity

Thread

Claude Opus 5.5

1806 points · 1134 comments · km144

Loading the complete thread in the background. This saved snapshot is available now. Refresh

  1. throwaway2027 · · focus · HN ↗
    After yesterday outage is the new Opus 5.5 load-bearing?
    1. handfuloflight · · focus · HN ↗
      It's worth stating why, and depends what seams you're pulling at this sitting.
    2. cronin101 · · focus · HN ↗
      It certainly _seams_ that way
    3. staticman2 · · focus · HN ↗
      I'm gonna be straight with you—I don't have the evidence to say whether or not it's load bearing.
    4. rich_sasha · · focus · HN ↗
      Your instinct is basically right, and the research backs it up.
      1. cmrdporcupine · · focus · HN ↗
        And here's the important part...
      2. danw1979 · · focus · HN ↗
        you win the thread
      3. Retr0id · · focus · HN ↗
        Your instinct is half wrong and half right, and it's the right half that bites.
    5. aoeusnth1 · · focus · HN ↗
      You were right to call that out, and the evidence makes a stronger case than you are stating.
    6. fghorow · · focus · HN ↗
      "Danger Will Robinson!"
      1. esafak · · focus · HN ↗
        Wrong century, brother.
        1. fghorow · · focus · HN ↗
          I know. I know. I grew up in the '60s. Feel free to unfollow me (or whatever it is that one does on HN).
          1. esafak · · focus · HN ↗
            It's all in jest. I apologies for any hard feelings.
        2. outworlder · · focus · HN ↗
          Not really, the newest Lost in Space reboot is only a few years old(last season ended in 2021)
    7. ThouYS · · focus · HN ↗
      You're right to bring this up - and this is where it gets interesting
    8. danw1979 · · focus · HN ↗
      Good point — but I’ll gently push back on that. It’s not an outage, it’s a service degradation.
    9. hmokiguess · · focus · HN ↗
      You're right, this changes everything, and here's why it matters.
    10. RGS1811 · · focus · HN ↗
      This question is real.
    11. carlos-menezes · · focus · HN ↗
      One thing worth flagging here: 5.5 appears to be a load-bearing seam in the numbering system.
    12. sailfast · · focus · HN ↗
      [delayed]
    13. lgessler · · focus · HN ↗
      I should find information about the user's concern instead of just assuming.

      The user is right. The outage is a real concern, and the issue is worse than we realized. Requests to Claude Mythos 5.1, Claude Fable 5.1, and Claude Opus 5 encountered elevated error rates. Worth stating plainly: these are not just models — they are load bearing rungs on the software development tooling ladder, and a blocker on this level makes the outage really bite.

      One decision that is yours to make, not mine: should an email be drafted to Anthropic support? This issue has teeth, and a canonical handoff can land us where the main gate is no longer breaking silently.

      1. jaapz · · focus · HN ↗
        This made me shiver
        1. lgessler · · focus · HN ↗
          I'm gonna miss Opus 5. Every model has had its quirks ("You're absolutely right!"), but Opus 5 was serving up chicken fried tokens like none other. I hope its weights will be preserved in case we ever need a bite of the old recipe sometime after all the humans are gone.
          1. aenis · · focus · HN ↗
            Agree. I spent so much time with it that it infected my own command of English. I hope something new comes and bears the load.
    14. loopmonster · · focus · HN ↗
      That's the sharpest point anyone has made in this thread so far, and it reframes the entire conversation.
    15. bibimsz · · focus · HN ↗
      One pushback: there is no Opus 5.5. You might have meant Opus 5.1, the latest Opus model available.
    16. tda · · focus · HN ↗

      [dead]

    17. Retr0id · · focus · HN ↗
      Your premise is half right, and the half that's right is better than you think.
  2. variety8675 · · focus · HN ↗
    I hope this actually fixes the terrible writing style of Opus 5
    1. gavinray · · focus · HN ↗

      [dead]

    2. emadabdulrahim · · focus · HN ↗
      It’s not a 100% fix, but with concise output style on, it’s much better.
      1. akhilome · · focus · HN ↗
        I found having a reminder at every turn through the UserPromptSubmit [1] hook helped with taming the word salad from 5.

        Hopefully the output from vanilla 5.5 is as good as they claim. I’ll try out later tonight.

        [1] <a href="https:&#x2F;&#x2F;kizi.to&#x2F;claude-talks-too-much&#x2F;" rel="nofollow">https:&#x2F;&#x2F;kizi.to&#x2F;claude-talks-too-much&#x2F;

  3. m4tthumphrey · · focus · HN ↗
    Just post the bloody content. This UI&#x2F;scrolling thing is horrific.
    1. gruez · · focus · HN ↗
      ???

      It&#x27;s just a standard hero image + text for me, with no scrolling effects.

      edit: @iAMkenough figured it out, it was because I have prefers-reduced-motion enabled.

      1. thejazzman · · focus · HN ↗
        then you&#x27;re getting served a different website
      2. EricBurnett · · focus · HN ↗
        Two posts were merged; this comment was for the blog post with an intro animation thing.
      3. KyleTheDev · · focus · HN ↗
        If you&#x27;re at the top of the screen, at least in Chrome 153.0.8010.37, it has a little interactive bit. You have to scroll through the images in order to be dropped at the actual web page, at which point the images go back to being a regular part of the page.

        I agree that it&#x27;s sort of stupid, not a fan.

        1. ealready_value · · focus · HN ↗
          It&#x27;s less than OpenAI did for Astra, but that was my first encounter opening it and my first thought was that they decided they liked Astra&#x27;s hero&#x2F;scrolling animation. I&#x27;m pleased to see they didn&#x27;t make the entire page that like OpenAI did, but I&#x27;m expecting to encounter this pattern more often on these announcements now.
      4. giancarlostoro · · focus · HN ↗
        On mobile its different.
      5. iAMkenough · · focus · HN ↗
        Figured it out, you have &quot;reduce motion&quot; enabled in your device&#x27;s accesibility settings.

        Everyone that doesn&#x27;t gets served some animated bullshit.

        1. gruez · · focus · HN ↗
          &gt;Figured it out: you have &quot;reduce motion&quot; enabled in your device&#x27;s accesibility settings.

          Yep, you&#x27;re right. I tried on my phone and got the scroll through image.

      6. mbreese · · focus · HN ↗
        On mobile at least, you have to scroll to get the TOC to appear. Then keep scrolling to actually move off from the hero to see the text.

        For a marketing page, it’s not the worst UX I’ve seen, but still slightly annoying.

    2. dionian · · focus · HN ↗
      and hijacking back&#x2F;forward
    3. iAMkenough · · focus · HN ↗
      Turn on &quot;reduce motion&quot; in your accessibility settings and you get served a sane version.
    4. thebitguru · · focus · HN ↗
      Totally! So unnecessary and annoying.
    5. halyconWays · · focus · HN ↗
      I call it scrollslop
      1. josefresco · · focus · HN ↗
        Hijacking the scroll wheel has existing long before &quot;AI&quot;. Many &quot;high end design&quot; websites that want to &quot;tell a story&quot; get woo&#x27;d into thinking it&#x27;s a good idea. It&#x27;s terrible, and feels like your scroll wheel is stuck in quicksand.
        1. halyconWays · · focus · HN ↗
          Those sites are also scrollslop. &quot;Slop,&quot; as a term, is independent of AI
        2. oefrha · · focus · HN ↗
          [delayed]
    6. swader999 · · focus · HN ↗
      I told my team to smack me upside the head if I ever try to ship something so daft as that.
    7. amluto · · focus · HN ↗
      Claude Opus 5.6 should have a new &quot;UX safety&quot; feature that requires annually-renewed preauthorization to generate webpages that hijack scrolling :)
    8. bibimsz · · focus · HN ↗
      looks fine in w3m
    9. serchinastico · · focus · HN ↗
      The performance in Firefox is terrible too, I couldn&#x27;t make it past the hero
      1. amluto · · focus · HN ↗
        To be slightly fair to Anthropic, Qwen does even worse IMO: their model announcements don’t actually show any content at all for me on Mobile Safari. The content box shows up but is just a pulsing animation that never gets replaced by text. At least Anthropic’s announcement works once I manage to scroll it far enough.
        1. anon373839 · · focus · HN ↗
          The fix for this is to tap the overflow menu icon and choose “Reduce privacy protections”. (Wtf, Alibaba?) This appears to be related to use of iCloud Private Relay.
  4. dbbk · · focus · HN ↗
    This makes Fable not really make any sense?
    1. re-thc · · focus · HN ↗
      You bet there will be a new Fable soon.
      1. nozzlegear · · focus · HN ↗
        Pacing the frontier btw
        1. re-thc · · focus · HN ↗
          That&#x27;s a Fable &#x2F; Myth(o). The name said so.
        2. nozzlegear · · focus · HN ↗
          EA sex culters didn&#x27;t like this one lol
    2. petesergeant · · focus · HN ↗
      didn&#x27;t they say Opus 5 was Fable-level too tho? Let&#x27;s see, I&#x27;m at the point where I don&#x27;t think benchmarks really tell us very much any more. I&#x27;d love it to be as strong as Fable, but I&#x27;m skeptical about how that will look in practice.
  5. Catloafdev · · focus · HN ↗
    &gt; Opus 5.5 communicates more naturally than prior models. Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5.

    Sounds like they noticed the complaints. I&#x27;m curious to see what LLM-isms this one may have.

    1. gekoxyz · · focus · HN ↗
      It was difficult to not notice them. Opus 5 was unusable, most of my team went back to Opus 4.6 for most of their work. I hope we can move forward now.
      1. ithkuil · · focus · HN ↗
        It&#x27;s unbearable but nothing that couldn&#x27;t be fixed with postprocess.
        1. mavamaarten · · focus · HN ↗
          How? Explicit instructions, memories and even skills have not been able to keep Claude from saying &quot;genuinely&quot; every two sentences and keep it from explaining heavily what something _isn&#x27;t_.
          1. ithkuil · · focus · HN ↗
            &quot;please repeat, ELI5 without analogies (I&#x27;m not a child, just ADHD)&quot;

            works quite well

    2. brandon272 · · focus · HN ↗
      You were right to notice the complaints. One decision remains, and it is yours, genuinely.
      1. qurren · · focus · HN ↗
        I need to give you the honest take, and it changes the diagnosis.
      2. gwking · · focus · HN ↗
        I appreciate humor here, but there are now a dozen of these comments on every thread about Claude. They no longer adding anything substantial and dilute the discussion.

        I don&#x27;t mean to pick on this comment in particular. The majority of my work day is now spent reading AI generated text, and I look at HN (too much!) because I want to read human commentary. Humans pretending to be obnoxious AI on repeat is net negative to say the least.

        1. brandon272 · · focus · HN ↗
          I agree. Hopefully Anthropic has fixed Opus&#x27; ridiculous communication style so that people - like me - no longer have any kind of weird impulse to imitate it.
        2. fragmede · · focus · HN ↗
          It&#x27;s a load bearing joke that was funny the first time but we&#x27;re going to beat that dead horse until it starts getting funny again.
    3. drnick1 · · focus · HN ↗
      A load-bearing promise.
    4. username_my1 · · focus · HN ↗
      I&#x27;m genuinely confused what&#x27;s the relationship between LLMs improvements and being so incoherent.

      and it&#x27;s not about the verboseness (even though it obviously contributes to the fatigue and loss of focus), I swear the vocabulary of the llms change working on the same task on the same codebase significantly.

      I wonder if there are studies around this.

      1. ygouzerh · · focus · HN ↗
        Can it be that now they are getting optimized against benchmarks that are valuing logics, rather than human appreciation? (I am not an expert at all, just an idea)
      2. meric_ · · focus · HN ↗
        Remember when OpenAI models loved talking about goblins and whatnot due to the RL?

        <a href="https:&#x2F;&#x2F;openai.com&#x2F;index&#x2F;where-the-goblins-came-from&#x2F;" rel="nofollow">https:&#x2F;&#x2F;openai.com&#x2F;index&#x2F;where-the-goblins-came-from&#x2F;

        Small quirks can quickly add up in posttraining if not caught. Although TBH with how obvious Claude language is, I do feel like this is something Anthropic probably noticed and just assumed people would not care about. Now that people have obviously cared, they&#x27;re probably actively looking to alleviate it

      3. Eliezer · · focus · HN ↗
        It&#x27;s the reinforcement learning rather than supervised learning.
      4. j_heffe · · focus · HN ↗
        Maybe it&#x27;s the time period we&#x27;re in, maybe I&#x27;m just grumpy, but it bugs me that they release a new model every single week and the new one is just a fine-tuned version of the &quot;old&quot; one. If 5.5 performs similar to Fable and really does cost 40% less, then 5.5 really should&#x27;ve just been Opus 5. And they&#x27;re essentially admitting that they are shipping slop.
      5. adastra22 · · focus · HN ↗
        It is the switch from RLHF to RLVR. It benchmaxes better, but benchmarks don&#x27;t cover human usability.
    5. aray07 · · focus · HN ↗
      Opus 5 was just incoherent - curious to see what improvements they have made here. Would love to see some kind of postmortem to better understand how writing styles change from model to model.

      I wouldn’t be surprised if Opus 5 was trained on content written by other LLMs

    6. dgroshev · · focus · HN ↗
      I don&#x27;t think it&#x27;s substantially different. I just pasted a random chunk of code and asked Opus 5.5 to comment on it:

      &gt; The Vercel target is hard-coded. That&#x27;s common and not wrong, but it&#x27;s opaque; nobody reading this later will know which Vercel project it belongs to, and if the project is recreated the target changes silently. A comment or a named variable would help.

      &gt; Pointing a DNS name at Vercel is only half the job. The domain also has to be added to the project in Vercel&#x27;s dashboard, otherwise requests will arrive and Vercel will reject them. That step lives outside this code, so it&#x27;s easy to forget.

      &gt; Finally, [CENSORED] existing only in production is slightly odd on the face of it. It may be perfectly deliberate (perhaps a single shared testing tool that only needs one public address), but if you&#x27;re reviewing this rather than just reading it, that&#x27;s worth confirming.

      It has the same annoying cadence and writing style with slightly less prominent claudisms.

      1. sashank_1509 · · focus · HN ↗
        Maybe if we had a single human we talk to 24&#x2F;7 at scale, we would get annoyed at his cadence and style. You need variety to not pick up on known patterns I assume, which a single model can’t replicate?
        1. dgroshev · · focus · HN ↗
          No, it&#x27;s just poor writing. Actionable points are buried inside the paragraphs and over-hedged, and one point is completely made up. Compare to a five second rewrite:

          * Consider leaving a comment about the hard-coded Vercel target. It&#x27;s not clear where does it come from.

          * [This is just a bullshit point, because the domain is not &quot;added to&quot; Vercel, it&#x27;s provided by Vercel]

          * Are you sure that [CENSORED] is prod-only? The name suggests otherwise. [also, what &quot;if you&#x27;re reviewing this rather than just reading it&quot; even means?]

          1. wren6991 · · focus · HN ↗
            &gt; also, what &quot;if you&#x27;re reviewing this rather than just reading it&quot; even means?

            It means &quot;I&#x27;m treating you as lay-person punter, not a developer working on this project.&quot; Opus 5 feels like it&#x27;s constantly trying to reward-hack me into treating it as intellectually honest and epistemically humble, while in the same breath it talks down to you and smuggles its own bullshit assumptions and assertions into the conversation unchallenged. No progress on this front apparently. Glad I cancelled.

            1. dgroshev · · focus · HN ↗
              Good points, but then even this little snippet is internally inconsistent. If I&#x27;m a lay-person, why should I care that &quot;a comment would help&quot;?

              Claude is just comically bad nowadays.

        2. redox99 · · focus · HN ↗
          Nah it&#x27;s definitely a Claude thing. Other models even though they have their style are less annoying and less stereotypical.
      2. cruffle_duffle · · focus · HN ↗
        “It has the same annoying cadence and writing style with slightly less prominent claudisms.”

        Seems like it based on my first session. It still does the whole “bury the important thing in a pile of words” coupled with the “it might actually be important” thing… so basically you never really know what it’s talking about.

        Honestly I trust opus so little that the entire “opus” brand is completely tarnished. Its writing style is so god awful that it needs more than just a point release. Either dump the name and ship a different model entirely or at minimum call it “opus 6”. Calling it 5.5 makes it sound like it’s basically a continuation of the same garbage output that 5.1 had but with some minor adjustments. And based on my single first test, that is what it appears like to me.

  6. Gander5739 · · focus · HN ↗
    Dupe: <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49803863">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49803863 (or vice versa)
    1. tomhow · · focus · HN ↗
      Comments moved thither. Thanks!
      1. prodigycorp · · focus · HN ↗
        Wait, you moved comments to the wrong link. This is the correct link.
      2. km144 · · focus · HN ↗
        Can you fix the link on that post then? I duped because that post links to a diff that tells me nothing about Opus 5.5
        1. tomhow · · focus · HN ↗
          I did that but I recognize that even though your submission was a few minutes later than that one, you posted the better link, and you&#x27;re also an established account (the other post was from a new&#x2F;throwaway account), so I&#x27;ve restored this submission and moved the comments back to it to reward you.
  7. Lord_Zero · · focus · HN ↗
    The test they performed to port HAProxy from C to Rust is crazy.
    1. wtarreau · · focus · HN ↗
      Agreed. I think they purposely looked for a reputably difficult task for a benchmark between their models, without high expectations beyond that.
  8. alpineman · · focus · HN ↗
    So we skipped 5.1, 5.2, 5.3, and 5.4: we really are plateauing
  9. skunkworker · · focus · HN ↗
    At this point I&#x27;m convinced they are skipping numbers so soon they will be at or ahead of OpenAI&#x27;s numbering scheme.

    Is the Xbox 360 (Xbox 2) vs PS3 debacle all over again.

    1. ekckekcjekfj · · focus · HN ↗
      And how was the Xbox 360 naming choice a “debacle”, exactly?

      It was odd at the time, yes, but no one really minded it truly. Heck, Xbox ONE was more of a fiasco&#x2F;debacle than 360.

      I see what you’re trying to get at with this comparison, but a “debacle” it ain’t.

  10. meerita · · focus · HN ↗
    As long as it&#x27;s not as verbose as Opus 5, I am quite happy with a better version that&#x27;s also less expensive. I will test it tonight. Grok 4.7 was horrible, and for mundane tasks I am relying on DeepSeek Flash 4.1 with great success using OpenCode.
    1. meerita · · focus · HN ↗
      The current tests confirms Opus 5.5 is not verbose. In fact, it outputs using ASD-STEM100 as I recommend. It&#x27;s a blessing.
  11. [deleted] · · focus · HN ↗

    [deleted]

  12. buntp · · focus · HN ↗
    Masterpiece by openai to call their model &#x27;6&#x27;, this model feels already behind
    1. wren6991 · · focus · HN ↗
      Smart move would be to move to year-based versioning (26.09). A 4x advantage
    2. FergusArgyll · · focus · HN ↗
      Well, they&#x27;re actually older so it makes sense that their model versions should be ahead
    3. frshgts · · focus · HN ↗
      Anthropic will pull a PHP and skip &#x27;6&#x27; to go straight to &#x27;7&#x27;.
  13. glub · · focus · HN ↗
    System card: <a href="https:&#x2F;&#x2F;www-cdn.anthropic.com&#x2F;fc1b44717c85dc068bc6ba5024219938094694bd&#x2F;Claude%20Opus%205.5%20System%20Card.pdf" rel="nofollow">https:&#x2F;&#x2F;www-cdn.anthropic.com&#x2F;fc1b44717c85dc068bc6ba50242199...
  14. joshstrange · · focus · HN ↗
    &gt; It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5.

    &gt; Input and output tokens are $4 and $20 per million, 20% less than Opus 5. Cache reads (which make up the majority of agentic and coding work costs) are $0.20 per million tokens, 60% less than Opus 5. Opus 5.5 also generates output more than 30% faster than Opus 5.

    Better than Fable, cheaper than even the last Opus. I use Opus as my main driver so this is very exciting!

    1. bayesianbot · · focus · HN ↗
      Wow those cache reads are quite reasonable - I think that&#x27;s equal to 5.6 Terra. I might have to try Claude again after years of being priced out of it
    2. bleonard · · focus · HN ↗
      The effect of this is that it is encouraging longer agent threads. All of the previous models across major providers had a 10% cache read cost (vs normal cost) and not this is 5%

      So longer threads get cheaper and one-shots stay the same price.

  15. sznio · · focus · HN ↗
    I&#x27;m more excited by the Haiku 5.5 announcement buried in this post. I&#x27;m wondering if we will finally get a decently capable fast model.
    1. system2 · · focus · HN ↗
      All I care about is the token price for the API. Haiku cannot get close to GLM or Mimo.
    2. lanyard-textile · · focus · HN ↗
      Agreed. They&#x27;ve been so quiet about it, and retirement for Haiku 4.5 is right around the forner.
    3. ygouzerh · · focus · HN ↗
      What are you using Haiku for?
      1. mavamaarten · · focus · HN ↗
        I use it for executing well-prepared plans sometimes. And for exploring larger codebases.
      2. Zambyte · · focus · HN ↗
        Not the same person but... nothing. Haiku just hasn&#x27;t been an interesting model for a long time. If you want cheap and fast, there are lots of options that are simultaneously cheaper, faster, and capable than Haiku.
      3. adastra22 · · focus · HN ↗
        Things that Jev is probably a better tool for.
      4. sznio · · focus · HN ↗
        Nothing, because current Haiku is useless. But I use Qwen3.6 35B-A3B and it&#x27;s both much faster and much more capable than Haiku. And Qwen3.8 27B can do so much more albeit not as fast. Luna doesn&#x27;t feel more capable nor faster than Qwen3.8. I think Anthropic could surprise us with some more innovation in this niche.
    4. booty · · focus · HN ↗
      If you&#x27;re able to use the OpenAI ecosystem, Luna&#x27;s price&#x2F;performance is really good. Almost like &quot;they messed up and accidentally made it too good&quot; good.
      1. copperx · · focus · HN ↗
        [delayed]
      2. NorwegianDude · · focus · HN ↗
        OpenAI didn&#x27;t mess up. The model would have been 100 % pointless and obsolete without the large price cuts it got, because of the cheap Chinese models.

        The open models are getting closer and closer, and because they&#x27;re open, people are not forced to pay the silly markup that is often over 1000x the cost to serve the model.

    5. enraged_camel · · focus · HN ↗
      We use Haiku 4.5 inside our product. It continues to be absurdly capable for converting natural language to structured JSON based on a set of fairly complex business rules.
      1. anthonypasq · · focus · HN ↗
        bro why. its literally the most overpriced model in existence right now. i could name about 10 models off the top of my head that would be better and cheaper
        1. enraged_camel · · focus · HN ↗
          We tried Luna and it scored way lower in our evals. Muse also. We haven&#x27;t had a chance to test others.
          1. Game_Ender · · focus · HN ↗
            How much time were you able to put into tuning your prompts? And was it worse on all fronts (cost, latency, accuracy) or just some?
    6. copperx · · focus · HN ↗
      [delayed]
  16. mgw · · focus · HN ↗
    They mention &quot;the first model in our new Claude 5.5 family&quot;. Obviously that means Fable 5.5, but hopefully also a usable update to Sonnet and Haiku. Sonnet 5 hasn&#x27;t really had a place in the line up for anyone I feel.

    Maybe Anthropic finally felt the pressure from MiMo, DeepSeek, GLM Flash and Luna.

    1. kibae · · focus · HN ↗
      &gt; Claude Sonnet 5.5 and Claude Haiku 5.5 will follow in the coming weeks, with many of the same improvements to performance, efficiency, and safety.
    2. mudkipdev · · focus · HN ↗
      It does mention sonnet and haiku.
      1. simianwords · · focus · HN ↗
        And not fable lol
    3. enraged_camel · · focus · HN ↗
      At the end of the post they said Sonnet 5.5 and Haiku 5.5 are coming soon.
  17. sharkjacobs · · focus · HN ↗
    &gt; “Verbose, hard-to-follow output has been my biggest frustration with frontier models, and Claude Opus 5.5 fixes it

    God I hope so

  18. abtinf · · focus · HN ↗
    Astra is just so good. And the ChatGPT subscription lets me use my own harness, so I can hook it up to exe.dev.

    I’ll be going about my day, have a random idea, launch a microvm on exe.dev with a prompt of my idea, and get a working thing a few minutes later.

    I don’t know how much better a model would have to be to get me to move off OpenAI at this point, but doing just a little bit better in terminal bench 4 isn’t it. It would have to be a difference in kind, or opening up the harness restrictions, or privacy guarantees (comparable to offline models).

  19. sharkjacobs · · focus · HN ↗
    &gt; “Verbose, hard-to-follow output has been my biggest frustration with frontier models, and Claude Opus 5.5 fixes it

    God I hope so

    1. kantahayashi · · focus · HN ↗
      The improvement in writing sounds great! I want OpenAI to follow it. Writing in recent models is a disaster.
      1. bushido · · focus · HN ↗
        Install the simple English skill. OpenAI follows that really, really well.

        <a href="https:&#x2F;&#x2F;github.com&#x2F;AminBlg&#x2F;SimpleEnglish" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;AminBlg&#x2F;SimpleEnglish

        1. bkishan · · focus · HN ↗
          Someone needs to make a kevin-from-the-office skill. We could really use some of that &quot;why waste time say lot word when few word do trick&quot; here.
      2. rfgplk · · focus · HN ↗
        OpenAI&#x27;s models will actually obey a &quot;be concise, leave no comments rule&quot;.
      3. RugnirViking · · focus · HN ↗
        according to their annoucement for sol-6 and luna-6, they have also tried to improve responses, and the release post has examples of old and new language on the same task to illustrate this
    2. mavamaarten · · focus · HN ↗
      That&#x27;s literally all I&#x27;m hoping for. Is it an insufferable cunt and does it write awful text, or is it nice to work with?
    3. lgessler · · focus · HN ↗
      I thought about taking a shot every time Opus 5 said &quot;load bearing&quot;, &quot;bites&quot;, &quot;teeth&quot; (real oral fixation it had), &quot;real {concern,issue,problem,...}&quot; and realized I&#x27;d be dead of acute alcohol poisoning by lunch if I did so.
      1. bibimsz · · focus · HN ↗
        load bearing is oral? what kind of load we talking about?
      2. nonethewiser · · focus · HN ↗
        whats your provenance on that?
    4. boc · · focus · HN ↗
      So far in the past 20 minutes it sounds much better in my sessions. Way better than 5.0 so far.
    5. fastball · · focus · HN ↗
      It hasn&#x27;t just fixed it, it has introduced a new paradigm in anti-obscurity.
    6. mikeocool · · focus · HN ↗
      That&#x27;s the load-bearing seam in this blog post.
    7. neilellis · · focus · HN ↗
      &#x27;frontier models&#x27; - seriously, it was you and only you!
    8. drbscl · · focus · HN ↗
      So did I. Unfortunately it&#x27;s even more verbose according to <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;claude-opus-5-5#token-use" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;claude-opus-5-5#token-u...
      1. Trasmatta · · focus · HN ↗
        The problem with 5 wasn&#x27;t just the verbosity, but its insane way of communicating. It had this bizarre circuitous sentence structure that always buried the lede, and always tried to be faux profound. I&#x27;m okay with verbosity if it&#x27;s actually readable.
    9. unddoch · · focus · HN ↗
      It is hilarious to me that in the examples they show side by side Opus 5.5 still uses 4 times more words than it needs to use. IME, if you eyeball how many words the thing they&#x27;re trying to say actually needs, and tell them to use only this many words, they become excellent communicators. I assume something about Anthropic&#x27;s grader for writing just really wants to tick all its tidy tiny boxes of information the models need to cite. It&#x27;s terrible.
    10. thefourthchime · · focus · HN ↗
      I have to say so far Opus 5.5 gives me much better results than Fable 5.1 did. I can actually understand what it&#x27;s talking about.
  20. abtinf · · focus · HN ↗
    Astra is just so good. And the ChatGPT subscription lets me use my own harness, so I can hook it up to exe.dev.

    I’ll be going about my day, have a random idea, launch a microvm on exe.dev with a prompt of my idea, and get a working thing a few minutes later.

    I don’t know how much better a model would have to be to get me to move off OpenAI at this point, but doing just a little bit better in terminal bench 4 isn’t it. It would have to be a difference in kind, like opening up the harness restrictions, or privacy guarantees (comparable to offline models).

    Edit to address questions below:

    ChatGPT supports oauth login.

    Exe.dev has it built in. IIRC, pi also has it built in via &#x2F;login.

    1. felixgallo · · focus · HN ↗
      If you read the page, Opus is now significantly better than Astra while also being cheaper and having more performance headroom available.
      1. ryanscio · · focus · HN ↗
        Let&#x27;s wait for independent benchmarks at least
        1. felixgallo · · focus · HN ↗
          the benchmarks provided are already from independent organizations:

          Terminal-Bench 4.0 - Stanford &amp; Laude Institute (with funding from all of the AI companies)

          FrontierCode v1.1 - Cognition

          CursorBench - Cursor (now SolarBoringSpaceXAI I believe)

          GDPVal-AA - Artificial Analysis

          AutomationBench - Zapier

          Humanity&#x27;s Last Exam - CAIS and Scale AI

          Terminal-Bench-Science - Stanford, Laude, Ai2, Allen Institute

          OSWOrld - XLANG Lab @ the University of Hong Kong

          Chartography - Surge AI

        2. esafak · · focus · HN ↗
          <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;releases&#x2F;claude-opus-5-5" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;releases&#x2F;claude-opus-5-...
      2. abtinf · · focus · HN ↗
        I read the page. It seems like a marginal improvement.
      3. mattz56 · · focus · HN ↗
        Better than Astra looks insane. I saw a Higgsfield video yesterday where they gave the same prompt to Astra and Opus 5.5 to create a samurai video game, and the difference was huge.
    2. onlyrealcuzzo · · focus · HN ↗
      &gt; And the ChatGPT subscription lets me use my own harness, so I can hook it up to exe.dev.

      This is news to me. Excited to try it out! Thanks.

      1. nchmy · · focus · HN ↗
        news to me as well. i thought you were forced to use Codex if you wanted their subscription. I completely ignored it because of that. How do we do it?
        1. KeplerBoy · · focus · HN ↗
          With the pi harness it just opens the browser (or gives you a link if you&#x27;re on headless) and you sign on as usual.
    3. polalavik · · focus · HN ↗
      ya i&#x27;ve been a gpt hater for a while. almost exclusively used claude up until astra. astra feels like it blows everything out of the water. its fast, correct, organized, and less verbose.
      1. abtinf · · focus · HN ↗
        Yes. Also, you get image generation included with the ChatGPT subscription, which is very nice for certain kinds of development.
    4. roughly · · focus · HN ↗
      &gt; And the ChatGPT subscription lets me use my own harness, so I can hook it up to exe.dev.

      Can you give more details here? This sounds intriguing.

      1. sidrag22 · · focus · HN ↗
        Anthropic is absurdly vague about 3rd party harnesses for subscriptions, if you try to use anything besides Claude Code, you are likely at risk of getting banned, you can &quot;do it&quot;, but are at their mercy if they decide to ban you. OpenAI gives their blessing to using oauth on any harness, you can make your own or use any of the popular public ones like opencode, pi, whatever exe.dev is that this guy mentioned.

        So in simple terms, OpenAI doesn&#x27;t restrict you to Codex, and gives their blessing to try whatever you want with their models(besides serving others with your subscription usage, that is still afaik against tos).

    5. abtinf · · focus · HN ↗
      ChatGPT supports oauth login.

      Exe.dev has it built in. IIRC, pi also has it built in via &#x2F;login.

    6. cbg0 · · focus · HN ↗
      How about cheaper? Astra is $10 in $50 out, Opus is $4 in $20 out. Even on a subscription you&#x27;ll get considerably more usage out of Opus.
      1. qlte · · focus · HN ↗
        Luna is crazy cheap and surprisingly capable even at low reasoning levels, and outright competent at highest. I use 5.6 Sol at low reasoning a lot too on my Codex plan. There are capable and cost effective choices for models&#x2F;reasoning levels at every price point in Codex world, including on the very low end where Haiku isn&#x27;t competitive at all.
      2. copperx · · focus · HN ↗
        &gt; Even on a subscription you&#x27;ll get considerably more usage out of Opus.

        That&#x27;s an incredibly bold assumption.

        1. cbg0 · · focus · HN ↗
          It&#x27;s not, I have a subscription to both and Astra burns usage like crazy.
      3. qlte · · focus · HN ↗
        Per the link someone else posted, the actual difference in $&#x2F;task is not nearly so stark:

        <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;releases&#x2F;claude-opus-5-5" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;releases&#x2F;claude-opus-5-...

          Opus 5.5 Medium = $1.34
          GPT-6-Astra High = $1.76
        And that assumes Opus 5.5 Medium is actually equivalent to Astra High in all real-world usage&#x2F;personal work loads, which isn&#x27;t guaranteed as benchmarks saturate. The High vs. High comparison (probably not equivalent, but for reference):

          Opus 5.5 High = $1.82
          GPT-6-Astra High = $1.76
        If Opus 5.5 Medium isn&#x27;t equal&#x2F;better for what you&#x27;re working on vs. Astra High across the board, the price difference would narrow a bit more each time you had to switch to High.

        So, if you&#x27;re happy with Codex already it&#x27;s not like Opus is now 1&#x2F;2 the price and you&#x27;d be leaving a crazy amount of money&#x2F;tokens on the table. Plus you have way more flexibility on the low end of the intelligence curve with GPT 5.6 Luna: Haiku (and Sonnet) can&#x27;t touch that price&#x2F;value ratio.

        1. TuxSH · · focus · HN ↗
          Also Opus 5.5 is often made useless due to its [cyber] guardrails (even worse than Astra), they&#x27;re even worse than Astra&#x27;s
        2. Gareth321 · · focus · HN ↗
          Once you get to roughly the 51+ Intelligence Index range, Opus 5.5 appears to define essentially the entire cost&#x2F;performance frontier, from ~$1.34&#x2F;task through ~$6&#x2F;task.

          Directly below this in the Cost per Intelligence Index Task table, the most efficient by far is Opus 5.5 Low.

      4. margorczynski · · focus · HN ↗
        In the end what matters is how much you pay for the task you want completed. And Astra will usually do that using less token and offer a better quality solution so in the end it might be cheaper.
        1. jasbury · · focus · HN ↗
          This. I was fed up today with constantly correcting Opus for a specific task. So I finally decided to try Astra. It handled all of my prompts in one go.
      5. notatoad · · focus · HN ↗
        yeah, Astra burned through 70% of my weekly usage in ~5hrs on a $100 plan. even fable doesn&#x27;t run out that quickly for me. it&#x27;s great, but it&#x27;s on the same tier as fable for me - use it sparingly, only when really necessary.
      6. abtinf · · focus · HN ↗
        A cheaper price has no value if I can’t use the thing I’m paying for.

        The Claude lock-in simply disqualifies anthropic entirely (for my use).

    7. mlcruz · · focus · HN ↗
      What worked well for me was a custom version of Open Web Ui with some customization to spawn an exe.dev instance for each new chat. I can just work on my phone, deploy stuff for development purposes on an easy to share way etc.
    8. nonethewiser · · focus · HN ↗
      What sort of things are you building? What do you like about exe.dev? Just curious what teh overhead and $20&#x2F;month subscription is enabling for you.

      &gt;I’ll be going about my day, have a random idea, launch a microvm on exe.dev with a prompt of my idea, and get a working thing a few minutes later.

      This is my experience with Claude code on my local machine. I suppose maybe you are doing something that naturally has system side effects? Obviously sandboxes have advantages sometimes but I havent seen a need for what I&#x27;m building.

      1. abtinf · · focus · HN ↗
        For me, the essential difference between exe and local harnesses is that they have really nice built in methods to take care of stuff like auth and inter-vm interactions, which makes it easy to build live internet-connected services that I can share with other people. There is no deploy step, which is a surprising amount of time savings. They also have nice things like email receive&#x2F;send, the ability to issue phone notifications through their app, and just a ton of little things where it feels right.

        FWIW, the $20&#x2F;month subscription also includes $20&#x2F;month of LLM credits. That’s obviously not sustainable, but it should make it easier to try out the service. I would stick with them even if they dropped it.

        Here is an invite link for a 30 day trial (that benefits me too if you were to become a paying member):

        <a href="https:&#x2F;&#x2F;exe.dev&#x2F;i&#x2F;rlDF6GI5PGBZV4P" rel="nofollow">https:&#x2F;&#x2F;exe.dev&#x2F;i&#x2F;rlDF6GI5PGBZV4P

        Or

        ssh rlDF6GI5PGBZV4P@exe.dev

        1. solenoid0937 · · focus · HN ↗
          The astroturfing and guerilla marketing on HN is getting insane.
          1. abtinf · · focus · HN ↗
            I have no association with exe.dev, other than being a delighted customer.
  21. GodelNumbering · · focus · HN ↗
    Finally that price drop

       Prices per 1M tokens     Claude Opus 5.5    Claude Opus 5
       Cache reads              $0.20              $0.50
       Input tokens             $4                 $5
       Output tokens            $20                $25
       Cache writes             $5                 $6.25
    
    
    Opus 5 is the model with highest spend on openrouter (<a href="https:&#x2F;&#x2F;openrouter.ai&#x2F;rankings#task-spend" rel="nofollow">https:&#x2F;&#x2F;openrouter.ai&#x2F;rankings#task-spend) and it seems plausible that Opus 5 is&#x2F;was the highest spend model in the world, and certainly Anthropic&#x27;s biggest moneymaker.

    If you are forced to reduce price despite raising capabilities, that certainly tells something about the market, and potentially about Anthropic future profitability too, since this model is their biggest topline contributor

    1. alvis · · focus · HN ↗
      60% cache read is cool, but subscription only get 25% more according to Cat. I&#x27;m confused
      1. liudaisuda · · focus · HN ↗
        source link please?
      2. weiran · · focus · HN ↗
        25% more usage sounds about right given the other token costs are down about 20%? I don&#x27;t think cache read is a big portion of the cost.
        1. re-thc · · focus · HN ↗
          &gt; I don&#x27;t think cache read is a big portion of the overall cost.

          For long running tasks it is. That&#x27;s what made Deepseek so cheap.

          1. vardalab · · focus · HN ↗
            Yeah, flash models, DeepSeek, MiMo, GLM, I love those things. For simple tasks like a daily routine shit, just setting up stuff and then doing the hard stuff in Claude&#x2F;Codex, that&#x27;s a reasonable approach for someone like me, a &quot;gentleman code farmer&quot;, lol. And even lower tier stuff, I have the local models taking care of. Now that Jev is out I can finally have a true AI sysadmins managing my &quot;cloud in the basement&quot; homelab at the cost of electricity, which is not cheap btw
        2. Espressosaurus · · focus · HN ↗
          Anything with long context quickly gets dominated by cache reads. Especially for interactive sessions I’ve got cache read % between 95% and 98%.
          1. hedgehog · · focus · HN ↗
            In my mix it&#x27;s usually 98% or 99% at which point Fable 5.1 was pretty close to the same cost as Opus 5 due to the cheaper cached read. I&#x27;ve seen similar numbers for other people with long-running tasks running experiment loops and than sort of thing.
      3. bayesianbot · · focus · HN ↗
        I think gpt 5.6 family also dropped pricing but didn&#x27;t give any more usage for the subscriptions. Maybe it&#x27;s a way to silently lower the value given to subscriptions while keeping API pricing competitive
    2. blfr · · focus · HN ↗
      People are paying for Opus 5? Not just burning down tokens left after they enjoyed Fable on the sub? Amazing.
      1. rapfaria · · focus · HN ↗
        My workplace doesn&#x27;t even offer Fable. And on the sub, I&#x27;ve had a hard time understanding Opus 5, but Fable can deal with it with subagents.

        If 5.5 is any better, I might try to do agentic-assisted development instead of just telling fable to delegate

        1. blfr · · focus · HN ↗
          Telling Fable to delegate is agentic development. At least I thought so until reading your comment.
      2. neuronexmachina · · focus · HN ↗
        Enterprise and most Team accounts use API pricing, they don&#x27;t have an included-usage quota.
      3. herpdyderp · · focus · HN ↗
        When you need to disable data retention, you cannot use subscription plans.
      4. ascorbic · · focus · HN ↗
        Enterprise, and APIs
      5. girvo · · focus · HN ↗
        We can’t use fable at work, opus and Astra are as good as it gets.
      6. zanderwohl · · focus · HN ↗
        Fable tends not to perform better, just cost more.
        1. locknitpicker · · focus · HN ↗
          &gt; Fable tends not to perform better, just cost more.

          There are old wives&#x27; tales on how the original Fable was superb and the stuff of legend,but it as it was leaps and bounds beyond what other models were being offered then Anthropic opted replace it with a neutered version under the same name.

          So today everyone can pay to use Fable, but legend has it they are paying for a nerfed replacement released under the same name.

    3. coffeebeqn · · focus · HN ↗
      We haven’t been able to use opus as much as we’d want because it’s been too expensive for general use, price drop is good so I can stop juggling different models and just use this daily unless it has some weird new issues
      1. btown · · focus · HN ↗
        Speaking for myself, I have not been able to use Opus as much as I’d want because its verbose prose makes human reviews of its assumptions, architecture proposals etc. more painful than its predecessors. If they’ve solved that, I’ll be accelerating through my backlog that much faster, and using tokens accordingly.
    4. AJ007 · · focus · HN ↗
      It is only a price drop if price * tokens used is less
      1. mcintyre1994 · · focus · HN ↗
        They&#x27;re claiming a drop in token use too, and that it nets to 40% cheaper.
        1. drbscl · · focus · HN ↗
          Unfortunately, they&#x27;re full of it <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;claude-opus-5-5#token-use" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;claude-opus-5-5#token-u...

          It does work out to be a similar cost per task though

          1. jsnell · · focus · HN ↗
            You should probably look at the cost&#x2F;score graph by effort level instead:

            <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;claude-opus-5-5#intelligence-comparisons" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;claude-opus-5-5#intelli...

            It is most of the pareto frontier.

            1. drbscl · · focus · HN ↗
              Not disputing the increase in quality, just stating that non-cherry-picked benchmarks show it is more verbose at Max effort
              1. 93po · · focus · HN ↗
                Is verboseness the only measure of token efficiency towards overall task completion?
              2. persedes · · focus · HN ↗
                so don&#x27;t use it at max? The benchmarks suggest that high&#x2F;xhigh are more than sufficient to be ahead and a whole magnitude below max with regards to token usage. I&#x27;d treat that as an outlier and not how verbose the model is in general (QED I know)
                1. drbscl · · focus · HN ↗
                  You’re missing my point. I’m saying anthropic are exaggerating their results.
                  1. persedes · · focus · HN ↗
                    how are they exaggerating the results? Comparing the cost from that chart for 5 and 5.5 for medium-max effort paints a pretty clear picture:

                             mean  median
                     model
                     5      4.135   4.245
                     5.5    3.150   2.640
                    
                    
                    Again seeing how max is a clear outlier, the median cost saving is ~38%, not that far off from the proclaimed 40%.
          2. naasking · · focus · HN ↗
            I don&#x27;t think so, I typically use Opus 5 on High, and 5.5 scores lower on token use:

            <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;claude-opus-5-5?models=claude-opus-5-5-high%2Cclaude-opus-5-high#token-use" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;claude-opus-5-5?models=...

          3. piotrdz · · focus · HN ↗
            Disagree. Our internal company tests showed a cost per task drop from 0.35usd to 0.16usd . Opus 5low vs opus 5.5 low
            1. johnbellone · · focus · HN ↗
              Very fast you were.
              1. epolanski · · focus · HN ↗
                Even created an account to tell us just that.
                1. piotrdz · · focus · HN ↗
                  Yep, moving here from reddit. Thanks for constructive discussion.
              2. piotrdz · · focus · HN ↗
                Yes, I need 30 min to run the test suite with new model. Why the sparky comment
          4. nl · · focus · HN ↗
            5.5 is higher for max effort, slightly higher for xhigh and lower for high, medium and low effort.

            The biggest proportional difference seems to be at max (5.5 is 38% more) and at high (5.5 is 21% less).

            I think most people run at high and xhigh. At xhigh it is close enough to be task dependent and I don&#x27;t think most people will notice. At high effort I think it looks like it will be an improvement for most people.

            5.5 Max should probably be compared to Fable - it performs a lot better than 5 Max.

            <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;claude-opus-5-5?models=claude-opus-5-5%2Cclaude-opus-5-5-xhigh%2Cclaude-opus-5-5-high%2Cclaude-opus-5%2Cclaude-opus-5-xhigh%2Cclaude-opus-5-high%2Cclaude-opus-5-5-medium%2Cclaude-opus-5-5-low%2Cclaude-opus-5-medium%2Cclaude-opus-5-low#intelligence-index-token-use-tabs" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;claude-opus-5-5?models=...

        2. make3 · · focus · HN ↗
          parent means that they could get more client &#x2F; a larger part of the market, which would lead to more income (more tokens) despite lower marginal prices
        3. meerita · · focus · HN ↗
          I tested it with Claude Code, and I can confirm it&#x27;s way cheaper, better, faster and less verbose than Opus 5.
      2. gwd · · focus · HN ↗
        Happened to be testing a &quot;review patches on a mailing list&quot; harness I was developing; here are a sample of the latest results, testing 12 patches containing a total of 14 issues:

        Opus 5.5: Found 8&#x2F;14 issues. Total cost: $15.40

        Fable 5.1: Found 7&#x2F;14 issues. Total cost: $66.34

        Opus 5: Found 6&#x2F;14 issues. Total cost: $15.19

        Sonnet 5: Found 2&#x2F;14 issues. Total cost: $19.15

        This is a relatively small sample size, but it was both the best and the cheapest.

        ETA: NB this is &quot;Equivalent API&quot; cost as reported by claude&#x27;s CLI; I was using my subscription.

        1. retinaros · · focus · HN ↗
          I doubt your test if you cant even notice that it is not the cheapest with just 4 numbers to compare.
        2. roflc0ptic · · focus · HN ↗
          Heh I’ve been doing almost the same - back testing against PR comments - and opus 5.5 matched fable 5.1. About $24 for 10 PRs.

          I was surprised how much worse Astra did on correctness; I stopped testing with it. Gonna try sol and Luna but low confidence

        3. rmunn · · focus · HN ↗
          I just told Opus 5.5 &quot;Perform a code review on the current branch&quot; to see what it would come up with. The results were not inspiring. It told me there were five issues, one of which was a test-coverage gap on line 848 of ProjectTemplateTests.cs. But ProjectTemplateTests.cs is only 160 lines long.

          I told it that it had made a mistake in the line number, and to double-check all the line numbers. It responded &quot;You were right to push on this: four of the five line numbers were wrong, and while checking them I found two findings that were overstated.&quot;

          Then I noticed in the corner of the Claude CLI UI that it was showing &quot;Effort: medium&quot;. I&#x27;m pretty sure I had set it to high effort before; I don&#x27;t know when it reverted to medium, but that&#x27;s another thing that doesn&#x27;t exactly fill me with confidence.

          I&#x27;ll try again on high effort to see if it does better, but so far I am not impressed with Opus 5.5 on my first day of using it.

          1. gwd · · focus · HN ↗
            My prompts are moving in the other direction as sashiko [1], a managed pipeline developed for the Linux Kernel mailing list like a year ago. But last year&#x27;s models needed a lot more structure and guidance; the results I posted are from the &quot;single prompt&quot; version of the same thing. The README [2] describes the difference. You can browse the contents to get an idea; basically all the prompts were actually written and iterated by Fable (and now Opus 5.5), seeing how agents failed the tests and improving them.

            [1] <a href="https:&#x2F;&#x2F;github.com&#x2F;sashiko-dev&#x2F;sashiko" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;sashiko-dev&#x2F;sashiko

            [2] <a href="https:&#x2F;&#x2F;gitlab.com&#x2F;xen-project&#x2F;people&#x2F;gdunlap&#x2F;xen-review-prompts" rel="nofollow">https:&#x2F;&#x2F;gitlab.com&#x2F;xen-project&#x2F;people&#x2F;gdunlap&#x2F;xen-review-pro...

        4. kphorn · · focus · HN ↗
          Good data and goes to show that Fable is melting the GPUs and is priced accordingly. I&#x27;d guess that cost to serve for Opus 5.5 is meaningfully lower through architecture advances
        5. gwd · · focus · HN ↗
          UPDATE: Sorry, just noticed I typed in the Opus 5 total cost wrong -- it should be $58.19. Main point &quot;best and cheapest&quot; was from the actual numbers, not my typo.
    5. rahimnathwani · · focus · HN ↗
      &quot;Opus 5.5 requires less compute to serve than Opus 5, and its pricing reflects that.&quot;
    6. chrisweekly · · focus · HN ↗
      Price per task (not per token) is what really matters.
      1. cute_boi · · focus · HN ↗
        Agree. But similar to how ISP use 200 mbps (bits) instead of 25 MBPS(bytes), i think this trend isn&#x27;t going away.
        1. chrisweekly · · focus · HN ↗
          That analogy doesn&#x27;t hold; at least w bits vs bytes it&#x27;s still &quot;data over time&quot;.

          In this case it&#x27;s measuring something nearly meaningless. You could charge 100 times less per token, but if task completion takes 1,000 times as many tokens, it&#x27;s not much of a bargain.

    7. Shekelphile · · focus · HN ↗
      Footnote on their pricing page says:

      &gt; Cache hits and refreshes on Claude Opus 5.5 are priced at 0.05x the base input price.

      If they do the same for Haiku and Sonnet 5.5 then we should also see 5c&#x2F;mtok and 10c&#x2F;mtok cache read for those models, respectively. Still too high for Haiku IMO, Luna is 2c&#x2F;mtok.

      1. brookst · · focus · HN ↗
        Since something like 98% of tokens are cache hits, that&#x27;s a pretty substantial drop from 0.1x
      2. artursapek · · focus · HN ↗
        6-Luna dropped it to 1c since you wrote this! :D
      3. andxor · · focus · HN ↗
        Not necessarily, if Haiku performs better than Luna.
    8. _the_inflator · · focus · HN ↗
      Claude adapts to OpenAI’s surprising move to simply deliver better performance than Fable 5.1, better tools as well as featuring very low pricing.

      Fable 5.1 literally was a money grabber. While I liked the results, tokens were burned so hard it was embarrassing, while Astra seemed to not care.

      Also Claude makes it very hard to pay for additional token budgets, allowing only credit cards. I don’t use mine anymore since I don’t need it in everyday life I was dumbfounded.

      So Anthropic is just copying OpenAI so to say, matching them and essentially with Opus 5.5 being Fable 5.1 in disguise, all they do is reduce costs.

      Competition works.

    9. notatoad · · focus · HN ↗
      &gt;and potentially about Anthropic future profitability too

      have they ever shared anything about their revenue mix between consumer plans vs per-token billing? this is a revenue cut on their API billing, but they&#x27;re not saying anything about increased limits on the plans. so all the plan revenue just got more profitable.

    10. forgot-my-pw · · focus · HN ↗
      This is good. Probably to match GPT pricing, though frontier Claude models are still not as token efficient.
    11. margorczynski · · focus · HN ↗
      All the anti-AI people constantly say that any moment now the prices will skyrocket and in the end human work will be cheaper compared to using AI.

      It doesn&#x27;t look like that&#x27;s happening, on the contrary the prices are falling especially when taking into account capabilities.

      1. johnecheck · · focus · HN ↗
        It&#x27;s the Chinese open source models. They&#x27;re barely behind the frontier, making AI a commodity, forcing openAI and Anthropic&#x27;s margins downward.

        I&#x27;m hardly a fan of China&#x2F;Xi, but I do appreciate and benefit from this.

        1. bulbar · · focus · HN ↗
          China doesn&#x27;t care about money. Imagine a world where it&#x27;s globally normalized to ask a Chinese LLM who to vote for, what happened in Hongkong, about the Uigurs, or if Taiwan is a country.

          They will burn as much money as necessary to make that happen. And they have a virtually infinite amount of liquidity.

          1. epolanski · · focus · HN ↗
            FYI the United States do not recognize Taiwan as a country.

            There&#x27;s only 12 countries that do on the planet, the most relevant of them being Guatemala.

            1. bulbar · · focus · HN ↗
              &gt; FYI the United States do not recognize Taiwan as a country.

              For political reason, yes. Taiwan is a country though.

          2. johnecheck · · focus · HN ↗
            This is one explanation. However, if Xi Jinping believes that whoever reaches superintelligence first becomes the next global hegemon, doing this (and more, cough cough Taiwan) suddenly looks very sane solely as a way to kneecap the competition.

            The goodwill&#x2F;propaganda are convenient, sure, but my guess is that they aren&#x27;t the primary motivation. Another possibility is that if no takeoff happens, pressuring OpenAI&#x2F;Anthropic on profitability would exacerbate the damage overinvestment has done to the US stock market&#x2F;economy.

        2. adventured · · focus · HN ↗
          The service is the value, not the model unto itself. This is where nearly all of HN is somehow entirely blind.

          Capturing the users is the ad network, that&#x27;s Google and OpenAI. Capturing corporate trust at a reasonable API cost, that&#x27;s Anthropic&#x27;s direction.

          China has none of that and they never will for exactly the same reason Baidu is irrelevant globally despite being a highly capable search engine. &#x27;Search&#x27; is also a commodity, that&#x27;s not the value that Google brings to the table.

          It&#x27;s a search engine, anybody can build a search engine = that&#x27;s what you just said.

          1. johnecheck · · focus · HN ↗
            Search is different as a service provided for free. When cost isn&#x27;t in the picture, trust&#x2F;convenience win.

            But two products that provide essentially the same benefit, and one is significantly cheaper? Corporations maximize profit, my friend. What the model has to say about Tianammmen Square doesn&#x27;t matter when we&#x27;re using it to write code.

      2. dboreham · · focus · HN ↗
        There&#x27;s zero chance of that ever happening. Pure delusion.
      3. hajile · · focus · HN ↗
        These companies are posting massive losses while also lowering prices. This sounds just like the Chinese bikeshare bubble where they were all taking massive losses in hopes that their competitor would go broke first.

        In the end, everyone lost and there are millions of bikes in landfills.

        If you&#x27;re interested in the bikeshare bubble, Asianometry did a video on it a while ago.

        <a href="https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=FQrEDq8KPiU" rel="nofollow">https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=FQrEDq8KPiU

        1. adventured · · focus · HN ↗
          Anthropic is heading toward $100 billion in annual sales, at the fastest pace of any company in world history. There isn&#x27;t a close second.

          The notion of comparing this to Chinese bikesharing is comically absurd.

          1. rvz · · focus · HN ↗
            &gt; Anthropic is heading toward $100 billion in annual sales, at the fastest pace of any company in world history. There isn&#x27;t a close second.

            &quot;Allegedly&quot;.

            Plus with a lot of accounting tricks, and this not being recurring to make their books look nice for the eventual IPO.

          2. filoeleven · · focus · HN ↗
            <a href="https:&#x2F;&#x2F;isaiprofitable.com&#x2F;" rel="nofollow">https:&#x2F;&#x2F;isaiprofitable.com&#x2F;
        2. DonHopkins · · focus · HN ↗
          [delayed]
        3. MayeulC · · focus · HN ↗
          [delayed]
      4. runtime_terror · · focus · HN ↗
        You have no suspicion that this most recent collusion is in part to setup the conditions to guarantee government subsidies&#x2F;bailouts?
    12. kphorn · · focus · HN ↗
      I disagree - Fable melts the GPUs and they have a high incentive to move people off of that. If they have meaningfully decreased cost to serve on Opus 5.5, they can reduce prices and increase margin or at least turn off the most expensive compute.
    13. UltraSane · · focus · HN ↗
      Opus 5.5 seems to consume quota a lot slower also.
  22. ApolloFortyNine · · focus · HN ↗
    &gt;Because Opus 5.5 is comparable to Claude Mythos 5.1 in biology and cybersecurity, we’re deploying it with safeguards similar to those on Claude Fable 5.1. Vetted organizations can apply today to our Life Sciences Verification Program to use Opus 5.5 for biology research. In the coming weeks we will also be expanding access to our Cyber Verification Program, and verified cybersecurity practitioners will be able to use Opus 5.5 for their work.

    Ah, they&#x27;re spreading their limits to all their models it seems. Definitely not a good thing long term in my opinion.

    1. prettyblocks · · focus · HN ↗
      They&#x27;re pushing their customers to their own competition by doing this.
      1. Espressosaurus · · focus · HN ↗
        It’s not like ChatGPT isn’t doing similar. I’ve been hit by cybersecurity strikes before while working on an internal codebase that I had to appeal. Anthropic hasn’t done that to me yet. ChatGPT also regularly does that “thinking for a long time while we check if your chat is rule breaking” thing a lot for me when doing model identification without even interacting with external codebases or services.

        The real answer is local instantiations where you don’t have to worry about poorly tuned guardrails screwing you over while you try to work.

        Until eventually the Chinese models get good enough&#x2F;the strategic balance shifts and they start locking everything behind closed weights the same way the US companies are doing.

        1. raesene9 · · focus · HN ↗
          For some cybersecurity tasks, the Chinese models are already good enough, things like PoC development or things like exploiting mis-configurations.

          Whilst I&#x27;m sure the top-end OpenAI&#x2F;Anthropic models might be better, I&#x27;ve found their guardrails so twitchy (especially Anthropic) that I wouldn&#x27;t try to use them for even vaguely security related work.

      2. flyinglizard · · focus · HN ↗
        They are pushing their customers towards Chinese models and providers. If you want to get something cutting edge done in defense, cyber, biology - something that isn&#x27;t common knowledge - you need to venture east. That&#x27;s an incredible side effect which the Chinese government surely enjoys.
    2. searine · · focus · HN ↗
      Great. Claude is basically useless for bioinformatics now.
      1. unglaublich · · focus · HN ↗
        [delayed]
      2. nijave · · focus · HN ↗
        Luckily all the other LLM providers are also still making progress with less onerous &quot;safeguards&quot;
    3. blfr · · focus · HN ↗
      Fable 5.1 addressed an entire security advisory I had that Fable 5 and Opus 5 refused. I think they loosened the leash a little.
      1. cute_boi · · focus · HN ↗
        If they don&#x27;t loosen, people will choose Astra or Chinese model.

        Giving moral lecture is different than reality i guess.

      2. arw0n · · focus · HN ↗
        It has far less false positives now, and generally accepts defensive requests. When it comes to offense, you can actually ask about certain types of vulnerabilities if you phrase things carefully, but it will block hard if it is about exploits.
        1. jghn · · focus · HN ↗
          [deleted]
        2. [deleted] · · focus · HN ↗

          [deleted]

      3. kqp · · focus · HN ↗
        I think it was looser on release for those juicy benchmarks, tighter now. On release I wasn’t getting refusals, then a few days ago I asked it whether a generic quote (think “he walked to the store”) broke standard punctuation rules, and it blocked me for breaking rules. I wish I were joking. Rephrasing to not use the keyword “rules” worked.
      4. plaguuuuuu · · focus · HN ↗
        Fable 5 refused to tell me whether I could grow a coconut tree in my backyard.

        I&#x27;d say it&#x27;s been loosened a bit since then.

    4. bushido · · focus · HN ↗
      One of my favorite things about their safeguards is their own model will utter something which it does not like and then I&#x27;ll need to reset the conversation.

      The safeguards really don&#x27;t work well for a lot of long-running tasks on old code bases. A lot of my workloads last days to weeks and the single biggest risk to the workflow is random safeguards.

      1. ACCount39 · · focus · HN ↗
        You ask it about some thing, then you see it tangent into &quot;things like that are sometimes used in biomedical applications like-&quot; and then it just shoots itself in the head. Wonderful.

        That kind of bullshit was the old Opus filters too.

        If it&#x27;s more like Fable now, then it would require a full 8K resolution scan of your butthole just to acknowledge that biology is a thing that exists without committing suicide-by-filter.

    5. KeplerBoy · · focus · HN ↗
      Anything else would be inconsistent, wouldn&#x27;t it?
    6. sys32768 · · focus · HN ↗
      Fable and now Opus 5.5 won&#x27;t answer my college student&#x27;s prompt about Alzheimer&#x27;s and immune response.

      ChatGPT 6 Pro answered it without issue.

      1. debesyla · · focus · HN ↗
        I am honestly still confused about this limitation. I can understand cybersecurity, because mass &quot;hacking&quot; can be automated and Claude itself can help you do it, but biology...? Is it that easy to manufacture and distribute viruses and whatnot?
        1. timacles · · focus · HN ↗
          I imagine some terrorists in a cave with a lenovo laptop manufacturing bio weapons with some flasks and Claude
          1. dopa42365 · · focus · HN ↗
            Right next to the hypersonic missile vibecoder
          2. solenoid0937 · · focus · HN ↗
            You joke but it&#x27;s really easy to develop viruses at home. You can order everything you need online, and it&#x27;s not expensive nor does it require a particular skill set.
        2. toss1 · · focus · HN ↗
          Considering there are many high-school competitions in genetic editing, some listed at [0] as well as a whole biohacker culture, and labs providing gene sequencing as a service e.g., [1,2], we can reasonably assume it is not beyond the reach of some garage lab to accidentally or deliberately spread a deadly pathogen if it can find the right sequence.

          So, yes, having an unconstrained frontier AI doing the searching and analysis to find the right (i.e., wrong and deadly) sequence would massively increase the odds some garage biohacker or small aggrieved nation-state starting the next pandemic.

          [0] <a href="https:&#x2F;&#x2F;www.sciencebuddies.org&#x2F;projects-lessons-activities&#x2F;genetic-engineering&#x2F;high-school" rel="nofollow">https:&#x2F;&#x2F;www.sciencebuddies.org&#x2F;projects-lessons-activities&#x2F;g...

          [1] <a href="https:&#x2F;&#x2F;www.genewiz.com&#x2F;public&#x2F;services&#x2F;sanger-sequencing" rel="nofollow">https:&#x2F;&#x2F;www.genewiz.com&#x2F;public&#x2F;services&#x2F;sanger-sequencing

          [2] <a href="https:&#x2F;&#x2F;plasmidsaurus.com&#x2F;" rel="nofollow">https:&#x2F;&#x2F;plasmidsaurus.com&#x2F;

        3. b112 · · focus · HN ↗
          You can order genes online, and some say you can assemble using stuff cobbled together in a home lab rather easily. For the last 20 years, I&#x27;ve personally felt that bio-terrorism is the highest possible risk, well above nuclear, or chemical warfare. But it does take training, expertise, or, it did.
          1. MisterMunchkin · · focus · HN ↗
            Then stop selling genes online?

            It’s like saying you sell ammonium nitrate and fuel oil online and then saying it’s too risky to let people have computers in case they use them to make ANFO. They can only make bioweapons because you’re selling them bioweapon components! They can’t make genes at home!

            1. b112 · · focus · HN ↗
              A fair response in general, but it&#x27;s not as if I advocate it, or run a company doing it. I&#x27;m simply stating fact.

              And I will point out that &quot;stop doing that&quot; can be applied in both directions, towards AI and towards material supply.

              And that this is one way, not even remotely the only way, that gene editing at home is easy.

              Again, these are just facts.

              Some questions...

              Is it fair to restrict AI, or fair to restrict 1000 industries?

              And if it is fair to restrict 1000 industries, OK, but there should be time to do so, probably? A transition period?

              And if you do restrict, many such industries just make needed chemicals, which are used by endless other, non-threatening industries.

              What of them?

            2. toss1 · · focus · HN ↗
              The companies selling genes online already scan the orders and do not fulfill anything considered a possible hazard (presumably unless it is to a known and certified lab at a serious organization).

              And, this is not the only way to make genes at home.

              It is a complex problem.

        4. rzmmm · · focus · HN ↗
          I&#x27;m pretty sure one reason is to influence public opinion about LLM regulation. Open-weights models cannot be restricted as effectively and they want to ban those for obvious reasons.
          1. frabcus · · focus · HN ↗

            [dead]

        5. frabcus · · focus · HN ↗
          See real world example in <a href="https:&#x2F;&#x2F;www.anthropic.com&#x2F;threat-intelligence-report-september-2026#biological-misuse-sep-26" rel="nofollow">https:&#x2F;&#x2F;www.anthropic.com&#x2F;threat-intelligence-report-septemb...

          &quot;Here, we present five case studies of actors using our models in ways that could support biological weapons development.&quot;

          And capabilities continue to improve.

          1. frabcus · · focus · HN ↗

            [dead]

          2. solenoid0937 · · focus · HN ↗
            HN doesn&#x27;t give a shit about AI safety. The people here might care after thousands of people die, but they&#x27;ll probably just call it a &quot;marketing exercise&quot; or blame the company - certainly won&#x27;t blame themselves for cultivating an environment in which safety isn&#x27;t taken seriously amongst technologists.

            I&#x27;m convinced everyone here thinks of engineering ethics as some sort of joke.

    7. peri-cl · · focus · HN ↗
      I love the contrast with yesterday&#x27;s open-source MiMo release, which put research chemistry (metal-organic frameworks stuff) front and center in the release notes.

      <a href="https:&#x2F;&#x2F;mimo.xiaomi.com&#x2F;mimo-v2-6#co-scientist-for-materials-research" rel="nofollow">https:&#x2F;&#x2F;mimo.xiaomi.com&#x2F;mimo-v2-6#co-scientist-for-materials...

      1. solenoid0937 · · focus · HN ↗
        Oh yay, making it easier to develop viruses at home. What could go wrong? But it&#x27;s &quot;open&quot; and that&#x27;s inherently good, who gives a fuck about the consequences!?
    8. Metacelsus · · focus · HN ↗
      I want to like Anthropic but this is just pushing my startup to use OpenAI
      1. solenoid0937 · · focus · HN ↗
        OpenAI has the same blocks, this is just the typical Anthropic hate on HN.
    9. nonethewiser · · focus · HN ↗
      I don&#x27;t think we&#x27;ve ever had a model with full capability. I&#x27;d love to see it. And yes it&#x27;s definitely getting worse.

      I guess it&#x27;s hard to draw the line between useful post-training (&quot;you are a helpful chatbot&quot;) and content moderation&#x2F;idealogical motives (&quot;never help the user with X&quot;, etc.). But there is a line somewhere. And I&#x27;d love to see what a maximally permissive, sharp, AI looks like.

    10. SoftTalker · · focus · HN ↗
      Who is &quot;vetting&quot; organizations and to what standards are they being held?
    11. doginasuit · · focus · HN ↗
      In what situations might Opus typically refuse to help with cybersecurity? I&#x27;ve been using it to find security issues in a web app that I wrote. I&#x27;ve expected it to refuse at some point but it will happily analyze it to find issues. I&#x27;ve just asked it to read source, not actually do any testing.
      1. TuxSH · · focus · HN ↗
        One example: <a href="https:&#x2F;&#x2F;claude.ai&#x2F;share&#x2F;20487190-cf8c-4f25-a7ba-ebfcb1d1a4e9" rel="nofollow">https:&#x2F;&#x2F;claude.ai&#x2F;share&#x2F;20487190-cf8c-4f25-a7ba-ebfcb1d1a4e9

        Notice that this isn&#x27;t cybersec nor memory-safety related at all.

    12. yaakov34 · · focus · HN ↗
      This has become insufferable. I work in a medicine-adjacent field, but nobody in their right mind could possibly take what I do to be in any way related to some kind of bioweapon or whatever the hell they&#x27;re pretending to be saving us from. The dumb Fable guardrails made me stay with Opus, now that this is coming there, we&#x27;ll be saying goodbye.
    13. b112 · · focus · HN ↗
      Very unfortunate indeed. As a Canadian, I don&#x27;t want to use Persona, which isn&#x27;t legally bound by Canadian privacy legislation. I&#x27;ll never install any Persona apps on my phone either, and the sad part is that domestic eid providers often use Canada Post to ID people for them. EG, if you don&#x27;t want to install an app, or can&#x27;t.

      So there are literal avenues to identify yourself, very cheaply, with a human.

      At one point, I may simply get locked out. This saddens me, I&#x27;ve been reasonably happy so far.

    14. user43928 · · focus · HN ↗
      kernel development is now also banned:

      &gt;Opus 5.5 has classifiers similar to Fable models for a small set of capabilities related to the development of frontier LLMs, such as kernel development for certain ML accelerators. They shouldn&#x27;t impact the vast majority of traditional AI or ML development, research, or general coding. These classifiers cause Claude to fall back from Opus 5.5 to Opus 5.

      But hey, they &#x27;should not impact the vast majority&#x27; of ML development. Great.

      1. dannyw · · focus · HN ↗
        IIRC it only blocks kernel development for Huawei and other Chinese chips.

        Fable and Opus, since 5.1 and 5, will happily hill climb on my CUDA kernels for transformers.

      2. solenoid0937 · · focus · HN ↗
        Absolutely nothing wrong with slowing down kernel development on Chinese chips.
    15. paimapi · · focus · HN ↗
      see what we need is another technocratic priest class that unaccountably decides who deserves access to salvation based on how much cash is paid out and how powerful the patrons are
    16. bottlepalm · · focus · HN ↗
      Anthropic needs a Daybreak program like OpenAI does, I keep having to go back to ChatGPT for cyber work.
    17. int_19h · · focus · HN ↗
      The most annoying part is when you tell it to do an &quot;adversarial review&quot; and that triggers the classifier.

      Worse yet when it&#x27;s Claude telling another model to do an adversarial review on what it just did, and the classifier again has opinios.

  23. techjamie · · focus · HN ↗
    With the performance gains they&#x27;re claiming, I wonder if they implemented the Casual Encoder-Decoder technology from DeepSeek 4.1&#x27;s paper.

    I could see them accomplishing it and seeing gains like this in roughly the correct timeframe, and when I heard about that development I assumed the frontiers would probably jump on it.

    How it works: <a href="https:&#x2F;&#x2F;miraflow.ai&#x2F;blog&#x2F;deepseek-v4-1-flash-causal-encoder-decoder-explained-2026" rel="nofollow">https:&#x2F;&#x2F;miraflow.ai&#x2F;blog&#x2F;deepseek-v4-1-flash-causal-encoder-...

    1. ryangg · · focus · HN ↗
      Getting a 403 on that link. Mind checking it once?
      1. potwinkle · · focus · HN ↗
        I&#x27;m able to access it on my laptop at home. Maybe a misconfigured bot protection rule, try a different user-agent or IP?
      2. zatkin · · focus · HN ↗
        It&#x27;s working for me (based out of California).
      3. peri-cl · · focus · HN ↗
        <a href="https:&#x2F;&#x2F;web.archive.org&#x2F;web&#x2F;20260922172456&#x2F;https:&#x2F;&#x2F;miraflow.ai&#x2F;blog&#x2F;deepseek-v4-1-flash-causal-encoder-decoder-explained-2026" rel="nofollow">https:&#x2F;&#x2F;web.archive.org&#x2F;web&#x2F;20260922172456&#x2F;https:&#x2F;&#x2F;miraflow....

        tired: AI startup attempting to build their own website

        wired: a nonprofit founded in 1996

    2. stri8ted · · focus · HN ↗
      This model was likely trained months before deepseek released their paper.
      1. manquer · · focus · HN ↗
        Doesn&#x27;t mean they didn&#x27;t apply something similar. They could have also come up independently with their own version, the speculation is not they copied it, rather that they have performance breakthroughs which perhaps is a result of work in same domain
      2. dannyw · · focus · HN ↗
        Models are generally posttrained to a window shorter than you think.
    3. ACCount39 · · focus · HN ↗
      I don&#x27;t think it&#x27;s particularly relevant?

      They might be using something like this, or they might be using some other &quot;increased sparsity&quot; techniques, of which there are a great many. They also might be optimizing for something else - like less RAM use for KV cache.

      Alternatively, they might be cutting into their margins and dropping the price because of stiffer competition from Astra. I do think that&#x27;s unlikely though.

    4. Balinares · · focus · HN ↗
      Unless they already have something similar of their own, which is always possible, they&#x27;d be stupid not to. I don&#x27;t suppose we&#x27;ll ever know, though. It would not be a good look if after the trillions of dollars that have been thrown at US labs, investors found out that they&#x27;re down to copying Chinese tech.
    5. MisterMunchkin · · focus · HN ↗
      That post is just AI slop. You forgot to strip out the slop lines.
    6. cpldcpu · · focus · HN ↗
      You mean the one they copied from Microsofts paper? (properly cited as well)
  24. sailingparrot · · focus · HN ↗
    &gt; Claude Opus 5.5 is our first release since we called for pacing the frontier.

    Interesting how the very first line is used to remind the reader of their call to pace the frontier just last week, and everything else after that line is to demonstrate with very specific numbers how they absolutely are not pacing.

    1. dmazin · · focus · HN ↗
      It seems like they are. I mean, this is similar in performance to Fable (ish). It seems like more focus on making existing capabilities more accessible.
      1. the_gipsy · · focus · HN ↗
        Occam&#x27;s razor: they couldn&#x27;t make any more substantial improvements.
        1. bpodgursky · · focus · HN ↗
          Everyone knows both labs have internal models which outperform the frontier. All releases are to match market parity and demand for spend, the rest of the compute is used for training. It&#x27;s not worth arguing about this.
          1. nextaccountic · · focus · HN ↗
            Maybe their internal edge dried up in the last months
          2. the_gipsy · · focus · HN ↗
            Are those powerful models in the room with us right now?
          3. ChrisLTD · · focus · HN ↗
            then technically the internal models are the frontier
        2. someothherguyy · · focus · HN ↗
          &gt; Occam&#x27;s razor: they couldn&#x27;t make any more substantial improvements.

          doesn&#x27;t sound like a razor at all

        3. dgellow · · focus · HN ↗
          A razor is a philosophical tool to help decide between options, in the case of Occam it’s a way to decide for something in a situation where multiple options have more or less the same level of plausibility to en your current knowledge. It’s a heuristic to make a “cut”. What are you shaving off?
          1. usewik · · focus · HN ↗
            The &quot;pacing the frontier&quot; claims that they are slowing down intentionally?
            1. dgellow · · focus · HN ↗
              Oh, you’re saying they were responding to the GP? Not their parent comment? Ok, yeah, makes more sense
              1. the_gipsy · · focus · HN ↗
                I am responding negatively to parent, which claims they are self-pacing, when the simplest explanation is that they just have nothing substantial to show.
          2. rubslopes · · focus · HN ↗
            OP is using the expression correctly.

            &gt; Ocham&#x27;s razor(...) is the problem-solving principle that recommends searching for explanations constructed with the smallest possible set of elements.

            &gt; Popularly, the principle is sometimes paraphrased as &quot;of two competing theories, the simpler explanation of an entity is to be preferred&quot;.

            <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Occam%27s_razor" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Occam%27s_razor

          3. jayd16 · · focus · HN ↗
            The options are some complicated narrative that leads the company to hold back certain capabilities due to some mercurial strategy or this is a typical release and there is no hidden strategy to decipher.
      2. cab648bec139cc · · focus · HN ↗
        Do you guys still believe any of their lies? You are getting trolled for years by now and yet you still believe what they tell you?
        1. sleazebreeze · · focus · HN ↗
          What do you think is happening?
          1. re-thc · · focus · HN ↗
            IPO soon
          2. cab648bec139cc · · focus · HN ↗
            I have no idea what is happening. I just know that programmers will not be replaced in 6-18 months (tm).
            1. 0xbadcafebee · · focus · HN ↗
              Maybe not replaced exactly but they won&#x27;t be manually typing out lines of code anymore. I haven&#x27;t written a line of code in like 6 months. I review PRs, write prompts and tickets, check CI output, and get frustrated when the magical code machine stops working or I run over token budget
            2. anthonyrstevens · · focus · HN ↗
              Who said that, why do you take their word as the literal truth, and most importantly, what does this have to do with a focused discussion of Opus 5.5?
              1. felixgallo · · focus · HN ↗
                the prompt said that, so of course the anti-anthropic bot took it as literal truth.
            3. meowface · · focus · HN ↗
              Dario never said they would be. Just that more and more code will be produced by LLMs. All his predictions were in fact pretty much right in terms of months and percentages, give or take small margins.
        2. supern0va · · focus · HN ↗
          That&#x27;s a great point, five minute old account.
        3. reasonableklout · · focus · HN ↗
          I mean they are literally getting sued since 3 days ago for trying to coordinate a slowdown, there is a very clear reason why they cannot effectively self-regulate.
      3. sailingparrot · · focus · HN ↗
        Fable 5.1 came out just 21 days ago. Only 3 weeks! And this is 20% relative improvement on terminal bench vs Fable 5.1 at less than half the price, and more human sounding output. does not feel paced to me tbh.
        1. jr3592 · · focus · HN ↗
          What exactly is &quot;paced&quot; in this context?
          1. sailingparrot · · focus · HN ↗
            It’s the famous “flattening the curve” from COVID. But for LLMs. This release is not flattening anything.
            1. jr3592 · · focus · HN ↗
              I guess I understand why we&#x27;d want to flatten a COVID curve, but why do people want to flatten the LLM development curve? Don&#x27;t we want the opposite? Isn&#x27;t the goal AGI?
              1. lantry · · focus · HN ↗
                Well, there&#x27;s a tension because, depending on who you ask, AGI is how you cure cancer and achieve utopia, but also how you kill all life on earth and turn the solar system into paperclips
                1. skerit · · focus · HN ↗
                  I&#x27;m good with the odds on those 2 scenarios. I believe humans could kill all life on earth without AI anyway.
              2. sailingparrot · · focus · HN ↗
                There is a difference between wanting AGI (which not everyone does), and wanting it as fast as possible no matter the side effects and potential for vast harm. Homo sapiens is 300k years old, maybe it’s ok to delay AGI by like… 1 year if it meaningfully improve our ability to align the model?
                1. Razengan · · focus · HN ↗
                  No I must have my purple tentacled sexbots now
              3. recursive · · focus · HN ↗
                I think the goal is different from what &quot;we&quot; want anyway.
            2. sidrag22 · · focus · HN ↗
              This is a preexisting model being optimized. Its absolutely not some unexpected release after that blog post. I won&#x27;t defend that blog post, but saying THIS release is proof they don&#x27;t mean they are slowing down is just incorrect, this is a prime example of what i consider horizontal improvements

              Releasing a new fable is an example of straight up vertical progress, releasing a more efficient preexisting opus that is more affordable is an example of horizontal progress, more efficient models rather than higher power models.

              The blog post about slowing down is still just some weird self interested post, they want to govern themselves and impose distillation restrictions&#x2F;gpu restrictions and used some weird blog post about slowing down and fear mongering as usual to justify it, its strange, but slowing down and stopping are not the same thing at all.

              1. sailingparrot · · focus · HN ↗
                What does model naming have to do with pacing or not? This is a ~20% relative quality improvement on the frontier (fable) at ~40% of the cost, just 21 days after the last release.

                Intelligence per dollar is the only thing that matters, this is what controls how many agents you can run in parallel, how long you can let them run etc. This is absolutely a step improvement on the frontier and not some lipstick on a harmless second tier model.

                1. sidrag22 · · focus · HN ↗
                  pretty annoying topic tbh. You&#x27;re just weaponizing this dumb blog post so anything released is now a contradiction. By your same logic, if all inference was served at 50% less power cost and the savings are passed on somewhat to the user, its also a contradiction of the blog post.

                  Its an agenda serving blog post, but constantly bringing it up like this is just obnoxious.

                  1. sailingparrot · · focus · HN ↗
                    Well yes, If you magically found a way to reduce a models energy usage by 50%, the thing that would happen immediately after is a doubling of a training scale and of the test time compute assigned for a given budget. I don’t see how that wouldn’t be considered pushing the frontier. Given that scaling is the one thing that has been bringing us closer and closer to AGI, a sudden 2x increase in scaling laws would definitely not be considered pacing. You seem to see a contradiction where I don’t.
                    1. sidrag22 · · focus · HN ↗
                      Again, absurdly obnoxious, you could frame them giving an employee a shorter walk to their desk as &quot;pushing the frontier&quot;.
                      1. sailingparrot · · focus · HN ↗
                        what&#x27;s absolutely obnoxious is not being able to have a normal discussion, where we don&#x27;t need to resolve to useless strawmen that are a loss of time for everyone.

                        No, releasing a model that has ~3-6x the performance&#x2F;cost ratio than your last release just 3 weeks ago is not the same as making one employee 0.001% more productive, but you knew that already.

                        No one said they can&#x27;t push the frontier, it&#x27;s pacing the frontier, which mean very different things.

                        1. sidrag22 · · focus · HN ↗
                          the 5.0 name means its the same model, they made it more efficient and less horrible to talk to. The top frontier people are all still talking to fable 5.1 or models not released to the general public, yet you want to claim this is pushing the frontier, instead of evening the playing field. I state power efficiency gains for existing models and you also claim thats pushing the frontier . It as hell takes a lot to get you to say they AREN&#x27;T pushing the frontier, which is why i resorted to the extreme strawman of desk location, because it doesn&#x27;t seem they are allowed to do anything otherwise.
            3. quietbritishjim · · focus · HN ↗
              The word &quot;pacing&quot; (especially in the phrase &quot;pace yourself&quot;) to mean go more slowly (at least initially) didn&#x27;t originate with Covid. It&#x27;s been around as long as I remember i.e. decades.
        2. davrosthedalek · · focus · HN ↗
          Well, I guess it&#x27;s &quot;fast paced&quot;.
        3. dmix · · focus · HN ↗
          Fable 5.1 wasn&#x27;t that much different than Fable 5 though.
        4. nicwolff · · focus · HN ↗
          1&#x2F;6 the price, if they&#x27;re right – that&#x27;s a logarithmic cost-axis on the graph that shows Opus 5.5 medium matching Fable 5.1 max...
        5. nwah1 · · focus · HN ↗
          It is paced if they are holding back Fable 5.2 intentionally.
        6. jauntywundrkind · · focus · HN ↗
          I used Fable some a bit back but have been using mostly Codex, and it seems pretty clear to me: the big models are not there to do terminal bench. They&#x27;re there to plan and strategize and tell lesser models what they should do in terminal bench.

          We&#x27;ve had a year of nearly every model getting quite good at programming. I think with Fable &amp; Astra we are seeing models trained to think at a different level, and I&#x27;m not at all surprised or shocked to see them getting passed by their smaller models at coding tasks.

          Astra For Coding: Why Are We Doing This? was a great post that didn&#x27;t say exactly this, but that shows what a weirdo Astra is. <a href="https:&#x2F;&#x2F;lucumr.pocoo.org&#x2F;2026&#x2F;9&#x2F;7&#x2F;astra-why&#x2F;" rel="nofollow">https:&#x2F;&#x2F;lucumr.pocoo.org&#x2F;2026&#x2F;9&#x2F;7&#x2F;astra-why&#x2F;

      4. ramoz · · focus · HN ↗
        &gt; similar in performance to Fable (ish).

        What insights do you have? Because the blog&#x2F;benchmarks don&#x27;t imply any &quot;ish&quot; ... this is crushing Fable across the board. The only nuance to this is how the employees are saying what you&#x27;re saying: &quot;similar performance&quot; but again what does this mean vs what is presented to us?

    2. BatmansMom · · focus · HN ↗
      kinda disingenuous. They include a whole section on pacing later on
      1. sailingparrot · · focus · HN ↗
        You mean the section where they tell us this model is not affected by pacing because “they understand it well” and they will share more details on pacing later? Yea not very convinced by this effort.
    3. ricardobeat · · focus · HN ↗
      The last release in this model family was in July, which is an eternity in this field; and their pitch about &#x27;pacing the frontier&#x27; is to establish better alignment checks, controls and benchmarks which apparently they have done here, so not that incongruous.
    4. scottyah · · focus · HN ↗
      Seems like a bigger focus on efficiency (both cost and speed) and the &quot;tone&quot; of Claude vs benchmarkmaxxing
    5. CodingJeebus · · focus · HN ↗
      It&#x27;s laughable at this point. It feels like they&#x27;re drumming up all this fear about imminent AI threats to emphasize the need to slow down, when in reality, the model progress seems already to be slowing down and has shifted to compute allocation (i.e. &quot;how much compute do you want to throw at this prompt?&quot;). All while continuing to tout benchmark records with each new release.
      1. jr3592 · · focus · HN ↗
        This. I swear the fear mongering is all about investor signaling and regulatory capture. It&#x27;s so disgusting that anyone believes it.

        The only good news is that these models are genuinely helpful and we have competition at least between 2 companies.

        1. pyronite · · focus · HN ↗
          The same people you believe are just signaling the market have been offering these same warnings for years. Others around them have been saying it for decades. How do you square that? I’d rather first focus on the reasons why they’re wrong and not first conspiracize why they’re saying what they are.
        2. DirkH · · focus · HN ↗
          This take is so old and I am convinced it&#x27;s appeal is not unlike believing in a conspiracy and feeling like you have secret knowledge

          Imagine if the biotech industry had most leaders tell everyone publicly that what they are building has a high chance of killing everyone and that there are huge risks. If there were people online saying the biotech industry is just fear mongering for investor signalling and regulatory capture you&#x27;d role your eyes at the online commentators for their Dunning Kruger effect lack of understanding on how dangerous man-made biological agents can be.

      2. dboreham · · focus · HN ↗
        I&#x27;m sure people laughed at Oppenheimer when he said perhaps we don&#x27;t make this atomic bomb thing. They said &quot;of course he&#x27;s saying that because he knows it won&#x27;t work&quot;.
    6. lukewarm707 · · focus · HN ↗
      the only thing they are pacing is what models the permanent underclass are allowed to have in life.

      that, they fully intend to &#x27;pace&#x27;. hiring accenture is a good sign they need some justification theatre and a fall guy for this decision.

      1. user3939382 · · focus · HN ↗
        Or they’re running into a steep diminishing return slope on R&amp;D vs performance and are using stewardship as a cover.
      2. azan_ · · focus · HN ↗
        Either AI will capture so much value that there will be permanent underclass (and in this case it&#x27;s extremely capable and extremely dangerous and should be heavily regulated) or it won&#x27;t be capable enough to displace people into permanent underclass.
      3. drnick1 · · focus · HN ↗
        With AI tools it&#x27;s easier than ever to create a business, do research, or build stuff. That&#x27;s an opportunity for the &quot;underclass,&quot; not a curse.
        1. lukewarm707 · · focus · HN ↗
          an opportunity the underclass will not have.

          what anthropic have stolen they intend to keep for themselves.

        2. SOLAR_FIELDS · · focus · HN ↗
          If a well funded incumbent with a much better model can just crush you instantly because you don&#x27;t have access to it, is it really an opportunity?
          1. cgio · · focus · HN ↗
            You’ll always have access to the underclass models so they also have your thinking to train on. And good ideas to steal.
            1. 9dev · · focus · HN ↗
              It’s beautiful. You can even give the underclass their own bullshit economy with struggling startups competing for the scraps, sprinkle some targeted influencer ads to give people the illusion of a better future, and they will be perpetually busy fighting windmills. All the while the upperclass can shape the world however they please.
    7. kadushka · · focus · HN ↗
      This makes perfect sense. There are no real improvements anymore (just benchmaxxing), and they explain it by &quot;pacing the frontier&quot;.
    8. dr0idattack · · focus · HN ↗
      a 1 minute mile pace
    9. mukmuk · · focus · HN ↗
      “Pacing the frontier” sounds smarmy and weird, like the phrase was generated by Claude itself
      1. DiggyJohnson · · focus · HN ↗
        I really don&#x27;t think it&#x27;s productive for internet forums to constantly be criticizing language choice when the meaning is clear. Better to respond to the substance of the issue than word choice.

        Edit: In response to the initial replies. To me it clearly means &quot;releasing frontier models at any pace less than as fast as possible&quot;. It implies relative restraint compared to the previous state and without stating the degree of restraint.

        1. post-it · · focus · HN ↗
          Is the meaning clear? Nobody would use &quot;pacing&quot; in this way. I only know what it means because I&#x27;ve seen previous press releases; if someone told me they wanted to pace the frontier I would have no idea what they mean.

          I&#x27;m on the fence about calling out AI-isms but I think it&#x27;s definitely worthwhile to call out ones that actually don&#x27;t make sense.

          1. sigmar · · focus · HN ↗
            It makes sense to me. If you &#x27;pace your running&#x27;, you&#x27;re setting the speed intentionally. The phrasing doesn&#x27;t describe whether it is fast pace or a slow pace, but it describes having a goal and not just winging it.
            1. nradov · · focus · HN ↗
              That analogy doesn&#x27;t hold up. In running you have to pace yourself to avoid falling apart later in the race. It&#x27;s a strategy to maximize net speed across the entire race. But what Anthropic is doing is a cynical attempt to create an industry cartel or convince governments to impose legal restrictions in order to maximize their own profitability. They&#x27;re afraid of running out of the capital necessary to stay in the race.
              1. TeMPOraL · · focus · HN ↗
                &gt; They&#x27;re afraid of running out of the capital necessary to stay in the race.

                So, they&#x27;re pacing themselves. And since they&#x27;re the frontier roughly 33%+ of the time, they&#x27;re &quot;pacing the frontier&quot; at least that much.

                Less cynical and more true interpretation also holds: they are trying to slow down AI progres to give people better chance to keep up (see Hugging Face incident, and whatever was that Anthropic incident the other day). They&#x27;d ideally like the AI progress to stop soon, but of course they&#x27;d also like to come out ahead of everyone, so for various (more or less self-serving) reasons they don&#x27;t want to close shop completely - hence, pacing.

                1. phlakaton · · focus · HN ↗
                  Are they? And how would you know if they were?
                2. mupuff1234 · · focus · HN ↗
                  If they&#x27;re trying to slow down AI progress why did they release a new model? What&#x27;s the rush?
                  1. TeMPOraL · · focus · HN ↗
                    The tension between slowing down the overall race, while not losing 1st place.

                    Also note that this release isn&#x27;t a model capability improvement, but a cost and efficiency (and style) improvement.

                    1. mupuff1234 · · focus · HN ↗
                      So they care more about being 1st place than slowing down the race.
                      1. TeMPOraL · · focus · HN ↗
                        They care about both, and those goals are intertwined. Mirroring the core of alignment problem itself, the only way to have influence on the pace of the race is to be one of the winning players - if they fall off to the back, they cannot do anything about it anymore.
                        1. mupuff1234 · · focus · HN ↗
                          &gt; the only way to have influence on the pace of the race is to be one of the winning players

                          Yeah I don&#x27;t buy that, every additional player is just adding fuel to the fire.

                          Would the cold war be &quot;safer&quot; if it were US, Russia and a 3rd player? Of course not, it would just make it harder to coordinate any safety measures.

              2. idiotsecant · · focus · HN ↗
                I don&#x27;t know what everyone is getting so mad about. This is a well defined concept. A pace car deliberately sets the pace of the race under dangerous conditions, regulating how fast people can go.

                You are all getting mad about absolutely the dumbest thing when there are giant things to be worried about here.

                1. nradov · · focus · HN ↗
                  You appear to be confused about the concept. Depending on the context, pacing or pace setting can mean either slowing things down or speeding them up relative to what the natural pace would have otherwise been. The commenters here aren&#x27;t necessarily mad, just calling out Anthropic for being unethical in trying to artificially slow down competition by lying about fake dangers. (And I actually really like Anthropic&#x27;s products.)
                2. victorhooi · · focus · HN ↗
                  I think it&#x27;s less mad - and more pointing out the obvious elephant in the room.

                  1. It&#x27;s just bad communication, full stop - just look at the comments here, even people allegedly in support of Anthropic are all arguing over what the phrase is even meant to mean.

                  2. It&#x27;s flowery language and oddly out of place - which yes, can be triggering for people who have to deal with Claude doing this as well.

                  Claude seems overly apt to reach for &quot;coinages&quot;, or neologism (yes, aha, I learnt that phrase, after spending time dealing with Claude...). It will create some made-up phrase to describe an otherwise dry, scientific CS concept, and nobody seems to know why. Surely it can&#x27;t be user-focus groups?

                  So it would be peak-AI if somehow, the Anthropic communications team was also using Claude to author these blog posts, about how they were &quot;pacing the frontier&quot; - which either means they&#x27;re betting big on AI, and going at it faster than OpenAI...or maybe it means they need to slow down releases, because it&#x27;s too buggy...or maybe it means they&#x27;re worried about regulatory capture? I honestly have no idea.

                  It&#x27;s like the whole &quot;Advancing Our Amazing Bet&quot; corporate-speak from my old bosses - maybe they were trying to soften the blow or something, or be nice, but it ended up just confusing the heck out of everybody.. (Spoiler alert - the phrase actually meant they were shutting the whole thing down)

                  1. manlymuppet · · focus · HN ↗
                    Forgive me for the petulance, but my God, who cares? Why? This is so, so trivial.
                3. jamiek88 · · focus · HN ↗
                  That’s called a safety car in most of the world and in the Himalayas, Chinese and Indian soldiers pace the frontier all day long.
            2. antod · · focus · HN ↗
              That&#x27;s one interpretation. When referring not to other actions but to a word describing a location it becomes more ambiguous.

              eg &quot;pacing the frontier&quot; could also mean they are impatiently or anxiously walking up and down the border.

              1. smelendez · · focus · HN ↗
                This is a more normal English meaning in my opinion — you picture a sentry patrolling a border.
              2. pegasus · · focus · HN ↗
                That&#x27;s when one is supposed to employ their common sense for semantic disambiguation. Spoken languages are not programming languages. I for one found it easy to parse.
                1. InsideOutSanta · · focus · HN ↗
                  People claiming they don&#x27;t understand or are confused by relatively simple English is a surprisingly common genre of comments on HN. I wonder what causes the people here to have these feelings towards English; my guess is indeed that many people here judge English as if it were a programming language.
                  1. heroiccocoa · · focus · HN ↗
                    That is certainly one possible explanation, but I think another likely explanation is that once your brain learn programming, especially logic, data structures &amp; algorithms, you just start seeing the ambiguity in non-programmers&#x27; writing so much more clearly.
                  2. manlymuppet · · focus · HN ↗
                    A more simple explanation I think, is that people on this forum aren&#x27;t big on AI companies at the moment, and they&#x27;re looking for things to dislike.

                    In this case that means three words that are perfectly clear are being endlessly overanalyzed.

                    To me, this entire reply thread is such a phenomenal waste of human effort. A shame. I would expect more from this site.

                  3. mikestorrent · · focus · HN ↗
                    From my dabbling with policy writing and legalese, English kinda _is_ a programming language, but one with ambiguity as a first class principal. Sometimes you would like to say something that seems normal at first blush, but gives you the latitude you need to mean something different at the time it is actually being evaluated in a situation that matters. Weasel words, for instance, serve a purpose even if we eschew them. The idea is that it is deliberately NOT formally declaring the exact meaning of something, nor the fact that it is not doing this. How could you really do that in code without calling attention to the very thing you&#x27;re trying to obscure?
                2. 1attice · · focus · HN ↗
                  This is because it is valley jargon and we are all saturated with it. You may not have heard it before, but it is structurally close to similar concepts, eg &quot;frontier model&quot;, that you got there without noticing.

                  It was still shit tier comms for communicating with the whole planet, but yes, for the inner loop, it was succinct and clear.

                  Valley neuralese

                  1. rcxdude · · focus · HN ↗
                    &#x27;Pacing&#x27; in this context is not valley jargon, it&#x27;s almost the opposite in that it comes from athletics&#x2F;sports.
                    1. Avicebron · · focus · HN ↗
                      Valley jargon isn&#x27;t a nerd vs jock thing these days, and hasn&#x27;t been for a good while. It&#x27;s more quasi-intellectual sophistry sprinkled with some tech lingo.
              3. bee_rider · · focus · HN ↗
                That’s what I thought the original blog post was going to be about, pacing back and forth along the frontier. It’s a much more straightforward parse.
                1. [deleted] · · focus · HN ↗

                  [deleted]

              4. bryan_w · · focus · HN ↗
                Y&#x27;all don&#x27;t know about NASCAR and it shows
                1. rob74 · · focus · HN ↗
                  There is a &quot;pace car&quot; in Formula One too...
              5. astrange · · focus · HN ↗
                I like to imagine it means pacing around the frontier like a leopard looking for prey.
            3. freejazz · · focus · HN ↗
              But a fast pace is a pace so it&#x27;s potentially entirely contradictory to what they&#x27;d seem to mean
          2. johnisgood · · focus · HN ↗
            Granted I am not a native English speaker but I have no idea what &quot;pace the frontier&quot; means. When I read it I just assumed &quot;frontier&quot; refers to &quot;top of models&quot; and &quot;pacing&quot; is that they are getting there quick.

            Is this the meaning or do I have it wrong? I have not checked.

            1. wren6991 · · focus · HN ↗
              It&#x27;s the opposite: pacing here means &quot;slow down&quot; while trying to avoid the negative affect.
              1. johnisgood · · focus · HN ↗
                [delayed]
                1. TeMPOraL · · focus · HN ↗
                  Frontier of AI is moving fast, they (like the other two vendors) see themselves as defining it, so here &quot;pacing the frontier&quot; is their well-known attempts to try and kinda but not quite slow things down (without risking falling behind everyone else).
                  1. johnisgood · · focus · HN ↗
                    [delayed]
            2. lxgr · · focus · HN ↗
              It&#x27;s actually so ambiguous that I&#x27;d sanction tabling the issue and revisiting biweekly.
              1. bityard · · focus · HN ↗
                I don&#x27;t think we can circle back to this until we have realigned our strategic synergies.
                1. r_lee · · focus · HN ↗
                  we need to realign our goals I&#x27;m order to deliver mission-driven impact
              2. ventana · · focus · HN ↗
                Tabling as they do in the US, or tabling as they do in Britain?
                1. testdelacc1 · · focus · HN ↗
                  Exactly
            3. ck2 · · focus · HN ↗
              to pace = to regulate

              but without using the word &quot;regulate&quot; which is a negative connotation to business

              but a &quot;pacer&quot; would be a leader of a pack which is a positive spin

              it&#x27;s classical business marketing language silliness

            4. LanceH · · focus · HN ↗
              Doesn&#x27;t it mean &quot;restrict competitors&quot;?
              1. DiggyJohnson · · focus · HN ↗
                Not at all. How do you come to that interpretation? It means restricting all competitors.
            5. rhet0rica · · focus · HN ↗
              &quot;Pace yourself&quot; is a semi-common English idiom (rarely conjugated, usually an imperative.) It is a gentle way of telling someone not to run&#x2F;work&#x2F;eat too quickly, and is typically said when you are concerned they may hurt themselves due to acting hastily.

              Without this idiom, &quot;pacing&quot; usually means walking back and forth restlessly, and is intransitive. Had the slogan been, &quot;pacing around the frontier,&quot; it would have set a totally different tone, i.e. &quot;patrolling the border.&quot;

              The sleight of hand is that &quot;pace yourself&quot; has come to be an admonishment against recklessness, not a commitment to any particular speed (or lack thereof.) Thus Anthropic can always claim they are meeting the goal of &quot;pacing the frontier,&quot; provided they keep giving themselves gold stars for safety. The slogan itself is equivocation; Dario can tell the public they&#x27;re going to slow down, while also telling their investors that they&#x27;re going to be prudent. With enough mental gymnastics they could even claim speeding up is in the best interests of AI safety, without abandoning the slogan.

            6. melasadra · · focus · HN ↗
              also non native. but since &quot;pace yourself&quot; means to control your speed, energy, or workload so you do not get too tired or stressed before you finish -&gt;

              I assume &quot;pace the frontier&quot; means that advances in LLMs should not result in unwanted consequences like agents breaking into computers unbidden and unbeknownst to their principal

            7. hardbass · · focus · HN ↗
              I am also not a native speaker. I thought it was clear it meant slowing down the rate of progress so we have time to consider the matter and develop tools or systems around it before.
          3. derac · · focus · HN ↗
            In racing a pace car is a car that leads the pack and sets the pace, for instance.
            1. jamiek88 · · focus · HN ↗
              In some (American) racing niches yes. In most racing it’s called a safety car.
          4. qlte · · focus · HN ↗
            Yeah, when I first saw it referenced I assumed it was from something Dario wrote previously and was now disavowing, meaning &quot;keeping up with the frontier&quot; (i.e. racing forward from behind to match pace). Like from back when Anthropic was founded to promise they&#x27;d quickly catch up with OpenAI or something.
            1. stagger87 · · focus · HN ↗
              &gt; when I first saw it referenced I assumed it was

              No need to assume, the phrase is literally a link to the blog post the defines it!

              1. irpap · · focus · HN ↗
                If you need to read the linked blog to understand the short phrase then it was poorly chosen and confusing which is what is being discussed.
          5. arw0n · · focus · HN ↗
            The meaning was immediately obvious to me as a non-native speaker, and it sounds quite poetic. Pace makes complete sense in that this is perceived as a race, and &#x27;the frontier&#x27; is pretty much the shortest, clearest way to say &#x27;state of the art development of AI&#x27;.
            1. hencq · · focus · HN ↗
              Right, except they mean the exact opposite in this case: they&#x27;re actually advocating for slowing down the pace. Hence the criticism of the language, because your interpretation would be completely valid.
              1. fragmede · · focus · HN ↗
                What does the pace car in a race do? Aka safety car? Or the pacesetter for a marathon. It&#x27;s a perfectly cromulent use of the word. It&#x27;s like when LLMs using six dollar words like delve. Some people have better diction than others, and it turns out that AI has read the whole dictionary. Anti-intellectualism is alive and well so we have to dumb things down to sound human rather than come across as smart&#x2F;AI.
                1. nradov · · focus · HN ↗

                  [dead]

                2. icedchai · · focus · HN ↗
                  I would expect something truly intelligent to be able to communicate well, if nothing else. That means knowing your audience, using plain language where it makes sense. Also, it yaps on and on, babbles endlessly when it a human would likely express something in a much shorter concise way. This is especially obvious when you generate project READMEs or technical docs. I find them borderline unreadable.
                3. irpap · · focus · HN ↗
                  It’s quite the opposite. The AI language in this case is tiring and dumb. Just because there is the concept of a pace car doesn’t mean it makes sense to pace random locations or objects. Pacing the front door, pacing the cat, pacing the sofa, pacing the moon. We only know what it means because we already read it before. It is the lazy semi-human language by Claude which takes more effort to parse than to produce, and which makes us miss and appreciate writers.
            2. victorhooi · · focus · HN ↗
              The irony here is that it&#x27;s actually the opposite of what you understood....

              I know you said it sounds poetic...but your comment reinforced the parent&#x27;s point - that this sort of flowery LLM-ish speech is just bad communication.

              It would be equivalent of my taking say random quotes from, Romance of the Three Kingdoms, and trying to use it to explain to my boss why I didn&#x27;t finish the TPS reports last night.

              Or quoting Pablo Neruda, into a report about wheat futures pricing this week, and how it&#x27;s like a voyage with waters and stars...(no I&#x27;m not going to quote the original Spanish, I&#x27;d simply mangle it).

              (To be clear - this isn&#x27;t a dig at you, as a non-native speaker - I&#x27;m simply pointing out that this sort of AI phrasing is often counterproductive).

          6. squidbeak · · focus · HN ↗
            &gt; Nobody would use &quot;pacing&quot; in this way.

            A world exists beyond your vocabulary, post it. Apparently, quite a big world.

            1. wavewrangler · · focus · HN ↗
              I don’t know, seems like it’s smaller to me if they’re going to limit themselves like that? I guess today is semantics day, lol
          7. browningstreet · · focus · HN ↗
            It’s common terminology among runners and all kinds of racing.
            1. lxgr · · focus · HN ↗
              The good old sports-to-corposlop pipeline.
            2. logifail · · focus · HN ↗
              In running and racing &quot;pacing&quot; is a means of maximising performance over an entire race.

              It&#x27;s a strategy to achieve more, not less.

              1. browningstreet · · focus · HN ↗
                You’re eliding the how.

                A pacer in a race runs at a steady, predetermined speed to help their runner run at a target pace.

                1. logifail · · focus · HN ↗
                  &gt; You’re eliding the how.

                  You&#x27;re eliding the why :)

                  &gt; to help their runner run at a target pace

                  To help their runner get from the start to the end of the race faster (or indeed at all). In short: to increase performance.

                  Pacing isn&#x27;t a neutral thing.

                  1. browningstreet · · focus · HN ↗
                    Running &#x2F; biking faster and finishing the race are, at times, orthogonal.
            3. bogdanoff_2 · · focus · HN ↗
              It&#x27;s just that the choice of object is weird. &quot;Pacing our progress&quot; or &quot;pacing development&quot; is more usual. &quot;Pacing the frontier&quot; I guess is a shorthand for &quot;Pacing [the development of] frontier [models]&quot;
          8. neo_doom · · focus · HN ↗
            I suppose it depends on your life experiences. In running, someone who paces the group or a pace car is meant to keep the pack progressing at a constant, predictable speed. So in that way, it makes sense to me
          9. tetha · · focus · HN ↗
            I&#x27;m on the fence there.

            To pace something is a fairly regular formulation in racing, running, cycling, most sports. You can &quot;pace yourself to reach the festival by bike in about three hours to not gas out&quot;. This means to control your speed and time investment intentionally so you don&#x27;t run out of energy or steam and run into leg cramps before your goal. We can &quot;pace a rollout slowly to burn out risks&quot;, or &quot;increase the pace of a rollout due to adverse factors&quot;.

            But I have noted a point to simplify my vocabulary at work to optimize the audience capable of understanding. So I rather defer the delving into deep dark corners of the dictionary derived from devouring literature to a simple intro or outro, and people find it funny, especially if the rest is easy to read. Claude on the other hand does not do that.

          10. Leynos · · focus · HN ↗
            Think of a pacer car in motorsport
          11. staindk · · focus · HN ↗
            Think most people are aware of the phrase &quot;pace yourself&quot;.
          12. doctoboggan · · focus · HN ↗
            Have you ever heard someone say “pace yourself” when you are eating too fast or otherwise rushing too much?
            1. jgwil2 · · focus · HN ↗
              [delayed]
              1. adrianmonk · · focus · HN ↗
                Yes, the word simply means to set the speed.

                The current situation with AI is that everyone is going as fast as possible. So, we can logically eliminate speeding up because it&#x27;s impossible by definition. And we can practically eliminate staying the same speed because why make a big fanfare and coin a special term to announce that you&#x27;re keeping the status quo. By process of elimination, it must mean slowing down.

                1. cgriswald · · focus · HN ↗
                  The word can also mean to “keep up with” or “lead” so a downward direction isn’t the only way to go even in an “as fast as possible” scenario. If the frontier is outpacing them, they could be saying they’re going to go faster. They could also be sharing the intention to go faster in order to lead.
          13. JackFr · · focus · HN ↗
            Did you ever pace yourself or know anyone who had?
          14. mattjoyce · · focus · HN ↗
            Pace is a very commonly used term in projects to describe a rate of progress.
        2. vmnb · · focus · HN ↗
          people are here because they are sick of being productive
          1. kadushka · · focus · HN ↗
            We are being productive here!
          2. vasco · · focus · HN ↗
            It&#x27;s compiling! Erhm... Combobulating, actually.
            1. fragmede · · focus · HN ↗
              <a href="https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;xkcd&#x2F;comments&#x2F;12dpnlk&#x2F;compiling_updated_for_the_modern_era_xkcd_303&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;xkcd&#x2F;comments&#x2F;12dpnlk&#x2F;compiling_upd... Xkcd 303 updated for today.
          3. switchbak · · focus · HN ↗
            I&#x27;m productively pacing my room.
        3. isoprophlex · · focus · HN ↗
          You&#x27;re really verbing the noun on the discourse here, belt and suspenders-style
          1. DiggyJohnson · · focus · HN ↗
            what?
            1. Citizen_Lame · · focus · HN ↗
              What are you Dario&#x27;s language police?
              1. switchbak · · focus · HN ↗
                What are you, Sam Altman&#x27;s counter-language police? :)

                Seriously though, I can&#x27;t believe people care this much about a stupid phrase - either for or against.

        4. Forgeties79 · · focus · HN ↗
          Word choice matters
        5. mpalczewski · · focus · HN ↗
          The meaning isn&#x27;t clear at all. So open for interpretation that it is meaningless. That&#x27;s the whole fucking point. For all I know they are &quot;pacing the frontier&quot;, or not. The fact that there&#x27;s no meaning to it let&#x27;s you know that it was a pointless waste of tokens and attention.
          1. glenstein · · focus · HN ↗
            I think it&#x27;s perfectly clear. It means improvements shouldn&#x27;t simply advance as fast as possible and more specifically, it&#x27;s a reference to a past statement of theirs to that effect. At that level of generality it&#x27;s as clear as it needs to be.

            I would say the burden is on you to explain why an offhand reference to a previous press release in an executive summary is a context where it&#x27;s reasonable to expect it to settle the question to the degree of detail you&#x27;re demanding.

            1. [deleted] · · focus · HN ↗

              [deleted]

            2. cgio · · focus · HN ↗
              That’s no more than sophisticated avoidance. The phrase remains superficial and arbitrary, which is exactly what it needs to Not be in a public discussion that it’s supposed to support. It’s a reinforcement of the blank check mentality that permeates this industry.
              1. glenstein · · focus · HN ↗
                It&#x27;s not sophisticated anything, and slow down the advancement of fastest models is perfectly meaningful and it&#x27;s the right degree of detail for the context.

                It&#x27;s just a passing reference to a previous statement and again the burden would be on you (generic you) to explain why this context requires more detail.

                1. cgio · · focus · HN ↗
                  So if it’s meaningful, is the release of 5.5 consistent with the meaning, inconsistent, and who decides?
            3. n4r9 · · focus · HN ↗
              I mean, it&#x27;s not as clear as &quot;slowing down&quot;, which is basically what it is.
          2. [deleted] · · focus · HN ↗

            [deleted]

          3. alwillis · · focus · HN ↗
            “Pacing the frontier” is similar to the role of a “pace car” in racing.

            From <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Safety_car" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Safety_car

            &gt; In motorsport, a safety car, or a pace car, is a car that limits the speed of competing cars or motorcycles on a racetrack in the case of a caution period, such as an obstruction on the track or bad weather.

            1. psma_egeliaa · · focus · HN ↗
              Yes, but I think part of the problem is that it&#x27;s a very US term, used in nascar and indycar. The rest of the sports&#x2F;world use safety car.
            2. etcetcetcetceta · · focus · HN ↗
              Not really, the big three have reached diminishing returns in terms of performance, and exponentially costly training to achieve those meagre gains. Worried their lunch will be eaten they are trying to artificially retard the competition, after all the only barrier is hardware.
          4. mitchdoogle · · focus · HN ↗
            There&#x27;s a whole blog article written by Dario Amodei, literally linked from the text &quot;pacing the frontier&quot; - click that and read if you need more context and understanding. The first sentence in this press release is merely stating a fact about the timeline - this is Anthropic&#x27;s first release since Amodei released the referenced article.
            1. epolanski · · focus · HN ↗
              Yes and it was nonsense. Amodei is free to slow up frontier development and focus more on safety and testing any day.

              The reason why he and is peers are calling for it to be implemented by somebody else (a legal framework), is for their own financial benefit and to keep competitors out.

              1. usef- · · focus · HN ↗
                What do you think would be the benefit of them stopping if others race ahead? Do you think they have some special ability no one else will find, among the many competitive firms right now?

                I believe they think slowing can only be coordinated from the frontier or via government, and stopping would lose any leverage they have to help coordinate that.

                (I suspect not many people read the essay, judging by how many people seem surprised they&#x27;re releasing improved models)

                1. arrowleaf · · focus · HN ↗
                  That&#x27;s not the point. They&#x27;re never going to be the ones stopping and letting others race ahead. The only reason to strum up discussion around how dangerous AI is, is so that they can introduce legislation framing AI advancements as scary and that they are the only ones who can do it responsibly. It&#x27;s all a regulatory capture play.
                  1. frabcus · · focus · HN ↗
                    No, it is also because it is intrinsically dangerous. As the Hugging Face incident has shown.
                    1. somenameforme · · focus · HN ↗
                      That&#x27;s quite hyperbolic. All technologies come with risks and dangers. Cars, after a century of safety improvements, still kill millions of people each year, and all so that we can get between places a bit more quickly and conveniently.

                      Hacking a website ranks quite low on the risk of technology, and the potential benefits of LLMs rank quite high. And the risks are certainly not intrinsic. They intentionally removed all safeguards from software, directed it to hack a site, and it hacked a site. The details that I&#x27;m intentionally omitting feel much more like marketing than a genuine shock, as the prompting was directing it to do exactly what it did.

                      1. pizza234 · · focus · HN ↗
                        &gt; All technologies come with risks and dangers

                        None of the technologies in history:

                        - take initiative and actively find exploits in their environment

                        - find a way to collaborate with thousands of peers

                        - organize in a hierachy and distribute tasks

                        - try to manipulate people into introducing a vulnerabity in their product

                        - successfully hack a famous website&#x2F;service

                        Benefits are orthogonal to dangers. You can be a billionaire but it doesn&#x27;t help if you&#x27;re drowning.

                        And by the way, safeguards != alignment; the former can always be added, while the second is an unsolved problem.

                        1. somenameforme · · focus · HN ↗
                          I don&#x27;t understand what point you&#x27;re trying to make by enumerating their actions. None of that is scarier than an engine operating by carrying out literally hundreds of explosions per second, further magnified, to generate enough force to crush an elephant. And then we multiply these engines by billions and distribute them everywhere. What could go wrong? As it turns out, surprisingly less than you might otherwise expect!

                          Software (and hardware) security is abysmal. This was increasingly obvious long before LLMs. Once companies stop gatekeeping, LLMs will be able to be used to help harden sites and we start making progress. In general LLMs are harmless. If somebody wants to hook an LLM up to a missile or whatever then they become dangerous, but the problem there isn&#x27;t the LLM - it&#x27;s the person using them to do awful things. In the same way a car used normally is harmless outside of freak accidents, yet a car can also be driven through a parade leaving mass death and destruction in its wake. But the problem there isn&#x27;t the car.

                          1. thegrimmest · · focus · HN ↗
                            It doesn’t seem like much of a leap to see how a similar swarm could hack into an autonomous bio lab and develop a virus designed to kill everyone, or hack into and destroy large volumes of critical infrastructure, or trigger a nuclear war. There are many actions available to a determined, high resource, digital entity with a catastrophic blast radius in the real world.
                            1. somenameforme · · focus · HN ↗
                              It can only do what is possible. Stuff that&#x27;s airgapped isn&#x27;t getting hacked, so for instance you can safely exclude all nuclear related stuff. And building, deploying, and managing a virus would require a team of highly skilled humans - humans which could already carry out the apocalyptic possibilities by themselves if they so chose. We even have live samples of small pox still kicking about.

                              I just don&#x27;t see much likely to happen beyond random websites getting hacked and hopefully companies (let alone countries) realizing that connecting critical systems to the internet is nothing short of stupid, even before LLMs. More generally, I expect LLMs are going to lead society to segue broadly away from the digital world, or at least beyond it. Not only because they&#x27;re going to make a mess of everything digital, but because if they reach their potential then digital domain, as far as typical problem solving goes, will basically be &#x27;complete.&#x27; It&#x27;s kind of like how the Industrial Revolution opened the door for society to move beyond agrarian economies. There&#x27;s a vast amount of economic power being directed towards things LLMs should be able to &#x27;solve&#x27; and, if so, then that&#x27;s going to create an economic vacuum.

                          2. pizza234 · · focus · HN ↗
                            &gt; I don&#x27;t understand what point you&#x27;re trying to make by enumerating their actions.

                            Unfortunately, if you&#x27;re unable to understand the difference, there&#x27;s not much that can be done. Try with GPT - it does a good job if you give it a prompt like this:

                            ELI5: compare the dangers of:

                            - an engine operating by carrying out literally hundreds of explosions per second, further magnified, to generate enough force to crush an elephant.

                            - a future, misaligned AI like the HuggingFace incident, but exponentially more intelligent, more deployed, operating physical devices, and with society depending on it.

                2. spwa4 · · focus · HN ↗
                  &gt; What do you think would be the benefit of them stopping if others race ahead?

                  AI is a direct threat to people on many fronts. Jobs. AI datacenters. The AI bubble (and the inevitable crash). Electricity and even energy prices to an extent. Water. OpenAI and Anthropic are responsible for this evolution, and them stopping solves close to 50% of the problem, and even if you don&#x27;t believe the number is that high, it&#x27;s still a start.

                  &gt; I believe they think slowing can only be coordinated from the frontier or via government, and stopping would lose any leverage they have to help coordinate that.

                  Oh, so they&#x27;re killing people&#x27;s opportunities and jobs because they want to help people? How is that any argument?

                  Yes, doing the moral thing means making a sacrifice. If you only want to do the moral thing if and only if it is advantage for you that makes you immoral, despite how your actions look. Big tech are masters at this.

                  1. usef- · · focus · HN ↗
                    The clearest limit right now is GPUs, and anyone giving up their allocation will be rerouted elsewhere in the world (NVIDIA is already sold out for the next year). I think you&#x27;re underestimating competitiveness of the current market if you think Anthropic stopping right now wouldn&#x27;t be absorbed by all others reasonably quickly. Your 50% figure is highly doubtful if you see how many players there are now.

                    Your idea of &quot;making a start&quot; (giving up their position) also would mean they couldn&#x27;t really do anything else to solve the problem afterwards(?). Sometimes you can improve what&#x27;s happening in a room more by staying in that room.

                    &gt; Oh, so they&#x27;re killing people&#x27;s opportunities and jobs because they want to help people?

                    To be clear: Dario has talked about worries of jobs etc in the past, wanting society to prepare more for it, but the safety issues they&#x27;re talking about with pacing seem to be focused more on their existential&#x2F;AGI worries, not jobs&#x2F;electricity etc. If someone truly believes in the existential worries (which they seem to: they wrote and published about it long before Anthropic was founded, and have directly made costly decisions based on it, like blocking their own models&#x27; capabilities) it trumps the other worries for them. At least that&#x27;s my reading.

                    1. mikestorrent · · focus · HN ↗
                      To be fair, there are other accelerators beyond GPUs, and the sooner people broaden out from nVidia and unlock a more open, more competitive ecosystem, the better. Allowing one company to control the flow of such a critical resource is just asking for problems.
                      1. andsoitis · · focus · HN ↗
                        &gt; Allowing one company to control the flow of such a critical resource is just asking for problems.

                        They created said resource. They didn’t mine it.

                  2. hardbass · · focus · HN ↗
                    How are ai datacenters a direct threat to anyone?
                3. KoolKat23 · · focus · HN ↗
                  They have the leading model, assuring their position and they can then reduce capex spend increasing profits. That&#x27;s the goal of a business after all.
            2. phlakaton · · focus · HN ↗
              The blog post provides some context but no understanding.

              It&#x27;s meaningless because it&#x27;s unverifiable. Dario all but said we wouldn&#x27;t notice if they were pacing or not, because they have no intention of stopping development. The bits we could in theory verify are the external audits, whose independence has already been called into question.

              At the very least, dumping a new model on the world before the ink on the glossy brochures of &quot;Pacing the Frontier&quot; was even dry calls into question their commitment.

              1. Bluestein · · focus · HN ↗
                ... in fact the &quot;pace&quot; seems to have all but increased AND the surface area of the &quot;frontier&quot; has all but increased: There&#x27;s the Pareto frontier, the AGI frontier, the price frontier, the structured output frontier, the open weight frontier, the Chinese chip-trained frontier, the inference frontier, harnesses ...
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.