‹ BackHN Continuity

Thread

Getting the most out of Opus 5.5 in Claude and Claude Code

232 points · 156 comments · saikatsg

  1. rdli · · focus · HN ↗
    It’s a really good model. Over the past few days, I give Opus some general directives to basically speed up our CI, and telling it I care both about billing minutes and wall clock time. I told it to create a plan after analyzing everything in our CI, run the plan by a Fable subagent, and then focus on low-risk, high-reward changes.

    9 hours later, I had 12 PRs ready to be merged, and the net result is CI time has dropped from ~10 minutes to ~4 minutes, and billing minutes have dropped around 60%. Less than an hour of my attention.

    1. Waterluvian · · focus · HN ↗
      Every time weird stuff happened this week it was because the model choice in VSCode got set back to “auto” and some other model was trying its best. 5.5 is what I just set it to. Even Fable feels worse for my use cases.
      1. AnotherGoodName · · focus · HN ↗
        Even Anthropic rates Fable lower than 5.5 on pretty much all benchmarks.

        "Why does Fable even exist" is a very very reasonable question right now.

        1. wyre · · focus · HN ↗
          Because Anthropic releases their different model level's at very different points in time, they seem to always have one model that is by far an away the best to use for everything. Haiku 4.5 is almost a year old. Sonnet is fine, but idk if it has any real benefits over Opus. It's only been since Fable has been released that you get to choose between Fable and Opus, but not with 5.5 there is no reason to use Fable.

          I feel like instead of releasing fable, they should have released it as Opus 5, then their next Opus release they would call Sonnet, and their next Sonnet release they would have called Haiku. I don't know if their pricing structure would have been able to support that, but Anthropic has always been the least competitive regarding token pricing.

          1. comradesmith · · focus · HN ↗
            Fable came with a new set of api prices, and a special allowance for subscriptions.

            If they did what you suggested either they eat a ton of additional costs, or send a signal to the market that they’re increasing costs more generally.

            Also Fable and Opus have different specialties so they really are best presented as different models

        2. satvikpendem · · focus · HN ↗
          It's a class of model not a static one. There'll be Fable 5.5 that's even better than Opus 5.5.
          1. Wowfunhappy · · focus · HN ↗
            Although we never had a Claude Opus 5.1. Fable went 5 → 5.1 and Opus went 5 → 5.5.
            1. satvikpendem · · focus · HN ↗
              Opus 5.5 is probably a distilled version of a 6 model.
              1. Wowfunhappy · · focus · HN ↗
                If they had 6, don't you think they'd be serving it? Even if it was too expensive for most users, some big companies would likely be interested.
                1. satvikpendem · · focus · HN ↗
                  > If they had 6, don't you think they'd be serving it?

                  No I do not. OpenAI has some hidden model that's apparently 5x or so better at certain benchmarks than GPT 6 but they're not and have no plans to release it. It is increasingly likely that these AI companies keep the best models for themselves and then release smaller, cheaper distilled models for everyone else, especially since the AI companies are vertically integrating into many fields.

                2. K0balt · · focus · HN ↗
                  You’re assuming the goal of the big AI companies is to serve models for money. That is not the goal. They have a lot of incentives not to share their best models.
        3. adastra22 · · focus · HN ↗
          You are making the mistake of classifying models on a single linear axis, or even a multi axis basis set of all benchmarks. That just isn’t true. Each model is unique in its skills and capabilities and the way it approaches problems, in a way that is not represented in benchmarks. Fable is better at reviewing things. I don’t know how to explain it well but it is true. I trust Fable to do thorough reviews (sometimes too thorough) and to present its information in a dense but ordered way. Its output is equivalent to what you used to get from security firms doing code reviews. Having Opus do the work, and have fable do reviews (of the plan and implementation) is a good combo.
        4. sutterd · · focus · HN ↗
          I find fable better at high level planning, as in planning a task without filling in all the details. Opus doesn’t seem to be great at this, but other models aren’t either.
          1. BobbyJo · · focus · HN ↗
            I find the same. Fable is better at hashing out a plan with some back and forth and opus 5.5 is better (cheaper certainly) at sterile execution.
            1. cromka · · focus · HN ↗
              My exact experience as well.
        5. cromka · · focus · HN ↗
          I think of Fable as more knowledgeable, smart, erudite. Opus is a skilled technician.

          Benchmarks do not measure the first aspect.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.