‹ BackHN Continuity

Thread

Claude Opus 5.5

1806 points · 1134 comments · km144

  1. abtinf · · focus · HN ↗
    Astra is just so good. And the ChatGPT subscription lets me use my own harness, so I can hook it up to exe.dev.

    I’ll be going about my day, have a random idea, launch a microvm on exe.dev with a prompt of my idea, and get a working thing a few minutes later.

    I don’t know how much better a model would have to be to get me to move off OpenAI at this point, but doing just a little bit better in terminal bench 4 isn’t it. It would have to be a difference in kind, like opening up the harness restrictions, or privacy guarantees (comparable to offline models).

    Edit to address questions below:

    ChatGPT supports oauth login.

    Exe.dev has it built in. IIRC, pi also has it built in via /login.

    1. cbg0 · · focus · HN ↗
      How about cheaper? Astra is $10 in $50 out, Opus is $4 in $20 out. Even on a subscription you'll get considerably more usage out of Opus.
      1. qlte · · focus · HN ↗
        Luna is crazy cheap and surprisingly capable even at low reasoning levels, and outright competent at highest. I use 5.6 Sol at low reasoning a lot too on my Codex plan. There are capable and cost effective choices for models/reasoning levels at every price point in Codex world, including on the very low end where Haiku isn't competitive at all.
      2. copperx · · focus · HN ↗
        > Even on a subscription you'll get considerably more usage out of Opus.

        That's an incredibly bold assumption.

        1. cbg0 · · focus · HN ↗
          It's not, I have a subscription to both and Astra burns usage like crazy.
      3. qlte · · focus · HN ↗
        Per the link someone else posted, the actual difference in $/task is not nearly so stark:

        <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;releases&#x2F;claude-opus-5-5" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;releases&#x2F;claude-opus-5-...

          Opus 5.5 Medium = $1.34
          GPT-6-Astra High = $1.76
        And that assumes Opus 5.5 Medium is actually equivalent to Astra High in all real-world usage&#x2F;personal work loads, which isn&#x27;t guaranteed as benchmarks saturate. The High vs. High comparison (probably not equivalent, but for reference):

          Opus 5.5 High = $1.82
          GPT-6-Astra High = $1.76
        If Opus 5.5 Medium isn&#x27;t equal&#x2F;better for what you&#x27;re working on vs. Astra High across the board, the price difference would narrow a bit more each time you had to switch to High.

        So, if you&#x27;re happy with Codex already it&#x27;s not like Opus is now 1&#x2F;2 the price and you&#x27;d be leaving a crazy amount of money&#x2F;tokens on the table. Plus you have way more flexibility on the low end of the intelligence curve with GPT 5.6 Luna: Haiku (and Sonnet) can&#x27;t touch that price&#x2F;value ratio.

        1. TuxSH · · focus · HN ↗
          Also Opus 5.5 is often made useless due to its [cyber] guardrails (even worse than Astra), they&#x27;re even worse than Astra&#x27;s
        2. Gareth321 · · focus · HN ↗
          Once you get to roughly the 51+ Intelligence Index range, Opus 5.5 appears to define essentially the entire cost&#x2F;performance frontier, from ~$1.34&#x2F;task through ~$6&#x2F;task.

          Directly below this in the Cost per Intelligence Index Task table, the most efficient by far is Opus 5.5 Low.

      4. margorczynski · · focus · HN ↗
        In the end what matters is how much you pay for the task you want completed. And Astra will usually do that using less token and offer a better quality solution so in the end it might be cheaper.
        1. jasbury · · focus · HN ↗
          This. I was fed up today with constantly correcting Opus for a specific task. So I finally decided to try Astra. It handled all of my prompts in one go.
      5. notatoad · · focus · HN ↗
        yeah, Astra burned through 70% of my weekly usage in ~5hrs on a $100 plan. even fable doesn&#x27;t run out that quickly for me. it&#x27;s great, but it&#x27;s on the same tier as fable for me - use it sparingly, only when really necessary.
      6. abtinf · · focus · HN ↗
        A cheaper price has no value if I can’t use the thing I’m paying for.

        The Claude lock-in simply disqualifies anthropic entirely (for my use).

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.