‹ BackHN Continuity

Thread

Claude Opus 5.5

1806 points · 1134 comments · km144

  1. abtinf · · focus · HN ↗
    Astra is just so good. And the ChatGPT subscription lets me use my own harness, so I can hook it up to exe.dev.

    I’ll be going about my day, have a random idea, launch a microvm on exe.dev with a prompt of my idea, and get a working thing a few minutes later.

    I don’t know how much better a model would have to be to get me to move off OpenAI at this point, but doing just a little bit better in terminal bench 4 isn’t it. It would have to be a difference in kind, like opening up the harness restrictions, or privacy guarantees (comparable to offline models).

    Edit to address questions below:

    ChatGPT supports oauth login.

    Exe.dev has it built in. IIRC, pi also has it built in via /login.

    1. felixgallo · · focus · HN ↗
      If you read the page, Opus is now significantly better than Astra while also being cheaper and having more performance headroom available.
      1. mattz56 · · focus · HN ↗
        Better than Astra looks insane. I saw a Higgsfield video yesterday where they gave the same prompt to Astra and Opus 5.5 to create a samurai video game, and the difference was huge.
    2. onlyrealcuzzo · · focus · HN ↗
      > And the ChatGPT subscription lets me use my own harness, so I can hook it up to exe.dev.

      This is news to me. Excited to try it out! Thanks.

    3. polalavik · · focus · HN ↗
      ya i've been a gpt hater for a while. almost exclusively used claude up until astra. astra feels like it blows everything out of the water. its fast, correct, organized, and less verbose.
    4. roughly · · focus · HN ↗
      > And the ChatGPT subscription lets me use my own harness, so I can hook it up to exe.dev.

      Can you give more details here? This sounds intriguing.

    5. [deleted] · · focus · HN ↗

      [deleted]

    6. cbg0 · · focus · HN ↗
      How about cheaper? Astra is $10 in $50 out, Opus is $4 in $20 out. Even on a subscription you'll get considerably more usage out of Opus.
      1. qlte · · focus · HN ↗
        Luna is crazy cheap and surprisingly capable even at low reasoning levels, and outright competent at highest. I use 5.6 Sol at low reasoning a lot too on my Codex plan. There are capable and cost effective choices for models/reasoning levels at every price point in Codex world, including on the very low end where Haiku isn't competitive at all.
      2. copperx · · focus · HN ↗
        > Even on a subscription you'll get considerably more usage out of Opus.

        That's an incredibly bold assumption.

        1. cbg0 · · focus · HN ↗
          It's not, I have a subscription to both and Astra burns usage like crazy.
      3. qlte · · focus · HN ↗
        Per the link someone else posted, the actual difference in $/task is not nearly so stark:

        <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;releases&#x2F;claude-opus-5-5" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;releases&#x2F;claude-opus-5-...

          Opus 5.5 Medium = $1.34
          GPT-6-Astra High = $1.76
        And that assumes Opus 5.5 Medium is actually equivalent to Astra High in all real-world usage&#x2F;personal work loads, which isn&#x27;t guaranteed as benchmarks saturate. The High vs. High comparison (probably not equivalent, but for reference):

          Opus 5.5 High = $1.82
          GPT-6-Astra High = $1.76
        If Opus 5.5 Medium isn&#x27;t equal&#x2F;better for what you&#x27;re working on vs. Astra High across the board, the price difference would narrow a bit more each time you had to switch to High.

        So, if you&#x27;re happy with Codex already it&#x27;s not like Opus is now 1&#x2F;2 the price and you&#x27;d be leaving a crazy amount of money&#x2F;tokens on the table. Plus you have way more flexibility on the low end of the intelligence curve with GPT 5.6 Luna: Haiku (and Sonnet) can&#x27;t touch that price&#x2F;value ratio.

        1. TuxSH · · focus · HN ↗
          Also Opus 5.5 is often made useless due to its [cyber] guardrails (even worse than Astra), they&#x27;re even worse than Astra&#x27;s
        2. Gareth321 · · focus · HN ↗
          Once you get to roughly the 51+ Intelligence Index range, Opus 5.5 appears to define essentially the entire cost&#x2F;performance frontier, from ~$1.34&#x2F;task through ~$6&#x2F;task.

          Directly below this in the Cost per Intelligence Index Task table, the most efficient by far is Opus 5.5 Low.

      4. margorczynski · · focus · HN ↗
        In the end what matters is how much you pay for the task you want completed. And Astra will usually do that using less token and offer a better quality solution so in the end it might be cheaper.
        1. jasbury · · focus · HN ↗
          This. I was fed up today with constantly correcting Opus for a specific task. So I finally decided to try Astra. It handled all of my prompts in one go.
      5. notatoad · · focus · HN ↗
        yeah, Astra burned through 70% of my weekly usage in ~5hrs on a $100 plan. even fable doesn&#x27;t run out that quickly for me. it&#x27;s great, but it&#x27;s on the same tier as fable for me - use it sparingly, only when really necessary.
      6. abtinf · · focus · HN ↗
        A cheaper price has no value if I can’t use the thing I’m paying for.

        The Claude lock-in simply disqualifies anthropic entirely (for my use).

    7. mlcruz · · focus · HN ↗
      What worked well for me was a custom version of Open Web Ui with some customization to spawn an exe.dev instance for each new chat. I can just work on my phone, deploy stuff for development purposes on an easy to share way etc.
    8. nonethewiser · · focus · HN ↗
      What sort of things are you building? What do you like about exe.dev? Just curious what teh overhead and $20&#x2F;month subscription is enabling for you.

      &gt;I’ll be going about my day, have a random idea, launch a microvm on exe.dev with a prompt of my idea, and get a working thing a few minutes later.

      This is my experience with Claude code on my local machine. I suppose maybe you are doing something that naturally has system side effects? Obviously sandboxes have advantages sometimes but I havent seen a need for what I&#x27;m building.

      1. abtinf · · focus · HN ↗
        For me, the essential difference between exe and local harnesses is that they have really nice built in methods to take care of stuff like auth and inter-vm interactions, which makes it easy to build live internet-connected services that I can share with other people. There is no deploy step, which is a surprising amount of time savings. They also have nice things like email receive&#x2F;send, the ability to issue phone notifications through their app, and just a ton of little things where it feels right.

        FWIW, the $20&#x2F;month subscription also includes $20&#x2F;month of LLM credits. That’s obviously not sustainable, but it should make it easier to try out the service. I would stick with them even if they dropped it.

        Here is an invite link for a 30 day trial (that benefits me too if you were to become a paying member):

        <a href="https:&#x2F;&#x2F;exe.dev&#x2F;i&#x2F;rlDF6GI5PGBZV4P" rel="nofollow">https:&#x2F;&#x2F;exe.dev&#x2F;i&#x2F;rlDF6GI5PGBZV4P

        Or

        ssh rlDF6GI5PGBZV4P@exe.dev

        1. solenoid0937 · · focus · HN ↗
          The astroturfing and guerilla marketing on HN is getting insane.
          1. abtinf · · focus · HN ↗
            I have no association with exe.dev, other than being a delighted customer.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.