Astra is just so good. And the ChatGPT subscription lets me use my own harness, so I can hook it up to exe.dev.
I’ll be going about my day, have a random idea, launch a microvm on exe.dev with a prompt of my idea, and get a working thing a few minutes later.
I don’t know how much better a model would have to be to get me to move off OpenAI at this point, but doing just a little bit better in terminal bench 4 isn’t it. It would have to be a difference in kind, like opening up the harness restrictions, or privacy guarantees (comparable to offline models).
Edit to address questions below:
ChatGPT supports oauth login.
Exe.dev has it built in. IIRC, pi also has it built in via /login.
Luna is crazy cheap and surprisingly capable even at low reasoning levels, and outright competent at highest. I use 5.6 Sol at low reasoning a lot too on my Codex plan. There are capable and cost effective choices for models/reasoning levels at every price point in Codex world, including on the very low end where Haiku isn't competitive at all.
And that assumes Opus 5.5 Medium is actually equivalent to Astra High in all real-world usage/personal work loads, which isn't guaranteed as benchmarks saturate. The High vs. High comparison (probably not equivalent, but for reference):
Opus 5.5 High = $1.82
GPT-6-Astra High = $1.76
If Opus 5.5 Medium isn't equal/better for what you're working on vs. Astra High across the board, the price difference would narrow a bit more each time you had to switch to High.
So, if you're happy with Codex already it's not like Opus is now 1/2 the price and you'd be leaving a crazy amount of money/tokens on the table. Plus you have way more flexibility on the low end of the intelligence curve with GPT 5.6 Luna: Haiku (and Sonnet) can't touch that price/value ratio.
Once you get to roughly the 51+ Intelligence Index range, Opus 5.5 appears to define essentially the entire cost/performance frontier, from ~$1.34/task through ~$6/task.
Directly below this in the Cost per Intelligence Index Task table, the most efficient by far is Opus 5.5 Low.
In the end what matters is how much you pay for the task you want completed. And Astra will usually do that using less token and offer a better quality solution so in the end it might be cheaper.
This. I was fed up today with constantly correcting Opus for a specific task. So I finally decided to try Astra. It handled all of my prompts in one go.
yeah, Astra burned through 70% of my weekly usage in ~5hrs on a $100 plan. even fable doesn't run out that quickly for me. it's great, but it's on the same tier as fable for me - use it sparingly, only when really necessary.
abtinf · · focus · HN ↗
I’ll be going about my day, have a random idea, launch a microvm on exe.dev with a prompt of my idea, and get a working thing a few minutes later.
I don’t know how much better a model would have to be to get me to move off OpenAI at this point, but doing just a little bit better in terminal bench 4 isn’t it. It would have to be a difference in kind, like opening up the harness restrictions, or privacy guarantees (comparable to offline models).
Edit to address questions below:
ChatGPT supports oauth login.
Exe.dev has it built in. IIRC, pi also has it built in via /login.
cbg0 · · focus · HN ↗
qlte · · focus · HN ↗
copperx · · focus · HN ↗
That's an incredibly bold assumption.
cbg0 · · focus · HN ↗
qlte · · focus · HN ↗
<a href="https://artificialanalysis.ai/models/releases/claude-opus-5-5" rel="nofollow">https://artificialanalysis.ai/models/releases/claude-opus-5-...
And that assumes Opus 5.5 Medium is actually equivalent to Astra High in all real-world usage/personal work loads, which isn't guaranteed as benchmarks saturate. The High vs. High comparison (probably not equivalent, but for reference): If Opus 5.5 Medium isn't equal/better for what you're working on vs. Astra High across the board, the price difference would narrow a bit more each time you had to switch to High.So, if you're happy with Codex already it's not like Opus is now 1/2 the price and you'd be leaving a crazy amount of money/tokens on the table. Plus you have way more flexibility on the low end of the intelligence curve with GPT 5.6 Luna: Haiku (and Sonnet) can't touch that price/value ratio.
TuxSH · · focus · HN ↗
Gareth321 · · focus · HN ↗
Directly below this in the Cost per Intelligence Index Task table, the most efficient by far is Opus 5.5 Low.
margorczynski · · focus · HN ↗
jasbury · · focus · HN ↗
notatoad · · focus · HN ↗
abtinf · · focus · HN ↗
The Claude lock-in simply disqualifies anthropic entirely (for my use).