GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
Unofficial Hacker News client; not affiliated with Y Combinator.
the_duke · · focus · HN ↗
Sol 6 was so bad that I switched over to Opus 5.5 exclusively.
Huge regression compared to Sol 5.6, often doing really dumb things. Same for Luna.
Even Astra is very unreliable for coding. Brilliant for vision, sometimes just great, but it also often does very stupid things.
I'm a bit sour on OpenAI right now and skeptical that 6.1 will be much different.
(Note: this is after preferring and shilling Codex/OpenAI models for the last half year)
jstummbillig · · focus · HN ↗
I mean Opus 5.5 is absolutely fantastic, unreasonably and unexpectedly so, but Astra was great and as far as I can tell SOTA until, when was it, 3 days ago, no?
(Sol 6 idk, have not used it much for coding really. Seemed to work just fine when Astra used it in Codex as subagents.)
the_duke · · focus · HN ↗
phoghed · · focus · HN ↗
Codex itself seems to have a regression. You can see clearly the token use changing significantly coincides with a score drop
nicce · · focus · HN ↗
copperx · · focus · HN ↗
Marha01 · · focus · HN ↗
Eridrus · · focus · HN ↗
Astra seems better though.
Showing one potentially saturated benchmark doesn't necessarily fill me with a lot of confidence in the coding results.
phoghed · · focus · HN ↗
Since like last December I haven’t had any issues getting work done with whatever the latest Anthropic or OpenAI models at the time were. Tooling and models have only gotten better since then.