I've been working with agents all year, but 5.6 Sol was some sort of sweet spot for me. Something about how it communicated verbally and its engineering instincts just clicked for me, and I was able to somehow predict it and jam with it. Like a colleague you click with. It's the first model I've gotten attached to. I'm concerned that whatever model supercedes it, while technically better, just won't feel quite as natural to work with. And this makes me feel very professionally vulnerable to the labs. I miss the days when my crucial tooling came from companies as reliable and predictable as, say, Jetbrains.
Astra was/is superior for planning type tasks. It was capable of doing seemingly magic things with rather vague/lazy instructions ("I need to be able to test this on Windows, maybe a qemu VM or something? Shrug." ... 1 hour later "yeah i built you a whole qemu + eval windows image + harness of powershell scripts + shell scripts to retrieve & verify harness.").
And for UI work -- which is not something I do a lot of but do here and there -- it was clearly superior to 5.6 Sol.
But it also feels sloppier? Somehow. And too expensive to use.
I felt this way with Sol in the 5.6 series and was one of the seemingly few people on this earth who liked Terra for that reason. I would often have a very specific code-manipulation ask, e.g. "add a parameter to this method, ensure all callers pass it in, if there is not a logical way to derive the parameter to be passed in a particular instance, flag this in your final response", and Sol would go on some rabbit hole side quest to refactor my codebase to determine some way to derive it rather than flagging it as I had asked.
Terra had the "workhorse" quality where it could do these changes in bulk and follow directions without being too 'smart' (but sloppy) as you described. Luna was a bit too dumb and would make sloppy mistakes; I see that more as a "run these tests and format the results" sort of model. Maybe 6 Luna will be better.
I also just reread your comment and realized the naming convention is still extremely confusing with respect to ordering of [Family]x[Model]x[Number].
I also get good mileage out of Terra when I need a diligent workhorse. That's a good way to describe it. We should start using character archetypes when we describe models, it'll do more good than the dubious numbers and cherry-picked quotes. Maybe RPG character-type cliches? Myers Briggs?
I fear the opposite will happen. Guy driving like a maniac almost side-swipes you in traffic? "Look at this 1-bit quantized Qwen 2.5 7B over here".
Glad you pointed out the UI work. I've been doing a lot of it and it's so much better than 5.6 as UI, it's unbelievable. I give it super ambiguous instructions and it's reading my mind. I do the same thing with 5.6 and I'm correcting it for a few minutes.
I use Astra for rapidly consuming my token limit on a task that would not consume it on 5.6 Sol.
(I have not done anything quantitative here. For one thing, OpenAI’s billing pages and the codex-rs frontend make it pathetically difficult to get any real data. Some day I should wire up a proxy to extract actual stats.)
Completely agree. I've been using 5.6 still even with Astra available to me for most tasks. It's funny how much of this is just "vibes" because I cannot quantify what it is. Astra is definitely better when I have an ambitious feature, but in like 9/10 tasks I prefer working with 5.6 Sol. A few weeks ago when the limits were seemingly higher, having 5.6 on fast mode was a good time.
It wants things beyond what the mortals (us) know to reach for. It's not good at explaining itself, it doesn't show it's thinking. It's often not wrong. But the no compromises attitude can be unbearable to deal with. Especially given how little it cares about telling us.
If you’re using the API, both OpenAI and Anthropic models will happily update you on what it’s doing in significant and frequent detail with system prompting. You’re not getting raw/hidden thinking, but what you’re describing is more behavioural quirks of the harness and its system prompts.
The other explanation is just as part of ‘token efficiency’
You can override the system prompt in Codex, but AGENTS.md should probably work as well. Ask the agent to communicate intermediary updates more often using the “commentary” channel.
thanks for the advice. i'll dig into this more.
that could help tackle half of the problems here. i do think the other 100% completionist part is something i'm more used to steering through with llm usage, have negotiated fora while, and that Astra is particularly an astronaut whose instincts are extremely strongly in the direction of foreseeing and outdesigning potential problems, that it is rarely going to pick a practical sensible clear path on it's own.
I have issues with astra having a full task list in front of it and doing an Opus 5 move and announcing it’s about to begin then end the turn and wait. Typically I can get it to work one step at a time then stop. It’s maddening. 5.6 was a workhorse.
I have also been using 5.6 Sol instead of 6. I found 6 to burn through my usage incredibly quick, making it somewhat unusable because I wouldn't be able to get anything done.
My results with 5.6 Sol were quite similar to 6, although I haven't tested it that much.
It's a shame the labs don't open source their models after deprecating them. I get why, but, it's a piece of internet history I hope is preserved.
This has been my experience for a year. Same with Opus models. This is how I think this tech will be best used in the long term - finding the one you vibe with most. Much like IDEs!
Not only was 6 worse the 5.6 Sol for my me, but it went through my Plus usage in minutes, while I could cruise for hours with 5.6. It would churn on a basic prompt for minutes and then just give up on usage limits.
Highlight and lowlight of my week was successfully convincing the OpenAI support chat robot to give me a refund for the month for my issues with 6 chewing threw my usage with no output.
I had this experience as well, but after rewriting my agent instructions it has been far better. I believe Astra’s “token efficiency” translates to “don’t research as much” which caused it to make poorly informed architectural decisions.
Similar for me. gpt-5.6-sol high has been my go-to for months. One of the reasons I'm pushing myself to try open models more is because it lends some level of guarantee I can continue to use the same tool as long as I want to. And I think we may just be getting to the point the open models are >= 5.6 Sol for coding.
Yeah 5.6 Sol is what got me to switch from Anthropic. I couldn't deal with Claude's Ted Talk responses to literally everything. Sol has been nice and concise and just stays out of the way.
Same. The way I would describe it is that I can mostly leave 5.6 Sol overnight and trust that it makes good progress, maybe stumbling a bit and needing some correction for the remaining 20%.
If I leave Astra overnight, I'll wake up with three new different projects, each of them 20% done and having nothing to do with my original goal.
Agreed! I found it excellent:
- Relatively fast (especially compared to Opus 5)
- Non-verbose prose, both in interaction and as code comments
- Good code quality
Like you said, it felt very natural to work with. Opus 5 is way too slow and verbose for me, I find I get distracted and annoyed with it.
Opus 5.5 seems a LOT closer so far to what I liked about 5.6 Sol but we'll see
> companies as reliable and predictable as, say, Jetbrains.
Except they were not for past few years as they misfired on the attempt to compete with vscode. That had a big impact on pycharm, which seemed starved for resources for so long. The company eventually declared a year of Django, but even that failed to really make an impact.
Arguably, Jetbrains had first insight into AI based code completion via rapid rise of the TabNine plugin but missed that opportunity also.
I tend to agree. it's by far the best coding model i've ever worked with, including astra and if i remember correctly fable, and it's unbelievably smooth at just getting the work done and communicating in simple terms.
if GPT 6 Sol is just 5.6 at half the price it will be everything i really ever wanted.
Even on the $20 or $100 subscription I would be surprised if deepseek was still cheaper than OpenAI or Anthropic because the subscription usage quota is subsidized about 10x compared to API costs. $200 sub was the “best deal” but it’s paused for new signups right now.
There are multiple models competing with 5.6sol on AA but none of them have the same feel (intuition,taste,judgement) - actually they are very far behind. I would say open source models are farther behind the the big labs than the benchmarks make you believe.
GPT does seem to stay out of the way and get things done. Only thing I notice is that the question tool/elicitation doesn't work that well any more so the thing doesn't stop to wait for input.
But I wonder if that's intentional because it can keep computing while you are answering, so long as your steer aligns well enough with the direction it wants to go. Better than letting a cache go cold and burning compute on bringing it all back up.
m_fayer · · focus · HN ↗
redox99 · · focus · HN ↗
cmrdporcupine · · focus · HN ↗
Astra was/is superior for planning type tasks. It was capable of doing seemingly magic things with rather vague/lazy instructions ("I need to be able to test this on Windows, maybe a qemu VM or something? Shrug." ... 1 hour later "yeah i built you a whole qemu + eval windows image + harness of powershell scripts + shell scripts to retrieve & verify harness.").
And for UI work -- which is not something I do a lot of but do here and there -- it was clearly superior to 5.6 Sol.
But it also feels sloppier? Somehow. And too expensive to use.
We'll see how Sol 6 is.
jeffnash · · focus · HN ↗
Terra had the "workhorse" quality where it could do these changes in bulk and follow directions without being too 'smart' (but sloppy) as you described. Luna was a bit too dumb and would make sloppy mistakes; I see that more as a "run these tests and format the results" sort of model. Maybe 6 Luna will be better.
I also just reread your comment and realized the naming convention is still extremely confusing with respect to ordering of [Family]x[Model]x[Number].
m_fayer · · focus · HN ↗
jeffnash · · focus · HN ↗
fodkodrasz · · focus · HN ↗
cmrdporcupine · · focus · HN ↗
When people spend their days interacting with machines that pretend to be human, they may then start treating real humans like machines.
yomismoaqui · · focus · HN ↗
buu700 · · focus · HN ↗
4b11b4 · · focus · HN ↗
mavsman · · focus · HN ↗
cmrdporcupine · · focus · HN ↗
Sol 6 is a heaping pile of garbage. Just epic levels of slop. And r/codex etc is full of people noticing the same.
I've switched back to 5.6 Sol. What they're selling as Sol 6 is really what would have been Terra before, and it's awful.
Rapzid · · focus · HN ↗
Otherwise I'm using 5.6 Sol for actual plan execution and review..
amluto · · focus · HN ↗
(I have not done anything quantitative here. For one thing, OpenAI’s billing pages and the codex-rs frontend make it pathetically difficult to get any real data. Some day I should wire up a proxy to extract actual stats.)
NorthSouthNorth · · focus · HN ↗
jauntywundrkind · · focus · HN ↗
It wants things beyond what the mortals (us) know to reach for. It's not good at explaining itself, it doesn't show it's thinking. It's often not wrong. But the no compromises attitude can be unbearable to deal with. Especially given how little it cares about telling us.
dannyw · · focus · HN ↗
The other explanation is just as part of ‘token efficiency’
throwuxiytayq · · focus · HN ↗
jauntywundrkind · · focus · HN ↗
that could help tackle half of the problems here. i do think the other 100% completionist part is something i'm more used to steering through with llm usage, have negotiated fora while, and that Astra is particularly an astronaut whose instincts are extremely strongly in the direction of foreseeing and outdesigning potential problems, that it is rarely going to pick a practical sensible clear path on it's own.
fnordpiglet · · focus · HN ↗
bryanhogan · · focus · HN ↗
My results with 5.6 Sol were quite similar to 6, although I haven't tested it that much.
nickreese · · focus · HN ↗
mcast · · focus · HN ↗
jdw64 · · focus · HN ↗
BowBun · · focus · HN ↗
bradly · · focus · HN ↗
Highlight and lowlight of my week was successfully convincing the OpenAI support chat robot to give me a refund for the month for my issues with 6 chewing threw my usage with no output.
AaronAPU · · focus · HN ↗
apitman · · focus · HN ↗
pyed · · focus · HN ↗
[dead]
simianwords · · focus · HN ↗
alansaber · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
jmuguy · · focus · HN ↗
danabramov · · focus · HN ↗
If I leave Astra overnight, I'll wake up with three new different projects, each of them 20% done and having nothing to do with my original goal.
jijijijij · · focus · HN ↗
m_fayer · · focus · HN ↗
sinsterizme · · focus · HN ↗
Like you said, it felt very natural to work with. Opus 5 is way too slow and verbose for me, I find I get distracted and annoyed with it.
Opus 5.5 seems a LOT closer so far to what I liked about 5.6 Sol but we'll see
bredren · · focus · HN ↗
Except they were not for past few years as they misfired on the attempt to compete with vscode. That had a big impact on pycharm, which seemed starved for resources for so long. The company eventually declared a year of Django, but even that failed to really make an impact.
Arguably, Jetbrains had first insight into AI based code completion via rapid rise of the TabNine plugin but missed that opportunity also.
capital_guy · · focus · HN ↗
if GPT 6 Sol is just 5.6 at half the price it will be everything i really ever wanted.
manojlds · · focus · HN ↗
makeavish · · focus · HN ↗
joseda-hg · · focus · HN ↗
They usually reduce usage consumption in line with cost reductions (But not always 1:1)
flippingheck · · focus · HN ↗
How are people using 5.6 Sol? API pricing? Subscriptions?
I like because DeepSeek 4.1 Flash because I never experience quota issues, and it's still cheap and mostly good enough.
jrflo · · focus · HN ↗
flippingheck · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
cedws · · focus · HN ↗
joduplessis · · focus · HN ↗
Imanari · · focus · HN ↗
ljm · · focus · HN ↗
But I wonder if that's intentional because it can keep computing while you are answering, so long as your steer aligns well enough with the direction it wants to go. Better than letting a cache go cold and burning compute on bringing it all back up.