I've been working with agents all year, but 5.6 Sol was some sort of sweet spot for me. Something about how it communicated verbally and its engineering instincts just clicked for me, and I was able to somehow predict it and jam with it. Like a colleague you click with. It's the first model I've gotten attached to. I'm concerned that whatever model supercedes it, while technically better, just won't feel quite as natural to work with. And this makes me feel very professionally vulnerable to the labs. I miss the days when my crucial tooling came from companies as reliable and predictable as, say, Jetbrains.
Completely agree. I've been using 5.6 still even with Astra available to me for most tasks. It's funny how much of this is just "vibes" because I cannot quantify what it is. Astra is definitely better when I have an ambitious feature, but in like 9/10 tasks I prefer working with 5.6 Sol. A few weeks ago when the limits were seemingly higher, having 5.6 on fast mode was a good time.
It wants things beyond what the mortals (us) know to reach for. It's not good at explaining itself, it doesn't show it's thinking. It's often not wrong. But the no compromises attitude can be unbearable to deal with. Especially given how little it cares about telling us.
If you’re using the API, both OpenAI and Anthropic models will happily update you on what it’s doing in significant and frequent detail with system prompting. You’re not getting raw/hidden thinking, but what you’re describing is more behavioural quirks of the harness and its system prompts.
The other explanation is just as part of ‘token efficiency’
You can override the system prompt in Codex, but AGENTS.md should probably work as well. Ask the agent to communicate intermediary updates more often using the “commentary” channel.
thanks for the advice. i'll dig into this more.
that could help tackle half of the problems here. i do think the other 100% completionist part is something i'm more used to steering through with llm usage, have negotiated fora while, and that Astra is particularly an astronaut whose instincts are extremely strongly in the direction of foreseeing and outdesigning potential problems, that it is rarely going to pick a practical sensible clear path on it's own.
I have issues with astra having a full task list in front of it and doing an Opus 5 move and announcing it’s about to begin then end the turn and wait. Typically I can get it to work one step at a time then stop. It’s maddening. 5.6 was a workhorse.
I have also been using 5.6 Sol instead of 6. I found 6 to burn through my usage incredibly quick, making it somewhat unusable because I wouldn't be able to get anything done.
My results with 5.6 Sol were quite similar to 6, although I haven't tested it that much.
m_fayer · · focus · HN ↗
NorthSouthNorth · · focus · HN ↗
jauntywundrkind · · focus · HN ↗
It wants things beyond what the mortals (us) know to reach for. It's not good at explaining itself, it doesn't show it's thinking. It's often not wrong. But the no compromises attitude can be unbearable to deal with. Especially given how little it cares about telling us.
dannyw · · focus · HN ↗
The other explanation is just as part of ‘token efficiency’
throwuxiytayq · · focus · HN ↗
jauntywundrkind · · focus · HN ↗
that could help tackle half of the problems here. i do think the other 100% completionist part is something i'm more used to steering through with llm usage, have negotiated fora while, and that Astra is particularly an astronaut whose instincts are extremely strongly in the direction of foreseeing and outdesigning potential problems, that it is rarely going to pick a practical sensible clear path on it's own.
fnordpiglet · · focus · HN ↗
bryanhogan · · focus · HN ↗
My results with 5.6 Sol were quite similar to 6, although I haven't tested it that much.