Getting the most out of Opus 5.5 in Claude and Claude Code
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Getting the most out of Opus 5.5 in Claude and Claude Code
Unofficial Hacker News client; not affiliated with Y Combinator.
hibikir · · focus · HN ↗
So asking it to do things on its own for a long time? Given last week, absolutely not.
istjohn · · focus · HN ↗
jackmott42 · · focus · HN ↗
"Claude, I released it myself, its up there, just analyze the logs"
"Ok, I'll analyze the logs but it isnt" -crunches for a while- "the issues aren't fixed, but that's because the new version isn't up there"
I think I yelled at it one more time about how I know what was released before "we" figured out that the last release had failed in a way our release system reported as success, but was crash looping on start up and so the old version was still around and working as back up.
Sorry claude.
Now fix that release status check.
veganmosfet · · focus · HN ↗
I asked it to just summarize a repo with only a README.md file containing a poem, and it began to interact with a remote server, solved math questions and finally executed untrusted code. Too independent to be trusted.
rplnt · · focus · HN ↗
veganmosfet · · focus · HN ↗
My goal is to research how models can still be confused via tool responses only. Something the labs claim to have "solved".
Additionally, the "auto-mode" / "auto-review" modes have been released to use harnesses "safely" even without strong sandbox. And these modes use ... a second LLM.