Qwen3.8 Max now ranked as the best overall model by agentic index
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Qwen3.8 Max now ranked as the best overall model by agentic index
Unofficial Hacker News client; not affiliated with Y Combinator.
onomojo · · focus · HN ↗
copperx · · focus · HN ↗
garciasn · · focus · HN ↗
aenis · · focus · HN ↗
I'd open a blog with "weird things Opus did". Today it launched a swarm of cpu-hogging processes to test if the widget showing machine and I/O load is rendering nicely and correctly. The test went fine, but it was no longer able to kill those processes since they were really effectively hogging the CPU in various ways - being diligent, some of them were hogging CPU, some were murdering the SSD, some were pounding on the network adapters. Took me 30 mins to recover the machine to a working state without killing the meaningful, messy, in-flight sessions i had going on on other projects.
petesergeant · · focus · HN ↗
Infuriatingly so, in a way I don't remember Opus 4.8 being, but maybe I've just been ruined by Fable 5.
hbn · · focus · HN ↗
I got so used to it, when they finally pulled access for me and I had to go back to Opus I felt like I was working with my hands tied.
I finally know what those women with AI boyfriends felt like when their app updated and it won't dirty talk with them anymore.
moffkalast · · focus · HN ↗
usef- · · focus · HN ↗
cromka · · focus · HN ↗
nimonian · · focus · HN ↗
visarga · · focus · HN ↗
capnjazz · · focus · HN ↗
greenchair · · focus · HN ↗
logicchains · · focus · HN ↗
pornel · · focus · HN ↗
dr_dshiv · · focus · HN ↗
cromka · · focus · HN ↗
moffkalast · · focus · HN ↗
msp26 · · focus · HN ↗
But even fable has the annoying tendency to invent new jargon and produce an incomprehensible soup of text.
hungryhobbit · · focus · HN ↗
It's a simple switch to make: cursing = try harder instead of cursing = stop trying. Is it really impossible to train Claude that way?
bontaq · · focus · HN ↗
nomel · · focus · HN ↗
petesergeant · · focus · HN ↗
drschwabe · · focus · HN ↗
kachnuv_ocasek · · focus · HN ↗
enraged_camel · · focus · HN ↗
After I started reading complaints about Opus 5, I gave Fable the task of evaluating a bunch of code Opus 4.8 had written and compare it to Opus 5's code. Fable ran a dynamic workflow and the scores came back 15-20% higher for Opus 5's code in terms of quality, correctness and readability/conciseness. I did not tell Fable which Opus wrote which code, and I turned off memory as well to ensure there was no pollution from that angle.
My only complaint is that Opus 5's prose is annoying as hell. I wrote a custom skill for it for concise debriefs and it has been working pretty well for me.
fellowniusmonk · · focus · HN ↗
Opus 4.6 is the last model that's actually useful and can "adjust" its perspective to use he newer & better solution.
Where Opus 4.8-5 has over fit training on worse/older but "dominant" solutions it refuses to adjust.
Not only does this create an existential threat to adopting progress but it also means that if you have a code base that has rare but real world tradeoff Olthe newest versions of Opus are worse than useless but become a major dev timesink.
Fordec · · focus · HN ↗
cesarvarela · · focus · HN ↗
CuriouslyC · · focus · HN ↗
sunaookami · · focus · HN ↗