3.8 Max is the most “grounded” model I think - talks generally normal, doesn’t go crazy and start doing things (I see you Gemini), has good design choices and isn’t overly nitpicky. But god it’s slow. And only available from Alibaba. Their token plan is stingy too. If I had to pick the “old reliable boring” LLM, a modern Claude 4.5 if you will, Qwen is my choice. Hopefully they don’t RL it to oblivion.
It’s absolutely down to their post-training RL, yeah. It’s where most of its strongest behaviour comes from, with regards to this kind of agentic behaviour
conception · · focus · HN ↗
rubslopes · · focus · HN ↗
What would that mean in this context?
pennomi · · focus · HN ↗
I swear I spend more time telling Claude not to do things than telling it what to do.
mdp2021 · · focus · HN ↗
But is that because of training, or can that be (also? mostly?) an effect of the "system prompt"?
girvo · · focus · HN ↗
vintermann · · focus · HN ↗
disgruntledphd2 · · focus · HN ↗
Personally, I think this is a bad idea, but someone's gotta build the Machine God I guess.