We're past the one model fits them all kind of LLM. Most of the recent release actually regress on world knowledge for example, but optimize for something different: tool usage, thinking process, and agentic approach. And yes, in my own usage, some usecases Opus beats Fable.
user43928 · · focus · HN ↗
Maybe Terminal Bench 4.0 and ExploitGym are reasonable.
Terminal Bench 4.0
ExploitGym DeepSWE v1.13abiton · · focus · HN ↗