I don't trust any of the benchmarks where Opus 5 surpasses Astra or Fable 5.1.Maybe Terminal Bench 4.0 and ExploitGym are reasonable.Terminal Bench 4.0 GPT 6 Astra 59.6 Claude Fable 5.1 55.1 Claude Opus 5 49.0 MiMo-V2.6-Pro 34.9 MiMo-V2.6-Flash 28.8 DeepSeek V4.1 Flash 26.8 MiMo-V2.5-Pro 1.5 ExploitGym GPT 6 Astra 42.4 Claude Fable 5.1 30.4 Claude Opus 5 22.1 MiMo-V2.6-Pro 17.8 MiMo-V2.6-Flash 6.0 MiMo-V2.5-Pro 0.1 DeepSWE v1.1 DeepSeek V4.1 Flash 74.2 Claude Opus 5 74.0 GPT 6 Astra 74.0 MiMo-V2.6-Pro 71.9 Claude Fable 5 70.0 MiMo-V2.6-Flash 67.9 MiMo-V2.5-Pro 19.0
Why not? In my own benchmark Opus 5 does in fact come out on top[1]1 - <a href="https://bench.killswitch-lang.org/" rel="nofollow">https://bench.killswitch-lang.org/
Yeah, for better or worse, writing style is practically uncorrelated with agentic performance, which is all the rage right now and the thing that most popular benchmarks currently prioritize.
user43928 · · focus · HN ↗
Maybe Terminal Bench 4.0 and ExploitGym are reasonable.
Terminal Bench 4.0
ExploitGym DeepSWE v1.1dom96 · · focus · HN ↗
1 - <a href="https://bench.killswitch-lang.org/" rel="nofollow">https://bench.killswitch-lang.org/
user43928 · · focus · HN ↗
cosmojg · · focus · HN ↗