I don't trust any of the benchmarks where Opus 5 surpasses Astra or Fable 5.1.Maybe Terminal Bench 4.0 and ExploitGym are reasonable.Terminal Bench 4.0 GPT 6 Astra 59.6 Claude Fable 5.1 55.1 Claude Opus 5 49.0 MiMo-V2.6-Pro 34.9 MiMo-V2.6-Flash 28.8 DeepSeek V4.1 Flash 26.8 MiMo-V2.5-Pro 1.5 ExploitGym GPT 6 Astra 42.4 Claude Fable 5.1 30.4 Claude Opus 5 22.1 MiMo-V2.6-Pro 17.8 MiMo-V2.6-Flash 6.0 MiMo-V2.5-Pro 0.1 DeepSWE v1.1 DeepSeek V4.1 Flash 74.2 Claude Opus 5 74.0 GPT 6 Astra 74.0 MiMo-V2.6-Pro 71.9 Claude Fable 5 70.0 MiMo-V2.6-Flash 67.9 MiMo-V2.5-Pro 19.0
can you recommend any benchmark websites that show up-to-date details like this?TerminaBench, DeepSwe sites are out of date.
Yeah it’s a shame a lot of these benchmarks are behind. My favourite was ‘SlopCodeBench’ [1] as I’m most interested in ai reinforcing its own bad decisions, but it’s not even up to current gen oai1: <a href="https://www.scbench.ai/" rel="nofollow">https://www.scbench.ai/
user43928 · · focus · HN ↗
Maybe Terminal Bench 4.0 and ExploitGym are reasonable.
Terminal Bench 4.0
ExploitGym DeepSWE v1.1novaleaf · · focus · HN ↗
TerminaBench, DeepSwe sites are out of date.
UnfitFootprint · · focus · HN ↗
1: <a href="https://www.scbench.ai/" rel="nofollow">https://www.scbench.ai/