I don't trust any of the benchmarks where Opus 5 surpasses Astra or Fable 5.1.Maybe Terminal Bench 4.0 and ExploitGym are reasonable.Terminal Bench 4.0 GPT 6 Astra 59.6 Claude Fable 5.1 55.1 Claude Opus 5 49.0 MiMo-V2.6-Pro 34.9 MiMo-V2.6-Flash 28.8 DeepSeek V4.1 Flash 26.8 MiMo-V2.5-Pro 1.5 ExploitGym GPT 6 Astra 42.4 Claude Fable 5.1 30.4 Claude Opus 5 22.1 MiMo-V2.6-Pro 17.8 MiMo-V2.6-Flash 6.0 MiMo-V2.5-Pro 0.1 DeepSWE v1.1 DeepSeek V4.1 Flash 74.2 Claude Opus 5 74.0 GPT 6 Astra 74.0 MiMo-V2.6-Pro 71.9 Claude Fable 5 70.0 MiMo-V2.6-Flash 67.9 MiMo-V2.5-Pro 19.0
can you recommend any benchmark websites that show up-to-date details like this?TerminaBench, DeepSwe sites are out of date.
<a href="https://artificialanalysis.ai/evaluations/terminalbench-4-0" rel="nofollow">https://artificialanalysis.ai/evaluations/terminalbench-4-0
user43928 · · focus · HN ↗
Maybe Terminal Bench 4.0 and ExploitGym are reasonable.
Terminal Bench 4.0
ExploitGym DeepSWE v1.1novaleaf · · focus · HN ↗
TerminaBench, DeepSwe sites are out of date.
1899-12-30 · · focus · HN ↗