DeepSWE is saturated now IMO, and is basically worthless. Lots of new models get around 74%. Shame too, because it was a pretty decent benchmark for a few months there.
> and google just started letting all their engineers use claude
That's misleading.
1. Having different models available is useful for A/B testing and helping improve Gemini itself.
2. They have an enterprise offering for Antigravity (their agentic coding platform), and they need to test that it works well with non-Gemini models too.
Also worth taking a look at is the mimo harness. It's a fork of opencode with some new modes added for long horizon tasks. One of the better open harnesses out there at the moment.
ricardobeat · · focus · HN ↗
Fable scores 70%, Kimi K3 69%, Astra 74% (all on max effort).
<a href="https://deepswe.datacurve.ai/blog/deepswe-v1-1">https://deepswe.datacurve.ai/blog/deepswe-v1-1
Cookingboy · · focus · HN ↗
Even flash reached 60.7% by step 12, and it's on step 16 now.
This is so exciting lmao.
arcanemachiner · · focus · HN ↗
brookst · · focus · HN ↗
markasoftware · · focus · HN ↗
ehsankia · · focus · HN ↗
That's misleading.
1. Having different models available is useful for A/B testing and helping improve Gemini itself.
2. They have an enterprise offering for Antigravity (their agentic coding platform), and they need to test that it works well with non-Gemini models too.
buffalobuffalo · · focus · HN ↗