‹ BackHN Continuity

Thread

Contrastive Language Models

176 points · 59 comments · erichocean

  1. vessenes · · focus · HN ↗
    >"CLM-8B also sets a new SOTA on challenging agentic coding benchmarks, including DeepSWE (81.6%) and Terminal Bench 2.1 (87.6%).

    I didn&#x27;t see any details on this on the announce page. And I don&#x27;t believe it. Astra x-high pass@1 on DeepSWE is 74% +&#x2F;- 3%. (<a href="https:&#x2F;&#x2F;deepswe.datacurve.ai">https:&#x2F;&#x2F;deepswe.datacurve.ai).

    That said, love seeing some of these new architectures get people exploring. But, surely somebody is incorrect here inre: those numbers.

    1. kuukyo · · focus · HN ↗
      They don&#x27;t report the pass@1 success rate. They sample multiple solutions from Opus&#x2F;Fable and CLM decides which one to submit, that&#x27;s why they get &gt;80%.
      1. vessenes · · focus · HN ↗
        Ah-ha. Interesting! Thanks.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.