‹ BackHN Continuity

Thread

Contrastive Language Models

176 points · 59 comments · erichocean

  1. vessenes · · focus · HN ↗
    >"CLM-8B also sets a new SOTA on challenging agentic coding benchmarks, including DeepSWE (81.6%) and Terminal Bench 2.1 (87.6%).

    I didn&#x27;t see any details on this on the announce page. And I don&#x27;t believe it. Astra x-high pass@1 on DeepSWE is 74% +&#x2F;- 3%. (<a href="https:&#x2F;&#x2F;deepswe.datacurve.ai">https:&#x2F;&#x2F;deepswe.datacurve.ai).

    That said, love seeing some of these new architectures get people exploring. But, surely somebody is incorrect here inre: those numbers.

    1. kuukyo · · focus · HN ↗
      They don&#x27;t report the pass@1 success rate. They sample multiple solutions from Opus&#x2F;Fable and CLM decides which one to submit, that&#x27;s why they get &gt;80%.
      1. abeppu · · focus · HN ↗
        Maybe I&#x27;m misunderstanding this but when would you ever use it this way? If you&#x27;re already willing to call Opus&#x2F;Fable, then isn&#x27;t the obvious comparison whether Opus&#x2F;Fable can choose among sampled solutions better or worse than their fast model? If you&#x27;re willing to pay many seconds for many code samples from a slow model, it&#x27;s contrived to imagine you care about picking between them in ms.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.