‹ BackHN Continuity

Thread

Beating GPT-5.6 Sol on retrieval with 100x cheaper open models

414 points · 114 comments · moonikakiss

  1. breadislove · · focus · HN ↗
    On what do you guys test the model. Its very dubious that there is no common retrieval benchmark such as browsecomp plus or similar tested. And what metric do you report?
    1. seahyinghang8 · · focus · HN ↗
      (founder of castform here) - we didn&#x27;t get to dive too deep into the dataset we were using for the retrieval in the blogpost for brevity, but we did link the training run (which shows the dataset) here: <a href="https:&#x2F;&#x2F;app.castform.com&#x2F;train&#x2F;a7a898f6-d802-4908-b044-acb812f14a48?tab=comp" rel="nofollow">https:&#x2F;&#x2F;app.castform.com&#x2F;train&#x2F;a7a898f6-d802-4908-b044-acb81...

      the page shows the exact trace of all the models we are comparing against and the aggregate scores

      we generated the question &amp; answer pair from gitlab product handbook (<a href="https:&#x2F;&#x2F;handbook.gitlab.com&#x2F;" rel="nofollow">https:&#x2F;&#x2F;handbook.gitlab.com&#x2F;) since the point is to show that you can generate training questions from raw data corpus (something a company already has today)

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.