‹ BackHN Continuity

Thread

Beating GPT-5.6 Sol on retrieval with 100x cheaper open models

414 points · 114 comments · moonikakiss

  1. aliljet · · focus · HN ↗
    There is a more serious question in here that's not being answered. How effective is the retrieval in finding buried needles in larger and larger haystacks. And there's a correlary question, how effective could you be in finding paired needles in that haystack where you need to hold a needle to unlock finding another needle.
    1. seahyinghang8 · · focus · HN ↗
      <founder of castform here> tldr: we generated synthetic training questions from the gitlab product handbook.

      totally agree that this larger corpus with harder to search information would be a good way to stress test - i'm sure we will encounter more interesting problems to solve. love to hear any suggestions of corpus to search against that is not just the public internet

      the training run link is also a little buried but here, you can see the comparison against the various models and their exact traces: <a href="https:&#x2F;&#x2F;app.castform.com&#x2F;train&#x2F;a7a898f6-d802-4908-b044-acb812f14a48?tab=comp" rel="nofollow">https:&#x2F;&#x2F;app.castform.com&#x2F;train&#x2F;a7a898f6-d802-4908-b044-acb81...

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.