‹ BackHN Continuity

Thread

Beating GPT-5.6 Sol on retrieval with 100x cheaper open models

427 points · 117 comments · moonikakiss

  1. JCharante · · focus · HN ↗
    I have done my own testing and found that smaller models can beat their larger siblings on fact retrieval from documents. I haven’t investigated it in depth with a large enough dataset but my guess is that larger models overthink it while smaller ones just do it. I would like if they compared this with 5.6 Luna instead.
    1. seahyinghang8 · · focus · HN ↗
      we actually have the test benchmark against luna too! it's just not in our title but you can see it in the first diagram below the title. luna does pretty well tbh but sol is just a tad bit better. but luna is way cheaper.

      if you want to dive down into the various traces of the benchmark, you can check this out: <a href="https:&#x2F;&#x2F;app.castform.com&#x2F;train&#x2F;a7a898f6-d802-4908-b044-acb812f14a48?tab=comp" rel="nofollow">https:&#x2F;&#x2F;app.castform.com&#x2F;train&#x2F;a7a898f6-d802-4908-b044-acb81...

      - founder of castform

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.