‹ BackHN Continuity

Thread

Beating GPT-5.6 Sol on retrieval with 100x cheaper open models

414 points · 114 comments · moonikakiss

  1. JCharante · · focus · HN ↗
    I have done my own testing and found that smaller models can beat their larger siblings on fact retrieval from documents. I haven’t investigated it in depth with a large enough dataset but my guess is that larger models overthink it while smaller ones just do it. I would like if they compared this with 5.6 Luna instead.
    1. barake · · focus · HN ↗
      Anecdotally, it feels like Opus, Fable, and Sol "get distracted" when you use them for writing code. Great at reasoning and coordination but they will go off on a tangent and refactor half the code base. I only use them for reasoning (of course) and coordinating subagents.
      1. CoolCold · · focus · HN ↗
        mind sharing hints/links on your harness/flow setup?

        I did several attempts with naive prompting, but spent more time babysitting than actual flow

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.