‹ BackHN Continuity

Thread

Beating GPT-5.6 Sol on retrieval with 100x cheaper open models

427 points · 117 comments · moonikakiss

  1. mrinterweb · · focus · HN ↗
    There is so much opportunity for purpose built models like this. Ideally a harness should spin up a subagent to offload to targeted models for specific tasks like this. I know this is not a novel idea. Claude code does some of this by handing off the "explore" agent work to haiku. I just love seeing that specialized LLMs are being developed.
    1. nikcub · · focus · HN ↗
      There has been an over-obsession with frontier models and benchmarks. Most of the work will be done by task specific models. You don't put Phds on the factory floor.
      1. bizzletk · · focus · HN ↗
        If the PhDs don't cost very much more than your equivalent of factory-technicians but still get the job done, why wouldn't you do that, at least in the blunt case before cost control rears up?
        1. polotics · · focus · HN ↗
          because they will take initiatives for localized improvements you don't want them to take
      2. dwaltrip · · focus · HN ↗
        This only works if the tasks are actually specific and don’t benefit from broad competency.

        IMO, this doesn’t match most things that people use LLMs for.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.