‹ BackHN Continuity

Thread

Beating GPT-5.6 Sol on retrieval with 100x cheaper open models

414 points · 114 comments · moonikakiss

  1. jillesvangurp · · focus · HN ↗
    Models without tools and harnesses are not really that useful. My observation is that the tool ux is driving most progress at this point. There are of course open source tools and harnesses but they require more effort to setup properly.

    The key challenge is to pick the right model for the right task or sub task and doing that automatically rather than manually. A big part of the problem here is that everybody is picking the most expensive and resource intensive models by default just in case they hit something that is a bit more difficult to get right. It's overkill. Most work people actually do is completely routine and would not have been a challenge for most of the mainstream OSS models.

    I'm starting to suffer a bit from model fatigue. There are announcements almost on a daily basis about this or that new model. I can't keep up with that and I don't have time to try them out or evaluate them. I don't want to waste brain cycles on which one to use. I just want to get shit done without micromanaging AI models.

    All this marketing BS and confusing naming isn't helping either. It seems a lot of that is just about tricking people into picking the expensive model so they'll burn through more tokens.

    1. drob518 · · focus · HN ↗
      I agree with this. Interestingly, I’ve been using Deepseek v4 Flash a lot these days but I definitely have to constrain it a lot with tests. Fortunately, I can have it write the tests. It’s extremely cost effective. Still trying to figure out whether the latest update last week that made it smarter actually translates into something I can see in the output. It’s not dumber, but it’s still an open question as to whether it’s “real world smarter.”
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.