‹ BackHN Continuity

Thread

Beating GPT-5.6 Sol on retrieval with 100x cheaper open models

415 points · 114 comments · moonikakiss

  1. breadislove · · focus · HN ↗
    On what do you guys test the model. Its very dubious that there is no common retrieval benchmark such as browsecomp plus or similar tested. And what metric do you report?
    1. krm01 · · focus · HN ↗
      Keeping track of any AI progress is becoming harder by the day, because there's ambiguity around common/clear/consistent benchmarks. Everything is constantly skewed into favourable directions.
      1. alansaber · · focus · HN ↗
        TBF gaming benchmarks is not something new to AI
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.