‹ BackHN Continuity

Thread

Strands Harness

150 points · 97 comments · zuckerborg0101

  1. johnmlussier · · focus · HN ↗
    I am increasingly hesitant to use non-native harnesses - model providers are now starting to train their agents for use within the harness. An eval like terminal bench can only capture so much data. I don't want to have to assess each harness every model release to make sure it's working as well as it can.
    1. avaer · · focus · HN ↗
      This matters less as models get better and everyone settles on the same overall harness architectures. The model matters more than the harness anyway.

      The bigger issue is that the use cases and harnesses for models is infinite, which is hard to compress into benchmark numbers that actually apply to you.

      Everyone is benchmaxxing, desperate to sell, and almost nobody except the labs is doing actual science on the results, so harnesses tend to be chosen on voodoo and hunches, like which company made it. There isn't necessarily a good alternative though, bearing the cost of being a harness researcher is probably not many people's goal.

      1. altcognito · · focus · HN ↗
        Agreed, especially since the more frontier models are able to accomplish in a vacuum, the more people will trust them. That being said, tool use is still really important for pulling in the right information.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.