‹ BackHN Continuity

Thread

Strands Harness

150 points · 97 comments · zuckerborg0101

  1. seizethecheese · · focus · HN ↗
    > With Fable 5, Strands harness cost 77% less than Claude Code and scored higher on Terminal Bench 2.1.

    Terminal Bench 2.1 is saturated. Many token saving techniques would save money and score basically the same running Fable 5 against Terminal Bench 2.1. (They claim a better score but don’t say how much better. I’d bet my favorite hat that it’s not statistically significant.)

    This is at least the fourth time I’ve seen a project hit front page with a “save money with same score on saturated benchmark” claim.

    1. strandstan · · focus · HN ↗
      The scores are in blog post's bar chart. For Terminal Bench 2.1, Strands harness (Fable 5) scored 69.7 while Claude Code (Fable 5) scored 61.8. This is on high effort.

      I hear you tho about saturation. We're working on a follow-up deep dive post with more harnesses, so could look into Terminal Bench 4.0?

      1. seizethecheese · · focus · HN ↗
        Then I&#x27;m really confused. Terminal Bench 2.1 scores on Artificial Analysis are like 80-90%. <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;evaluations&#x2F;terminalbench-2-1" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;evaluations&#x2F;terminalbench-2-1
        1. stefan_lec · · focus · HN ↗
          AA isn’t testing different harnesses there, they’re testing different models on the same common harness:

          “We run Terminal-Bench 2.1 with the Terminus 2 agent harness in an e2b sandbox and report pass@1 averaged over 3 repeats per task.”

          This is a different agent harness than those Strands was testing with, and there’s no guarantee the pass criteria matches up the same either.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.