‹ BackHN Continuity

Thread

GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price

1066 points · 953 comments · crorella

  1. minimaxir · · focus · HN ↗
    > Cached input costs just $0.10 per million tokens—95% less than standard input pricing and 50% less than GPT‑6 Sol’s cached input pricing

    This is the actual big announcement. 50% cheaper cache than GPT-6 Sol will get you far more mileage on Codex.

    1. TuxSH · · focus · HN ↗
      Exactly half as expensive as Opus 5.5 in every API pricing metric
      1. dom96 · · focus · HN ↗
        Based on my benchmark[1] it is the same price as Opus 5.5 and just as capable.

        1 - <a href="https:&#x2F;&#x2F;bench.killswitch-lang.org" rel="nofollow">https:&#x2F;&#x2F;bench.killswitch-lang.org

        1. zeroonetwothree · · focus · HN ↗
          Opus 5 scoring higher than 5.5 makes me question of the value of this benchmark to real world usage
          1. dom96 · · focus · HN ↗
            Well, it is genuine.

            Opus 5.5 fails the &quot;understanding&quot; tasks which Opus 5 passes. I feed it a script which takes two numbers and prints the max of the two numbers. Opus 5.5 thinks it prints 1&#x2F;0 instead of the max numbers. Opus 5 gets it right.

            Here are the outputs from both: <a href="https:&#x2F;&#x2F;gist.github.com&#x2F;dom96&#x2F;b5bce82b6e6c1ebd5271ed70ad941b49" rel="nofollow">https:&#x2F;&#x2F;gist.github.com&#x2F;dom96&#x2F;b5bce82b6e6c1ebd5271ed70ad941b....

            Looking at that Opus 5.5 fails to deduce that the &quot;hack statement&quot; is actually an if statement in disguise, but Opus 5 gets this right. I feel like this is a pretty good test and shows Opus 5&#x27;s greater intelligence for what it&#x27;s worth.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.