‹ BackHN Continuity

Thread

GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price

1066 points · 953 comments · crorella

  1. minimaxir · · focus · HN ↗
    > Cached input costs just $0.10 per million tokens—95% less than standard input pricing and 50% less than GPT‑6 Sol’s cached input pricing

    This is the actual big announcement. 50% cheaper cache than GPT-6 Sol will get you far more mileage on Codex.

    1. joshstrange · · focus · HN ↗
      > 50% cheaper cache than GPT-6 Sol will get you far more mileage on Codex.

      Cache doesn't help you much when you are compacting every 5 minutes...

      I was shocked at how quickly I ran out my $100/mo subscription with a single agent (sol medium).

      1. onlyrealcuzzo · · focus · HN ↗
        If you're compacting every 5 minutes, you have a workflow problem - period.

        No LLM will be cost effective if it's compacting this often. You have to find a way around it.

        1. ngruhn · · focus · HN ↗
          Context window is only 275k or something. And honestly compaction is not that bad in Codex. I often don't even notice I went through 5 compactions in a session.
          1. jeremyjh · · focus · HN ↗
            I don’t usually have a problem doing a complete task in that context size. OMP does make a lot of use of rewind which may be helping - basically forks itself and sends back a summary after a long tangent. Coding takes use a Luna max agent.

            I’ve also found compaction not to be a problem when it does happen.

            1. threecheese · · focus · HN ↗
              How do you trigger this? I've been messing with OMP lately for funsies.
              1. jeremyjh · · focus · HN ↗
                Its in settings under Tools->Checkpoint/Rewind. I don't know why its not enabled by default.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.