‹ BackHN Continuity

Thread

GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price

1066 points · 953 comments · crorella

  1. minimaxir · · focus · HN ↗
    > Cached input costs just $0.10 per million tokens—95% less than standard input pricing and 50% less than GPT‑6 Sol’s cached input pricing

    This is the actual big announcement. 50% cheaper cache than GPT-6 Sol will get you far more mileage on Codex.

    1. joshstrange · · focus · HN ↗
      > 50% cheaper cache than GPT-6 Sol will get you far more mileage on Codex.

      Cache doesn't help you much when you are compacting every 5 minutes...

      I was shocked at how quickly I ran out my $100/mo subscription with a single agent (sol medium).

      1. onlyrealcuzzo · · focus · HN ↗
        If you're compacting every 5 minutes, you have a workflow problem - period.

        No LLM will be cost effective if it's compacting this often. You have to find a way around it.

        1. ngruhn · · focus · HN ↗
          Context window is only 275k or something. And honestly compaction is not that bad in Codex. I often don't even notice I went through 5 compactions in a session.
          1. sally_glance · · focus · HN ↗
            Same for me, I started wondering if maybe workflows using compaction instead of clear + markdown memory would be more efficient. Writing a plan or tasks to a file often has the next session repeat part of the exploration, compaction seems to keep most relevant context.
          2. SyneRyder · · focus · HN ↗
            Sounds like that's the problem then, 275k is a tiny context window. I regularly have sessions that go to 450k or even up to 700k for an unattended overnight Claude Opus session.

            Apparently OpenAI makes you manually setup their 1 Million context window, and it seems to be only documented on X:

            <a href="https:&#x2F;&#x2F;x.com&#x2F;thsottiaux&#x2F;status&#x2F;2089082893804896524" rel="nofollow">https:&#x2F;&#x2F;x.com&#x2F;thsottiaux&#x2F;status&#x2F;2089082893804896524

            There&#x27;s at least a forum thread about it here:

            <a href="https:&#x2F;&#x2F;community.openai.com&#x2F;t&#x2F;why-does-codex-report-a-258-400-token-context-window-for-gpt-5-6-sol&#x2F;1394346&#x2F;5" rel="nofollow">https:&#x2F;&#x2F;community.openai.com&#x2F;t&#x2F;why-does-codex-report-a-258-4...

            1. gf000 · · focus · HN ↗
              But that 250k context worth way more than 1M in terms of how well it&#x27;s utilized, so actually I do like codex trying to keep you at that sweet spot.
            2. rrvsh · · focus · HN ↗
              Its really not tiny; you can&#x27;t compare Claude to GPT, they have honestly diverged enough that as the other reply said, 256k GPT is about equal to 1M Claude. The compaction is slightly annoying, and you can turn it up to 1M as you said if you truly need everything in context, but otherwise it&#x27;s perfectly serviceable
              1. kaoD · · focus · HN ↗
                [delayed]
            3. bjord · · focus · HN ↗
              &gt; unattended overnight claude opus session

              yes, exactly

              1. SyneRyder · · focus · HN ↗
                Not sure I understand if this was meant as a slight against Claude? Or agreement?

                These are often my best sessions - they&#x27;re unattended overnight, because by then we have the specification figured out, and I can just leave Claude to build out the rest, making good choices if it does find gaps in the spec. I regularly go to sleep &amp; wake up to an entirely new application completed. Claude never uses compacting in my sessions.

                I haven&#x27;t used GPT as much as I should have, so I&#x27;m prepared to be incorrect &amp; out of date. It just intuitively feels like I wouldn&#x27;t get the same from a 275K context window - maybe it uses lots of subagents? Even Deepseek &amp; GLM have 1 Million context windows now, so it &quot;feels&quot; strange for people to actually prefer the 275K window. But that&#x27;s just my intuition.

                1. bjord · · focus · HN ↗
                  neither, actually, just that unattended &quot;oneshot&quot; sessions are incredibly token inefficient

                  if you talk about them (in which you lean on an LLM as a sort-of independent employee) and conservative, chunk-based usage (in which you use the LLM as more of an extension of yourself), you&#x27;re comparing apples to oranges

                  a predefined spec obviously reduces that gap but how much is highly dependent on the level of detail

          3. jeremyjh · · focus · HN ↗
            I don’t usually have a problem doing a complete task in that context size. OMP does make a lot of use of rewind which may be helping - basically forks itself and sends back a summary after a long tangent. Coding takes use a Luna max agent.

            I’ve also found compaction not to be a problem when it does happen.

            1. threecheese · · focus · HN ↗
              How do you trigger this? I&#x27;ve been messing with OMP lately for funsies.
              1. jeremyjh · · focus · HN ↗
                Its in settings under Tools-&gt;Checkpoint&#x2F;Rewind. I don&#x27;t know why its not enabled by default.
          4. onlyrealcuzzo · · focus · HN ↗
            If it&#x27;s compacting every 5 mins, you&#x27;re going to notice it in your cache miss ratio and your costs...
            1. rrvsh · · focus · HN ↗
              It doesn&#x27;t - try it first
          5. Benjamin_Dobell · · focus · HN ↗
            The context window is configurable. I&#x27;ve been using ~600k for months. No, not API pricing, on a Codex sub.

            ~&#x2F;.codex&#x2F;config.toml

              model = &quot;gpt-6.1-sol&quot;
              model_context_window = 700000
              model_auto_compact_token_limit = 630000
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.