‹ BackHN Continuity

Thread

GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price

1066 points · 953 comments · crorella

  1. minimaxir · · focus · HN ↗
    > Cached input costs just $0.10 per million tokens—95% less than standard input pricing and 50% less than GPT‑6 Sol’s cached input pricing

    This is the actual big announcement. 50% cheaper cache than GPT-6 Sol will get you far more mileage on Codex.

    1. joshstrange · · focus · HN ↗
      > 50% cheaper cache than GPT-6 Sol will get you far more mileage on Codex.

      Cache doesn't help you much when you are compacting every 5 minutes...

      I was shocked at how quickly I ran out my $100/mo subscription with a single agent (sol medium).

      1. codewithcheese · · focus · HN ↗
        you can config codex to compact at a higher context limit
      2. redox99 · · focus · HN ↗
        If you run out of sol medium with $100 you're doing something wrong. Astra destroys your usage, I get 1 day of usage with Astra, but 6 sol is almost unlimited and I only use xhigh.
        1. shimman · · focus · HN ↗
          "You're holding it wrong." Is hardly a retort from a real paying customer having problems with their paid services.

          This is why these companies are struggling to make money, they're chastising their customers just like they've been chastising the human race.

          1. trio8453 · · focus · HN ↗
            > "You're holding it wrong." Is hardly a retort from a real paying customer having problems with their paid services.

            It's very appropriate in the cases when you're holding it wrong. The fact that you're paying doesn't mean that you can't make mistakes or waste resources.

            1. shimman · · focus · HN ↗
              I don't find it appropriate at all, especially regarding technology that workers deeply hate and are skeptical of.

              If this is how you want to get people on your side, I can understand why the entire country/human race are against these companies.

              1. trio8453 · · focus · HN ↗
                Sides? Hate? This is all very emotional. Try to put the facts down plainly and see how ridiculous it is --

                It's a product and if you're using it incorrectly, we can either

                1. say so

                2. pretend that you don't so to get/keep you on "our side"? or not say is because you're skeptical or hate it? (how does that last bit even follow logically?)

                How is 2 better in any way for anyone involved?

            2. crossroadsguy · · focus · HN ↗
              [delayed]
          2. Anonasty · · focus · HN ↗
            Literally the prompting and task definition is main variable how LLM's performs. There are literally millions of examples of vibe coders and new AI adopters who run out of tokens since they don't know how the LLM's work.
        2. jorblumesea · · focus · HN ↗
          yeah I use sol constantly and have done maybe $15 of spend in the past week. it's solid and cheaper. this is at least 4-5 investigations, prs, whatever per day.
        3. Aeolun · · focus · HN ↗
          It’s only nearly unlimited if you haven’t just used a banked reset. After a banked reset your weekly usage gets cut by about 80% (not the week you need to wait to get your normal limits back though). ChatGPT has given me a really good reason to cancel.
          1. threecheese · · focus · HN ↗
            Can you elaborate? I've been getting great usage out of my $200/mo plan, and thought I'd try a reset (first time) which was expiring just for giggles. Am I going to get only 20% of it effectively?

            I overused Astra in order to drain my weekly, figuring I'd have the reset. (not wastefully, I did get more work done)

            1. Aeolun · · focus · HN ↗
              I can’t say what will happen to you, but yes, that has been my experience. It is better to wait for your normal full limit to return, because if you use a banked reset you get only 1/5th of the tokens but you still have to wait the full week afterwards for it to reset. 20% would be fine if it didn’t also reset the date your normal reset fires.
              1. seunosewa · · focus · HN ↗
                Banked resets do expire if you don't use them, so use them anyway.
              2. nkmnz · · focus · HN ↗
                Did you “earn” that reset on a lower tier?
              3. edot · · focus · HN ↗
                Proof? Like, do you have logs or something? Not calling you a liar but this seems not correct based on my usage.
        4. mattkenefick · · focus · HN ↗
          How do you get 1 day of usage with Astra?

          I create a lot, but I can make a full month with Astra on the current Pro plan. What are you doing to spend that much?

          1. redox99 · · focus · HN ↗
            Currently spending a lot of tokens programming the AI for my videogame.

            1 day is kind of generous, it probably lasts like 12 hours of running non stop.

      3. apitman · · focus · HN ↗
        You have a lot of control over compaction, both directly by changing compaction settings, and indirectly by how you structure your codebase/docs so agents use less tokens.
      4. antonvs · · focus · HN ↗
        Try Gemini. It’s so cheap I often use my personal AI Pro account for corporate work, and most of the time it doesn’t matter.
        1. ChickeNES · · focus · HN ↗
          Gemini is dumb as hell though, it's not like for like
          1. Foobar8568 · · focus · HN ↗
            cheerleader hallucinating agent. That's Gemini.
          2. Marha01 · · focus · HN ↗
            Gemini 3.8 Flash is actually pretty good.
          3. antonvs · · focus · HN ↗
            I doubt you've tried it recently, or perhaps you confused the search engine version of Gemini for the frontier models.

            I've been using Gemini on development of a DNN training pipeline, and there's no way you can describe it as "dumb as hell". That description just makes it clear that you're not talking about the technical capabilities of the models, but about some sort of fanboy comparison from a parallel hype universe.

      5. _davide_ · · focus · HN ↗
        As a reference i burn 1% percent for every 40 minutes of sol on average
      6. onlyrealcuzzo · · focus · HN ↗
        If you're compacting every 5 minutes, you have a workflow problem - period.

        No LLM will be cost effective if it's compacting this often. You have to find a way around it.

        1. ngruhn · · focus · HN ↗
          Context window is only 275k or something. And honestly compaction is not that bad in Codex. I often don't even notice I went through 5 compactions in a session.
          1. sally_glance · · focus · HN ↗
            Same for me, I started wondering if maybe workflows using compaction instead of clear + markdown memory would be more efficient. Writing a plan or tasks to a file often has the next session repeat part of the exploration, compaction seems to keep most relevant context.
          2. SyneRyder · · focus · HN ↗
            Sounds like that's the problem then, 275k is a tiny context window. I regularly have sessions that go to 450k or even up to 700k for an unattended overnight Claude Opus session.

            Apparently OpenAI makes you manually setup their 1 Million context window, and it seems to be only documented on X:

            <a href="https:&#x2F;&#x2F;x.com&#x2F;thsottiaux&#x2F;status&#x2F;2089082893804896524" rel="nofollow">https:&#x2F;&#x2F;x.com&#x2F;thsottiaux&#x2F;status&#x2F;2089082893804896524

            There&#x27;s at least a forum thread about it here:

            <a href="https:&#x2F;&#x2F;community.openai.com&#x2F;t&#x2F;why-does-codex-report-a-258-400-token-context-window-for-gpt-5-6-sol&#x2F;1394346&#x2F;5" rel="nofollow">https:&#x2F;&#x2F;community.openai.com&#x2F;t&#x2F;why-does-codex-report-a-258-4...

            1. gf000 · · focus · HN ↗
              But that 250k context worth way more than 1M in terms of how well it&#x27;s utilized, so actually I do like codex trying to keep you at that sweet spot.
            2. rrvsh · · focus · HN ↗
              Its really not tiny; you can&#x27;t compare Claude to GPT, they have honestly diverged enough that as the other reply said, 256k GPT is about equal to 1M Claude. The compaction is slightly annoying, and you can turn it up to 1M as you said if you truly need everything in context, but otherwise it&#x27;s perfectly serviceable
              1. kaoD · · focus · HN ↗
                [delayed]
            3. bjord · · focus · HN ↗
              &gt; unattended overnight claude opus session

              yes, exactly

              1. SyneRyder · · focus · HN ↗
                Not sure I understand if this was meant as a slight against Claude? Or agreement?

                These are often my best sessions - they&#x27;re unattended overnight, because by then we have the specification figured out, and I can just leave Claude to build out the rest, making good choices if it does find gaps in the spec. I regularly go to sleep &amp; wake up to an entirely new application completed. Claude never uses compacting in my sessions.

                I haven&#x27;t used GPT as much as I should have, so I&#x27;m prepared to be incorrect &amp; out of date. It just intuitively feels like I wouldn&#x27;t get the same from a 275K context window - maybe it uses lots of subagents? Even Deepseek &amp; GLM have 1 Million context windows now, so it &quot;feels&quot; strange for people to actually prefer the 275K window. But that&#x27;s just my intuition.

                1. bjord · · focus · HN ↗
                  neither, actually, just that unattended &quot;oneshot&quot; sessions are incredibly token inefficient

                  if you talk about them (in which you lean on an LLM as a sort-of independent employee) and conservative, chunk-based usage (in which you use the LLM as more of an extension of yourself), you&#x27;re comparing apples to oranges

                  a predefined spec obviously reduces that gap but how much is highly dependent on the level of detail

          3. jeremyjh · · focus · HN ↗
            I don’t usually have a problem doing a complete task in that context size. OMP does make a lot of use of rewind which may be helping - basically forks itself and sends back a summary after a long tangent. Coding takes use a Luna max agent.

            I’ve also found compaction not to be a problem when it does happen.

            1. threecheese · · focus · HN ↗
              How do you trigger this? I&#x27;ve been messing with OMP lately for funsies.
              1. jeremyjh · · focus · HN ↗
                Its in settings under Tools-&gt;Checkpoint&#x2F;Rewind. I don&#x27;t know why its not enabled by default.
          4. onlyrealcuzzo · · focus · HN ↗
            If it&#x27;s compacting every 5 mins, you&#x27;re going to notice it in your cache miss ratio and your costs...
            1. rrvsh · · focus · HN ↗
              It doesn&#x27;t - try it first
          5. Benjamin_Dobell · · focus · HN ↗
            The context window is configurable. I&#x27;ve been using ~600k for months. No, not API pricing, on a Codex sub.

            ~&#x2F;.codex&#x2F;config.toml

              model = &quot;gpt-6.1-sol&quot;
              model_context_window = 700000
              model_auto_compact_token_limit = 630000
      7. AmazingTurtle · · focus · HN ↗
        you can actually leverage 400k and 1M contexts in codex with very little code changes to the harness. note that excess context past the.. 250k or 400k mark (i don&#x27;t remember) is charged at 2x the price.
      8. manmal · · focus · HN ↗
        Your tool calls (MCPs?) are very likely too wasteful. Apply some filtering logic on the offending tool’s output. Either a wrapper CLI, or just tell codex how to filter.
      9. Gareth321 · · focus · HN ↗
        &gt; Cache doesn&#x27;t help you much when you are compacting every 5 minutes...

        It&#x27;s crazy on Codex. I sometimes get just 2-3 turns before it compacts. It has forced me to use persistent project documentation for everything. Maybe that&#x27;s not a bad thing but unless it reads all the documentation after every compaction (and uses half its cache), it goes off the rails. By comparison, Opus 5.5 is a breath of fresh air. It takes FAR longer to hit the cache limit and that means it keeps useful information in working memory far longer. I think this alone has resulted in a massive productivity and efficiency increase for me.

        1. RugnirViking · · focus · HN ↗
          iirc you can still turn the compaction limit up in codex, though they don&#x27;t make it easy. It costs way more when you use &quot;large context&quot; though, more than the ~256k that codex allows by default. You can also use the large context via the api directly
      10. exfalso · · focus · HN ↗
        what. I use Astra xhigh, sometimes max, never ran out of tokens on the 100$ thing. I&#x27;m using pi though which is by definition harder better faster stronger than claude code&#x2F;codex.
      11. jmalicki · · focus · HN ↗
        Use more subagents.

        The longer your chat gets, the slower and more expensive it gets.

        Subagents are expensive but they scale way closer to O(n) than O(n^2).

        Have some agents make bug reports&#x2F;feature requests&#x2F;roadmaps (linear is very AI friendly), others coordinate, others work on grinding out an individual ticket.

        If there is a good ticket-level description, it&#x27;s a waste of time IMO to have a main agent do it, that should be an agent with fresh context that will do it better faster (the shorter the context, the better models are at using the context they&#x27;re given).

        1. jaktet · · focus · HN ↗
          Subagents will inherit the context window at the point in which they are spawned, but it sounds like you&#x27;re more referring to orchestrating&#x2F;conducting&#x2F;managing multiple agents?
          1. jmalicki · · focus · HN ↗
            Both... even subagents inheriting the context window doesn&#x27;t cost a huge amount if the context window was never that large, but yes orchestrating&#x2F;conducting&#x2F;managing multiple agents is even better though higher thought cost (but the newer claude agents are really good at this in my experience, part of why I am using Claude a lot lately despite the models being more expensive that ChatGPT&#x27;s for the same performance when taken alone).

            Whenever I see my main agent do a compaction, that to me is a clear sign I didn&#x27;t have it delegate bounded tasks enough.

            Still, I see no evidence Codex or Claude Code inherit full context of the main agent in subagents, I&#x27;ve always seen them be prompted, but this is something high priority on my list of unknowns to understand better...

      12. KetoManx64 · · focus · HN ↗
        Do you just keep one conversation going for all projects? That&#x27;s the only way I&#x27;ve seen other people in my company burn through all their tokens.

        Everyone else that uses memory files and a new conversation for each new sub project&#x2F;feature rarely hit their weekly allotments.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.