‹ BackHN Continuity

Thread

GPT-6 Sol and Luna

1779 points · 855 comments · OfficialTurkey

  1. jeffnash · · focus · HN ↗
    At this point, the deciding factors for me between Claude Code 20x and Codex Pro 20x are:

    1/ Usage limits: downstream of input/output cost, but resets and obscure windows and odd 20x plan / 5x plan != 4x usage math throw a wrench into it. Winner right now is Codex by a mile, especially when you factor in the fact that ChatGPT usage (even 6 Astra Pro) is essentially unmetered on the 20x plan. Always a bummer when asking if I should see a doctor about a rash means I can't code as much. It's also is a godsend if you use an MCP like oracle to automate the process of calling the Pro model on particularly tough problems, giving better planning results or deeper code analysis without burning usage.

    2/ Context window in the harness. Claude Code wins on this. There used to be a toml file workaround for Codex to extend the GPT context window to 1m, but this stopped working on the plans and only on per-token billing (ETA: noname120 pointed out this is no longer the case and it can be enabled again [1]). 252k is just not enough. Codex's compaction is very good, fwiw, but it happens so frequently that even a model as powerful as Astra sometimes loses the plot on long-running tasks.

    3/ Ability to use the plan outside of the official harness. Codex wins. Anthropic does shit like bills requests as extra usage if it sees a hermes.md in a commit.

    I've subscription hopped a bunch, and at times I've had both, but I keep coming back to Codex because it wins on 2/3.

    ETA: apparently I haven't been Keeping Up With the Altmans and new 20x signups have been disabled for a few weeks. I am grandfathered in, which makes the comparison above pretty much moot.

    [1]<a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49806060">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49806060

    1. glub · · focus · HN ↗
      &gt; Usage limits [...] Winner right now is Codex by a mile

      This hasn&#x27;t been the case since around July. If you measure usage in raw api costs, Anthropic is actually giving more on $200 than OpenAI now. This includes resets. Usage allocation difference would be humiliating for codex subs were it not for resets. But fixing usage limits with resets is ugly, and they&#x27;re not good for your mental well-being.

      &gt; Context window in the harness

      Codex now allows 1M for subs with config params. But generally speaking, you shouldn&#x27;t really be using 1M context. If you accidentally send a request with say, ~700k context already accumulated in a session which is outside cache TTL, you&#x27;re paying full cost of these 700k tokens.

      &gt; I&#x27;ve subscription hopped a bunch

      OpenAI actually has a new strategy to prevent subscription hopping after their 2-3 month-long marketing push to get claude-folks to switch over:

      you can&#x27;t buy a $200 sub anymore. So if you cancel, you won&#x27;t be able to get back in. Hostage situation, essentially.

      EDIT: re: usage limits, oh-my-pi maintainer has been tracking this - <a href="https:&#x2F;&#x2F;nitter.xitter.cc&#x2F;_can1357&#x2F;status&#x2F;2090075496948060372" rel="nofollow">https:&#x2F;&#x2F;nitter.xitter.cc&#x2F;_can1357&#x2F;status&#x2F;2090075496948060372

      1. rudedogg · · focus · HN ↗
        I’ve been a Claude user, switched to Codex expecting usage limits to be more loose but I can’t even get through a basic sysadmin task on the $20 plan using Sol medium before I hit the 5hr one.

        I think I’m gonna move back to a Claude plan. I could barely hit the $200 limit if I went non-stop on programming tasks.

        1. qlte · · focus · HN ↗
          I do the bulk of work on Sol Medium&#x2F;Low and don&#x27;t have that experience on the $20 plan. If you said Astra I&#x27;d agree it&#x27;s easy to burn through the 5 hours even on the lower reasoning levels.

          Do you have &#x2F;fast enabled by any chance?

          1. rudedogg · · focus · HN ↗
            I don’t think so, I’ve seen it suggest I try it.

            Yes, I was considering the $100 plan, but I hit the 5hr limit in an hour, thought about how even at $100 I cant go non-stop on a single agent running Sol Medium and decided I should probably get back on Claude

            1. hirvi74 · · focus · HN ↗
              Sorry if I am misunderstanding you, but I am pretty sure the $100 plan doesn’t have a 5hr usage limit. So, if that was what was preventing you from going non-stop, it might be worth it.

              I am considering the plan myself. I just don’t know if I want to fork out $100 per month for something I will make $0 off of.

              1. rudedogg · · focus · HN ↗
                Thanks for the info, I didn’t know that about the 5hr limit
        2. glub · · focus · HN ↗
          I think OpenAI essentially executed a bait-and-switch here, and they&#x27;ve lost a lot of goodwill with me, like Anthropic did, before them.

          When they started the aggressive campaign, entire X (including myself, sadly) was full of posts about how &quot;unlimited&quot; codex usage is even on a $20 plan. Sam Altman was posting something in line of &quot;we love our users, unlike Anthropic&quot;. Got my network to get codex subs because of the value compared to claude.

          Then they gradually reduced the limits to the point where even $200 plan only lasts you just 1-2 days and $20 is basically unusable, then the hostage thing.

          1. ivm · · focus · HN ↗
            &gt; even $200 plan only lasts you just 1-2 days

            I’m working on 2–3 apps at once and barely manage to use 70–80% of the weekly quota, with everything being done on Sol Xhigh and, lately, all the planning on Astra. Still have a reset stashed too.

          2. malshe · · focus · HN ↗
            I remember Tibo Sottiaux telling people on X how OI doesn&#x27;t believe in 5 hour limit just a day or two before OI adopted it.
            1. phyrex · · focus · HN ↗
              tbf that&#x27;s only for the pro plan, not the two max plans
          3. Numerlor · · focus · HN ↗
            I think the models getting dumber impacted that too, after a couple weeks both sol and Luna felt notably worse to me than they did at release
          4. slopinthebag · · focus · HN ↗
            &gt; Then they gradually reduced the limits to the point where even $200 plan only lasts you just 1-2 days

            what? im on the $100 plan and ive literally never run out of usage, and thats mostly running Astra high.

            maybe its the harness

        3. hadlock · · focus · HN ↗
          I&#x27;ve run into hitting limits on the personal plan perhaps twice since the beginning of the year. But also I don&#x27;t use the personal plan for coding tasks between 7am-noon M-F.
        4. boardwaalk · · focus · HN ↗
          similar here: I tried Codex $20&#x2F;mo on a trial and I ran out of 5hr usage mid way through a medium complexity task on a medium size model twice and gave up there. I don’t recall the equiv Claude plan being anything like that. Anecdata, but not great for OAI if they actually want to retain people on a trial.
          1. cromka · · focus · HN ↗
            You don&#x27;t get Fable on Claude 20 USD plan. You get Sol on equivalent Codex plan.
            1. istjohn · · focus · HN ↗
              You meant Astra, not Sol, I think. But Opus 5.5 is slightly better than Fable and Astra now.
              1. cromka · · focus · HN ↗
                Yes, it&#x27;s Astra and Sol before.

                Opus 5.5 is better in benchmarks, but has substantially less parameters so is world knowledge cannot compare against Fable or Astra.

        5. this_user · · focus · HN ↗
          Astra is barely usable even on the $100 plan. And that is if it doesn&#x27;t just burn through 80% of your weekly quota in a couple of hours by continually expanding the scope of the task you gave it - while not noticing the failing tests that are right in front of it.

          Opus is at least actually usable even on the small plan. The main downside is its insane writing style, but 5.5 seems to address that somewhat. Otherwise, you can just use your $20 OpenAI plan to have Luna de-slop Opus&#x27; prose, which seems to work fine.

          1. Huppie · · focus · HN ↗
            I have a Claude Code hook that calls codex for a code review on commit time (Codex is set to Astra Medium) and it&#x27;s been pretty good in general. It sometimes hits the 5hr limit but most of the time it provides really good feedback and because it&#x27;s a completely different model it&#x27;s mostly complementary to what Fable&#x2F;Opus do themselves. IMHO it&#x27;s been $20 well spent.

            ...but the few times I&#x27;ve tried to use codex for a moderately difficult task it burned through its limit extremely quickly.

        6. cromka · · focus · HN ↗
          But you don&#x27;t get Fable on Claude 20 USD plan, then why compare it Sol on Codex 20 USD?
          1. rudedogg · · focus · HN ↗
            Sol is their middle model. Luna is smallest. And Astra is big, their Fable equivalent.
            1. matheusmoreira · · focus · HN ↗
              My code review benchmark put Sol 5.6 on the same performance tier as Fable 5.

              <a href="https:&#x2F;&#x2F;www.matheusmoreira.com&#x2F;articles&#x2F;code-reviewing-lone-lisp-with-sol-and-fable" rel="nofollow">https:&#x2F;&#x2F;www.matheusmoreira.com&#x2F;articles&#x2F;code-reviewing-lone-...

          2. sisyphus15 · · focus · HN ↗
            Sol is OpenAI&#x27;s Opus, and Astra is OpenAI&#x27;s Fable. Both pricing-wise, and performance-wise.
        7. joquarky · · focus · HN ↗
          On the $20 plan, you can&#x27;t use Sol for much more than planning and review. Luna xhigh for the rest. Have Sol write the plan specifically for Luna so it adds more direction and validation to the plan.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.