‹ BackHN Continuity

Thread

GPT-6 Sol and Luna

1779 points · 855 comments · OfficialTurkey

  1. simonw · · focus · HN ↗
    GPT-6 Luna being half the price of GPT-5.6 Luna is a really big deal.

    Here&#x27;s GPT-6 Luna pelicans: <a href="https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F40d129fc140faca378b9c9f4f16c6ec2" rel="nofollow">https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=ht...

    And GPT-6 Sol: <a href="https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Fbe7ae25af2634b68bc34b7b7aaf02cb2" rel="nofollow">https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=ht...

    Scroll to the bottom for the GPT-6 Sol max one: <a href="https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Fbe7ae25af2634b68bc34b7b7aaf02cb2#response-5" rel="nofollow">https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=ht...

    For comparison, here are the pelicans I got for GPT-6 Astra: <a href="https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Ff789d2784fc6c5b870cc80f0b7cd9d01" rel="nofollow">https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=ht... - I still like the Astra Max one best.

    Here&#x27;s a comparison grid showing all of the GPT-6 and GPT-5.6 pelicans at all effort levels: <a href="https:&#x2F;&#x2F;static.simonwillison.net&#x2F;static&#x2F;2026&#x2F;gpt-6-and-5.6.html" rel="nofollow">https:&#x2F;&#x2F;static.simonwillison.net&#x2F;static&#x2F;2026&#x2F;gpt-6-and-5.6.h...

    The grid is actually really interesting, because it shows that the 5.6 family default to brighter colors than the 6 family.

    1. gizmodo59 · · focus · HN ↗
      6-luna is at the pareto for most of the tasks! I dont know how they make money here but its insane value from a closed source model. I&#x27;d go further and say it makes no sense (privacy, sovereignty etc aside) to use many other models as its not only expensive but also many providers don&#x27;t have that much GPUs to serve at a significant volume. <a href="https:&#x2F;&#x2F;openrouter.ai&#x2F;rankings?view=month#top-models" rel="nofollow">https:&#x2F;&#x2F;openrouter.ai&#x2F;rankings?view=month#top-models 5.6 luna is already the most used model this month.
      1. sieve · · focus · HN ↗
        My OpenCode Go stats for the last 30d:

        Cached Read: ~6,500M

        Input: ~150M

        Output: ~20M

        Approx $40 worth of usage across DeepSeek V4 Flash + MuseSpark Contributor 1.3. And a bit of both the GLM models. This is covered in a $10 subscription.

        If I were to use Luna&#x27;s API pricing:

        $0.02 x 6,500 = $130

        $0.20 x 150 = $30

        $1.20 x 20 = $24

        So $184. And this is assuming smaller coding sessions (&lt;272K) beyond which Luna pricing doubles.

        --

        Cost wise, these models are nice for small stuff. Translations etc. Any model that does not provide multiple Mtoks of cached reads per cent is not very useful to me for coding workflows.

        1. nearbuy · · focus · HN ↗
          This isn&#x27;t right. You&#x27;re comparing cost per token, but DeepSeek V4 Flash uses more tokens. Artificial Analysis found GPT 6 Luna to be significantly cheaper than DeepSeek: <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;comparisons?compare=deepseek-v4-1-flash%2Cgpt-6-luna" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;comparisons?compare=dee...
          1. sieve · · focus · HN ↗
            I do not (generally) trust benchmarks. I only trust what a model does with MY code.

            Forget DS. I asked MiMo 2.6 yesterday to explain ML&#x2F;LLMs to me succinctly and the pointed it at Karpathy&#x27;s micrograd code. It produced a C implementation called `xor_mlp`, a tiny model that learnt how `xor` worked. I then asked it to produce a model that can play tictactoe without losing (mostly). It did. It supervised the training process and produced a compiled version with multiple switches. The pi-dev session is still running, so here are actual stats

            ↑45k ↓35k R1.0M CH99.4% $0.019 4.2%&#x2F;1.0M (auto) - (opencode-go) mimo-v2.6-flash • high

            And here is Luna on the same workflow (I had to poke and prod a bit to get what I wanted):

            ↑141 ↓34k R1.0M W43k CH95.3% $0.072 4.2%&#x2F;1.1M (auto) (opencode-go) gpt-5.6-luna • high

            I expect similar results from DS41F&#x2F;MS13. Closer to MiMo costs than Luna.

            So the &quot;significantly cheaper&quot; thing may not really hold, more so when Luna has to actually read my codebase to do the stuff that I want rather than rely on world knowledge. The 8-10x cache read cost differential itself will kill the token budget.

            1. nearbuy · · focus · HN ↗
              With GPT-6 Luna (which is what the parent comment was talking about), that would come to 3.2¢, assuming GPT-6 used the same number of tokens.

              I don&#x27;t think you can guess more precisely than an order of magnitude from trying each once on one task.

              1. sieve · · focus · HN ↗
                DS is VERY talkative. Luna is less so. Still do not think, based on this little experiment, that Luna could beat DS in price: API-to-API. As part of a Plus&#x2F;Pro plan? Sure.
                1. ducktoysleftout · · focus · HN ↗
                  Its style of writing is part of the fun. Seeing reasoning traces fly by that each start with “Hmm…” is pretty amusing in my opinion especially if you try to vocalize it in your mind.
                  1. Thanemate · · focus · HN ↗
                    In a discussion about cost effectiveness, how subjectively fun the writing feels like to the reader isn&#x27;t a factor, except maybe if we were working on writing comedy.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.