‹ BackHN Continuity

Thread

GPT-6 Sol and Luna

1779 points · 855 comments · OfficialTurkey

  1. simonw · · focus · HN ↗
    GPT-6 Luna being half the price of GPT-5.6 Luna is a really big deal.

    Here&#x27;s GPT-6 Luna pelicans: <a href="https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F40d129fc140faca378b9c9f4f16c6ec2" rel="nofollow">https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=ht...

    And GPT-6 Sol: <a href="https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Fbe7ae25af2634b68bc34b7b7aaf02cb2" rel="nofollow">https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=ht...

    Scroll to the bottom for the GPT-6 Sol max one: <a href="https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Fbe7ae25af2634b68bc34b7b7aaf02cb2#response-5" rel="nofollow">https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=ht...

    For comparison, here are the pelicans I got for GPT-6 Astra: <a href="https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Ff789d2784fc6c5b870cc80f0b7cd9d01" rel="nofollow">https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=ht... - I still like the Astra Max one best.

    Here&#x27;s a comparison grid showing all of the GPT-6 and GPT-5.6 pelicans at all effort levels: <a href="https:&#x2F;&#x2F;static.simonwillison.net&#x2F;static&#x2F;2026&#x2F;gpt-6-and-5.6.html" rel="nofollow">https:&#x2F;&#x2F;static.simonwillison.net&#x2F;static&#x2F;2026&#x2F;gpt-6-and-5.6.h...

    The grid is actually really interesting, because it shows that the 5.6 family default to brighter colors than the 6 family.

    1. gizmodo59 · · focus · HN ↗
      6-luna is at the pareto for most of the tasks! I dont know how they make money here but its insane value from a closed source model. I&#x27;d go further and say it makes no sense (privacy, sovereignty etc aside) to use many other models as its not only expensive but also many providers don&#x27;t have that much GPUs to serve at a significant volume. <a href="https:&#x2F;&#x2F;openrouter.ai&#x2F;rankings?view=month#top-models" rel="nofollow">https:&#x2F;&#x2F;openrouter.ai&#x2F;rankings?view=month#top-models 5.6 luna is already the most used model this month.
      1. sieve · · focus · HN ↗
        My OpenCode Go stats for the last 30d:

        Cached Read: ~6,500M

        Input: ~150M

        Output: ~20M

        Approx $40 worth of usage across DeepSeek V4 Flash + MuseSpark Contributor 1.3. And a bit of both the GLM models. This is covered in a $10 subscription.

        If I were to use Luna&#x27;s API pricing:

        $0.02 x 6,500 = $130

        $0.20 x 150 = $30

        $1.20 x 20 = $24

        So $184. And this is assuming smaller coding sessions (&lt;272K) beyond which Luna pricing doubles.

        --

        Cost wise, these models are nice for small stuff. Translations etc. Any model that does not provide multiple Mtoks of cached reads per cent is not very useful to me for coding workflows.

        1. gizmodo59 · · focus · HN ↗
          It’s not direct token to token pricing and everyone misses it. The cost is how much tokens to complete something multiplied by token pricing. I can have a model at .0001 per million tokens but it’s so inefficient that it takes 10B tokens to complete a task means it’s expensive.
          1. sieve · · focus · HN ↗
            I am not designing rockets. Most of my work is bog standard hobbyist stuff: compilers, vms, sandboxes, system tools of various kinds, SSGs, markup languages, plain text ledgers etc. Even Gemma&#x2F;Qwen running locally can manage this.

            Frankly, I have no idea what people do with Opus&#x2F;Fable etc. I don&#x27;t think anything I do needs something that charges $50&#x2F;M for output tokens.

            1. apatheticonion · · focus · HN ↗
              Can confirm. I have been using DeepSeek since forever and it&#x27;s so good I was able to write a compiler and native desktop applications with it. I use it as a coding assistant in my IDE so the results end up at the same quality I would write by hand.

              I recently started a job that only uses Claude models. Opus and Sonnet are so slow you have no choice but to do multiple tasks in parallel. You create a git worktree, set off an agent to do something, another worktree, set out an agent - then play video games for 20 minutes until they complete the task (poorly).

              You can&#x27;t really do &quot;guide coding&quot; like you can with DeepSeek-style flash models because Claude is too slow.

              I think the idea with slow frontier models is to end up with &quot;software factories&quot;, where you just write tickets and send them to a harness that delegates work to agents&#x2F;subagents. Your job is to prompt and review (and eventually just prompt).

              Mathematically and assuming token prices&#x2F;efficiency remains constant, the collective US AI industry needs to increase token usage by 15x before 2030 (3.5 years from now) to satisfy investors. With companies already implementing token limits, the only place from here is for frontier models to replace staff entirely to expand budgets for tokens. The only way to do that is to demonstrate the efficacy of software factories and headless agentic workflows.

              Objectively, I have set up a software factory and I do see the utility of it, though I did it with DeepSeek and prices are 1% that of frontier models - which doesn&#x27;t bode well for investors looking for an eventual return.

              Heck, my old M1 MBP 32gb running Qwen 3.6 35b a3b sipping 10w when generating tokens is good enough for a lot of my guide-coding work - it&#x27;s just a bit slow so I use DeepSeek instead. When hardware prices come down, I honestly wouldn&#x27;t see a need to subscribe to any service, I&#x27;d just grow my own tokens at home.

              1. tripzilch · · focus · HN ↗
                &gt; running Qwen 3.6 35b a3b sipping 10w when generating tokens is good enough for a lot of my guide-coding work

                can you tell more about how you&#x27;re using it? like, what harness? or also in the IDE?

                I found Qwen3.6 35B&#x2F;A3B to make slightly too many mistakes (already in its harness&#x27; tool use, hence my question), maybe it gets the job done, but it will also sometimes generate a bit of a mess (e.g. editing&#x2F;creating files in the wrong folders) and fixing&#x2F;solving its own mistakes takes time (or tokens) ..

                1. apatheticonion · · focus · HN ↗
                  Guide coding is more forgiving because the diffs are small enough that you can just try something and veto it if it&#x27;s not good. Qwen 3.6 a3b makes more mistakes than DeepSeek but it&#x27;s free. I&#x27;d imagine the next ~32gb-vram-class MoE model from Qwen will close the gap.

                  The real deciding factor for me is inference speed.

                  I use VSCode insiders with their BYO model configuration. I stage little bits at a time in a tight prompt-review-prompt-review workflow. I occasionally use dedicated harnesses (DeepSeek harness, Codex, etc) because they have better tools and outcomes than VSCode&#x27;s built-in harness. When building visual applications, desktop harnesses are better because it&#x27;s a bit easier to send screenshots to the agent.

                  I love Zed editor but its AI review features are lacking compared to VSCode, but I got back to it frequently and am ready to switch over when it solves that.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.