GPT-6 Luna being half the price of GPT-5.6 Luna is a really big deal.
Here's GPT-6 Luna pelicans: <a href="https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F40d129fc140faca378b9c9f4f16c6ec2" rel="nofollow">https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
And GPT-6 Sol: <a href="https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Fbe7ae25af2634b68bc34b7b7aaf02cb2" rel="nofollow">https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
Scroll to the bottom for the GPT-6 Sol max one: <a href="https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Fbe7ae25af2634b68bc34b7b7aaf02cb2#response-5" rel="nofollow">https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
For comparison, here are the pelicans I got for GPT-6 Astra: <a href="https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Ff789d2784fc6c5b870cc80f0b7cd9d01" rel="nofollow">https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - I still like the Astra Max one best.
Here's a comparison grid showing all of the GPT-6 and GPT-5.6 pelicans at all effort levels: <a href="https://static.simonwillison.net/static/2026/gpt-6-and-5.6.html" rel="nofollow">https://static.simonwillison.net/static/2026/gpt-6-and-5.6.h...
The grid is actually really interesting, because it shows that the 5.6 family default to brighter colors than the 6 family.
6-luna is at the pareto for most of the tasks! I dont know how they make money here but its insane value from a closed source model. I'd go further and say it makes no sense (privacy, sovereignty etc aside) to use many other models as its not only expensive but also many providers don't have that much GPUs to serve at a significant volume. <a href="https://openrouter.ai/rankings?view=month#top-models" rel="nofollow">https://openrouter.ai/rankings?view=month#top-models 5.6 luna is already the most used model this month.
Approx $40 worth of usage across DeepSeek V4 Flash + MuseSpark Contributor 1.3. And a bit of both the GLM models. This is covered in a $10 subscription.
If I were to use Luna's API pricing:
$0.02 x 6,500 = $130
$0.20 x 150 = $30
$1.20 x 20 = $24
So $184. And this is assuming smaller coding sessions (<272K) beyond which Luna pricing doubles.
--
Cost wise, these models are nice for small stuff. Translations etc. Any model that does not provide multiple Mtoks of cached reads per cent is not very useful to me for coding workflows.
It’s not direct token to token pricing and everyone misses it. The cost is how much tokens to complete something multiplied by token pricing. I can have a model at .0001 per million tokens but it’s so inefficient that it takes 10B tokens to complete a task means it’s expensive.
I am not designing rockets. Most of my work is bog standard hobbyist stuff: compilers, vms, sandboxes, system tools of various kinds, SSGs, markup languages, plain text ledgers etc. Even Gemma/Qwen running locally can manage this.
Frankly, I have no idea what people do with Opus/Fable etc. I don't think anything I do needs something that charges $50/M for output tokens.
Can confirm. I have been using DeepSeek since forever and it's so good I was able to write a compiler and native desktop applications with it. I use it as a coding assistant in my IDE so the results end up at the same quality I would write by hand.
I recently started a job that only uses Claude models. Opus and Sonnet are so slow you have no choice but to do multiple tasks in parallel. You create a git worktree, set off an agent to do something, another worktree, set out an agent - then play video games for 20 minutes until they complete the task (poorly).
You can't really do "guide coding" like you can with DeepSeek-style flash models because Claude is too slow.
I think the idea with slow frontier models is to end up with "software factories", where you just write tickets and send them to a harness that delegates work to agents/subagents. Your job is to prompt and review (and eventually just prompt).
Mathematically and assuming token prices/efficiency remains constant, the collective US AI industry needs to increase token usage by 15x before 2030 (3.5 years from now) to satisfy investors. With companies already implementing token limits, the only place from here is for frontier models to replace staff entirely to expand budgets for tokens. The only way to do that is to demonstrate the efficacy of software factories and headless agentic workflows.
Objectively, I have set up a software factory and I do see the utility of it, though I did it with DeepSeek and prices are 1% that of frontier models - which doesn't bode well for investors looking for an eventual return.
Heck, my old M1 MBP 32gb running Qwen 3.6 35b a3b sipping 10w when generating tokens is good enough for a lot of my guide-coding work - it's just a bit slow so I use DeepSeek instead. When hardware prices come down, I honestly wouldn't see a need to subscribe to any service, I'd just grow my own tokens at home.
I use Claude Sonnet and ChatGPT via the web UI. I often use Claude to come up with specs for my ideas. This is becoming less and less useful. DS4/MS13/MiMo are almost there for these use cases as well.
I dogfood everything I produce, and the models are good at collaborating with me on a spec and then turning it into code.
If Sonnet/ChatGPT suddenly became unavailable due to Anthropic/OpenAI suddenly not being able to subsidize the freemium/loss-leader experience, I probably would not miss them. Google/BraveAI already give you the AI experience during search (when you are looking for stuff to buy, or something particular). Claude/ChatGPT still have a minor edge in this use case for me right now.
I haven't tried GPT Luna yet, but anything Claude takes forever in my experience.
In guided coding sessions, DeepSeek is so fast I rarely have the time to look away from my screen. I've noticed DS 4.1 is a little slower, but claude is still orders of magnitude slower.
e.g.
- "Turn this SQL CREATE TABLE statement into a repository class and generate models for it"
- "Add error handling / timeout to this function"
- Highlight text "Using this as a reference, repeat for all other files in this folder"
> running Qwen 3.6 35b a3b sipping 10w when generating tokens is good enough for a lot of my guide-coding work
can you tell more about how you're using it? like, what harness? or also in the IDE?
I found Qwen3.6 35B/A3B to make slightly too many mistakes (already in its harness' tool use, hence my question), maybe it gets the job done, but it will also sometimes generate a bit of a mess (e.g. editing/creating files in the wrong folders) and fixing/solving its own mistakes takes time (or tokens) ..
Same thing for me. on an M5 max, Qwen 3.6 35b gives me between 150 and 200 tps using splash as inference engine.
More than enough for guided code sessions, at 100% privacy. And i can use obliverated models if i am trying to harden my own app, something i cannot do with cloud providers.
I feel like ornith1.5 35B/A3B is an overall stronger model on the same architecture, so a drop-in replacement untill qwen3.8/qwen4 is released. Using the 8bit quant on my M4 max gets around 80tok/sec output/decode on an empty context, dropping down to 35ish on nearly full one.
Guide coding is more forgiving because the diffs are small enough that you can just try something and veto it if it's not good. Qwen 3.6 a3b makes more mistakes than DeepSeek but it's free. I'd imagine the next ~32gb-vram-class MoE model from Qwen will close the gap.
The real deciding factor for me is inference speed.
I use VSCode insiders with their BYO model configuration. I stage little bits at a time in a tight prompt-review-prompt-review workflow. I occasionally use dedicated harnesses (DeepSeek harness, Codex, etc) because they have better tools and outcomes than VSCode's built-in harness. When building visual applications, desktop harnesses are better because it's a bit easier to send screenshots to the agent.
I love Zed editor but its AI review features are lacking compared to VSCode, but I got back to it frequently and am ready to switch over when it solves that.
It sounds like you’re still writing code by hand and reading and reviewing code.
For that any decent model from the past year will do.
If you want to forget how to write code and not read generated code, then you need a very good frontier model, ideally one from 6-12 months in the future.
I used to read the code till around May. Now I don't. Instead I validate behavior. And have multiple LLMs verify that the code implements my handwritten spec.
MiMo 2.5/2.6, MuseSpark 1.3, DeepSeek V4/4.1 Flash and GLM 5.3 Flash are perfectly capable of following my spec and then poking holes in the implementation till there are none left.
> It sounds like you’re still writing code by hand and reading and reviewing code.
This is such a naive, baseless opinion.
Nowadays any AI coding assistant service supports or can be used with sub-agent orchestration frameworks.
If you are in the business of software factories, you can use the cheapest models and even local models to handle some if not all tasks in the orchestration chain.
Adding tests or executing tests (unit, integration, UI, you name it) doesn't require a cutting edge frontier model. Neither does refactoring. Neither does identifying call stacks. Neither does planning a changeset.
You have your specialized subagents, you put together a small orchestrator subagent that handles feedback loops and handoffs,and you throw it at tasks.
For the past couple of months, most of the code I write is not code per se, it's subtask orchestrators. And unlike the old "only Opus is passable" days, the cheapest models do get the job done.
simonw · · focus · HN ↗
Here's GPT-6 Luna pelicans: <a href="https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F40d129fc140faca378b9c9f4f16c6ec2" rel="nofollow">https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
And GPT-6 Sol: <a href="https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Fbe7ae25af2634b68bc34b7b7aaf02cb2" rel="nofollow">https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
Scroll to the bottom for the GPT-6 Sol max one: <a href="https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Fbe7ae25af2634b68bc34b7b7aaf02cb2#response-5" rel="nofollow">https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
For comparison, here are the pelicans I got for GPT-6 Astra: <a href="https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Ff789d2784fc6c5b870cc80f0b7cd9d01" rel="nofollow">https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - I still like the Astra Max one best.
Here's a comparison grid showing all of the GPT-6 and GPT-5.6 pelicans at all effort levels: <a href="https://static.simonwillison.net/static/2026/gpt-6-and-5.6.html" rel="nofollow">https://static.simonwillison.net/static/2026/gpt-6-and-5.6.h...
The grid is actually really interesting, because it shows that the 5.6 family default to brighter colors than the 6 family.
gizmodo59 · · focus · HN ↗
sieve · · focus · HN ↗
Cached Read: ~6,500M
Input: ~150M
Output: ~20M
Approx $40 worth of usage across DeepSeek V4 Flash + MuseSpark Contributor 1.3. And a bit of both the GLM models. This is covered in a $10 subscription.
If I were to use Luna's API pricing:
$0.02 x 6,500 = $130
$0.20 x 150 = $30
$1.20 x 20 = $24
So $184. And this is assuming smaller coding sessions (<272K) beyond which Luna pricing doubles.
--
Cost wise, these models are nice for small stuff. Translations etc. Any model that does not provide multiple Mtoks of cached reads per cent is not very useful to me for coding workflows.
gizmodo59 · · focus · HN ↗
sieve · · focus · HN ↗
Frankly, I have no idea what people do with Opus/Fable etc. I don't think anything I do needs something that charges $50/M for output tokens.
apatheticonion · · focus · HN ↗
I recently started a job that only uses Claude models. Opus and Sonnet are so slow you have no choice but to do multiple tasks in parallel. You create a git worktree, set off an agent to do something, another worktree, set out an agent - then play video games for 20 minutes until they complete the task (poorly).
You can't really do "guide coding" like you can with DeepSeek-style flash models because Claude is too slow.
I think the idea with slow frontier models is to end up with "software factories", where you just write tickets and send them to a harness that delegates work to agents/subagents. Your job is to prompt and review (and eventually just prompt).
Mathematically and assuming token prices/efficiency remains constant, the collective US AI industry needs to increase token usage by 15x before 2030 (3.5 years from now) to satisfy investors. With companies already implementing token limits, the only place from here is for frontier models to replace staff entirely to expand budgets for tokens. The only way to do that is to demonstrate the efficacy of software factories and headless agentic workflows.
Objectively, I have set up a software factory and I do see the utility of it, though I did it with DeepSeek and prices are 1% that of frontier models - which doesn't bode well for investors looking for an eventual return.
Heck, my old M1 MBP 32gb running Qwen 3.6 35b a3b sipping 10w when generating tokens is good enough for a lot of my guide-coding work - it's just a bit slow so I use DeepSeek instead. When hardware prices come down, I honestly wouldn't see a need to subscribe to any service, I'd just grow my own tokens at home.
sieve · · focus · HN ↗
I dogfood everything I produce, and the models are good at collaborating with me on a spec and then turning it into code.
If Sonnet/ChatGPT suddenly became unavailable due to Anthropic/OpenAI suddenly not being able to subsidize the freemium/loss-leader experience, I probably would not miss them. Google/BraveAI already give you the AI experience during search (when you are looking for stuff to buy, or something particular). Claude/ChatGPT still have a minor edge in this use case for me right now.
jklmnopqrstuvw · · focus · HN ↗
apatheticonion · · focus · HN ↗
In guided coding sessions, DeepSeek is so fast I rarely have the time to look away from my screen. I've noticed DS 4.1 is a little slower, but claude is still orders of magnitude slower.
e.g.
- "Turn this SQL CREATE TABLE statement into a repository class and generate models for it"
- "Add error handling / timeout to this function"
- Highlight text "Using this as a reference, repeat for all other files in this folder"
tripzilch · · focus · HN ↗
can you tell more about how you're using it? like, what harness? or also in the IDE?
I found Qwen3.6 35B/A3B to make slightly too many mistakes (already in its harness' tool use, hence my question), maybe it gets the job done, but it will also sometimes generate a bit of a mess (e.g. editing/creating files in the wrong folders) and fixing/solving its own mistakes takes time (or tokens) ..
jorgeleo · · focus · HN ↗
More than enough for guided code sessions, at 100% privacy. And i can use obliverated models if i am trying to harden my own app, something i cannot do with cloud providers.
Phemist · · focus · HN ↗
Phemist · · focus · HN ↗
apatheticonion · · focus · HN ↗
The real deciding factor for me is inference speed.
I use VSCode insiders with their BYO model configuration. I stage little bits at a time in a tight prompt-review-prompt-review workflow. I occasionally use dedicated harnesses (DeepSeek harness, Codex, etc) because they have better tools and outcomes than VSCode's built-in harness. When building visual applications, desktop harnesses are better because it's a bit easier to send screenshots to the agent.
I love Zed editor but its AI review features are lacking compared to VSCode, but I got back to it frequently and am ready to switch over when it solves that.
unsupp0rted · · focus · HN ↗
For that any decent model from the past year will do.
If you want to forget how to write code and not read generated code, then you need a very good frontier model, ideally one from 6-12 months in the future.
sieve · · focus · HN ↗
MiMo 2.5/2.6, MuseSpark 1.3, DeepSeek V4/4.1 Flash and GLM 5.3 Flash are perfectly capable of following my spec and then poking holes in the implementation till there are none left.
cluckindan · · focus · HN ↗
locknitpicker · · focus · HN ↗
This is such a naive, baseless opinion.
Nowadays any AI coding assistant service supports or can be used with sub-agent orchestration frameworks.
If you are in the business of software factories, you can use the cheapest models and even local models to handle some if not all tasks in the orchestration chain.
Adding tests or executing tests (unit, integration, UI, you name it) doesn't require a cutting edge frontier model. Neither does refactoring. Neither does identifying call stacks. Neither does planning a changeset.
You have your specialized subagents, you put together a small orchestrator subagent that handles feedback loops and handoffs,and you throw it at tasks.
For the past couple of months, most of the code I write is not code per se, it's subtask orchestrators. And unlike the old "only Opus is passable" days, the cheapest models do get the job done.
darkwater · · focus · HN ↗