All that keeps jumping out at me is how they've set it to refuse giving users thinking tokens and prompts for full reasoning in output. Just drives me further away; I may not stop using Claude completely for now, but I'll be moving even more of my primary workload to Chinese providers. That's where openness and freedom is now at.
What Chinese models/providers are you using for this? I'm hitting Claude's weekly limits much sooner than I used to with roughly the same workload, so I'm interested in trying alternatives, especially ones with strong coding/agentic performance.
In my experience that tactic works well if the codebase is limited in size, or well maintained and separated. Otherwise I do notice a difference also letting fable do the execution, not just the planning for complex tasks.
In my experience it never works well on any real work. In fact, I'd go the opposite, plan with the dumb model and execute with the smart model because at least the model writing the code and solving the emergent problems is capable.
In my experience (and I've been trying this a bunch): smart planner + dumb executor produces worse code with higher spend than simply using the smart planner to do both.
It's easy to understand why:
- If the planner has truly thought the issue through, properly designed the solution, solved all of the emergent problems, then the final "write" of the code is just a few more output tokens.
- If the planner has NOT truly planned the issue completely, then you're letting a substantially dumber and less capable model make significant decisions, and trusting its problem solving, without having a better model check it.
If you're highly cost conscious (paying for your own tokens and not making any money) then you have no choice but to trade your time and effort for tricks like this to save money by lowering the quality of your output.
But if your employer is paying for tokens: just use the smarter model. You save your time preventing re-work and reducing code review, you save your employer money (primarily from the cost of your own labor and reduced rework), and you get a better output every time (Opus 5.5 mogs Deepseek 4.1 flash in every single way except cost).
The cost is 20-40x less for Deepseek Flash v4.1. If you are just comparing to Sonnet or you aren't paying (your case) then your advice makes perfect sense.
I also agree that its a big mistake to have a flash model implement without a strong model reviewing.
I have Opus plan, Deepseek implement the code, and then review with Opus [1]. In this workflow I am saving a lot of money by having Deepseek do the implementation. Note that the review back-and-forth is fully automated [2], so it doesn't take any extra attention from me.
Deepseek Flash v4.1 is only "40X cheaper" if you do not account for the time of the engineer reading the output. If Opus 5.5 high requires 1/2 of the actual engineer time, and the engineer costs $100-$200/hr, then Deepseek v4.1 is actually the more expensive model to use.
I have tried your workflow many times, and simply letting Opus do the implementation costs much less than wasting hundreds of millions of tokens letting deepseek and opus go back and forth and back and forth. And bonus, my project finishes in 5 minutes instead of 20.
Bingo, DeepSeek (v4.1) is horribly overhyped. In all my personal benchmarks it sits below Glm5.3 Flash. Waaaay below Qwen3.8-Flash-Next a model less than half it's size.
No, the only open weight model that really makes sense for me is Qwen3.8-Flash-Next, but it is mainly because I can run it locally with reasonable speed (prefill between 650-1400t/s generation between 22-50t/s depending on number of slots/users I configure).
This is the first model that truly competes with Opus 4.8. I'd say it may be better than Opus 4.6 on programming.
But it is very verbose when it comes to reasoning tokens. The more difficult the task the more verbose it is. Certain very hard tasks that take opus 4.8 400k tokens take Qwen3.8-Flash-Next 2M tokens... But it finishes them.
And what you loose on the generation speed you get back on input caching you can keep on for weeks.
skeledrew · · focus · HN ↗
rajeevk · · focus · HN ↗
arcanemachiner · · focus · HN ↗
A common tactic is to used a big brain model like Opus for planning and reviewing, and a cheaper model for execution.
lukan · · focus · HN ↗
criley2 · · focus · HN ↗
In my experience (and I've been trying this a bunch): smart planner + dumb executor produces worse code with higher spend than simply using the smart planner to do both.
It's easy to understand why:
- If the planner has truly thought the issue through, properly designed the solution, solved all of the emergent problems, then the final "write" of the code is just a few more output tokens.
- If the planner has NOT truly planned the issue completely, then you're letting a substantially dumber and less capable model make significant decisions, and trusting its problem solving, without having a better model check it.
If you're highly cost conscious (paying for your own tokens and not making any money) then you have no choice but to trade your time and effort for tricks like this to save money by lowering the quality of your output.
But if your employer is paying for tokens: just use the smarter model. You save your time preventing re-work and reducing code review, you save your employer money (primarily from the cost of your own labor and reduced rework), and you get a better output every time (Opus 5.5 mogs Deepseek 4.1 flash in every single way except cost).
gregwebs · · focus · HN ↗
I also agree that its a big mistake to have a flash model implement without a strong model reviewing.
I have Opus plan, Deepseek implement the code, and then review with Opus [1]. In this workflow I am saving a lot of money by having Deepseek do the implementation. Note that the review back-and-forth is fully automated [2], so it doesn't take any extra attention from me.
criley2 · · focus · HN ↗
I have tried your workflow many times, and simply letting Opus do the implementation costs much less than wasting hundreds of millions of tokens letting deepseek and opus go back and forth and back and forth. And bonus, my project finishes in 5 minutes instead of 20.
Roark66 · · focus · HN ↗
No, the only open weight model that really makes sense for me is Qwen3.8-Flash-Next, but it is mainly because I can run it locally with reasonable speed (prefill between 650-1400t/s generation between 22-50t/s depending on number of slots/users I configure).
This is the first model that truly competes with Opus 4.8. I'd say it may be better than Opus 4.6 on programming.
But it is very verbose when it comes to reasoning tokens. The more difficult the task the more verbose it is. Certain very hard tasks that take opus 4.8 400k tokens take Qwen3.8-Flash-Next 2M tokens... But it finishes them.
And what you loose on the generation speed you get back on input caching you can keep on for weeks.
It really depends on the workload.