Absolutely! DeepSeek-V4-Flash-0731 has become my daily driver. It's pretty amazing what it can do for what it costs at deepinfra.com (I don't use deepseek as a provider since they train on your data [at least their honest about it]). GLM-5.1 was my daily driver before that and Kimi K2.5 before that.
My primary use is AI coding agent. Its vastly cheaper than Kimi K3 and I haven't found a scenario where I really need Kimi K3 versus smaller models. GLM-5.3 Flash is good but there is series of bugs in the vllm middleware that prevent GLM models from getting all of their reasoning content returned to them that impairs inference quality. A lot of inference providers use vllm which makes it hard to find a good provider for GLM. I've been using friendli.ai but using GLM-5.3 Flash from them is more expensive then using DS V4 Flash from deepinfra.com simply because deepinfra.com is so cheap. The DS V4 Flash cost at together.ai is similar to the GLM-5.3 Flash from friendli.ai or at least that's what I found in my benchmarks a week ago: <a href="https://www.linkedin.com/posts/joshheitzman_i-ran-a-fuller-round-of-benchmarks-on-my-activity-7506040800166723584-WN3v" rel="nofollow">https://www.linkedin.com/posts/joshheitzman_i-ran-a-fuller-r...
4.1 consistently surprises me in capability for the price. And I don't think I'm the only one. It's been dominating the leaderboard at OpenRouter, and I just got an email today from Fireworks saying they were _raising_ the price by about 30%. I'll probably switch, because their infra doesn't support being the highest-cost, but it's still telling.
I tried it a few times and liked the speed, but often found it ended up looping, i.e. repeating the same token sequence (e.g. the same sequence of 5 paragraphs) over and over again until it hit the max output limit. This doesn't end up happening every session, but does every now and then.
My impression of DSv4.1-flash was very positive aside from this. But that was enough for me to stick with GLM-5.3(-flash), which both gave me consistently great results
I was using a vibe coded bare bones harness.
I was wondering if this was normal from DSv4.1-flash, or if its my harnesses fault.
I've had that looping issue with open models too. But never 4.1. I wonder if it's a model + harness combo? But yeah, one loop issue and I'm done with a model forever.
Yeah, that's what I'd been leaning towards.
No mcp, but I'll see if I can reproduce and debug it, since other people don't seem to have that problem as badly as I've experienced it (and the idea of having a nasty bug like that bothers me).
No mcp support.
I'll try copying deepseek harness's basic tool call formats as a starting point.
lwansbrough · · focus · HN ↗
joshheitzman · · focus · HN ↗
kingforaday · · focus · HN ↗
joshheitzman · · focus · HN ↗
pimeys · · focus · HN ↗
I just had like four big sessions going today, paid about $8 in tokens. I see no reason to pay more, this is more than I need for intelligence.
pkulak · · focus · HN ↗
celrod · · focus · HN ↗
My impression of DSv4.1-flash was very positive aside from this. But that was enough for me to stick with GLM-5.3(-flash), which both gave me consistently great results
I was using a vibe coded bare bones harness. I was wondering if this was normal from DSv4.1-flash, or if its my harnesses fault.
pkulak · · focus · HN ↗
pimeys · · focus · HN ↗
So if you use MCP a lot, simplify the params, be more lenient on validation and rework the errors.
It is quite good with shell.
celrod · · focus · HN ↗
No mcp support. I'll try copying deepseek harness's basic tool call formats as a starting point.