I sometimes switch mid-session. Depends what I’m doing. Sometimes a model is slow or it’s not giving me what I want and I don’t want to dump the context. But if there’s a natural break, I’ll start a new session with a new model.
Between GLM 5.3 Flash on my legacy Z.ai coding plan, and Qwen 3.8 Flash locally on my DGX Spark-like, I'm barely using my Anthropic/OpenAI subscriptions, likely to cancel them soon.
It's been slow like molasses on the coding plan. I ended up using it more on fireworks. But DeepSeek is so much cheaper because the cache cost is much better.
Cache cost is the primary metric, imo. Most people look at output cost, but you only pay that once. Cache cost you pay every turn and it grows each turn as the context length grows.
monksy · · focus · HN ↗
surgical_fire · · focus · HN ↗
It's an excellent workhorse. When I am running out of my GLM quota I switch GLM-5.3-flash to DS-4.1-flash.
finnjohnsen2 · · focus · HN ↗
surgical_fire · · focus · HN ↗
finnjohnsen2 · · focus · HN ↗
surgical_fire · · focus · HN ↗
I find that it makes coding and review at the same time less prone to errors and cheaper
drob518 · · focus · HN ↗
monksy · · focus · HN ↗
drob518 · · focus · HN ↗
girvo · · focus · HN ↗
vardalab · · focus · HN ↗
girvo · · focus · HN ↗
The legacy plan I have is so good as to be basically unlimited usage for my workloads, so I’m kind of stuck with it til they stop renewing it haha
drob518 · · focus · HN ↗
pmoriarty · · focus · HN ↗
[1] - <a href="https://www.youtube.com/watch?v=P4dTq4X8bqk" rel="nofollow">https://www.youtube.com/watch?v=P4dTq4X8bqk