Between GLM 5.3 Flash on my legacy Z.ai coding plan, and Qwen 3.8 Flash locally on my DGX Spark-like, I'm barely using my Anthropic/OpenAI subscriptions, likely to cancel them soon.
It's been slow like molasses on the coding plan. I ended up using it more on fireworks. But DeepSeek is so much cheaper because the cache cost is much better.
Cache cost is the primary metric, imo. Most people look at output cost, but you only pay that once. Cache cost you pay every turn and it grows each turn as the context length grows.
monksy · · focus · HN ↗
girvo · · focus · HN ↗
vardalab · · focus · HN ↗
girvo · · focus · HN ↗
The legacy plan I have is so good as to be basically unlimited usage for my workloads, so I’m kind of stuck with it til they stop renewing it haha
drob518 · · focus · HN ↗
pmoriarty · · focus · HN ↗
[1] - <a href="https://www.youtube.com/watch?v=P4dTq4X8bqk" rel="nofollow">https://www.youtube.com/watch?v=P4dTq4X8bqk