Between GLM 5.3 Flash on my legacy Z.ai coding plan, and Qwen 3.8 Flash locally on my DGX Spark-like, I'm barely using my Anthropic/OpenAI subscriptions, likely to cancel them soon.
It's been slow like molasses on the coding plan. I ended up using it more on fireworks. But DeepSeek is so much cheaper because the cache cost is much better.
Cache cost is the primary metric, imo. Most people look at output cost, but you only pay that once. Cache cost you pay every turn and it grows each turn as the context length grows.
monksy · · focus · HN ↗
girvo · · focus · HN ↗
vardalab · · focus · HN ↗
drob518 · · focus · HN ↗