It costs 20x more than the Chinese models I use. I just don’t need them anymore. Sure I’d use them if forced to for a job, but I don’t pay them outside of that anymore.
And my job won’t even pay for Claude now because it’s so ruinously expensive.
I was under the impression that the cache write fee was added to both the input and output costs (except in cases where the cache write is explicitly disabled via e.g. DISABLE_PROMPT_CACHING). The output becomes part of the context, after all; if they don't (for some reason, due to disaggregated inference perhaps) then I'd expect output tokens get charged both output then input+cache_write on the subsequent completion request.
The pricing model confuses me though (I presume by design, Hanlon be damned).
MisterMunchkin · · focus · HN ↗
And my job won’t even pay for Claude now because it’s so ruinously expensive.
throwa356262 · · focus · HN ↗
Mimo 2.6 Pro: 0.04/0.4/0.87
Sonnet 5.5: 0.2/2/10
Opus 5.5: Sonnet prices times 2
What I dont understand is their cache writes ($2.5). Why is that not covered by input cost?
lcampbell · · focus · HN ↗
The pricing model confuses me though (I presume by design, Hanlon be damned).