Yes, for fun I tried OVH AI Endpoint and they do not have cache read at all. They bill you every time you send a prompt regardless if you are hit cache or not. One agent session was like 80M input and 300K output and I paid 30$ for that. Or rather I interrupted it and let my local Qwen finish it because cost was getting radicoulous.
_ache_ · · focus · HN ↗
in/out ($) Gemini : 1.5 / 9.0 | Qwen 3.8: 0.15 / 0.47
That is a massive cost reduction.
Refs: <a href="https://www.alibabacloud.com/help/en/model-studio/model-pricing#china-beijing-h4" rel="nofollow">https://www.alibabacloud.com/help/en/model-studio/model-pric... <a href="https://runware.ai/gemini-omni" rel="nofollow">https://runware.ai/gemini-omni
killingtime74 · · focus · HN ↗
tidbeck · · focus · HN ↗
npodbielski · · focus · HN ↗