‹ BackHN Continuity

Thread

Qwen 3.8 Omni Flash

346 points · 138 comments · jjcm

  1. _ache_ · · focus · HN ↗
    If the performances are comparable, and there is no evidence it's not.

    in/out ($) Gemini : 1.5 / 9.0 | Qwen 3.8: 0.15 / 0.47

    That is a massive cost reduction.

    Refs: <a href="https:&#x2F;&#x2F;www.alibabacloud.com&#x2F;help&#x2F;en&#x2F;model-studio&#x2F;model-pricing#china-beijing-h4" rel="nofollow">https:&#x2F;&#x2F;www.alibabacloud.com&#x2F;help&#x2F;en&#x2F;model-studio&#x2F;model-pric... <a href="https:&#x2F;&#x2F;runware.ai&#x2F;gemini-omni" rel="nofollow">https:&#x2F;&#x2F;runware.ai&#x2F;gemini-omni

    1. killingtime74 · · focus · HN ↗
      You can&#x27;t just look at the per token cost, but how many tokens it takes on average to do a task. The difference can be massive.
      1. tidbeck · · focus · HN ↗
        Also cache write&#x2F;read cost + cache efficiency.
        1. npodbielski · · focus · HN ↗
          Yes, for fun I tried OVH AI Endpoint and they do not have cache read at all. They bill you every time you send a prompt regardless if you are hit cache or not. One agent session was like 80M input and 300K output and I paid 30$ for that. Or rather I interrupted it and let my local Qwen finish it because cost was getting radicoulous.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.