‹ BackHN Continuity

Thread

Best LLM for every budget, updated daily

184 points · 113 comments · terryds

  1. hsnewman · · focus · HN ↗
    I'm sure that local LLM will be far cheaper
    1. qwerpy · · focus · HN ↗
      Wouldn't be so sure, at least not for a GPU-based system. Quick math for a 5090 running at around 500W generating 100 tokens per second (reasonable for Qwen 3.8-27B) is around 2-3 kWh for a million tokens, which is around $0.60 for some mix of off/on peak electricity rates.

      This is in the same ballpark for that same model on openrouter (<a href="https:&#x2F;&#x2F;openrouter.ai&#x2F;qwen&#x2F;qwen3.8-27b" rel="nofollow">https:&#x2F;&#x2F;openrouter.ai&#x2F;qwen&#x2F;qwen3.8-27b), highly dependent on input&#x2F;output mix. Deepseek is a much more capable model that you can&#x27;t run locally on normal hardware, and their rates are insanely cheap ($0.04 &#x2F; $1 per 1M).

      And so far I haven&#x27;t considered the cost of the hardware. I happen to have a gaming PC that can be put to use on inference when not gaming, but given these numbers I don&#x27;t think I would buy new hardware to do inference at home. Unless my math is wrong, it seems you&#x27;re way better off paying for Deepseek than running Qwen or some other locally runnable model yourself. Of course if you have specific privacy requirements or prefer something unique about a particular model you can run locally, the equation changes.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.