‹ BackHN Continuity

Thread

Intelligence per Watt: Measuring Intelligence Efficiency of Local AI

169 points · 65 comments · pythonic_hell

  1. api · · focus · HN ↗
    Unless I misread it, are they saying local GPUs use less energy?

    That’s surprising, almost unbelievable, due to batching. Local is usually not batched.

    1. zozbot234 · · focus · HN ↗
      You can most definitely batch local models and do unattended inference on a 24/7 basis to maximize utilization on local hardware too. The limits are usually set by some combination of memory utilization for KV cache (particularly on small dGPUs) and overall thermals/power limits (particularly on iGPUs with unified RAM/VRAM). (If you're not near thermal limits, the main alternative to batching is to use MTP or speculative decoding in order to raise arithmetic intensity and speed with the same memory utilization. But batching requests is generally viewed as preferable.) Newer models, especially from DeepSeek, do a nice job of reducing KV cache memory impact for any given context length and/or amount of parallel sessions, so batching on local hw really ought to be quite feasible.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.