Intelligence per Watt: Measuring Intelligence Efficiency of Local AI
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Intelligence per Watt: Measuring Intelligence Efficiency of Local AI
Unofficial Hacker News client; not affiliated with Y Combinator.
scottcha · · focus · HN ↗
Some of the items like model routing, if you do it per request instead of per session, can break down on the cloud from an energy and cost POV since one of the best things you can do for both is to maintain the KV cache which both reduces time component of energy and the quite expensive prefill energy.
I am keen on the future where we have local/cloud hybrid serving which is cache aware. I do think that could be the best use of energy resources for AI.