‹ BackHN Continuity

Thread

Intelligence per Watt: Measuring Intelligence Efficiency of Local AI

169 points · 65 comments · pythonic_hell

  1. kaziava · · focus · HN ↗
    this is the metric i've been waiting for. we run everything local (ollama + neo4j) for compliance reasons, so 'quality per watt' is literally our budget line. one data point from our setup: qwen2.5:3b on an m2 macbook handles nl-to-cypher for simple graph schemas at ~3-5s per answer, and the energy cost is a rounding error compared to shipping the same queries to a frontier api. the hard part was never the model though, it was parsing pdfs locally without a vision model. would love to see parsing/ocr covered in future benchmarks.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.