‹ BackHN Continuity

Thread

Best LLM for every budget, updated daily

184 points · 113 comments · terryds

  1. Xeoncross · · focus · HN ↗
    If you have a 24-64GB mac, consider running Qwen3.8 27B locally at night. It's a bit slower to run locally, but if you're sleeping it's less of a problem.

    Depending on your memory, you'll need to use the weaker Q4 versions but they still perform well.

    It ranks higher than GPT-5.3 Codex (xhigh) or Claude Opus 4.6 (max) so is great for pairing with <a href="https:&#x2F;&#x2F;github.com&#x2F;kunchenguid&#x2F;gnhf" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;kunchenguid&#x2F;gnhf for nightly experimentation, cleanup, or recommendation lists for in the morning.

    1. vardalab · · focus · HN ↗
      I have all sorts of local compute, and local models fairly capable the Frontier models still way more capable&#x2F;faster and local electricity consumption is something else. Good thing it is getting colder around here.

      I often pair them up, and I have an Astra or Sol work as a supervisor and reviewer while Qwen 27B FP8 or Qwen 3.8 Flash Next implements things. I mostly do it as an experiment, just to see what kind of level of autonomy I can get, and they are slow to getting a decent reviewed outcome despite Qwen27B running at 100+ tps and 3.5-4K prefill rates and Qwen3.8 Next at 40 tps and 1-2K prefill. I&#x27;ve been also using similar approach more with OMP, not just the straight Pi harness. And OMP seems to be slower because it has more guardrails. OMP has an interesting feature where you can assign a better LLM as an advisor, wehere it just sort of monitors the progress and injects guidance. And it definitely helps, but one has to be careful. It actually turns out to be expensive if the cache reads are expensive. I learned it the hard way. Where on Fireworks&#x27; API, the cache rates for GLM 5.3 flash are quite a bit more expensive than for DeepSeek, and a simple runs ended up costing me three bucks in oversight. So a better way is to have a Frontier model running a separate tmux pane and just directing it to Wake up every 10 minutes, take a peek at what&#x27;s going on, review the milestones give feedback and then sleep. This turns out to be pretty decent cost saving strategy when quote needs to be stretched. Paradoxically, OpenAI tightening up their quota allowance once they released Astra actually pushed me into all these sorts of experiments, and it&#x27;s actually been interesting. I&#x27;ve been exploring all these smaller flash models, and it&#x27;s been nice. I do like using local LLMs for chore type tasks that are just mostly information gathering, post-session reviews, stuff like that.

      1. tehjoker · · focus · HN ↗
        What do you mean by &quot;local electricity consumption is something else&quot;? Doesn&#x27;t an M3 Pro for example draw about as much power as a bright incandescent lightbulb for a maxed out gpu workload (~100W)? That&#x27;s less than a tenth of what a frontier model will use in the datacenter (which I believe are racks of BlackWell or Vera Lynn GPUs, each using 500W+).
        1. vardalab · · focus · HN ↗
          Because M3 Pro is not a real thing as far as actual agentic workflows go. I run dual R9700 boxes those idle at 150W Because of a Ryzen AM5, and I run Spark boxes, which are decent But still idle at 45-50 watts a pop. My favorite, to be honest, is a 5090 with a Qwen 27B because that one is good for quick hitters and flies, But again, the box itself idles at 140 watts. So it&#x27;s really the idle power that I don&#x27;t like And it&#x27;s too much of an inconvenience to power boxes down and power them back on, so they just end up running and sucking electricity. I am working on getting some sort of a smarter proxy setup where I would give boxes time to wake up and go to the cloud while they&#x27;re waking up. So the whole point of having these things running in the background is that you do want them to be almost always available. So power consumption is definitely an issue.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.