If you have a 24-64GB mac, consider running Qwen3.8 27B locally at night. It's a bit slower to run locally, but if you're sleeping it's less of a problem.
Depending on your memory, you'll need to use the weaker Q4 versions but they still perform well.
It ranks higher than GPT-5.3 Codex (xhigh) or Claude Opus 4.6 (max) so is great for pairing with <a href="https://github.com/kunchenguid/gnhf" rel="nofollow">https://github.com/kunchenguid/gnhf for nightly experimentation, cleanup, or recommendation lists for in the morning.
I still haven't found a use case for Qwen3.8 27B that Qwen 3.6 35b A3b (MoE) is better at. I can get at most 40 tokens/second with 27B, but I get around 90 tokens/second with the MoE and it seems to be a more capable model.
I guess I should still keep experimenting though. Maybe I'm just not using a dense model correctly.
Xeoncross · · focus · HN ↗
Depending on your memory, you'll need to use the weaker Q4 versions but they still perform well.
It ranks higher than GPT-5.3 Codex (xhigh) or Claude Opus 4.6 (max) so is great for pairing with <a href="https://github.com/kunchenguid/gnhf" rel="nofollow">https://github.com/kunchenguid/gnhf for nightly experimentation, cleanup, or recommendation lists for in the morning.
seanmcdirmid · · focus · HN ↗
I guess I should still keep experimenting though. Maybe I'm just not using a dense model correctly.