For everyone self-hosting models, optimising for cost is an anti-feature. It makes the results worse for no benefit (except a little speed).
What I'd love to see is a harness that deeply optimises for the best results obtainable out of non-frontier models. Many of these have 1M context windows, and most of it remains unused and under utilised in these harnesses, in my opinion.
Speed is actually the thing stopping me using local models more. Qwen 3.8 27b is surprisingly capable, but spends a lot of time/tokens brute forcing problems until it gets it right. The longer task times mean that to fully utilize my attention I need more tasks running concurrently, and the increased context switching just feels bad and leads to me making worse decisions.
nojs · · focus · HN ↗
What I'd love to see is a harness that deeply optimises for the best results obtainable out of non-frontier models. Many of these have 1M context windows, and most of it remains unused and under utilised in these harnesses, in my opinion.
jakkos · · focus · HN ↗
solarkraft · · focus · HN ↗