My biggest frustration with the frontier AI companies isn't what they're announcing, but that the announced-thing that exists ~6 months later is severely nerfed to reduce compute spend. It doesn't resemble the demo in any way. For example, this was what the 4o voice capability sounded like in 2024(!) <a href="https://www.youtube.com/watch?v=vgYi3Wr7v_g" rel="nofollow">https://www.youtube.com/watch?v=vgYi3Wr7v_g. What exists today pales in comparison.
Moments of brilliance yet unpredictable versus reasonably good (lets be fair: 27b models are pretty damn good now) and predictable.
I'm an old school engineer. I like predictable boring technology that consistently gets the job done over hotness. A 27b model at high quant is "hot" enough.
Speaking in terms of teams: I hate working with hotshots. They ruin teams. And they're usually inconsistent and bad for morale.
How do you afford to do this if you want something resembling the best that's out there right now? The hardware needed to run beefy open source models is like $15,000 to $50,000+ for a robust local multi-GPU rig, and even its performance might lag behind.
I don't do that. I use my 2020 top end gaming PC with its 3090. I just ordered 128GB RAM for it. I'll use a single-chat/slot runtime like Strata or Freetoken. And then I'll be able to run either a blazing fast 8-bit Qwen 3.8 27b or a 4 or 5 bit Qwen 3.8 Flash Next. That's good enough for me for development.
And then I'll have my current gaming rig with less RAM but better CPU and GPU run a reasonable smart tool agent for handling Home Assistant Voice Assist. Downside there: when I'm gaming, no voice assist. That may piss the wife off. We'll see.
At least, this way, I don't have to worry about god damn usage limits. I can knock myself out.
mvkel · · focus · HN ↗
sleight42 · · focus · HN ↗
Thanks. I'll stick with self-hosting.
mvkel · · focus · HN ↗
navigate8310 · · focus · HN ↗
mvkel · · focus · HN ↗
sleight42 · · focus · HN ↗
I'm an old school engineer. I like predictable boring technology that consistently gets the job done over hotness. A 27b model at high quant is "hot" enough.
Speaking in terms of teams: I hate working with hotshots. They ruin teams. And they're usually inconsistent and bad for morale.
calderwoodra · · focus · HN ↗
kelseydh · · focus · HN ↗
sleight42 · · focus · HN ↗
And then I'll have my current gaming rig with less RAM but better CPU and GPU run a reasonable smart tool agent for handling Home Assistant Voice Assist. Downside there: when I'm gaming, no voice assist. That may piss the wife off. We'll see.
At least, this way, I don't have to worry about god damn usage limits. I can knock myself out.