My biggest frustration with the frontier AI companies isn't what they're announcing, but that the announced-thing that exists ~6 months later is severely nerfed to reduce compute spend. It doesn't resemble the demo in any way. For example, this was what the 4o voice capability sounded like in 2024(!) <a href="https://www.youtube.com/watch?v=vgYi3Wr7v_g" rel="nofollow">https://www.youtube.com/watch?v=vgYi3Wr7v_g. What exists today pales in comparison.
How do you afford to do this if you want something resembling the best that's out there right now? The hardware needed to run beefy open source models is like $15,000 to $50,000+ for a robust local multi-GPU rig, and even its performance might lag behind.
I don't do that. I use my 2020 top end gaming PC with its 3090. I just ordered 128GB RAM for it. I'll use a single-chat/slot runtime like Strata or Freetoken. And then I'll be able to run either a blazing fast 8-bit Qwen 3.8 27b or a 4 or 5 bit Qwen 3.8 Flash Next. That's good enough for me for development.
And then I'll have my current gaming rig with less RAM but better CPU and GPU run a reasonable smart tool agent for handling Home Assistant Voice Assist. Downside there: when I'm gaming, no voice assist. That may piss the wife off. We'll see.
At least, this way, I don't have to worry about god damn usage limits. I can knock myself out.
mvkel · · focus · HN ↗
sleight42 · · focus · HN ↗
Thanks. I'll stick with self-hosting.
kelseydh · · focus · HN ↗
sleight42 · · focus · HN ↗
And then I'll have my current gaming rig with less RAM but better CPU and GPU run a reasonable smart tool agent for handling Home Assistant Voice Assist. Downside there: when I'm gaming, no voice assist. That may piss the wife off. We'll see.
At least, this way, I don't have to worry about god damn usage limits. I can knock myself out.