It's becoming more clear that for big enterprises to really adopt AI, they need to use open models. Especially if they want to own their own intelligence, which they should.
I spent the last 2 days building basic AI agents to automate some mundane supply chain workflows for a large company. Those seemingly boring workflows had bank statements, supplier IDs and other sensitive information.
For me it was all alarm bells, there is no way they can afford to give closed models access to this data. I was compelled to figure out an open model based solution for them, which made me realize that this is probably the only way for enterprises going forward.
As long as you don't have "realtime" workloads, owning the GPUs quickly becomes the economical option. The main cost problems is in e.g. chat applications where the workload is spikey, and users expect an near-instant response, for which you need to scale the GPUs to the highest spikes of the workload.
You really don't need to be large. $100k can buy you a lot of compute and it's less than hiring an engineer. With that kind of money you can build an LLM server for a dozen people.
One engineer's salary to accelerate a team of twelve is so cheap you can't afford not to.
Open models on-prem is the future, not a single doubt in my mind.
I'm old enough to remember when my company had everything on-prem (both analytical and operational databases and servers) due to cost and security concerns. Nowadays we have everything on GCP.
The biggest problem we had with on-prem was maintenance as it took a lot of staff and time to ensure decent reliability.
Zaraif13 · · focus · HN ↗
I spent the last 2 days building basic AI agents to automate some mundane supply chain workflows for a large company. Those seemingly boring workflows had bank statements, supplier IDs and other sensitive information.
For me it was all alarm bells, there is no way they can afford to give closed models access to this data. I was compelled to figure out an open model based solution for them, which made me realize that this is probably the only way for enterprises going forward.
tomp · · focus · HN ↗
So if you're running open models on AWS GPUs, you might as well run Claude (which AWS supports, and doesn't share any data with Anthropic).
Same with Azure/OpenAI.
schnitzelstoat · · focus · HN ↗
hobofan · · focus · HN ↗
unrented7977 · · focus · HN ↗
One engineer's salary to accelerate a team of twelve is so cheap you can't afford not to.
Open models on-prem is the future, not a single doubt in my mind.
schnitzelstoat · · focus · HN ↗
The biggest problem we had with on-prem was maintenance as it took a lot of staff and time to ensure decent reliability.