They’ve both been working on this for a while, I’m surprised they still haven’t launched.
Domains, auth, and databases are probably the stickiest software products of all time. I suspect it’s a combo of security concerns and the risk of eroding goodwill among developers (their most important customer base rn) if they move up the stack too quickly. Once you know that the goal is to become cloud vendors, train on/compete with their own customers, and own the entire software stack e2e it’s hard to really feel grateful that they’re dragging it out but planning on doing it anyway.
For those of us working in infra/SaaS outside these companies it’s pretty clear that the only viable path that doesn’t involve getting cannibalized is training your own agents/models. The new coding agent-infra-data business model is a path towards full commoditization and undifferentiated prompting in 3-4 vertically integrated walled gardens. If you ever start making real money on pure software infra they’ll just be able to eat you alive by implementing something similar, training on their own tools rather than yours, and integrating it into their stack (which all of your customers are on already).
Also, you should only ever use $20-200/mo subscriptions on work you want them to train against or you don’t mind automate yourself out of. Think a little bit about what you’re teaching them to do when you use their products, especially if they’re the primary interface you’re working in. People are going to start caring about this a lot when the AI companies feel safe enough to begin the “extinguish” phase of AI coding. Train your own models!
> I do not see enough people discussing this and trying to position their businesses in a way to avoid destruction at the hands of the labs.
What can most even really do about it? Assuming they pull off what they're saying they will: it won't matter what you do, what vertical, what moat you think you have. It will be a concentration of capital and power we've never seen before, and frankly that's terrible for the world.
We need to build collaborative post-training pipelines to augment models with new custom capabilities without catastrophic forgetting/excessive loss of capability in other domains.
Currently what makes this too difficult for anybody but frontier labs is the lack of access to the full distribution of workloads/data they use to RL multiple separate envs/evals without regressing more than they advance in general capabilities. Because they are the primary buyers/builders of that stuff, they have no incentive or reason to allow anybody to replay or resample it but themselves. And it's all so very expensive to do in aggregate so there's not really demand for other product shapes except more openly available, collaborative/bundled RL that you could plug your specific workloads into (who the frontier labs would obviously not want to support with their business).
However, if someone were to build a collaborative rollout platform that you could use to train private workloads (ie create useful IP that doesn't just become profit for other labs/come from what they already can do), the upfront investment would be amortized over the very large number of potential buyers once it gets into the 5-8 figure range, who essentially have no other choice if they want to remain competitive in the technology industry.
Until openai/anthropic ipo the concentration of capital/spending and ndas/loss of employability is too concentrated for the best researchers to really do this without rocking the boat. And a lot of also-ran ai/saas have the same risk due to the lack of capitalized acquisition opportunities or AI vendors to partner with.
If the plan is ultimately to drive you out of business if you ever build anything profitable with their products, and monetize your knowledge without fairly rewarding or explaining their intent to do so, you might as well defect early and build what you inevitably would need anyway.
bananaflag · · focus · HN ↗
copperx · · focus · HN ↗
weitendorf · · focus · HN ↗
Domains, auth, and databases are probably the stickiest software products of all time. I suspect it’s a combo of security concerns and the risk of eroding goodwill among developers (their most important customer base rn) if they move up the stack too quickly. Once you know that the goal is to become cloud vendors, train on/compete with their own customers, and own the entire software stack e2e it’s hard to really feel grateful that they’re dragging it out but planning on doing it anyway.
For those of us working in infra/SaaS outside these companies it’s pretty clear that the only viable path that doesn’t involve getting cannibalized is training your own agents/models. The new coding agent-infra-data business model is a path towards full commoditization and undifferentiated prompting in 3-4 vertically integrated walled gardens. If you ever start making real money on pure software infra they’ll just be able to eat you alive by implementing something similar, training on their own tools rather than yours, and integrating it into their stack (which all of your customers are on already).
Also, you should only ever use $20-200/mo subscriptions on work you want them to train against or you don’t mind automate yourself out of. Think a little bit about what you’re teaching them to do when you use their products, especially if they’re the primary interface you’re working in. People are going to start caring about this a lot when the AI companies feel safe enough to begin the “extinguish” phase of AI coding. Train your own models!
dennisy · · focus · HN ↗
I do not see enough people discussing this and trying to position their businesses in a way to avoid destruction at the hands of the labs.
girvo · · focus · HN ↗
What can most even really do about it? Assuming they pull off what they're saying they will: it won't matter what you do, what vertical, what moat you think you have. It will be a concentration of capital and power we've never seen before, and frankly that's terrible for the world.
weitendorf · · focus · HN ↗
Currently what makes this too difficult for anybody but frontier labs is the lack of access to the full distribution of workloads/data they use to RL multiple separate envs/evals without regressing more than they advance in general capabilities. Because they are the primary buyers/builders of that stuff, they have no incentive or reason to allow anybody to replay or resample it but themselves. And it's all so very expensive to do in aggregate so there's not really demand for other product shapes except more openly available, collaborative/bundled RL that you could plug your specific workloads into (who the frontier labs would obviously not want to support with their business).
However, if someone were to build a collaborative rollout platform that you could use to train private workloads (ie create useful IP that doesn't just become profit for other labs/come from what they already can do), the upfront investment would be amortized over the very large number of potential buyers once it gets into the 5-8 figure range, who essentially have no other choice if they want to remain competitive in the technology industry.
Until openai/anthropic ipo the concentration of capital/spending and ndas/loss of employability is too concentrated for the best researchers to really do this without rocking the boat. And a lot of also-ran ai/saas have the same risk due to the lack of capitalized acquisition opportunities or AI vendors to partner with.
If the plan is ultimately to drive you out of business if you ever build anything profitable with their products, and monetize your knowledge without fairly rewarding or explaining their intent to do so, you might as well defect early and build what you inevitably would need anyway.
you could try to help build that!