I don't think any new languages will easily beat the existing ones, simply because of the mass of training data that is available. An LLM will have a much easier time one-shotting Java than even Rust today because it doesn't have to look up much and the ecosystem was stable for many years, so the internet is filled with content that is still up2date. Building a new language (or even altering existing ones) will take much longer to get into the models.
LLMs that have to constantly fix their code and look up libraries or features are much slower, more expensive and more prone to create inefficient and slightly wrong code. Sure, an LLM can just create a new web framework from scratch, but you usually don't want to spend your tokens on that.
I also think that with LLMs, new languages will have more trouble building an ecosystem, because until now, the popular libraries had a "proof of work" that kinda also made sure that they were well-trodden and maintained. Now, anyone can just push their new generated library and we have no clue on how to compare (and tbh reading LLM-generated READMEs is also not pleasant).
> I don't think any new languages will easily beat the existing ones, simply because of the mass of training data that is available.
Seems like a solvable problem though by generating synthetic data that’s guaranteed to be accurate through linters, compilers and tests.
If it’s ~9 figures to train a frontier model, it seems like training on a new language could be a rounding error if it was a priority.
I could see the appeal of a new agentic-friendly language that’s focused on minimizing tokens. The standard library could be massive with no concern of making the language easily readable or learnable.
> Seems like a solvable problem though by generating synthetic data that’s guaranteed to be accurate through linters, compilers and tests.
This makes sense and is very insightful, thank you. But with that solution it seems that only LLM companies will be in the position to create new languages.
mqus · · focus · HN ↗
LLMs that have to constantly fix their code and look up libraries or features are much slower, more expensive and more prone to create inefficient and slightly wrong code. Sure, an LLM can just create a new web framework from scratch, but you usually don't want to spend your tokens on that.
I also think that with LLMs, new languages will have more trouble building an ecosystem, because until now, the popular libraries had a "proof of work" that kinda also made sure that they were well-trodden and maintained. Now, anyone can just push their new generated library and we have no clue on how to compare (and tbh reading LLM-generated READMEs is also not pleasant).
awb · · focus · HN ↗
Seems like a solvable problem though by generating synthetic data that’s guaranteed to be accurate through linters, compilers and tests.
If it’s ~9 figures to train a frontier model, it seems like training on a new language could be a rounding error if it was a priority.
I could see the appeal of a new agentic-friendly language that’s focused on minimizing tokens. The standard library could be massive with no concern of making the language easily readable or learnable.
merelydev · · focus · HN ↗
This makes sense and is very insightful, thank you. But with that solution it seems that only LLM companies will be in the position to create new languages.