I don't think any new languages will easily beat the existing ones, simply because of the mass of training data that is available. An LLM will have a much easier time one-shotting Java than even Rust today because it doesn't have to look up much and the ecosystem was stable for many years, so the internet is filled with content that is still up2date. Building a new language (or even altering existing ones) will take much longer to get into the models.
LLMs that have to constantly fix their code and look up libraries or features are much slower, more expensive and more prone to create inefficient and slightly wrong code. Sure, an LLM can just create a new web framework from scratch, but you usually don't want to spend your tokens on that.
I also think that with LLMs, new languages will have more trouble building an ecosystem, because until now, the popular libraries had a "proof of work" that kinda also made sure that they were well-trodden and maintained. Now, anyone can just push their new generated library and we have no clue on how to compare (and tbh reading LLM-generated READMEs is also not pleasant).
"LLMs dont create anything new, if programmers stop reading the code technology will be forever frozen to 2022, no new programming languages, operating systems, concurrency primitives, databases, networking protocols, UI frameworks everything will be based on the training data and future generations will forget about all the primitives we now take for granted.
If someone creates a new programming language/ framework or new better way to do async or whatever, no one will use it because it is not in the training data and it wont take off because everyone is using LLMs. It will be like using the same Lego pieces over and over."
Mmmm idk. I used an LLM to write assembler in my made up API so I don’t think what you say is true. The value in LLMs is precisely that they are not just regurgitating training data, but rather inferring concepts extracted from trained data. If a programming language used concepts completely disconnected from existing paradigms you’re probably right… but that would also be quite challenging for humans to learn to use, since by intrinsic construction it would also be widely separated from human language.
Esoteric languages like BrainFuck are esoteric and difficult precisely because they go out of their way to eschew conceptual links to existing languages or paradigms.
So if you invented a new type of esoteric language with arbitrary syntax and strange operators (not sure how you’d do that, exactly, iirc all fundamental binary operators are known) it might be impossible to use with an LLM even if the user manual was in context… but aside from that, languages and the underlying concepts are extremely generalizable.
Yeah. I have my own file formats for music, pixel art, and levels in a little game I’m building with the kids. LLMs are really good at understanding these proprietary formats that exist nowhere else in the world except on my old laptop.
> I used an LLM to write assembler in my made up API so I don’t think what you say is true.
I think you are underestimating the amount of data/context that is required to use a battle tested general purpose language, for both humans and LLMs. The ecosystem requires official docs, stack overflow answers, blog posts, tutorials, existing source code, subreddits, issue/PR discussions of undocumented features, obscure mailing list threads with rare insights, books, youtube videos, benchmarks, tests suits... and the ecosystem of libraries for the language that also need their own official docs, stack overflow answers...
You also need the collective audit by the community and assurance that this language has been used in production by countless others.
In the age of intelligence on tap, any new language not specifically designed for human-only use will be born into the world with an attendant plague of documentation. And languages are shockingly generalizable. There are few things in any language that cannot be done in any other. Even those are computationally equivalent to some other set of instructions.
What might be the case though, is that new languages will use more context to think about until they are well represented in the training set.
The frontier models are exceptionally good at one shoting rust. I find them to be even better at writing rust than python (and the dataset for python must surely be bigger). I think there might be a threshold of enough data for a programming language to make it useful for an agente
I've been playing around with a new DSL, and I've had good luck adding a "describe" feature, eg,
newdsl describe some_feature
This returns the docs for that language feature. So instead of a giant SKILL.md containing the language spec, it exposes a way for an LLM to introspect.
DSLs climb the ladder of abstraction and constrain the solution space at the language level, which IMO provides tighter feedback to both humans and machines.
Tremendous amounts of the Java out there in the training corpus would be based on outdated patterns, especially anything pre-Java 8 or pre-Java 21 (Pattern Matching, Project Loom). I don’t think this negatively impacts the quality of LLM Java code but it sort of calls into question whether this is simply a case of more is better.
My personal experience is that LLMs are quite good at one-shotting complex solutions in both Rust and Go, and tend to be idiomatic.
> I don't think any new languages will easily beat the existing ones, simply because of the mass of training data that is available.
Seems like a solvable problem though by generating synthetic data that’s guaranteed to be accurate through linters, compilers and tests.
If it’s ~9 figures to train a frontier model, it seems like training on a new language could be a rounding error if it was a priority.
I could see the appeal of a new agentic-friendly language that’s focused on minimizing tokens. The standard library could be massive with no concern of making the language easily readable or learnable.
> Seems like a solvable problem though by generating synthetic data that’s guaranteed to be accurate through linters, compilers and tests.
This makes sense and is very insightful, thank you. But with that solution it seems that only LLM companies will be in the position to create new languages.
mqus · · focus · HN ↗
LLMs that have to constantly fix their code and look up libraries or features are much slower, more expensive and more prone to create inefficient and slightly wrong code. Sure, an LLM can just create a new web framework from scratch, but you usually don't want to spend your tokens on that.
I also think that with LLMs, new languages will have more trouble building an ecosystem, because until now, the popular libraries had a "proof of work" that kinda also made sure that they were well-trodden and maintained. Now, anyone can just push their new generated library and we have no clue on how to compare (and tbh reading LLM-generated READMEs is also not pleasant).
[deleted] · · focus · HN ↗
[deleted]
merelydev · · focus · HN ↗
"LLMs dont create anything new, if programmers stop reading the code technology will be forever frozen to 2022, no new programming languages, operating systems, concurrency primitives, databases, networking protocols, UI frameworks everything will be based on the training data and future generations will forget about all the primitives we now take for granted.
If someone creates a new programming language/ framework or new better way to do async or whatever, no one will use it because it is not in the training data and it wont take off because everyone is using LLMs. It will be like using the same Lego pieces over and over."
<a href="https://news.ycombinator.com/item?id=49854536">https://news.ycombinator.com/item?id=49854536
K0balt · · focus · HN ↗
Esoteric languages like BrainFuck are esoteric and difficult precisely because they go out of their way to eschew conceptual links to existing languages or paradigms.
So if you invented a new type of esoteric language with arbitrary syntax and strange operators (not sure how you’d do that, exactly, iirc all fundamental binary operators are known) it might be impossible to use with an LLM even if the user manual was in context… but aside from that, languages and the underlying concepts are extremely generalizable.
christophilus · · focus · HN ↗
merelydev · · focus · HN ↗
I think you are underestimating the amount of data/context that is required to use a battle tested general purpose language, for both humans and LLMs. The ecosystem requires official docs, stack overflow answers, blog posts, tutorials, existing source code, subreddits, issue/PR discussions of undocumented features, obscure mailing list threads with rare insights, books, youtube videos, benchmarks, tests suits... and the ecosystem of libraries for the language that also need their own official docs, stack overflow answers...
You also need the collective audit by the community and assurance that this language has been used in production by countless others.
K0balt · · focus · HN ↗
What might be the case though, is that new languages will use more context to think about until they are well represented in the training set.
reinhash · · focus · HN ↗
williamcotton · · focus · HN ↗
DSLs climb the ladder of abstraction and constrain the solution space at the language level, which IMO provides tighter feedback to both humans and machines.
fcarraldo · · focus · HN ↗
My personal experience is that LLMs are quite good at one-shotting complex solutions in both Rust and Go, and tend to be idiomatic.
awb · · focus · HN ↗
Seems like a solvable problem though by generating synthetic data that’s guaranteed to be accurate through linters, compilers and tests.
If it’s ~9 figures to train a frontier model, it seems like training on a new language could be a rounding error if it was a priority.
I could see the appeal of a new agentic-friendly language that’s focused on minimizing tokens. The standard library could be massive with no concern of making the language easily readable or learnable.
merelydev · · focus · HN ↗
This makes sense and is very insightful, thank you. But with that solution it seems that only LLM companies will be in the position to create new languages.