> Before approving construction, I would want communities of humans to understand why the design works and what justifies confidence in its safety. I would hope that we all would.
Until very recently, I pored over every single line of code Claude generated with razor sharp scrutiny. I would usually catch issues with every response. I'm catching fewer problems these days. Maybe the model is just getting better, and maybe I'm being less careful while under pressure to ship more and more often. But model capability is obviously growing. Even back in March, you could tell it "give me a function that adds two numbers" and you could be 100% confident that it would write the correct function. There was almost no point in looking at the code. Since then, the complexity floor of problems in the category "this is so simple that the model couldn't possibly get it wrong" is rising, and with it, my cognitive surrender to the model is increasing too. Why check it? It's obviously going to be correct.
If AI designs a terawatt fusion plant, then of course we're going to meticulously pore over every detail to ensure safety, reliability, efficiency, whatever. If we find no flaws in the design whatsoever, will we be less careful about the second one? The third one? What about the ten thousandth one? Will "a nuclear fusion plant" become something that models couldn't possibly get wrong?
Terence Tao is arguing that the human involvement in research is crucial, but doesn't convincingly justify why, in my opinion. He says that "human agency is a value of fundamental importance" and that we will need to build "thriving human communities that can understand [AI ideas] together" - not for the sake of correctness, which AI may surpass us on, but for, I guess, the possibility of reclaiming human meaning and purpose. I don't disagree with this at all, but it's not an argument, it's a statement of values. Unfortunately, the stark reality is that if AI does surpass humans, it will become the economically dominant strategy to not verify them and not double check them, but to just do whatever they say. This seems like a great way to raise p(doom). But as the models get better and better, and as I'm scrutinizing Claude's output less and less... I just hope that there are more Terence Taos out there than people like me.
LLMs dont create anything new, if programmers stop reading the code technology will be forever frozen to 2022, no new programming languages, operating systems, concurrency primitives, databases, networking protocols, UI frameworks everything will be based on the training data and future generations will forget about all the primitives we now take for granted.
If someone creates a new programming language/ framework or new better way to do async or whatever, no one will use it because it is not in the training data and it wont take off because everyone is using LLMs. It will be like using the same Lego pieces over and over.
False dichotomy. Chess/Go can still be played between two humans and there is allot of value in that because humans compare each other to other humans, when you see a skillful Grandmaster play you know they are good compared to yourself or the average human, that is why people still watch, play chess/go and train hard to get good. Programming is different because you are creating something not necessarily trying to win a game.
Most programming tasks are exactly like that. Is this agent able to complete this task? Is this agent able to optimize a kernel beyond previous attempts?
Of course some are subjective and that's where progress is harder, like "Is this website pretty?". But for tasks that can be objectively measured, LLMs will go beyond human level, just like with Chess and Go.
That's why RL is so important when training LLMs.
My point is that LLMs depend on training data so the code they produce will be stuck in 2022, no new languages, techniques beyond that because new techniques are not in the training data (at least not enough of it for training because most coders are now using LLMs).
Chess/Go continues to progress because it is primarily a human vs human activity, people will always be learning to play chess and chess will continue to develop.
> Pre training data is in large part synthetic these days
How much of that data can lead to innovation? Can you predict all innovation map it out on paper.
> Computer Chess progress has nothing to do with human vs human activity.
The point is that humans will always be learning chess because it primarily a human vs human activity they will be contributing games to the chess database, unlike with programmers who are stopping to code and only prompting, generating code stuck in 2022.
> AlphaGo Zero used no human game data at all.
Sure, but that instance of AlphaGo is still dependent on its training, its intelligence, so it is a question of is that the best and only way to win a game of Go. Just a few weeks ago, a Go Grandmaster found a way to beat one of the strongest Go AIs.
So a specific instance of an LLM might be the smartest based on what we know and need today but that is not the limit of how far we can go, this is why it is important for humans to always have an intimate connection with the code, math, science, chess etc for progress to continue.
> Sure, but that instance of AlphaGo is still dependent on its training, its intelligence, so it is a question of is that the best and only way to win a game of Go.
If this were true then it would be impossible for these models to ever exceed the top human level as there would exist no training data that allows them to exceed the top human level.
However, despite there being no training data on ability to beat the top humans these models have achieved it.
> this is why it is important for humans to always have an intimate connection with the code, math, science, chess etc for progress to continue.
This is just you wanting to remain relevant rather than actually based on evidence.
> If this were true then it would be impossible for these models to ever exceed the top human level as there would exist no training data that allows them to exceed the top human level.
Of course AI exceeds humans at chess, I never denied that. I am saying because chess is primarily a human vs human game, humans will always be learning and playing chess, their games will add to the chess knowledge base, AI also adds to this knowledge base. But programming is not primarily a human vs human activity so there is a risk programmers will forget how to code and all software will be stuck in 2022 because of the training data, this stifles innovation.
> I am saying because chess is primarily a human vs human game, humans will always be learning and playing chess, their games will add to the chess knowledge base
I guess I'm contesting that idea you are putting forward that the data from the games these humans are playing, which are at a vastly lower level that the top AIs are meaningfully important for helping the AIs to improve at the top level.
Would more people learning their times tables be helpful for top mathematicians in their fields to get better at the frontier of maths? Probably not right. Same applies here.
Would AIs advance at the same rate for the top level of chess in a world where humans completely stopped playing chess vs the world we have today. I would say they would as the human level data is of limited value to the frontier which is dominated by AI and AI game data, you are claiming that it does.
> But programming is not primarily a human vs human activity so there is a risk programmers will forget how to code and all software will be stuck in 2022 because of the training data, this stifles innovation.
Does it? Or will AI be able to run its own experiments and find better/more efficient abstractions that propagate because they are better and this will find its way into training data for future AI.
> I guess I'm contesting that idea you are putting forward that the data from the games these humans are playing, which are at a vastly lower level that the top AIs are meaningfully important for helping the AIs to improve at the top level.
I never said human games are meaningfully important for training AI. Human games are still important for the advancement of chess, maybe Magnus Carlson can learn from games between two Super AIs but most humans still learn from games by humans, Grandmasters are continuously developing the opening, middle-game and end-game systems, adding to the chess knowledge base. Every serious chess player still reviews and study games by prominent Grandmasters, every serious chess player documents their own games, writing down every move. All rated games are recorded and added to the chess database that every player can review and study.
> Human games are still important for the advancement of chess
Are they? Why?
For a human vs human game sure but at the very top level? No of course not because it's all done by AI.
> Only if it is in the training data.
This is trivially not true, as how has AI managed to become better than humans if the knowledge of how to do so never existed in the training data.
We are well past AI can't do X unless X is in the training data. If your claim were true then AI could never surpass human expertise in any field because by definition all the available training data will at best be at the current human frontier and not beyond.
> For a human vs human game sure but at the very top level? No of course not because it's all done by AI.
Glad that you finally agree,this is what I was saying the whole time.
> This is trivially not true, as how has AI managed to become better than humans if the knowledge of how to do so never existed in the training data.
AI can do more work, faster, AI it only needs sufficient compute and data. But that does not mean it is more intelligent than humans, it still uses the same code, algorithms, frameworks, protocols etc etc that are in the training data, sourced from human open source code on the web.
pyridines · · focus · HN ↗
Until very recently, I pored over every single line of code Claude generated with razor sharp scrutiny. I would usually catch issues with every response. I'm catching fewer problems these days. Maybe the model is just getting better, and maybe I'm being less careful while under pressure to ship more and more often. But model capability is obviously growing. Even back in March, you could tell it "give me a function that adds two numbers" and you could be 100% confident that it would write the correct function. There was almost no point in looking at the code. Since then, the complexity floor of problems in the category "this is so simple that the model couldn't possibly get it wrong" is rising, and with it, my cognitive surrender to the model is increasing too. Why check it? It's obviously going to be correct.
If AI designs a terawatt fusion plant, then of course we're going to meticulously pore over every detail to ensure safety, reliability, efficiency, whatever. If we find no flaws in the design whatsoever, will we be less careful about the second one? The third one? What about the ten thousandth one? Will "a nuclear fusion plant" become something that models couldn't possibly get wrong?
Terence Tao is arguing that the human involvement in research is crucial, but doesn't convincingly justify why, in my opinion. He says that "human agency is a value of fundamental importance" and that we will need to build "thriving human communities that can understand [AI ideas] together" - not for the sake of correctness, which AI may surpass us on, but for, I guess, the possibility of reclaiming human meaning and purpose. I don't disagree with this at all, but it's not an argument, it's a statement of values. Unfortunately, the stark reality is that if AI does surpass humans, it will become the economically dominant strategy to not verify them and not double check them, but to just do whatever they say. This seems like a great way to raise p(doom). But as the models get better and better, and as I'm scrutinizing Claude's output less and less... I just hope that there are more Terence Taos out there than people like me.
merelydev · · focus · HN ↗
If someone creates a new programming language/ framework or new better way to do async or whatever, no one will use it because it is not in the training data and it wont take off because everyone is using LLMs. It will be like using the same Lego pieces over and over.
redox99 · · focus · HN ↗
merelydev · · focus · HN ↗
redox99 · · focus · HN ↗
Of course some are subjective and that's where progress is harder, like "Is this website pretty?". But for tasks that can be objectively measured, LLMs will go beyond human level, just like with Chess and Go.
That's why RL is so important when training LLMs.
merelydev · · focus · HN ↗
Chess/Go continues to progress because it is primarily a human vs human activity, people will always be learning to play chess and chess will continue to develop.
redox99 · · focus · HN ↗
Computer Chess progress has nothing to do with human vs human activity. AlphaGo Zero used no human game data at all.
merelydev · · focus · HN ↗
How much of that data can lead to innovation? Can you predict all innovation map it out on paper.
> Computer Chess progress has nothing to do with human vs human activity.
The point is that humans will always be learning chess because it primarily a human vs human activity they will be contributing games to the chess database, unlike with programmers who are stopping to code and only prompting, generating code stuck in 2022.
> AlphaGo Zero used no human game data at all.
Sure, but that instance of AlphaGo is still dependent on its training, its intelligence, so it is a question of is that the best and only way to win a game of Go. Just a few weeks ago, a Go Grandmaster found a way to beat one of the strongest Go AIs.
So a specific instance of an LLM might be the smartest based on what we know and need today but that is not the limit of how far we can go, this is why it is important for humans to always have an intimate connection with the code, math, science, chess etc for progress to continue.
jpleyden98 · · focus · HN ↗
If this were true then it would be impossible for these models to ever exceed the top human level as there would exist no training data that allows them to exceed the top human level.
However, despite there being no training data on ability to beat the top humans these models have achieved it.
> this is why it is important for humans to always have an intimate connection with the code, math, science, chess etc for progress to continue.
This is just you wanting to remain relevant rather than actually based on evidence.
merelydev · · focus · HN ↗
Of course AI exceeds humans at chess, I never denied that. I am saying because chess is primarily a human vs human game, humans will always be learning and playing chess, their games will add to the chess knowledge base, AI also adds to this knowledge base. But programming is not primarily a human vs human activity so there is a risk programmers will forget how to code and all software will be stuck in 2022 because of the training data, this stifles innovation.
jpleyden98 · · focus · HN ↗
I guess I'm contesting that idea you are putting forward that the data from the games these humans are playing, which are at a vastly lower level that the top AIs are meaningfully important for helping the AIs to improve at the top level.
Would more people learning their times tables be helpful for top mathematicians in their fields to get better at the frontier of maths? Probably not right. Same applies here.
Would AIs advance at the same rate for the top level of chess in a world where humans completely stopped playing chess vs the world we have today. I would say they would as the human level data is of limited value to the frontier which is dominated by AI and AI game data, you are claiming that it does.
> But programming is not primarily a human vs human activity so there is a risk programmers will forget how to code and all software will be stuck in 2022 because of the training data, this stifles innovation.
Does it? Or will AI be able to run its own experiments and find better/more efficient abstractions that propagate because they are better and this will find its way into training data for future AI.
merelydev · · focus · HN ↗
I never said human games are meaningfully important for training AI. Human games are still important for the advancement of chess, maybe Magnus Carlson can learn from games between two Super AIs but most humans still learn from games by humans, Grandmasters are continuously developing the opening, middle-game and end-game systems, adding to the chess knowledge base. Every serious chess player still reviews and study games by prominent Grandmasters, every serious chess player documents their own games, writing down every move. All rated games are recorded and added to the chess database that every player can review and study.
jpleyden98 · · focus · HN ↗
Are they? Why?
For a human vs human game sure but at the very top level? No of course not because it's all done by AI.
> Only if it is in the training data.
This is trivially not true, as how has AI managed to become better than humans if the knowledge of how to do so never existed in the training data.
We are well past AI can't do X unless X is in the training data. If your claim were true then AI could never surpass human expertise in any field because by definition all the available training data will at best be at the current human frontier and not beyond.
merelydev · · focus · HN ↗
Glad that you finally agree,this is what I was saying the whole time.
> This is trivially not true, as how has AI managed to become better than humans if the knowledge of how to do so never existed in the training data.
AI can do more work, faster, AI it only needs sufficient compute and data. But that does not mean it is more intelligent than humans, it still uses the same code, algorithms, frameworks, protocols etc etc that are in the training data, sourced from human open source code on the web.