Vote on which of Hacker News' challenges for AI have been met
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Vote on which of Hacker News' challenges for AI have been met
Unofficial Hacker News client; not affiliated with Y Combinator.
ben_w · · focus · HN ↗
Sure, sure, what LLMs make still isn't "efficient bug-free code": my prediction is falsified because while LLMs can write and train new models with machine learning, ML is fundamentally not advanced enough to throw arbitraty new tasks at like this.
FabCH · · focus · HN ↗
In your case, the comment you link to says „business tasks“ and you expanded it now to „arbitrary new tasks“. Those are not the same. An LLM today sure can do many many many business-speak conversion tasks.
tripleee · · focus · HN ↗
Not reliably, and not without supervision. That's the main point. I'm trying really hard to figure out a workflow that doesn't require me to review the code and I just don't see how it's possible (yet)
You either need a comprehensive test suite (which requires understanding the code in order to create) or you need to review the actual implementation code to make sure it does the right thing
user43928 · · focus · HN ↗
I've stopped reviewing the code in my mobile app project months ago. I now only look at files changed and lines count in MRs. Functionality is best verified via manual QA testing.
I know that people here are going to doubt the quality of my project and say that it is impossible, but they are clueless and have evidently not build a project in this way. Experiences from eg. corporate backend work are hardly relevant.
It is clear to me that the fewer consequential mistakes people find during code review, their attention to code reviews is going to go down, to the point of also skipping them.
I expect that for most development, not reviewing the code will be the standard by March of next year. Only critical code like authentication will be reviewed.
suddenlybananas · · focus · HN ↗
user43928 · · focus · HN ↗
It's a paid app, and after 6 months of work I expect to publish it this month.
In other words, you will have to take my word for its quality.
tripleee · · focus · HN ↗
user43928 · · focus · HN ↗
One thing that I do is have each change reviewed by another model. If I implement with Codex, I would have Claude do the review, or the other way around.
I did not benchmark this against a review from a subagent with the same model, so I don't know if that in particular helps.