Great write-up. People keep saying "we can just write tests" or more recently "we can use formal verification," thinking these are sufficient safeguards we can use and then relegate all the implementation to LLMs. But the fact is that probabilistic guessing machines can't save them. People can't escape the need to actually understand the things they are building.
It does somewhat, or at least used to for a few months. Nowadays it kinda seems that the models have internalized something like TLA and are thinking in it in parallel to thinking in the language they’re writing, so it doesn’t help as much. This is all educated guesses from me, I’ve been telling models to do TLA back in the stone age around February and stopped seeing improvements when telling them to start with specs.
I’m however pretty sure that if you push a good model hard enough on a code base complex enough it’ll find stuff it wouldn’t have otherwise, the Specula folks have some experience with this.
adamddev1 · · focus · HN ↗
nonethewiser · · focus · HN ↗
baq · · focus · HN ↗
I’m however pretty sure that if you push a good model hard enough on a code base complex enough it’ll find stuff it wouldn’t have otherwise, the Specula folks have some experience with this.
<a href="https://github.com/specula-org/Specula" rel="nofollow">https://github.com/specula-org/Specula