Parsing Expression Grammar vs. Regexes: Building Org Parser in Lisp, Export HTML
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Parsing Expression Grammar vs. Regexes: Building Org Parser in Lisp, Export HTML
Unofficial Hacker News client; not affiliated with Y Combinator.
overthenexttwod · · focus · HN ↗
LLMs have changed the equation. It used to take me 2-3 hours to write a parser for a moderately complex grammar. Now, if I hand an LLM a loosely written, BNF-ish grammar and ask for a recursive-descent parser, it finishes the job in five minutes. At this point, writing them by hand is hard to justify. The model does it substantially faster, and lately, often better than I do.
Which makes me wonder: what’s the appeal of PEGs or parser generators now? They used to make sense when hand-writing wasn't practical, but what compelling reasons are left to use them today?
dig1 · · focus · HN ↗
The reason why it became second nature for you is this hand-rolling, manual work. I observed that with math - unless you manually practice problem solving (integrals, differential equations), you might know the mechanics but you'll have hard time solving them. And that knowledge of mechanics will also fade away at some point.
> LLMs have changed the equation. It used to take me 2-3 hours to write a parser for a moderately complex grammar. Now, if I hand an LLM a loosely written, BNF-ish grammar and ask for a recursive-descent parser, it finishes the job in five minutes. At this point, writing them by hand is hard to justify
So, what is the difference between LLMs generated "manual" parsers vs those with parser generators, beside that with LLM that will be done in five minutes + plus some $$ and with parser generators you'll get that for free and under a second?
> Which makes me wonder: what’s the appeal of PEGs or parser generators now? They used to make sense when hand-writing wasn't practical, but what compelling reasons are left to use them today?
The appeal is predictability - use i.e. Bison, feed it with grammar and you'll get always the same output. Not so much with LLMs.