‹ BackHN Continuity

Thread

Parsing Expression Grammar vs. Regexes: Building Org Parser in Lisp, Export HTML

126 points · 20 comments · jjba23

  1. overthenexttwod · · focus · HN ↗
    Over the years, I have implemented parsers numerous times, both at work and for side projects, so writing recursive-descent parsers from scratch has become second nature. Once you understand the mechanics, writing a parser by hand is straightforward and offers distinct advantages, particularly much greater flexibility with error handling and reporting. Because of that, I had always viewed PEGs and parser generators as tools primarily for people who couldn’t hand-roll their own because of their circumstances or skill level. (If you look at projects that are neither understaffed nor underskilled you'll find that hand-written recursive-descent parsers are very common: Clang, Go, Rust, TypeScript, Swift and Lua all have hand-written parsers.)

    LLMs have changed the equation. It used to take me 2-3 hours to write a parser for a moderately complex grammar. Now, if I hand an LLM a loosely written, BNF-ish grammar and ask for a recursive-descent parser, it finishes the job in five minutes. At this point, writing them by hand is hard to justify. The model does it substantially faster, and lately, often better than I do.

    Which makes me wonder: what’s the appeal of PEGs or parser generators now? They used to make sense when hand-writing wasn't practical, but what compelling reasons are left to use them today?

    1. genxy · · focus · HN ↗
      Maintenance when you don't have access to an LLM?

      How small of a model can complete the operation you described above? If writing a parser still requires a software forge with 2T of vram and a petabyte of training data, then I still see value in PEGs.

      Maybe the smaller local models, given a structured grammar can use a PEG to generate a parser.

    2. torginus · · focus · HN ↗
      Agreed, and let me add that Pratt parsers are both powerful, and efficient, and the code is fairly easy to understand and write by hand (while being slightly more powerful than recursive descent).

      No LLMs needed, and you can jump straight to coding, and you don't have to go through precedence shenanigans.

    3. UncleEntity · · focus · HN ↗
      > Which makes me wonder: what’s the appeal of PEGs or parser generators now?

      I just had the robots write a PEG parser generator...

      Which can do analysis on the grammars which, I suspect, a hand (or LLM) written one can&#x27;t do so you don&#x27;t end up chasing infinite recursion, dead rules and whatnot. It also got shoehorned into the regex engine (<a href="https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;1210.4992" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;1210.4992) for my toy Java 1.0 compiler to loop back to TFA.

      For my APL interpreter they had to do the &#x27;handwritten&#x27; parser (a Pratt parser a sibling comment brings up) as you need to combine parsing and evaluation since there&#x27;s no way for the parser to tell what it&#x27;s looking at because APL syntax is just weird.

      Horses for courses, as they say.

    4. dig1 · · focus · HN ↗
      &gt; Over the years, I have implemented parsers numerous times, both at work and for side projects, so writing recursive-descent parsers from scratch has become second nature. Once you understand the mechanics, writing a parser by hand is straightforward

      The reason why it became second nature for you is this hand-rolling, manual work. I observed that with math - unless you manually practice problem solving (integrals, differential equations), you might know the mechanics but you&#x27;ll have hard time solving them. And that knowledge of mechanics will also fade away at some point.

      &gt; LLMs have changed the equation. It used to take me 2-3 hours to write a parser for a moderately complex grammar. Now, if I hand an LLM a loosely written, BNF-ish grammar and ask for a recursive-descent parser, it finishes the job in five minutes. At this point, writing them by hand is hard to justify

      So, what is the difference between LLMs generated &quot;manual&quot; parsers vs those with parser generators, beside that with LLM that will be done in five minutes + plus some $$ and with parser generators you&#x27;ll get that for free and under a second?

      &gt; Which makes me wonder: what’s the appeal of PEGs or parser generators now? They used to make sense when hand-writing wasn&#x27;t practical, but what compelling reasons are left to use them today?

      The appeal is predictability - use i.e. Bison, feed it with grammar and you&#x27;ll get always the same output. Not so much with LLMs.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.