‹ BackHN Continuity

Thread

Lightweight PDF parser with layout, tables, formulas and bounding boxes

81 points · 7 comments · beatrizalmeidaf

Loading the complete thread in the background. This saved snapshot is available now. Refresh

  1. beatrizalmeidaf · · focus · HN ↗

    [dead]

    1. thatcherc · · focus · HN ↗
      This looks fantastic! The table and formula extraction features are especially interesting. My immediate question is: can this be integrated into Zotero? Most of the PDFs I read are research papers and extracting tables and formulas directly from my zotero collection would be super handy.
  2. archeantus · · focus · HN ↗
    Great work, thanks for sharing.
  3. phenomen · · focus · HN ↗
    I currently use <a href="https:&#x2F;&#x2F;github.com&#x2F;firecrawl&#x2F;anydoc" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;firecrawl&#x2F;anydoc in my PDF pipelines. In most cases it performs well. I&#x27;ll test your lib to compare.
    1. coinfused · · focus · HN ↗
      Do you know if this one can also crop the PDF to a region of interest before converting to markdown?
  4. JaumeGar · · focus · HN ↗
    Really nice work, thanks for sharing. The table and formula extraction look great.
  5. antonly · · focus · HN ↗
    I just tried this and the results look very off. I put in the MiniSat[1] Paper, and the result appears catastrophically wrong: <a href="https:&#x2F;&#x2F;imgur.com&#x2F;a&#x2F;i3RHwcQ" rel="nofollow">https:&#x2F;&#x2F;imgur.com&#x2F;a&#x2F;i3RHwcQ

    [1]: <a href="http:&#x2F;&#x2F;minisat.se&#x2F;downloads&#x2F;MiniSat.pdf" rel="nofollow">http:&#x2F;&#x2F;minisat.se&#x2F;downloads&#x2F;MiniSat.pdf

  6. pcthrowaway · · focus · HN ↗
    I literally just automated some transcription of bank and credit card statements, using `pdftotext` from the homebrew `poppler` package, and Papero is unfortunately vastly inferior to pdftotext for this format (banks in Canada seemingly only allow you to get historic statements as PDFs).

    Papero seemingly fails to determine any structure here.. for reference, this is the bash script I used to process credit card statements from Scotiabank. While this doesn&#x27;t generate structured data (which I could certainly do with more work), it generates text files which preserve the layout, making it greppable, which is good enough for my needs right now.

        for f in *.pdf; do
          f=&quot;${f%.pdf}&quot;
          if [[ -f &quot;${f}&quot;.txt ]]; then
            continue
          fi
          # e.g. &quot;Statement Period Mar 7, 2024 - Apr 4, 2024&quot;
          period=$(pdftotext -f 1 -l 1 -nopgbrk -layout -x 331 -y 9 -W 277 -H 15 &quot;${f}.pdf&quot; - | xargs | sed &#x27;s&#x2F;Statement Period &#x2F;&#x2F;&#x27;)
          startdate=${period%%-*}
          enddate=${period#*-}
          newname=$(gdate -d &quot;${startdate}&quot; +%F)_$(gdate -d &quot;${enddate}&quot; +%F)
          mv &quot;${f}.pdf&quot; ${newname}.pdf
          pdftotext -nopgbrk -f 1 -l 1 -x 70 -W 300 -y 276 -H 600 -layout &quot;${newname}.pdf&quot;
          pdftotext -nopgbrk -f 3 -l 3 -x 70 -W 300 -y 170 -H 720 -layout &quot;${newname}.pdf&quot; - &gt;&gt; ${newname}.txt
        done
    
    Extracting the tabular data here should be pretty straightforward as well, I just haven&#x27;t needed to do it.

    LLMs would absolutely eat this kind of task up (writing a script that can turn it into structured data), but I don&#x27;t have local LLMs set up and don&#x27;t really wanna send financial records to big AI.

  7. vitaliinemudrii · · focus · HN ↗

    [dead]

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.