‹ BackHN Continuity

Thread

Using Opus 5.5 to discover a new eyewitness record of the dodo

226 points · 79 comments · benbreen

  1. jamienk · · focus · HN ↗
    My dad died and had many many notebooks of his journals with very hard-to-read handwriting. Is it worth the effort to scan all of these so I can feed them in and go to work. Seems like so much minutia is out there, ready to be meta-understood.
    1. komali2 · · focus · HN ↗
      I set up an OCR flow using local models on all my many tens of journals stretching back the last 30 years.

      I would say it's about 80% accurate, which means it's missing enough key words to make a lot of it uselessly unintelligible. I can easily compare the images against text I turn up in a grep which is nice if I'm looking for something.

      Allegedly Claude set up a system for retraining for my handwriting, but it would require me to manually revise several hundred pages by hand so I don't think I'll ever do it.

      <a href="https:&#x2F;&#x2F;github.com&#x2F;508-dev&#x2F;journal-ocr" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;508-dev&#x2F;journal-ocr

      1. jiggawatts · · focus · HN ↗
        Accurate text OCR from bad handwriting is still very much a &quot;bleeding edge&quot; frontier capability that isn&#x27;t practical with local models.

        GPT 6.1 and Gemini Flash 3.8 both do pretty well, their OCR of your sample image is only &quot;wrong&quot; in the sense that the original has typos and they corrected some inadvertently and&#x2F;or filled in gaps where you had &quot;unintelligible&quot; in the canonical text.

        If you have the budget and want the best possible results, you need to run each image through multiple models and then combine the outputs into a final &quot;merge these&quot; prompt. Better scanning helps too, your sample image is rotated and you used a phone in low light. Try a DSLR or a flatbed scanner and process only one page at a time instead of two at once.

        1. jamienk · · focus · HN ↗
          &#x27;run each image through multiple models and then combine the outputs into a final &quot;merge these&quot; prompt&#x27; &lt;&lt; How to do this? This would be an amazing workflow to get documented. This could be the start of a full-service company &quot;send us a bunch of notebooks, get back HIGH QUALITY text version&quot;
          1. komali2 · · focus · HN ↗
            That&#x27;s the workflow in the GitHub repo I linked, I guess the person you&#x27;re replying to didn&#x27;t look. You could just change the calls to be to frontier models instead of local ones if you want.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.