‹ BackHN Continuity

Thread

Using Opus 5.5 to discover a new eyewitness record of the dodo

226 points · 79 comments · benbreen

  1. jamienk · · focus · HN ↗
    My dad died and had many many notebooks of his journals with very hard-to-read handwriting. Is it worth the effort to scan all of these so I can feed them in and go to work. Seems like so much minutia is out there, ready to be meta-understood.
    1. komali2 · · focus · HN ↗
      I set up an OCR flow using local models on all my many tens of journals stretching back the last 30 years.

      I would say it's about 80% accurate, which means it's missing enough key words to make a lot of it uselessly unintelligible. I can easily compare the images against text I turn up in a grep which is nice if I'm looking for something.

      Allegedly Claude set up a system for retraining for my handwriting, but it would require me to manually revise several hundred pages by hand so I don't think I'll ever do it.

      <a href="https:&#x2F;&#x2F;github.com&#x2F;508-dev&#x2F;journal-ocr" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;508-dev&#x2F;journal-ocr

      1. jiggawatts · · focus · HN ↗
        Accurate text OCR from bad handwriting is still very much a &quot;bleeding edge&quot; frontier capability that isn&#x27;t practical with local models.

        GPT 6.1 and Gemini Flash 3.8 both do pretty well, their OCR of your sample image is only &quot;wrong&quot; in the sense that the original has typos and they corrected some inadvertently and&#x2F;or filled in gaps where you had &quot;unintelligible&quot; in the canonical text.

        If you have the budget and want the best possible results, you need to run each image through multiple models and then combine the outputs into a final &quot;merge these&quot; prompt. Better scanning helps too, your sample image is rotated and you used a phone in low light. Try a DSLR or a flatbed scanner and process only one page at a time instead of two at once.

        1. staticman2 · · focus · HN ↗
          &gt; you need to run each image through multiple models and then combine the outputs into a final &quot;merge these&quot; prompt.

          I haven&#x27;t tested this recently but my possibly dated experience is frontier LLMs can&#x27;t figure out which model is correct or incorrect if there&#x27;s disagreement on vision recognition.

          Have you found otherwise?

          (Edit: I see you gave an anecdote about merging terrible results. My experience is with merging overall accurate results).

          1. komali2 · · focus · HN ↗
            They can&#x27;t, which is why I take the average of all results, and then have a final sanity check that considers the context, so if the text contains a bunch of references to &quot;The Disposessed,&quot; and two models record &quot;Shevek thinks Sabul is a propertarian,&quot; and three record &quot;Shrek thinks Sabul is a propertarian,&quot; the final sanity check records Shevek.

            I don&#x27;t think this is viable for critical record OCR. I think the only way to do that is one pass with a frontier model and then a mechanical turk manual review passthrough with good compensation that allows for a slow and methodical approach. Plus of course much better scanning than a phone camera.

            I basically kept rabbit holing this problem and finally settled on &quot;80% and done is better than sitting on this problem for 4 years waiting to have time and equipment for a 99.99% solution.&quot; Crank the pictures between Claude code sessions, run the local LLMs when I&#x27;m asleep, done, now I can free text search years of journals plus I have photo backups now finally of them.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.