Using Opus 5.5 to discover a new eyewitness record of the dodo
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Using Opus 5.5 to discover a new eyewitness record of the dodo
Unofficial Hacker News client; not affiliated with Y Combinator.
jamienk · · focus · HN ↗
komali2 · · focus · HN ↗
I would say it's about 80% accurate, which means it's missing enough key words to make a lot of it uselessly unintelligible. I can easily compare the images against text I turn up in a grep which is nice if I'm looking for something.
Allegedly Claude set up a system for retraining for my handwriting, but it would require me to manually revise several hundred pages by hand so I don't think I'll ever do it.
<a href="https://github.com/508-dev/journal-ocr" rel="nofollow">https://github.com/508-dev/journal-ocr
jiggawatts · · focus · HN ↗
GPT 6.1 and Gemini Flash 3.8 both do pretty well, their OCR of your sample image is only "wrong" in the sense that the original has typos and they corrected some inadvertently and/or filled in gaps where you had "unintelligible" in the canonical text.
If you have the budget and want the best possible results, you need to run each image through multiple models and then combine the outputs into a final "merge these" prompt. Better scanning helps too, your sample image is rotated and you used a phone in low light. Try a DSLR or a flatbed scanner and process only one page at a time instead of two at once.
jamienk · · focus · HN ↗
jiggawatts · · focus · HN ↗
Roughly:
Get API keys for multiple vendors or just use OpenRouter (but availability of frontier models tends to be limited). Alternatively, Azure Foundry has everything except Google models, so just two subscriptions is enough.
Run the same prompt and same input image through each of your chosen models.
Then feed the smartest model the original image together with the collected output texts. Use a prompt along the lines of "Merge these attempts to OCR together into an corrected and improved combined version, taking special care to exactly preserve the original's typos, etc, etc..."
You can do this manually, it's just fiddly. It's not hard to automate, most of the "code" is English instructions!
The downside of this approach is the cost: even the "light" frontier models are a few cents per page, which is not so bad until you're doing this 5x or 10x times per page and suddenly scanning a notebook can set you back tens of dollars, more than buying a good novel at a book store.
I picked up on this technique back when GPT 4 was released. People noticed that it could translate ancient Akkadian, but only if you ran the prompt through 4x times and merged. I tried this with a few random samples I found online and the merged translations were generally better than the "official" ones, even thought the individual attempts were unreadable gibberish.
There are already scripts/tools floating around for this!
Look into OpenRouter Fusion, Consensus AI, Multi-Model Debate, etc... or just whip up something yourself.
komali2 · · focus · HN ↗
komali2 · · focus · HN ↗
staticman2 · · focus · HN ↗
I haven't tested this recently but my possibly dated experience is frontier LLMs can't figure out which model is correct or incorrect if there's disagreement on vision recognition.
Have you found otherwise?
(Edit: I see you gave an anecdote about merging terrible results. My experience is with merging overall accurate results).
komali2 · · focus · HN ↗
I don't think this is viable for critical record OCR. I think the only way to do that is one pass with a frontier model and then a mechanical turk manual review passthrough with good compensation that allows for a slow and methodical approach. Plus of course much better scanning than a phone camera.
I basically kept rabbit holing this problem and finally settled on "80% and done is better than sitting on this problem for 4 years waiting to have time and equipment for a 99.99% solution." Crank the pictures between Claude code sessions, run the local LLMs when I'm asleep, done, now I can free text search years of journals plus I have photo backups now finally of them.
onetrickwolf · · focus · HN ↗
I scanned quite a few documents many years ago and it's been fun trying new tools every couple of years to see how good they are getting. I would say it's still not 100% there, but if you have them scanned you can basically just just keep trying and compare the results.
I think the only disappointment in recent years has been that storage has gotten more expensive instead of cheaper. I was waiting for SSDs to get cheap enough to justify moving all my documents to a fast flash array to process and search through them faster, but that doesn't seem like it will happen anytime soon.
nater5000 · · focus · HN ↗
How much data are you working with? lol
Seems like you're doing something professional-grade if this is the case. I imagine most people, like the OP, can basically use whatever machine they have laying around and never be concerned with storage size/speed.
I'm also curious if you've tapped into cloud computing. Not that using the cloud is cheap, but I suppose if I was in a position where I'm concerned with the cost of SSDs for a processing task, then I'd be exploring all of my options and I'd be surprised if the cloud wouldn't be an "easy" solution.