I know someone who works in law and deals particularly with an area of US benefits and healthcare law. One of their workflows for lower-level employees at their firm involves taking in documents from healthcare plans and organizations, analyzing them for certain kinds of data, and then importing that data into an internal system they use to analyze and provide guidance on plans. The internal system can contain hundreds of documents for an individual client. All of the documents have the same information (roughly) but in totally diverse formats and styles. Once it's in the system, it's easy to compare and analyze across documents and the research process is much faster.
They recently bought a Claude subscription and began using Claude to do the initial read of the documents and output JSON they can import into their internal systems. The work still must be reviewed by an attorney - Claude is nowhere near making the kinds of judgments a lawyer would make about this content - but it has increased their throughput from 2-3 documents an hour to 8-10 documents an hour by killing the busy work.
LLMs have great advantages for this kind of work - but not for decision-making. I just don't see OpenAI ever admitting that.
(I've left some details intentionally vague because this is a very specific area of law and I don't want my friends to be identified without their consent.)
"Write a python script that breaks down this PDF by X feature" would not hallucinate anything in the PDF. Certainly you could trivially double check that all text in the extracted JSON was in the text layer of the PDF.
How much experience do you have with LLMs exactly? It would be consistent with my experience if Claude stuck in a line of python that just emits a JSON literal with no justification, potentially buried in a large program where an untrained person might not notice it. I don't even trust them if the output consists of structured data paired with source images from the PDF, because I've experienced LLMs fabricating the source rectangles to match the output. I only use tools like this by asking for programs, because as you note LLMs are good at that, and the verification process consists of tool calls to legitimate PDF manipulation tools so I have some confidence everything is above board. Even then I only do this for hobbies, not anything that matters.
Lawyer here. I used to trust Claude as hallucinations are near non-existent now. However for large volume tasks such as due diligence exercises, they still happen.
We also tried Legora's tabular review, there were also numerous halucinated provisions in our due diligence exercise.
And when they do, you can train them or fire them, and they learn not to do it.
LLMs change not a whit, and there's no one to take responsibility for the failure (and thus no way to fix it).
As the new variation on the old theme has it, "A computer can never be held accountable, and so very many people are trying to get them make management decisions."
You can’t train people to never make a mistake, particularly when doing highly repetitive work like this. You must build your systems to account for that regardless.
Yes, exactly. Humans are non-deterministic as well, just in different ways. A tired human can make all sorts of errors for example, regardless of how much training they've had.
That's a recipe for disaster in my experience. I tried it (with Claude) on a simple tabular bank statement PDF, and it transposed two amounts, placinh each against the other's description. And the bot assured me the result was cotrect. The chance of a human checker catching such corruption is low.
Doing similar-ish things with Claude, it's helpful to have something to ground it.
For instance, if you can say:
"Refer to the database schema in x.sql as your source of truth for the database structure we want to import int. Do not invent data, tables or columns that do not exist. Carefully match all output against this database schema and do not create output that doesn't exist if it does not match the schema, simply skip it."
You will end up with a far better result in my experience.
But it gets it right like 99% of the time so human attention can be put towards catching the 1%, not entering data from one table to another and then catching that human’s mistakes.
Few things are. It also doesn’t require standards. “Stamp this diff” culture is everywhere even before AI. A stamp is literally easier than anything else.
Whether that is useful measurement I suppose depends on the circumstances.
Yes, remember that these are effectively random PDFs in various different designs and formats, some of them not editable or even OCR'd.
It took a human attorney 20-30 minutes on average to manually copy-paste data from these PDFs into a spreadsheet (while also fixing any errors they found in the document and re-checking for quality).
Now, the AI copies everything into the spreadsheet in a small amount of time, and then the human reviews it. It takes maybe ~5-7 minutes to scroll to the appropriate pages in the document, read the lines vs the spreadsheet, and make corrections. So you've gone from 2-3 items an hour to ~8-10 items an hour.
Maybe you could pay someone to develop an OCR/ML application that could do this. But that project would never be profitable, even with the time savings. At the cost of a couple Claude subscriptions, it makes sense.
I'm doing some public court records processing for bankruptcy cases (interested mostly to seek out corruption in big national cases), and yes, the "variousness" of random PDFs is exactly the issue. Trying to get the cost for a whole case down to a minimum.
Sample is around 300 court dates, shy under 1k files.
Does the human find enough bugs that they stay on guard, or just rubber stamp everything without really looking at it? It’s hard to stay vigilant when stuff looks plausible.
This is what bag scanners at airports do - the hit rate is so low and the job so boring the software projects fake contraband onto the imagery. Fail to spot the knuckledusters and expect a chat with the manager.
> Maybe you could pay someone to develop an OCR/ML application that could do this. But that project would never be profitable, even with the time savings. At the cost of a couple Claude subscriptions, it makes sense.
A better use of these Claude subscription would be to develop the app (which it can pretty much do at that point) and you could iterate to make the workflow even more efficient than your current one.
Nobody working there has the requisite experience to do this in a reasonable amount of time. These are not particularly tech-savvy folks, Claude use aside.
Yes. And it might not even be worth it, as the AI agents gets cheaper and cheaper.
Keep in mind that the task is fixed, so as the frontier of AI advances, you can switch to a cheaper trailing edge system and still get the same or even better performance for this task.
> These are not particularly tech-savvy folks, Claude use aside.
The difference between a tech-savyy person, and a non-tech-savyy person has always been mostly in the later's head, but this is even more true now that we have pocket assistants who can answer pretty much all of our questions in a language tuned to our level of understanding.
Thanks for sharing the details. Does the attorney check that the AI copied the data accurately? Or is it just assumed to be correct?
Your experience mirrors my own. AI is great for parsing data that can take up a huge amount of time. My only concern is whether or not it’s done accurately. I wouldn’t use it for anything where mistakes cause serious consequences.
ivraatiems · · focus · HN ↗
They recently bought a Claude subscription and began using Claude to do the initial read of the documents and output JSON they can import into their internal systems. The work still must be reviewed by an attorney - Claude is nowhere near making the kinds of judgments a lawyer would make about this content - but it has increased their throughput from 2-3 documents an hour to 8-10 documents an hour by killing the busy work.
LLMs have great advantages for this kind of work - but not for decision-making. I just don't see OpenAI ever admitting that.
(I've left some details intentionally vague because this is a very specific area of law and I don't want my friends to be identified without their consent.)
refurb · · focus · HN ↗
You’ve accurately stated that AI isn’t as rigorous as a trained attorney. Doesn’t that mean that every single datapoint must be confirmed by a human?
How is that quicker than just using a human to read the content and make the call? Data entry savings?
juiceland · · focus · HN ↗
cromka · · focus · HN ↗
margalabargala · · focus · HN ↗
"Write a python script that breaks down this PDF by X feature" would not hallucinate anything in the PDF. Certainly you could trivially double check that all text in the extracted JSON was in the text layer of the PDF.
jeffbee · · focus · HN ↗
terminalcommand · · focus · HN ↗
rayiner · · focus · HN ↗
NateEag · · focus · HN ↗
LLMs change not a whit, and there's no one to take responsibility for the failure (and thus no way to fix it).
As the new variation on the old theme has it, "A computer can never be held accountable, and so very many people are trying to get them make management decisions."
IanCal · · focus · HN ↗
enraged_camel · · focus · HN ↗
NateEag · · focus · HN ↗
But they do learn and improve.
The models don't (yet).
hollerith · · focus · HN ↗
It might be that the models learn and improve in this sense faster than a human child does.
juiceland · · focus · HN ↗
LLM output is nondeterministic and humans take responsibility for the failure the same way they take responsibility of a photocopy is too dark.
NateEag · · focus · HN ↗
You mean, they notice it's too dark right after making it, change the settings, do it again, and give you the good copy?
Because yes, that's my experience of humans.
podocarp · · focus · HN ↗
stevesimmons · · focus · HN ↗
eru · · focus · HN ↗
chrisjj · · focus · HN ↗
egorfine · · focus · HN ↗
mcmcvane · · focus · HN ↗
[dead]
chrisjj · · focus · HN ↗
margalabargala · · focus · HN ↗
chrisjj · · focus · HN ↗
ferngodfather · · focus · HN ↗
For instance, if you can say:
"Refer to the database schema in x.sql as your source of truth for the database structure we want to import int. Do not invent data, tables or columns that do not exist. Carefully match all output against this database schema and do not create output that doesn't exist if it does not match the schema, simply skip it."
You will end up with a far better result in my experience.
Gotta treat it like a child.
chrisjj · · focus · HN ↗
ivraatiems · · focus · HN ↗
But now it's comparing already filled columns on a spreadsheet, not copy-pasting every single thing from an (often uncopyable) PDF.
chrisjj · · focus · HN ↗
ivraatiems · · focus · HN ↗
juiceland · · focus · HN ↗
edmundsauto · · focus · HN ↗
lolakutty · · focus · HN ↗
edmundsauto · · focus · HN ↗
Whether that is useful measurement I suppose depends on the circumstances.
eru · · focus · HN ↗
ivraatiems · · focus · HN ↗
It took a human attorney 20-30 minutes on average to manually copy-paste data from these PDFs into a spreadsheet (while also fixing any errors they found in the document and re-checking for quality).
Now, the AI copies everything into the spreadsheet in a small amount of time, and then the human reviews it. It takes maybe ~5-7 minutes to scroll to the appropriate pages in the document, read the lines vs the spreadsheet, and make corrections. So you've gone from 2-3 items an hour to ~8-10 items an hour.
Maybe you could pay someone to develop an OCR/ML application that could do this. But that project would never be profitable, even with the time savings. At the cost of a couple Claude subscriptions, it makes sense.
barrenko · · focus · HN ↗
Sample is around 300 court dates, shy under 1k files.
At best I'm building a claude skills file.
newAccount2025 · · focus · HN ↗
eru · · focus · HN ↗
And Claude should write down the mistake in a sealed envelope, so it doesn't make into the database.
A review that doesn't find the mistake counts as invalid.
k4tsu · · focus · HN ↗
camdenreslink · · focus · HN ↗
eru · · focus · HN ↗
stymaar · · focus · HN ↗
A better use of these Claude subscription would be to develop the app (which it can pretty much do at that point) and you could iterate to make the workflow even more efficient than your current one.
ivraatiems · · focus · HN ↗
eru · · focus · HN ↗
Keep in mind that the task is fixed, so as the frontier of AI advances, you can switch to a cheaper trailing edge system and still get the same or even better performance for this task.
stymaar · · focus · HN ↗
The difference between a tech-savyy person, and a non-tech-savyy person has always been mostly in the later's head, but this is even more true now that we have pocket assistants who can answer pretty much all of our questions in a language tuned to our level of understanding.
ricky54 · · focus · HN ↗
tkgally · · focus · HN ↗
refurb · · focus · HN ↗
Your experience mirrors my own. AI is great for parsing data that can take up a huge amount of time. My only concern is whether or not it’s done accurately. I wouldn’t use it for anything where mistakes cause serious consequences.