‹ BackHN Continuity

Thread

Astra for Law

589 points · 689 comments · vertigoruntime

  1. ivraatiems · · focus · HN ↗
    I know someone who works in law and deals particularly with an area of US benefits and healthcare law. One of their workflows for lower-level employees at their firm involves taking in documents from healthcare plans and organizations, analyzing them for certain kinds of data, and then importing that data into an internal system they use to analyze and provide guidance on plans. The internal system can contain hundreds of documents for an individual client. All of the documents have the same information (roughly) but in totally diverse formats and styles. Once it's in the system, it's easy to compare and analyze across documents and the research process is much faster.

    They recently bought a Claude subscription and began using Claude to do the initial read of the documents and output JSON they can import into their internal systems. The work still must be reviewed by an attorney - Claude is nowhere near making the kinds of judgments a lawyer would make about this content - but it has increased their throughput from 2-3 documents an hour to 8-10 documents an hour by killing the busy work.

    LLMs have great advantages for this kind of work - but not for decision-making. I just don't see OpenAI ever admitting that.

    (I've left some details intentionally vague because this is a very specific area of law and I don't want my friends to be identified without their consent.)

    1. jorvi · · focus · HN ↗
      LLMs are still absolutely horrid at analyzing PDFs so the results they are getting must be chock full of errors..
      1. Gareth321 · · focus · HN ↗
        Are they? I've had excellent success. The confusing part of this is that there are two types of PDF. The first is a "normal" digital PDF. The second is a scanned PDF. The first can essentially be read like a document. LLMs have no issues with this. It's the second kind of PDF where the constraint becomes the vision capability, and this is very impressive with Astra. I've had no issues with either. I imagine there could be issues with unusually dense and/or misaligned text on scanned PDFs, but I have not tested this.

        The bottom line, though, is that PDF OCR is usually regarded as a solved problem. LLMs won't usually do the recognition itself. It will farm it out to established tools which are very good.

        1. gf000 · · focus · HN ↗
          Well, I would argue about the first part. Even if they contain "native" text that can be extracted, in most cases their order will be messed up and it is often crucial for correct parsing.

          So in many cases the visual way is the only one that works correctly, the textual one is just a shortcut that may be walkable in certain cases.

        2. Otterly99 · · focus · HN ↗
          It depends on what you called solved.

          If the goal is to only extract the unstructured text from the document, it is definitely solved. Extracting a more natural structure like paragraph separation, tables, header, footers (what is referred as document intelligence) is much more complicated and not fully solved, but I would say almost.

          1. jorvi · · focus · HN ↗
            Yup, this.

            It is actually one of my test cases: take the weekly discount PDFs of all the big supermarkets and process each of them, creating a nice table per supermarkt, converting discounts like 1+1 and only listing discounts that are interesting value. I then share that with a bunch of people.

            All models fail this, even the really expensive ones. Even with harness, examples and proper insistent instruction, they'll mix up items and their related discount, which category the item should be in, which page they are on, skipping over items etc.

            As said above, you can OCR it, but at that point you're not processing a PDF, you're processing an image.

            And yes, I know the underlying raw PDF data is messy, but that's why it's such a good test.

        3. QuantumGood · · focus · HN ↗
          The more data in the PDF, the more nines you need in the OCR accuracy. Plenty of 5/S, O/0 and other issues exist at frequencies that cause problems.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.