‹ BackHN Continuity

Thread

Using LLMs to trace alchemical knowledge and decode 17th century letters

174 points · 44 comments · benbreen

  1. loufe · · focus · HN ↗
    I've been using AI to help with my genealogical research, and it's been fantastic. I am loving pushing the family history back and catching mistakes that someone who many others share as a common ancestor have made.
    1. vintermann · · focus · HN ↗
      I'm very interested in this topic. I have been using AI for genealogy, but not in ways that have been able to push family history back or catch mistakes (yet).

      So how do you use the models more specifically? I have so far used them to write a - in some ways - better genealogy program, with some features I've sorely missed especially with respect to DNA genealogy. I've also tried to use them to help with the tedium of transcribing horrible handwriting, but the results have been so poor and with such a high level of hallucinations that I've not tried that for a while.

      1. grey-area · · focus · HN ↗
        Maybe they’re just more tolerant of hallucinations that you are.
      2. thomasfromcdnjs · · focus · HN ↗
        A couple ideas (that I have used with success);

        - ask it to generate Lean 4 proofs based off centimorgans

        - extract subject, object, subjects from all source texts and create a graph database (maybe someone mentioned a red dog in their oral history, and someone else mentioned a sick dog in a eulogy, ask ai to explore weakly correlated connections, sometimes pays off)

        1. vintermann · · focus · HN ↗
          > ask it to generate Lean 4 proofs based off centimorgans

          That's not confidence inspiring. It'll give you proofs I'm sure, but with what assumptions? You need a full blown model for crossover, and even that would hardly help, because you don't know what assumptions went into the segment match. What size segments are filtered out, how many mismatches are ignored, how interpolation is done, how no-calls are treated etc. Let alone the cM value for services which don't give segments! And then there's that the question for instance Bettinger's shared cM project asks (given this relationship, what shared cM level do you expect to see?) is actually very different from the question we most often ask (given this shared cM level, what relationship should I expect?). When you get into the distant cousin range, those questions can be very different indeed. And we're not even getting into endogamy.

          They say the first principle is that you must not fool yourself, and you're the easiest person to fool. I'm worried that AI can be really powerful self-fooling tools!!

          1. thomasfromcdnjs · · focus · HN ↗
            aha I appreciate the rant, and you obviously know what you are talking about.

            I have used the Lean 4 approach and I wasn't proposing it solves the problem, but it has reduced the search space for me in certain circumstances. At best it's a nice way for agents to keep in their context windows potential relations, I made it create enormous amounts of hypothetical nodes and then visualized it so I could talk with my relatives about potential paths.

            As a side note, I believe, if we had a better DNA site that had great API access (and exposed genomes), I think agents could probably solve genealogy, even as far as complex endogamics.

            (I don't think centimorgans is that useful past 4 generations kind of thing)

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.