‹ BackHN Continuity

Thread

Training a 4B model to produce 81% faster query plans than Postgres

702 points · 144 comments · polyphilz

  1. 2001zhaozhao · · focus · HN ↗
    Engineer: "HELP, our production DB is frozen on this query that worked fine before!"

    Infra: "Hmm, let's check... Well would you look at that, it seems like your LLM query planner usually works and produces fast queries, but this time when you changed a variable name to trigger query rebuild, it happened to hallucinate and miss an index, would you mind re-running the LLM a few times until you get a faster query?"

    1. malisper · · focus · HN ↗
      Funnily enough, you could replace "LLM query planner" with just "query planner" and this comment would still hold true
      1. Tanjreeve · · focus · HN ↗
        That bug is fixable and verifiable. The LLM you cross your fingers till the next time the same thing happens.
        1. egeozcan · · focus · HN ↗
          > fixable and verifiable

          By people with a specific skill set. LLMs generation can also be fixed and verified by people with a certain skill set, and non-deterministic computing doesn't automatically mean unpredictable. When people say that the LLMs are a black box, it means unpredictability in unknown situations.

          You do structured output, input validation, output validation, lower temperature, limit decisions, RL, etc. to increase predictability to near certainty. It's just statistics after all. Or you can as well generate the code to do the job.

          It's just that the required skill set is a different one to do those things, and unusual in the context of DB administration.

          1. tempfile · · focus · HN ↗
            None of the things you mention are guaranteed to increase the probability of correctness. You can run the LLM output through as many deterministic programs as you like, but "the query plan runs in acceptable time" is not something you can verify with such a tool. Nobody knows how the LLM does it, so they cannot know how to make the LLM do it better.
            1. malisper · · focus · HN ↗
              Even if the query plan was not generated by an llm, you can't verify it will run in an acceptable time. This is one of the biggest unsolved problems in databases
              1. Tanjreeve · · focus · HN ↗
                So what is an LLM bringing to the table if it's not just a fun experiment?
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.