‹ BackHN Continuity

Thread

Training a 4B model to produce 81% faster query plans than Postgres

702 points · 144 comments · polyphilz

  1. refibrillator · · focus · HN ↗
    “81% faster query plans than Postgres”…on an 8 GB dataset that fits entirely in memory, with shared_buffers constrained to a fraction of that, queries warmed before measuring, and read-only SELECTs.

    I would be cautious about over fitting, it’s tough to say if those query plans would really be more optimal than Postgres heuristics at scale and with a bit more realistic OLTP workloads.

    In any case, such is life with profile guided optimization. Many of us appreciate how database workloads can drift over time and with scale.

    Kudos to the author for getting their hands dirty and writing up their experiments.

    1. sandeepkd · · focus · HN ↗
      I had the similar feelings, the setup is biased for certain outcomes it feels. At times I feel like that I am in an eternal questioning mode but then again I find it to be a better choice to be critical and skeptical for technology related things.

      Couple things that I found interesting

      1. Inefficiencies/limitations of the query planner in certain cases are known for a long time, its a trade off. This is the reason why hints exists and one can provide their own plan too. DBAs have been doing that for a while now.

      2. A SQL database by design is a resilient unit in itself just dependent on CPU and memory/disk. The availability for the DB is heavily dependent on this factor. Everything built on the top derives their availability and reliability from this. Adding an LLM in between is more cost for sure, question is if its really brining the benefits which are worth the trade off

      1. treebeard901 · · focus · HN ↗
        If the LLM is writing the SQL it might as well build the query plan too.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.