‹ BackHN Continuity

Thread

Training a 4B model to produce 81% faster query plans than Postgres

702 points · 144 comments · polyphilz

  1. refibrillator · · focus · HN ↗
    “81% faster query plans than Postgres”…on an 8 GB dataset that fits entirely in memory, with shared_buffers constrained to a fraction of that, queries warmed before measuring, and read-only SELECTs.

    I would be cautious about over fitting, it’s tough to say if those query plans would really be more optimal than Postgres heuristics at scale and with a bit more realistic OLTP workloads.

    In any case, such is life with profile guided optimization. Many of us appreciate how database workloads can drift over time and with scale.

    Kudos to the author for getting their hands dirty and writing up their experiments.

    1. dnautics · · focus · HN ↗
      I think in principle you could clone your database in prod and at least test to see if your most difficult + common queries are indeed faster after running through the LLM optimizer?
      1. vidarh · · focus · HN ↗
        Frankly, just a general extension to feed a query log to a batch job to do offline optimisation of common actual reoccurring query shapes based on a query log might well be worth it.
      2. locknitpicker · · focus · HN ↗
        > I think in principle you could clone your database in prod and at least test to see if your most difficult + common queries are indeed faster after running through the LLM optimizer?

        That is the responsibility of whoever thought it would be a good idea to write this article. It's their responsibility to show that their idea has merit, and that their results are significant. I mean, don't they have a vested interest in manipulating and cherry-picking their results to inflate their relevance?

        This is why academic papers are peer reviewed.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.