‹ BackHN Continuity

Thread

Training a 4B model to produce 81% faster query plans than Postgres

702 points · 144 comments · polyphilz

  1. refibrillator · · focus · HN ↗
    “81% faster query plans than Postgres”…on an 8 GB dataset that fits entirely in memory, with shared_buffers constrained to a fraction of that, queries warmed before measuring, and read-only SELECTs.

    I would be cautious about over fitting, it’s tough to say if those query plans would really be more optimal than Postgres heuristics at scale and with a bit more realistic OLTP workloads.

    In any case, such is life with profile guided optimization. Many of us appreciate how database workloads can drift over time and with scale.

    Kudos to the author for getting their hands dirty and writing up their experiments.

    1. dragontamer · · focus · HN ↗
      With a 4B parameter model that probably ran through 8GBs of RAM multiple times to run.

      At a certain point we should seriously talk about CUDA accelerating Postgres instead.

      1. soerxpso · · focus · HN ↗
        I would think it's possible to make it so that the 4B model only needs to be called during an initial phase, and then the same queries it constructed can just be re-used with values replaced, unless you're generating a lot of unique on-the-fly query shapes.
        1. setr · · focus · HN ↗
          With query hints finally being added it’d probably be doable as an extension
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.