‹ BackHN Continuity

Thread

Training a 4B model to produce 81% faster query plans than Postgres

702 points · 144 comments · polyphilz

  1. kingjimmy · · focus · HN ↗
    Aren't optimizations suppose to be deterministic?
    1. tintor · · focus · HN ↗
      They are not. Choice among several query plans depends on various summary statistics about the data, which might not be the most recent.
      1. cogman10 · · focus · HN ↗
        Including the input parameters.

        It's not unusual for us to end up with bad query plans because the shape of our data can vary pretty greatly. In many cases, a Foo has 1 Bar. But in some cases, a Foo has a million Bars. That can cause the query optimizer to treat lookups on the bar table as if there are few elements there (causing a scan instead of a seek).

        For the general case, the optimizer gets it right. However, the fringe case is one that causes the entire system to crash. It's a bit akin to how an insertion sort can be faster than quick sort when n is small. The optimizer might make a bad assumption about the size of n which makes it pick an expensive n lookup when log(n) is available (but slower for small n).

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.