‹ BackHN Continuity

Thread

RIP, vector database

397 points · 116 comments · razin

  1. blakeashleyjr · · focus · HN ↗
    This sounds like the Postgres vs. InnoDB argument 10 years later. Postings pointed at physical location (the ANN slot), so every SPFresh rebalance rewrote every index touching that doc. InnoDB solved this by pointing secondary indexes at the PK and eating an extra lookup on read. Curious what that extra lookup costs you when it's an S3 GET instead of a B-tree hop.

    "Updating one vector can move hundreds of attributes and their indexes" is basically Uber's 2016 Postgres write amplification post, but for search. Same fix too: stop pointing indexes at where the row lives.

    So ANN becomes a secondary index that points at a doc ID, and vector search now needs a hop to complete. Do clusters keep their own copy of the vectors so the search itself stays local, and only result fetch pays the indirection? Otherwise cold p99 seems like it gets worse.

    1. alfiedotwtf · · focus · HN ↗
      I haven’t read it, but “ stop pointing indexes at where the row lives” sounded interesting. So if not the row, what does the index point to instead?
      1. ddorian43 · · focus · HN ↗
        It points to the full primary key (which rarely changes).
        1. alfiedotwtf · · focus · HN ↗
          Ah, thanks. But weird though because I would have thought that’s exactly what it was pointing to!
    2. peterpanhead · · focus · HN ↗
      Postgres is a database server, InnoDB is a storage engine. Comparing apples to oranges. MySQL and MariaDB both can utilize different storage engines that bring different things to the table.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.