> This write amplification is large enough that our efforts to tune indexing throughput have started to hit diminishing returns.
> don't key on the ANN address. That is precisely the change turbopuffer v3 makes. As you can imagine, it is not a trivial change.
This is a direct parallel to how Postgres and Mysql built indexes.
Your design choice went from a Postgres design pattern to a Mysql one. The difference is the reindexing cost vs the lookup cost - Postgres optimized for lookup and Mysql does for indexing on writes. Or more accurately, Postgres was better with good schema design using joins & mysql was optimized for a bad design with less normalization where many indexes exist for the same table.
Postgres always points an index to a row-id within postgres which is an arbitrary value which changes on each update.
Mysql, always assuming the storage engine is pluggable, points to the primary index entry and adds an extra indirection to the lookup.
This means that you point the mysql index to a stable id, so unless you go update the primary key for a row, you won't have to update the indexes for all the attribute lookups you might have made to data.
I don't do databases any more that much, but the design for NIMBLE file format has a lot of quirks which are relevant to this specific idea (wide tables).
But the old Uber post about switching from Postgres to Mysql to prevent index amplification[1] is a direct mirror to this post.
> Your design choice went from a Postgres design pattern to a Mysql one. The difference is the reindexing cost vs the lookup cost - Postgres optimized for lookup and Mysql does for indexing on writes. Or more accurately, Postgres was better with good schema design using joins & mysql was optimized for a bad design with less normalization where many indexes exist for the same table
You are right that MySQL does better when you have lots of indexes, but I don't think the tradeoff is that the overall Postgres architecture is better with good schema design.
Having secondary indexes point the primary key enables things like undo logging, which obviates the need for vacuums - vacuums being the most painful part of Postgres. On top of that your primary key index will be mostly cached so the cost of the indirection is much smaller than it may first appear
I think OP is just alluding to the fact that Postgres needs to do less work to go from secondary index to table data, since the tid is a direct pointer to the exact page and slotted entry while MySQL needs a b-tree walk.
> primary key index will be mostly cached so the cost of the indirection is much smaller than it may first appear
Not sure I follow. If it's in-memory you save having to read from disk, but you still have to walk the b-tree to go from PK to data.
MySQL was generally (pre 8) optimized for point queries on primary keys. So rows are stored in the PK index, the PK index is a clustered index. Everything more or less falls out of this.
Sure, I'm only taking about innodb. The join strategies supported before 8.0 (5.7 and lower) were basically nested loop and that's it. Which only really works well for a couple of query patterns: point queries with a handful of joined tables all using indexed FKs, or a big sorted scan over one single table and everything joined to it, with pagination through where clause and sorting all on the primary key, and possibly a straight_join to make sure it doesn't start the execution plan anywhere other than the spine table.
> only really works well for a couple of query patterns
To be fair this probably covered probably 9x% of production queries, like high-selectivity index-covered filters and sorts (with a limit). Even today most planners will still use an inner loop.
It's been interesting to watch how different planners have evolved over the years. SQL Server 7 in 1998 launched with features that MySQL and Postgres wouldn't catch up to until the late 2010s, but they had other features and quality-of-life improvements that made these edge-case optimizations hardly noticable.
> I think OP is just alluding to the fact that Postgres needs to do less work to go from secondary index to table data, since the tid is a direct pointer to the exact page and slotted entry while MySQL needs a b-tree walk.
Yes, this is true, but they framed this as "the Postgres approach is better when you have a good schema design", but that's not true. There are plenty of ways the MySQL approach is better even when you have a really good schema.
> Not sure I follow. If it's in-memory you save having to read from disk, but you still have to walk the b-tree to go from PK to data.
The point I was trying to make is that going to disk is going to be orders of magnitude slower than doing an in-memory B-tree traversal. Because of that, the cost of doing an extra b-tree traversal to find the page you're looking for is a relatively small cost compared to reading the page in the first place
gopalv · · focus · HN ↗
> don't key on the ANN address. That is precisely the change turbopuffer v3 makes. As you can imagine, it is not a trivial change.
This is a direct parallel to how Postgres and Mysql built indexes.
Your design choice went from a Postgres design pattern to a Mysql one. The difference is the reindexing cost vs the lookup cost - Postgres optimized for lookup and Mysql does for indexing on writes. Or more accurately, Postgres was better with good schema design using joins & mysql was optimized for a bad design with less normalization where many indexes exist for the same table.
Postgres always points an index to a row-id within postgres which is an arbitrary value which changes on each update.
Mysql, always assuming the storage engine is pluggable, points to the primary index entry and adds an extra indirection to the lookup.
This means that you point the mysql index to a stable id, so unless you go update the primary key for a row, you won't have to update the indexes for all the attribute lookups you might have made to data.
I don't do databases any more that much, but the design for NIMBLE file format has a lot of quirks which are relevant to this specific idea (wide tables).
But the old Uber post about switching from Postgres to Mysql to prevent index amplification[1] is a direct mirror to this post.
[1] - <a href="https://www.uber.com/us/en/blog/postgres-to-mysql-migration/" rel="nofollow">https://www.uber.com/us/en/blog/postgres-to-mysql-migration/
malisper · · focus · HN ↗
You are right that MySQL does better when you have lots of indexes, but I don't think the tradeoff is that the overall Postgres architecture is better with good schema design.
Having secondary indexes point the primary key enables things like undo logging, which obviates the need for vacuums - vacuums being the most painful part of Postgres. On top of that your primary key index will be mostly cached so the cost of the indirection is much smaller than it may first appear
tomnipotent · · focus · HN ↗
> primary key index will be mostly cached so the cost of the indirection is much smaller than it may first appear
Not sure I follow. If it's in-memory you save having to read from disk, but you still have to walk the b-tree to go from PK to data.
barrkel · · focus · HN ↗
sroussey · · focus · HN ↗
barrkel · · focus · HN ↗
tomnipotent · · focus · HN ↗
To be fair this probably covered probably 9x% of production queries, like high-selectivity index-covered filters and sorts (with a limit). Even today most planners will still use an inner loop.
It's been interesting to watch how different planners have evolved over the years. SQL Server 7 in 1998 launched with features that MySQL and Postgres wouldn't catch up to until the late 2010s, but they had other features and quality-of-life improvements that made these edge-case optimizations hardly noticable.
malisper · · focus · HN ↗
Yes, this is true, but they framed this as "the Postgres approach is better when you have a good schema design", but that's not true. There are plenty of ways the MySQL approach is better even when you have a really good schema.
> Not sure I follow. If it's in-memory you save having to read from disk, but you still have to walk the b-tree to go from PK to data.
The point I was trying to make is that going to disk is going to be orders of magnitude slower than doing an in-memory B-tree traversal. Because of that, the cost of doing an extra b-tree traversal to find the page you're looking for is a relatively small cost compared to reading the page in the first place