> This write amplification is large enough that our efforts to tune indexing throughput have started to hit diminishing returns.
> don't key on the ANN address. That is precisely the change turbopuffer v3 makes. As you can imagine, it is not a trivial change.
This is a direct parallel to how Postgres and Mysql built indexes.
Your design choice went from a Postgres design pattern to a Mysql one. The difference is the reindexing cost vs the lookup cost - Postgres optimized for lookup and Mysql does for indexing on writes. Or more accurately, Postgres was better with good schema design using joins & mysql was optimized for a bad design with less normalization where many indexes exist for the same table.
Postgres always points an index to a row-id within postgres which is an arbitrary value which changes on each update.
Mysql, always assuming the storage engine is pluggable, points to the primary index entry and adds an extra indirection to the lookup.
This means that you point the mysql index to a stable id, so unless you go update the primary key for a row, you won't have to update the indexes for all the attribute lookups you might have made to data.
I don't do databases any more that much, but the design for NIMBLE file format has a lot of quirks which are relevant to this specific idea (wide tables).
But the old Uber post about switching from Postgres to Mysql to prevent index amplification[1] is a direct mirror to this post.
MySQL was often just much easier to start with for a greater number of people for the average complexity and need for the vast majority of projects.
This in no way made Postgres any less amazing and cool quietly all those years - if anything it's what really let it step into the forefront the past few years.
As someone who has worked with more than a few databases and MySQL a lot, I'm quietly content learning and beginning with postgres every chance I can get now, to see where and how long "Postgres for everything" can work in a project, simply from there being fewer pieces to build, maintain, integrate, and let the bottlenecks reveal themselves instead of prematurely optimizing for them.
gopalv · · focus · HN ↗
> don't key on the ANN address. That is precisely the change turbopuffer v3 makes. As you can imagine, it is not a trivial change.
This is a direct parallel to how Postgres and Mysql built indexes.
Your design choice went from a Postgres design pattern to a Mysql one. The difference is the reindexing cost vs the lookup cost - Postgres optimized for lookup and Mysql does for indexing on writes. Or more accurately, Postgres was better with good schema design using joins & mysql was optimized for a bad design with less normalization where many indexes exist for the same table.
Postgres always points an index to a row-id within postgres which is an arbitrary value which changes on each update.
Mysql, always assuming the storage engine is pluggable, points to the primary index entry and adds an extra indirection to the lookup.
This means that you point the mysql index to a stable id, so unless you go update the primary key for a row, you won't have to update the indexes for all the attribute lookups you might have made to data.
I don't do databases any more that much, but the design for NIMBLE file format has a lot of quirks which are relevant to this specific idea (wide tables).
But the old Uber post about switching from Postgres to Mysql to prevent index amplification[1] is a direct mirror to this post.
[1] - <a href="https://www.uber.com/us/en/blog/postgres-to-mysql-migration/" rel="nofollow">https://www.uber.com/us/en/blog/postgres-to-mysql-migration/
phoghed · · focus · HN ↗
TIL I should have been using mysql the whole time
j45 · · focus · HN ↗
This in no way made Postgres any less amazing and cool quietly all those years - if anything it's what really let it step into the forefront the past few years.
As someone who has worked with more than a few databases and MySQL a lot, I'm quietly content learning and beginning with postgres every chance I can get now, to see where and how long "Postgres for everything" can work in a project, simply from there being fewer pieces to build, maintain, integrate, and let the bottlenecks reveal themselves instead of prematurely optimizing for them.