I think what we're seeing with every database company providing new full-text search capabilities is an example of AI coding productivity showing up in the real world.
It started with paradeDB and pg_search <a href="https://www.paradedb.com/blog/introducing-search">https://www.paradedb.com/blog/introducing-search
Timescale has pg_textsearch <a href="https://github.com/timescale/pg_textsearch" rel="nofollow">https://github.com/timescale/pg_textsearch
Neon and Databricks have Lakebase Search <a href="https://docs.databricks.com/aws/en/oltp/projects/lakebase-search" rel="nofollow">https://docs.databricks.com/aws/en/oltp/projects/lakebase-se...
Now PlanetScale.
AFAIK all of these are implementations of the BM25 algorithm. You can just tell an agent to read about BM25 and implement it in your system of choice. Cool to see. Seems like there's still a lot of juice to be squeezed out of how it's architected and integrated into each system, but you can't help but wonder if this will lead to aggressive commodification
BM25 is the easy part. It's probably a dozen lines of code, maybe two. The real work is in the design of the index that enables you to write that dead-simple function -- the in-memory data structures, the on-disk data structures, keeping them in sync, fault tolerance, batching, and a bunch of other things.
If you ask Claude to "implement BM25" you will not get what you want. I see a whole lot of this: people that don't know what they're doing get garbage results out of LLMs.
andrenotgiant · · focus · HN ↗
It started with paradeDB and pg_search <a href="https://www.paradedb.com/blog/introducing-search">https://www.paradedb.com/blog/introducing-search
Timescale has pg_textsearch <a href="https://github.com/timescale/pg_textsearch" rel="nofollow">https://github.com/timescale/pg_textsearch
Neon and Databricks have Lakebase Search <a href="https://docs.databricks.com/aws/en/oltp/projects/lakebase-search" rel="nofollow">https://docs.databricks.com/aws/en/oltp/projects/lakebase-se...
Now PlanetScale.
AFAIK all of these are implementations of the BM25 algorithm. You can just tell an agent to read about BM25 and implement it in your system of choice. Cool to see. Seems like there's still a lot of juice to be squeezed out of how it's architected and integrated into each system, but you can't help but wonder if this will lead to aggressive commodification
Boxxed · · focus · HN ↗
If you ask Claude to "implement BM25" you will not get what you want. I see a whole lot of this: people that don't know what they're doing get garbage results out of LLMs.