I think what we're seeing with every database company providing new full-text search capabilities is an example of AI coding productivity showing up in the real world.
It started with paradeDB and pg_search <a href="https://www.paradedb.com/blog/introducing-search">https://www.paradedb.com/blog/introducing-search
Timescale has pg_textsearch <a href="https://github.com/timescale/pg_textsearch" rel="nofollow">https://github.com/timescale/pg_textsearch
Neon and Databricks have Lakebase Search <a href="https://docs.databricks.com/aws/en/oltp/projects/lakebase-search" rel="nofollow">https://docs.databricks.com/aws/en/oltp/projects/lakebase-se...
Now PlanetScale.
AFAIK all of these are implementations of the BM25 algorithm. You can just tell an agent to read about BM25 and implement it in your system of choice. Cool to see. Seems like there's still a lot of juice to be squeezed out of how it's architected and integrated into each system, but you can't help but wonder if this will lead to aggressive commodification
There is a lot of truth to this, but it's also very much down to domain experts being able to do this to move faster.
Planetscale (assuming they used a agentic development practice) will have pulled this off, to the level of performance that they have, because they have a team of very highly experienced Postgres developers. Their knowlage of Postgres internals will have given them the insights needed to steer the models to a plan that used the architecture as described in the post. That's not something a model can do on its own*
World experts + LLMs = moving mountains.
(* we're obviously seeing something a little different from inside the research teams in the labs. They are showing that the models, when you burn the level of tokens only they can, are able to do novel things from the models own insights.)
Seems like a lot of this knowledge was encoded into the blog post. I wonder if given this post and access to a planet scale instance to compare with, how close an agentic agent could get.
The folks at Springbird are giving it a shot: <a href="https://github.com/TeamSpringbird/stannum" rel="nofollow">https://github.com/TeamSpringbird/stannum
I think we can frame it as LLMs materializing existing potential. It seems like there needs to be an underlying potential to tap into, without which, the results could be slop.
> There is a lot of truth to this, but it's also very much down to domain experts being able to do this to move faster.
Yeah, I don't think I could tell Qwen3.8 (my LLM of choice) to study up on bm25 and then implement full text search in the couchdb instances I maintain without studying both bm25 and couchdb internals myself.
I am exploring this exact area of search and analytics for vanilla postgres as replicas. Guess what, the LLM came up with this exact conclusion of using ctids as docids, all by itself. It was surreal for me to read the blog above , when I hit that paragraph about ctids.
andrenotgiant · · focus · HN ↗
It started with paradeDB and pg_search <a href="https://www.paradedb.com/blog/introducing-search">https://www.paradedb.com/blog/introducing-search
Timescale has pg_textsearch <a href="https://github.com/timescale/pg_textsearch" rel="nofollow">https://github.com/timescale/pg_textsearch
Neon and Databricks have Lakebase Search <a href="https://docs.databricks.com/aws/en/oltp/projects/lakebase-search" rel="nofollow">https://docs.databricks.com/aws/en/oltp/projects/lakebase-se...
Now PlanetScale.
AFAIK all of these are implementations of the BM25 algorithm. You can just tell an agent to read about BM25 and implement it in your system of choice. Cool to see. Seems like there's still a lot of juice to be squeezed out of how it's architected and integrated into each system, but you can't help but wonder if this will lead to aggressive commodification
samwillis · · focus · HN ↗
Planetscale (assuming they used a agentic development practice) will have pulled this off, to the level of performance that they have, because they have a team of very highly experienced Postgres developers. Their knowlage of Postgres internals will have given them the insights needed to steer the models to a plan that used the architecture as described in the post. That's not something a model can do on its own*
World experts + LLMs = moving mountains.
(* we're obviously seeing something a little different from inside the research teams in the labs. They are showing that the models, when you burn the level of tokens only they can, are able to do novel things from the models own insights.)
cjonas · · focus · HN ↗
awesome_dude · · focus · HN ↗
dukepiki · · focus · HN ↗
geraneum · · focus · HN ↗
nozzlegear · · focus · HN ↗
Yeah, I don't think I could tell Qwen3.8 (my LLM of choice) to study up on bm25 and then implement full text search in the couchdb instances I maintain without studying both bm25 and couchdb internals myself.
copperx · · focus · HN ↗
ffsm8 · · focus · HN ↗
bddicken · · focus · HN ↗
throwaway7783 · · focus · HN ↗
I am exploring this exact area of search and analytics for vanilla postgres as replicas. Guess what, the LLM came up with this exact conclusion of using ctids as docids, all by itself. It was surreal for me to read the blog above , when I hit that paragraph about ctids.
I am no postgres internals expert.
what · · focus · HN ↗
Or it’s read countless articles on doing the same thing.
throwaway7783 · · focus · HN ↗
awesome_dude · · focus · HN ↗
LLMs are still playing word association - humans have something a little more complex going on where the word has meaning.