‹ BackHN Continuity

Thread

Tin: full-text search for Postgres

230 points · 98 comments · ksec

  1. andrenotgiant · · focus · HN ↗
    I think what we're seeing with every database company providing new full-text search capabilities is an example of AI coding productivity showing up in the real world.

    It started with paradeDB and pg_search <a href="https:&#x2F;&#x2F;www.paradedb.com&#x2F;blog&#x2F;introducing-search">https:&#x2F;&#x2F;www.paradedb.com&#x2F;blog&#x2F;introducing-search

    Timescale has pg_textsearch <a href="https:&#x2F;&#x2F;github.com&#x2F;timescale&#x2F;pg_textsearch" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;timescale&#x2F;pg_textsearch

    Neon and Databricks have Lakebase Search <a href="https:&#x2F;&#x2F;docs.databricks.com&#x2F;aws&#x2F;en&#x2F;oltp&#x2F;projects&#x2F;lakebase-search" rel="nofollow">https:&#x2F;&#x2F;docs.databricks.com&#x2F;aws&#x2F;en&#x2F;oltp&#x2F;projects&#x2F;lakebase-se...

    Now PlanetScale.

    AFAIK all of these are implementations of the BM25 algorithm. You can just tell an agent to read about BM25 and implement it in your system of choice. Cool to see. Seems like there&#x27;s still a lot of juice to be squeezed out of how it&#x27;s architected and integrated into each system, but you can&#x27;t help but wonder if this will lead to aggressive commodification

    1. samwillis · · focus · HN ↗
      There is a lot of truth to this, but it&#x27;s also very much down to domain experts being able to do this to move faster.

      Planetscale (assuming they used a agentic development practice) will have pulled this off, to the level of performance that they have, because they have a team of very highly experienced Postgres developers. Their knowlage of Postgres internals will have given them the insights needed to steer the models to a plan that used the architecture as described in the post. That&#x27;s not something a model can do on its own*

      World experts + LLMs = moving mountains.

      (* we&#x27;re obviously seeing something a little different from inside the research teams in the labs. They are showing that the models, when you burn the level of tokens only they can, are able to do novel things from the models own insights.)

      1. cjonas · · focus · HN ↗
        Seems like a lot of this knowledge was encoded into the blog post. I wonder if given this post and access to a planet scale instance to compare with, how close an agentic agent could get.
        1. awesome_dude · · focus · HN ↗
          The easiest way to find that out is to TIAS
        2. dukepiki · · focus · HN ↗
          The folks at Springbird are giving it a shot: <a href="https:&#x2F;&#x2F;github.com&#x2F;TeamSpringbird&#x2F;stannum" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;TeamSpringbird&#x2F;stannum
      2. geraneum · · focus · HN ↗
        I think we can frame it as LLMs materializing existing potential. It seems like there needs to be an underlying potential to tap into, without which, the results could be slop.
      3. nozzlegear · · focus · HN ↗
        &gt; There is a lot of truth to this, but it&#x27;s also very much down to domain experts being able to do this to move faster.

        Yeah, I don&#x27;t think I could tell Qwen3.8 (my LLM of choice) to study up on bm25 and then implement full text search in the couchdb instances I maintain without studying both bm25 and couchdb internals myself.

        1. copperx · · focus · HN ↗
          Oh, you could.
          1. ffsm8 · · focus · HN ↗
            Yeah, he totally could. Wherever the output is actually useable or a dumbsterfire would be opaque for him, however
      4. bddicken · · focus · HN ↗
        THIS++
      5. throwaway7783 · · focus · HN ↗
        LLMs are becoming a world expert in everything.

        I am exploring this exact area of search and analytics for vanilla postgres as replicas. Guess what, the LLM came up with this exact conclusion of using ctids as docids, all by itself. It was surreal for me to read the blog above , when I hit that paragraph about ctids.

        I am no postgres internals expert.

        1. what · · focus · HN ↗
          &gt;all by itself

          Or it’s read countless articles on doing the same thing.

          1. throwaway7783 · · focus · HN ↗
            I mean, that goes without saying for LLMs. It is an approximation of human knowledge after all.
            1. awesome_dude · · focus · HN ↗
              In the same way that a google search is .

              LLMs are still playing word association - humans have something a little more complex going on where the word has meaning.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.