‹ BackHN Continuity

Thread

Tin: full-text search for Postgres

230 points · 98 comments · ksec

  1. andrenotgiant · · focus · HN ↗
    I think what we're seeing with every database company providing new full-text search capabilities is an example of AI coding productivity showing up in the real world.

    It started with paradeDB and pg_search <a href="https:&#x2F;&#x2F;www.paradedb.com&#x2F;blog&#x2F;introducing-search">https:&#x2F;&#x2F;www.paradedb.com&#x2F;blog&#x2F;introducing-search

    Timescale has pg_textsearch <a href="https:&#x2F;&#x2F;github.com&#x2F;timescale&#x2F;pg_textsearch" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;timescale&#x2F;pg_textsearch

    Neon and Databricks have Lakebase Search <a href="https:&#x2F;&#x2F;docs.databricks.com&#x2F;aws&#x2F;en&#x2F;oltp&#x2F;projects&#x2F;lakebase-search" rel="nofollow">https:&#x2F;&#x2F;docs.databricks.com&#x2F;aws&#x2F;en&#x2F;oltp&#x2F;projects&#x2F;lakebase-se...

    Now PlanetScale.

    AFAIK all of these are implementations of the BM25 algorithm. You can just tell an agent to read about BM25 and implement it in your system of choice. Cool to see. Seems like there&#x27;s still a lot of juice to be squeezed out of how it&#x27;s architected and integrated into each system, but you can&#x27;t help but wonder if this will lead to aggressive commodification

    1. Boxxed · · focus · HN ↗
      BM25 is the easy part. It&#x27;s probably a dozen lines of code, maybe two. The real work is in the design of the index that enables you to write that dead-simple function -- the in-memory data structures, the on-disk data structures, keeping them in sync, fault tolerance, batching, and a bunch of other things.

      If you ask Claude to &quot;implement BM25&quot; you will not get what you want. I see a whole lot of this: people that don&#x27;t know what they&#x27;re doing get garbage results out of LLMs.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.