‹ BackHN Continuity

Thread

RIP, vector database

397 points · 116 comments · razin

  1. gopalv · · focus · HN ↗
    > This write amplification is large enough that our efforts to tune indexing throughput have started to hit diminishing returns.

    > don't key on the ANN address. That is precisely the change turbopuffer v3 makes. As you can imagine, it is not a trivial change.

    This is a direct parallel to how Postgres and Mysql built indexes.

    Your design choice went from a Postgres design pattern to a Mysql one. The difference is the reindexing cost vs the lookup cost - Postgres optimized for lookup and Mysql does for indexing on writes. Or more accurately, Postgres was better with good schema design using joins & mysql was optimized for a bad design with less normalization where many indexes exist for the same table.

    Postgres always points an index to a row-id within postgres which is an arbitrary value which changes on each update.

    Mysql, always assuming the storage engine is pluggable, points to the primary index entry and adds an extra indirection to the lookup.

    This means that you point the mysql index to a stable id, so unless you go update the primary key for a row, you won't have to update the indexes for all the attribute lookups you might have made to data.

    I don't do databases any more that much, but the design for NIMBLE file format has a lot of quirks which are relevant to this specific idea (wide tables).

    But the old Uber post about switching from Postgres to Mysql to prevent index amplification[1] is a direct mirror to this post.

    [1] - <a href="https:&#x2F;&#x2F;www.uber.com&#x2F;us&#x2F;en&#x2F;blog&#x2F;postgres-to-mysql-migration&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.uber.com&#x2F;us&#x2F;en&#x2F;blog&#x2F;postgres-to-mysql-migration&#x2F;

    1. FLeXMurphy · · focus · HN ↗
      I find it amusing people started quoting LLM output and are responding to it. Hopefully the original authors end up having the LLM respond back.
      1. 0c3ca83 · · focus · HN ↗
        Many of the commenters on this site are also obviously LLMs. I&#x27;d imagine that quite a few of the entities quoting aren&#x27;t necessarily people. Keep an eye on where they slide mentions of other products that a marketing team would like to promote.
        1. FLeXMurphy · · focus · HN ↗
          Personally I haven&#x27;t seen it too often; the other aspect is that HN is a forum for startups to pitch shit to each other, so this has been happening with or without marketers (f.e. &quot;i&#x27;m working on a similar thing&quot;).

          That LLMs are taking over the comments section is something that was already flagged, and Lobste.rs and others have started solving it by having gated registrations. HN should do this but it is unlikely to until it is too late.

          1. senderista · · focus · HN ↗
            lobste.rs has always been invite-only.
          2. awesome_dude · · focus · HN ↗
            Sorry, how does a gated registration stop someone creating an account, then handing it over to an LLM?

            Serious question, am I misreading what&#x27;s being said?

            1. lukan · · focus · HN ↗
              It stops the automation part. One LLM spammer can be dealt with.
              1. unglaublich · · focus · HN ↗
                One can focus on generating 10 other accounts before it&#x27;s flagged. Now you have 10 others to deal with.
                1. lukan · · focus · HN ↗
                  But that does not work on lobsters (or anywhere else with gated registration), because you cannot just create 10 accounts. You can create 1 account after 1 invite. If you abuse, you loose the invite (and maybe the person inviting you the right to invite others).
                  1. awesome_dude · · focus · HN ↗
                    I think, if i have understood the answers correctly that it&#x27;s stopping a quick proliferation - but not a long slow infiltration?
                    1. leviathant · · focus · HN ↗
                      The spam bots I regularly deal with on reddit are using accounts that are several years old. I&#x27;m not sure if they are hacked, purchased, or something else.
          3. alfiedotwtf · · focus · HN ↗
            I don’t get it though… what’s the point of using an LLM just to comment here? Like what does it gain the person doing it
            1. 0c3ca83 · · focus · HN ↗
              Marketing and influcence.

              For example, consider this prompt -- &quot;Find topics that would be relevant to people interested in &lt;company&gt; and post a topical post that mentions &lt;product&gt;&quot;.

              Also, influencing public opinion on certain topics, such as Israel or Palestine, the current administration, the democrats, the various wars that are ongoing, or AI itself.

            2. porkshoe · · focus · HN ↗
              Some bots are trying to influence conversations on some topic, and they comment on neutral topics to build up karma or whatever.

              Also, some people are weirdos.

              1. qlte · · focus · HN ↗
                Disproportionate number of the &quot;weirdos&quot; category on this particular site vs astroturf marketing bots on Reddit&#x2F;etc.

                I&#x27;ve lost count of how many times I&#x27;ve clicked into a bio from a flagged, obviously LLM written comment to find some variation of &quot;Building new AI tools for agentic devops&quot;...

            3. mikestew · · focus · HN ↗
              There are HN accounts that have been shadow-banned for years, and yet they keep posting despite few people ever seeing their posts. They&#x27;re not LLMs AFAICT, but they just keep posting into the void.

              So why use an LLM for commenting? See above, people are weird.

            4. rapidaneurism · · focus · HN ↗
              Steelmaning it: perhaps people who cannot write to save their life are tired of grammar Nazis correcting them? (Admittedly since the llms I have noticed less and less people being pedantic about these things, perhaps it is like the ill fitting cupboard door signaling that this is handmade)
            5. friendzis · · focus · HN ↗
              &gt; Like what does it gain the person doing it

              Time &#x2F; quantity.

              Pre-LLMs you needed a whole &quot;marketing&quot; agency to astroturf on a meaningful scale. Post-LLMs a single person can run multiple astroturfing campaigns in parallel.

            6. somat · · focus · HN ↗
              LLM&#x27;s when leaned on too heavily, have a tendency to destroy critical thought, or perhaps more accurately replace it.

              It is similar to a spell checker, while spell checkers improve spelling in general, they do not improve a persons spelling ability. Instead acting as a crutch. No need to spell well when the machine will do it for you.

              Some people really like expanding their thoughts via LLM prompt. Some so much it acts like a big crutch, no need to think coherently, the machine will do it for you. So they use the LLM for everything.

              As a related tangent something is messed up in my web browser spell checker, it gives the red squiggles indicating a misspelling, but refuses to give suggested corrections. I would fix it but... My spelling ability has never been better than it is right now.

            7. james_marks · · focus · HN ↗
              Build up karma to expand the reach of future promotional posts.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.