RIP, vector database
Thread
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
RIP, vector database
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
OutOfHere · · focus · HN ↗
UPDATE: It loads now, but it didn't when it was first posted. Traffic load on the server does matter.
syndacks · · focus · HN ↗
uproarchat · · focus · HN ↗
alexjplant · · focus · HN ↗
OutOfHere · · focus · HN ↗
WarcrimeActual · · focus · HN ↗
phoghed · · focus · HN ↗
jjgreen · · focus · HN ↗
_kidlike · · focus · HN ↗
swedishPerson1 · · focus · HN ↗
[dead]
throwawy0352 · · focus · HN ↗
wilj · · focus · HN ↗
And what's with the throwaway account for this one comment? Is this becoming reddit with throwaway shills now?
phoghed · · focus · HN ↗
throwawy0352 · · focus · HN ↗
The reason is that I have no account at HN and rarely comment. I create a new account each time because I don't remember, or care about my previous account.
I could have made an account named john2042 and you would not think twice. Instead I let people know upfront what type of account this is. Quite the opposite of what a true shill would do.
jasonmp85 · · focus · HN ↗
[dead]
sreekanth850 · · focus · HN ↗
polynomial · · focus · HN ↗
[dead]
ijidak · · focus · HN ↗
sreekanth850 · · focus · HN ↗
peterpanhead · · focus · HN ↗
sreekanth850 · · focus · HN ↗
_kidlike · · focus · HN ↗
peterpanhead · · focus · HN ↗
blakeashleyjr · · focus · HN ↗
"Updating one vector can move hundreds of attributes and their indexes" is basically Uber's 2016 Postgres write amplification post, but for search. Same fix too: stop pointing indexes at where the row lives.
So ANN becomes a secondary index that points at a doc ID, and vector search now needs a hop to complete. Do clusters keep their own copy of the vectors so the search itself stays local, and only result fetch pays the indirection? Otherwise cold p99 seems like it gets worse.
alfiedotwtf · · focus · HN ↗
ddorian43 · · focus · HN ↗
alfiedotwtf · · focus · HN ↗
peterpanhead · · focus · HN ↗
gopalv · · focus · HN ↗
> don't key on the ANN address. That is precisely the change turbopuffer v3 makes. As you can imagine, it is not a trivial change.
This is a direct parallel to how Postgres and Mysql built indexes.
Your design choice went from a Postgres design pattern to a Mysql one. The difference is the reindexing cost vs the lookup cost - Postgres optimized for lookup and Mysql does for indexing on writes. Or more accurately, Postgres was better with good schema design using joins & mysql was optimized for a bad design with less normalization where many indexes exist for the same table.
Postgres always points an index to a row-id within postgres which is an arbitrary value which changes on each update.
Mysql, always assuming the storage engine is pluggable, points to the primary index entry and adds an extra indirection to the lookup.
This means that you point the mysql index to a stable id, so unless you go update the primary key for a row, you won't have to update the indexes for all the attribute lookups you might have made to data.
I don't do databases any more that much, but the design for NIMBLE file format has a lot of quirks which are relevant to this specific idea (wide tables).
But the old Uber post about switching from Postgres to Mysql to prevent index amplification[1] is a direct mirror to this post.
[1] - <a href="https://www.uber.com/us/en/blog/postgres-to-mysql-migration/" rel="nofollow">https://www.uber.com/us/en/blog/postgres-to-mysql-migration/
phoghed · · focus · HN ↗
TIL I should have been using mysql the whole time
woadwarrior01 · · focus · HN ↗
__s · · focus · HN ↗
There's a bunch of internal types like decimal vs newdecimal, binlog started out statement based until they realized uuid generation is random so added data replication on top. CDC offset started as filepos before GTID was made so offset could survive failover
There were aspects of the design I appreciated (logical slots in postgres have a bunch of drawbacks avoided by just appending to 2nd serial log which has an expiry date instead of tracking clients' offsets), but developing against protocol you learn to not try build a consistent mental model
nixon_why69 · · focus · HN ↗
jgalt212 · · focus · HN ↗
j45 · · focus · HN ↗
This in no way made Postgres any less amazing and cool quietly all those years - if anything it's what really let it step into the forefront the past few years.
As someone who has worked with more than a few databases and MySQL a lot, I'm quietly content learning and beginning with postgres every chance I can get now, to see where and how long "Postgres for everything" can work in a project, simply from there being fewer pieces to build, maintain, integrate, and let the bottlenecks reveal themselves instead of prematurely optimizing for them.
FLeXMurphy · · focus · HN ↗
0c3ca83 · · focus · HN ↗
FLeXMurphy · · focus · HN ↗
That LLMs are taking over the comments section is something that was already flagged, and Lobste.rs and others have started solving it by having gated registrations. HN should do this but it is unlikely to until it is too late.
senderista · · focus · HN ↗
awesome_dude · · focus · HN ↗
Serious question, am I misreading what's being said?
lukan · · focus · HN ↗
unglaublich · · focus · HN ↗
lukan · · focus · HN ↗
awesome_dude · · focus · HN ↗
leviathant · · focus · HN ↗
alfiedotwtf · · focus · HN ↗
0c3ca83 · · focus · HN ↗
porkshoe · · focus · HN ↗
Also, some people are weirdos.
qlte · · focus · HN ↗
I've lost count of how many times I've clicked into a bio from a flagged, obviously LLM written comment to find some variation of "Building new AI tools for agentic devops"...
mikestew · · focus · HN ↗
So why use an LLM for commenting? See above, people are weird.
rapidaneurism · · focus · HN ↗
friendzis · · focus · HN ↗
Time / quantity.
Pre-LLMs you needed a whole "marketing" agency to astroturf on a meaningful scale. Post-LLMs a single person can run multiple astroturfing campaigns in parallel.
somat · · focus · HN ↗
It is similar to a spell checker, while spell checkers improve spelling in general, they do not improve a persons spelling ability. Instead acting as a crutch. No need to spell well when the machine will do it for you.
Some people really like expanding their thoughts via LLM prompt. Some so much it acts like a big crutch, no need to think coherently, the machine will do it for you. So they use the LLM for everything.
As a related tangent something is messed up in my web browser spell checker, it gives the red squiggles indicating a misspelling, but refuses to give suggested corrections. I would fix it but... My spelling ability has never been better than it is right now.
james_marks · · focus · HN ↗
Culonavirus · · focus · HN ↗
bmacho · · focus · HN ↗
matwood · · focus · HN ↗
I feel like we've already reached the point where HN users have discovered 100 of the last 5 LLM commenters. It's the new way to disagree by not having to engage with the argument at all. Just say that a certain sentence structure or word means it's an LLM and move on.
FLeXMurphy · · focus · HN ↗
However, such a policy requires enforcing otherwise its like the rest of guidelines - vapid shit. Maybe now that HN is infused with LLM-Powered Moderation™, dang can do a bit better.
malisper · · focus · HN ↗
You are right that MySQL does better when you have lots of indexes, but I don't think the tradeoff is that the overall Postgres architecture is better with good schema design.
Having secondary indexes point the primary key enables things like undo logging, which obviates the need for vacuums - vacuums being the most painful part of Postgres. On top of that your primary key index will be mostly cached so the cost of the indirection is much smaller than it may first appear
tomnipotent · · focus · HN ↗
> primary key index will be mostly cached so the cost of the indirection is much smaller than it may first appear
Not sure I follow. If it's in-memory you save having to read from disk, but you still have to walk the b-tree to go from PK to data.
barrkel · · focus · HN ↗
sroussey · · focus · HN ↗
barrkel · · focus · HN ↗
tomnipotent · · focus · HN ↗
To be fair this probably covered probably 9x% of production queries, like high-selectivity index-covered filters and sorts (with a limit). Even today most planners will still use an inner loop.
It's been interesting to watch how different planners have evolved over the years. SQL Server 7 in 1998 launched with features that MySQL and Postgres wouldn't catch up to until the late 2010s, but they had other features and quality-of-life improvements that made these edge-case optimizations hardly noticable.
malisper · · focus · HN ↗
Yes, this is true, but they framed this as "the Postgres approach is better when you have a good schema design", but that's not true. There are plenty of ways the MySQL approach is better even when you have a really good schema.
> Not sure I follow. If it's in-memory you save having to read from disk, but you still have to walk the b-tree to go from PK to data.
The point I was trying to make is that going to disk is going to be orders of magnitude slower than doing an in-memory B-tree traversal. Because of that, the cost of doing an extra b-tree traversal to find the page you're looking for is a relatively small cost compared to reading the page in the first place
throwaway7783 · · focus · HN ↗
mattashii · · focus · HN ↗
That's neither here nor there. Heap-oriented tables can have undo logging, too; and index-oriented tables don't strictly require undo logging.
It's just that you're much more likely to want something like undo logging for index-oriented tables, because maintaining uniqueness for rows that are not in-place updated in the index becomes very expensive as more and more versions of the same key value may need to be checked for visibility and liveness, and removing old deleted versions becomes a maintenance hassle, too. It can be done without undo logging, but apparently that wasn't a sufficiently robust (or performant) design.
SigmundA · · focus · HN ↗
Not having true clustered indexes in PG is something I miss coming from MSSQL, it helps performance when the majority of access is always primary index avoid indirection from index lookup then tuple lookup and it also saves space if its the only index.
Derekcai · · focus · HN ↗
drewlanenga · · focus · HN ↗
benesch · · focus · HN ↗
gk1 · · focus · HN ↗
tveita · · focus · HN ↗
anuptalwalkar · · focus · HN ↗
I built a corrective memory layer for our agents which is using filtering, hybrid/ranked fusion search and strongly typed predicates to provide the LLMs context to correct themselves in case of errors.
Small plug, if anyone wants to try it out- <a href="https://polign.com/recall" rel="nofollow">https://polign.com/recall
I struggled quite a bit relying on pure vector DBs, so this is a welcome change. You still need vectors to reach close enough areas to fetch the context though.
Tsarp · · focus · HN ↗
[dead]
ActorNightly · · focus · HN ↗
I see something like taking an auto encoder, cutting it in half to get the discrete latent space, and then mapping documents across that latent space in terms of threshold values.
When storing documents, you just compute their latent space representation, and for each value in the latent space, you have a map of threshold -> document.
When doing a query, you simply map the query into the space, and then filter each value on the activation threshold.
Then when you do a query, that gets mapped to latent space. Then you sequentially check every v
Meanwile do
That way to index documents, you just have to compute their latent space threshold values.
tyromaniac · · focus · HN ↗
thefxperson · · focus · HN ↗
marekgalovic · · focus · HN ↗
We've realized this a long time ago at TopK and built a flexible serverless search engine from scratch. Supports dense/sparse vectors, late interaction, lexical search, indexed regex, filtering, and custom scoring in one query.
- <a href="https://www.topk.io/blog/vector-dbs-are-the-wrong-abstraction-how-we-built-a-new-search-database-from-scratch" rel="nofollow">https://www.topk.io/blog/vector-dbs-are-the-wrong-abstractio... - <a href="https://www.topk.io/blog/topk-embed-v1" rel="nofollow">https://www.topk.io/blog/topk-embed-v1
Tsarp · · focus · HN ↗
orliesaurus · · focus · HN ↗
newtonianrules · · focus · HN ↗
childintime · · focus · HN ↗
Ultimately this system will encompass the whole OS, of course, but the DB might be the best place to start.
tyre · · focus · HN ↗
If so, then no. It is not time for that.
cogman10 · · focus · HN ↗
The best I got is if you are trying to do an old-school style video game asset/save game storage. But even then, the value in just using sqlite or even parquet is really high.
There's so many really good data formats that deciding on a new one at this point seems pretty silly. Particularly because what you sign up for when you make a new one is losing any and all tools that could be used to work with and diagnose that data.
Ostatnigrosh · · focus · HN ↗
childintime · · focus · HN ↗
dymk · · focus · HN ↗
pessimizer · · focus · HN ↗
combobyte · · focus · HN ↗
childintime · · focus · HN ↗
look at the sibling comment, people are already doing this. with time only more people will.
an era is ending. there is a time and a season for everything.
dymk · · focus · HN ↗
eatonphil · · focus · HN ↗
childintime · · focus · HN ↗
Data loss is the obvious concern. Are you perhaps reusing parts of sqlite or postgres?
eatonphil · · focus · HN ↗
nemothekid · · focus · HN ↗
That's interesting. Maybe to decrease latency the LLM could "cache" it's build of it's database and reuse in between instances. It could host this artifact on a "hub" of git trees and then any new use cases that come up, can be added to this git tree. Then it can possibly be reused in different use cases.
childintime · · focus · HN ↗
It tends to make sure you understand the system fully, as no foreign concepts need to be imported and deferred to. That should make your organization run better.
Just like Rust does, btw.
Thx.
bijowo1676 · · focus · HN ↗
childintime · · focus · HN ↗
tschellenbach · · focus · HN ↗
imnotr0b0t · · focus · HN ↗
vhiremath4 · · focus · HN ↗
gravitronic · · focus · HN ↗
turbopuffer is founded by some of the smartest people I ever worked with in past jobs. I strongly doubt they used an LLM in the writing of this article.
_peregrine_ · · focus · HN ↗
croemer · · focus · HN ↗
> Object storage as the source of truth gave the economics, and tiered NVMe SSD/memory caches gave the performance.
mediaman · · focus · HN ↗
croemer · · focus · HN ↗
dolebirchwood · · focus · HN ↗
infamouscow · · focus · HN ↗
nightfly · · focus · HN ↗
Cute heading, followed by wordy opening sentence that feels like it's repeating stuff even when it's not
dolebirchwood · · focus · HN ↗
infamouscow · · focus · HN ↗
ironqcold · · focus · HN ↗
real_faxenoff · · focus · HN ↗
At first, I tried all those popular vector databases and was disappointed with their performance. In the end, the best and fastest solution turned out to be building a multi-database system on SQLite, compiled with everything related to multi-client operations removed. Only exclusive mode was left. Everything is as binary as possible. The index is completely separate — an IVF with pre-training — and is built on the GPU (250K vectors are built, processed, and saved in 4 seconds). Right now, my biggest problem is frequent data changes, and I need to implement optimizations to reduce recalculations.
So far, I haven’t seen any vector database implementations that are heading in the right direction. Maybe only Lancedb looks promising, but it’s too heavy for my needs.
sebastienburel · · focus · HN ↗
[dead]
alexpadula · · focus · HN ↗
xer · · focus · HN ↗
[dead]
croemer · · focus · HN ↗
benesch · · focus · HN ↗
As you can see from the dates we're a few weeks behind our publishing schedule. Hard to pull ourselves away from the perf hacking to write the dev log entries.
croemer · · focus · HN ↗
DevKoala · · focus · HN ↗
pjmlp · · focus · HN ↗
akras14 · · focus · HN ↗
The post itself says ANN becomes “just another” secondary index, and that v3 should make vector search faster along with text and regex.
“RIP, vector-primary index” would have been accurate. “RIP, vector database” is marketing.
andypbl · · focus · HN ↗
[dead]
drpython · · focus · HN ↗
- You create a new problem to solve an already tough problem (information retrieval) - You bill people for storing vectors and searching when they are already paying for information storage and search - You overcomplicate a retrieval problem that has well-defined solutions - All the BS HYPE
The LLM wiki is one good option. At WAIC, I saw a company change the filesystem to store the document index as you save documents so they never go out of sync.
Think harder boys.
magesh_magi1 · · focus · HN ↗
peterpanhead · · focus · HN ↗