Better Vector Search for Long Documents: Chunking Inside Manticore Search
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Better Vector Search for Long Documents: Chunking Inside Manticore Search
Unofficial Hacker News client; not affiliated with Y Combinator.
entrope · · focus · HN ↗
For languages like English, there's also usually a lot of redundancy within a text, so 512 tokens might not give a very clear indication of the context. Lots of documents have similar introductions (like "#include <foo.h>\n") that make short contexts and truncation particularly harmful.
Also, "Nothing in the document past that point can ever be retrieved, and nothing anywhere told you." This is user-hostile behavior, even if they didn't want to admit to users that the auto-embedding support was poor.
Finally, the paragraph later on about truncation being "what you already have" reads like Claude talking to the developer, not like a vendor talking to users. But sure, maybe this is a good default for a database searching page titles, chat logs and Xeets?
PaulHoule · · focus · HN ↗
The answer is "make the context window as large as you reasonably can" because you have to have enough (con)text in the window for the system to decide what the words mean and if you don't have enough of it you won't get it right.