Benchmarking retrieval for agents on messy real-world company knowledge
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Benchmarking retrieval for agents on messy real-world company knowledge
Unofficial Hacker News client; not affiliated with Y Combinator.
emil_sorensen · · focus · HN ↗
polotics · · focus · HN ↗
What is your opinion of:
- BEAM
- MemEval
- MemoryArena
- any other you can find on eg. HuggingFace
How much value do you see in Karpathy's gist on the Episodic/Procedural/Semantic split (aka. btw. Doxa/Koine/Gnosis) If you do see value what parts of your design matches?
As you are using real company data, will you at least make one step towards reproducibility by publishing some extraction pipeline so other companies can run their own comparison?