Hister: A private search engine for the pages you visit and the files you keep
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Hister: A private search engine for the pages you visit and the files you keep
Unofficial Hacker News client; not affiliated with Y Combinator.
taude · · focus · HN ↗
I have it up on GitHub, but I don't think anyone should use my implementation.
Loosely, what I built:
* On each of my machines I have a cron job running that looks at all my web browser history (usualy it's inspecting the brower's SQLlite across firefox and chrome). If it matches my rule list: hacker news stories, certain reddits, etc. it'll grab the page, convert to markdown and drop in my Obsidian Vault incoming.
* It has a whole de-duping architecture since I might open the same page on multiple machines. Uses the CloudFlare SQLITE D1 storage for tracking the processed links.
* it'll then trigger the LLM to do some Karpathy wiki style taxonomy assignment to the articles, organize them, create an index etc.
It's then available for my "bot" stuff to do writings for me.... I will probably write more about it at some point. I'm not certain it's totally useful and not just a yak-shave on hoarding knowledge.
Ai-drafted article on this [1]
Example AI-Drafted article based on some discussions the other day on Ollma vs LLama.cpp [2]
[1] <a href="https://taude.xyz/posts/how-archivore-turns-browsing-into-a-wiki/" rel="nofollow">https://taude.xyz/posts/how-archivore-turns-browsing-into-a-...
[2] <a href="https://taude.xyz/posts/skip-ollama-run-llama-cpp-directly-on-a-mac/" rel="nofollow">https://taude.xyz/posts/skip-ollama-run-llama-cpp-directly-o...
skinfaxi · · focus · HN ↗
rolandog · · focus · HN ↗
taude · · focus · HN ↗
EDIT: it's also the type of thing that feels very personally customized for my needs. I encourage you to build something similar on the idea. Much like how Karpathy Wiki was suggestive and not a runtime to just use...
pwython · · focus · HN ↗
As far as the need for private search, well, I've already searched for or visited those pages, so...
taude · · focus · HN ↗
The biggest win was the realization that both firefox and chrome maintain all the links you visit in a very queryable SQLite database. I've been poking at that for a lot of custom tools, like WHAT JIRA tickets am I paying attention to this week, etc....
devsda · · focus · HN ↗
I think browsers can play a part in building a local search index for URLs based on those keywords the page declares and cross verify/accept only those that are in prominently visible content, or may be delegate to an external engine(like LLMs) via an extension etc. This is particularly useful for cases where full text indexing is not feasible or desirable.
I doubt Google will ever add such feature in chrome though.
marginalia_nu · · focus · HN ↗
Exoristos · · focus · HN ↗
marginalia_nu · · focus · HN ↗
The poor data quality is a problem for anyone who wants to use the tags, which has seen everyone almost universally reaching for other solutions. Search engines have preferred anchor tags, bookmarking solutions have applied user tagging.
With keyword tags it's been a vicious circle of poor data quality and neglect since day one. Even in documents from the early 1990s when people were really trying, the data quality is inconsistent at best.
z3t4 · · focus · HN ↗
verdverm · · focus · HN ↗
ditto, it's an experiment in near-vibe coding, which also uses Typesense for queries using BM-25 & RAG with fusion. I have the web search/fetch/crawl features persisting raw intermediate values (api responses, search result lists) because I might re-use them one day... at least good for auditability if I need to
related, it is using Hister author's prior project SearXNG as one of the search providers