Hister: A private search engine for the pages you visit and the files you keep
Thread
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
Hister: A private search engine for the pages you visit and the files you keep
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
jammaloo · · focus · HN ↗
<a href="https://news.ycombinator.com/item?id=49351802">https://news.ycombinator.com/item?id=49351802
evilduck · · focus · HN ↗
tamimio · · focus · HN ↗
billbrown · · focus · HN ↗
tamimio · · focus · HN ↗
361994752 · · focus · HN ↗
cobertos · · focus · HN ↗
This has an MCP server specifically so a workflow like that would work for you. This is just made to gold the data, and I'm a human accessible way should your AI fail you
361994752 · · focus · HN ↗
itsdesmond · · focus · HN ↗
rglullis · · focus · HN ↗
You didn't solve the problem, you are just trading pain points.
randomblock1 · · focus · HN ↗
361994752 · · focus · HN ↗
pixl97 · · focus · HN ↗
bradrn · · focus · HN ↗
jval43 · · focus · HN ↗
Nobody seems to remember it, even though it was a headline feature. Was removed in 2013, I think due to technical constraints.
Will definitely try this.
xd1936 · · focus · HN ↗
Edit: Found it. Thanks Claude.
<a href="https://github.com/ssnangua/google-chrome-comic-hd/blob/main/english/images/19.jpg" rel="nofollow">https://github.com/ssnangua/google-chrome-comic-hd/blob/main...
<a href="https://dfir.blog/history-index-files-removed-from-chrome-v30/" rel="nofollow">https://dfir.blog/history-index-files-removed-from-chrome-v3...
jval43 · · focus · HN ↗
varispeed · · focus · HN ↗
I think due to shareholders wanting new sportcars. The offline pages don't show Google Ads.
wtallis · · focus · HN ↗
thereforegrin · · focus · HN ↗
TeMPOraL · · focus · HN ↗
corney91 · · focus · HN ↗
iririririr · · focus · HN ↗
93po · · focus · HN ↗
iririririr · · focus · HN ↗
when it launched it was fast because it had LESS features.
jval43 · · focus · HN ↗
The risks were clear from day one but Google and Chrome were great to both users and devs, and it stayed like that for a long time.
iririririr · · focus · HN ↗
ericol · · focus · HN ↗
I even hacked it for a company I was working for at the time: I installed it in a machine that had a lot of pdfs from some other client company of them.
I don't recall the exact details, but I did some sort of proxying between the Microsoft Web Server that came with NT? at the time and G Desktop, and then an entire team of first support agents had almost instant search across all those docs. Good times.
TeMPOraL · · focus · HN ↗
It's like there's a timer or a cache somewhere, with aggressive limit, saying "only scan these top 100 results and then abort if it takes more than 10 ms to findf anything", or something. Making this completely useless.
That's on top of occasional deletion of browsing history past few weeks.
vbarrielle · · focus · HN ↗
TeMPOraL · · focus · HN ↗
negura · · focus · HN ↗
arantius · · focus · HN ↗
A while back I looked up a manual for an old Singer sewing machine, which turned out to be the model 15. So I typed [sin 15] in the address bar and -- yep, first result! A week ago I looked into the game Metroid Dread. Typed [met dre] and two relevant pages from my history came up, immediately.
I don't know what you mean by "100% unique match" but, this feature works for me, and in fact I regularly use it and find it very worthwhile.
arantius · · focus · HN ↗
I've been learning the eBay developer APIs. Recently (i.e. within the past few hours) I looked up the "Browse API" specifically. My first or second try failed so I typed just [ebay brows] and didn't find it! But when I finished [ebay browse] that document (within ebay.com and titled "Browse API | eBay Develo...") did show up!
I can't reproduce now so I can only guess but it's possible this was a timing thing. I was typing several things, and changing/re-typing "as soon as" I did or did not see what I wanted to show up. But there was definitely a moment where I'd expect "that page I just had open an hour ago" to be at the top, but instead search suggestions (ugh, search in the address bar ...) were all that was showing.
witrak · · focus · HN ↗
Lio · · focus · HN ↗
RobGR · · focus · HN ↗
asciimoo · · focus · HN ↗
Hister builds a personal search index from pages you visit, bookmarks, browser history, local files, and crawled websites. It stores extracted content with offline result previews, so information remains searchable even when the original page changes or disappears. It supports full text and semantic search, can run entirely on your own machine, and includes a web interface, command line tools, and an MCP endpoint for assistant integrations.
Website: <a href="https://hister.org/" rel="nofollow">https://hister.org/
Tiny read-only demo: <a href="https://demo.hister.org/" rel="nofollow">https://demo.hister.org/
Ps.: It looks like our name conflicts with a registered trademark in the US. The owner of the other project has asked us to change it, so we’ll probably need to comply sooner or later.
Name suggestions are welcome! Ideally, the new name should be relatively short, sound good, and have an available .org domain.
Thanks!
Capricorn2481 · · focus · HN ↗
adfm · · focus · HN ↗
asciimoo · · focus · HN ↗
adfm · · focus · HN ↗
juliend2 · · focus · HN ↗
_kidlike · · focus · HN ↗
ιστορία, if you wanna copy paste.
and yes, the English word comes from the Greek word!
golem14 · · focus · HN ↗
TaLiTr · · focus · HN ↗
golem14 · · focus · HN ↗
corndoge · · focus · HN ↗
<a href="https://archivebox.io/" rel="nofollow">https://archivebox.io/
What does Hister do differently? Search seems like a major differentiator, I'm wondering if leveraging the existing archivebox project for archival and implementing good search on top would be more efficient
asciimoo · · focus · HN ↗
corndoge · · focus · HN ↗
dbliss · · focus · HN ↗
asciimoo · · focus · HN ↗
nottorp · · focus · HN ↗
Looks like you can even set authentication up so you can run it at home but connect while you're away too...
dwedge · · focus · HN ↗
pava0 · · focus · HN ↗
dwedge · · focus · HN ↗
danielrmay · · focus · HN ↗
Anonymous106 · · focus · HN ↗
Beijinger · · focus · HN ↗
So? Where are you based? For what class was the trademark filed? When was it filed?
I doubt that he has any leverage, but I don't know the background.
asdfqwertzxcv · · focus · HN ↗
Beijinger · · focus · HN ↗
How does it sync via several computers?
al_hag · · focus · HN ↗
michaelje · · focus · HN ↗
jamienk · · focus · HN ↗
yeahdef · · focus · HN ↗
mxuribe · · focus · HN ↗
swyx · · focus · HN ↗
Cider9986 · · focus · HN ↗
john_strinlai · · focus · HN ↗
do you have a link to that trademark?
the only one i see containing "hister" is <a href="https://tmsearch.uspto.gov/search/search-results/98575798" rel="nofollow">https://tmsearch.uspto.gov/search/search-results/98575798 which is "dead" and "abandoned".
"This trademark application was refused, dismissed, or invalidated by the Office and this application is no longer active."
the only other similar-ish ones ("historix") appear to be in different domains (e.g. translation services).
if you're tied to the name, it's probably worthwhile digging deeper.
mxuribe · · focus · HN ↗
* chronilog.org ...as in, a log of one's chronicles.
* And if you will include this into KDE, then can use a 'k' instead, such as kronilog.org :-)
Both seem to be available. ;-)
asciimoo · · focus · HN ↗
consumer451 · · focus · HN ↗
pidgeon_lover · · focus · HN ↗
skripp11 · · focus · HN ↗
elektor · · focus · HN ↗
Question for you: For the less tech savvy of us on here, is there any chance Hister can be can hosted on something like Pikapods? <a href="https://www.pikapods.com/" rel="nofollow">https://www.pikapods.com/
asciimoo · · focus · HN ↗
elektor · · focus · HN ↗
cyanlimetea · · focus · HN ↗
[dead]
mircea · · focus · HN ↗
E.g.: for this submission I would want both <a href="https://news.ycombinator.com/item?id=49743097">https://news.ycombinator.com/item?id=49743097 and <a href="https://github.com/asciimoo/hister" rel="nofollow">https://github.com/asciimoo/hister captured.
asciimoo · · focus · HN ↗
iririririr · · focus · HN ↗
asciimoo · · focus · HN ↗
asimovDev · · focus · HN ↗
brunousedtowrit · · focus · HN ↗
[dead]
hipjiveguy · · focus · HN ↗
Zizizizz · · focus · HN ↗
Seekfold (seek and manifold)
Seekdex (seek and index)
6510 · · focus · HN ↗
fooqux · · focus · HN ↗
smellf · · focus · HN ↗
culi · · focus · HN ↗
<a href="https://filmot.com/" rel="nofollow">https://filmot.com/
It lets you search YouTube transcripts. If you could somehow integrate video transcripts into this tool, I would be extremely interested in trying it out
asciimoo · · focus · HN ↗
culi · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
kadoban · · focus · HN ↗
hefner1456 · · focus · HN ↗
aaron695 · · focus · HN ↗
[dead]
usernomdeguerre · · focus · HN ↗
Put another way, i can't write an extractor for Reuters and then point a config to it from my current hister binary?
cyanlimetea · · focus · HN ↗
[dead]
g_host · · focus · HN ↗
I ran it for most of this year but encountered some problems with it I couldn't fix and thus have not had it hooked up to anything since June when I finally couldn't take it anymore.
In short, I serve a good number of apps from an Nginx reverse proxy. Maybe 25% of them are exposed to the WWW while everything else is limited to the LAN but I still get valid TLS for all of it.
Hister, though, kept breaking my whole reverse proxy and I could never figure out EXACTLY why so I could fix it. After running fine for a few days, it would hog the whole server and everything else proxied by Nginx would become unreachable. I tried tuning the config for it to no avail.
One day when I'm less lazy, I'll probably hook it back up via it's LAN IP to every machine I've got again. I REALLY liked that I could log my browsing history from any machine anywhere in the world without a VPN and I was really disappointed when I had to disable its config in Nginx.
I still use it a lot to go find stuff I flagged as important quickly.
I'm curious if this is something you've heard of before, or if I've got a one off problem here.
I even ported the config to a brand new VM with NGINX and still had the same problem.
dwedge · · focus · HN ↗
It might be something similar if all of your sites are subdomains. Try incognito at the same time next time
g_host · · focus · HN ↗
I've got some applications used daily by friends all over the world and the services would all become unavailable to them when this started happening. Only fix I found was restarting NGINX service and then it could happen again an hour later or 3 days later. Once I removed the proxy config for Hister from the service, the issue never happened again.
I'll repro the issue and get the details intoa. Github Issue this weekend.
asciimoo · · focus · HN ↗
g_host · · focus · HN ↗
rurban · · focus · HN ↗
traktorn · · focus · HN ↗
ggm · · focus · HN ↗
brimwats · · focus · HN ↗
darkwater · · focus · HN ↗
And for the name, what about "Historex" (although already taken as well) o "Histearch", mesh-up of "history" and "search"?
zuminator · · focus · HN ↗
amai · · focus · HN ↗
scoot · · focus · HN ↗
"The HYSTER trademark is filed in the category of Education and Entertainment Services" [1], so you can safely ignore any demands to rename, as it doesn't conflict. Embarrassing for them that their lawyers don't understand even the basics of trademark law
[1] <a href="https://www.trademarkia.com/hyster-77843354" rel="nofollow">https://www.trademarkia.com/hyster-77843354
scoot · · focus · HN ↗
Or more likely do, but hope that you don't...
ptaffs · · focus · HN ↗
jsmo · · focus · HN ↗
SimplGy · · focus · HN ↗
sbeckeriv · · focus · HN ↗
I like the search ui. my projects become functional but never polished. <a href="https://github.com/sbeckeriv/memoir" rel="nofollow">https://github.com/sbeckeriv/memoir
kilroy123 · · focus · HN ↗
It's badly needed, and so far it's working well for me.
pkamb · · focus · HN ↗
Is there any site/project that works as a fully customizable personal front-end to all other SERPs?
When I search for something, I always want a link to the best Wikipedia result. This should always be in the same place and have a giant icon/picture.
Then there could be easily clickable links to the SERP pages for Google, DDG, etc. for that query.
A big link to route it to your favorite LLM.
Seems like you could have a really useful "homepage" for all searches that sat in front of all the other sites. It could be local only and would not require indexing the web. Also wouldn't be a files search thing, as Hister appears to be.
taude · · focus · HN ↗
I have it up on GitHub, but I don't think anyone should use my implementation.
Loosely, what I built:
* On each of my machines I have a cron job running that looks at all my web browser history (usualy it's inspecting the brower's SQLlite across firefox and chrome). If it matches my rule list: hacker news stories, certain reddits, etc. it'll grab the page, convert to markdown and drop in my Obsidian Vault incoming.
* It has a whole de-duping architecture since I might open the same page on multiple machines. Uses the CloudFlare SQLITE D1 storage for tracking the processed links.
* it'll then trigger the LLM to do some Karpathy wiki style taxonomy assignment to the articles, organize them, create an index etc.
It's then available for my "bot" stuff to do writings for me.... I will probably write more about it at some point. I'm not certain it's totally useful and not just a yak-shave on hoarding knowledge.
Ai-drafted article on this [1]
Example AI-Drafted article based on some discussions the other day on Ollma vs LLama.cpp [2]
[1] <a href="https://taude.xyz/posts/how-archivore-turns-browsing-into-a-wiki/" rel="nofollow">https://taude.xyz/posts/how-archivore-turns-browsing-into-a-...
[2] <a href="https://taude.xyz/posts/skip-ollama-run-llama-cpp-directly-on-a-mac/" rel="nofollow">https://taude.xyz/posts/skip-ollama-run-llama-cpp-directly-o...
skinfaxi · · focus · HN ↗
rolandog · · focus · HN ↗
taude · · focus · HN ↗
pwython · · focus · HN ↗
As far as the need for private search, well, I've already searched for or visited those pages, so...
taude · · focus · HN ↗
The biggest win was the realization that both firefox and chrome maintain all the links you visit in a very queryable SQLite database. I've been poking at that for a lot of custom tools, like WHAT JIRA tickets am I paying attention to this week, etc....
devsda · · focus · HN ↗
I think browsers can play a part in building a local search index for URLs based on those keywords the page declares and cross verify/accept only those that are in prominently visible content, or may be delegate to an external engine(like LLMs) via an extension etc. This is particularly useful for cases where full text indexing is not feasible or desirable.
I doubt Google will ever add such feature in chrome though.
marginalia_nu · · focus · HN ↗
Exoristos · · focus · HN ↗
marginalia_nu · · focus · HN ↗
The poor data quality is a problem for anyone who wants to use the tags, which has seen everyone almost universally reaching for other solutions.
With keyword tags it's been a vicious circle of poor data quality and neglect since day one. Even in documents from the early 1990s when people were really trying, the data quality is inconsistent at best.
z3t4 · · focus · HN ↗
verdverm · · focus · HN ↗
ditto, it's an experiment in near-vibe coding, which also uses Typesense for queries using BM-25 & RAG with fusion. I have the web search/fetch/crawl persisting raw intermediate values (api responses, search result lists) because I might use them one day...
tombert · · focus · HN ↗
[1] <a href="https://git.brucewillis.sexy/~tombert/fs_index" rel="nofollow">https://git.brucewillis.sexy/~tombert/fs_index I promise, safe for work, despite the URL.
testycool · · focus · HN ↗
tombert · · focus · HN ↗
I like it.
kanzure · · focus · HN ↗
asciimoo · · focus · HN ↗
MomsAVoxell · · focus · HN ↗
Every single web page I’ve found interesting, since the advent of the Web, I have printed to PDF and stored locally for my own personal reference.
Something like 80,000+ files - my own copy of my own Internet - indexable, searchable.
Available offline. Something to read when I am far out to sea.
There is no need to involve third parties in your Internet history - no matter how trustworthy they seem to want to appear.
Print to PDF, and you’ve got everything you need, safe and sound.
computator · · focus · HN ↗
The more "modern" the site, the worse it is. Surprisingly, government websites often print correctly since they've done the least amount of work to make the site modern looking.
jiehong · · focus · HN ↗
Even full page screenshot doesn’t always capture the non visible part of the page (below the viewport).
MomsAVoxell · · focus · HN ↗
Reader mode.
nottorp · · focus · HN ↗
MomsAVoxell · · focus · HN ↗
nottorp · · focus · HN ↗
MomsAVoxell · · focus · HN ↗
krackers · · focus · HN ↗
computator · · focus · HN ↗
How do other people handle this dilemma?
Even solution I can think of involves are a great amount of extra work.
BrokenCogs · · focus · HN ↗
computator · · focus · HN ↗
applfanboysbgon · · focus · HN ↗
itsdesmond · · focus · HN ↗
edoceo · · focus · HN ↗
perlgeek · · focus · HN ↗
Eduard · · focus · HN ↗
Zephyrix · · focus · HN ↗
One thing that can be helpful when reasoning about things like this is figuring out what your actual threat model is. What does system compromise look like to you? Data exfiltration, arbitrary code execution, something else?
invalidator · · focus · HN ↗
Unfortunately there's no one-size-fits-all solution for this yet, but there are a lot of groups attacking it from different angles:
Qubes GrapheneOS Firejail Bubblewrap Android/iOS app permissions Landlock App Sandbox etc
EvanAnderson · · focus · HN ↗
We should be doing this with all software, regardless of provenance, anyway. Least privilege applies to servers just as much as it does to users. Even if the software isn't untrustworthy you can be it has vulnerabilities.
My first go-to is network segmentation because I spend most of my time doing networking work. For every vendor who has shit-talked me to Customers ("Wah, wah! Your networking vendor is making this so much harder because they want us to enumerate our traffic!") I have concrete examples I can cite when attacks were stopped by network segmentation (preventing shellcode from downloading a payload, preventing C2 communication, firing off alerts when unexpected network traffic starts coming out of a host, etc).
Beyond network segmentation, I am very suspicious of software that needs to run as a privileged user. So many attacks get easier when privilege escalation in the host OS is already done for you.
nobody42 · · focus · HN ↗
Apart from high-overhead solutions like VMs and containers, there are seamless and maintenance-free solutions (after the initial setup):
- systemd service hardening [0] [1]
pretty powerful, but it's a blacklist approach - whack-a-mole
- AppArmor [2]
Whitelist, proactive approach. Contrary to SElinux, it's not a programming language, and could be grasped pretty quickly. I made a tool to easily convert AA logs into usable rules. [3]
[0] <a href="https://github.com/alegrey91/systemd-service-hardening" rel="nofollow">https://github.com/alegrey91/systemd-service-hardening
[1] <a href="https://github.com/desbma/shh" rel="nofollow">https://github.com/desbma/shh
[2] <a href="https://presentations.nordisch.org/apparmor/" rel="nofollow">https://presentations.nordisch.org/apparmor/
[3] <a href="https://github.com/nobody43/apparmor-suggest" rel="nofollow">https://github.com/nobody43/apparmor-suggest
TeMPOraL · · focus · HN ↗
For me: consider this a form of paranoia and ignore it, while worrying more about cleanup costs of non-vetted packages.
Like, even if there is 5% chance that a program I download will start downloading global python or node packages, odds are within a year I'll deal with couple that have mutually incompatible requirements and are impossible to run without more VM surgery than I have patience for, and that I'll discover this only after a botched installation bricks software that used to work before.
But that's solvable with less extra work. Just throwaway containers. With no hand-wringing about read-only access or isolating it from network, because my threat modeling doesn't consider loss of privacy or any data leak from my personal local side to be realistic or impactful threat event - OTOH, it assigns great magnitudes to loss of personal time.
torvald · · focus · HN ↗
GlacierFox · · focus · HN ↗
Vaslo · · focus · HN ↗
ifh-hn · · focus · HN ↗
jamienk · · focus · HN ↗
Can I add NOTES about pages? This might be a good spot to do that...? Maybe the interface can be in a web page instead of terminal?
Before Google took off there was a vibrant ecosystem of FOSS dev around search, all different little aspects of it. Then after Google people stopped fiddling with search, search became "solved" or maybe "must be coded by the big boys". Shame.
Thank you for this, looooong time coming
asciimoo · · focus · HN ↗
Notes are not supported yet, only labels. But it is a useful addition, added to my TODO.
jjice · · focus · HN ↗
I also have it index my Obsidian notes, which is another little bonus for global search.
I did need to build up quite a few exclusion rules early on, but it's been hands off since.
roschdal · · focus · HN ↗
w10-1 · · focus · HN ↗
phyzome · · focus · HN ↗
machomaster · · focus · HN ↗
<a href="https://github.com/gildas-lormeau/SingleFile" rel="nofollow">https://github.com/gildas-lormeau/SingleFile
"SingleFile helps you to save a complete web page into a single HTML file. SingleFile is a Web Extension (and a CLI tool) compatible with Chrome, Firefox (Desktop and Mobile), Microsoft Edge, Safari, Vivaldi, Brave, Waterfox, Yandex browser, and Opera."
diarrhea · · focus · HN ↗
gglanzani · · focus · HN ↗
mirashii · · focus · HN ↗
autoexec · · focus · HN ↗
jevogel · · focus · HN ↗
machomaster · · focus · HN ↗
1. Leaves all kinds of css, image file, which might be hundreds for each saved page. SingleFile saves everything in a convenient, working, single file. 2. is not really saving all the info.
I got burned when I was using Mozilla/Internet Explorer to save mhtml files, which ended up not being supported anymore. Additionally, turns out, they were saving the original versions of pages, without the Javascript changes, meaning that all the opened threads of comments I wanted to save were never saved! Never again!
phyzome · · focus · HN ↗
machomaster · · focus · HN ↗
If you check what kind of stuff SingleFile does, you will find a list of things Firefox is not doing.
colbertw08 · · focus · HN ↗
liberian · · focus · HN ↗
[dead]
etamponi · · focus · HN ↗
wrxd · · focus · HN ↗
etamponi · · focus · HN ↗
Invictus0 · · focus · HN ↗
rao-v · · focus · HN ↗
I built myself a little extension last year that tracks what information I was looking at, but focused on generating "new info" recaps for the day / week.
I realized that I open / quick view a lot of pages and close them, which is a strong signal that I don't care about that specific page, and it shouldn't be a source of "new insights" that I learnt that day (since I probably don't care about that topic).
I'd love to re-try a simpler version of that project that builds on Hister as a backend actually.
kenanfyi · · focus · HN ↗
Adding a timer constraint would be really nice and I guess it might be quite easy to do it since it's only a concern in the browser extension.
devdoshi · · focus · HN ↗
alphabet9000 · · focus · HN ↗
<a href="https://web.archive.org/web/20170808120657/http://www.rotten.com/library/bio/misc/nostradamus/" rel="nofollow">https://web.archive.org/web/20170808120657/http://www.rotten...
valcarvalho · · focus · HN ↗
saimiam · · focus · HN ↗
<a href="https://www.history.co.uk/articles/did-nostradamus-really-predict-the-rise-of-adolf-hitler" rel="nofollow">https://www.history.co.uk/articles/did-nostradamus-really-pr...
jambalaya8 · · focus · HN ↗
Notkel · · focus · HN ↗
mattjbarnes · · focus · HN ↗
<a href="https://stashpad.ai/" rel="nofollow">https://stashpad.ai/
ferrule · · focus · HN ↗
alwillis · · focus · HN ↗
[1]: <a href="https://findanyfile.app" rel="nofollow">https://findanyfile.app
rochansinha · · focus · HN ↗
febed · · focus · HN ↗
Yashjain413 · · focus · HN ↗
One of the most useful use cases for me, especially since I work in GTM, is keeping track of new ways to get replies from prospects, whether through cold outbound or things like SEO/GEO optimization. I read at least 2-3 articles a day on this, and it genuinely helps me figure out which ideas are worth trying because I can now keep track of everything I’ve read.
noisy_boy · · focus · HN ↗
cardboardguru · · focus · HN ↗
IndiaInfraNotes · · focus · HN ↗
[dead]
Yehoshaphat · · focus · HN ↗
soapdog · · focus · HN ↗
frumiousirc · · focus · HN ↗
szamski · · focus · HN ↗
pidgeon_lover · · focus · HN ↗
(For file search, I use voidtools' Everything, and I'm not sure why anyone other than Microsoft would want to mix local file results and web results)
elestor · · focus · HN ↗
ozim · · focus · HN ↗
Running this would be nice but it still takes time and still there is no ROI for me.
Of course there will be people who find it useful but I am pretty much done with building knowledge bases or having todo lists.
Stuff that I need to do or remember - everything else if I forget nothing happens and it doesn’t impact my life or work.
amai · · focus · HN ↗
adrienconrath · · focus · HN ↗
It can index all your data (emails, whatsapp/imessage, gdrive, finance, health, etc) and notably also allows you to install a chrome extension to also capture and index pages you see on the web.
It builds a search index and connected graph of all this data.
I have just open sourced: <a href="https://github.com/omnesis-dev/Omnesis" rel="nofollow">https://github.com/omnesis-dev/Omnesis
Unified-Mentor · · focus · HN ↗
[dead]
itsmeduncan · · focus · HN ↗
[dead]
dang · · focus · HN ↗
Of course, it's impossible to know for sure what was LLM processed or not, but some of your posts (like this one) have been getting classified that way.
smilliken · · focus · HN ↗
chrisss395 · · focus · HN ↗
Does Hister handle this well? If not, can anyone suggest other options?
javatextbook · · focus · HN ↗
1vuio0pswjnm7 · · focus · HN ↗
For person using resource-constrained computers where CPU, memory and storage space is limited
URLs from the local forward proxy log are extracted periodically stored in compressed files (URL logs)
(I also store post-data)
The compression method used is old and unpopular: recursive pairing
Compression ratio is better than gzip but worse than zstd, compression/decompression speed better than zstd but worse than gzip
More recently a method was developed to search these compressed files
Size of compression utility: 42.3K static binary
Size of search utility: 102.4K static binary
No Java
Limitations include basic regex only (no back-references) and files must be line-oriented
No decompression step is needed. IME, this search is very fast. If it is slow then this means the keyword is too common: refine the search
With minor modification (insert a newline at the top of file) I can also search inside compressed tar files
As a www user with underpowered computers doing relatively small jobs, these old, unpopular methods have proven to be fast and reliable for me
When I'm searching more than just URL strings, e.g., dates, titles, filetypes, etc., I reformat the data into SQL and store it in a text file
Instead of storing large SQL database files, I compress the text file instead using recusrive pairing
I can then keyword search the compressed SQL; the output is piped into sqlite3 to create a "results" SQL database, e.g., in memory
For me, the speed of sqlite3 in creating relatively small databases is excellent
Then I can full-text search the "results.db" using SQL query language
1vuio0pswjnm7 · · focus · HN ↗