‹ BackHN Continuity

Thread

Hister: A private search engine for the pages you visit and the files you keep

741 points · 202 comments · bookofjoe

  1. 1vuio0pswjnm7 · · focus · HN ↗
    A different approach

    For person using resource-constrained computers where CPU, memory and storage space is limited

    URLs from the local forward proxy log are extracted periodically and stored in compressed files (URL logs)

    (I also store post-data)

    The compression method used is old and unpopular: recursive pairing

    Compression ratio is better than gzip but worse than zstd, compression/decompression speed better than zstd but worse than gzip

    More recently a method was developed to search these compressed files

    Size of compression utility: 42.3K static binary

    Size of search utility: 102.4K static binary

    No Java

    Limitations include basic regex only (no back-references) and files must be line-oriented

    No decompression step is needed. IME, this search is very fast. If it is slow then this means the keyword is too common: refine the search

    With minor modification (insert a newline at the top of file) I can also search inside compressed tar files

    As a www user with underpowered computers doing relatively small jobs, these old, unpopular methods have proven to be fast and reliable for me

    When I'm searching more than just URL strings, e.g., dates, titles, etc., I reformat the data into SQL and store it in a text file

    Instead of storing large SQL database files, I store the text file compressed with recusrive pairing

    I can then search the compressed SQL using basic regex; the output is piped into sqlite3 to create a "results" SQL database, e.g., in memory

    For me, the speed of sqlite3 in creating relatively small databases is excellent

    Then I can query the "results.db" using SQL

    1. 1vuio0pswjnm7 · · focus · HN ↗
      *recursive
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.