Let me add a few more layers just for some fun and profit :-D
Here[0] is an archive.org page which archives an archive.is page which archives the original article about internet censorship of an archive site that might require an archive of an archive site to read (given that people of spain cant now view archive.is in the first place or might have some difficulties doing so)
(If someone is perhaps interested and wants to see me talk more about what this is, then please read the blogpost that I had made for more info: <a href="https://smileplease.mataroa.blog/blog/htmlpipe-and-how-we-can-use-it-for-archive/" rel="nofollow">https://smileplease.mataroa.blog/blog/htmlpipe-and-how-we-ca...)
Thanks, I will try to look more into this. I think it might be related to number of connections of piping servers but I am a little perplexed about it on how this is possible on an archived link.
Can you please tell me more details so that I can hopefully try to fix it in the future. Did you try to open up the archive.org link of the website or the website itself (which is just a gh page: <a href="https://serjaimelannister.github.io/htmlpipe/?https://ppng.io/2YdN9_3" rel="nofollow">https://serjaimelannister.github.io/htmlpipe/?https://ppng.i...)
I would like to know more so that I can hopefully fix that for the future, have a nice day :-D
Been using the archive page firefox extension for a year or two. Yet now it looks like a total war against this resource was waged and it's no longer a good resource for archiving.
archive.is is the best substitute. For some reason people here won't accept that. They think the internet is still a calm cooperative place like the olden days.
My office's "net nanny" blocks archive.is as a "Russian site". However, archive.ph (use the same link across all the various sites) is not blocked (yet).
"Anyone know of any good reliable substitutes?"
For me, archive.today, archive.is, archive.md, archive.ph, etc. are _not reliable_ for a number of reasons
But some archive.today users who comment on HN cannot seem to accept that archive.today may not work for everybody else
NB. Archive.today is not a "substitute for archive.org". Archive.today does not do www crawls
As for archive.org, I know of a number of alternatives but each is generally less reliable and/or less comprehensive than archive.org
Comman Crawl, i.e., downloads from data.commoncrawl.org, is reasonably reliable but not as comprehensive as archive.org. CC is not a reasonable substitute for archive.org's CDX service. The CC CDX endpoint, index.commoncrawl.org, historically has been easily overwhelmed and unreliable
As for archive.today alternatives (no crawls, only user-submitted URLs), ghostarchive.org seems well-designed but not used much. No CAPTCHA, HTTPS and Javascript are optional and HAR files are provided. Whether it gets blocked like archive.today sites I do not know
NB. Archive.today users may be using archive.today not as an archive but as a lazy man's solution for "paywalls" (Javascript annoyances)
Where that's the case, comparsions to archive.org or other archives that are derived from crawls are inappropriate
peri-cl · · focus · HN ↗
Lio · · focus · HN ↗
An article about internet censorship of an archive site that requires an archive site to read!
That’s Brilliant! :D
Imustaskforhelp · · focus · HN ↗
Here[0] is an archive.org page which archives an archive.is page which archives the original article about internet censorship of an archive site that might require an archive of an archive site to read (given that people of spain cant now view archive.is in the first place or might have some difficulties doing so)
[0]: <a href="https://web.archive.org/web/20260920074416/https://serjaimelannister.github.io/htmlpipe/?https://ppng.io/2YdN9_3" rel="nofollow">https://web.archive.org/web/20260920074416/https://serjaimel...
(If someone is perhaps interested and wants to see me talk more about what this is, then please read the blogpost that I had made for more info: <a href="https://smileplease.mataroa.blog/blog/htmlpipe-and-how-we-can-use-it-for-archive/" rel="nofollow">https://smileplease.mataroa.blog/blog/htmlpipe-and-how-we-ca...)
1vuio0pswjnm7 · · focus · HN ↗
Imustaskforhelp · · focus · HN ↗
Can you please tell me more details so that I can hopefully try to fix it in the future. Did you try to open up the archive.org link of the website or the website itself (which is just a gh page: <a href="https://serjaimelannister.github.io/htmlpipe/?https://ppng.io/2YdN9_3" rel="nofollow">https://serjaimelannister.github.io/htmlpipe/?https://ppng.i...)
I would like to know more so that I can hopefully fix that for the future, have a nice day :-D
[deleted] · · focus · HN ↗
[deleted]
paul7986 · · focus · HN ↗
Anyone know of any good reliable substitutes?
GHanku · · focus · HN ↗
[dead]
mitxela · · focus · HN ↗
Tangurena2 · · focus · HN ↗
1vuio0pswjnm7 · · focus · HN ↗
For me, archive.today, archive.is, archive.md, archive.ph, etc. are _not reliable_ for a number of reasons
But some archive.today users who comment on HN cannot seem to accept that archive.today may not work for everybody else
NB. Archive.today is not a "substitute for archive.org". Archive.today does not do www crawls
As for archive.org, I know of a number of alternatives but each is generally less reliable and/or less comprehensive than archive.org
Comman Crawl, i.e., downloads from data.commoncrawl.org, is reasonably reliable but not as comprehensive as archive.org. CC is not a reasonable substitute for archive.org's CDX service. The CC CDX endpoint, index.commoncrawl.org, historically has been easily overwhelmed and unreliable
As for archive.today alternatives (no crawls, only user-submitted URLs), ghostarchive.org seems well-designed but not used much. No CAPTCHA, HTTPS and Javascript are optional and HAR files are provided. Whether it gets blocked like archive.today sites I do not know
NB. Archive.today users may be using archive.today not as an archive but as a lazy man's solution for "paywalls" (Javascript annoyances)
Where that's the case, comparsions to archive.org or other archives that are derived from crawls are inappropriate
ccgreg · · focus · HN ↗
Use our Parquet index.
It's also worth noting that archive.org downloads all of our crawl data and adds it to the IA Wayback Machine.
1vuio0pswjnm7 · · focus · HN ↗
1vuio0pswjnm7 · · focus · HN ↗