web-archives

Article tag · 6 matching pages

Login-only material missing from public web archives | Dead Internet Theory

Dead Internet Theory: Public web archives capture what an anonymous visitor can read, so anything behind a login quietly sits outside the record. The Yahoo Groups shutdown shows how a twenty-year archive of communities escaped public capture, and how authorized members saved what they could.

Streaming media missed by page-level archiving | Dead Internet Theory

Dead Internet Theory: Streaming audio and video is not a file a crawler can copy, so page-level archives save the player and miss the stream. The British Library's One & Other case shows what was lost, what survived, and what a complete capture has to include.

Robots exclusions and gaps in historical archive coverage | Dead Internet Theory

Dead Internet Theory: How robots.txt, a file written for search-engine crawlers, has decided which pages get archived and which stored captures can be seen, and what a missing capture actually does and does not prove.

Personalized pages without a single canonical version to archive | Dead Internet Theory

Dead Internet Theory: Many pages are built separately for each visitor, so one URL can stand for dozens of different versions. A web archive captures only the version its anonymous crawler happened to get, and reading that capture later takes context the archive does not store.

Archive replay errors that make preserved content appear lost | Dead Internet Theory

Dead Internet Theory: Why a saved page can open blank, broken, or behind an error wall even though it was captured, and how to tell a replay failure from a real archive gap.

Ancient Web: The Nuclear Weapon Archive Started as a 1994 University Site

A nuclear-weapons reference site that began on a university server in 1994 survived shutdowns, migrations, mirrors, and three decades of the Web.