Digital Extinction

Article tag · 50 matching pages

Dynamic pages that web crawlers capture only as empty shells | Dead Internet Theory

Dead Internet Theory: Why content assembled in the browser arrives too late for many crawlers, so archives store a page skeleton instead of the page — and what running a real browser does and does not recover.

Login-only material missing from public web archives | Dead Internet Theory

Dead Internet Theory: Public web archives capture what an anonymous visitor can read, so anything behind a login quietly sits outside the record. The Yahoo Groups shutdown shows how a twenty-year archive of communities escaped public capture, and how authorized members saved what they could.

Form-driven databases that cannot be preserved by following links | Dead Internet Theory

Dead Internet Theory: Records that live behind search forms have no links for crawlers to follow, so most web archives overlook them. This article explains how form-driven databases get lost — the Haddon catalogue is the documented case — and why exports and static pages preserve them where crawls cannot.

Streaming media missed by page-level archiving | Dead Internet Theory

Dead Internet Theory: Streaming audio and video is not a file a crawler can copy, so page-level archives save the player and miss the stream. The British Library's One & Other case shows what was lost, what survived, and what a complete capture has to include.

Robots exclusions and gaps in historical archive coverage | Dead Internet Theory

Dead Internet Theory: How robots.txt, a file written for search-engine crawlers, has decided which pages get archived and which stored captures can be seen, and what a missing capture actually does and does not prove.

CAPTCHAs and anti-bot barriers that also block preservation | Dead Internet Theory

Dead Internet Theory: Why the CAPTCHAs and browser checks that keep scrapers off sites also keep preservation crawlers out, what those barriers leave uncaptured, and how authorized access can still save material without dismantling the defenses.

Personalized pages without a single canonical version to archive | Dead Internet Theory

Dead Internet Theory: Many pages are built separately for each visitor, so one URL can stand for dozens of different versions. A web archive captures only the version its anonymous crawler happened to get, and reading that capture later takes context the archive does not store.

Geographic access restrictions and uneven archive coverage | Dead Internet Theory

Dead Internet Theory: Web servers decide what to send based on where a request comes from, and crawlers are requests too. That leaves archives crawling from a single country with a lopsided record of the web, which national archives and regional captures only partly repair.

Orphaned files reachable by a known URL but absent from every index | Dead Internet Theory

Dead Internet Theory: A file can keep answering at its address long after the links and index entries pointing to it vanish. This explains how pages end up reachable by a known URL but absent from every index, and what records still lead back to them.

Archive replay errors that make preserved content appear lost | Dead Internet Theory

Dead Internet Theory: Why a saved page can open blank, broken, or behind an error wall even though it was captured, and how to tell a replay failure from a real archive gap.

Forum migrations that break post identifiers and citation links | Dead Internet Theory

Dead Internet Theory: Why switching forum software can invalidate every post address, what broken permalinks cost readers and contributors, and how ID mapping and redirects keep the links alive.

CMS replacements that discard old revision histories | Dead Internet Theory

Dead Internet Theory: A site migration can preserve the current page while quietly erasing the older versions that explain how the page changed over time.

Cloud document deletion and the loss of publicly shared research | Dead Internet Theory

Dead Internet Theory: A cloud document can function like a public webpage until its owner deletes it, loses the account, or changes sharing permissions.

Code-hosting shutdowns and the survival of repository history | Dead Internet Theory

Dead Internet Theory: A software repository is more than its latest ZIP file, and a hosting shutdown can erase the history around code even when the code itself survives.

Package removal and the reproducibility of old software projects | Dead Internet Theory

Dead Internet Theory: An old project may have all of its source code intact and still become impossible to rebuild when one dependency disappears from a package registry.

Lost firmware files and the repairability of discontinued hardware | Dead Internet Theory

Dead Internet Theory: Old hardware may remain physically repairable while becoming digitally stranded because the exact firmware needed to recover it is no longer available.

Vanished API documentation for discontinued products | Dead Internet Theory

Dead Internet Theory: Source code can show that an old API existed, but without its reference documentation developers may lose the rules needed to understand or reproduce its behavior.

Government website transitions and missing historical publications | Dead Internet Theory

Dead Internet Theory: Government websites are reorganized whenever administrations, agencies, and publishing systems change, and older publications can become difficult to distinguish from genuinely deleted material.

Local newspaper closures and inaccessible reporting archives | Dead Internet Theory

Dead Internet Theory: When a local newspaper closes, the loss is not only future reporting; decades of past reporting can become difficult to search, cite, or retrieve.

Public records lost through database replacement rather than deliberate deletion | Dead Internet Theory

Dead Internet Theory: A public database can lose fields, attachments, searchability, or historical context during replacement even when nobody intended to delete a record.