All posts

How to Research a Defunct Company Using Web Archives

July 23, 2026

When a company shuts down, its website usually vanishes within months — the domain lapses, gets parked, or is bought by someone selling something unrelated. But the questions do not vanish with it. Journalists investigating a failed startup, lawyers handling a dispute, researchers studying an industry, or former customers trying to document what they were promised all need to know what a dead company actually built, claimed, and sold. Web archives are usually the only place left to look.

Start with the domain history

Before diving into snapshots, establish the timeline. The Wayback Machine's calendar view for the company's homepage shows when captures start, when they get dense, and when they stop — a rough proxy for the company's lifespan and public activity. A sudden gap in captures, or a switch to a parking page, often pinpoints the shutdown window more precisely than any press coverage does.

Watch for domain reuse. If the captures from 2015 show a logistics startup and the captures from 2019 show an online casino, the domain changed hands. Note the transition date and treat everything after it as a different entity — a surprisingly common source of confusion in defunct-company research.

Enumerate what was captured

The homepage is rarely where the useful material lives. Use the CDX API (web.archive.org/cdx/search/cdx?url=example.com/*&output=json) to list every URL the Archive captured on the domain. Look for the high-value directories: /about/ and /team/ for who was involved and when, /pricing/ for what they charged, /blog/ and /press/ for what they announced, /docs/ or /help/ for how the product actually worked, and /terms/ or /privacy/ for the legal entities behind the site.

Team pages deserve special attention. Comparing /team/ or /about/ snapshots across years shows leadership turnover — who joined, who quietly disappeared from the page, and when. Cross-reference names against LinkedIn and press coverage to build a staffing timeline the company never published.

Fill gaps from outside the domain

The company's own site is only one source. Search the Archive for the company name across texts and media collections — trade publications, conference talks, and podcast episodes often survive when the corporate site does not. Archived versions of third-party sites (review platforms, job boards, partner pages, app store listings) preserve claims and details that the company's own marketing omitted. Crunchbase, AngelList, and press-release wires were heavily crawled and frequently retain funding and headcount data.

Government records are the durable backstop: SEC filings for anything that raised publicly, state business registries for incorporation and dissolution dates, and court records for the disputes that often accompany a shutdown. Archived pages help you know what to ask these systems for.

Arkibber helps most in the enumeration and triage phase — searching across the Archive's collections for a company name, filtering by media type and date range, and keeping the promising items organized while you work through a domain's capture history. Defunct-company research is inherently a many-sources problem, and losing track of what came from where is the fastest way to undermine the work.

Keep the evidence honest

One discipline matters above all: distinguish between what a company said and what was true. An archived pricing page proves the company advertised a price on a given date — not that anyone paid it. An archived team page proves someone was listed as CTO — not the terms of their employment. Record your findings with snapshot dates and archived URLs, label inferences as inferences, and your reconstruction will hold up when someone else checks it.

How to Find Old Newspapers and Magazines on the Internet Archive
Is Archive.org Down? How to Check and What to Do