An AI Dataset Is a Supply Chain, Not a Bag of Files
Revised and substantially expanded July 17, 2026. This essay discusses child sexual abuse material only at the level necessary to address dataset governance, research safety, and institutional responsibility; it does not reproduce or link to the material itself.
In late 2023, Stanford Internet Observatory research documented the presence of known child sexual abuse material (CSAM) in LAION-5B, a web-scale index widely used in image-model research. The finding was rightly disturbing. It also exposed a structural weakness that extends far beyond a single dataset or content category: the AI research ecosystem had learned to distribute data at internet scale without building an equally mature system for establishing provenance, communicating hazards, and containing downstream harm.
The original version of this post described the incident as a warning that harmful material can “go unnoticed” in very large datasets. That is true, but incomplete. Scale is not the root cause. Scale merely makes weak governance consequential.
A dataset used to train or evaluate an AI system is not a passive collection of files. It is the product of a supply chain: sources are selected; content is acquired or indexed; metadata is created; filters and transformations are applied; versions are published; mirrors and derivatives proliferate; models are trained; and those models become dependencies of still other systems. Every stage creates claims, obligations, and opportunities for failure.
When illegal or profoundly harmful material is discovered, the event should therefore be handled as a data-supply-chain incident. The response cannot end with deleting a few records from the latest copy. The responsible organization must determine what entered the chain, which artifacts derived from it, who received those artifacts, what models may have incorporated the data, what legal and victim-safety obligations apply, and what evidence is necessary to demonstrate containment.