Reading Superfund Archival Sources at Scale: AI-Assisted Methods for Environmental History

Ella Howard, Wentworth Institute of Technology
Mehmet Ergezer, Wentworth Institute of Technology

Established under the Comprehensive Environmental Response, Compensation, and Liability Act of 1980, Superfund allows the Environmental Protection Agency to address contaminated sites. Forty-five years of investigations, evaluations, remediations, and follow-up studies have yielded a documentary archive containing hundreds of thousands of files. Historians can play a meaningful role in communicating the history and impact of Superfund, but only if they can navigate legacy federal file management systems. As part of an ongoing interdisciplinary research project, we explore how computational tools can help scholars navigate repositories. We argue that large language models (LLMs) can serve as research infrastructure to organize, compare, and surface patterns across document collections. The corpus includes more than 150,000 Superfund-related PDF documents downloaded from federal websites, grouped by waste site and indexed for retrieval and analysis. Our experimental computational pipeline uses retrieval-augmented generation (RAG) to produce short answers constrained to retrieved sources paired with page-level citations. To reduce cross-site leakage, we use single-site vector indexes. Our early findings reveal that LLMs can identify conceptually related passages missed by traditional keyword searching. However, document genre strongly shapes extraction quality, and model interpretation weakens when documents lack explicit factual statements. Critical questions addressed by this study include: What categories of regulatory documents are most useful for large-scale historical analysis? Can computational tools help historians surface case studies that might otherwise remain buried in institutional archives? How might large-scale document analysis complement traditional close reading and archival interpretation? This paper presents early findings from an ongoing project; additional corpus expansion, model testing, and case identification are planned prior to the November conference.

No extended abstract or paper available

 Presented in Session 125. Beyond the Census I