|
|
Sam Sam Hwang, Seoul National University
Christian Møller Dahl, University of Southern Denmark
Torben Johansen, University of Southern Denmark
Munir Squires, University of British Columbia
Historical census records enable researchers to track individual outcomes over time, but linking individuals across census rounds is particularly challenging for minority and immigrant populations due to transcription errors in handwritten names. We develop a machine learning approach that improves name transcription in historical U.S. census records, addressing the specific challenges of transcribing unfamiliar names from dense tabular formats. Focusing on the 1940 census, we find that independent human transcribers disagree on names in 30 percent of records, with disagreement rates rising for foreign-born individuals and non-English speakers. Our machine transcriptions increase linking rates by 121 percent for records where human transcribers disagree, while simultaneously improving match quality by 28 percent. These gains help expand sample sizes for traditionally under-linked groups — including the foreign-born, non-white residents, and those with no formal schooling — where each additional linked record is particularly valuable for statistical inference. Preliminary results for the 1930 census are being prepared for presentation at the conference.
No extended abstract or paper available
Presented in Session 159. Building and Interpreting Censuses I