How to Deal with Machine Learning Bias in Economic History

Julius Koschnick, University of Southern Denmark
Torben Johansen, University of Southern Denmark
Christian Vedel, University of Southern Denmark

Recent advances in machine learning (ML) have the potential to transform economic history by enabling large-scale analysis of historical sources. However, ML predictions often contain systematic, non-random errors that bias downstream econometric inference. These risks are especially acute in historical settings, where archaic language, anachronisms, and selective source survival exacerbate prediction bias. This paper argues that standard validation is insufficient and proposes adopting debiasing frameworks that estimate and correct prediction bias using gold-standard annotations produced by historically trained experts. We review the growing literature using ML in economic history and develop a taxonomy of three common applications—automated annotation, missing-data imputation, and embedding-based measurement—together with tailored bias-mitigation strategies. We argue that ML should complement, not replace, historical domain expertise. Integrating ML within structured debiasing frameworks preserves scalability while ensuring credible inference and critical historical evaluations in quantitative economic history.

See extended abstract

 Presented in Session 165. Microdata Methods