Challenges and Insights from the First Full Count Linking of Censuses for England and Wales, 1851-1921

Guillaume Proffit, University of Cambridge
Alexis Litvine, University of Cambridge
Emma Diduch, University of Cambridge

In this paper, we adapt an unsupervised Bayesian census linking method to British censuses between 1851 and 1921. We split the process into several phases, which allows us to reduce the amount of blocking required and minimize false negatives. We also combine immutable and mutable characteristics to disambiguate cases with a high level of similarity. To overcome some of the limitations of traditional string comparators for nominal matching, we train a custom embeddings model to project the universe of name strings in such a way that geometric distance (as measure by cosine simimarity) between two names reflects the true similarity between two name strings. We then leverage those embeddings as a string comparator in our census linking pipeline, and cluster names within that embeddings space for use as a blocking key: we show that these improve both precision and recall of our linking. Historical census linking pipelines are making a growing use of mutable/transient characteristics as linking criteria: a number of recent papers show that the risk of erroneous links and selection bias is compensated by the greater recall that they enable. We show that occupations, as reported in British historical censuses follow that trend: handled carefully, the signal they bring is overall beneficial to census linking. Finally, by incorporating the entire set of marriage civil registration indexes in several phases of the linking process, we are able to halve the gender bias compared to existing methods in the field.

No extended abstract or paper available

 Presented in Session 177. Building and Interpreting Censuses II