Parsing Historical Job Titles via LLMs to Analyze Social Mobility in Late Imperial China

Yue Yu, The Hong Kong University of Science and Technology
Jun Chen, Renmin University of China
Michael Chung, Hong Kong University of Science and Technology
Cameron Campbell, The Hong Kong University of Science and Technology

The digitization of Chinese historical career records has provided millions of data points for the study of social mobility. However, the unstructured nature of Chinese bureaucratic job titles—where an official’s functional office is conflated with honorary ranks, duty assignments, and prestige titles—has created a substantial barrier to accurate quantitative measurement, which requires a clear distinction between substantive power and symbolic status. We address this bottleneck by fine-tuning a Small Language Model (Qwen3-8B) on thousands of annotated pairs of job titles in the Qing dynasty to decompose these strings into a structured historical schema. We demonstrate that this specialist model outperforms commercial Large Language Models (LLMs) in both accuracy and cost-efficiency. By applying this parser to the full China Government Employee Database-Qing Jinshenlu (CGED-Q JSL) with over four million records, our empirical analysis reveals a sharp “inflation of honors” following the Taiping Rebellion (post-1864), quantitatively identifying the demographic sub-groups that leveraged financial capital to acquire symbolic status as the dynasty, desperate for revenue, resorted to the sale of honors. Our approach can be applied to other sources, including from historical Korea, Japan, and Vietnam, where job titles or other information about status are written in Chinese characters.

See extended abstract

 Presented in Session 70. Wrangling Occupational and Denominational Data