From Mechanisms to Community Indicators: Measurement, Data Integration, and Calibration

Manjun Yu Chair
 
Monday, Aug 3: 8:30 AM - 10:20 AM
6438 
Contributed Papers 
Thomas M. Menino Convention & Exhibition Center 
Room: CC-106 
This session focuses on how Social Statistics turns theory and messy real-world data into reliable, actionable community indicators. Talks connect mechanistic ideas to distributional implications, and develop/validate measurement pipelines using administrative records, digital traces, and human-in-the-loop AI systems. Topics include improving income/poverty estimates under imputation, reconciling migration flows with accounting constraints, and new distributional measures for public opinion polarization—supporting "community in action" through stronger statistical infrastructure.

Main Sponsor

Social Statistics Section

Presentations

Advancing Knowledge to Advance Society: Specific Functions and Probability Distributions

To advance the growth of social science knowledge, the key is to mathematize the mechanisms and link them to probability distributions. This paper describes the four stages in mathematizing mechanisms and the link of fourth-stage specific functions to probability distributions, providing illustrations of both theoretical and empirical work, drawn largely from the study of status, justice, power, identity, and happiness. As shown, when an idea appears as an input-outcome relation, denoted X→Y, where X is a personal quantitative characteristic (e.g., income, beauty, skill) and Y a personal outcome (e.g., self-esteem, status, happiness), its mathematization leads to two sets of testable implications, via the specific function and via the probability distributions of X and Y, plus substantial additional empirical analysis. Testable predictions reach farflung topical domains – theft, gifts, classrooms, marital cohesion, conversation, salary secrecy, religious vocations, migration, polarization, intergroup relations, proportions integrationist and segregationist, whistleblowers, differential mourning for fathers and mothers, and whether parents buy more toys at birthdays or Christmas. 

Keywords

deductive theory

testable predictions, including novel predictions

inequality in personal characteristics (e.g., income) and inequality in personal sociobehavioral outcomes (e.g., happiness)

specific functions

probability distributions 

Speaker

Guillermina Jasso, New York University

AI-Driven Environmental Scanning for Education Policy and Practice

This project develops an AI enabled environmental scanning system to help the National Comprehensive Center (NCC) monitor emerging education issues. Funded by the U.S. Department of Education, the system automates collection of public information from news outlets and state, territorial, and tribal education agency websites. Using LLM based topic extraction aligned to a domain–theme–topic hierarchy, the system organizes unstructured text into actionable categories for technical assistance (TA) decision makers. A human in the loop process ensures quality, with reviewers refining prompts and schema definitions to stabilize outputs. The system processes information from about 60 states (including territories) and is expected to handle roughly 1,000 items per month. Early results show recurring patterns across states-for example, reports related to issues such as teacher shortages-illustrating how the system surfaces trends that support TA leads in identifying needs and coordinating assistance. By reducing manual scanning and providing structured, near real time insights, this approach strengthens the NCC's ability to detect priorities and improve cross center coordination. 

Keywords

Human-in-the-loop

Landscape scan

Trend detection

Automation 

Speaker

Atsushi Miyaoka

Co-Author

Gizem Korkmaz, Westat

Building Geolocated Business Inventories from Digital Trace Data: A Tobacco Retail Use Case

Digital trace data have become central to social science research, yet restricted platform access and data quality limitations complicate efforts to build reliable data collection pipelines. This study examines social media marketing by tobacco retailers, a task hindered by the lack of geolocation information in platform data, which prevents researchers from linking online marketing to local contexts. To overcome this limitation, we develop an innovative approach that identifies tobacco shops/vendors and their physical locations, and systematically links these entities to their online presence.

Using tobacco retailers as a case study, we present a generalizable framework that integrates Google Maps and Yelp listings and uses a hybrid record linkage strategy to construct geolocated, entity-level inventories. Associated social media accounts are identified and validated through manual review, enabling integration with population surveys to support spatial and inferential analyses. More broadly, this framework informs the development of sampling frames and analytic datasets in settings where traditional data sources are incomplete and platform access is rapidly evolving. 

Keywords

Digital trace data

Entity resolution

Record linkage

Geospatial data integration

Social media marketing

Survey data integration 

Speaker

Sara Lafia, NORC at the University of Chicago

Co-Author(s)

Yoonsang Kim, NORC at The University of Chicago
Simon Page, NORC at the University of Chicago
Mateusz Borowiecki, NORC at the University of Chicago
Chandler Carter, NORC at the University of Chicago
Sherry Emery
Anna Kostygina, NORC at the University of Chicago

Measuring Public Opinion: The Wasserstein Bipolarization Index

The extent to which the American public is politically polarized is of great interest in the lay and academic communities. To study opinion polarization, political scientists and public opinion researchers examine the distribution of respondents on survey items, using visual comparison of histograms and/or measures such as variances and bimodality coefficients. We prove these measures fail to align with prevailing conceptualizations of polarization put forth in the literature. To remedy this situation we specify several properties a measure of polarization consistent with these conceptualizations should possess: in particular, it should increase as a distribution spreads away from a center toward the poles and/or as clustering below or above this center increases. We then propose a p-Wasserstein bipolarization index that satisfies these properties and measures the distance between the distribution of an item and a most polarized distribution with all mass concentrated on the lower and upper endpoints of the scale, using the index to examine bipolarization in attitudes toward governmental COVID-19 vaccine mandates across 11 countries. 

Keywords

Polarization

Wasserstein Distance

Public Opinion

COVID-19 Vaccination Mandates

Political Methodology 

Speaker

Hane Lee, University of Chicago

Co-Author

Michael Sobel, Columbia University

Reconciling Bilateral Migration Flows with Net Migration Totals

The Human Migration Database (HMigD) at the Max Planck Institute for Demographic Research estimates international migration flows by integrating diverse data sources within a hierarchical Bayesian framework. A persistent challenge is that estimated bilateral flows are often inconsistent with external, high-quality country-level net migration totals, and reconciling these without modifying the underlying model is difficult. This paper introduces an information-projection (I-projection) approach as a post-processing step: starting from the posterior over bilateral flows, we compute the Kullback–Leibler–minimizing update that satisfies net-migration accounting constraints to obtain a consistent joint distribution. Unlike earlier solutions limited to independent Poisson flows, this method handles arbitrary joint distributions, preserves dependence structure, and provides evidence lower bound (ELBO)-like guarantees. We demonstrate scalable implementation and show that incorporating net-migration constraints improves HMigD estimates. 

Keywords

Migration Statistics

Constrained Estimation

Data Reconciliation

Information Projection

Kullback–Leibler Divergence 

Speaker

Boris Barron, Home Address

Co-Author(s)

Maciej Danko, Max Planck Institute for Demographic Research
Emilio Zagheni, Max Planck Institute for Demographic Research