Monday, Aug 3: 8:30 AM - 10:20 AM
6438
Contributed Papers
Thomas M. Menino Convention & Exhibition Center
Room: CC-106
This session focuses on how Social Statistics turns theory and messy real-world data into reliable, actionable community indicators. Talks connect mechanistic ideas to distributional implications, and develop/validate measurement pipelines using administrative records, digital traces, and human-in-the-loop AI systems. Topics include improving income/poverty estimates under imputation, reconciling migration flows with accounting constraints, and new distributional measures for public opinion polarization—supporting "community in action" through stronger statistical infrastructure.
Main Sponsor
Social Statistics Section
Presentations
To advance the growth of social science knowledge, the key is to mathematize the mechanisms and link them to probability distributions. This paper describes the four stages in mathematizing mechanisms and the link of fourth-stage specific functions to probability distributions, providing illustrations of both theoretical and empirical work, drawn largely from the study of status, justice, power, identity, and happiness. As shown, when an idea appears as an input-outcome relation, denoted X→Y, where X is a personal quantitative characteristic (e.g., income, beauty, skill) and Y a personal outcome (e.g., self-esteem, status, happiness), its mathematization leads to two sets of testable implications, via the specific function and via the probability distributions of X and Y, plus substantial additional empirical analysis. Testable predictions reach farflung topical domains – theft, gifts, classrooms, marital cohesion, conversation, salary secrecy, religious vocations, migration, polarization, intergroup relations, proportions integrationist and segregationist, whistleblowers, differential mourning for fathers and mothers, and whether parents buy more toys at birthdays or Christmas.
Keywords
deductive theory
testable predictions, including novel predictions
inequality in personal characteristics (e.g., income) and inequality in personal sociobehavioral outcomes (e.g., happiness)
specific functions
probability distributions
This project develops an AI enabled environmental scanning system to help the National Comprehensive Center (NCC) monitor emerging education issues. Funded by the U.S. Department of Education, the system automates collection of public information from news outlets and state, territorial, and tribal education agency websites. Using LLM based topic extraction aligned to a domain–theme–topic hierarchy, the system organizes unstructured text into actionable categories for technical assistance (TA) decision makers. A human in the loop process ensures quality, with reviewers refining prompts and schema definitions to stabilize outputs. The system processes information from about 60 states (including territories) and is expected to handle roughly 1,000 items per month. Early results show recurring patterns across states-for example, reports related to issues such as teacher shortages-illustrating how the system surfaces trends that support TA leads in identifying needs and coordinating assistance. By reducing manual scanning and providing structured, near real time insights, this approach strengthens the NCC's ability to detect priorities and improve cross center coordination.
Keywords
Human-in-the-loop
Landscape scan
Trend detection
Automation
Digital trace data have become central to social science research, yet restricted platform access and data quality limitations complicate efforts to build reliable data collection pipelines. This study examines social media marketing by tobacco retailers, a task hindered by the lack of geolocation information in platform data, which prevents researchers from linking online marketing to local contexts. To overcome this limitation, we develop an innovative approach that identifies tobacco shops/vendors and their physical locations, and systematically links these entities to their online presence.
Using tobacco retailers as a case study, we present a generalizable framework that integrates Google Maps and Yelp listings and uses a hybrid record linkage strategy to construct geolocated, entity-level inventories. Associated social media accounts are identified and validated through manual review, enabling integration with population surveys to support spatial and inferential analyses. More broadly, this framework informs the development of sampling frames and analytic datasets in settings where traditional data sources are incomplete and platform access is rapidly evolving.
Keywords
Digital trace data
Entity resolution
Record linkage
Geospatial data integration
Social media marketing
Survey data integration
The extent to which the American public is politically polarized is of great interest in the lay and academic communities. To study opinion polarization, political scientists and public opinion researchers examine the distribution of respondents on survey items, using visual comparison of histograms and/or measures such as variances and bimodality coefficients. We prove these measures fail to align with prevailing conceptualizations of polarization put forth in the literature. To remedy this situation we specify several properties a measure of polarization consistent with these conceptualizations should possess: in particular, it should increase as a distribution spreads away from a center toward the poles and/or as clustering below or above this center increases. We then propose a p-Wasserstein bipolarization index that satisfies these properties and measures the distance between the distribution of an item and a most polarized distribution with all mass concentrated on the lower and upper endpoints of the scale, using the index to examine bipolarization in attitudes toward governmental COVID-19 vaccine mandates across 11 countries.
Keywords
Polarization
Wasserstein Distance
Public Opinion
COVID-19 Vaccination Mandates
Political Methodology
The Human Migration Database (HMigD) at the Max Planck Institute for Demographic Research estimates international migration flows by integrating diverse data sources within a hierarchical Bayesian framework. A persistent challenge is that estimated bilateral flows are often inconsistent with external, high-quality country-level net migration totals, and reconciling these without modifying the underlying model is difficult. This paper introduces an information-projection (I-projection) approach as a post-processing step: starting from the posterior over bilateral flows, we compute the Kullback–Leibler–minimizing update that satisfies net-migration accounting constraints to obtain a consistent joint distribution. Unlike earlier solutions limited to independent Poisson flows, this method handles arbitrary joint distributions, preserves dependence structure, and provides evidence lower bound (ELBO)-like guarantees. We demonstrate scalable implementation and show that incorporating net-migration constraints improves HMigD estimates.
Keywords
Migration Statistics
Constrained Estimation
Data Reconciliation
Information Projection
Kullback–Leibler Divergence