AI, LLMs, Machine Learning, and Survey Research

James Wagner Chair
University of Michigan
 
Tuesday, Aug 4: 10:30 AM - 12:20 PM
6442 
Contributed Papers 
Thomas M. Menino Convention & Exhibition Center 
Room: CC-106 

Main Sponsor

Survey Research Methods Section

Presentations

Evaluating Methods for Estimating Influence Effects in Social Networks: A Simulation Study

The contagion effect, also known as the peer effect, plays a significant role in social net-
works. It refers to the phenomenon in which one person or group can influence the behavior
of other individuals with whom they share social connections. However, accurately estimating
the contagion effect is challenging due to the confounding impact of peer selection. Previous
studies have shown that this can be viewed as a problem of omitted variable bias. In an attempt
to address this issue, simulation studies were conducted to compare the effectiveness of various
methodologies in estimating peer influence within social networks, including latent variable
models and machine learning techniques. Our research demonstrates that the performance
of various approaches vary depending on the social network scenario. Overall, the latent
space approach exhibits the best performance. This analysis provides valuable insights for re-
searchers studying estimation techniques and emphasizes the importance of understanding the
strengths, limitations, and objectives of these methods before using them for inference-based
estimation. 

Keywords

Contagious effect

Latent factor

Latent space

Node2Vec

SDNE 

Speaker

Brisilda Ndreka, National Institute of Health

Co-Author(s)

Dipak Dey, University of Connecticut
Ran Xu, UCONN

Mixed Integer Programming for Feature Selection in Scalar-on-Function Regression

Feature selection is a critical challenge in model selection, particularly for functional data, where appropriate statistical methodologies remain underdeveloped. This study investigates the application of Mixed Integer Programming (MIP) combined with information criteria for best feature subset selection in scalar-on-function regression (i.e., regression models where predictors are curves). Utilizing the computational power of an optimization tool uniquely allows us to employ combinatorics in feature selection, identifying the true best subset of features by comparing which minimizes the residuals the most. Transforming the functional regression problem into a classic linear model framework with grouped variables allows the use of model selection criteria such as Bayesian Information Criterion (BIC), in combination with MIP. In simulation studies, we compared our MIP method to alternative approaches and found that it consistently identifies truly active features, while not overselecting inactive features. 

Keywords

Functional Data Analysis

Mixed Integer Programming

Model Selection 

Speaker

Asha Pantula

Co-Author(s)

Luca Frigato, Università di Torino
Ana Kenney, UC Irvine
Marzia Cremona, Universite Laval

A Generative AI Approach for Integrating Synthetic Respondents with Probability-Based Human Panels

The growing demand for rapid, cost-effective, and scalable research solutions is driving interest in synthetic data for social science research. Synthetic respondents offer efficiency and scalability while reducing respondent burden, but their use must preserve data quality. This paper introduces a novel approach combining probability-based samples with synthetic respondents designed to mirror human respondents. Leveraging generative AI, we fine-tune large language models on NORC's AmeriSpeak panel data to create realistic synthetic panelists that emulate both aggregate response patterns and nuanced respondent behaviors. We compare finetuning strategies against context engineering-only approaches to optimize predictive validity. Synthetic responses are integrated with human data using dynamic models that adjust based upon predictive accuracy, ensuring insights remain grounded in authentic responses. We additionally implement ongoing validation protocols to assess bias, variance, and representativeness. We present findings from a pilot study, including comparative analyses of methods, integration strategies, and validation outcomes, and note implications for survey methodology. 

Keywords

Synthetic data

Survey methodology

Probability-based samples

Large language models

Data integration and fusion 

Speaker

Brandon Sepulvado

Co-Author(s)

Leah Christian, NORC
Joshua Y. Lerner, NORC at the University of Chicago
Soubhik Barari
Lilian Huang
Natalie Wang, NORC at the University of Chicago
Sabrina Sedovic, NORC at the University of Chicago

Testing the Limits of Generative AI: Do LLMs Mimic Human Evaluations in Factorial Survey Experiments

With the rapid development of large language models (LLMs), the debate over their ability to complement or replace traditional surveys by generating "synthetic samples" has intensified. The pre-trained models draw on vast amounts of training data, potentially reflecting nuanced attitudes and behaviors that are difficult to capture through surveys alone. Recent research examined whether LLMs can mimic participants' behavior in psychological experiments, with mixed results. However, the use of LLMs at the intersection of psychological experiments and surveys – namely factorial survey experiments (FSE) – remains underexplored in the current literature, despite their potential for studying attitude and behavioral intentions. In this study, we investigate whether and to what extent LLMs can mimic human evaluations in an FSE on earnings fairness. We compare results from probability and non-probability samples with multiple synthetic samples generated by different LLMs, in which personas are matched to the characteristics of the human respondents. Our findings contribute to the growing literature on synthetic samples and thereby inform future applications of LLMs in survey research. 

Keywords

Large Language Models

Factorial Survey Experiments

Probability Sample

Non-Probability Sample

Synthetic Sample 

Speaker

Sophie Hensgen, Institute for Employment Research

Co-Author(s)

Emma Fössing, Institute for Employment Research
Tobias Holtdirk, LMU
Joseph Sakshaug, German Institute for Employment Research

Evaluating LLM-Based Survey Questionnaire Translations with Human Evaluators and Cosine Similarities

Large language models (LLMs) may advance multilingual data collection instruments and tackle barriers in researching linguistic minorities, by assisting with language translations. However, the quality of LLM-based survey questionnaire translation and the method for such quality evaluation are yet to be understood.

This study compares human expert evaluation and automated similarity metrics for assessing LLM-based translation. We translated a survey questionnaire with 35 questions across three topics (socio-demographics, social networks, and cognitive health) from English (source) into four target languages (Spanish, Chinese, Korean, and Vietnamese) using commercial and open-source LLMs. Multiple human experts recorded the translation quality of each question, including the existence and type of translation errors, and the level of post-editing necessary to use the translated version. For the automated metric, we used cosine similarity scores between the source and target languages using the sBERT model. We will examine concordance between human and machine judgement of the translation quality by the type of translation error. 

Keywords

Large language models

Cosine similarity

Artificial intelligence

Survey methodology

Linguistic minorities 

Speaker

Mi Huynh

Co-Author(s)

Sunghee Lee, University of Michigan
Yajuan Si, University of Michigan
Mengyao Hu
Stephanie Morales
Jay Kim, University of Michigan
Felix Baez-Santiago, University of Michigan
Beining Niu, University of Michigan
Caleb Crouch, University of Michigan
Mengdi Ji, University of Michigan

Machine Learning Techniques for Survey Nonresponse Adjustment

Nearly all surveys have suffered a secular decline in response rates in the last two decades, raising growing concerns about potential nonresponse bias. In recent years NORC has helped pioneer the use of big data classifiers – machine learning models which use address-level data to target subpopulations of interest at the time of sample design. These address-level big data remain available after sampling for further modelling when survey data collection is complete. We conduct an empirical analysis of the use of these big data sources at the nonresponse adjustment stage of weighting. Machine learning models are used to predict response propensities, and the performance of the resulting nonresponse adjusted weights are compared to weights derived from more traditional methods of nonresponse adjustment. 

Keywords

Machine Learning

Big Data

Survey Non-Response

Survey Statistics 

Speaker

Noah Bassel, NORC

Co-Author

F. Jay Breidt, NORC at The University of Chicago

Innate hesitancies toward AI and work: flat-binomial composites to separate unsureness from feelings

As AI matures, the impact it might have on how society functions is being actively pondered on. In this article, through uniform-binomial mixtures, we shed quantitative light on this matter, showing topics that unite and divide the population at an unobserved, latent level. Simultaneously, we analyze Likert scale answers we collected in collaboration with Gallup, surveying 5835 US residents on how they sense guidelines on remote work or a four-day work week or spending time on work outside of scheduled hours could impact their lives. We find the tradeoff between uncertainty and feeling switches across topics. For some, such as a four-day week or returning to in-person work, uncertainty and feeling covary, while for others, such as spending time outside of work, they anitivary.Available are GitHub codes and interactive Shiny dashboards. Follow-up of two recent works: Bhaduri, M. (2025) Altered tradeoffs between uncertainty and feeling in inherent indecisions: uniform-binomial mixture densities to gauge the impact of shifts in work operations on employees. Societal Impacts.
Bhaduri, M. (2025) Are individuals who are positive about artificial intelligence also more unsure? Patterns. 

Keywords

sample surveys

Likert scale

uniform-binomial mixtures

clustering

sheltering effect

demographic diversity and persistence 

Speaker

Moinak Bhaduri