Tuesday, Aug 4: 10:30 AM - 12:20 PM
6442
Contributed Papers
Thomas M. Menino Convention & Exhibition Center
Room: CC-106
Main Sponsor
Survey Research Methods Section
Presentations
The contagion effect, also known as the peer effect, plays a significant role in social net-
works. It refers to the phenomenon in which one person or group can influence the behavior
of other individuals with whom they share social connections. However, accurately estimating
the contagion effect is challenging due to the confounding impact of peer selection. Previous
studies have shown that this can be viewed as a problem of omitted variable bias. In an attempt
to address this issue, simulation studies were conducted to compare the effectiveness of various
methodologies in estimating peer influence within social networks, including latent variable
models and machine learning techniques. Our research demonstrates that the performance
of various approaches vary depending on the social network scenario. Overall, the latent
space approach exhibits the best performance. This analysis provides valuable insights for re-
searchers studying estimation techniques and emphasizes the importance of understanding the
strengths, limitations, and objectives of these methods before using them for inference-based
estimation.
Keywords
Contagious effect
Latent factor
Latent space
Node2Vec
SDNE
Feature selection is a critical challenge in model selection, particularly for functional data, where appropriate statistical methodologies remain underdeveloped. This study investigates the application of Mixed Integer Programming (MIP) combined with information criteria for best feature subset selection in scalar-on-function regression (i.e., regression models where predictors are curves). Utilizing the computational power of an optimization tool uniquely allows us to employ combinatorics in feature selection, identifying the true best subset of features by comparing which minimizes the residuals the most. Transforming the functional regression problem into a classic linear model framework with grouped variables allows the use of model selection criteria such as Bayesian Information Criterion (BIC), in combination with MIP. In simulation studies, we compared our MIP method to alternative approaches and found that it consistently identifies truly active features, while not overselecting inactive features.
Keywords
Functional Data Analysis
Mixed Integer Programming
Model Selection
The growing demand for rapid, cost-effective, and scalable research solutions is driving interest in synthetic data for social science research. Synthetic respondents offer efficiency and scalability while reducing respondent burden, but their use must preserve data quality. This paper introduces a novel approach combining probability-based samples with synthetic respondents designed to mirror human respondents. Leveraging generative AI, we fine-tune large language models on NORC's AmeriSpeak panel data to create realistic synthetic panelists that emulate both aggregate response patterns and nuanced respondent behaviors. We compare finetuning strategies against context engineering-only approaches to optimize predictive validity. Synthetic responses are integrated with human data using dynamic models that adjust based upon predictive accuracy, ensuring insights remain grounded in authentic responses. We additionally implement ongoing validation protocols to assess bias, variance, and representativeness. We present findings from a pilot study, including comparative analyses of methods, integration strategies, and validation outcomes, and note implications for survey methodology.
Keywords
Synthetic data
Survey methodology
Probability-based samples
Large language models
Data integration and fusion
With the rapid development of large language models (LLMs), the debate over their ability to complement or replace traditional surveys by generating "synthetic samples" has intensified. The pre-trained models draw on vast amounts of training data, potentially reflecting nuanced attitudes and behaviors that are difficult to capture through surveys alone. Recent research examined whether LLMs can mimic participants' behavior in psychological experiments, with mixed results. However, the use of LLMs at the intersection of psychological experiments and surveys – namely factorial survey experiments (FSE) – remains underexplored in the current literature, despite their potential for studying attitude and behavioral intentions. In this study, we investigate whether and to what extent LLMs can mimic human evaluations in an FSE on earnings fairness. We compare results from probability and non-probability samples with multiple synthetic samples generated by different LLMs, in which personas are matched to the characteristics of the human respondents. Our findings contribute to the growing literature on synthetic samples and thereby inform future applications of LLMs in survey research.
Keywords
Large Language Models
Factorial Survey Experiments
Probability Sample
Non-Probability Sample
Synthetic Sample
Large language models (LLMs) may advance multilingual data collection instruments and tackle barriers in researching linguistic minorities, by assisting with language translations. However, the quality of LLM-based survey questionnaire translation and the method for such quality evaluation are yet to be understood.
This study compares human expert evaluation and automated similarity metrics for assessing LLM-based translation. We translated a survey questionnaire with 35 questions across three topics (socio-demographics, social networks, and cognitive health) from English (source) into four target languages (Spanish, Chinese, Korean, and Vietnamese) using commercial and open-source LLMs. Multiple human experts recorded the translation quality of each question, including the existence and type of translation errors, and the level of post-editing necessary to use the translated version. For the automated metric, we used cosine similarity scores between the source and target languages using the sBERT model. We will examine concordance between human and machine judgement of the translation quality by the type of translation error.
Keywords
Large language models
Cosine similarity
Artificial intelligence
Survey methodology
Linguistic minorities
Nearly all surveys have suffered a secular decline in response rates in the last two decades, raising growing concerns about potential nonresponse bias. In recent years NORC has helped pioneer the use of big data classifiers – machine learning models which use address-level data to target subpopulations of interest at the time of sample design. These address-level big data remain available after sampling for further modelling when survey data collection is complete. We conduct an empirical analysis of the use of these big data sources at the nonresponse adjustment stage of weighting. Machine learning models are used to predict response propensities, and the performance of the resulting nonresponse adjusted weights are compared to weights derived from more traditional methods of nonresponse adjustment.
Keywords
Machine Learning
Big Data
Survey Non-Response
Survey Statistics
As AI matures, the impact it might have on how society functions is being actively pondered on. In this article, through uniform-binomial mixtures, we shed quantitative light on this matter, showing topics that unite and divide the population at an unobserved, latent level. Simultaneously, we analyze Likert scale answers we collected in collaboration with Gallup, surveying 5835 US residents on how they sense guidelines on remote work or a four-day work week or spending time on work outside of scheduled hours could impact their lives. We find the tradeoff between uncertainty and feeling switches across topics. For some, such as a four-day week or returning to in-person work, uncertainty and feeling covary, while for others, such as spending time outside of work, they anitivary.Available are GitHub codes and interactive Shiny dashboards. Follow-up of two recent works: Bhaduri, M. (2025) Altered tradeoffs between uncertainty and feeling in inherent indecisions: uniform-binomial mixture densities to gauge the impact of shifts in work operations on employees. Societal Impacts.
Bhaduri, M. (2025) Are individuals who are positive about artificial intelligence also more unsure? Patterns.
Keywords
sample surveys
Likert scale
uniform-binomial mixtures
clustering
sheltering effect
demographic diversity and persistence