Testing the Limits of Generative AI: Do LLMs Mimic Human Evaluations in Factorial Survey Experiments
Tuesday, Aug 4: 11:20 AM - 11:35 AM
2547
Contributed Papers
Thomas M. Menino Convention & Exhibition Center
With the rapid development of large language models (LLMs), the debate over their ability to complement or replace traditional surveys by generating "synthetic samples" has intensified. The pre-trained models draw on vast amounts of training data, potentially reflecting nuanced attitudes and behaviors that are difficult to capture through surveys alone. Recent research examined whether LLMs can mimic participants' behavior in psychological experiments, with mixed results. However, the use of LLMs at the intersection of psychological experiments and surveys – namely factorial survey experiments (FSE) – remains underexplored in the current literature, despite their potential for studying attitude and behavioral intentions. In this study, we investigate whether and to what extent LLMs can mimic human evaluations in an FSE on earnings fairness. We compare results from probability and non-probability samples with multiple synthetic samples generated by different LLMs, in which personas are matched to the characteristics of the human respondents. Our findings contribute to the growing literature on synthetic samples and thereby inform future applications of LLMs in survey research.
Large Language Models
Factorial Survey Experiments
Probability Sample
Non-Probability Sample
Synthetic Sample
Main Sponsor
Survey Research Methods Section
You have unsaved changes.