Testing the Limits of Generative AI: Do LLMs Mimic Human Evaluations in Factorial Survey Experiments

Sophie Hensgen Speaker
Institute for Employment Research
 
Emma Fössing Co-Author
Institute for Employment Research
 
Tobias Holtdirk Co-Author
LMU
 
Joseph Sakshaug Co-Author
German Institute for Employment Research
 
Tuesday, Aug 4: 11:20 AM - 11:35 AM
2547 
Contributed Papers 
Thomas M. Menino Convention & Exhibition Center 
With the rapid development of large language models (LLMs), the debate over their ability to complement or replace traditional surveys by generating "synthetic samples" has intensified. The pre-trained models draw on vast amounts of training data, potentially reflecting nuanced attitudes and behaviors that are difficult to capture through surveys alone. Recent research examined whether LLMs can mimic participants' behavior in psychological experiments, with mixed results. However, the use of LLMs at the intersection of psychological experiments and surveys – namely factorial survey experiments (FSE) – remains underexplored in the current literature, despite their potential for studying attitude and behavioral intentions. In this study, we investigate whether and to what extent LLMs can mimic human evaluations in an FSE on earnings fairness. We compare results from probability and non-probability samples with multiple synthetic samples generated by different LLMs, in which personas are matched to the characteristics of the human respondents. Our findings contribute to the growing literature on synthetic samples and thereby inform future applications of LLMs in survey research.

Keywords

Large Language Models

Factorial Survey Experiments

Probability Sample

Non-Probability Sample

Synthetic Sample 

Main Sponsor

Survey Research Methods Section