Evaluating LLM-Based Survey Questionnaire Translations with Human Evaluators and Cosine Similarities

Mi Huynh Speaker
 
Sunghee Lee Co-Author
University of Michigan
 
Yajuan Si Co-Author
University of Michigan
 
Mengyao Hu Co-Author
 
Stephanie Morales Co-Author
 
Jay Kim Co-Author
University of Michigan
 
Felix Baez-Santiago Co-Author
University of Michigan
 
Beining Niu Co-Author
University of Michigan
 
Caleb Crouch Co-Author
University of Michigan
 
Mengdi Ji Co-Author
University of Michigan
 
Tuesday, Aug 4: 11:35 AM - 11:50 AM
3660 
Contributed Papers 
Thomas M. Menino Convention & Exhibition Center 
Large language models (LLMs) may advance multilingual data collection instruments and tackle barriers in researching linguistic minorities, by assisting with language translations. However, the quality of LLM-based survey questionnaire translation and the method for such quality evaluation are yet to be understood.

This study compares human expert evaluation and automated similarity metrics for assessing LLM-based translation. We translated a survey questionnaire with 35 questions across three topics (socio-demographics, social networks, and cognitive health) from English (source) into four target languages (Spanish, Chinese, Korean, and Vietnamese) using commercial and open-source LLMs. Multiple human experts recorded the translation quality of each question, including the existence and type of translation errors, and the level of post-editing necessary to use the translated version. For the automated metric, we used cosine similarity scores between the source and target languages using the sBERT model. We will examine concordance between human and machine judgement of the translation quality by the type of translation error.

Keywords

Large language models

Cosine similarity

Artificial intelligence

Survey methodology

Linguistic minorities 

Main Sponsor

Survey Research Methods Section