Evaluating LLM-Based Survey Questionnaire Translations with Human Evaluators and Cosine Similarities
Jay Kim
Co-Author
University of Michigan
Tuesday, Aug 4: 11:35 AM - 11:50 AM
3660
Contributed Papers
Thomas M. Menino Convention & Exhibition Center
Large language models (LLMs) may advance multilingual data collection instruments and tackle barriers in researching linguistic minorities, by assisting with language translations. However, the quality of LLM-based survey questionnaire translation and the method for such quality evaluation are yet to be understood.
This study compares human expert evaluation and automated similarity metrics for assessing LLM-based translation. We translated a survey questionnaire with 35 questions across three topics (socio-demographics, social networks, and cognitive health) from English (source) into four target languages (Spanish, Chinese, Korean, and Vietnamese) using commercial and open-source LLMs. Multiple human experts recorded the translation quality of each question, including the existence and type of translation errors, and the level of post-editing necessary to use the translated version. For the automated metric, we used cosine similarity scores between the source and target languages using the sBERT model. We will examine concordance between human and machine judgement of the translation quality by the type of translation error.
Large language models
Cosine similarity
Artificial intelligence
Survey methodology
Linguistic minorities
Main Sponsor
Survey Research Methods Section
You have unsaved changes.