Statistical Methods and Case Studies in Text Analysis

Carol Haney Chair
Qualtrics
 
Tuesday, Aug 4: 8:30 AM - 10:20 AM
6437 
Contributed Papers 
Thomas M. Menino Convention & Exhibition Center 
Room: CC-107C 

Main Sponsor

Section on Text Analysis

Presentations

Reproducing Expert Judgement with Shortened Surveys

Patient-reported outcome measures allow us to measure constructs that are often not directly observable. The items are summarized through a simple unweighted summed score or through more advanced methods, including factor analysis, and item response theory. Shortening questionnaires reduces respondent burden, cost, and can increase data quality. However, shortened forms may have lower precision in latent trait estimation, but can they be used for accurate prediction of a diagnosis or expert judgment? We use a Markov Chain Monte Carlo algorithm to find shortened forms that maintain diagnostic accuracy. We demonstrate its efficiency through simulation studies, and apply it to a screener for alcohol use disorder. 

Keywords

Patient Reported Outcome Measures

Shortening Surveys

Markov Chain Monte Carlo 

Speaker

Daphna Harel, New York University

Co-Author

Klint Kanopka, NYU

Dynamic Topic Modeling with a Higher-Order Hypergraphical Representation

Most traditional topic models represent text using multinomial likelihood. These models treat documents as collections of independent word counts, combining word occurrence and repetition into a single probabilistic mechanism. However, this formulation obscures higher-order word interaction patterns at the document level and limits model expressiveness. To address these limitations, we propose a higher-order, hypergraphical text representation. In this model, each document induces a hyperedge connecting all co-occurring words, and repetition intensity is encoded as a node weight. This construction yields a novel hypergraph-induced multinomial distribution that decouples word occurrence from repetition intensity via support-dependent normalization. Building on this framework, we develop a dynamic topic modeling approach based on low-rank factorization that can accommodate evolving topic semantics over time. Under suitable initialization and regularity conditions, we establish local convergence guarantees and derive non-asymptotic error bounds. Numerical experiments on the synthetic datasets and real-world corpora demonstrate the advantages of our approach over existing methods. 

Keywords

Topic model

Dynamic modeling

Hypergraph

Text representation

Nonconvex optimization

Low-rank factorization 

Speaker

Hanjia Gao, University of California, Irvine

Co-Author(s)

Hanwen Ye
Qing Nie, University of California, Irvine
Annie Qu, University of California At Irvine

Are the Gospels and Acts historical? Testing names for goodness-of-fit before, during, and after

Are the personal names in the Gospels and Acts historically grounded? In a 2024 study published in the Journal for the Study of the Historical Jesus, we applied goodness-of-fit tests to compare name distributions in the Gospels and Acts with a historical reference database of over 2,200 attested individuals from Palestine (4 BCE–73 CE). The results showed that Gospel and Acts name frequencies fit this historical reference distribution at least as well as those found in the works of Josephus, and significantly better than name distributions derived from ancient fictional texts and modern historical novels.
Because name distributions vary over time, we extend this analysis by applying the same goodness-of-fit methodology to reference datasets from periods preceding (330–5 BCE) and following (74–135 CE; 136–200 CE) the time of the Gospels and Acts. We find that the Gospel and Acts names fit the contemporary period only, and fail to fit both earlier and later periods. This temporal specificity provides a discriminative test between historical and fictional name generation and is consistent with the names reflecting the onomastic environment of their purported time of composition. 

Keywords

goodness-of-fit

goodness of fit

Bible

Gospel

name distribution

onomastic 

Speaker

Jason Wilson, Biola University

Co-Author

Luuk Van de Weghe, Independent Scholar

Lying More and Less with AI: How AI Shapes Commitment and Honesty in Random Outcome Task Reporting

This paper investigates how generative AI alters the effectiveness of honesty commitments
in unmonitored work settings. Building on previous work that ex-ante honesty oaths reduce lying and shirking by between 5-10%, we conducted a between-subjects experiment in which participants composed brief honesty statements (e.g., oaths, honor codes) under varying AI-assistance conditions before self-reporting outcomes from an incentivized coin-flip task. Unassisted honesty statements reduced both aggregate misreporting and maximal outcome inflation (up to 14%), and shirking by 10%. Optional collaboration with a large language model assistant greatly amplified these effects. However, when participants blindly delegated statement composition to AI, honesty dissipated, and shirking rose sharply. We analyzed the semantic and linguistic structure of the honesty statements themselves using language AI and contemporary natural language processing. This text analyses revealed substantial differences between solely human-composed and AI-assisted statements. Human-composed statements primarily emphasized personal honesty commitments, integrity, and individualized moral ownership, whereas AI-assisted statements exhibited more formalized, procedural, and institutionalized ethical language. AI-assisted statements were also longer, more lexically complex, semantically convergent, and more procedurally framed than human. We interpret this shift as a movement from personal moral ownership toward proceduralized ethical framing. Importantly, stronger proceduralization is associated with behavioral delegation patterns and elevated misreporting. Collectively, our findings suggest that AI does not simply strengthen or weaken honesty commitments. Rather, AI changes the cognitive and linguistic architecture through which honesty is framed, and AI delegation may hollow out personal moral ownership and transform ethical reasoning. 

Keywords

dishonesty

artificial intelligence

generative AI

large language models

lying

oaths 

Speaker

J Jobu Babin, University of Northern Iowa

Co-Author

Haritima Chauhan, DePaul University

Integrating data science with qualitative coding

With the ubiquity of data science and AI tools, research teams face growing client demands for speedy analyses and quick-turnaround deliverables. However, those tools are not as easily adapted to qualitative coding, which often involves highly nuanced information and unstructured or semi-structured data. Therefore, many qualitative researchers continue to perform the labor-intensive work of manually coding what can amount to thousands of pages of text-based data, straining time and project budgets.

This presentation offers a successful case study for integrating data science tools into qualitative coding to increase efficiency without compromising analytic rigor. We describe the use of large language models to support cleaning and coding qualitative data, and produce structured output that can be seamlessly imported into qualitative analysis platforms such as NVivo for deeper, researcher led coding and interpretation. We also address how to evaluate whether a project is an appropriate candidate for this hybrid approach, discussing when it adds value and efficiency and when it does not. 

Keywords

qualitative research

data science

AI tools

qualitative coding

large language models 

Speaker

Lindsay Giesen, Westat

Co-Author(s)

Rashi Saluja
Gizem Korkmaz, Westat