05: Are machine learning methods for supervised classification truly better learners than classical LDA

Oluwafunmibi Fasanya Speaker
Louisiana State University Health Sciences Center
 
Kent Eskridge Co-Author
University of Nebraska, Statistics Department
 
Tuesday, Aug 4: 2:00 PM - 3:50 PM
2749 
Contributed Posters 
Thomas M. Menino Convention & Exhibition Center 
The Increasing availability of large datasets with limited documentation or anonymized participant poses a serious threat to valid data analyses. Analyzing such data as an independent measurement without accounting for its longitudinal structure can bias results in some situations. To characterize this bias, we conducted simulations and analyzed repeated measures data as independent observations using Linear Discriminant Analysis (LDA) and machine learning (ML) algorithms: SVM, Random Forest, Neural Network, and K-Nearest Neighbor. We evaluated bias assuming independence for longitudinal datasets with compound symmetry & first-order autoregressive correlation structures, varying serial correlations (0–0.99), between-variable correlations (0.2–0.9), and variable variances. Results showed that when between-variable correlations were low, variances high, and serial correlation strong, ML models overestimated accuracy compared to LDA. LDA gave accurate results irrespective of the serial correlation levels among the observation. These results highlight the importance of understanding data generation and correctly modeling the data structure for valid inference, and reliable predictions.

Keywords

Longitudinal Data

Machine Learning

Bias

Linear Discriminant Analysis

Compound Symmetry

AR(1) 

Main Sponsor

Biometrics Section