05: Are machine learning methods for supervised classification truly better learners than classical LDA
Kent Eskridge
Co-Author
University of Nebraska, Statistics Department
Tuesday, Aug 4: 2:00 PM - 3:50 PM
2749
Contributed Posters
Thomas M. Menino Convention & Exhibition Center
The Increasing availability of large datasets with limited documentation or anonymized participant poses a serious threat to valid data analyses. Analyzing such data as an independent measurement without accounting for its longitudinal structure can bias results in some situations. To characterize this bias, we conducted simulations and analyzed repeated measures data as independent observations using Linear Discriminant Analysis (LDA) and machine learning (ML) algorithms: SVM, Random Forest, Neural Network, and K-Nearest Neighbor. We evaluated bias assuming independence for longitudinal datasets with compound symmetry & first-order autoregressive correlation structures, varying serial correlations (0–0.99), between-variable correlations (0.2–0.9), and variable variances. Results showed that when between-variable correlations were low, variances high, and serial correlation strong, ML models overestimated accuracy compared to LDA. LDA gave accurate results irrespective of the serial correlation levels among the observation. These results highlight the importance of understanding data generation and correctly modeling the data structure for valid inference, and reliable predictions.
Longitudinal Data
Machine Learning
Bias
Linear Discriminant Analysis
Compound Symmetry
AR(1)
Main Sponsor
Biometrics Section
You have unsaved changes.