Interval Estimation, Hypothesis Testing, and Power Analysis for Fβ scores
Qi Liu
Co-Author
Vanderbilt University
Yu Shyr
Co-Author
Vanderbilt University Medical Center
Monday, Aug 3: 11:20 AM - 11:35 AM
2053
Contributed Papers
Thomas M. Menino Convention & Exhibition Center
Machine learning and artificial intelligence are increasingly applied to medical diagnostics and clinical decision-making. The F1 score and its generalized form, the Fβ score, are commonly used to evaluate diagnostic performance because they effectively balance precision and sensitivity, particularly in the presence of class imbalance. Despite their popularity, rigorous statistical inference and power analysis for these scores remain limited. In this study, we develop methods for interval estimation, hypothesis testing, and power and sample size calculation for both single and comparative F1 and Fβ scores. Extensive simulations demonstrate the accuracy and robustness of the proposed methods. We further showcase their practical utility through real-world biomedical classification applications. These methods enable principled evaluation and comparison of classifiers using F1 and Fβ scores, providing reliable uncertainty quantification and informed sample size planning. Finally, we extend the proposed framework from independent classifier comparisons to settings with correlated classifiers.
F1 score
Class imbalance
Precision
Sensitivity
Sample size calculation
Main Sponsor
Biopharmaceutical Section
You have unsaved changes.