Interval Estimation, Hypothesis Testing, and Power Analysis for Fβ scores

Chih-Yuan Hsu Speaker
Vanderbilt University Medical Center
 
Qi Liu Co-Author
Vanderbilt University
 
Yu Shyr Co-Author
Vanderbilt University Medical Center
 
Monday, Aug 3: 11:20 AM - 11:35 AM
2053 
Contributed Papers 
Thomas M. Menino Convention & Exhibition Center 
Machine learning and artificial intelligence are increasingly applied to medical diagnostics and clinical decision-making. The F1 score and its generalized form, the Fβ score, are commonly used to evaluate diagnostic performance because they effectively balance precision and sensitivity, particularly in the presence of class imbalance. Despite their popularity, rigorous statistical inference and power analysis for these scores remain limited. In this study, we develop methods for interval estimation, hypothesis testing, and power and sample size calculation for both single and comparative F1 and Fβ scores. Extensive simulations demonstrate the accuracy and robustness of the proposed methods. We further showcase their practical utility through real-world biomedical classification applications. These methods enable principled evaluation and comparison of classifiers using F1 and Fβ scores, providing reliable uncertainty quantification and informed sample size planning. Finally, we extend the proposed framework from independent classifier comparisons to settings with correlated classifiers.

Keywords

F1 score

Class imbalance

Precision

Sensitivity

Sample size calculation 

Main Sponsor

Biopharmaceutical Section