32: New Dogs Learn Old Tricks: Reducing Bias in Small-Sample Machine Learning via Design of Experiments

Jace Ritchie Speaker
 
Alan Wisler Co-Author
 
Monday, Aug 3: 2:00 PM - 3:50 PM
3764 
Contributed Posters 
Thomas M. Menino Convention & Exhibition Center 
In a world of big data, machine learning (ML) practice tends to rely on asymptotic assumptions to achieve low bias when estimating generalizable test error. However, for small-sample problems, ML is being used even though these assumptions no longer hold. This is problematic, as the common practice of using the combination of k-fold cross-validation and grid search to tune hyperparameters results in high bias between the validated and true generalization errors. We propose using optimal design of experiments principles to fit a response surface to a space-filling design that can be used to generate an optimal hyperparameter set. By performing Monte Carlo simulations on real datasets, we show that this approach generates hyperparameter sets with similar performance to grid search while also drastically decreasing the discrepancy between validation and generalization error. This decreased bias will aid practitioners in accurately assessing model performance without a significant reduction in predictive accuracy.

Keywords

Optimal Experimental Design

Small Sample Machine Learning

Hyperparameter Tuning 

Main Sponsor

Section on Statistical Learning and Data Science