Two-phase validation via principal components to improve efficiency in multi-model estimation from error-prone biomedical databases

Cole Manschot Speaker
Merck
 
Wednesday, Aug 5: 11:35 AM - 11:55 AM
Topic-Contributed Paper Session 
Thomas M. Menino Convention & Exhibition Center 
Two-phase sampling offers a cost-effective way to validate error-prone measurements in biomedical databases. Inexpensive or easy-to-obtain information is collected for the entire study in Phase I. Then, a subset of patients undergoes cost-intensive validation (eg, expert chart review) to collect more accurate data in Phase II. Critically, any Phase I variables can be used to strategically select Phase II, often enriched for a particular model of interest, and then missing data methods like multiple imputation can be used for estimation. However, when balancing primary and secondary analyses, competing models and priorities can result in poorly defined objectives for the most informative Phase II sampling criterion.
Extreme tail sampling (ETS), wherein patients with the smallest and largest values of a particular quantity (like a covariate or residual), can offer great statistical efficiency in two-phase studies when focusing on a single analytic objective by targeting observations with the biggest contributions to the Fisher information. We propose an intuitive, easy-to-use approach that extends ETS to balance and prioritize explaining the largest amount of variability across multiple models of interest. Using principal components, we succinctly summarize the inherent variability of all models' error-prone covariates. Then, we sample patients with the most extreme principal components for validation.
Through simulations and an application to the National Health and Nutrition Examination Survey (NHANES), the proposed sampling strategy offered simultaneous efficiency gains across multiple models of interest. Its advantages persisted across various real-world scenarios: different covariance structures, measurement error severities, validation proportions, and shared versus distinct model outcomes.
When design a validation study , concentrating on a single model may be short-sighted. Allocating resources more broadly, while still strategically, could balance multiple analytical goals simultaneously. Employing dimension reduction before sampling will allow this strategy to scale up well to big-data applications with many error-prone covariates.
Our proposed sampling strategy is implemented in the open-source R package, auditDesignR.

Keywords

Extreme tail sampling

Linear regression

Measurement error

NHANES

Principal components analysis

Source document verification