Targeted Validation, Broader Benefit: Advancing Public Health with Two-Phase Designs

Lucy D'Agostino McGowan Chair
Wake Forest University
 
Sarah Lotspeich Organizer
Wake Forest University
 
Wednesday, Aug 5: 10:30 AM - 12:20 PM
1559 
Topic-Contributed Paper Session 
Thomas M. Menino Convention & Exhibition Center 
Room: CC-211 
This session highlights recent innovations in two-phase studies, especially extreme-tail and related targeted sampling strategies, to emphasize how design choices yield broader public health benefits. Our speakers will showcase diverse applications: reducing bias in community-level food access metrics, extending causal inference to new populations despite incomplete covariate overlap, balancing efficiency across multiple models in pharmaceutical and observational studies, and improving cost-effective evaluation of cancer screening tests. Collectively, these talks demonstrate how modern applications and extensions of extreme tail sampling and two-phase designs advance beyond traditional paradigms to meet the needs of complex health research. In line with the JSM theme, Communities in Action: Advancing Society, the session underscores how innovations in extreme-tail and targeted two-phase sampling can generate more accurate, equitable, and actionable insights for the communities we aim to serve.

Applied

Yes

Main Sponsor

Biometrics Section

Co Sponsors

ENAR
Section on Statistics in Epidemiology

Presentations

A Novel Framework for Addressing Disease Under-Diagnosis Using EHR Data

Effective treatment of medical conditions begins with an accurate diagnosis. However, many conditions are often underdiagnosed, either being overlooked or diagnosed after significant delays. Electronic Health Records (EHRs) contain extensive patient health information, offering an opportunity to probabilistically identify underdiagnosed individuals. The rationale is that both diagnosed and underdiagnosed patients may display similar health profiles in EHR data, distinguishing them from condition-free patients. Thus, EHR data can be leveraged to develop models that assess an individual's risk of having a condition. To date, this opportunity has largely remained unexploited, partly due to the lack of suitable statistical methods. The key challenge is the positive-unlabeled EHR data structure, which consists of data for diagnosed ("positive") patients and the remaining ("unlabeled") that include underdiagnosed patients and many condition-free patients. Therefore, data for patients who are unambiguously condition-free, essential for developing risk assessment models, is unavailable. To overcome this challenge, we propose ascertaining condition statuses for a small subset of unlabeled patients. We develop a novel statistical method for building accurate models using this supplemented EHR data to estimate the probability that a patient has the condition of interest. Building on the developed risk prediction model, we further study the potential factors that may contribute to under-diagnosis. We establish the asymptotic properties of the proposed methods. Numerical simulation studies and real data applications are also conducted to assess the performance of the proposed methods. 

Keywords

Risk Prediction

Electronic Health Records

Unlabeled Data

Partial Validation 

Speaker

Weidong Ma, Department of Biostatistics, Epidemiology and Informatics, University of Pennsylvania Perelman School of Medicine

Strategically validating error-prone food access measures using map-based software with efficient, equity-focused analyses in mind

Quantifying neighborhood food environments and understanding their relationships with residents' health is a public health priority. Using simple, error-prone food access metrics (like the shortest straight-line routes) introduces measurement error and leads to bias in downstream statistical models, but measuring the more-accurate, map-based ones (like the shortest map-based driving routes) for entire studies is often implausible. Fortunately, adopting a two-phase design can harness the best of both metrics by combining the error-prone values for the entire study and the more-accurate ones for a chosen subset. Fortunately, this subset to be validated with map-based food access measures can be strategically chosen to not only reduce bias but further improve statistical efficiency in modeling relationships between neighborhood health and the food environment. Using simulations and data for the Piedmont Triad Region of North Carolina, we evaluate various validation sampling designs as we quantify the associations of diabetes and obesity with neighborhood-level access to healthy foods. 

Keywords

Measurement error

Partial validation

Health equity

Food environment

Extreme tail sampling

Driving distance 

Speaker

Sarah Lotspeich, Wake Forest University

Two-phase designs for cost-effective evaluation of cancer screening tests

Screening tests are crucial for detecting diseases at preclinical stages when timely intervention can prevent progression to more severe conditions. Advances in technology have facilitated the development of screening tests based on novel markers, but evaluation of their performance using biospecimens from large cohorts can be a logistical and financial challenge. Two-phase designs offer a cost-effective solution by allowing inference when expensive marker measurements are performed on only a carefully selected subsample. While traditional two-phase designs are often focused on estimating associations between a marker and outcome, they can effectively be extended to evaluate the clinical performance of a test, such as the estimation of positive predictive value (PPV, the risk in test positives) and complementary negative predictive value (cNPV, the risk in test negatives). We propose a novel two-phase design for efficiently evaluating the risk stratification utility of screening tests in distinguishing between high- and low-risk individuals for both current asymptomatic disease and future disease development. Using biospecimens from screening studies, our methodology is developed to accommodate cohorts that include both pre-existing cases at an initial screening visit and new cases identified during follow-up. We demonstrate the efficiency gains of our proposed design compared to other subsampling schemes through simulation and illustrate its application in the motivating study evaluating the p16/ki-67 dual-stain test for managing HPV-positive women in cervical cancer screening. Data from Kaiser Permanente Northern California on a screening cohort are used along with stored biospecimen samples in this analysis.  

Keywords

screening test evaluation

risk stratification

human papillomavirus (HPV) and cervical cancer

two-phase design

mixture model 

Speaker

Fangya Mao

Co-Author(s)

Richard Cook, University of Waterloo
Thomas Lorey, Kaiser Permanente Northern California
Nicolas Wentzensen, National Cancer Institute
Li Cheung, National Cancer Institute

Two-phase validation via principal components to improve efficiency in multi-model estimation from error-prone biomedical databases

Two-phase sampling offers a cost-effective way to validate error-prone measurements in biomedical databases. Inexpensive or easy-to-obtain information is collected for the entire study in Phase I. Then, a subset of patients undergoes cost-intensive validation (eg, expert chart review) to collect more accurate data in Phase II. Critically, any Phase I variables can be used to strategically select Phase II, often enriched for a particular model of interest, and then missing data methods like multiple imputation can be used for estimation. However, when balancing primary and secondary analyses, competing models and priorities can result in poorly defined objectives for the most informative Phase II sampling criterion.
Extreme tail sampling (ETS), wherein patients with the smallest and largest values of a particular quantity (like a covariate or residual), can offer great statistical efficiency in two-phase studies when focusing on a single analytic objective by targeting observations with the biggest contributions to the Fisher information. We propose an intuitive, easy-to-use approach that extends ETS to balance and prioritize explaining the largest amount of variability across multiple models of interest. Using principal components, we succinctly summarize the inherent variability of all models' error-prone covariates. Then, we sample patients with the most extreme principal components for validation.
Through simulations and an application to the National Health and Nutrition Examination Survey (NHANES), the proposed sampling strategy offered simultaneous efficiency gains across multiple models of interest. Its advantages persisted across various real-world scenarios: different covariance structures, measurement error severities, validation proportions, and shared versus distinct model outcomes.
When design a validation study , concentrating on a single model may be short-sighted. Allocating resources more broadly, while still strategically, could balance multiple analytical goals simultaneously. Employing dimension reduction before sampling will allow this strategy to scale up well to big-data applications with many error-prone covariates.
Our proposed sampling strategy is implemented in the open-source R package, auditDesignR. 

Keywords

Extreme tail sampling

Linear regression

Measurement error

NHANES

Principal components analysis

Source document verification 

Speaker

Cole Manschot, Merck

Discussant

Speaker

Sebastien Haneuse, Harvard T.H. Chan School of Public Health