Causal Inference, Applied Modeling, and Specialized Statistical Methods

Wenjin Zhang Chair
 
Thursday, Aug 6: 8:30 AM - 10:20 AM
6057 
Contributed Papers 
Thomas M. Menino Convention & Exhibition Center 
Room: CC-252A 

Main Sponsor

IMS

Presentations

Asymptotic Distribution of Robust Effect Size Index

The Robust Effect Size Index (RESI) is a recently proposed standardized index to quantify effect magnitude across models, with growing applications in high-dimensional settings such as neuroimaging. Existing confidence interval (CI) construction relies on computationally intensive bootstrap methods. We establish a general theorem for the asymptotic distribution of RESI using a Taylor expansion, applicable to a broad class of models. Simulations under linear and logistic settings show that RESI and its CI have smaller bias and more reliable coverage than commonly used effect sizes such as Cohen's d and f. When combined with robust covariance estimation, our method provides valid inference under model misspecification, a critical aspect in neuroimaging where spatial dependence and heterogeneous variances across regions are common. Moreover, our approach substantially reduces computation time, achieving up to a 50-fold speedup compared to bootstrap procedures. Building on our earlier work on confidence sets for imaging studies, this paper offers a scalable and reliable alternative for RESI inference, significantly enhancing its applicability to neuroimaging and other complex data type 

Keywords

Semiparametric

Generalized linear model

Hypothesis testing 

Speaker

Xinyu Zhang

Co-Author(s)

Rachael Muscatello, Department of Psychiatry & Behavioral Sciences, Vanderbilt University Medical Center
Megan Jones
Blythe Corbett, Department of Psychiatry & Behavioral Sciences, Vanderbilt University Medical Center
Simon Vandekar, Vanderbilt University Medical Center

Distributional Discontinuity Design

We introduce distributional discontinuity design, a framework for studying distributional causal effects for a scalar outcome at the boundary of a discontinuity in treatment assignment (a generalization of the regression discontinuity design). Our causal estimand is the Wasserstein distance between limiting conditional outcome distributions above and below the treatment discontinuity; a single scale-interpretable measure of distribution shift. We show that this weakly bounds the average treatment effect, where equality holds if and only if the treatment effect is purely additive. Moreover, we show that the Wasserstein distance can be decomposed into squared differences in L-moments, thereby quantifying the contribution from location, scale, skewness, etc. to the overall distributional distance. This decomposition provides a novel way of encoding the heterogeneity in the treatment effect. Next, we extend this framework to distributional kink designs by evaluating the Wasserstein derivative at a deterministic policy kink; this describes the flow of probability mass through the kink. In both settings, we allow the treatment assignment to be either sharp or fuzzy. Notably, we derive 

Keywords

Regression Discontinuity Design

Regression Kink Design

Optimal Transport



Wasserstein Distance

Quantile Treatment Effects 

Speaker

Kyle Schindl, Iowa State University

Co-Author

Larry Wasserman, Carnegie Mellon University

Estimation of Out-of-Sample Sharpe Ratio for High Dimensional Portfolio Optimization

Out-of-sample Sharpe ratio is a key metric for portfolio optimization, but in high-dimensional problems naive plug-in evaluation using the sample covariance is inconsistent. We develop an estimator of the out-of-sample Sharpe ratio using only in-sample observations based on random matrix theory. In the Markowitz mean–variance framework with p/n→c∈(0,∞), where p is the portfolio dimension and n is the number of samples or time points. We propose to correct the sample covariance by a regularization matrix and provide a consistent estimator of its Sharpe ratio. The new estimator works well under either of the following conditions: (a) bounded covariance spectrum, (b) arbitrary number of diverging spikes when c<1 and (c) fixed number of diverging spikes with weak requirement on their diverging speed when c≥1. We can also extend the results to construct global minimum variance portfolio and correct out-of-sample efficient frontier. We demonstrate the effectiveness of our approach through comprehensive simulations and real data experiments. Our results highlight the potential of this methodology as a useful tool for portfolio optimization in high dimensional settings. 

Keywords

Efficient frontier recovery

High dimensionality

Portfolio allocation

Ridge regularization

Spiked covariance structure 

Speaker

Xuran Meng, University of Michigan

Co-Author(s)

Yuan Cao, University of HongKong
Weichen Wang, University of HongKong

Prototype Selection using Topological Data Analysis

Prototype selection methods compress a training set, but the existing taxonomy of condensation, edition, hybrid, competence-based, optimization-based, and clustering-based families does not include methods that operate on the multi-scale topological structure of the data. In this talk I will present three new prototype selection methods based on tools from topological data analysis (TDA) that serve as a potential foundation for a new taxonomy of topological-based selection methods. These methods are either persistence-based such as the Topological Prototype Selector (TPS) and Boundary-Conscious Topological Prototype Selector (BoundaryTPS) or graph-based with the Mapper Prototype Selector (MPS). TPS uses two sequential Rips filtrations to retain boundary-relevant and interior-typical points. BoundaryTPS is a single-stage variant whose vertex-weighted filtration concentrates retention near the decision boundary. MPS uses the Mapper graph to generate synthetic observations. We evaluate all methods against seven classical prototype selection baselines across fifteen real datasets. Our analysis shows that the persistence-based topological methods occupy a different operating point in the prototype-selection design space than existing methods. BoundaryTPS achieves the lowest mean Friedman rank on H1 persistence-diagram preservation and is significantly better than five of the seven baselines (Nemenyi, α = 0.05). TPS ranks third on the same endpoint. Both methods are more stable under fold perturbation than any chained-decision selector tested, and both inherit the source set's class proportions without label-aware machinery. Graph-based methods in contrast perform much better on downstream classification tasks and are among the top ranked against other methods. Empirically, all methods scale at worst sub-quadratically in sample size.  

Keywords

Prototype Selection

Instance Selection

Prototype Generation

Mapper Algorithm

Topological Data Analysis

Persistent Homology 

Speaker

Jordan Eckert, Auburn University

Co-Author(s)

Elvan Ceyhan, Auburn University
Henry Schenck, Auburn University

The Likelihood Ratio Test for the Multiple Signals Model for Signal Detection from fMRI Brain Images

We developed a likelihood ratio test for detecting multiple signals in functional magnetic resonance imaging (fMRI) brain images, extending classical single-signal detection methods. Using the framework of reproducing kernel Hilbert spaces, we derive the test statistic \( Y_{{\text{max}}} \) for the multiple signals model, generalizing prior approaches that utilized global maxima for signal detection. A computationally robust alternative, \( U_{{\text{max}}} \), is also introduced. The performance of these test statistics is evaluated through extensive simulation studies under varying image resolutions, sample sizes, and signal configurations. The results demonstrate that \( Y_{{\text{max}}} \) offers superior sensitivity compared to traditional \( X^2_{{\text{max}}} \) statistic, particularly under complex multi-signal scenarios. Feature importance analysis using Random Forest regression reveals that image resolution significantly impacts test outcomes. These findings provide a robust statistical tool for improved activation detection in neuroimaging applications. 

Keywords

Gaussian Random Field, Scale Space, Multiple Signals Model, Likelihood Ratio Test, Reproducing Kernel Hilbert Space, fMRI, Signal Detection, Simulation Study 

Speaker

Khalil Shafie, University of Northern Colorado

Co-Author

Annie Lu, National Taiwan University

Uniform k-tuple Partially Rank-Ordered Set Sampling

Ranked Set Sampling (RSS), introduced by McIntyre, and other
related methods, such as Partially Rank-Ordered Set Sampling
(PROSS), have shown that inclusion of a ranking mechanism produ-
ces estimators with lower variance than their simple random sample
(SRS)-based counterparts. Like RSS, PROSS takes only one measure-
ment from each partially ranked-ordered set. We propose a sampling
plan called Uniform k-Tuple Partially Rank-Ordered (UKPRSS) where a
measurement is collected from each group of a partially rank-
ordered set. This article demonstrates estimators from UKPRSS have
lower variance than their SRS counterparts. In addition, there is a
reduction in the number units needing to be screened when com-
pared to PROSS. Estimation of the mean and distribution function
are investigated theoretically. 

Keywords

Ranked Set Sampling

Partially Rank-Ordered Set Sampling

k-tuple Ranked Set Sampling

k-tuple Partially Rank-Ordered Set Sampling 

Speaker

Kaushik Ghosh, University of Nevada-Las Vegas

Co-Author

Marvin Javier

Which Covariates to Adjust for? Specification-robust Causal Inference in Observational Studies

In observational causal inference, domain knowledge often leaves multiple covariate adjustments plausible, yet which sets satisfy ignorability is untestable. Different adjustment sets can yield conflicting estimates of the average treatment effect, and standard remedies (adjusting for their union or intersection, or reporting the union or convex hull of confidence intervals) can fail or produce intervals whose width does not vanish with sample size. We propose a specification-robust procedure that returns a single point estimate and a confidence interval that is valid as long as at least one candidate adjustment set is valid and has width shrinking at the parametric $n^{-1/2}$ rate. Our approach mirrors how trimming and overlap weighting handle overlap violations:~We shift the target to a reweighted population, closest in KL-divergence to the original population, for which credible, specification-robust inference is feasible. We also provide diagnostic plots to assess the population shift and an extension to protect any function of the covariates used for reweighting, similar to calipers in matching. Synthetic and real-data examples demonstrate that our procedure provides substantially tighter confidence intervals than the convex hull while maintaining nominal coverage.  

Keywords

Causal inference

Covariate adjustment

Ignorability violations

Multiverse analysis

Observational studies

Robustness 

Speaker

Aditya Ghosh, Stanford University

Co-Author

Dominik Rothenhaeusler, Stanford University