Wednesday, Aug 5: 10:30 AM - 12:20 PM
6455
Contributed Papers
Thomas M. Menino Convention & Exhibition Center
Room: CC-108
Main Sponsor
Quantum Computing in Statistics and Machine Learning Interest Group
Presentations
Gaussian process models(GP) are widely used for surrogate modeling because they provide flexible regression with uncertainty quantification, but their use is limited by high-dimensional inputs and cubic computational cost. Most existing methods address dimension reduction(DR) and scalability separately, often through multi-stage or approximate procedures that weaken uncertainty propagation. We propose a unified, fully Bayesian framework that jointly performs input DR and GP modeling within a single hierarchical model. DR is achieved through an orthonormal projection matrix with priors on the Stiefel manifold, enabling coherent posterior learning via Hamiltonian Monte Carlo with geodesic dynamics. We extend this framework to Deep GP with built-in DR to model complex nonlinear systems. To scale inference to large datasets, we incorporate Vecchia sparse covariance approximations, reducing computational complexity from cubic to near-linear while preserving predictive accuracy and calibrated uncertainty. Extensive numerical studies show the Bayesian method with Vecchia scaling gives better predictions and more reliable uncertainty, providing a robust alternative to existing methods.
Keywords
Bayesian inference
Dimension reduction
Deep Gaussian processes
Hamiltonian Monte Carlo
Uncertainty quantification
Vecchia approximation
Machine-generated probability predictions are essential to modern classification tasks such as image classification. A model is well calibrated when its predicted probabilities correspond to observed event frequencies. Despite the importance of multicategory recalibration, existing approaches are limited by comparing calibration between models rather than assessing the calibration of a single model, requiring under-the-hood access to model access, and producing outputs that are difficult for human analysts to understand. To address these limitations, we propose Multicategory Linear Log Odds (MCLLO) recalibration, which incorporates a likelihood ratio hypothesis test to assess calibration, doesn't require internal model access, and yields interpretable results. We demonstrate the effectiveness of MCLLO through simulations and three case studies involving image classification via convolutional neural network, obesity analysis via random forest, and ecology via regression modeling. We compare MCLLO to four comparator recalibration techniques using both our hypothesis test and the Expected Calibration Error to show that our method works well alone and in concert with other methods.
Keywords
Multiclass Classification
Calibration
Machine Learning
Confidence Scores
Likelihood Ratio Test
Synthetic data generation is a valuable tool for the analysis of health data with varied use cases, such as privacy protection and data supplementation for rare events. We present a pilot study evaluating the use of quantum-enhanced generative adversarial networks (GANs) to generate synthetic records based on structured fields in the FDA Adverse Event Reporting System (FAERS) to evaluate if these enhancements can provide deeper or equivalent insights for health data using smaller samples. The FAERS tracks adverse drug reactions, including rare reactions and reactions associated with rarely prescribed drugs. The model is implemented in a simulated quantum environment run on a classical computer using open-source tools. The presentation will conclude with an evaluation based on data fidelity and data utility in comparison to classical GAN models. Further, limitations of quantum simulations will be discussed.
Keywords
pharmacovigilance
generative adversarial networks
FAERS