Problems in Modern Inference: Distribution Free Prediction and Goodness of Fit Testing
Description
This dissertation studies two important directions in modern statistical inference: distribution-free prediction and hypothesis testing. These directions are increasingly important in an era where complex machine-learning algorithms are widely deployed, but their statistical behavior is often difficult to characterize. In such settings, classical guarantees based on correctly specified parametric models can be unreliable or too restrictive. Distribution-free prediction aims to construct prediction sets that are valid without assumptions on the fitted model or the underlying distribution, typically under exchangeability, while modern goodness-of-fit testing seeks flexible procedures that can detect structured deviations from complex null models while maintaining rigorous type-I error control. The central theme of this dissertation is to develop methods that preserve finite-sample or approximate frequentist guarantees while remaining useful in realistic settings.
The first part of the dissertation contributes to conformal prediction. Conformal prediction constructs prediction sets around the output of a fitted model with coverage guarantees that do not rely on the model being correct. Standard full and split conformal methods, however, assume exchangeability between training and test data. We study two challenges that arise in practice. First, we consider deviations from exchangeability caused by covariate shift between the training and test distributions, focusing on the setting where the shift is explained by a discrete covariate. In this group-weighted setting, we show that existing guarantees for weighted conformal prediction can be substantially sharpened, especially when the number of groups is large. We then extend this perspective to label-weighted conformal prediction for settings with rare classes, showing how to obtain statistically efficient guarantees for macro-coverage. Second, we address the tradeoff between statistical efficiency and computation faced by full conformal prediction or other conformal prediction approaches. We develop a tournament-correction framework that converts a broad class of approximations to full conformal prediction into procedures with rigorous distribution-free coverage guarantees.
The second part of the dissertation contributes to flexible goodness-of-fit testing. We develop approximate co-sufficient sampling via Bayes (aCSS-B), a method for generating artificial data sets that are approximately exchangeable with the observed data under a composite null hypothesis. This enables valid \(p\)-values using essentially arbitrary test statistics, allowing the test statistic to be tailored to the scientific alternative of interest rather than restricted to generic likelihood-based summaries. We prove approximate exchangeability for the proposed samples and derive approximate type-I error control for the resulting tests.
Files
Final thesis.pdf
Files
(1.2 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:77359d1b7245ea250e62de404f45eaea
|
1.2 MB | Preview Download |