The V-fold jackknife for semiparametric inference: variance estimation, confidence intervals, and simultaneous confidence bands

Generative AI & LLMs
Published: arXiv: 2607.22493v1
Authors

Yi Li Ashkan Ertefaie Mark van der Laan

Abstract

For decades, the bootstrap has been a default tool for statistical inference because of its broad applicability and minimal analytic requirements. Although its validity is well understood for smooth parametric estimators, its theoretical properties for many modern semiparametric and machine-learning estimators remain largely unstudied. Nevertheless, bootstrap procedures are often used routinely in such settings, even when their validity is unknown and their computational cost is substantial. We develop the $V$-fold jackknife as a computationally efficient and theoretically justified alternative for semiparametric inference. It requires only $V$ leave-fold-out refits and uses the empirical dispersion of jackknife pseudo-values to quantify uncertainty, without deriving or evaluating an influence function. For regular asymptotically linear estimators of pathwise differentiable parameters, we show that, for fixed $V$, the Studentized $V$-fold jackknife statistic converges to a $t$-distribution with $V-1$ degrees of freedom, giving valid confidence intervals even though the jackknife variance estimator does not converge in probability. When $V\to\infty$, we establish consistency of the variance estimator at rate $V^{-1/2}$, allowing $V$ to diverge slowly, for example at rate $\log n$. We also develop simultaneous confidence bands based on the correct componentwise-Studentized limiting distribution. Finally, we extend the theory to generalized asymptotically linear estimators with diverging influence-function variance and slower-than-$\sqrt n$ convergence; scale invariance of Studentization eliminates the need to know the effective convergence rate. Simulations on the average treatment effect, Kaplan--Meier survival curve, and highly adaptive lasso dose-response curves confirm reliable inference, including where influence-function-based standard errors are anti-conservative or unstable.

Paper Summary

Problem
The main problem addressed in this research paper is the limitations of the bootstrap method for statistical inference, particularly in modern applications involving high-dimensional nuisance estimation, adaptive model selection, and machine learning algorithms. The ordinary nonparametric bootstrap is often unclear in these settings, and its validity is not well understood.
Key Innovation
The key innovation of this work is the development of the V-fold jackknife as a computationally efficient and theoretically justified alternative for semiparametric inference. The V-fold jackknife procedure requires only V leave-fold-out refits and uses the empirical dispersion of jackknife pseudo-values to quantify uncertainty, without requiring analytic derivation or numerical evaluation of an influence function.
Practical Impact
This research has significant practical implications for statistical inference in various fields, including medicine, economics, and social sciences. The V-fold jackknife provides a reliable and computationally efficient method for constructing confidence intervals and simultaneous confidence bands, which are essential for making informed decisions in these fields. The method is particularly useful for estimating parameters in high-dimensional settings, where the ordinary bootstrap may not be valid.
Analogy / Intuitive Explanation
Imagine you're trying to estimate the average height of a population, but you only have a small sample of people to work with. The bootstrap method is like taking many random samples from your original sample, and then using the average height of each sample to estimate the population average. However, this method can be problematic if the original sample is not representative of the population. The V-fold jackknife is like taking a smaller number of larger samples, and then using the average height of each sample to estimate the population average. This method is more reliable and efficient, especially when dealing with high-dimensional data.
Paper Information
Categories:
stat.ME math.ST stat.CO stat.ML
Published Date:

arXiv ID:

2607.22493v1

Quick Actions