arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

廉价而强大的监督子空间检验:PLS 的逐分量推断

Cheap and Powerful Tests for Supervised Subspaces: Per-Component Inference for PLS

Paweł Lenartowicz, Hubert Plisiecki

arXiv 2609.36307首次发表:更新:

发表机构

Society for Open Science (Stowarzyszenie na Rzecz Otwartej Nauki); Centre for Brain Research, Jagiellonian University; IDEAS Research Institute(开放科学协会; 雅盖隆大学脑研究中心; IDEAS研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对 PLS 等监督子空间推断昂贵或缺失的问题,提出基于留出 OLS 重拟合的两种检验,具有更高功效且成本低,并发布多语言绑定库。

AI 中文摘要

偏最小二乘(PLS)回归在高维 X 中提取少数与结果对齐的方向,并广泛应用于应用科学领域,但对所得拟合的推断要么昂贵、有偏且不被鼓励,要么缺失。我们将推断简化为对监督子空间进行留出 OLS 重拟合,这是 PLS、监督 PCA 和线性探针共享的原语,并提供两种基于留出相关性的检验:一种 Nadeau-Bengio 校正的渐近 t 检验作为快速近似,以及一种具有相当功效的置换检验,在结果-预测变量独立且行独立同分布的情况下有限样本有效。留出预测在监督跨度的任何正交重基下不变,因此诸如 varimax 的可解释基继承联合声明但不继承逐轴 p 值;逐分量声明来自对 PLS 提取顺序的固定序列检验。我们在合成几何、两个近红外化学计量数据集和跨语言词嵌入回归上验证;精确检验也适用于监督 PCA 和岭探针。所提出的检验比 CV-置换-Q^2 具有更高功效,且成本仅为其一小部分。对 n 和 X 谱的预运行检查说明近似何时安全。我们发布了一个带有 Python、R 和 Julia 绑定的 Rust 库,以及一个 Python 文本流水线。

英文摘要

Partial Least Squares (PLS) regression extracts a few outcome-aligned directions in a high-dimensional X and is widely used across applied science, but inference on the resulting fit is either expensive, biased and discouraged, or absent. We reduce inference to held-out OLS refits of the supervised subspace, a primitive shared by PLS, supervised PCA, and linear probes, and supply two tests using held-out correlations: a Nadeau-Bengio corrected asymptotic t-test as a fast approximation, and a permutation test with comparable power, finite-sample valid under outcome-predictor independence and iid rows. Held-out predictions are unchanged under any orthogonal rebasing of the supervised span, so an interpretable basis such as varimax inherits the joint claim but not a per-axis p-value; per-component claims come from a fixed-sequence test on the PLS extraction order. We validate on synthetic geometries, two NIR chemometric datasets, and cross-lingual word-embedding regressions; the exact test also transfers to supervised PCA and a ridge probe. The proposed tests have more power than CV-permutation-Q^2, at a fraction of its cost. A pre-run check on n and the spectrum of X says when the approximation is safe. We release a Rust library with Python, R, and Julia bindings, plus a Python text pipeline.

Comments41 pages, 5 figures, 28 tables. Accepted at NeurIPS 2026. Code and results: https://doi.org/10.5281/zenodo.23004298

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑