arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于机器学习模型的序列条件独立性检验

Sequential Conditional Independence Testing with Machine Learning Models

Angel Reyero-Lobo, Michele Meziu, Sebastian Uriel Arias, Peter Grünwald

arXiv 2610.11388首次发表:更新:

发表机构

Université de Toulouse; Inria Paris-Saclay; CWI; Leiden University(图卢兹大学; 巴黎萨克雷国家信息与自动化研究所; 荷兰数学与计算机科学研究学会; 莱顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对条件独立性检验问题,解释了直接检验可交换性的e-变量实践功效更高的现象,探索中间零假设以减少误差,并提供适应三重鲁棒性的估计误差界,实现快速收敛。

AI 中文摘要

条件独立性检验是科学发现中的普遍问题。广泛采用的模型-X假设将建模负担从输出对输入的依赖转移到输入内部的依赖。对数最优e-变量已在该场景下得到研究,但如何将机器学习模型纳入其设计仍不明确。其他方法直接检验可交换性,得到的e-变量理论上功效较低,但实践中功效更高。我们通过将误差分解为零假设放大、近似和估计误差来解释这一现象。该分解表明,GRO e-变量估计因近似和估计误差更差而可被超越,我们探索模型-X条件独立性与可交换性之间的中间零假设以减少这些误差。此外,模型-X假设通常仅在估计误差范围内成立,使精确的I类错误保证失效。我们提供了适应三重鲁棒性结果的估计误差界,实现了快速收敛速率。

英文摘要

Conditional independence testing is a ubiquitous problem in scientific discovery. The widely employed model-X assumption shifts the modelling burden from the dependence of the output on the inputs to the dependencies within the inputs. Log-optimal e-variables have been studied in this setting, but it remains unclear how to incorporate machine learning models into their design. Other approaches test exchangeability directly, yielding an e-variable with lower power in theory but, surprisingly, higher power in practice. We explain this phenomenon by decomposing the error into null enlargement, approximation, and estimation error. The decomposition shows that GRO e-variable estimates can be beaten because of their worse approximation and estimation errors, and we explore intermediate null hypotheses between model-X conditional independence and exchangeability to reduce these errors. Moreover, the model-X assumption often only holds up to an estimation error, invalidating exact type-I error guarantees. We provide estimation error bounds that accommodate triple robustness results, achieving fast convergence rates.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑