arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15332stat.MEstat.ML

GFCM:一种用于因果发现的尾部敏感混合型条件独立性检验方法

GFCM: A Tail-Sensitive Mixed-Type Conditional Independence Test for Causal Discovery

Pavel Averin, Theodoros Moysiadis, Ioannis Katakis

AI总结:

本文提出GFCM,一种适用于混合型数据、对协方差之外的尺度与尾部依赖敏感的因果发现条件独立性检验方法,可在PC框架内保持校准与低骨架错误,恢复协方差方法遗漏的尾部边。

AI中文摘要:

基于约束的因果发现方法(如PC和FCI)依赖于其条件独立性检验。偏相关系数与广义协方差测度(Generalised Covariance Measure, GCM)仅能检测残差的条件协方差,因此会遗漏均值的非线性部分、尺度以及尾部中的依赖关系。而检测更多依赖关系的检验方法要么在PC框架内存在偏差、不具备可扩展性,要么仅适用于连续数据,要么并非针对尾部设计。本文提出的广义特征协方差测度(Generalised Feature Covariance Measure, GFCM)是一种有效、对协方差之外的依赖关系敏感、在PC框架内具有鲁棒性且适用于混合型数据的检验方法。该方法在一组可配置的条件均值为零的残差特征(包括中心矩和条件分位数指示符)上运行GCM模板,将这些特征分块池化后通过柯西规则(Cauchy rule)组合,同时在回归阶段采用增长节点样条作为冗余项。本文的贡献包括:(i)一项中心性结果,该结果使得尺度特征满足奈曼正交性,而非中心版本则存在偏差;(ii)均值-分位数构造在PC框架内产生的定向不对称性及其修正方法;(iii)在Phi-忠实性理论下,结合GFCM的PC方法能够恢复该检验类别的CPDAG(部分有向无环图);(iv)在合成数据、半合成尾部注入数据以及随机有向无环图(DAG)上进行的、针对协方差之外依赖关系敏感的条件独立性检验的基准测试。在经过尺寸校正的功效下,GFCM能够恢复协方差家族遗漏的尺度和尾部边,并且在PC处理的深层条件集合问题中保持功效;随着样本量n增长,GFCM保持校准状态,而增强型GCM(FFCI)和偏copula检验则无法做到这一点,且GFCM可直接处理混合型数据。在PC框架的大规模测试中,GFCM在保持校准状态的检验方法中获得了最低的骨架结构汉明距离(SHD),而其他方法则会产生过多的错误边。GFCM的有效性依赖于加性冗余项,其尾部优势在模拟数据和半合成数据上得到了验证,因为目前不存在同时具备重尾和已知结构的完全真实基准数据。

英文摘要:

Constraint-based causal discovery like PC and FCI depends on its conditional independence test. Partial correlation and the Generalised Covariance Measure (GCM) detect only the conditional covariance of residuals, so they miss dependence in the mean's nonlinear part, the scale, and the tails. Tests that detect more are biased inside PC, not scalable, only continuous, or not aimed at the tails. Our Generalised Feature Covariance Measure (GFCM) is valid, sensitive beyond covariance, robust inside PC, and applicable to mixed-type data. It runs the GCM template on a configurable set of residual features with conditional mean zero (centered moments and conditional quantile indicators), pooled in blocks and combined by the Cauchy rule, with a growing-knot spline nuisance at regression cost. We contribute (i) a centering result making the scale feature Neyman orthogonal, where the uncentered version is biased; (ii) the orientation asymmetry the mean-quantile construction creates inside PC, and its fix; (iii) a Phi-faithfulness theory under which PC with GFCM recovers the CPDAG of the set's detection class; and (iv) a benchmark of CI tests sensitive beyond covariance on synthetic data, semi-synthetic tail injections, and PC discovery on random DAGs. Under size-corrected power, GFCM recovers the scale and tail edges the covariance family misses and alone keeps power at the deep conditioning sets PC issues. It stays calibrated as n grows, whereas FFCI, the boosted GCM, and the partial copula test do not, and it handles mixed-type data directly. Inside PC at scale it attains the lowest skeleton SHD among tests that stay calibrated, while the others inflate false edges. Validity rests on an additive nuisance, and the tail advantage is shown on simulated and semi-synthetic data, as no fully real benchmark with both heavy tails and known structure exists.

补充信息

↑