发表机构
Universidad de Castilla-La Mancha; Università degli Studi di Milano Bicocca(卡斯蒂利亚-拉曼恰大学; 米兰比可卡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对联邦环境下LiNGAM因果发现的局限,提出FedRCD系列算法,克服FedISHC在近对称噪声下的失效问题,实验揭示其排序机制及边缘标准化的影响。
AI 中文摘要
本文研究线性非高斯无环模型(LiNGAM)在联邦环境中的应用,这类因果模型可突破马尔可夫等价的限制。然而,在诸多领域中数据稀缺,且因GDPR等法规约束,集中不同客户端的数据以增加样本量并不可取,联邦环境为平衡隐私保护与因果发现精度提供了可行方案。遗憾的是,LiNGAM设定下的标准集中式估计器DirectLiNGAM无法直接联邦化。高阶累积量张量为解决该问题提供了途径:它们仅依赖所涉及变量的联合分布,且可在独立样本组间精确相加,因此在水平、垂直及混合分区中仅需一轮通信即可完成。但当前沿此思路提出的联邦方法FedISHC在近对称噪声下会失效。为克服上述局限,本文提出FedRCD系列因果发现算法,并研究了三种在通信轮次与代数噪声间进行权衡的变体;其中两种是集中式高阶累积量(HC)和HC-LiNGAM算法的精确联邦对应版本,且单轮变体还能在任意粒度(从单个观测值到整个客户端)上有效支持精确弃权(不执行)。数值实验表明,在实际部署的典型样本量下,整个基于累积量的联邦系列算法实际上并非按分数编码的总体不对称性对变量排序,而是按有向无环图(DAG)沿其有向路径诱导的方差阶梯排序,这是变量可排序性的累积量对应形式。边缘标准化会使所有累积量方法的排序接近随机,而该协议下无法联邦化的尺度不变DirectLiNGAM则不受影响。
英文摘要
In this paper we study linear non-Gaussian acyclic models (LiNGAM) when used in federated environments. These causal models allow one to go beyond Markov equivalence. However, in many domains data are scarce, and increasing the sample size by centralising data from different clients is not advisable due to regulations such as the GDPR. The federated environment offers an attractive option to balance privacy and causal discovery accuracy. Unfortunately, the standard centralised estimator in the LiNGAM setting, i.e., DirectLiNGAM, cannot be straightforwardly federated. Higher-order cumulant tensors offer a way around this obstacle: they depend only on the joint distribution of the variables involved and add exactly across independent sample groups, so a single communication round suffices in horizontal, vertical, and hybrid partitions. However, FedISHC, i.e., the current federated method along these lines, breaks down under near-symmetric noise. To overcome the above limitation, we introduce the FedRCD family of causal discovery algorithms, and investigate three variants that trade off communication rounds against algebraic noise; two of them are exact federated counterparts of the centralised high-order cumulant (HC) and HC-LiNGAM algorithms, and the single-round variants further effectively support exact unlearning at any granularity, from a single observation to a whole client. Numerical experiments show that at sample sizes typical of real deployments, the entire cumulant-based federated family does not actually rank variables by the population asymmetry that the scores encode at zero. It ranks them by a variance ladder induced by the DAG along its directed paths, the cumulant counterpart of varsortability. Marginal standardisation collapses every cumulant method to near-random ordering, while scale-invariant DirectLiNGAM, not federable under this protocol, is unaffected.
CommentsAccepted at the 12th International Conference on Probabilistic Graphical Models (PGM 2026)