arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.19914cs.LG

多源Wasserstein分布鲁棒图学习

Multi-Source Wasserstein Distributionally Robust Graph Learning

Chuansen Peng, Yifan Xia, Jinshan Zhong, Xiaojing Shen

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对图信号处理中异质源域数据融合难题,提出MS-WDRO多源Wasserstein分布鲁棒图学习框架,经实验验证其在图恢复等任务上优于多个基线方法。

中文摘要 AI 辅助

从图信号中推断网络拓扑是图信号处理的核心,应用于神经科学、传感器网络和社交网络等领域。在实际场景中,目标域样本稀缺而异质的源域数据丰富,融合这些源数据极具挑战性:欧氏平均适用于同质源,但随着源间差异增大,性能会急剧下降,将不同几何结构坍缩为膨胀且有偏的共识。我们利用Wasserstein度量的分布保持特性来应对异质性,同时保留每个源的固有几何结构。我们提出MS-WDRO,一种多源Wasserstein分布鲁棒图学习框架,该框架通过加权Wasserstein重心(一种几何上合理的名义分布)融合异质源,随后在其周围构建歧义球以对冲残留不确定性。最小化最坏情况风险可得到易处理的正则化拉普拉斯估计器,该估计器可通过可证收敛的ADMM方案高效求解。我们建立了非渐近保证:经验重心的有限样本收敛界、证明朴素聚合次优的池化偏差下界,以及仅随源数量呈对数依赖、以参数速率衰减的样本外超额风险界。为校准控制鲁棒性、稀疏性和源融合的超参数,我们将求解器展开为可微架构并进行端到端训练,实现了超越交叉验证的数据自适应校准,同时保留可解释性。在合成基准和多站点ABIDE I神经影像数据集上的实验表明,MS-WDRO在图恢复、样本效率和下游诊断效用方面始终优于七个基线方法,且在样本稀缺的场景中增益最大。

英文摘要

Reconstructing complex network topologies from data is a fundamental challenge in cybernetics and graph signal processing, with applications in neuroscience, sensor, and social networks. In practice, target-domain samples are scarce while heterogeneous source-domain data are abundant. Fusing these sources is challenging: Euclidean averaging works for homogeneous sources but degrades sharply as inter-source divergence grows, collapsing distinct geometries into an inflated, biased consensus. We exploit the Wasserstein metric's distribution-preserving properties to counter heterogeneity while preserving each source's intrinsic geometry. We propose MS-WDRO, a multi-source Wasserstein distributionally robust graph learning framework that fuses heterogeneous sources via their weighted Wasserstein barycenter, a geometrically principled nominal distribution, then builds an ambiguity ball around it to hedge residual uncertainty. Minimizing worst-case risk yields a tractable regularized Laplacian estimator solved efficiently via a provably convergent ADMM scheme. We establish non-asymptotic guarantees: a finite-sample concentration bound for the empirical barycenter, a pooling bias lower bound proving naive aggregation is suboptimal, and an out-of-sample excess risk bound decaying at a parametric rate with only logarithmic dependence on source count. To calibrate hyperparameters governing robustness, sparsity, and source fusion, we unroll the solver into a differentiable architecture trained end-to-end, achieving data-adaptive calibration beyond cross-validation while retaining interpretability. Experiments on synthetic benchmarks and the multi-site ABIDE~I neuroimaging dataset show MS-WDRO consistently outperforms seven baselines in graph recovery, sample efficiency, and downstream diagnostic utility, with the largest gains in the sample-scarce regime.

发表机构

  • School of Mathematics, Sichuan University(四川大学数学学院)
  • School of Statistics and Data Science, Southwestern University of Finance and Economics(西南财经大学统计与数据科学学院)

机构由 AI 辅助整理,请以论文原文为准。

↑