arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.16676math.STcs.LGmath.PRstat.MLstat.TH

稀疏图上消息传递中深度的价值:凯斯滕 - 斯蒂古姆二分法

The Value of Depth in Message Passing on Sparse Graphs: A Kesten-Stigum Dichotomy

Aseem Raj Baranwal

中文总结 AI 辅助

研究图神经网络在稀疏图上的深度,基于稀疏上下文随机块模型,证明深度由凯斯滕 - 斯蒂古姆比率\(\kappa\)决定,低于阈值误差几何收敛,高于阈值深度有几何生产性,还得出局部分类器下限及模拟相关结论。

中文摘要 AI 辅助

我们研究了图神经网络在稀疏图上所需的深度,采用其最纯粹的统计形式:在平均度\(\Delta = O(1)\)的稀疏上下文随机块模型(CSBM)上进行节点分类,其局部弱极限是带广播标签的泊松高尔顿 - 沃森树。先前工作得出了一个消息传递分类器\(h_\ell\),它从距离\(k\leq\ell\)的每个顶点聚合衰减证据\(2\operatorname{artanh}(\gamma^k t(X_v))\)。我们证明深度的值由单个数字凯斯滕 - 斯蒂古姆比率\(\kappa=\gamma^2\Delta\)决定。低于阈值\((\kappa<1)\)时,误差序列呈几何速率柯西收敛;高于阈值\((\kappa>1)\)时,深度具有几何生产性。没有任何深度的局部分类器能超越孤立根设定的通用下限\(e^{-\Delta}\Phi(-\zeta)\),而第一层有明确的总变差量帮助。模拟表明成对规则的误差曲线在\(\ell\)中轻微非单调,存在最优有限深度,BP饱和更快。

英文摘要

How deep does a graph neural network need to be on a sparse graph? We study its purest statistical form: node classification on the sparse contextual stochastic block model (CSBM) with average degree $Δ=O(1)$, whose local weak limit is a broadcast-labelled Poisson Galton-Watson tree. Prior work derived a message-passing classifier $h_\ell$ that aggregates from each vertex at distance $k\le\ell$ the attenuated evidence $2\operatorname{artanh}(γ^k t(X_v))$, with $γ$ the edge signal and $t$ a bounded likelihood-ratio transform of the feature. We prove that the value of depth is governed by a single number, the Kesten-Stigum ratio $κ=γ^2Δ$. Below the threshold ($κ<1$), the error sequence is Cauchy at a geometric rate, $|\mathcal{E}(\ell)-\mathcal{E}(\ell')|\le Cκ^{(\ell+1)/3}$ for all $\ell'>\ell$, so all layers beyond depth $O(\log(1/ε))$ change the error by less than $ε$; conversely, under mild regularity each sufficiently deep layer still flips the decision with probability at least $cκ^{\ell/2}$, the empirically sharp exponent. Above the threshold ($κ>1$), depth is geometrically productive: $\mathcal{E}(\ell)$ is driven to a branching-process floor of order at most $1/(κ-1)$ at any geometric rate $κ^{-s\ell}$, $s<1$ (this bound has content only for $κ>17$). No local classifier of any depth beats the universal floor $e^{-Δ}Φ(-ζ)$ set by isolated roots ($ζ$ the feature signal-to-noise ratio), while the first layer provably helps by an explicit total-variation amount. Simulations with an exact belief-propagation baseline on the same trees show that the pairwise rule's error curve is mildly non-monotone in $\ell$, so an optimal finite depth exists (an exact instance is certified in the appendix), while BP saturates strictly faster, at an effective per-layer ratio below $κ$ that we identify.

补充信息

↑