HAWKEYE:更深一层的观测——面向时序链接预测的感知内聚性的结构通道
HAWKEYE: Seeing One Layer Deeper -- A Cohesion-Aware Structural Channel for Temporal Link Prediction
- The University of Memphis(孟菲斯大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
HAWKEYE作为时序链接预测模型的即插即用结构通道,通过维护2跳内聚桥特征提升性能,在6个数据集及tgbl-subreddit上显著优于基线,可高效处理大规模图数据。
AI中文摘要:
当前最先进的时序链接预测(TLP)模型本质上是多通道信息聚合器,它们将交互历史通道、时间编码通道和结构通道结合起来。前两者已被不断优化,而结构通道仍是粗糙的事后补充——DyGFormer将其编码为1-2位的邻居共现计数。我们首先进行了一项测量:在稀疏时序图上,经典的1跳共同邻居信号接近随机(判别AUC≈0.50),因为两个节点几乎从不共享直接邻居;真正具有判别性的信号位于更深的1跳——即2跳内聚桥,其判别AUC在二分图和非二分图上可达0.73-0.98。受此启发,我们提出HAWKEYE,这是一个感知内聚性的结构通道,它增量维护经典的k族内聚性指标(度→k核→k桁架)并形成2跳内聚桥特征。HAWKEYE可作为时序图模型原生结构通道的即插即用替换,无需改变主干。将HAWKEYE替换到DyGFormer中,在6个经多种子验证的数据集(uci、enron、USLegis、CanParl、reddit、mooc)上,测试AP/MRR较共现通道提升了0.6至10.8个点。在二分图推荐基准tgbl-subreddit上,一项3种子的单遍仅结构消融实验显示,HAWKEYE几乎将基线测试MRR翻倍(0.103±0.003→0.204±0.005,所有3种子提升10.1个点);该流处理管道可在5分钟内处理6700万边的tgbl-flight。我们进一步确定了其适用场景:增益与图的无训练2跳判别AUC相关,在退化或饱和图上消失——这是可预测的边界。所有代码、数据和绘图脚本均已公开。
英文摘要:
State-of-the-art temporal-link-prediction (TLP) models are, in essence, multi-channel information aggregators: they combine an interaction-history channel, a time-encoding channel, and a structure channel. The first two have been refined relentlessly; the structure channel remains a crude afterthought -- DyGFormer encodes it as a 1--2-bit neighbour-cooccurrence count. We begin with a measurement: on sparse temporal graphs the classical 1-hop common-neighbour signal is near-random (discriminative AUC $\approx 0.50$), because two nodes almost never share a direct neighbour; the genuinely discriminative signal lies one hop deeper -- the 2-hop cohesive bridge, whose discAUC reaches 0.73--0.98, on both bipartite and non-bipartite graphs. Motivated by this, we propose HAWKEYE, a cohesion-aware structural channel that incrementally maintains the classical k-family of cohesiveness indicators (degree $\to$ k-core $\to$ k-truss) and forms 2-hop cohesive-bridge features. HAWKEYE is a drop-in replacement for a temporal-graph model's native structure channel, with no change to the backbone. Swapping HAWKEYE into DyGFormer improves test AP/MRR over the cooccurrence channel by +0.6 to +10.8 points across six multi-seed-validated datasets (uci, enron, USLegis, CanParl, reddit, mooc). On the bipartite recommendation benchmark tgbl-subreddit, a 3-seed single-pass struct-only ablation shows HAWKEYE nearly doubling the baseline test MRR (0.103$\pm$0.003 $\to$ 0.204$\pm$0.005, +10.1 points across all three seeds); the streaming pipeline scales to the 67M-edge tgbl-flight in five minutes per pass. We further characterise when it helps: the gain tracks a graph's training-free 2-hop discAUC and vanishes on degenerate or saturated graphs -- a predictable boundary. All code, data, and figure-generation scripts are released.