发表机构
Konkuk University(建国大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对连续时间动态图(CTDG)的负采样评估存在偏差的问题,提出全实体排序方法,经多数据集实验验证其可消除采样影响,建议作为CTDG架构比较的主要评估方式。
AI 中文摘要
连续时间动态图(CTDG)中的下一个目的地预测通常会将观察到的交互与采样的负目的地进行排序。所得分数取决于负分布和研究人员选择的候选数量。我们表明,非均匀负分布会改变贝叶斯最优排序,而即使是均匀抽取的有限候选集也会破坏模型排序和测量的模块效应。随时间变化的源-目的地历史成员关系以及使用此信息的模型操作会直接将采样器的影响传递到评估分数。我们通过对重复和新的正样本与已见和未见负样本进行析因评估、仅基于对历史成员关系的最小评分器以及受控表示干预来研究此机制。在 LastFM、MOOC、Reddit 和 Wikipedia 四个数据集上的六个模型中,至少有一对模型在三个数据集上的预期 Uniform-20 指标与全目录之间的相对顺序发生了变化。同一模块的测量效应也会随着候选集大小和训练目标而改变幅度和方向。这些结果表明,来自负采样基准的模型优越性和 ablation 结论取决于所声明的候选配置。全实体排序会评估固定目录中的每个目的地,消除负选择自由度和采样变化,同时保留原始 CTDG 评分器。因此,我们建议将全实体排序作为具有可枚举固定目的地目录的 CTDG 基准上架构比较的主要证据。
英文摘要
Next-destination prediction in continuous-time dynamic graphs (CTDGs) commonly ranks an observed interaction against sampled negative destinations. The resulting score is conditional on both the negative distribution and the number of candidates chosen by the researcher. We show that a non-uniform negative distribution changes the Bayes-optimal ranking, while even a finite candidate set drawn uniformly can destabilize model rankings and measured module effects. Time-varying source-destination history membership and model operations that use this information directly transmit the sampler's influence to the evaluation score. We examine this mechanism using a factorial evaluation of repeated and new positives against seen and unseen negatives, a minimal scorer based solely on pair-history membership, and controlled representation interventions. Across six models on LastFM, MOOC, Reddit, and Wikipedia, at least one model pair changes relative order between the expected Uniform-20 metric and the full catalog on three of the four datasets. The measured effect of the same module also changes in magnitude and direction with the candidate-set size and training objective. These results establish that model-superiority and ablation conclusions from sampled-negative benchmarks are conditional on the stated candidate configuration. All-entity ranking evaluates every destination in a fixed catalog, eliminating negative-selection freedom and sampling variation while retaining the original CTDG scorer. We therefore recommend all-entity ranking as the primary evidence for architecture comparisons on CTDG benchmarks with an enumerable, fixed destination catalog.