发表机构
School of Data Science, The Chinese University of Hong Kong, Shenzhen; Huawei Technologies Co., Ltd.(香港中文大学(深圳)数据科学学院; 华为技术有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
PreGress是首个原生排序预训练与提示学习框架,通过多任务预训练与轻量提示模块,在低开销下实现了多节点排序任务的优异性能。
AI 中文摘要
节点排序是图信息检索中的基础问题,用于衡量节点的相对重要性,支撑影响力分析、推荐、基于图的检索增强生成等大量应用。然而,基于图的排序指标的精确计算在大规模场景下往往存在计算成本过高的问题。现有的基于图神经网络(GNN)的排序方法可实现可扩展的近似计算,但通常针对特定排序标准定制,且每个下游任务都需要重新训练,这限制了其迁移性与效率。近期的图预训练方法旨在实现跨任务的知识迁移,但其学习目标与节点排序严重不匹配,导致对面向排序的应用的适应性欠佳。为解决这些局限,我们提出PreGress,首个支持广泛节点排序任务的原生排序预训练与提示学习框架。PreGress采用精心设计的目标进行多任务预训练,包括度中心性预测与属性重构,以联合捕获结构与属性信息。为支持异构排序标准,我们设计了轻量级、任务特定的提示模块,可在不重新训练整个模型的情况下,将冻结的排序主干适配至下游任务。在6个公开图数据集、2个真实世界的查询-物品基准(Yelp2018与MovieLens-100K)以及受控的五标准图访问研究上的实验表明,该框架在低任务特定状态开销下具备优异的排序质量。
英文摘要
Node ranking is a fundamental problem in graph information retrieval, measuring the relative importance of nodes and supporting a wide range of applications such as influence analysis, recommendation, and graph-based retrieval augmented generation. However, exact computation of graph-based ranking measures is often computationally prohibitive at scale. Existing GNN-based ranking methods provide scalable approximations, but they are typically tailored to individual ranking criteria and require retraining for each downstream task, which limits their transferability and efficiency. Recent graph pre-training approaches aim to enable knowledge transfer across tasks, yet their learning objectives are largely misaligned with node ranking, resulting in suboptimal adaptability to ranking-oriented applications. To address these limitations, we propose PreGress, the first ranking-native pre-training and prompting framework for supporting a wide range of node ranking tasks. PreGress performs multi-task pre-training using our carefully designed objectives, including degree centrality prediction and attribute reconstruction, to jointly capture structural and attribute information. To support heterogeneous ranking criteria, we design lightweight, task-specific prompt modules that adapt a frozen ranking backbone to downstream tasks without full retraining. Experiments on six public graphs and two real-world query-to-item benchmarks---Yelp2018 and MovieLens-100K---together with a controlled five-criterion graph-access study demonstrate strong ranking quality with low task-specific state overhead.