arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.09561cs.CL

面向引文图谱的自动演化树生成

Towards Automatic Evolution Tree Generation from Citation Graphs

  • Georgia Institute of Technology(佐治亚理工学院)
  • Emory University(埃默里大学)

机构由 AI 辅助整理,请以论文原文为准。

Zexing Zhao, Yuntong Hu, Liang Zhao

AI总结:

针对现有分类学归纳方法局限于叶节点且忽略时间的问题,提出EvoTree分阶段框架,通过图感知编码器、层次聚类和时间微调生成演化树,在首个覆盖11个AI子领域的基准上取得最优性能。

AI中文摘要:

综述仍是研究者掌握人工智能子领域内方法谱系的主要途径,但其扩展性难以适应当前的论文发表速度。现有的分类学归纳方法大多局限于叶节点且不考虑时间因素;它们倾向于将过渡性论文强行归入成熟的叶节点,并可能在祖先与后代之间造成拓扑倒置。我们提出EvoTree,一种分阶段框架,将概念骨干学习与时间细化解耦:基于图感知编码器与基于分布的层次聚类生成稳定的分类学骨干;随后进行时间微调,在单调路径约束下将边缘论文重新挂接到内部节点;最后通过大语言模型(LLM)处理为概念打标签,而不改变拓扑结构。我们发布了首个针对该任务、覆盖11个人工智能子领域的带标注基准。EvoTree在所有基线中取得了最高的归一化互信息(NMI)和引文方向准确率,并在标注基准上实现了最佳的概念纯度,且是唯一在标注集上具有非平凡边缘论文检测能力的方法。

英文摘要:

Surveys remain the primary way researchers grasp the lineage of methods within an AI subfield, but they scale poorly against the current rate of publication. Existing taxonomy-induction methods are largely leaf-bound and time-agnostic; they tend to force transitional papers into mature leaves and can create topological inversions between ancestors and descendants. We propose EvoTree, a staged framework that decouples conceptual backbone learning from temporal refinement: a graph-aware encoder with distribution-based hierarchical clustering yields a stable taxonomy backbone; temporal fine-tuning then re-attaches marginal papers to internal nodes under monotonic-path constraints; a final LLM pass labels concepts without altering the topology. We release the first annotated benchmark for this task across 11 AI subfields. EvoTree attains the highest NMI and citation-direction accuracy among all baselines and the best concept purity on the annotated benchmark, and is the only method with non-trivial marginal-paper detection on the annotated set.

↑