arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.09149cs.DB

在不断演变的学术数据中进行分类法维护:可靠性、效率和成本效益

Taxonomy Maintenance In The Wild Over Evolving Scholarly Data: Reliability, Efficiency, and Cost-Effectiveness

Daomin Ji, Hui Luo, Zhifeng Bao, Junhao Gan, Zi Huang

首次发表
浏览论文内容

中文总结 AI 辅助

研究科学出版物增长致学术分类法过时问题,提出GIST框架。该框架基于专家证据归纳结构,集成部分分类法,学习映射,用核心集选择增量更新,结合生成器与检索模块。实验表明GIST性能优且成本低。

中文摘要 AI 辅助

科学出版物的快速增长使学术分类法迅速过时。我们研究在实际应用中的分类法维护,这是一个超越静态构建的新问题,通过不断使分类法适应不断演变的学术知识库(如arXiv)中给定研究主题。我们提出了GIST,一个用于维护不断演变的分类法的强大框架。与纯粹以大语言模型为中心的方法不同,GIST通过从论文的“相关工作”部分提取部分层次结构,将结构归纳建立在专家策划的证据基础上。它在几何盒嵌入空间中将这些部分分类法集成到一个统一的全局分类法中,其中盒包含编码了“是-a”关系的归纳偏差。为了将语义与几何结构联系起来,GIST学习词嵌入和盒嵌入之间的双向映射。为了进行高效的增量更新,GIST使用新颖性感知的核心集选择,用有代表性历史信号和新证据更新模型,避免昂贵的完全重新训练。为了在特定用户令牌预算下处理高速论文流,GIST进一步将假设概念生成器与具有成本效益的证据检索模块相结合。在真实世界的arXiv数据集上的实验表明,GIST优于现有基线,在最强基线的基础上,节点F1和边F1分别提高了11.0%和13.1%,同时只需要其运行时间的9.6%和货币成本的12.7%。

英文摘要

The rapid growth of scientific publications makes scholarly taxonomies quickly obsolete. We study taxonomy maintenance in the wild, a new problem that moves beyond static construction by continuously adapting taxonomies to evolving scholarly repositories, such as arXiv, for a given research topic. We propose GIST, a robust framework for maintaining evolving taxonomies. Unlike purely LLM-centric approaches, GIST grounds structure induction in expert-curated evidence by extracting partial hierarchies from the "Related Work" sections of papers. It integrates these partial taxonomies into a unified global taxonomy in a geometric box-embedding space, where box containment encodes the inductive bias of is-a relations. To connect semantics with geometric structure, GIST learns a bidirectional mapping between word embeddings and box embeddings. For efficient incremental updates, GIST uses novelty-aware coreset selection to update the model with representative historical signals and new evidence, avoiding costly full retraining. To handle high-velocity paper streams under user-specific token budgets, GIST further combines a hypothesized concept generator with a cost-effective evidence retrieval module. Experiments on real-world arXiv datasets show that GIST outperforms state-of-the-art baselines, improving Node F1 and Edge F1 by 11.0% and 13.1% over the strongest baseline while requiring only 9.6% of its runtime and 12.7% of its monetary cost.

发表机构

  • RMIT University(皇家墨尔本理工大学)
  • The University of Queensland(昆士兰大学)
  • University of Wollongong(伍伦贡大学)
  • The University of Melbourne(墨尔本大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑