发表机构
Skoltech; HSE University; Lomonosov Moscow State University; ITMO University; AIRI(斯科尔科沃科学技术学院; 高等经济大学; 莫斯科国立罗蒙诺索夫大学; ITMO大学; AIRI人工智能研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出Drift Inspector系统,通过LLM提取原子贡献声明并聚类,测量研究领域随时间的变化,应用于EMNLP和ACL选集,揭示从经典NLP向LLM能力的漂移。
AI 中文摘要
科学摘要将贡献与背景、动机和元语言混合在一起,因此按原样读取这些摘要的工具无法区分一个领域所产生的内容与其所讨论的内容。我们提出了Drift Inspector,一个开源系统,用于在原子贡献声明(ACCs)层面测量和探索研究领域随时间的变化:这些声明是从每个摘要中提取的、去语境化的、承载贡献的命题,由大语言模型(LLM)在分析前提取。该系统将这些声明跨年份聚类成一个交互式地图,其中每个趋势都可以追溯到其背后的声明和论文。应用于EMNLP六年的数据,它显示了该领域从经典NLP任务转向LLM时代的能力,如推理和多模态——这一趋势是关键词或整篇摘要计数所模糊的。发布的数据不仅限于EMNLP:相同的流程已处理了完整的ACL选集(346k条声明,80k篇摘要,423个会议)。提取经过人工验证,聚类结果对照外部手动构建的分类法进行了检查。
英文摘要
Scientific abstracts mix contributions with background, motivation, and meta-language, so tools that read them as-is cannot separate what a field produces from what it discusses. We present Drift Inspector, an open-source system for measuring and exploring how a research field changes over time at the level of Atomic Contribution Claims (ACCs): decontextualized, contribution-bearing propositions an LLM extracts from each abstract before analysis. The system clusters these claims across years into an interactive map where every trend traces back to the claims and papers behind it. Applied to six years of EMNLP, it shows the field shifting away from classic NLP tasks toward LLM-era capabilities such as reasoning and multimodality -- a movement that keyword or whole-abstract counts blur. The released data extend beyond EMNLP: the same pipeline has processed the full ACL Anthology (346k claims, 80k abstracts, 423 venues). Extraction is human-validated and clustering checked against an external manually constructed taxonomy.
CommentsAccepted to EMNLP 2026 System Demonstrations. 11 pages. Live demo, code and data: https://hamyrappy.github.io/drift-inspector