arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

WikiSTAR:一个揭示科学维基百科文章隐藏历史的系统

WikiSTAR: A System for Shedding Light on the Hidden History of Scientific Wikipedia Articles

Omer Ehrlich, Nitzan Barzilay, Rona Aviram, Tom Hope

arXiv 2607.12441首次发表:更新:

发表机构

The Hebrew University of Jerusalem; Ben-Gurion University of the Negev; Allen Institute for AI (Ai2)(耶路撒冷希伯来大学; 内盖夫本-古里安大学; 艾伦人工智能研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

WikiSTAR系统利用大语言模型分类器标记编辑类型,通过交互式视图追溯科学维基百科文章修订历史,揭示知识发展,经用户研究验证其能发现新模式、问题并实现新分析,还发布了系统、代码和基准。

AI 中文摘要

维基百科在塑造公众对科学的理解方面发挥着关键作用,其可公开获取的修订历史是科学知识随时间演变的独特记录。然而,有科学意义的修订被大量日常编辑所掩盖,每篇文章的科学历史都被隐藏。我们展示了WikiSTAR(文章修订的科学追踪),这是一个用于探索文章修订历史中有科学意义变化的交互式系统。使用带有专家设计的多标签分类法的大语言模型分类器,WikiSTAR首先标记编辑类型,如技术术语的添加、新研究发现以及科学叙述的变化。然后,通过交互式视图,可以以任何粒度追溯文章的完整修订历史——从揭示科学内容何时以及在哪些部分被添加或完善的总体趋势,到单个编辑——展示科学知识以前所未有的规模发展。在一项用户研究中,来自三个领域的专家发现WikiSTAR揭示了新模式和研究问题,并实现了以前不切实际的分析。我们发布了我们的系统、代码和一个人工标注的基准。

英文摘要

Wikipedia plays a key role in shaping public understanding of science, and its openly accessible revision history is a unique record of how scientific knowledge evolves over time. Yet scientifically meaningful revisions are obscured by the sheer volume of routine edits, leaving each article's scientific history hidden. We present WikiSTAR (Scientific Tracking of Article Revisions), an interactive system for exploring scientifically meaningful changes across an article's revision history. Using an LLM classifier with an expert-designed multi-label taxonomy, WikiSTAR first tags edit types such as the addition of technical terms, new research findings, and changes in scientific narrative. Then, through interactive views, an article's full revision history can be traced at any granularity - from aggregate trends that reveal when and in which sections scientific content was added or refined, down to individual edits - showing how scientific knowledge develops at a scale previously impossible. In a user study, experts from three domains found that WikiSTAR surfaced new patterns and research questions and enabled previously impractical analyses. We release our system, code and a human-annotated benchmark.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑