发表机构
Old Dominion University; Internet Archive; Virginia Tech(老道明大学; 互联网档案馆; 弗吉尼亚理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究系统比较了六种输入格式及63种组合在学术文档URL提取中的性能,构建了2,338个标注URL的基准数据集,发现TEXTWAL+LaTeX组合最优,并验证了2015年后URL密度激增,强调格式选择的重要性。
AI 中文摘要
学术文档中的URL链接到丰富的外部资源,如数据集、软件、出版物和网站。提取这些URL在许多下游任务(如链接腐烂分析、网络爬虫和构建知识图谱)的数据准备阶段至关重要。然而,现有研究往往低估这一阶段,通常仅从单一格式(通常是从PDF直接转换的文本)中提取URL。我们提出了一项系统性研究,评估跨六种输入格式(带注释层的文本、LaTeX、HTML、XML、Markdown和从PDF转换的PNG)的URL提取性能。为支持评估,我们构建了一个基准数据集,包含来自200篇arXiv论文的2,338个手动标注的URL,这些论文跨越33年,涵盖广泛领域。除了评估单一文件格式外,我们还比较了63种复合输入格式组合。我们的广泛评估表明,TEXTWAL在单一格式输入中表现最佳,而TEXTWAL+LaTeX在整体URL提取性能上达到最优。对于链接到开放获取数据集和软件的URL,也观察到同样的趋势。为进一步验证这些发现,我们将特定格式的URL提取流程应用于跨越33年的364,744篇arXiv论文的纵向随机样本。我们观察到2015年后URL密度急剧增加,以及不同文件格式在时间上的URL提取存在显著差异。总体而言,我们的研究强调了在学术文档中选择适当格式进行URL提取的重要性。数据集和代码公开可用,网址为:this https URL。
英文摘要
URLs in scholarly documents link to rich external resources such as datasets, software, publications, and websites. Extracting these URLs is crucial in the data preparation stage of many downstream tasks, such as link rot analysis, web crawling, and building knowledge graphs. However, existing studies often downplay this phase, simply extracting URLs from a single format, usually text directly converted from PDFs. We present a systematic study evaluating URL extraction across six input formats (text with annotation layer, LaTeX, HTML, XML, Markdown, and PNG converted from PDF). To support the evaluation, we compiled a benchmark dataset consisting of 2,338 manually annotated URLs from 200 arXiv papers spanning a wide range of domains over a 33-year period. In addition to evaluating individual file formats, we also compared 63 composite input-format combinations. Our extensive evaluations indicate that TEXTWAL achieves the best performance among single-format inputs, while TEXTWAL+LaTeX achieves the best overall URL extraction performance. The same trend is observed for URLs linking to open-access datasets and software. To further validate these findings, we apply our format-specific URL extraction pipelines to a longitudinal random sample of 364,744 arXiv papers spanning 33 years. We observe a sharp increase in URL density after 2015, along with remarkable differences in URL extraction across file formats over time. Overall, our study highlights the importance of selecting an appropriate format for URL extraction from scholarly documents. The dataset and code are publicly available at: https://github.com/lamps-lab/arxiv-url-bench .
CommentsPeer-reviewed and accepted to JCDL 2026; 25 pages, 8 figures, and 10 tables