arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在大型网络爬取语料库中对文本包含进行稳健、可扩展的检测

FindMyText: Robust, Scalable Detection of Text Containment in Large Web-Crawled Corpora

Lars Henry Berge Olsen, Pierre Lison, Martin Jullum, Mark Anderson

arXiv 2607.10020首次发表:更新:

发表机构

Norwegian Computing Center(挪威计算中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究在大型网络爬取语料库中检测文本包含问题,核心方法是基于文档指纹识别技术构建FindMyText工具,通过扩展机制捕获匹配指纹序列,主要贡献是能可靠检测近乎逐字副本,在三个数据集上优于其他方法。

AI 中文摘要

我们展示了FindMyText,一个开源的Python包,旨在有效评估给定文本是否部分或全部出现在文本语料库中。该工具基于文档指纹识别的先前技术构建,但通过一种新颖机制进行扩展,以明确捕获匹配指纹序列。通过识别此类链,该工具能更可靠地检测给定文本的近乎逐字副本,而非仅仅是文本相似性。这使得FindMyText特别适合验证语料库中版权材料的存在。利用分布式、基于磁盘的索引框架,该系统可扩展到大型网络爬取数据集。使用新的基准评估文本包含方法,我们表明FindMyText在三个数据集(ArXiv论文、维基百科和通用网页内容)上优于其他方法。

英文摘要

We present FindMyText, an open-source Python package designed to efficiently assess whether a given text appears, in part or in full, within a text corpus. The tool builds on prior techniques for document fingerprinting, but extends them with a novel mechanism to explicitly capture sequences of matching fingerprints. By identifying such chains, the tool can more reliably detect near-verbatim copies of a given text rather than mere textual similarities. This makes FindMyText particularly suited for verifying the presence of copyrighted material in a corpus. Leveraging a distributed, disk-based indexing framework, the system scales to large web-crawled datasets. Using a new benchmark for evaluating text containment methods, we show that FindMyText outperforms alternative approaches across three datasets (ArXiv papers, Wikipedia, and generic web content).

Comments6 pages + references and appendices

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑