发表机构
Norwegian Computing Center(挪威计算中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究在大型网络爬取语料库中检测文本包含问题,核心方法是基于文档指纹识别技术构建FindMyText工具,通过扩展机制捕获匹配指纹序列,主要贡献是能可靠检测近乎逐字副本,在三个数据集上优于其他方法。
AI 中文摘要
我们展示了FindMyText,一个开源的Python包,旨在有效评估给定文本是否部分或全部出现在文本语料库中。该工具基于文档指纹识别的先前技术构建,但通过一种新颖机制进行扩展,以明确捕获匹配指纹序列。通过识别此类链,该工具能更可靠地检测给定文本的近乎逐字副本,而非仅仅是文本相似性。这使得FindMyText特别适合验证语料库中版权材料的存在。利用分布式、基于磁盘的索引框架,该系统可扩展到大型网络爬取数据集。使用新的基准评估文本包含方法,我们表明FindMyText在三个数据集(ArXiv论文、维基百科和通用网页内容)上优于其他方法。
英文摘要
We present FindMyText, an open-source Python package designed to efficiently assess whether a given text appears, in part or in full, within a text corpus. The tool builds on prior techniques for document fingerprinting, but extends them with a novel mechanism to explicitly capture sequences of matching fingerprints. By identifying such chains, the tool can more reliably detect near-verbatim copies of a given text rather than mere textual similarities. This makes FindMyText particularly suited for verifying the presence of copyrighted material in a corpus. Leveraging a distributed, disk-based indexing framework, the system scales to large web-crawled datasets. Using a new benchmark for evaluating text containment methods, we show that FindMyText outperforms alternative approaches across three datasets (ArXiv papers, Wikipedia, and generic web content).
Comments6 pages + references and appendices