arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27343cs.CL

从碎片化的历史文本复用中恢复成对的散文级再发布与复用:一项关于18世纪书籍与报纸的工作流研究

Pair-Level Essay-Scale Republication and Reuse from Fragmented Historical Text Reuse: A Workflow Study on Eighteenth-Century Books and Newspapers

Ke Shu, Kira Hinderks, Eetu Mäkelä, Mikko Tolonen

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对碎片化历史文本复用的成对证据整合难题,以18世纪休谟的散文为研究对象,提出分阶段规则工作流,在ECCO数据集上验证了方法的有效性,为历史文本复用的恢复提供了实用方案。

中文摘要 AI 辅助

本文针对从碎片化文本复用证据中恢复散文级再发布与复用的问题展开研究,该场景的核心挑战在于成对证据的整合,而非单纯的片段检索。研究以18世纪苏格兰哲学家大卫·休谟的散文为核心候选集,涵盖来自ECCO(18世纪藏书在线)的书籍及历史报纸。由于输入由碎片化的复用匹配结果构成,而非完整的文档对,且正样本覆盖范围本就不完整,因此我们将该任务形式化为将成对证据整合为合理的传播关系,并对比三类方法族:分阶段的基于规则的工作流、基线方法(一个决策树和两种直接LLM设置)以及自动化规则适配。在带标签的ECCO-ECCO切片上,仅成对特征聚合就在主要带标签切片上达到了0.948的F1值,而最终工作流在测试的规则阶段中展现出最强的整体精确率-召回率权衡。在完整的ECCO-ECCO候选空间中,直接LLM基线标记出多达14886对为再发布对,而最终工作流仅标记771对,在这种直接提示设置中,其表现为高召回率候选扩展器,而非精确率可控的部署分类器。在ECCO-报纸数据集上,人工审计确认所有176个预测正样本均为真实的再发布或复用案例,而期数重复及来源端的多重性揭示了额外的来源结构。在不完整的真实值条件下,可审计的成对证据整合为生成用于历史检查的紧凑候选空间提供了实用途径。

英文摘要

This paper addresses the recovery of essay-scale republication and reuse from fragmented text-reuse evidence, a setting whose central challenge is pair-level evidence consolidation and not fragment retrieval alone. The study focuses on a candidate set centered on essays by eighteenth-century Scottish philosopher David Hume, spanning books from ECCO (Eighteenth Century Collections Online) and historical newspapers. Because the input consists of fragmented reuse hits instead of clean document pairs, and positive coverage is inherently incomplete, we formulate the task as pair-level evidence consolidation into plausible transmission relations and compare three methodological families: a staged rule-based workflow, baselines (a decision tree and two direct LLM settings), and automated rule adaptation. On labeled ECCO--ECCO slices, pair-level feature aggregation alone already reaches 0.948 F1 on the main labeled slice, while the final workflow gives the strongest overall precision-recall trade-off among the tested rule stages. On the full ECCO--ECCO candidate universe, direct LLM baselines flag up to 14,886 pairs as reprints compared to 771 for the final workflow, behaving in this direct-prompt setup as high-recall candidate expanders rather than precision-controlled deployment classifiers. On ECCO--Newspaper, manual audit confirms all 176 predicted positives as genuine cases of republication or reuse, while issue duplication and source-side multiplicity reveal additional provenance structure. Under incomplete ground truth, auditable pair-level evidence consolidation provides a practical way to produce compact candidate spaces for historical inspection.

发表机构

  • University of Helsinki(赫尔辛基大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑