AI 中文总结
本研究将细粒度互文性提取设为LLM智能体任务,构建专家裁决的古典史书互文基准,验证12种LLM性能,扩展至二十四史揭示引文的总体稳定与个案漂移规律并发布相关资源。
AI 中文摘要
互文性的计算方法已从字符串匹配发展到神经检索,但其输出(相似度得分和平行段落列表)仅能识别文本间的复用位置,无法刻画复用的方式或原因。我们将细粒度互文性提取重新定义为智能体任务:大型语言模型(LLM)完整读取两个文本单元后,需通过受限工具接口,将每项复用提议锚定到双方文本的精确字符跨度,并依据包含形式、方面、来源标记、功能、立场的五维复用类型学进行标注。我们通过《论语》与《汉书》的穷尽式比较验证该方法,三名领域专家将多模型候选集裁决整合为含2533对互文的基准。基于此标准,我们研究了12种LLM,报告其精度为56%-93%、可比质量下51倍的成本差异,以及置信度的校准效果。专家一致性呈现可靠性梯度:文本表面清晰的维度标注一致,需意图推断的维度则存在争议,明确了此类标注的适用范围。将经验证的提取器扩展至全部二十四史(65380次比较、5766对互文),可恢复相似度得分无法表达的语料库层面结构。引文的解释性构成在18个世纪中无系统性变化,但同一段落的引用越来越少采用字面形式。总体稳定与个案漂移的情况符合文化吸引力理论的预期。我们发布了提取协议与专家裁决基准。
英文摘要
Computational approaches to intertextuality have advanced from string matching to neural retrieval, yet their outputs, similarity scores and parallel-passage lists, identify where texts reuse one another without characterizing how or why. We recast fine-grained intertextuality extraction as an agentic task in which a large language model (LLM) reads two text units in full and, through a constrained tool interface, must ground each proposed reuse in exact character spans on both sides and label it under a five-dimension typology of reuse (form, aspect, source-marking, function, stance). We validate the approach on an exhaustive comparison of the Analects with the Book of Han, where three domain experts adjudicate a pooled multi-model candidate set into a benchmark of 2,533 intertextual pairs. Against this standard we study twelve LLMs, reporting precision (56%-93%), a 51$\times$ cost spread at comparable quality, and how well their confidence is calibrated. Expert agreement traces a reliability gradient: dimensions legible on the textual surface are annotated consistently, while those requiring inference of intent are contested, delimiting the claims such annotation supports. Scaling the validated extractor to the full Twenty-Four Histories (65,380 comparisons, 5,766 pairs) recovers corpus-level structure a similarity score cannot express. The interpretive composition of citation shows no systematic change across eighteen centuries, yet the same passage is quoted less and less literally. Stability in the aggregate with drift in the individual case is what a cultural-attraction account expects. We release the extraction protocol and the expert-adjudicated benchmark.
Comments9 pages, 4 figures, 3 tables