检索凯伦·布里克森《七个哥特故事》中的圣经互文引用
Retrieving Biblical Intertextual References in Karen Blixen's Seven Gothic Tales
浏览论文内容
中文总结 AI 辅助
针对布里克森作品中的圣经互文检索,构建189条引用基准,比较多种检索方法,微调DFM-large显著提升性能,并提出模型作为启发式共同读者的定位。
中文摘要 AI 辅助
识别互文引用是文学研究的核心问题,但当源材料通过改写、典故、历史语言和翻译进行转化时,计算上具有很大难度。我们通过凯伦·布里克森《七个哥特故事》中的圣经互文性来研究这一问题。借助评注版评论,我们构建了一个包含189条注释引用的基准数据集,并针对历史上合理的丹麦语旧约和新约译本的全部31,170节经文进行检索评估。我们比较了TF-IDF和BM25与多语言及丹麦语句子编码器的性能,考察了语言规范化的效果,并使用硬负样本和五折交叉验证对丹麦语编码器进行微调。我们分析了自动推导的词汇重叠层(分别代表引文、改写和典故)上的性能表现。经语言规范化的BM25提供了强大的零样本基线,总体R@10达到0.365,并在前十位排名最高的经文中检索到所有引文。最佳的零样本稠密模型总体得分0.360,与基线相当,但在典故检索上表现更好。对DFM-large进行微调后,其总体R@10从0.265提升至0.508,在典故上的性能从0.138提升至0.339,提高了一倍以上。然而,仅针对编辑注释进行评估低估了模型在学术上的实用性:一位文学学者判定30个选定的排名第一预测中被视为假阳性的结果中有7个是有意义的额外引用。这些发现既展示了计算互文检索的潜力,也揭示了其认知上的局限性。我们不应将学术注释视为穷尽的,也不应将模型输出视为发现,而是提出将检索模型作为启发式共同读者,用于恢复已记录的引用并生成供专家主导的细读的候选。
英文摘要
Identifying intertextual references is central to literary scholarship, but computationally difficult when source material is transformed through paraphrase, allusion, historical language, and translation. We investigate this problem through biblical intertextuality in Karen Blixen's Seven Gothic Tales. Drawing on the commentary to a critical edition, we construct a benchmark of 189 annotated references and evaluate retrieval against all 31,170 verses of historically plausible Danish Old and New Testament translations. We compare TF-IDF and BM25 with multilingual and Danish sentence encoders, examine the effect of linguistic normalization, and fine-tune a Danish encoder using hard negatives and five-fold cross-validation. We analyze performance across automatically derived lexical-overlap strata representing quotations, paraphrases, and allusions. Linguistically normalized BM25 provides a strong zero-shot baseline, attaining an overall R@10 of 0.365 and retrieving every quotation within its ten highest-ranked verses. The best zero-shot dense model achieves a comparable overall score of 0.360 while performing better on allusions. Fine-tuning DFM-large raises its overall R@10 from 0.265 to 0.508 and more than doubles its performance on allusions, from 0.138 to 0.339. However, evaluation against editorial annotations alone understates the model's scholarly usefulness: a literary scholar judged seven of 30 selected rank-one predictions counted as false positives to be meaningful additional references. These findings show both the potential and the epistemic limits of computational intertextual retrieval. Rather than treating scholarly annotations as exhaustive or model outputs as discoveries, we propose retrieval models as heuristic co-readers that recover documented references and generate candidates for expert-led close reading.