发表机构
Deezer Research; LORIA; Université Sorbonne Paris Nord; CNRS; LIPN; IDIAP(Deezer研究院; 洛里亚实验室; 巴黎北索邦大学; 法国国家科学研究中心; 巴黎第十三大学计算机科学实验室; IDIAP研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对文学文本引语归属的效率与精度难题,提出基于编码器的联合评分方案,在Project Dialogism小说语料库上实现94.5%准确率,速度远超同类方法,且发布了修改后的ModernBookNLP工具。
AI 中文摘要
将文学文本中的引语归属到其说话者仍是一个未解决的挑战。标准方法会独立预测每个引语的说话者提及,虽高效但准确性仍有限。相比之下,大语言模型(LLM)方法能取得出色性能,但其计算成本限制了在大规模文学分析中的应用。我们提出一种基于编码器的高效方案,可在共享的大上下文窗口内解决多个引语归属问题。利用我们的新方案“联合评分(joint scoring)”,我们在Project Dialogism小说语料库(PDNC)上取得了最先进(SOTA)的性能,该语料库包含22部英文小说中超过35000个人工标注的引语。我们的最佳模型整体归属准确率达到94.5%,在A100 GPU上处理小说的速度比同类标准方法快20倍,比基于LLM的方法快1000倍以上。对模型表示的分析表明,联合评分通过保留长距离回指消解信号,改进了具有挑战性的归属示例,我们发现预训练编码器中已存在这类信息。为便于推广应用,我们发布了ModernBookNLP,它是BookNLP的修改分支,用我们的最佳系统替换了其引语归属模型,可在该https链接获取。
英文摘要
Attributing quotations to their speakers in literary texts remains an open challenge. Standard methods, which independently predict a speaker mention for each quotation, are efficient but still limited in accuracy. In contrast, large language model (LLM) approaches achieve strong performance, but their computational cost limits their use in large-scale literary analysis. We propose an encoder-based efficient formulation that resolves multiple quotation attributions within a shared, large context window. Using our new formulation, \textit{joint scoring}, we report state-of-the-art (SOTA) performance on the Project Dialogism Novel Corpus (PDNC), comprising more than 35,000 manually annotated quotations from 22 English novels. Our best model reaches 94.5\% overall attribution accuracy while processing novels $20\times$ faster than comparable standard methods and more than $1000\times$ faster than LLM-based approaches on an A100 GPU. An analysis of models' representations suggests that joint scoring improves on challenging attribution examples by preserving long-range anaphora resolution signal, an information that we found already present in pretrained encoders. To facilitate adoption, we release ModernBookNLP, a modified fork of BookNLP that replaces its quotation attribution model with our best system available at https://github.com/gasmichel/ModernBookNLP_QA/.