重叠、独特与冲突:大语言模型能否提取它们能识别的信息?
Overlap, Unique and Conflict: Can LLMs Extract What They Can Recognize?
浏览论文内容
中文总结 AI 辅助
提出重叠-独特-冲突(OUC)跨叙述提取任务,构建约2.2万叙述对基准,评估14个开源LLM发现独特信息易提取而重叠与冲突难,微调可显著提升性能但仍有挑战。
中文摘要 AI 辅助
理解多视角的替代性叙述需要识别其信息在不同来源之间如何一致、冲突或不同。现有的跨文本关系研究主要集中于对预定义文本对之间的关系进行分类,如蕴含或矛盾,而非直接从完整叙述中提取此类信息。为填补这一空白,我们提出了重叠-独特-冲突(OUC)提取任务,这是一项跨叙述任务,旨在从两个叙述中提取所有重叠、冲突和独特的子句。为支持本研究,我们构建了一个包含约2.2万叙述对和14万OUC实例的基准,涵盖事实性、论证性和政治性话语。通过评估14个开源大语言模型(0.6B-35B),我们发现独特信息的提取远易于重叠和冲突信息:最强模型Gemma-4-31B在重叠任务上仅达到61.13%的F1分数,在冲突任务上为48.58%,而在独特任务上超过75%。进一步的诊断分析表明,这一困难并非仅源于关系识别,而是源于未能从完整叙述中配对并提取相应子句,尤其是在较小模型中。尽管如此,通过任务特定监督学习这些提取任务显著缩小了差距:微调后的Qwen-3-8B相比其基线提升了15-28个绝对百分点,并在多个任务上超越了规模约为其四倍的模型(如Qwen-3.6-35B)。即便如此,重叠和冲突的提取仍远未达到令人满意的水平,使得跨叙述子句提取仍是一个开放挑战。
英文摘要
Understanding multi-perspective alternative narratives requires identifying how their information agrees, conflicts, or differs across sources. Existing work on cross-text relations largely focuses on categorizing relations between predefined text pairs, such as entailment or contradiction, rather than directly extracting such information from full narratives. To address this gap, we introduce Overlap-Unique-Conflict (OUC) extraction, a cross-narrative task that extracts all overlapping, conflicting, and unique clauses from two narratives. To support this study, we construct a benchmark of approximately 22K narrative pairs and 140K OUC instances spanning factual, argumentative, and political discourse. Evaluating 14 open-source LLMs (0.6B-35B), we find that unique information is far easier to extract than overlap and conflict: the strongest model, Gemma-4-31B, reaches only 61.13% F1-score on overlap and 48.58% on conflict, against more than 75% on unique. Further diagnostic analysis reveals that this difficulty does not stem from relation recognition alone, but rather from a failure to pair and extract the corresponding clauses from full narratives, especially in smaller models. Nevertheless, learning these extractions with task-specific supervision narrows the gap considerably: a fine-tuned Qwen-3-8B gains 15-28% absolute over its baseline and surpasses models roughly four times its size (e.g., Qwen-3.6-35B) on several tasks. Even so, overlap and conflict remain well below satisfactory, leaving cross-narrative clause extraction an open challenge.
发表机构
- University of Central Florida(中佛罗里达大学)
机构由 AI 辅助整理,请以论文原文为准。