arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

评估LLM能否可靠地连接DOTs?

Evaluating Whether LLMs Can Reliably Connect the DOTs?

Eftekhar Hossain, John Salvador, Santu Karmaker

arXiv 2609.38406首次发表:更新:

发表机构

University of Central Florida(中佛罗里达大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文构建约9.2K实例的多领域叙事填充基准,评估20个开源LLM,发现模型规模不能可靠预测填充质量,Gemma-2-2B表现最佳,显式推理收益有限。

AI 中文摘要

现实世界的信息获取往往充满噪声且碎片化。从这些碎片中构建连贯的叙述,要求模型在更广泛的故事线中重建缺失的片段,这通常被称为文本填充(text infilling),同时保持与局部上下文和全局故事线的一致性。尽管许多大型语言模型(LLMs)将文本填充作为预训练目标,但它们在现实世界叙事填充上的实际表现仍未得到充分探索。在本文中,我们通过引入一个包含约9.2K个实例的多领域叙事填充基准来填补这一空白,该基准通过遮蔽四种叙事类型(百科全书文本、常识故事、新闻文章和视觉叙事)中的一到三个句子构建而成。利用该基准,我们评估了20个指令微调的开源LLM,参数规模从1.5B到70B不等,涵盖不同级别的指令具体性和推理指导。输出使用标准自动指标和涵盖五个叙事维度的定性框架进行评估。结果表明,模型规模并不能可靠预测填充质量:Gemma-2-2B获得了最高的定性得分(4.02/5),优于其十倍以上规模的模型,包括DeepSeek-Qwen-32B(3.77/5,6.6%)和LLaMA-3.3-70B(3.71/5,8.3%)。我们进一步发现,显式推理带来的收益有限,因为思维链推理仅产生了边际改进(+0.6%)。此外,对于当前LLMs中的叙事填充,短叙事和领域特征比填充位置本身更能成为任务难度的强预测因子。

英文摘要

Access to real-world information is often noisy and fragmented. Constructing a coherent narrative from such fragments requires models to reconstruct missing spans within a broader storyline, commonly referred to as text infilling, while preserving consistency with both the local context and the global storyline. Despite using text infilling as a pre-training objective in many Large Language Models (LLMs), their actual performance on real-world narrative infilling remains underexplored. In this paper, we address this gap by introducing a multi-domain benchmark of ~9.2K instances for narrative infilling, constructed by masking one to three sentences across four narrative types: encyclopedic text, commonsense stories, news articles, and visual narratives. Using this benchmark, we evaluate 20 instruction-tuned open-source LLMs ranging from 1.5B to 70B parameters across varying levels of instruction specificity and reasoning guidance. Outputs are assessed using standard automatic metrics and a qualitative framework covering five narrative dimensions. Results show that model scale does not reliably predict infilling quality: Gemma-2-2B achieves the highest qualitative score (4.02/5), outperforming models over ten times larger, including DeepSeek-Qwen-32B (3.77/5, 6.6%) and LLaMA-3.3-70B (3.71/5, 8.3%). We further find that explicit reasoning offers limited benefits as chain-of-thought reasoning yields only a marginal improvement (+0.6%). Additionally, short narratives and domain characteristics emerge as stronger predictors of task difficulty than infill position alone for narrative infilling in current LLMs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑