arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19070cs.CL

字里行间:大语言模型能否发现文本背后的问题?

Reading Between the Lines: Can LLMs Discover the Question Behind the Text?

  • University of Bucharest(布加勒斯特大学)

机构由 AI 辅助整理,请以论文原文为准。

Claudiu Creanga, Liviu P. Dinu

中文总结 AI 辅助

本文提出“问题考古学”任务,通过新数据集评估大语言模型推断文本背后作者意图的能力,发现当前模型超越人类,为AI理解人类交流提供新基准。

中文摘要 AI 辅助

本文引入了“问题考古学”,这是一个专门的评估任务,旨在推断出促使一篇完整文本被创作的单一、真实的“起源问题”。与针对任何合理问题的问题生成任务,或对话语层面行为进行建模的话语框架不同,我们的任务评估的是模型对作者意图的把握。我们提出了一个新的数据集,其中包含委托撰写的文本及其原始研究问题和合理的干扰项。我们对专有模型(如Gemini Flash和Pro)以及开源模型(如Mistral和Qwen)的评估显示,该任务取得了显著进展,较新版本的模型优于早期版本,而基于BERT的模型表现不佳。值得注意的是,我们的研究结果表明,当前的大语言模型在此任务上超越了人类表现,表明其对作者意图具有高级理解。这一能力对人工智能在需要细致解读人类交流的任务中的作用具有重要意义。因此,我们的工作为未来模型提供了一个新的框架和具有挑战性的基准。

英文摘要

This paper introduces ``question archaeology'', a specific evaluation task focused on inferring the single, authentic "genesis question" that motivated the creation of a complete text. Distinct from question generation, which targets any plausible question, or discourse frameworks that model utterance-level acts, our task assesses a model's grasp of authorial intent. We present a new dataset of commissioned texts paired with their original research questions and plausible distractors. Our evaluation of both proprietary models, like Gemini Flash and Pro, as well as open source models like Mistral and Qwen, reveals significant progress in this task, with the newer versions outperforming the earlier ones, while BERT-based models performed poorly. Notably, our findings indicate that current LLMs surpass human performance on this task, suggesting advanced understanding of authorial intent. This capability has important implications for AI's role in tasks requiring nuanced interpretation of human communication. Our work thus provides a new framework and a challenging benchmark for future models.

↑