arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06842cs.CLcs.AI

XYBench:大语言模型能否对带有误解的查询做出务实回应?

XYBench: Can LLMs Respond Pragmatically to Queries with Misconceptions?

Akhila Yerukola, Jena D. Hwang, Mingqian Zheng, Jenna Godsey, Hyunwoo Kim, Valentina Pyatkin, Jennifer Hu, Maarten Sap

首次发表
浏览论文内容

中文总结 AI 辅助

XYBench提出含8,115个带误解查询的基准,评估LLM识别误解并提供务实方案的能力,发现最强模型仍主要回答字面请求,在务实重定向上远逊于人类。

中文摘要 AI 辅助

当非专家用户向大语言模型(LLM)寻求帮助时,他们的查询往往可能带有误解(例如,“如何用正则表达式解析XML?”)。这类情况通常被称为XY问题,此时LLM必须识别出误解(“正则表达式很脆弱”)并富有意义地引导用户走向一个务实的解决方案,以解决请求中隐含的根本问题(“使用XML解析器”)。我们引入了XYBench,这是一个包含8,115个此类查询的基准测试,这些查询来自技术领域(StackOverflow/StackExchange)和日常领域(WikiHow及一个手工精选的子集)。我们设计了一种评估范式,基于合作回应理论,从三个标准评估模型回应:(a)务实解决方案的存在性,(b)对其的强调程度,以及(c)对误解的识别。我们的实验表明,即使是最强的大语言模型也主要回答字面请求(0.75至0.92),而较少回答预期意图(0.33至0.71),同时在识别误解方面大幅落后于人类(最多63%对比79%至90%)。此外,模型在多项选择设置中压倒性地偏好务实回应,却始终未能生成此类回应。消融实验表明,在生成时提供明确的用户意图有所帮助;然而,仍存在巨大差距,这表明务实重定向是当前大语言模型中一项根本未发展成熟的能力。

英文摘要

When non-expert users ask LLMs for assistance, their queries can often have misconceptions (e.g., "How do I parse XML with regex?"). In such cases, often referred to as the XY-problem, LLMs must identify the misconception ("regex are fragile") and meaningfully direct the user toward a pragmatic solution that will address the root problem implicit in the request ("use an XML parser"). We introduce XYBench, a benchmark of 8,115 such queries, drawn from technical (StackOverflow/StackExchange) and everyday (WikiHow and a manually-curated subset) domains. We design an evaluation paradigm that assesses model responses along three criteria grounded in cooperative response theory: (a) presence and (b) emphasis on pragmatic solutions, and (c) identification of misconceptions. Our experiments show that even the strongest LLMs predominantly answer the literal request (0.75--0.92) and far less often the intended one (0.33--0.71), while substantially lagging behind humans at identifying misconceptions (at most 63% vs. 79--90%). Further, models overwhelmingly prefer pragmatic responses in a multiple choice setting yet consistently fail to generate them. Oracle ablation experiments show that providing explicit user intent at generation time helps; however a large gap remains, suggesting pragmatic redirection is a fundamentally underdeveloped capability in current LLMs.

发表机构

  • Carnegie Mellon University(卡内基梅隆大学)
  • Allen Institute for AI(艾伦人工智能研究所)
  • NVIDIA(英伟达)
  • Johns Hopkins University(约翰斯·霍普金斯大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑