XYBench:大语言模型能否对带有误解的查询做出务实回应?
XYBench: Can LLMs Respond Pragmatically to Queries with Misconceptions?
浏览论文内容
中文总结 AI 辅助
XYBench提出含8,115个带误解查询的基准,评估LLM识别误解并提供务实方案的能力,发现最强模型仍主要回答字面请求,在务实重定向上远逊于人类。
中文摘要 AI 辅助
当非专家用户向大语言模型(LLM)寻求帮助时,他们的查询往往可能带有误解(例如,“如何用正则表达式解析XML?”)。这类情况通常被称为XY问题,此时LLM必须识别出误解(“正则表达式很脆弱”)并富有意义地引导用户走向一个务实的解决方案,以解决请求中隐含的根本问题(“使用XML解析器”)。我们引入了XYBench,这是一个包含8,115个此类查询的基准测试,这些查询来自技术领域(StackOverflow/StackExchange)和日常领域(WikiHow及一个手工精选的子集)。我们设计了一种评估范式,基于合作回应理论,从三个标准评估模型回应:(a)务实解决方案的存在性,(b)对其的强调程度,以及(c)对误解的识别。我们的实验表明,即使是最强的大语言模型也主要回答字面请求(0.75至0.92),而较少回答预期意图(0.33至0.71),同时在识别误解方面大幅落后于人类(最多63%对比79%至90%)。此外,模型在多项选择设置中压倒性地偏好务实回应,却始终未能生成此类回应。消融实验表明,在生成时提供明确的用户意图有所帮助;然而,仍存在巨大差距,这表明务实重定向是当前大语言模型中一项根本未发展成熟的能力。
英文摘要
When non-expert users ask LLMs for assistance, their queries can often have misconceptions (e.g., "How do I parse XML with regex?"). In such cases, often referred to as the XY-problem, LLMs must identify the misconception ("regex are fragile") and meaningfully direct the user toward a pragmatic solution that will address the root problem implicit in the request ("use an XML parser"). We introduce XYBench, a benchmark of 8,115 such queries, drawn from technical (StackOverflow/StackExchange) and everyday (WikiHow and a manually-curated subset) domains. We design an evaluation paradigm that assesses model responses along three criteria grounded in cooperative response theory: (a) presence and (b) emphasis on pragmatic solutions, and (c) identification of misconceptions. Our experiments show that even the strongest LLMs predominantly answer the literal request (0.75--0.92) and far less often the intended one (0.33--0.71), while substantially lagging behind humans at identifying misconceptions (at most 63% vs. 79--90%). Further, models overwhelmingly prefer pragmatic responses in a multiple choice setting yet consistently fail to generate them. Oracle ablation experiments show that providing explicit user intent at generation time helps; however a large gap remains, suggesting pragmatic redirection is a fundamentally underdeveloped capability in current LLMs.
发表机构
- Carnegie Mellon University(卡内基梅隆大学)
- Allen Institute for AI(艾伦人工智能研究所)
- NVIDIA(英伟达)
- Johns Hopkins University(约翰斯·霍普金斯大学)
机构由 AI 辅助整理,请以论文原文为准。