arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.00711cs.SEcs.AI

ClarifyCodeBench: 评估大语言模型在代码生成中澄清模糊需求的能力

ClarifyCodeBench: Evaluating LLMs on Clarifying Ambiguous Requirements for Code Generation

Zheng Fang, Dongming Jin, Yihong dong, Yongmin Li, Kechi Zhang, Zhi Jin, Ge Li

首次发表
浏览论文内容

中文总结 AI 辅助

提出ClarifyCodeBench基准,通过人工标注的模糊需求、澄清问题和指标,评估LLMs主动澄清需求的能力,发现代码生成能力与需求澄清能力脱钩,且多模糊场景下性能下降。

中文摘要 AI 辅助

大语言模型已成为编程助手。然而,代码生成的效果受限于输入需求的质量,这些需求常常模糊、不完整或未充分指定。虽然LLMs擅长一次性代码合成,但它们主动澄清意图的能力——作为稳健软件工程的关键特质——仍未被充分探索。现有基准大多忽略了这一交互瓶颈,假设提示完美指定,未能反映需求获取的迭代本质。为填补这一空白,我们引入了ClarifyCodeBench,一个新颖的交互式基准,用于评估LLMs解决需求模糊性的能力。ClarifyCodeBench基于真实世界的编程任务构建,具有高质量的人工标注,包括N种独特的模糊类型、相关的澄清问题以及对应的真实答案。此外,我们形式化了两个严格的指标来评估交互质量:轮次折扣关键问题率(惩罚低效提问)和最优轮次遵守度(衡量获取过程的精确性)。我们使用ClarifyCodeBench对六个最先进的LLMs进行了系统评估。实证结果揭示了三个关键见解:1)能力解耦:强大的代码生成性能并不自然转化为有效的需求澄清;2)推理悖论:增加的计算思维提升了代码正确性,但在识别模糊性方面收益甚微;3)多模糊上限:随着模糊密度增加,LLMs的澄清性能急剧下降,揭示了处理复杂真实世界规范时的显著瓶颈。我们的工作强调了未来AI4SE研究从静态合成向交互式获取转变的必要性。

英文摘要

Large Language Models have emerged as programming assistants. However, the efficacy of code generation is constrained by the quality of input requirements, which are frequently ambiguous, incomplete, or underspecified. While LLMs excel at one-shot code synthesis, their ability to proactively clarify intent remains underexplored, as a critical trait for robust software engineering. Existing benchmarks largely overlook this interactive bottleneck, assuming perfectly specified prompts that do not reflect the iterative nature of requirement elicitation. To bridge this gap, we introduce ClarifyCodeBench, a novel interactive benchmark for evaluating LLMs' capability in resolving requirement ambiguity. Constructed from real-world programming tasks, ClarifyCodeBench features high-quality manual annotations, including N unique ambiguity types, associated clarification questions, and corresponding ground-truth answers. Furthermore, we formalize two rigorous metrics to assess the interaction quality: Turn-discounted Key Question Rate, which penalizes inefficient questioning, and Optimal Round Adherence, which measures the precision of the elicitation process. We conduct a systematic evaluation of six state-of-the-art LLMs using ClarifyCodeBench. Our empirical results yield three critical insights: 1) Capability Decoupling: Strong code generation performance does not inherently translate to effective requirement clarification; 2) The Reasoning Paradox: While increased computational thinking enhances code correctness, it yields marginal gains in identifying ambiguities; 3) The Multi-ambiguity Ceiling: LLMs' clarification performance degrades sharply as the density of ambiguities increases, revealing a significant bottleneck in handling complex, real-world specifications. Our work underscores the necessity for future AI4SE research to transition from static synthesis to interactive elicitation.

发表机构

  • Peking University(北京大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑