发表机构
Peking University; Wuhan University(北京大学; 武汉大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
CONTRA是一种无需训练的方法,通过生成候选问题并基于语义和执行筛选,在LLM代码生成中实现选择性澄清,提升澄清效果并减少不必要提问。
AI 中文摘要
编码智能体可以生成看似正确但实现的行为并非用户意图的代码。当智能体通过自身假设静默地解决需求不明确的问题时,可能会出现这种不匹配。随着后续开发建立在这些假设之上,纠正由此产生的行为的成本可能越来越高。早期澄清有助于防止此类不匹配,但不必要的问题会打断开发者并减缓开发速度。现有方法难以在识别关键澄清问题的同时避免不必要的问题。因此,我们提出CONTRA,一种无需训练的方法,将广泛的问题发现与基于语义和执行的问題筛选相结合。CONTRA首先生成候选问题,并过滤掉那些与所需行为无关或已被需求解决的问题。对于每个剩余问题,它生成基于两个合理答案的程序,并检查在共享输入上是否存在稳定的行为差异。然后,它利用交互历史在合格问题中进行选择或停止提问。在ClarifyCodeBench上的实验表明,CONTRA在所有四种编码智能体上取得了最高的F1分数,比最佳基线的宏平均F1高出13.88个百分点。在相同的LLM和评估协议下,CONTRA在澄清召回率和F1上也优于编码工具Claude Code和OpenHands。为支持实际使用,我们还实现了CONTRA作为Claude Code插件,将选择性澄清集成到日常开发中。
英文摘要
Coding agents can generate code that appears correct but implements behavior the user never intended. This mismatch can arise when an agent silently resolves underspecified requirements through its own assumptions. As subsequent development builds on these assumptions, correcting the resulting behavior can become increasingly costly. Early clarification can help prevent such mismatches, but unnecessary questions can interrupt developers and slow down development. Existing methods struggle to identify key clarification questions while avoiding unnecessary ones. Therefore, we propose CONTRA, a training-free method that combines broad question discovery with semantic and execution-based question qualification. CONTRA first generates candidate questions and filters out those unrelated to required behavior or already resolved by the requirement. For each remaining question, it generates programs conditioned on two plausible answers and checks for stable behavioral differences on shared inputs. It then uses the interaction history to select among qualified questions or stop asking. Experiments on ClarifyCodeBench show that CONTRA achieves the highest F1 with all four coding agents, exceeding the best baseline macro-average F1 by 13.88 percentage points. With the same LLM and evaluation protocol, CONTRA also achieves higher clarification recall and F1 than the coding harnesses Claude Code and OpenHands. To support practical use, we also implement CONTRA as a Claude Code plugin that integrates selective clarification into everyday development.
Comments15 pages. Code: https://github.com/fangz-cs/Contra