arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DevIntent:大语言模型生成的代码在多大程度上违反开发者意图?

DevIntent: How Much Does LLM-Generated Code Violate Developer Intent?

Susana Haing, Natan Vidra, Spurthi Setty

arXiv 2608.07614首次发表:更新:

发表机构

Anote AI(Anote AI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对LLM生成代码违反开发者隐式意图的问题,构建含49个问题的基准并提出IVR指标,发现两款主流LLM生成代码虽通过率高但超半数违反意图,揭示现有基准局限性。

AI 中文摘要

当给定模糊提示时,大语言模型(LLM)生成的代码可能违反开发者的隐式意图,而标准基准仅衡量代码是否通过其明确的测试。我们引入了意图违反率(IVR)和一个源自HumanEval+的含49个问题的试点基准,每个问题从澄清后的提示中剥离隐式约束并将其编码为隐藏约束测试。IVR用于衡量LLM生成的解决方案通过明确(可见)测试但未通过捕捉未陈述意图的隐藏约束测试的比例。评估Claude Sonnet 4.6和OpenAI GPT 4.1时,发现两者均通过了超过92%的明确测试,但在超过一半的问题中违反了意图(分别为54.5%和63.5%),且呈现出两种模型一致的系统性双峰模式。研究结果表明,通过率夸大了生成代码反映开发者意图的程度。

英文摘要

Code generated by LLMs can violate a developer's implicit intentions when given an ambiguous prompt, yet standard benchmarks measure only whether code passes its stated test. We introduce the Intent Violation Rate (IVR) and a 49-problem pilot benchmark derived from HumanEval+. Each problem strips implicit constraints from a clarified prompt and encodes them as hidden constraint tests. IVR measures the fraction of LLM-generated solutions that pass the stated (visible) tests yet fail hidden constraint tests that capture unstated intent. Evaluating Claude Sonnet 4.6 and OpenAI GPT 4.1, we find both pass over 92\% of stated tests yet violate intent in over half of problems (54.5\% and 63.5\%), following a systematic, bimodal pattern consistent across both models. Out findings indicate that pass rates overstate how well generated code reflects developer intent.

Comments8 pages, 2 figures, 8 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑