效率幻觉:形式化并测量基于LLM的代码优化中的行为校准
Efficiency Hallucination: Formalizing and Measuring Behavioral Calibration in LLM-Based Code Optimization
- Google(谷歌)
- Columbia University(哥伦比亚大学)
- Vanderbilt University(范德堡大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对LLM在代码优化中的“效率幻觉”问题,提出基于分类惩罚的验证框架,在九个模型上测试,将正确弃权率从0%提升至44.4%,且不牺牲次优代码编辑率,实现无需训练的校准。
AI中文摘要:
将大型语言模型(LLM)集成到自动化代码优化中引入了一个关键可靠性风险,我们将其称为“效率幻觉”:LLM倾向于在已优化的代码上发出非功能性修改,并附带未经证实的性能声明。这一现象由“评估陷阱”驱动,即二元基准测试激励不必要的修改,而非安全地弃权(不执行)。我们提出了一个使用分类惩罚方法的验证框架,并利用EffiBench在九个模型(GPT、Claude、Gemini)上进行了180次优化运行评估。在标准提示下,模型对最优代码的过度编辑率达到100%。我们的防护措施将正确的弃权(不执行)率从0%提升至44.4%,同时在对次优代码的编辑上保持100%的编辑率,且零错误弃权(不执行)。校准并不均匀:GPT-5.4 Mini接近完美的弃权(不执行)表现,简单代码比复杂代码更容易被可靠识别。我们的框架提供了一种无需训练的机制,可在生产部署前缓解LLM的过度自信。
英文摘要:
The integration of Large Language Models (LLMs) into automated code optimization introduces a critical reliability risk we term the Efficiency Hallucination: an LLM's tendency to issue non-functional mutations with unsubstantiated performance claims on already-optimized code. This is driven by the Evaluation Trap, wherein binary benchmarks incentivize unnecessary modifications over safely abstaining. We present a validation framework using classification penalty methods, evaluated across 180 optimization runs on nine models (GPT, Claude, Gemini) using EffiBench. Under standard prompts, models exhibit a 100% over-edit rate on optimal code. Our guardrail raises correct abstention from 0% to to 44.4%, preserving a 100% edit rate on sub-optimal code with zero false abstentions. Calibration is uneven: GPT-5.4 Mini approaches near-perfect abstention, and simple code is recognized more reliably than complex code. Our framework offers a training-free mechanism to mitigate LLM overconfidence before deployment in production.