行锚定反馈可降低人工智能代码编辑中的令牌成本并提高正确性
Line-Anchored Feedback Cuts Token Costs and Improves Correctness in AI Code Editing
浏览论文内容
中文总结 AI 辅助
研究探讨生成式AI代码编辑中令牌成本等问题,通过比较整体提示与行锚定导出方式,发现行锚定反馈可降低令牌成本,减少强模型花费,提升弱模型正确性,在不同规模文件上有显著效果,且减轻编辑负担时正确性收益更大。
中文摘要 AI 辅助
生成的令牌是生成式人工智能(GAI)代码编辑成本、延迟和能耗的直接驱动因素。我们发现反馈格式对这三者都有影响。我们比较了两种提供相同请求更改的方式:整体提示(对照)与FileMark的结构化、行锚定导出(处理)。FileMark是一个用于对任何文件进行内联注释的VSCodium扩展。在配对实验中,行锚定使生成的令牌减少了22%(Claude Opus)和58%(Claude Sonnet),在100行或更多行的文件上达到24%-80%,七个模型中有四个在多重检验校正后生成的令牌显著减少。在模型有提升空间的情况下,正确性有所提高:汇总提高2.0分,五个本地模型中的三个提高5到7分。一项探索性实验表明,当编辑应用负担减轻时,正确性收益会进一步增加:在锚定条件下,1 Hundred+行文件上的本地模型正确性大约提高两倍。行锚定反馈减少了更强模型的花费,并提高了较弱模型的正确性。
英文摘要
Generated tokens are a direct driver of the cost, latency, and energy of generative AI (GAI) code editing. We show the format of feedback is a lever on all three. We compare two deliveries of the same requested changes: a holistic prompt (control) versus the structured, line-anchored export of FileMark (treatment). FileMark is a VSCodium extension for inline comments on any file. In a paired experiment line anchoring cut generated tokens by 22% (Claude Opus) and 58% (Claude Sonnet), reaching 24%-80% on files of 100 lines or more, with four of seven models generating significantly fewer tokens after multiple-testing correction. Correctness rose where models had headroom: +2.0 points pooled and +5 to +7 points for three of five local models. An exploratory experiment in which the harness, not the GAI model, applies function-level patches shows the correctness benefit grows further when the edit-application burden is lifted: local-model correctness on 100+ line files roughly triples under anchoring. Line-anchored feedback reduces what stronger models spend and improves what weaker models get right.