推理努力,而非工具访问,决定了智能体代码生成的首试可靠性:一项观察性研究
Reasoning effort, not tool access, buys first-try reliability in agentic code generation: an observational study
浏览论文内容
中文总结 AI 辅助
本研究通过90次独立智能体运行构建同一应用,发现推理努力(从高到极高)将首试完美运行率从28%提升至89%,而测试工具虽增加成本却未改善功能得分或可靠性。
中文摘要 AI 辅助
智能体编码助手正被赋予额外能力,如基于浏览器的测试工具和面向设计的系统提示,其假设是更多能力能带来更好的软件。本研究直接检验了这一假设。90次独立智能体运行根据同一详细规格构建了同一应用——一个实时回顾板,每次运行根据固定的14项功能标准(最高42分)和视觉质量评审进行评分。这些运行跨越了多个模型世代、两个智能体框架、两个推理努力级别、一个测试工具和两个面向设计的提示。能力层级占主导地位:前沿模型接近上限,而低成本本地模型得分在24到37分之间。标准层面分析揭示了运行总数所隐藏的内容。容器部署是主要缺陷,在44%的运行中首试失败,其失败率在不同模型世代间急剧变化,而平均总分变化不到1分。测试工具使成本增加42%至68%,但未改善功能得分或可靠性,即使在界面可见标准上也是如此。将推理努力从高提升到极高,首试完美运行率从28%提高到89%,纠正提示减少约五倍,成本增加9%至29%。面向设计的提示提高了视觉质量(5分制中4.5分对比3.0分),但未提升功能,其指令的一段释义复现了全部提升。实际教训是匹配修复与失败:大多数首试失败源于弱推理,这可以通过更强的模型或更多努力来预防,而非检查工具能捕获的可见缺陷。
英文摘要
Agentic coding assistants are increasingly given extra capabilities, such as browser based testing tools and design oriented system prompts, on the assumption that more capability yields better software. This study tested that assumption directly. Ninety independent agent runs built the same application, a real time retrospective board, from one detailed specification, each scored on a fixed 14 criterion functional rubric (42 point maximum) and a visual quality review. The runs spanned several model generations, two agent harnesses, two reasoning effort levels, a testing tool, and two design oriented prompts. Capability tier dominated: frontier models clustered near the ceiling while a low cost local model fell to 24 to 37 points. A criterion level analysis revealed what run totals conceal. Container deployment was the dominant defect, failing first try in 44 percent of runs, with its failure rate shifting sharply across model generations while mean totals moved less than a point. The testing tool raised cost by 42 to 68 percent without improving functional score or reliability, even on interface visible criteria. Raising reasoning effort from High to xHigh lifted first try perfect runs from 28 percent to 89 percent and cut corrective prompts about five fold, for 9 to 29 percent more cost. A design oriented prompt raised visual quality, 4.5 versus 3.0 on a 5 point scale, without lifting function, and a one paragraph paraphrase of its directive reproduced the entire lift. The practical lesson is to match the fix to the failure: most first run failures came from weak reasoning, which a stronger model or more effort prevents, not from visible flaws a checking tool would catch.
发表机构
- TrendAI
机构由 AI 辅助整理,请以论文原文为准。