发表机构
PointFive(PointFive)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究通过预注册基准测试发现,提示词措辞和智能体框架会大幅影响大推理编码智能体的成本,部分提示指令可成倍增加推理成本却不提升任务成功率,框架选择的成本差异更显著。
AI 中文摘要
用作编码智能体的大推理模型会因推理、工具调用和多轮智能体交互产生成本,但提示词措辞对该成本的因果效应尚未被系统测量。我们在6个大推理模型、2个真实智能体框架和24个带有隐藏评估器的确定性编码任务上开展了一项预注册基准测试。在包含筛选、压力测试、保留、复制和跨服务商研究的4643次有效运行中,我们发现提示词表述方式可在不提升正确性的情况下成倍增加推理成本。要求模型开发并比较多种方法是最一致的浪费性指令,在所有模型中可使推理标记增加2.4至7.4倍;通用的“深度思考”提示也会使推理过程增加1.6至2.2倍;而指定范围、验收标准和停止条件的有界效率模板成本中性,可使推理量减半。框架选择的影响更大:相同的模型-任务-提示词三元组在Claude Code下每次成功的成本比pi高5至30倍,主要原因是静态前缀更大和交互轮次更多。误导性的架构提示比无关文字代价高得多,且服务商侧缓存会减少计费成本但不改变行为,因此不得将其视为效率提升。对Kimi-K3和Claude Sonnet 5的复制研究保留了主要效应方向,同时揭示了模型对思考和确定性提示的特定敏感性。总体而言,提示词措辞和框架设计会实质性影响智能体成本,且通常不会带来任务成功率的提升。
英文摘要
Coding-agent efficiency cannot be characterized by token count or model price alone. End-to-end cost and task success depend jointly on prompt semantics, inference effort, harness policy, model, task difficulty, tool use, context management, and provider accounting. Controlled experiments show that prompt wording can change reasoning and verification behavior without changing the task, that additional inference effort can help on difficult tasks but can also add cost without benefit, and that the value of an efficiency intervention can change when the harness changes. These results show that prompt, effort, and harness are interacting experimental factors rather than independent controls. We model efficiency as cost per successful task induced by the agent trajectory. Token and cache counts are measurements of that trajectory, not sufficient optimization targets. Agent evaluations should therefore measure success and end-to-end cost while controlling the system variables that determine how the trajectory is produced.