发表机构
Stony Brook University; LMU Munich; Munich Center for Machine Learning(石溪大学; 慕尼黑大学; 慕尼黑机器学习中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
QuoteBench针对LLM编码智能体的命令路径故障,通过56个任务实验发现匹配分数掩盖执行缺陷,披露边界可恢复部分性能,提出评估需报告多维度配置而非仅匹配分数。
AI 中文摘要
大语言模型(LLM)编码智能体通过可能序列化、包装和重新解析模型输出的接口发出Bash命令。仅匹配执行分数无法区分命令生成错误与生成后引入的故障。QuoteBench通过对来自14个事件衍生家族的56个单样本任务进行精确最终状态验证,测量这一边界,在围绕一个故意未转义的附加解析器处跨越生成契约与执行传输。在插值点转义可重现每个重放回复的原始路径结果,因此在已披露边界下的任何恢复都必须来自模型改变其生成。在8个同窗口配置中,通过附加解析器重放相同回复使成功率降低55.4至73.2个百分点;对于6种配置,披露可恢复30.4至60.7个百分点,其余两种配置则为零或略为负。原始生成在前沿几乎饱和;边界适应仍是区分模型的关键。GPT-5.6-sol的匹配差距为-3.6点,掩盖了-64.3点的损害和+60.7点的补偿。部署配置会重新排序模型:26个可比对中有1个反转明确无误,另有4个处于单任务边际。对发出命令的智能体的评估应报告模型配置、生成契约、执行路径、操作点和最终状态验证器,而非将匹配分数视为模型的固有属性。
英文摘要
LLM coding agents issue Bash commands through interfaces that may serialize, wrap, and reparse model output. Matched execution scores alone cannot distinguish command-generation errors from failures introduced after generation. QuoteBench measures this boundary with exact final-state validation on 56 one-shot tasks from 14 incident-derived families, crossing the generation contract with the execution transport around one deliberately unescaped added parser. Escaping at the interpolation point reproduces each replayed reply's raw-path outcome, so any recovery under a disclosed boundary must come from the model changing its generation. Across eight same-window configurations, replaying the same reply through the added parser lowers success by 55.4 to 73.2 percentage points; disclosure recovers 30.4 to 60.7 points for six configurations, and zero or slightly negative for the other two. Raw generation is nearly saturated at the frontier; boundary adaptation is what still separates models. GPT-5.6-sol's matched gap of -3.6 points hides -64.3 points of damage and +60.7 points of compensation. The deployment configuration reorders models: one reversal among 26 comparable pairs is unambiguous and four more sit on single-task margins. Evaluations of command-issuing agents should report the model configuration, generation contract, execution path, operating point, and final-state validator rather than treat a matched score as an intrinsic model property.
Comments29 pages, 5 figures. Project page: https://quotebench.lsamc.website/