arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

QuoteBench:匹配分数如何掩盖命令路径故障

QuoteBench: How Matched Scores Can Hide Command-Path Failures

Shangao Li, Yao Zhang, Volker Tresp, Yuanyuan Yang

arXiv 2608.13547首次发表:更新:

发表机构

Stony Brook University; LMU Munich; Munich Center for Machine Learning(石溪大学; 慕尼黑大学; 慕尼黑机器学习中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

QuoteBench针对LLM编码智能体的命令路径故障,通过56个任务实验发现匹配分数掩盖执行缺陷,披露边界可恢复部分性能,提出评估需报告多维度配置而非仅匹配分数。

AI 中文摘要

大语言模型(LLM)编码智能体通过可能序列化、包装和重新解析模型输出的接口发出Bash命令。仅匹配执行分数无法区分命令生成错误与生成后引入的故障。QuoteBench通过对来自14个事件衍生家族的56个单样本任务进行精确最终状态验证,测量这一边界,在围绕一个故意未转义的附加解析器处跨越生成契约与执行传输。在插值点转义可重现每个重放回复的原始路径结果,因此在已披露边界下的任何恢复都必须来自模型改变其生成。在8个同窗口配置中,通过附加解析器重放相同回复使成功率降低55.4至73.2个百分点;对于6种配置,披露可恢复30.4至60.7个百分点,其余两种配置则为零或略为负。原始生成在前沿几乎饱和;边界适应仍是区分模型的关键。GPT-5.6-sol的匹配差距为-3.6点,掩盖了-64.3点的损害和+60.7点的补偿。部署配置会重新排序模型:26个可比对中有1个反转明确无误,另有4个处于单任务边际。对发出命令的智能体的评估应报告模型配置、生成契约、执行路径、操作点和最终状态验证器,而非将匹配分数视为模型的固有属性。

英文摘要

LLM coding agents issue Bash commands through interfaces that may serialize, wrap, and reparse model output. Matched execution scores alone cannot distinguish command-generation errors from failures introduced after generation. QuoteBench measures this boundary with exact final-state validation on 56 one-shot tasks from 14 incident-derived families, crossing the generation contract with the execution transport around one deliberately unescaped added parser. Escaping at the interpolation point reproduces each replayed reply's raw-path outcome, so any recovery under a disclosed boundary must come from the model changing its generation. Across eight same-window configurations, replaying the same reply through the added parser lowers success by 55.4 to 73.2 percentage points; disclosure recovers 30.4 to 60.7 points for six configurations, and zero or slightly negative for the other two. Raw generation is nearly saturated at the frontier; boundary adaptation is what still separates models. GPT-5.6-sol's matched gap of -3.6 points hides -64.3 points of damage and +60.7 points of compensation. The deployment configuration reorders models: one reversal among 26 comparable pairs is unambiguous and four more sit on single-task margins. Evaluations of command-issuing agents should report the model configuration, generation contract, execution path, operating point, and final-state validator rather than treat a matched score as an intrinsic model property.

Comments29 pages, 5 figures. Project page: https://quotebench.lsamc.website/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑