arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.25647cs.AI

基于轨迹的LLM智能体早期结果预测的测试驱动可靠性审计:目标特定的校准迁移在单一基准内持续存在

Testing-Driven Reliability Audit of Trajectory-Based Early Outcome Prediction for LLM Agents: Target-Specific Calibration Transfer Persists Within a Single Benchmark

YanZe Cao

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过留一智能体出审计等方法,发现LLM智能体早期结果预测的校准迁移错误并非广泛存在,而是集中在特定目标组合上,且不跨基准泛化。

中文摘要 AI 辅助

基于轨迹预测早期结果可以通过在结果变得足够可预测时终止运行来降低智能体评估的费用,前提是预测器的置信度得到适当校准。当预测器应用于其从未训练过的智能体时,校准存在风险,但尚不清楚这种迁移失败是广泛存在于整个智能体系统中,还是集中在特定的目标智能体/头组合中。使用公开的SWE-bench Verified轨迹和冻结的双头早期结果预测流水线,我们进行了留一智能体出校准审计、共享预测器留二智能体出对照、预言机先验修正,以及针对训练队列、任务重采样、任务半区、刀切法和阈值的稳健性测试。固定脚手架TerminalBench分析作为预注册边界测试。广泛的同预测器成对异质性未得到支持;中位成对修正差距分别为0.0180(SUCCESS头,45对)和0.0385(FAILURE头,35对),且预注册的异质性标准在两个头上均未满足。两个特定组合,gpt-5-mini/SUCCESS和claude-opus-4.6/FAILURE,表现出持续的校准迁移错误(中位修正差距分别为0.1377和0.1107),在任何冻结对照下均无符号反转。TerminalBench未建立跨基准复制:成功目标产生零决策(INDETERMINATE),失败目标未满足预注册的持续性标准。因此,强目标特定的校准迁移错误可以存在于单个冻结环境中,但证据并未确立该错误是模型固有的或跨基准普遍存在的。

英文摘要

Predicting early outcomes based on trajectory can decrease the expenses associated with agent evaluation by terminating a run once the outcome becomes sufficiently predictable, assuming that the predictor's confidence is properly calibrated. Calibration is at risk when a predictor is applied to an agent on which it was never trained, but it is not known whether such transfer failures are broad across agent systems or concentrated in specific target agent/head combinations. Using public SWE-bench Verified trajectories and a frozen dual-head early-outcome prediction pipeline, we ran a leave-one-agent-out calibration audit, a shared-predictor leave-two-agents-out control, oracle prior correction, and a robustness battery over training cohorts, task resampling, task halves, jackknife, and thresholds. Fixed-scaffold TerminalBench analysis served as a pre-registered boundary test. Broad same-predictor pairwise heterogeneity was not supported; the median pairwise corrected-gap differences were 0.0180 (SUCCESS head, 45 pairs) and 0.0385 (FAILURE head, 35 pairs), and the pre-registered heterogeneity criterion was not met on either head. Two specific combinations, gpt-5-mini/SUCCESS and claude-opus-4.6/FAILURE, showed persistent calibration-transfer errors (median corrected gaps 0.1377 and 0.1107) without a sign reversal under any frozen control. TerminalBench did not establish cross-benchmark replication: the success target produced zero decisions (INDETERMINATE), and the failure target did not satisfy the pre-registered persistence criterion. Therefore, a strong target-specific calibration-transfer error can exist within one frozen environment, but the evidence does not establish that the error is intrinsic to the model or general across benchmarks.

发表机构

  • Lizhi College, Shaanxi University of Technology(陕西理工大学砺志学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑