arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

StateM:通过工具链扩展在Terminal-Bench 2.1上达到95.3%的原始准确率,或实现15美元的前沿运行

StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling

Ziheng Qin, Yaxin Lu, Zhangyang Atlas Wang, Kai Wang

arXiv 2608.15089首次发表:更新:

AI 中文总结

本研究提出智能体原生运行时StateM,通过工具链扩展优化智能体执行系统,在Terminal-Bench 2.1等基准上显著提升GPT-5.6、DeepSeek-V4 Flash等模型的任务准确率,大幅降低运行成本,验证了其有效性与泛化性。

AI 中文摘要

长程智能体即使其底层模型能够解决各个组成步骤,仍可能出现失败情况。它们可能会丢失可变状态的跟踪、无法重新激活早期执行的经验、跳过已知流程或过早停止。我们押注于工具链扩展,以在不改变智能体模型权重的情况下改进其执行系统。我们推出StateM,这是一种智能体原生运行时,围绕持久状态、阶段局部上下文、经检查的转换、可恢复的执行手册(runbooks)以及智能体和用户可共同检查的程序化实践来组织执行。在Terminal-Bench 2.1上,StateM将GPT-5.5 xhigh的准确率提升至92.1%,而参考值为83.1%,GPT-5.6 Sol Ultra为91.9%;执行手册无需修改即可迁移至GPT-5.6。使用GPT-5.6 Sol xhigh,StateM在445次试验中达到95.3%的原始准确率,且在全部89个任务中至少成功一次。冻结配置将GPT-5.6 Luna的准确率从76.7%提升至85.4%,高于84.9%的Sol xhigh参考值。使用相同的运行时、执行手册结构和黄金规则,不到38美元的适配将DeepSeek-V4 Flash在标准超时下的准确率从82.7%提升至88.1%,在88任务的通用核心上提升至89.1%;仅扩展剩余的延迟敏感任务即可达到报告的GPT-5.6 Sol 88.8%的最高结果。最终评分API使用成本约为15美元,而GPT参考值为574.68美元;DeepSeek总支出为52.22美元。在BusinessBench上,基于开发集构建的特定家族执行手册在保留集上分别获得0.55宏观和1.34微观分数提升,两个机制匹配的家族提升10.04点。当任务共享执行结构时,具体规则具有泛化性,而控制方法适用范围广泛。StateM将选定的事后分析结果转化为持久、可执行的前提和实践,通过有状态控制使学习到的控制明确且可执行。代码见此http URL。

英文摘要

Long-horizon agents can fail even when their underlying models can solve the constituent steps. They may lose track of mutable state, fail to reactivate lessons from earlier executions, skip known procedures, or stop prematurely. We bet on harness scaling to improve the execution system around an agent without changing its model weights. We introduce StateM, an agent-native runtime that organizes execution around durable states, phase-local context, checked transitions, recoverable runbooks, and versioned procedural practices that agents and users can inspect together. On Terminal-Bench 2.1, StateM raises GPT-5.5 xhigh to 92.1\%, versus 83.1\% reference and GPT-5.6 Sol Ultra at 91.9\%. The runbook transfers unchanged to GPT-5.6. With GPT-5.6 Sol xhigh, StateM reaches 95.3\% raw accuracy across 445 trials and succeeds on all 89 tasks at least once. The frozen profile raises GPT-5.6 Luna from 76.7 to 85.4\%, above the 84.9\% Sol xhigh reference. Using the same runtime, runbook structure, and golden rules, less than \$38 of adaptation raises DeepSeek-V4 Flash from 82.7 to 88.1\% under standard timeouts and to 89.1\% on an 88-task common core. Extending only the remaining latency-sensitive task matches the reported 88.8\% GPT-5.6 Sol max result. Final-score API usage is about \$15 versus \$574.68 for the GPT reference; total DeepSeek expenditure is \$52.22. On BusinessBench, family-specific runbooks built on development sets yield held-out gains of 0.55 macro and 1.34 micro points; two mechanism-matched families improve by 10.04 points. Concrete rules generalize when tasks share execution structure, while the control methodology applies broadly. StateM turns selected postmortem findings into persistent, executable preconditions and practices, making learned controls explicit and enforceable through stateful controls. Code at github.com/henryqin1997/statem.

CommentsHarness Scaling, Semi-Self-Evolving Agent

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑