arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

寻找合适搭配:智能体任务中的模型-执行框架交互

Finding the Right Fit: Model-Harness Interactions across Agent Tasks

Yixuan Li, Yiyun Zhou, Yao Long Teng, Fuchao Yang, Yanchen Deng, Zhiyi Lyu, Xuyu Dong, Feng Chen, Bo An

arXiv 2610.00917首次发表:更新:

发表机构

College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究评估66种模型-执行框架配置在三个基准上的表现,发现模型排名随框架反转,强调模型、框架和任务需协同评估,并发布全部6,204条轨迹数据。

AI 中文摘要

选择智能体系统意味着同时选择语言模型及其执行框架。我们探究当环境改变时,一个强大的模型、执行框架或配对是否仍能保持其优势。我们评估了66种配置:四个可配置的执行框架(OpenHands、DeepSeek Harness、PI和openJiuwen)与五个模型在TUA-Bench、ALE-CLI和Terminal-Bench 4上的配对,以及原生的Codex-GPT和Claude Code-Claude配对。模型排名在不同执行框架间发生反转。在Terminal-Bench 4上,Claude在OpenHands中领先GPT 7.94分,但在PI中却落后30.16分。对于五个模型中的四个,最佳执行框架从一个基准到另一个基准发生变化,但有些配对保持稳定:openJiuwen在全部三个基准上给予Kimi最高分,领先幅度为5.61至11.11分。模型自身的供应商执行框架并不总是其最佳选择,更高的成本也不一定能带来更高的分数。在Terminal-Bench 4上,GPT在PI下的得分高于在DSH下,而每任务成本不到后者的四分之一。匹配的轨迹揭示了适配性差异的原因。模型几乎自行启动所有修复,因此很大程度上取决于执行框架是否以模型可用的形式将失败交还。GPT在PI的简洁脚手架下表现最佳,而经常发出格式错误工具调用的Kimi则在openJiuwen中表现最好。我们认为模型、执行框架和任务应一起评估,并在以下网址发布执行框架适配器、评估代码以及全部6,204条评分轨迹:此https URL和此https URL。

英文摘要

Choosing an agent system means choosing both a language model and the harness through which it acts. We ask whether a strong model, harness, or pairing stays strong when the setting changes. We evaluate 66 configurations: four configurable harnesses (OpenHands, DeepSeek Harness, PI, and openJiuwen) paired with five models on TUA-Bench, ALE-CLI, and Terminal-Bench 4, plus the native Codex-GPT and Claude Code-Claude pairings. Model rankings reverse across harnesses. On Terminal-Bench 4, Claude leads GPT by 7.94 points in OpenHands but trails it by 30.16 points in PI. For four of the five models, the best harness changes from one benchmark to another, yet some pairings hold: openJiuwen gives Kimi its highest score on all three benchmarks, by 5.61 to 11.11 points. A model's own vendor harness is not reliably its best, and higher cost does not reliably buy a higher score. On Terminal-Bench 4, GPT scores higher under PI than under DSH at less than a quarter of the cost per task. Matched trajectories suggest why fit varies. Models start almost all repairs themselves, so much depends on whether the harness hands failures back in a form the model can use. GPT does best with PI's lean scaffold, while Kimi, which often issues malformed tool calls, does best in openJiuwen. We argue that the model, the harness, and the task should be evaluated together, and we release the harness adapters, evaluation code, and all 6,204 scored trajectories at https://github.com/liyix/finding-the-right-fit and https://huggingface.co/datasets/yixuanli97/finding-the-right-fit.

Comments19 pages, 9 figures, 6 tables. Code: https://github.com/liyix/finding-the-right-fit. Data: https://huggingface.co/datasets/yixuanli97/finding-the-right-fit

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑