arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

截图还是工具?在混合GUI-MCP计算机使用智能体中引出工具使用并管理多模态上下文

Screenshots or Tools? Eliciting Tool Use and Managing Multimodal Context in Hybrid GUI-MCP Computer-Use Agents

Siqi Fan, Minghao Li, Xiaoqian Ma, Wenhui Tan, Xiusheng Huang, Juntong Wu, Liujie Zhang, Shuo Shang, Weihang Chen

arXiv 2608.03327首次发表:更新:

AI 中文总结

该研究针对混合GUI-MCP计算机使用智能体,发现工具使用存在采用缺口,通过调整工具奖励与上下文压缩优化,提升了智能体性能并降低输入成本。

AI 中文摘要

混合计算机使用智能体可通过截图或调用文本工具执行操作。研究发现,工具的存在并不决定其使用方式。在OSWorld-MCP基准(309个任务)的同一GUI-MCP框架下,相同的MCP工具使推理模型性能提升4.0个百分点,却使非推理模型性能下降5.9个百分点(各进行5次运行,均超出2倍标准差)。两者的区别在于工具决策行为:非推理策略会忽略、错误命名工具,或围绕工具错误终止;推理模型避免了这些错误,但仅在55/309个任务中调用工具,占可通过工具完成任务的23.9%,研究将这一缺口称为采用缺口。两类问题的共同原因是:模型已有更廉价的执行路径,且从未接受过使用该路径的训练。多回合强化学习探针证实了这一点:在动作层面,密集工具奖励使电子表格采用率从0.03提升至0.33,并延续到贪婪解码,但保留的准确率并未随之提升,行为可引导,能力却不行,瓶颈在于工具调用语义。在上下文层面,成功的工具调用往往使下一张截图冗余,丢弃该截图并将图像历史减半可减少约三分之一的输入token,且准确率损失很小;在相同观察规则下重新训练可消除该损失,压缩后的智能体达到37.8%的性能,而未压缩操作点为33.0%,输入成本仅为后者的53%,并在预注册的退化子集上将富-瘦差距缩小至零。当模型选择并整合工具时,工具会发挥作用,但当前混合智能体未充分利用此类选择。

英文摘要

Hybrid computer-use agents can act through screenshots or call text tools. We find that having a tool available does not settle which way the effect goes. Under one identical GUI-MCP harness on the OSWorld-MCP benchmark (309 tasks), the same MCP tools improve a reasoning model by +4.0pp and degrade a non-reasoning model by -5.9pp (5 runs each, both beyond 2 SE). What separates the two is tool-decision behavior. The non-reasoning policy ignores, misnames, or falsely terminates around tools. The reasoning model avoids these failures, yet still calls a tool on only 55/309 tasks, 23.9% of the tool-reachable ones. We call this shortfall the adoption gap. Both levels of the problem share one cause: the model already has a cheaper route and is never trained to take it. Multi-turn RL probes that cause. At the action level, a dense tool bonus raises spreadsheet adoption 0.03 -> 0.33 and carries into greedy decoding, but held-out accuracy does not follow. Behavior is steerable; competence is not. The bottleneck lies in tool-call semantics. At the context level, a successful tool call often makes the next screenshot redundant. Dropping it and halving image history cuts input tokens by about a third, at a small accuracy cost. Retraining under the same observation rule removes that cost. The compressed agent then reaches 37.8% against 33.0% for the uncompressed operating point, at 53% of the input cost, and closes the rich-lean gap on a pre-registered degraded subset to zero. Tools help when the model chooses and integrates them, and current hybrid agents leave many such choices unused. Code and checkpoints: https://github.com/redai-infra/hybrid-routing-agent

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑