发表机构
Hanyang University(汉阳大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对本地计算机使用智能体的推理时扩展开展系统实证研究,评估多款模型在OSWorld基准上的表现,揭示其边际收益递减、失败模式转变等规律,提出高效本地智能体的设计方向。
AI 中文摘要
在本地部署自主计算机使用智能体(Computer-Use Agents, CUAs)对于隐私、成本效率和实际可用性愈发重要,但在严格硬件约束下提升其性能仍具挑战性。尽管近期研究表明,推理时扩展可通过执行过程中的额外计算提升前沿计算机使用智能体的性能,但其对资源受限的本地模型的有效性却鲜为人知。本文针对本地CUAs在上下文、时间、结构和并行维度上的推理时扩展开展系统实证研究,在OSWorld基准上评估Qwen3-VL-8B/30B-A3B、UI-TARS-1.5-7B和OpenCUA-7B。结果显示,额外计算常产生边际收益递减,同时改变失败模式:上下文扩展提供历史依据以提升轨迹稳定性和任务准确性,但其收益随token成本增加而饱和,失败类型从重复或停滞轨迹转向过早假成功;时间扩展类似地减少最大步数停滞,但未显著提升任务成功率,表明更长的时间范围常延长错误轨迹而非纠正;结构分解会在本地两阶段智能体中引入规划和格式开销,而并行扩展以大量计算成本部分缓解这些失败。总体而言,研究结果表明,高效的本地CUAs需要选择性计算分配、感知失败的控制机制,以及围绕本地模型能力和局限性设计的智能体框架。
英文摘要
Deploying autonomous computer-use agents (CUAs) locally is increasingly important for privacy, cost efficiency, and practical usability, yet improving their performance under strict hardware constraints remains challenging. While recent studies show that inference-time scaling can improve frontier computer-use agents through additional computation during execution, its effectiveness for resource-constrained local models remains poorly understood. We present a systematic empirical study of inference-time scaling in local CUAs across contextual, temporal, structural, and parallel dimensions. We evaluate Qwen3-VL-8B/30B-A3B, UI-TARS-1.5-7B, and OpenCUA-7B on the OSWorld benchmark. Our results show that additional computation often yields diminishing returns while changing failure modes. Contextual scaling provides historical grounding that improves trajectory stability and task accuracy, but its gains saturate as token cost increases and failures shift from repetitive or stalled trajectories toward premature false successes. Temporal scaling similarly reduces max-step stalls, yet does not substantially improve task success, indicating that longer horizons often extend erroneous trajectories rather than correct them. We further find that structural decomposition can introduce planning and formatting overhead in local two-stage agents, while parallel scaling partially mitigates these failures at a substantial computational cost. Overall, our findings suggest that efficient local CUAs require selective compute allocation, failure-aware control mechanisms, and agentic frameworks designed around the capabilities and limitations of local models.