OSWorld-Human: Benchmarking the Efficiency of Computer-Use Agents
OSWorld-Human: 评估计算机使用代理的效率基准
机构 * OpenAI ; Anthropic ; Google DeepMind(谷歌DeepMind) ; ByteDance(字节跳动) ; Agent S2 ; GTA1 ; Lei ; Jedi
AI总结 本文研究了计算机使用代理在OSWorld基准上的时间性能,发现大模型调用导致高延迟,并构建了包含人类轨迹的OSWorld Human数据集,评估发现最佳代理仍需更多步骤。