OSWorld-Science:用于学习和使用科学软件的计算器使用代理基准
OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software
浏览论文内容
中文总结 AI 辅助
OSWorld-Science是一个针对科学软件的计算器使用代理基准,包含146个任务,评估VLM在分子绘图、病理分析等领域的表现,发现当前最先进模型仍面临挑战。
中文摘要 AI 辅助
科学软件为基于视觉语言模型(VLM)的计算器使用代理提出了一个高要求的测试:完成研究工作流程需要解释专业界面、操作科学对象并产生可验证的结果。因此,我们引入了OSWorld-Science,一个结合了科学上有意义的任务、基于工件的评估和高效代理框架的基准和评估环境,用于研究科学领域的计算器使用。该基准包含12个VLM和跨多个科学领域和软件配置的146个高质量任务,涵盖分子绘图和逆合成、病理图像分析、统计计算和物理模拟等工作流程。任务通过专家提议和迭代的人机协同设计开发,选择依据科学价值和难度。特定任务的基于执行的评估器检查应用状态和生成的工件,包括分子结构、分割掩码、图表和数值结果,并为不完整的结果授予部分分数。我们的特殊框架集成了模型适配器、交互循环控制和轨迹记录,以支持模型和交互策略的比较。我们的结果表明,当前最先进的VLM配备强大的框架在解决科学领域的关键问题方面仍面临挑战。我们还跨多语言、推理努力、上下文长度和其他因素分析了基准测试结果,并得出了几个重要的结论和方向,以协助未来的发展。总体而言,我们提供了一个集成框架,将专家定义的科学目标与可验证的软件结果联系起来,使得在科学工作流程中系统评估代理能力和框架设计成为可能。
英文摘要
Scientific software presents a demanding test for computer-using agents based on visual language models (VLMs): completing a research workflow requires interpreting specialized interfaces, manipulating scientific objects, and producing verifiable results. We thus introduce OSWorld-Science, a benchmark and evaluation environment that combines scientifically meaningful tasks, artifact-based evaluation, and an efficient agent harness for studying computer use in the scientific domain. The benchmark contains 12 VLMs and 146 high-quality tasks across several scientific domains and software configurations, covering workflows such as molecular drawing and retrosynthesis, pathology image analysis, statistical computing, and physical simulation. Tasks are developed through expert proposals and iterative human--AI co-design, with selection guided by scientific value and difficulty. Task-specific execution-based evaluators inspect application states and generated artifacts, including molecular structures, segmentation masks, plots, and numerical results, and award partial credit for incomplete outcomes. Our special harness integrates model adapters, interaction-loop control, and trajectory logging to support comparisons of models and interaction strategies. Our results show that current state-of-the-art VLMs with a strong harness still face challenges in addressing key questions in the scientific domains. We also analyze the benchmarking results across multi-linguistics, reasoning efforts, context length and other factors and derive several important conclusions and directions to assist future development. Overall, we provide an integrated framework connecting expert-defined scientific goals to verifiable software outcomes, enabling systematic evaluation of both agent capabilities and harness design in scientific workflows.
发表机构
- Tsinghua University(清华大学)
- University of California, Los Angeles(加利福尼亚大学洛杉矶分校)
- RIKEN AIP(理化学研究所人工智能中心)
- University of Waterloo(滑铁卢大学)
- Carnegie Mellon University(卡内基梅隆大学)
- Yale University(耶鲁大学)
- Northwestern University(西北大学)
- Zhejiang University(浙江大学)
- University of California, Berkeley(加利福尼亚大学伯克利分校)
- University of Illinois Chicago(伊利诺伊大学芝加哥分校)
- Boston University(波士顿大学)
- The University of Hong Kong(香港大学)
- Stanford University(斯坦福大学)
- New York University(纽约大学)
- The University of Tokyo(东京大学)
机构由 AI 辅助整理,请以论文原文为准。