发表机构
CosmosMind; Peking University; Tsinghua University; HKUST; ModCraft; Renmin University of China(CosmosMind; 北京大学; 清华大学; 香港科技大学; ModCraft; 中国人民大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出 SWE-PolyVision 基准,含 92 个真实仓库任务和 402 张图像,评估 11 个编码模型在三种视觉访问模式下跨图像溯因推理能力,揭示视觉证据可用性与成功修复之间的差距。
AI 中文摘要
当前的多模态软件工程基准将图像作为附加上下文提供,但并未测试智能体能否将分布在多张图像中的证据整合为经过验证的仓库级修复。我们提出了 SWE-PolyVision,这是一个可执行的基准,包含来自 36 个开源组织的 92 个真实任务,其中 48 个为公开任务,44 个为私有保留任务。该发布包含 402 张静态图像和 6 个视频,每个任务至少有两个视觉输入。每个任务将一个固定的修复前仓库与一个隔离的验证器配对,并在三种访问模式(仅文本、原生视觉和工具介导视觉)中受支持的条件进行评估。在 11 个编码模型上,视觉访问改变了可解决的任务,但效果取决于模型和任务。两个可追溯的原生视觉案例展示了互补的视觉和文本线索如何导致源代码定位的、经过验证的修复;受控干预表明这种转换在输入之间尚不稳定。因此,SWE-PolyVision 将多图像证据的可用性与其在仓库级修复中的成功使用区分开来,而不将补丁成功单独视为显式推理的证据。
英文摘要
Current multimodal software-engineering benchmarks expose images as additional context, but do not test whether an agent can integrate evidence distributed across images into a verified repository-level repair. We present SWE-PolyVision, an executable benchmark of 92 real tasks from 36 open-source organizations, with 48 public tasks and 44 private holdouts. The release contains 402 static images and 6 videos, with at least two visual inputs per task. Each task pairs a fixed pre-fix repository with an isolated verifier and is evaluated under the supported conditions among three access modes: Text-only, Native Vision, and Tool-mediated Vision. Across eleven coding models, visual access changes which tasks are solved, but effects depend on both model and task. Two trace-linked Native Vision cases illustrate how complementary visual and textual clues can lead to source-localized, verified repairs; controlled interventions show that this conversion is not yet stable across inputs. SWE-PolyVision thus separates the availability of multi-image evidence from its successful use in repository-level repair, without treating patch success alone as proof of explicit reasoning.