arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.15205cs.SEcs.AI

MM-IssueLoc:用于评估多模态仓库级问题定位中视觉证据的可控基准

MM-IssueLoc: A Controlled Benchmark for Evaluating Visual Evidence in Multimodal Repository-Level Issue Localization

Shaoxiong Zhan, Shi Hu, Boyu Feng, Hai Lin, Andrew Gong, Zhengda Zhou, Jiaying Zhou, Yunyun Hou, Hao Su, Hai-Tao Zheng

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对仓库级问题定位多为文本任务,视觉证据作用不明的情况,引入MM-IssueLoc基准和评估协议,含多语言实例与标注等。评估LLM和检索系统,发现现有系统距可靠多模态定位有差距,该基准让视觉证据成评估变量,助于后续研究。

中文摘要 AI 辅助

实际仓库问题通常包含截图、错误对话框等视觉证据,但仓库级问题定位大多仅作为文本任务评估。现有多模态SE基准评估端到端修复,混淆了定位与补丁合成,无法明确视觉输入的作用。我们引入MM-IssueLoc,一个用于带视觉证据的仓库级定位的可控基准和评估协议。它包含多种语言的实例及相关标注,提供不同层级的金标准标签、配对评估和基于VCE的诊断。评估结果显示现有系统离可靠的多模态仓库定位仍有差距,MM-IssueLoc使视觉证据成为明确评估变量,利于后续研究。

英文摘要

Real repository issues routinely include visual evidence such as screenshots, error dialogs, rendered UI states, and logs, yet repository-level issue localization is evaluated mostly as a text-only task. Existing multimodal SE benchmarks evaluate end-to-end repair, entangling localization with patch synthesis and obscuring whether visual input helped, hurt, or was ignored. We introduce \textbf{MM-IssueLoc}, a controlled benchmark and evaluation protocol for repository-level localization with visual evidence. MM-IssueLoc contains 652 issue-PR instances across 23 languages, with annotations for 7 image categories and 4 relevance levels. It provides file-level and function-level gold labels, paired text-only and with-image evaluation, and VCE-based diagnostics that convert images into structured textual evidence. We evaluate LLM-based and retrieval-based systems, including MM-IssueLoc-VL-Emb as a controlled multimodal retriever. Results show that existing systems remain far from reliable multimodal repository localization: the strongest agent reaches 38.96 file Acc@5 and 22.45 function Acc@10, while the strongest retriever reaches 33.86 function Acc@10. Cross-benchmark comparisons show that high localization scores on text-dominant SWE benchmarks do not transfer cleanly to multimodal issue localization. MM-IssueLoc turns visual evidence into an explicit evaluation variable, enabling future work to test whether systems improve by using visual evidence for localization, rather than by relying on text-only cues or downstream patch-generation effects.

发表机构

  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

↑