arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.14075cs.SE

VisualRepair:用于视觉软件问题修复的动态工具调用和区域聚焦

VisualRepair: Dynamic Tool Calling and Region Focusing for Visual Software Issue Repair

Jingyu Xiao, Zhongyi Zhang, Haoran Hou, Yuxuan Wan, Yuan Jiang, Yintong Huo, Michael R. Lyu

首次发表
浏览论文内容

中文总结 AI 辅助

针对现代软件视觉信息利用难题,提出VisualRepair框架,通过图像类型感知工具调用和动态测试时区域聚焦两模块,在SWE-bench基准测试中性能超现有方法,有效提升自动视觉软件问题修复能力。

中文摘要 AI 辅助

随着大语言模型的出现,自动程序修复取得了显著进展。然而,现代软件系统中丰富的图形用户界面使得有效利用错误截图中的视觉信息对于多模态场景下理解错误和生成准确修复至关重要。实际问题报告包含多种视觉附件,且错误截图有大量无关区域。为此提出VisualRepair框架,包含图像类型感知工具调用和动态测试时区域聚焦两个核心模块。在SWE-bench多模态基准测试上的实验表明,VisualRepair性能优于现有方法,突出了类型感知视觉理解和区域聚焦定位在自动视觉软件问题修复中的有效性。

英文摘要

Automated Program Repair (APR) has witnessed significant progress with the advent of Large Language Models (LLMs). However, as modern software systems increasingly expose rich graphical user interfaces, effectively leveraging visual information from bug screenshots has become essential for understanding bugs and generating accurate fixes in multimodal scenarios. Real-world issue reports frequently contain heterogeneous visual attachments including UI screenshots, IDE snapshots, GIFs, and text-centric images, each with distinct visual patterns and domain-specific semantics that impose substantial perceptual demands on MLLMs. Furthermore, bug screenshots often contain large expanses of uninformative and bug-irrelevant regions, distracting the model's attention and limiting patch diversity. To address these challenges, we propose VisualRepair, an MLLM-based framework for visual software issue repair comprising two core modules: Image Type-aware Tool Calling (ITTC), which classifies input images and dynamically invokes a tailored tool-calling chain for robust visual interpretation, and Dynamic Test-time Region Focusing (DTRF), which grounds multiple bug-related region candidates and refines them via an adaptive zoom-in and zoom-out strategy to improve fault localization and promote diverse patch generation. Extensive experiments on the SWE-bench Multimodal benchmark demonstrate that VisualRepair consistently outperforms state-of-the-art approaches. VisualRepair resolves 196 and 25 instances on the test and dev sets, respectively, surpassing the best baseline by 10 and 11 instances. These results highlight the effectiveness of type-aware visual understanding and region-focused localization for automated visual software issue repair.

补充信息

↑