发表机构
Shanghai Jiao Tong University; Southeast University; South China University of Technology; Microsoft Research Asia; Fudan University; East China Normal University(上海交通大学; 东南大学; 华南理工大学; 微软亚洲研究院; 复旦大学; 华东师范大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文推出推理驱动视觉编辑相关基准 RISEBench++ 及无需训练的智能体框架 RISE-Agent,评估58种视觉编辑方法后发现该领域仍存重大挑战,最强模型准确率仅56.6%。
AI 中文摘要
大型多模态模型(LMM)在视觉理解与生成领域已取得显著进展,但在视觉编辑任务中仍面临挑战,尤其体现在难以遵循复杂指令、保持外观一致性及支持灵活输入格式等方面。为研究该差距,本文推出首个用于评估推理驱动视觉编辑(Reasoning-Informed viSual Editing,RISE)的基准 RISEBench,并将其扩展为更全面、细粒度的基准 RISEBench++。RISEBench++ 将任务分类扩展为涵盖六个推理维度的层级体系:时间推理、因果推理、空间推理、逻辑推理、反事实推理,以及整合多轮编辑中多种推理类型的混合推理;这些维度进一步分解为12个子类别和65个细粒度任务类型。本文还将输入格式扩展为支持多图像条件输入,并将基准规模扩大至1000个人工标注测试用例,同时发布了英文和中文版本。此外,本文改进了评估框架,通过人工评审和 LMM 作为评审的方法,对指令遵循性、外观一致性和视觉合理性进行评估,以获得更可靠、校准的判断。除基准构建外,本文还推出 RISE-Agent,这是一个无需训练的智能体框架,整合了推理驱动规划、工具增强执行及验证器引导优化,在各类 RISE 任务中表现优于多数现有强基线方法。本文评估了58种视觉编辑方法,包括34个开源模型、19个闭源模型和5个智能体方法。结果显示,推理驱动视觉编辑仍存在重大挑战,即便评估中表现最强的 GPT-Image-2.5 Sunburst 也仅达到56.6%的准确率。RISEBench++ 揭示了当前编辑模型的局限性,提供了相关见解,并指明了推理感知视觉编辑的未来方向。
英文摘要
Large Multi-modality Models (LMMs) have made significant progress in visual understanding and generation, but still face challenges in visual editing, particularly in following complex instructions, preserving appearance consistency, and supporting flexible input formats. To study this gap, we introduce RISEBench, the first benchmark for evaluating Reasoning-Informed viSual Editing (RISE), and extend it to RISEBench++, a more comprehensive and fine-grained benchmark for this emerging task. RISEBench++ extends the taxonomy into a hierarchical scheme spanning six reasoning dimensions: Temporal, Causal, Spatial, Logical, and Counterfactual Reasoning, together with Hybrid Reasoning integrating multiple reasoning types across multi-turn edits. These dimensions are further decomposed into 12 subcategories and 65 fine-grained task types. We expand input formats to include multi-image conditioning and scale the benchmark to 1000 human-annotated test cases, released in English and Chinese. We also improve our evaluation framework, assessing Instruction Reasoning, Appearance Consistency, and Visual Plausibility with human judges and an LMM-as-a-judge approach for more reliable and calibrated judgements. Beyond benchmarking, we introduce RISE-Agent, a training-free agentic framework integrating reasoning-driven planning, tool-augmented execution, and verifier-guided refinement, outperforming most strong existing approaches across diverse RISE tasks. We evaluate 58 visual editing approaches, including 34 open-source models, 19 closed-source models, and 5 agentic methods. The results reveal substantial challenges in reasoning-based visual editing, with even the strongest evaluated approach, GPT-Image-2.5 Sunburst, achieving only 56.6% accuracy. RISEBench++ highlights the limitations of contemporary editing models, provides insights, and indicates future directions for reasoning-aware visual editing.