arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06075cs.CVcs.AI

面向智能体图像编辑的领域 grounded 候选选择:以阴影去除为例

Domain-Grounded Candidate Selection for Agentic Image Editing: A Shadow Removal Case

Shilin Hu, Jingyi Xu, Dimitris Samaras, Hieu Le

首次发表
浏览论文内容

中文总结 AI 辅助

该研究以阴影去除为例,提出领域 grounded 的智能体候选选择流程,结合物理原理约束商用视觉-语言模型的生成,在 ShadowRemovalRefine 基准上使 CDD 降低至少 47%,证明经典低级视觉先验仍具实用价值。

中文摘要 AI 辅助

商用视觉-语言模型正在重塑计算机视觉,其视觉先验的广度足以与特定任务系统相媲美,这引发了一个自然的问题:它们是否降低了对经典的、基于物理的低级视觉的需求?我们以阴影去除这一问题展开研究,该问题受场景几何、光照、材料和遮挡物影响,而配对的阴影与无阴影数据难以大规模收集。我们发现,直接使用商用生成式编辑器可生成保留表面纹理和局部外观的干净无阴影编辑结果,但这会带来新的失败模式:同一编辑器可能会再生场景内容、生成幻觉对象,或将阴影误判为材料或几何结构,从而产生看似合理但物理错误的编辑结果。我们通过智能体候选选择流程解决这一问题:编辑器生成引导探针,评估器筛选主要失败项,必要时重试,采样多个候选结果,对其进行过滤,并选择在阴影去除与场景保留之间取得平衡的最终结果。将该流程基于阴影形成的物理原理进行 grounded 处理,可使其更可靠:提示生成器和评估器将阴影视为由光线遮挡引起的光照效果,而非材料或对象结构,这可显著提升质量和一致性。在 ShadowRemovalRefine 基准上,我们的面向物理的流程取得了 0.0075 的 CDD 值,与最强的现有方法相比,CDD 至少降低了 47%。这些结果表明,商用视觉-语言模型并未取代经典的低级视觉先验;相反,此类先验对于约束和引导物理上欠约束的生成仍具实用性。

英文摘要

Commercial vision-language models are reshaping computer vision, with visual priors broad enough to rival task-specific systems. This raises a natural question: do they reduce the need for classic, physics-informed low-level vision? We study this through shadow removal, a problem shaped by scene geometry, illumination, materials, and occluders, where paired shadow and shadow-free data are hard to collect at scale. We find that a commercial generative editor, used directly, can produce clean shadow-free edits that preserve surface texture and local appearance. However, this comes with a new failure mode: the same editor can regenerate scene content, hallucinate objects, or misread a shadow as material or geometry, producing plausible but physically wrong edits. We address this with an agentic candidate-selection pipeline: the editor generates a guided probe, an evaluator screens for major failures, retries when needed, samples multiple candidates, filters them, and selects a final result balancing shadow removal against scene preservation. Grounding this process in shadow-formation physics makes it more reliable: prompting the generator and evaluator to treat shadows as illumination effects caused by light occlusion, not material or object structure, measurably improves quality and consistency. On the ShadowRemovalRefine benchmark, our physics-oriented pipeline achieves a CDD of 0.0075, reducing CDD by at least 47% over the strongest prior method. These results suggest that commercial vision-language models do not replace classic low-level vision priors; instead, such priors remain useful for constraining and steering physically underconstrained generation.

发表机构

  • Stony Brook University(石溪大学)
  • UNC Charlotte(北卡罗来纳大学夏洛特分校)

机构由 AI 辅助整理,请以论文原文为准。

↑