arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

不同的变化需要不同的推理:用于鲁棒变化描述的变化类型专用专家

Different Changes Require Different Reasoning: Change-Type-Specialized Experts for Robust Change Captioning

Jiyoung Park, InJae Oh, Jung Uk Kim

arXiv 2609.01136首次发表:更新:

发表机构

Kyung Hee University(庆熙大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有变化描述方法忽略不同变化类型差异的问题,提出MEDIC框架,通过类型专用记忆专家提升描述精度,在多数据集上性能优于现有方法。

AI 中文摘要

变化描述是指生成自然语言描述以解释一对图像之间变化的任务。尽管不同的变化类型(例如颜色变化、物体添加)具有不同的视觉线索,且需要专门的推理过程,但现有方法往往忽略了这些区别。为解决这一局限,我们提出了图像变化多专家诊断框架(MEDIC),该框架通过显式建模变化类别引入变化类型感知。MEDIC采用类型专用记忆专家,可根据输入动态检索与类型相关的视觉模式,此设计使每个专家能捕获其变化类型内的多样变化,同时聚焦于最具信息性的视觉线索。MEDIC通过将输入软路由至各类型专用专家,并为每个变化类别学习专用表示,从而生成更精确且具类型感知的变化描述。大量实验表明,所提出的MEDIC在多样且具有挑战性的数据集上始终优于现有方法,代码可在GitHub获取。

英文摘要

Change captioning is the task of generating natural language descriptions that explain the changes between a pair of images. Although different change types (e.g., color shifts, object additions) exhibit distinct visual cues and require specialized reasoning processes, existing methods often overlook these distinctions. To address this limitation, we propose Multi-Expert Diagnosis for Image Change (MEDIC), a novel framework that introduces change-type awareness by explicitly modeling change categories. MEDIC employs type-specialized memory experts that dynamically retrieve type-relevant visual patterns conditioned on the input. This design enables each expert to capture diverse variations within its change type while focusing on the most informative visual cues. By softly routing inputs across type-specialized experts and learning dedicated representations for each change category, MEDIC generates more precise and type-aware change descriptions. Extensive experiments demonstrate that the proposed MEDIC consistently outperforms existing methods across diverse and challenging datasets. The code is available at \href{https://github.com/VisualAIKHU/MEDIC}{GitHub}.

CommentsAccepted to ECCV 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑