AI 中文总结
提出FISA框架,基于MLLM失败案例生成保留答案的增强图像,经自检验与双重保真过滤后,可提升视觉问答性能,且兼容文本自增强、数据效率更优。
AI 中文摘要
多模态大语言模型(MLLM)在视觉-语言任务中已取得显著性能,但其进展高度依赖大规模高质量多模态数据,这类数据标注成本高昂。自增强为模型扩展自身训练数据提供了无需外部监督的可行替代方案,但现有MLLM自增强方法大多以文本为中心,图像增强仍未得到充分探索,且通常依赖通用或手工设计的变换,与模型实际能力缺陷的关联性较弱。本文提出故障感知图像自增强(Failure-informed Image Self-Augmentation,FISA),这是一种用于MLLM自我改进的框架,它基于模型自身的失败案例构建增强图像。该方法生成具有视觉挑战性但保留答案的图像复杂变体,通过自检验验证其效用,并应用双重保真度过滤以避免语义失真。在视觉问答基准上的实验表明,所提方法在分布内和分布外设置中均能持续提升性能;进一步实验验证了FISA与现有文本自增强方法的兼容性、合成样本相较于通用图像增强基线的优越数据效率,以及所提过滤策略的实际有效性。
英文摘要
Multimodal large language models (MLLMs) have achieved remarkable performance across vision-language tasks, but their progress depends heavily on large-scale, high-quality multimodal data that are costly to annotate. Self-augmentation offers a promising alternative by enabling models to expand their own training data without external supervision. However, existing MLLM self-augmentation methods are largely text-centric, while image augmentation remains underexplored and typically relies on generic or handcrafted transformations that are weakly aligned with the model's actual incapability. We propose Failure-informed Image Self-Augmentation (\textbf{FISA}), a framework for MLLM self-improvement that constructs augmented images from the model's own failure cases. Our method generates visually challenging yet answer-preserving image complications, verifies their utility through self-examination, and applies dual fidelity filtering to avoid semantic distortion. Experiments on visual question answering benchmarks show that the proposed method consistently improves performance across both in-distribution and out-of-distribution settings. Further experiments validate the compatibility of FISA with existing textual self-augmentation approaches, the superior data efficiency of the synthesized samples over generic image augmentation baselines, and the practical effectiveness of the proposed filtering strategy.