用于移除大型视觉语言模型中虚假关联的感知分区式遗忘
Partition-Aware Unlearning for Removing Spurious Correlations in Large Vision-Language Models
- Wayne State University(韦恩州立大学)
- IIIT Delhi(德里印度信息技术学院)
- Institute for AI and Data Science (AIDaS), Wayne State University(韦恩州立大学人工智能与数据科学研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出PURGE框架,通过结构化数据集构建和感知分区式遗忘,减少大型视觉语言模型的虚假关联诱导错误,同时保持或提升整体性能。
AI中文摘要:
大型视觉语言模型(LVLMs)在众多多模态任务中表现出强劲性能;然而,它们常利用虚假的物体-背景关联,导致预测由上下文捷径而非与物体相关的视觉证据驱动。尽管人们对幻觉和鲁棒性评估的兴趣日益浓厚,现有基准难以控制模型预测是否基于目标物体或由相关背景线索诱导。本研究中,我们提出PURGE(Partition-aware Unlearning for Removing spurious-correlation Generated Errors,用于移除虚假关联生成错误的感知分区式遗忘),这是一个用于构建、基准测试和缓解LVLMs中虚假关联诱导故障的框架。该框架包含两部分:(1)结构化数据集构建,我们开发了三种互补的结构化数据构建策略,按与物体相关的证据和虚假背景线索对样本进行分区,以实现对捷径依赖的可控诊断;(2)感知分区式遗忘,利用这些分区选择性移除虚假的物体-背景关联,同时保留基于物体的推理。我们在多个LVLMs(包括LLaVA-1.6-7B、Qwen3-VL-8B-Instruct、Qwen3.5-9B,以及作为视觉语言编码器的CLIP)上,在多样化的基准套件(包括CHAIR、POPE、Causal-HalBench、MM-SpuBench、AMBER、MMHal和Waterbirds)上对PURGE框架进行评估。结果表明,PURGE在大多数评估设置中持续减少幻觉和虚假关联驱动的错误,同时保持或提升整体性能,为构建更可靠的LVLMs提供了可复用的评估协议和有效的缓解框架。
英文摘要:
Large Vision-Language Models (LVLMs) achieve strong performance across many multimodal tasks; however, they often exploit spurious object-background correlations, resulting in predictions driven by contextual shortcuts rather than object-relevant visual evidence. Despite growing interest in hallucination and robustness evaluation, existing benchmarks provide limited control over whether model predictions are grounded in the target object or induced by correlated background cues. In this work, we introduce PURGE (\underline{P}artition-aware \underline{U}nlearning for \underline{R}emoving spurious-correlation \underline{G}enerated \underline{E}rrors), a framework for constructing, benchmarking, and mitigating spurious-correlation-induced failures in LVLMs. The framework consists of: -- (1) Structured dataset construction wherein we develop three complementary structured data construction strategies that partition examples by object-relevant evidence and spurious background cues, enabling controlled diagnosis of shortcut reliance; and -- (2) Partition-aware unlearning, which uses these partitions to selectively remove spurious object-background associations while preserving object-based reasoning. We evaluate the \algo~framework across multiple LVLMs, including LLaVA-1.6-7B, Qwen3-VL-8B-Instruct, and Qwen3.5-9B, together with CLIP as a vision-language encoder, on a diverse suite of benchmarks, including CHAIR, POPE, Causal-HalBench, MM-SpuBench, AMBER, MMHal, and Waterbirds. Our results show that PURGE consistently reduces hallucinations and spurious-correlation-driven errors while maintaining or improving overall performance in most evaluated settings, providing both a reusable evaluation protocol and an effective mitigation framework for more reliable LVLMs.