V-ITI: Mitigating Hallucinations in Multimodal Large Language Models via Visual Inference-Time Intervention
V-ITI: 通过视觉推理时间干预缓解多模态大语言模型中的幻觉
机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) ; School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院) ; Baidu Inc.(百度公司)
专题命中 其他VLM :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI
AI总结 V-ITI通过视觉推理时间干预框架,有效缓解多模态大语言模型中的视觉相关幻觉问题,同时保持任务性能。