Mitigating Object and Action Hallucinations in Multimodal LLMs via Self-Augmented Contrastive Alignment
通过自增强对比对齐缓解多模态大语言模型中的对象和动作幻觉
机构 * Graduate Institute of Communication Engineering, National Taiwan University(国家交通大学通信工程研究所) ; NVIDIA
专题命中 GUI与屏幕智能体 :MLLM(abstract);分类 cs.CV、cs.AI、cs.LG
AI总结 SANTA框架通过自增强对比对齐方法,有效缓解多模态大语言模型中的对象和动作幻觉问题。
Comments IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2026. Project page: https://kpc0810.github.io/santa/