Disentangling Hallucinations: Orthogonal Semantic Projection for Robust Interpretability
解构幻觉:正交语义投影实现鲁棒可解释性
机构 * U2IS, ENSTA, Institut Polytechnique de Paris(巴黎综合理工学院ENSTA学院U2IS实验室) ; ISIR, Université Sorbonne, Pierre et Marie Curie(索邦大学皮埃尔和玛丽·居里分校ISIR实验室) ; AMIAD, Pôle Recherche(AMIAD研究部)
专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI
AI总结 针对视觉语言模型解释中的语义幻觉问题,提出线性语义归因(LSA)理论框架,并引入正交语义投影(OSP)方法,通过正交化查询向量消除共享特征干扰,最小化幻觉。
Comments 41 pages in total. 5 figures, and 2 tables in the main paper; 10 figures and 17 tables in the appendix