arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2026-02-11 至 2026-02-11 共收录 4 信号源:cs.CV, cs.AI, cs.LG

1. 幻觉与鲁棒性 4 篇

2602.09825 2026-02-11 cs.CV 79%

SAKED: Mitigating Hallucination in Large Vision-Language Models via Stability-Aware Knowledge Enhanced Decoding

SAKED: 通过稳定性感知知识增强解码缓解大视觉-语言模型中的幻觉

Zhaoxu Li, Chenqi Kong, Peijun Bao, Song Xia, Yi Tu, Yi Yu, Xinghao Jiang, Xudong Jiang

机构 * ROSE Lab, School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore(ROSE实验室,电子工程学院,南洋理工大学,新加坡) ROSE Lab, Interdisciplinary Graduate Programme, Nanyang Technological University, Singapore(ROSE实验室,跨学科研究生项目,南洋理工大学,新加坡) School of Physical and Mathematical Sciences, Nanyang Technological University, Singapore(物理与数学科学学院,南洋理工大学,新加坡) Shanghai Jiao Tong University, China(上海交通大学,中国)

专题命中 幻觉与鲁棒性 :vision-language model(title,abstract);分类 cs.CV

AI总结 SAKED通过引入稳定性感知知识增强解码方法,有效缓解大视觉-语言模型中的幻觉问题,无需训练即可集成至不同架构中,实现最佳性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06218 2026-02-11 cs.CV cs.LG 62%

Cross-Modal Redundancy and the Geometry of Vision-Language Embeddings

跨模态冗余与视觉-语言嵌入的几何学

Grégoire Dhimoïla, Thomas Fel, Victor Boutin, Agustin Picard

机构 * Brown University(布朗大学) ENS Paris Saclay(巴黎萨克雷大学) IRT Saint Exupéry(IRT圣埃克苏佩里) Kempner Institute, Harvard University(哈佛大学凯姆纳研究所)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV、cs.LG

AI总结 本文通过等能假设和对齐稀疏自编码器,揭示了视觉-语言模型中跨模态对齐的几何结构,发现稀疏双模态原子承载了跨模态对齐信号,单模态原子解释了模态差距,去除单模态原子可消除差距而不影响性能。

Comments Published as a conference paper at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16596 2026-02-11 cs.CV cs.AI 62%

SHIELD: Suppressing Hallucinations In LVLM Encoders via Bias and Vulnerability Defense

SHIELD:通过偏差和脆弱性防御抑制LVLM编码器中的幻觉

Yiyang Huang, Liang Shi, Yitian Zhang, Yi Xu, Yun Fu

机构 * Department of Electrical and Computer Engineering, Northeastern University(电气与计算机工程系,东北大学) Khoury College of Computer Science, Northeastern University(计算机科学学院,东北大学)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV、cs.AI

AI总结 SHIELD通过减少统计偏差、对抗固有偏差和解决脆弱性,有效抑制LVLM编码器中的对象幻觉。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09541 2026-02-11 cs.CV 57%

Scalpel: Fine-Grained Alignment of Attention Activation Manifolds via Mixture Gaussian Bridges to Mitigate Multimodal Hallucination

Scalpel: 通过混合高斯桥梁实现细粒度注意力激活流形对齐以缓解多模态幻觉

Ziqiang Shi, Rujie Liu, Shanshan Yu, Satoshi Munakata, Koichi Shirahata

机构 * Fujitsu Research & Development Center Co.,LTD.(Fujitsu 研究与开发中心有限公司) Fujitsu Limited(Fujitsu 有限公司)

专题命中 幻觉与鲁棒性 :vision-language model(abstract);分类 cs.CV

AI总结 Scalpel通过高斯混合模型和熵最优传输减少多模态幻觉,实现注意力激活流形的细粒度对齐,提升视觉-语言模型的输出一致性。

Comments WACV 2026 (It was accepted in the first round, with an acceptance rate of 6%.)

详情

展开后加载摘要…

URL PDF HTML 收藏