arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2025-11-13 至 2025-11-13 共收录 1 信号源:cs.CV, cs.AI, cs.LG

1. 文档图表理解 1 篇

2411.07722 2025-11-13 cs.AI 70%

Is Cognition Consistent with Perception? Assessing and Mitigating Multimodal Knowledge Conflicts in Document Understanding

Zirui Shao, Feiyu Gao, Zhaoqing Zhu, Chuwei Luo, Hangdi Xing, Zhi Yu, Qi Zheng, Ming Yan, Jiajun Bu

机构 * Zhejiang Key Laboratory of Accessible Perception and Intelligent Systems, Zhejiang University(浙江可感知智能系统重点实验室,浙江大学) Alibaba Group(阿里巴巴集团) Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and DataSecurity(杭州高新技术区(滨江)区块链与数据安全研究院)

专题命中 文档图表理解 :multimodal large language model(abstract);MLLM(abstract);分类 cs.AI

Comments Accepted at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏