arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2026-02-23 至 2026-02-23 共收录 3 信号源:cs.CV, cs.AI, cs.LG

1. 其他VLM 3 篇

2602.17869 2026-02-23 cs.CV 70%

Learning Compact Video Representations for Efficient Long-form Video Understanding in Large Multimodal Models

学习紧凑的视频表示以在大规模多模态模型中实现高效的长视频理解

Yuxiao Chen, Jue Wang, Zhikang Zhang, Jingru Yi, Xu Zhang, Yang Zou, Zhaowei Cai, Jianbo Yuan, Xinyu Li, Hao Yang, Davide Modolo

机构 * Amazon AGI(亚马逊人工智能研究院)

专题命中 其他VLM :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV

AI总结 本文提出了一种端到端的长视频理解框架,结合自适应视频采样器和空间时间视频压缩器,以高效处理长视频的冗余问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18262 2026-02-23 cs.CL cs.AI cs.LG 62%

Simplifying Outcomes of Language Model Component Analyses with ELIA

用ELIA简化语言模型组件分析的结果

Aaron Louis Eidt, Nils Feldhus

机构 * Technische Universität Berlin(柏林技术大学) Fraunhofer Heinrich Hertz Institute(弗劳恩霍夫海因里希·赫茨研究所) BIFOLD – Berlin Institute for the Foundations of Learning and Data(柏林学习与数据基础研究所)

专题命中 其他VLM :vision-language model(abstract);分类 cs.AI、cs.LG

AI总结 ELIA通过交互式界面和AI生成的自然语言解释,简化了语言模型组件分析,帮助非专家理解复杂模型。

Comments EACL 2026 System Demonstrations. GitHub: https://github.com/aaron0eidt/ELIA

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04587 2026-02-23 cs.CL cs.AI cs.CY 57%

VILLAIN at AVerImaTeC: Verifying Image-Text Claims via Multi-Agent Collaboration

VILLAIN 在 AVerImaTeC 上:通过多智能体协作验证图像-文本声明

Jaeyoon Jung, Yejun Yoon, Kunwoo Park

机构 * School of AI Convergence, Soongsil University(人工智能融合学院,顺世大学) MAUM AI Inc.(MAUM人工智能公司) Department of Intelligent Semiconductors, Soongsil University(智能半导体系,顺世大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.AI

AI总结 VILLAIN 通过多智能体协作验证图像-文本声明,实现事实核查系统的高效准确验证。

Comments A system description paper for the AVerImaTeC shared task at the Ninth FEVER Workshop (co-located with EACL 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏