arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-16 至 2025-09-16 共收录 10 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 10 篇

2501.15688 2025-09-16 cs.CL cs.AI cs.LG 84%

Transformer-Based Multimodal Knowledge Graph Completion with Link-Aware Contexts

Haodi Ma, Dzmitry Kasinets, Daisy Zhe Wang

机构 * Department of Computer and Information Science and Engineering, University of Florida(计算机与信息科学与工程系,佛罗里达大学)

专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00284 2025-09-16 cs.RO cs.AI 79%

LightEMMA: Lightweight End-to-End Multimodal Model for Autonomous Driving

Zhijie Qiao, Haowei Li, Zhong Cao, Henry X. Liu

机构 * Department of Civil and Environmental Engineering, University of Michigan(土木与环境工程系,密歇根大学) University of Michigan Transportation Research Institute(密歇根大学交通研究所)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11247 2025-09-16 cs.CV 74%

Contextualized Multimodal Lifelong Person Re-Identification in Hybrid Clothing States

Robert Long, Rongxin Jiang, Mingrui Yan

机构 * University of Padua(帕多瓦大学) Heilongjiang University of Science and Technology(黑龙江科技大学)

专题命中 图文多模态 :multimodal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01064 2025-09-16 cs.CV cs.AI 73%

Fighting Fire with Fire (F3): A Training-free and Efficient Visual Adversarial Example Purification Method in LVLMs

Yudong Zhang, Ruobing Xie, Yiqing Huang, Jiansheng Chen, Xingwu Sun, Zhanhui Kang, Di Wang, Yu Wang

机构 * Tsinghua University, Tencent(清华大学,腾讯) Tencent(腾讯) University of Science and Technology Beijing(北京科技大学) Tencent, University of Macau(腾讯,澳门大学) Tsinghua University(清华大学)

专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted by ACM Multimedia 2025 BNI track (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.16915 2025-09-16 cs.CV cs.LG 70%

Multilingual Diversity Improves Vision-Language Representations

Thao Nguyen, Matthew Wallingford, Sebastin Santy, Wei-Chiu Ma, Sewoong Oh, Ludwig Schmidt, Pang Wei Koh, Ranjay Krishna

机构 * University of Washington(华盛顿大学) Allen Institute for Artificial Intelligence(人工智能研究院)

专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV

Comments NeurIPS 2024 Spotlight paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16146 2025-09-16 cs.CV cs.AI cs.CL cs.LG 67%

Steering LVLMs via Sparse Autoencoder for Hallucination Mitigation

Zhenglin Hua, Jinghan He, Zijun Yao, Tianxu Han, Haiyun Guo, Yuheng Jia, Junfeng Fang

机构 * School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院) Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University)(东南大学新一代人工智能技术及其交叉应用关键实验室) Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所基础模型研究中心) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系) Wuhan University of Technology(武汉理工大学) National University of Singapore(新加坡国立大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to Findings of EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11895 2025-09-16 cs.CV cs.AI 62%

Integrating Prior Observations for Incremental 3D Scene Graph Prediction

Marian Renz, Felix Igelbrink, Martin Atzmueller

机构 * DFKI Niedersachsen(德克萨斯联合研究所(北莱茵威斯特法伦)) Cooperative and Autonomous Systems, DFKI Niedersachsen(合作与自主系统,DFKI北莱茵威斯特法伦) German Research Center for Artificial Intelligence(德国人工智能研究中心) Semantic Information Systems, Osnabrück University(语义信息系统,奥斯纳布吕克大学)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted at 24th International Conference on Machine Learning and Applications (ICMLA'25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11961 2025-09-16 cs.CL 57%

Spec-LLaVA: Accelerating Vision-Language Models with Dynamic Tree-Based Speculative Decoding

Mingxiao Huo, Jiayi Zhang, Hewei Wang, Jinfeng Xu, Zheyu Chen, Huilin Tai, Yijun Chen

机构 * Carnegie Mellon University(卡内基梅隆大学) University of Nottingham(诺丁汉大学) The University of Hong Kong(香港大学) The Hong Kong Polytechnic University(香港理工大学) Columbia University(哥伦比亚大学) University of California, Berkeley(加州大学伯克利分校)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CL

Comments 7pages, accepted by ICML TTODLer-FM workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05479 2025-09-16 cs.CV 57%

LATTE: Learning to Think with Vision Specialists

Zixian Ma, Jianguo Zhang, Zhiwei Liu, Jieyu Zhang, Juntao Tan, Manli Shu, Juan Carlos Niebles, Shelby Heinecke, Huan Wang, Caiming Xiong, Ranjay Krishna, Silvio Savarese

机构 * University of Washington(华盛顿大学) Salesforce Research(Salesforce研究)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

Journal ref EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11065 2025-09-16 cs.SE cs.PL 50%

ViScratch: Using Large Language Models and Gameplay Videos for Automated Feedback in Scratch

Yuan Si, Daming Li, Hanyuan Shi, Jialu Zhang

专题命中 图文多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏