arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-19 至 2025-09-19 共收录 48 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 9 篇

2509.15160 2025-09-19 cs.HC cs.CL cs.GR 57%

An Evaluation-Centric Paradigm for Scientific Visualization Agents

Kuangshi Ai, Haichao Miao, Zhimin Li, Chaoli Wang, Shusen Liu

机构 * Univ. Notre Dame(内布拉斯加大学) LLNL(劳伦斯利弗莫尔国家实验室) Vanderbilt Univ.(范德比大学)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CL

Journal ref 1st Workshop on GenAI, Agents, and the Future of VIS (IEEE VIS Conference 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14171 2025-09-19 cs.CL 57%

AssoCiAm: A Benchmark for Evaluating Association Thinking while Circumventing Ambiguity

Yifan Liu, Wenkuan Zhao, Shanshan Zhong, Jinghui Qin, Mingfu Liang, Zhongzhan Huang, Wushao Wen

机构 * Sun Yat-sen University(中山大学) Guangdong University of Technology(广东工业大学) Northwestern University(西北大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL

Comments Accepted by EMNLP 2025 main track

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22159 2025-09-19 cs.RO cs.CV 57%

ForceVLA: Enhancing VLA Models with a Force-aware MoE for Contact-rich Manipulation

Jiawen Yu, Hairuo Liu, Qiaojun Yu, Jieji Ren, Ce Hao, Haitong Ding, Guangyu Huang, Guofan Huang, Yan Song, Panpan Cai, Cewu Lu, Wenqiang Zhang

机构 * Fudan University(复旦大学) Shanghai Jiao Tong University(上海交通大学) Shanghai Innovation Institute(上海创新研究院) Shanghai AI Lab(上海人工智能实验室) National University of Singapore(新加坡国立大学) Shanghai University(上海大学) Xi’an Jiaotong University(西安交通大学) Noematrix Intelligence(Noematrix智能科技)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.13835 2025-09-19 cs.CL 57%

RAcQUEt: Unveiling the Dangers of Overlooked Referential Ambiguity in Visual LLMs

Alberto Testoni, Barbara Plank, Raquel Fernández

机构 * Amsterdam UMC, Department of Medical Informatics(阿姆斯特丹大学医学中心,医学信息学系) Center for Information and Language Processing, LMU Munich(信息与语言处理中心,慕尼黑大学) Munich Center for Machine Learning (MCML), Munich(慕尼黑机器学习中心(MCML)) Institute for Logic, Language and Computation (ILLC), University of Amsterdam(逻辑、语言与计算研究所(ILLC),阿姆斯特丹大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL

Comments Accepted to EMNLP 2025 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14446 2025-09-19 q-bio.NC 50%

Mouse vs. AI: A Neuroethological Benchmark for Visual Robustness and Neural Alignment

Marius Schneider, Joe Canzano, Jing Peng, Yuchen Hou, Spencer LaVere Smith, Michael Beyeler

专题命中 多模态评测 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态Agent 1 篇

2509.14480 2025-09-19 cs.CL cs.AI cs.MA 81%

Process-Supervised Reinforcement Learning for Interactive Multimodal Tool-Use Agents

Weiting Tan, Xinghua Qu, Ming Tu, Meng Ge, Andy T. Liu, Philipp Koehn, Lu Lu

机构 * Johns Hopkins University(约翰霍普金斯大学)

专题命中 多模态Agent :multimodal(title);multi-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态训练与对齐 6 篇

2509.14735 2025-09-19 cs.CL 88%

Decoupled Proxy Alignment: Mitigating Language Prior Conflict for Multimodal Alignment in MLLM

Chenkun Tan, Pengyu Wang, Shaojun Zhou, Botian Jiang, Zhaowei Li, Dong Zhang, Xinghao Wang, Yaqian Zhou, Xipeng Qiu

机构 * Fudan University(复旦大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(title,abstract);分类 cs.CL

Comments Accepted by Findings of EMNLP2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14383 2025-09-19 cs.RO cs.CV 83%

RLBind: Adversarial-Invariant Cross-Modal Alignment for Unified Robust Embeddings

Yuhong Lu

机构 * Samueli School of Engineering, Electrical and Computer Engineering, UCLA(UCLA电气与计算机工程学院萨姆利学校)

专题命中 多模态训练与对齐 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments This paper is submitted to IEEE International Conference on Robotics and Automation (ICRA) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18042 2025-09-19 cs.CV cs.AI 81%

VLM-E2E: Enhancing End-to-End Autonomous Driving with Multimodal Driver Attention Fusion

Pei Liu, Haipeng Liu, Haichao Liu, Xin Liu, Jinxin Ni, Jun Ma

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Li Auto Inc.(Li汽车公司) the School of Aeronautics and Astronautics, Xiamen University(厦门大学航空航天学院) The Hong Kong University of Science and Technology(香港科技大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15096 2025-09-19 cs.CV 79%

OmniSegmentor: A Flexible Multi-Modal Learning Framework for Semantic Segmentation

Bo-Wen Yin, Jiao-Long Cao, Xuying Zhang, Yuming Chen, Ming-Ming Cheng, Qibin Hou

机构 * VCIP, CS, Nankai University(视觉计算研究所,计算机科学,南开大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14739 2025-09-19 cs.CV 57%

FMGS-Avatar: Mesh-Guided 2D Gaussian Splatting with Foundation Model Priors for 3D Monocular Avatar Reconstruction

Jinlong Fan, Bingyu Hu, Xingguang Li, Yuxiang Yang, Jing Zhang

机构 * HangZhou Dianzi University(杭州电子大学) Shenzhen Polytechnic University(深圳职业技术大学) WuHan University(武汉大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14084 2025-09-19 cs.CV 57%

AD-DINOv3: Enhancing DINOv3 for Zero-Shot Anomaly Detection with Anomaly-Aware Calibration

Jingyi Yuan, Jianxiong Ye, Wenkang Chen, Chenqiang Gao

机构 * School of Intelligent Systems Engineering, Sun Yat-Sen University(智能系统工程学院,中山大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 其他多模态 6 篇

2411.14982 2025-09-19 cs.CV cs.CL 81%

Large Multi-modal Models Can Interpret Features in Large Multi-modal Models

Kaichen Zhang, Yifei Shen, Bo Li, Ziwei Liu

机构 * S-Lab, NTU, Singapore(新加坡国立大学S实验室) LMMs-Lab Team(多模态模型实验室团队) Microsoft Research Asia(微软亚洲研究院)

专题命中 其他多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.00046 2025-09-19 eess.IV cs.CV cs.LG 79%

Mixture of Multicenter Experts in Multimodal AI for Debiased Radiotherapy Target Delineation

Yujin Oh, Sangjoon Park, Xiang Li, Pengfei Jin, Yi Wang, Jonathan Paly, Jason Efstathiou, Annie Chan, Jun Won Kim, Hwa Kyung Byun, Ik Jae Lee, Jaeho Cho, Chan Woo Wee, Peng Shu, Peilong Wang, Nathan Yu, Jason Holmes, Jong Chul Ye, Quanzheng Li, Wei Liu, Woong Sub Koom, Jin Sung Kim, Kyungsang Kim

机构 * Center for Advanced Medical Computing and Analysis (CAMCA), Department of Radiology, Massachusetts General Hospital (MGH) and Harvard Medical School(先进医学计算与分析中心(CAMCA)、放射科、麻省总医院(MGH)和哈佛医学院) Department of Radiation Oncology, Yonsei University College of Medicine(燕京大学医学院放射肿瘤科) Institute for Innovation in Digital Healthcare, Yonsei University(数字医疗创新研究所、燕京大学) Department of Radiation Oncology, Massachusetts General Hospital(麻省总医院放射肿瘤科) Department of Radiation Oncology, Gangnam Severance Hospital(江南松云医院放射肿瘤科) Department of Radiation Oncology, Yongin Severance Hospital(永兴松云医院放射肿瘤科) School of Computing, University of Georgia(佐治亚大学计算机学院) Department of Radiation Oncology, Mayo Clinic(梅奥诊所放射肿瘤科) Kim Jaechul Graduate School of AI, Korea Advanced Institute of Science and Technology(金 Jaechul人工智能研究生院、韩国科学技术院)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV

Comments 12 pages, 5 figures, 4 tables, 1 supplementary material

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.11960 2025-09-19 cs.CL cs.LG 57%

Fast Multipole Attention: A Scalable Multilevel Attention Mechanism for Text and Images

Yanming Kang, Giang Tran, Hans De Sterck

机构 * University of Waterloo(滑铁卢大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14600 2025-09-19 cs.LG physics.bio-ph 56%

TICA-Based Free Energy Matching for Machine-Learned Molecular Dynamics

Alexander Aghili, Andy Bruce, Daniel Sabo, Razvan Marinescu

机构 * Baskin Engineering, University of California - Santa Cruz, Santa Cruz, United States(加州大学圣克鲁兹分校贝斯基工程)

专题命中 其他多模态 :multi-modal(abstract,comments)

Comments Proceedings of the ICML 2025 Workshop on Multi-modal Foundation Models and Large Language Models for Life Sciences, Vancouver, Canada. 2025. Copyright 2025 by the author(s). 4 Pages 5 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14557 2025-09-19 math.OC 50%

Data-Driven Contextual Optimization with Gaussian Mixtures: Flow-Based Generalization, Robust Models, and Multistage Extensions

YoungChul Yoon, Grani A. Hanasusanto, Yijie Wang

专题命中 其他多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20035 2025-09-19 cs.HC 50%

Cam-2-Cam: Exploring the Design Space of Dual-Camera Interactions for Smartphone-based Augmented Reality

Brandon Woodard, Melvin He, Mose Sakashita, Jing Qian, Zainab Iftikhar, Joseph J. LaViola

专题命中 其他多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏