arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-08 至 2025-09-08 共收录 38 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 5 篇

2509.05205 2025-09-08 eess.AS cs.SD 79%

MEAN-RIR: Multi-Modal Environment-Aware Network for Robust Room Impulse Response Estimation

Jiajian Chen, Jiakang Chen, Hang Chen, Qing Wang, Yu Gao, Jun Du

机构 * University of Science and Technology of China(科学技术大学) AI Research Center, Midea Group (Shanghai) Co.,Ltd.(美的集团(上海)有限公司人工智能研究中心)

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 eess.AS

Comments Accepted by ASRU 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04605 2025-09-08 cs.CL 79%

Spoken in Jest, Detected in Earnest: A Systematic Review of Sarcasm Recognition -- Multimodal Fusion, Challenges, and Future Prospects

Xiyuan Gao, Shekhar Nayak, Matt Coler

机构 * Campus Fryslân, University of Groningen(格罗宁根大学弗里桑校区)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

Comments 20 pages, 7 figures, Submitted to IEEE Transactions on Affective Computing

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04606 2025-09-08 cs.CL cs.AI cs.CV 75%

Sample-efficient Integration of New Modalities into Large Language Models

Osman Batur İnce, André F. T. Martins, Oisin Mac Aodha, Edoardo M. Ponti

机构 * University of Edinburgh(爱丁堡大学) Instituto de Telecomunicações(电信研究所) Instituto Superior Técnico, Universidade de Lisboa(里斯本大学技术学院) Unbabel

专题命中 音频语音多模态 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Pre-print

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22809 2025-09-08 cs.CL cs.AI cs.HC 62%

First Steps Towards Overhearing LLM Agents: A Case Study With Dungeons & Dragons Gameplay

Andrew Zhu, Evan Osgood, Chris Callison-Burch

机构 * University of Pennsylvania(宾夕法尼亚大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.AI

Comments 9 pages, 5 figures. COLM 2025 Workshop on AI Agents

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04736 2025-09-08 cs.CV 61%

WatchHAR: Real-time On-device Human Activity Recognition System for Smartwatches

Taeyoung Yeon, Vasco Xu, Henry Hoffmann, Karan Ahuja

机构 * Northwestern University(西北大学) University of Chicago(芝加哥大学)

专题命中 音频语音多模态 :multimodal(abstract,comments);分类 cs.CV

Comments 8 pages, 4 figures, ICMI '25 (27th International Conference on Multimodal Interaction), October 13-17, 2025, Canberra, ACT, Australia

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视频多模态 7 篇

2509.04751 2025-09-08 cs.IR cs.LG 89%

Multimodal Foundation Model-Driven User Interest Modeling and Behavior Analysis on Short Video Platforms

Yushang Zhao, Yike Peng, Li Zhang, Qianyi Sun, Zhihui Zhang, Yingying Zhuang

专题命中 视频多模态 :multimodal(title,abstract);multimodal foundation model(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.05265 2025-09-08 cs.AI 79%

MMoE: Robust Spoiler Detection with Multi-modal Information and Domain-aware Mixture-of-Experts

Zinan Zeng, Sen Ye, Zijian Cai, Heng Wang, Yuhan Liu, Haokai Zhang, Minnan Luo

机构 * Xi’an Jiaotong University(西安交通大学)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04714 2025-09-08 cs.SI 78%

ThumbnailTruth: A Multi-Modal LLM Approach for Detecting Misleading YouTube Thumbnails Across Diverse Cultural Settings

Wajiha Naveed, Zartash Afzal Uzmi, Zafar Ayyub Qazi

专题命中 视频多模态 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04957 2025-09-08 cs.CV cs.MM cs.SD eess.AS 67%

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper

Gehui Chen, Guan'an Wang, Xiaowen Huang, Jitao Sang

机构 * School of Computer Science Technology, Beijing Jiaotong University Beijing China Beijing Key Laboratory of Traffic Data Mining Key Laboratory of Big Data \& Artificial Intelligence in Transportation, Ministry of Education Beijing China Technology, Beijing Jiaotong University Key Laboratory of Big Data \& Artificial Intelligence in Transportation, Ministry of Education

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV、cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15298 2025-09-08 cs.CV 57%

TPA: Temporal Prompt Alignment for Fetal Congenital Heart Defect Classification

Darya Taratynova, Alya Almsouti, Beknur Kalmakhanbet, Numan Saeed, Mohammad Yaqub

机构 * Department of Machine Learning Mohamed bin Zayed University of Artificial Intelligence (MBZUAI) Abu Dhabi, UAE(机器学习系,Mohamed bin Zayed人工智能大学(MBZUAI),阿布扎比,阿联酋) Department of Computer Vision Mohamed bin Zayed University of Artificial Intelligence (MBZUAI) Abu Dhabi, UAE(计算机视觉系,Mohamed bin Zayed人工智能大学(MBZUAI),阿布扎比,阿联酋)

专题命中 视频多模态 :image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.03496 2025-09-08 cs.HC cs.AI 57%

Evaluating Temporal Patterns in Applied Infant Affect Recognition

Allen Chang, Lauren Klein, Marcelo R. Rosales, Weiyang Deng, Beth A. Smith, Maja J. Matarić

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

Comments 8 pages, 6 figures, 10th International Conference on Affective Computing and Intelligent Interaction (ACII 2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18724 2025-09-08 cs.CL 57%

Large Language Models with Temporal Reasoning for Longitudinal Clinical Summarization and Prediction

Maya Kruse, Shiyue Hu, Nicholas Derby, Yifu Wu, Samantha Stonbraker, Bingsheng Yao, Dakuo Wang, Elizabeth Goldberg, Yanjun Gao

机构 * University of Colorado Anschutz Medical Campus(科罗拉多大学安舒茨医学校园) University of Colorado Boulder(科罗拉多大学波德分校) Northeastern University(东北大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 跨模态检索 4 篇

2509.04376 2025-09-08 cs.CV 70%

AnomalyLMM: Bridging Generative Knowledge and Discriminative Retrieval for Text-Based Person Anomaly Search

Hao Ju, Hu Zhang, Zhedong Zheng

机构 * Faculty of Science and Technology and Institute of Collaborative Innovation, University of Macau(科技学院和协同创新研究所,澳门大学) CSIRO Data61(CSIRO数据61)

专题命中 跨模态检索 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06020 2025-09-08 cs.AI cs.CV 62%

ArtRAG: Retrieval-Augmented Generation with Structured Context for Visual Art Understanding

Shuai Wang, Ivona Najdenkoska, Hongyi Zhu, Stevan Rudinac, Monika Kackovic, Nachoem Wijnberg, Marcel Worring

机构 * University of Amsterdam(阿姆斯特丹大学) College of Business(商学院) Economics, University of Johannesburg(经济系,约翰内斯堡大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04836 2025-09-08 cs.RO 50%

COMMET: A System for Human-Induced Conflicts in Mobile Manipulation of Everyday Tasks

Dongping Li, Shaoting Peng, John Pohovey, Katherine Rose Driggs-Campbell

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Department of Electrical and Computer Engineering(电气与计算机工程系) ZJU-UIUC Institute(浙大-伊利诺伊大学联合学院)

专题命中 跨模态检索 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.09647 2025-09-08 cs.RO 50%

Object Instance Retrieval in Assistive Robotics: Leveraging Fine-Tuned SimSiam with Multi-View Images Based on 3D Semantic Map

Taichi Sakaguchi, Akira Taniguchi, Yoshinobu Hagiwara, Lotfi El Hafi, Shoichi Hasegawa, Tadahiro Taniguchi

机构 * Ritsumeikan University(立命馆大学) Soka University(早稻田大学) Kyoto University(京都大学)

专题命中 跨模态检索 :multimodal(abstract)

Comments See website at https://emergentsystemlabstudent.github.io/MultiViewRetrieve/. Accepted to IROS2024

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 多模态生成 1 篇

2503.19065 2025-09-08 cs.CV 79%

WikiAutoGen: Towards Multi-Modal Wikipedia-Style Article Generation

Zhongyu Yang, Jun Chen, Dannong Xu, Junjie Fei, Xiaoqian Shen, Liangbing Zhao, Chun-Mei Feng, Mohamed Elhoseiny

机构 * King Abdullah University of Science and Technology(国王阿卜杜勒阿齐兹大学科学与技术大学) Lanzhou University(兰州大学) Meta AI The University of Sydney(悉尼大学) IHPC, A*STAR(IHPC,A*STAR)

专题命中 多模态生成 :multi-modal(title);multimodal(abstract);分类 cs.CV

Comments ICCV 2025, Project in https://wikiautogen.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态评测 9 篇

2509.04844 2025-09-08 cs.MM cs.AI cs.IR 84%

REMOTE: A Unified Multimodal Relation Extraction Framework with Multilevel Optimal Transport and Mixture-of-Experts

Xinkui Lin, Yongxiu Xu, Minghao Tang, Shilong Zhang, Hongbo Xu, Hao Xu, Yubin Wang

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院)

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI、cs.MM

Comments ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04658 2025-09-08 cs.RO 82%

Surformer v2: A Multimodal Classifier for Surface Understanding from Touch and Vision

Manish Kansana, Sindhuja Penchala, Shahram Rahimi, Noorbakhsh Amiri Golilarz

机构 * Department of Computer Science(计算机科学系) Engineering Mississippi State University Mississippi State, USA(工程学硕士州大学密西西比州) Department of Computer Science The University of Alabama Tuscaloosa, USA(计算机科学系阿拉巴马大学塔斯卡洛osa)

专题命中 多模态评测 :multimodal(title,abstract);multi-modal(abstract)

Comments 6 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04469 2025-09-08 cs.CL cs.AI 81%

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing

David Berghaus, Armin Berger, Lars Hillebrand, Kostadin Cvejoski, Rafet Sifa

机构 * Fraunhofer IAIS(弗劳恩霍夫人工智能研究所) Lamarr Institute(拉马尔研究所)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04823 2025-09-08 cs.SI cs.CL 79%

Evaluating Cognitive-Behavioral Fixation via Multimodal User Viewing Patterns on Social Media

Yujie Wang, Yunwei Zhao, Jing Yang, Han Han, Shiguang Shan, Jie Zhang

机构 * State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences(中国科学院人工智能安全国家重点实验室,计算技术研究所) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15862 2025-09-08 cs.LG cs.CY 78%

Quantifying Holistic Review: A Multi-Modal Approach to College Admissions Prediction

Jun-Wei Zeng, Jerry Shen

机构 * Vanke Meisha Academy(万科梅沙书院)

专题命中 多模态评测 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04502 2025-09-08 cs.CL cs.AI 76%

VaccineRAG: Boosting Multimodal Large Language Models' Immunity to Harmful RAG Samples

Qixin Sun, Ziqin Wang, Hengyuan Zhao, Yilin Li, Kaiyou Song, Linjiang Huang, Xiaolin Hu, Qingpei Guo, Si Liu

专题命中 多模态评测 :multimodal(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04757 2025-09-08 cs.CV cs.AI 62%

MCANet: A Multi-Scale Class-Specific Attention Network for Multi-Label Post-Hurricane Damage Assessment using UAV Imagery

Zhangding Liu, Neda Mohammadi, John E. Taylor

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 34 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04468 2025-09-08 cs.CL cs.AI 62%

Evaluating Large Language Models for Financial Reasoning: A CFA-Based Benchmark Study

Xuan Yao, Qianteng Wang, Xinbo Liu, Ke-Wei Huang

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.10833 2025-09-08 eess.IV cs.CV 57%

Automatic segmentation of Organs at Risk in Head and Neck cancer patients from CT and MRI scans

Sébastien Quetin, Andrew Heschl, Mauricio Murillo, Rohit Murali, Piotr Pater, George Shenouda, Shirin A. Enger, Farhad Maleki

机构 * Medical Physics Unit, Department of Oncology, McGill University(麦吉尔大学医学物理单位) Montreal Institute for Learning Algorithms, Mila(蒙特利尔学习算法研究所) Department of Computer Science, University of Calgary(卡尔加里大学计算机科学系) University of British Columbia(不列颠哥伦比亚大学) Division of Radiation Oncology, Department of Oncology, McGill University(麦吉尔大学放射肿瘤学部) McGill University Health Centre(麦吉尔大学健康中心) Lady Davis Institute for Medical Research, Jewish General Hospital(Lady Davis医学研究所,犹太通用医院) Department of Diagnostic Radiology, McGill University(麦吉尔大学诊断放射学部) Department of Radiology, University of Florida(佛罗里达大学放射学部)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态Agent 4 篇

2408.02544 2025-09-08 cs.CL 83%

Caution for the Environment: Multimodal LLM Agents are Susceptible to Environmental Distractions

Xinbei Ma, Yiting Wang, Yao Yao, Tongxin Yuan, Aston Zhang, Zhuosheng Zhang, Hai Zhao

机构 * School of Computer Science(计算机学院) Key Laboratory of Shanghai Education Commission for Intelligent Interaction and Cognitive Engineering(智能交互与认知工程重点实验室) Shanghai Jiao Tong University(上海交通大学) Shanghai Key Laboratory of Trusted Data Circulation and Governance in Web3(Web3可信数据流通与治理上海市重点实验室) GenAI, Meta(Meta GenAI)

专题命中 多模态Agent :multimodal(title,abstract);MLLM(abstract);分类 cs.CL

Comments ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.17422 2025-09-08 cs.RO cs.CV 79%

Multimodal LLM Guided Exploration and Active Mapping using Fisher Information

Wen Jiang, Boshu Lei, Katrina Ashton, Kostas Daniilidis

机构 * University of Pennsylvania(宾夕法尼亚大学) Archimedes, Athena RC(阿基米德、阿提卡RC)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04908 2025-09-08 cs.AI cs.CL cs.CV cs.HC 75%

SparkUI-Parser: Enhancing GUI Perception with Robust Grounding and Parsing

Hongyi Jing, Jiafu Chen, Chen Rao, Ziqiang Dang, Jiajie Teng, Tianyi Chu, Juncheng Mo, Shuo Fang, Huaizhong Lin, Rui Lv, Chenguang Ma, Lei Zhao

专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.15876 2025-09-08 cs.HC cs.AI cs.GR 57%

AI-in-the-loop: The future of biomedical visual analytics applications in the era of AI

Katja Bühler, Thomas Höllt, Thomas Schulz, Pere-Pau Vázquez

机构 * Vienna Research Center for Visual Computing(维也纳视觉计算研究中心) VRVis GmbH(VRVis公司) Delft University of Technology(代尔夫特理工大学) University of Bonn(波恩大学) Lamarr Institute for Machine Learning and Artificial Intelligence(拉马尔机器学习与人工智能研究所) Universitat Politècnica de Catalunya(巴塞罗那理工大学)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

Comments Accepted for publication in IEEE Computer Graphics & Applications

Journal ref K. Bühler, T. Hollt, T. Schultz and P. Vazquez, "AI-in-The-Loop: The Future of Biomedical Visual Analytics Applications in the Era of AI" in IEEE Computer Graphics and Applications, vol. 45, no. 02, pp. 90-99, March-April 2025

详情

展开后加载摘要…

URL PDF HTML 收藏