arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-30 至 2025-09-30 共收录 142 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 34 篇

2507.04952 2025-09-30 cs.CL cs.SE 70%

ArtifactsBench: Bridging the Visual-Interactive Gap in LLM Code Generation Evaluation

Chenchen Zhang, Yuhang Li, Can Xu, Jiaheng Liu, Ao Liu, Changzhi Zhou, Ken Deng, Dengpeng Wu, Guanhua Huang, Kejiao Li, Qi Yi, Ruibin Xiong, Shihui Hu, Yue Zhang, Yuhao Jiang, Zenan Xu, Yuanxing Zhang, Wiggin Zhou, Chayse Zhou, Fengzong Lian

机构 * Tencent Hunyuan Team(腾讯文脉团队)

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22991 2025-09-30 cs.CL cs.AI cs.CV cs.IR cs.LG 67%

ADAM: A Diverse Archive of Mankind for Evaluating and Enhancing LLMs in Biographical Reasoning

Jasin Cekinmez, Omid Ghahroodi, Saad Fowad Chandle, Dhiman Gupta, Ehsaneddin Asgari

机构 * Qatar Computing Research Institute(卡塔尔计算研究所) Princeton University(普林斯顿大学) Virginia Tech(弗吉尼亚理工大学) Amity University(阿米蒂大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25032 2025-09-30 cs.RO cs.AI cs.CV 62%

AIRoA MoMa Dataset: A Large-Scale Hierarchical Dataset for Mobile Manipulation

Ryosuke Takanami, Petr Khrapchenkov, Shu Morikuni, Jumpei Arima, Yuta Takaba, Shunsuke Maeda, Takuya Okubo, Genki Sano, Satoshi Sekioka, Aoi Kadoya, Motonari Kambara, Naoya Nishiura, Haruto Suzuki, Takanori Yoshimoto, Koya Sakamoto, Shinnosuke Ono, Hu Yang, Daichi Yashima, Aoi Horo, Tomohiro Motoda, Kensuke Chiyoma, Hiroshi Ito, Koki Fukuda, Akihito Goto, Kazumi Morinaga, Yuya Ikeda, Riko Kawada, Masaki Yoshikawa, Norio Kosuge, Yuki Noguchi, Kei Ota, Tatsuya Matsushima, Yusuke Iwasawa, Yutaka Matsuo, Tetsuya Ogata

机构 * The University of Tokyo(东京大学) AI Robot Association (AIRoA)(人工智能机器人协会) Toyota Motor Corporation(丰田汽车公司) Telexistence, Inc.(Telexistence公司) National Institute of Advanced Industrial Science and Technology (AIST)(国家先进工业科学与技术研究院) Waseda University(早稻田大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24120 2025-09-30 cs.CL 57%

EduVidQA: Generating and Evaluating Long-form Answers to Student Questions based on Lecture Videos

Sourjyadip Ray, Shubham Sharma, Somak Aditya, Pawan Goyal

机构 * Indian Institute of Technology, Kharagpur(印度理工学院,克哈格普尔) Panjab University, Chandigarh(旁遮普大学,昌迪加尔)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL

Comments EMNLP 2025 (Main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23999 2025-09-30 cs.CV cs.LG 57%

TREAT-Net: Tabular-Referenced Echocardiography Analysis for Acute Coronary Syndrome Treatment Prediction

Diane Kim, Minh Nguyen Nhat To, Sherif Abdalla, Teresa S. M. Tsang, Purang Abolmaesumi, and Christina Luong

机构 * University of British Columbia, Vancouver, BC, Canada(不列颠哥伦比亚大学) Vancouver Coastal Health(温哥华海岸健康) Vancouver General Hospital(温哥华总医院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Comments 11 pages, 2 figures, MICCAI ASMUS 2025 paper

Journal ref Simplifying Medical Ultrasound (ASMUS 2025), LNCS 16165, Springer, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23243 2025-09-30 cs.CV 57%

Increasing the Diversity in RGB-to-Thermal Image Translation for Automotive Applications

Kaili Wang, Leonardo Ravaglia, Roberto Longo, Lore Goetschalckx, David Van Hamme, Julie Moeyersoms, Ben Stoffelen, Tom De Schepper

机构 * imec imec-IPI-Ghent University(IMEC-IPI-根特大学)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV

Comments Accepted in IEEE Sensors 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23111 2025-09-30 cs.RO cs.AI 57%

Liaohe-CobotMagic-PnP: an Imitation Learning Dataset of Intelligent Robot for Industrial Applications

Chen Yizhe, Wang Qi, Hu Dongxiao, Jingzhe Fang, Liu Sichao, Zixin An, Hongliang Niu, Haoran Liu, Li Dong, Chuanfen Feng, Lan Dapeng, Liu Yu, Zhibo Pang

机构 * School of Information Science and Engineering, Shandong Normal University(信息科学与工程学院,山东师范大学) School of Advanced Technology, Xi’an Jiaotong-Liverpool University(先进技术学院,西安交通大学利物浦大学) Liaohe Laboratory, Innovation Institute of Intelligent Robotics Shenyang Co., Ltd.(辽河实验室,沈阳智能机器人创新研究院有限公司) School of Information Engineering, Shenyang University of Chemical Technology(信息工程学院,沈阳化工大学) Shenyang Institute of Automation, Chinese Academy of Sciences(沈阳自动化研究所,中国科学院) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

Comments Accepted to IAI 2025 (International Conference on Industrial Artificial Intelligence), Shenyang, China, Aug 21 - 24, 2025. Preprint (before IEEE copyright transfer)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24217 2025-09-30 cs.LG cs.NA math.NA 50%

MDD-Thinker: Towards Large Reasoning Models for Major Depressive Disorder Diagnosis

Yuyang Sha, Hongxin Pan, Gang Luo, Caijuan Shi, Jing Wang, Kefeng Li

专题命中 多模态评测 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23789 2025-09-30 cs.LG cs.CR 50%

Visual CoT Makes VLMs Smarter but More Fragile

Chunxue Xu, Yiwei Wang, Yujun Cai, Bryan Hooi, Songze Li

机构 * Southeast University, China(东南大学) University of California, Merced, USA(加州大学梅德福分校) The University of Queensland, Australia(昆士兰大学) National University of Singapore, Singapore(新加坡国立大学)

专题命中 多模态评测 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18686 2025-09-30 cs.LG cs.IT math.IT math.PR stat.ML 50%

A Unified MDL-based Binning and Tensor Factorization Framework for PDF Estimation

Mustafa Musab, Joseph K. Chege, Arie Yeredor, Martin Haardt

机构 * Communications Research Laboratory(通信研究中心) Ilmenau University of Technology(伊门瑙技术大学) School of Electrical Engineering(电气工程学院) Tel Aviv University(特拉维夫大学)

专题命中 多模态评测 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态Agent 13 篇

2509.24855 2025-09-30 cs.AI 79%

PhysicsMinions: Winning Gold Medals in the Latest Physics Olympiads with a Coevolutionary Multimodal Multi-Agent System

Fangchen Yu, Junchi Yao, Ziyi Wang, Haiyuan Wan, Youling Huang, Bo Zhang, Shuyue Hu, Dongzhan Zhou, Ning Ding, Ganqu Cui, Lei Bai, Wanli Ouyang, Peng Ye

机构 * Shanghai AI Laboratory(上海人工智能实验室) CUHK-Shenzhen(香港中文大学(深圳)) CUHK(香港中文大学) UESTC(电子科技大学) Tsinghua University(清华大学) DUT(大连理工大学)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24314 2025-09-30 cs.AI 79%

MedMMV: A Controllable Multimodal Multi-Agent Framework for Reliable and Verifiable Clinical Reasoning

Hongjun Liu, Yinghao Zhu, Yuhui Wang, Yitao Long, Zeyu Lai, Lequan Yu, Chen Zhao

机构 * New York University(纽约大学) NYU Shanghai(纽约大学上海分校) The University of Hong Kong(香港大学) Zhejiang University(浙江大学)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

Comments 25 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12839 2025-09-30 cs.AI 79%

From An LLM Swarm To A PDDL-Empowered HIVE: Planning Self-Executed Instructions In A Multi-Modal Jungle

Kaustubh Vyas, Damien Graux, Yijun Yang, Sébastien Montella, Chenxin Diao, Wendi Zhou, Pavlos Vougiouklis, Ruofei Lai, Yang Ren, Keshuang Li, Jeff Z. Pan

机构 * Huawei Technologies Ltd., UK(华为技术有限公司,英国) University of Edinburgh, UK(爱丁堡大学,英国)

专题命中 多模态Agent :multi-modal(title,abstract);分类 cs.AI

Comments Published as a conference paper at ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20148 2025-09-30 cs.AI 77%

MineAnyBuild: Benchmarking Spatial Planning for Open-world AI Agents

Ziming Wei, Bingqian Lin, Zijian Jiao, Yunshuang Nie, Liang Ma, Yuecheng Liu, Yuzheng Zhuang, Xiaodan Liang

机构 * Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区) Shanghai Jiao Tong University(上海交通大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 多模态Agent :multimodal(abstract);multi-modal(abstract);MLLM(abstract);分类 cs.AI

Comments Accepted by NeurIPS 2025 Datasets and Benchmarks Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23517 2025-09-30 cs.CV cs.AI 76%

Evaluating point-light biological motion in multimodal large language models

Akila Kadambi, Marco Iacoboni, Lisa Aziz-Zadeh, Srini Narayanan

机构 * Psychiatry and Biobehavioral Sciences, UCLA(乌尔拉克大学精神病学与生物行为科学系) Brain and Creativity Institute, USC(美国大学脑与创造力研究所) Google DeepMind, Zurich(谷歌深度Mind瑞士分公司)

专题命中 多模态Agent :multimodal(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24413 2025-09-30 cs.RO cs.HC 71%

DynaMIC: Dynamic Multimodal In-Context Learning Enabled Embodied Robot Counterfactual Resistance Ability

Tianqiang Yan, Ziqiao Lin, Sicheng Wang, Tianwei Zhang, Zhenglong Sun

机构 * Faculty of Information Technology, Monash University(墨尔本大学信息技术学院) School of Science and Engineering, the Chinese University of Hong Kong-Shenzhen(香港中文大学(深圳)科学与工程学院) Shenzhen Institute of Artificial Intelligence and Robotics for Society, the Chinese University of Hong Kong-Shenzhen(深圳人工智能与机器人研究院)

专题命中 多模态Agent :multimodal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25185 2025-09-30 cs.CV 70%

PixelCraft: A Multi-Agent System for High-Fidelity Visual Reasoning on Structured Images

Shuoshuo Zhang, Zijian Li, Yizhen Zhang, Jingjing Fu, Lei Song, Jiang Bian, Jun Zhang, Yujiu Yang, Rui Wang

机构 * Microsoft Research(微软研究院) Tsinghua University(清华大学) Hong Kong University of Science and Technology(香港理工大学)

专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24133 2025-09-30 cs.CV cs.CL 62%

Generalist Scanner Meets Specialist Locator: A Synergistic Coarse-to-Fine Framework for Robust GUI Grounding

Zhecheng Li, Guoxian Song, Yiwei Wang, Zhen Xiong, Junsong Yuan, Yujun Cai

机构 * University of California, San Diego(加州大学圣地亚哥分校) ByteDance(字节跳动) University of California, Merced(加州大学默塞德分校) University of Southern California(南加州大学) University at Buffalo(布法罗大学) The University of Queensland(昆士兰大学)

专题命中 多模态Agent :cross-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11790 2025-09-30 cs.AI 57%

Visualizing Thought: Conceptual Diagrams Enable Robust Planning in LMMs

Nasim Borazjanizadeh, Roei Herzig, Eduard Oks, Trevor Darrell, Rogerio Feris, Leonid Karlinsky

机构 * Xero Inc.(Xero公司) Berkeley AI Research, UC Berkeley(伯克利人工智能研究实验室,伯克利大学) MIT–IBM Watson AI Lab(麻省理工–IBM沃森人工智能实验室)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07978 2025-09-30 cs.AI quant-ph 57%

Agents for self-driving laboratories applied to quantum computing

Shuxiang Cao, Zijian Zhang, Mohammed Alghadeer, Simone D Fasciati, Michele Piscitelli, Mustafa Bakr, Peter Leek, Alán Aspuru-Guzik

机构 * 1 Clarendon Laboratory, Department of Physics, University of Oxford, Oxford, OX1 3PU, UK 2 Department of Computer Science, University of Toronto, Toronto, ON M5S 2E4, Canada 3 Vector Institute for Artificial Intelligence, Toronto, ON, M5G 1M1, Canada 4 Department of Chemistry, University of Toronto, Toronto, ON M5S 3H6, Canada 5 Department of Materials Science \& Engineering, University of Toronto, Toronto, ON M5S 3E4, Canada 6 Department of Chemical Engineering \& Applied Chemistry, University of Toronto, Toronto, ON M5S 3E5, Canada 7 Canadian Institute for Advanced Research (CIFAR), Toronto, ON M5G 1M1, Canada

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23698 2025-09-30 cs.CL 57%

VIVA+: Human-Centered Situational Decision-Making

Zhe Hu, Yixiao Ren, Guanzhong Liu, Jing Li, Yu Yin

机构 * Department of Computing, The Hong Kong Polytechnic University(香港理工大学计算机系) Research Centre for Data Science & Artificial Intelligence(数据科学与人工智能研究中心) Department of Computer and Data Sciences, Case Western Reserve University(凯斯西储大学计算机与数据科学系)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL

Comments EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23087 2025-09-30 cs.LG 50%

Unleashing Flow Policies with Distributional Critics

Deshu Chen, Yuchen Liu, Zhijian Zhou, Chao Qu, Yuan Qi

机构 * Fudan University(复旦大学) INFLY TECH (Shanghai) Co., Ltd(INFLY TECH(上海)有限公司)

专题命中 多模态Agent :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22900 2025-09-30 cs.CR cs.SE 50%

Towards Context-aware Mobile Privacy Notice: Implementation of A Deployable Contextual Privacy Policies Generator

Haochen Gong, Zhen Tao, Shidong Pan, Zhenchang Xing, Xiaoyu Sun

专题命中 多模态Agent :multimodal(abstract)

Comments Accepted by ASE 2025, Tool Demonstration Track

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态训练与对齐 24 篇

2509.24896 2025-09-30 cs.CV 86%

DAM: Dual Active Learning with Multimodal Foundation Model for Source-Free Domain Adaptation

Xi Chen, Hongxun Yao, Zhaopan Xu, Kui Jiang

机构 * Harbin Institute of Technology(哈尔滨工业大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);multimodal foundation model(title);分类 cs.CV

Comments 5 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24693 2025-09-30 q-bio.NC 86%

Brain Harmony: A Multimodal Foundation Model Unifying Morphology and Function into 1D Tokens

Zijian Dong, Ruilin Li, Joanna Su Xian Chong, Niousha Dehestani, Yinghui Teng, Yi Lin, Zhizhou Li, Yichi Zhang, Yapei Xie, Leon Qi Rong Ooi, B. T. Thomas Yeo, Juan Helen Zhou

专题命中 多模态训练与对齐 :multimodal(title,abstract);multimodal foundation model(title)

Comments NeurIPS 2025. The first two authors contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24298 2025-09-30 cs.HC cs.AI cs.CL cs.CY cs.MM 85%

Bridging the behavior-neural gap: A multimodal AI reveals the brain's geometry of emotion more accurately than human self-reports

Changde Du, Yizhuo Lu, Zhongyu Huang, Yi Sun, Zisen Zhou, Shaozheng Qin, Huiguang He

机构 * State Key Laboratory of Brain Cognition and Brain-inspired Intelligence Technology, Institute of Automation, Chinese Academy of Sciences(脑认知与脑启发智能技术重点实验室,自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学) School of Future Technology, University of Chinese Academy of Sciences(未来技术学院,中国科学院大学) State Key Laboratory of Cognitive Neuroscience and Learning, Beijing Normal University(认知神经科学与学习国家重点实验室,北京师范大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CL、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22697 2025-09-30 cs.CV cs.AI cs.LG 84%

Learning Hyperspectral Images with Curated Text Prompts for Efficient Multimodal Alignment

Abhiroop Chatterjee, Susmita Ghosh

机构 * Jadavpur University(贾瓦帕尔大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted at the IEEE/CVF International Conference on Computer Vision (ICCV 2025), Workshop on Curated Data for Efficient Learning

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24505 2025-09-30 cs.CV 83%

Robust Multimodal Semantic Segmentation with Balanced Modality Contributions

Jiaqi Tan, Xu Zheng, Fangyu Li, Yang Liu

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) HKUST(GZ)(香港科技大学(广州)) INSAIT

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24776 2025-09-30 cs.CV cs.AI 81%

VTPerception-R1: Enhancing Multimodal Reasoning via Explicit Visual and Textual Perceptual Grounding

Yizhuo Ding, Mingkang Chen, Zhibang Feng, Tong Xiao, Wanying Qu, Wenqi Shao, Yanwei Fu

机构 * Fudan University(复旦大学) Shanghai AI Laboratory(上海人工智能实验室) The University of Hong Kong(香港大学) Shenzhen University(深圳大学) University of Science and Technology of China(中国科学技术大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24734 2025-09-30 cs.LG cs.AI cs.CV 81%

A TRIANGLE Enables Multimodal Alignment Beyond Cosine Similarity

Giordano Cicchetti, Eleonora Grassucci, Danilo Comminiello

机构 * Department of Information Engineering, Electronics, and Telecommunications(信息工程、电子与电信系)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏