arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-30 至 2025-09-30 共收录 13 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 13 篇

2509.24855 2025-09-30 cs.AI 79%

PhysicsMinions: Winning Gold Medals in the Latest Physics Olympiads with a Coevolutionary Multimodal Multi-Agent System

Fangchen Yu, Junchi Yao, Ziyi Wang, Haiyuan Wan, Youling Huang, Bo Zhang, Shuyue Hu, Dongzhan Zhou, Ning Ding, Ganqu Cui, Lei Bai, Wanli Ouyang, Peng Ye

机构 * Shanghai AI Laboratory(上海人工智能实验室) CUHK-Shenzhen(香港中文大学(深圳)) CUHK(香港中文大学) UESTC(电子科技大学) Tsinghua University(清华大学) DUT(大连理工大学)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24314 2025-09-30 cs.AI 79%

MedMMV: A Controllable Multimodal Multi-Agent Framework for Reliable and Verifiable Clinical Reasoning

Hongjun Liu, Yinghao Zhu, Yuhui Wang, Yitao Long, Zeyu Lai, Lequan Yu, Chen Zhao

机构 * New York University(纽约大学) NYU Shanghai(纽约大学上海分校) The University of Hong Kong(香港大学) Zhejiang University(浙江大学)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

Comments 25 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12839 2025-09-30 cs.AI 79%

From An LLM Swarm To A PDDL-Empowered HIVE: Planning Self-Executed Instructions In A Multi-Modal Jungle

Kaustubh Vyas, Damien Graux, Yijun Yang, Sébastien Montella, Chenxin Diao, Wendi Zhou, Pavlos Vougiouklis, Ruofei Lai, Yang Ren, Keshuang Li, Jeff Z. Pan

机构 * Huawei Technologies Ltd., UK(华为技术有限公司,英国) University of Edinburgh, UK(爱丁堡大学,英国)

专题命中 多模态Agent :multi-modal(title,abstract);分类 cs.AI

Comments Published as a conference paper at ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20148 2025-09-30 cs.AI 77%

MineAnyBuild: Benchmarking Spatial Planning for Open-world AI Agents

Ziming Wei, Bingqian Lin, Zijian Jiao, Yunshuang Nie, Liang Ma, Yuecheng Liu, Yuzheng Zhuang, Xiaodan Liang

机构 * Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区) Shanghai Jiao Tong University(上海交通大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 多模态Agent :multimodal(abstract);multi-modal(abstract);MLLM(abstract);分类 cs.AI

Comments Accepted by NeurIPS 2025 Datasets and Benchmarks Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23517 2025-09-30 cs.CV cs.AI 76%

Evaluating point-light biological motion in multimodal large language models

Akila Kadambi, Marco Iacoboni, Lisa Aziz-Zadeh, Srini Narayanan

机构 * Psychiatry and Biobehavioral Sciences, UCLA(乌尔拉克大学精神病学与生物行为科学系) Brain and Creativity Institute, USC(美国大学脑与创造力研究所) Google DeepMind, Zurich(谷歌深度Mind瑞士分公司)

专题命中 多模态Agent :multimodal(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24413 2025-09-30 cs.RO cs.HC 71%

DynaMIC: Dynamic Multimodal In-Context Learning Enabled Embodied Robot Counterfactual Resistance Ability

Tianqiang Yan, Ziqiao Lin, Sicheng Wang, Tianwei Zhang, Zhenglong Sun

机构 * Faculty of Information Technology, Monash University(墨尔本大学信息技术学院) School of Science and Engineering, the Chinese University of Hong Kong-Shenzhen(香港中文大学(深圳)科学与工程学院) Shenzhen Institute of Artificial Intelligence and Robotics for Society, the Chinese University of Hong Kong-Shenzhen(深圳人工智能与机器人研究院)

专题命中 多模态Agent :multimodal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25185 2025-09-30 cs.CV 70%

PixelCraft: A Multi-Agent System for High-Fidelity Visual Reasoning on Structured Images

Shuoshuo Zhang, Zijian Li, Yizhen Zhang, Jingjing Fu, Lei Song, Jiang Bian, Jun Zhang, Yujiu Yang, Rui Wang

机构 * Microsoft Research(微软研究院) Tsinghua University(清华大学) Hong Kong University of Science and Technology(香港理工大学)

专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24133 2025-09-30 cs.CV cs.CL 62%

Generalist Scanner Meets Specialist Locator: A Synergistic Coarse-to-Fine Framework for Robust GUI Grounding

Zhecheng Li, Guoxian Song, Yiwei Wang, Zhen Xiong, Junsong Yuan, Yujun Cai

机构 * University of California, San Diego(加州大学圣地亚哥分校) ByteDance(字节跳动) University of California, Merced(加州大学默塞德分校) University of Southern California(南加州大学) University at Buffalo(布法罗大学) The University of Queensland(昆士兰大学)

专题命中 多模态Agent :cross-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11790 2025-09-30 cs.AI 57%

Visualizing Thought: Conceptual Diagrams Enable Robust Planning in LMMs

Nasim Borazjanizadeh, Roei Herzig, Eduard Oks, Trevor Darrell, Rogerio Feris, Leonid Karlinsky

机构 * Xero Inc.(Xero公司) Berkeley AI Research, UC Berkeley(伯克利人工智能研究实验室,伯克利大学) MIT–IBM Watson AI Lab(麻省理工–IBM沃森人工智能实验室)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07978 2025-09-30 cs.AI quant-ph 57%

Agents for self-driving laboratories applied to quantum computing

Shuxiang Cao, Zijian Zhang, Mohammed Alghadeer, Simone D Fasciati, Michele Piscitelli, Mustafa Bakr, Peter Leek, Alán Aspuru-Guzik

机构 * 1 Clarendon Laboratory, Department of Physics, University of Oxford, Oxford, OX1 3PU, UK 2 Department of Computer Science, University of Toronto, Toronto, ON M5S 2E4, Canada 3 Vector Institute for Artificial Intelligence, Toronto, ON, M5G 1M1, Canada 4 Department of Chemistry, University of Toronto, Toronto, ON M5S 3H6, Canada 5 Department of Materials Science \& Engineering, University of Toronto, Toronto, ON M5S 3E4, Canada 6 Department of Chemical Engineering \& Applied Chemistry, University of Toronto, Toronto, ON M5S 3E5, Canada 7 Canadian Institute for Advanced Research (CIFAR), Toronto, ON M5G 1M1, Canada

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23698 2025-09-30 cs.CL 57%

VIVA+: Human-Centered Situational Decision-Making

Zhe Hu, Yixiao Ren, Guanzhong Liu, Jing Li, Yu Yin

机构 * Department of Computing, The Hong Kong Polytechnic University(香港理工大学计算机系) Research Centre for Data Science & Artificial Intelligence(数据科学与人工智能研究中心) Department of Computer and Data Sciences, Case Western Reserve University(凯斯西储大学计算机与数据科学系)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL

Comments EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23087 2025-09-30 cs.LG 50%

Unleashing Flow Policies with Distributional Critics

Deshu Chen, Yuchen Liu, Zhijian Zhou, Chao Qu, Yuan Qi

机构 * Fudan University(复旦大学) INFLY TECH (Shanghai) Co., Ltd(INFLY TECH(上海)有限公司)

专题命中 多模态Agent :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22900 2025-09-30 cs.CR cs.SE 50%

Towards Context-aware Mobile Privacy Notice: Implementation of A Deployable Contextual Privacy Policies Generator

Haochen Gong, Zhen Tao, Shidong Pan, Zhenchang Xing, Xiaoyu Sun

专题命中 多模态Agent :multimodal(abstract)

Comments Accepted by ASE 2025, Tool Demonstration Track

详情

展开后加载摘要…

URL PDF HTML 收藏