arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-07 至 2025-10-07 共收录 11 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 11 篇

2504.09532 2025-10-07 cs.RO cs.AI 86%

Humanoid Agent via Embodied Chain-of-Action Reasoning with Multimodal Foundation Models for Zero-Shot Loco-Manipulation

Congcong Wen, Geeta Chandra Raju Bethala, Yu Hao, Niraj Pudasaini, Hao Huang, Shuaihang Yuan, Baoru Huang, Anh Nguyen, Mengyu Wang, Anthony Tzes, Yi Fang

机构 * Embodied AI and Robotics (AIR) Lab, New York University, New York, USA and NYUAD Center for Artificial Intelligence and Robotics, New York University Abu Dhabi, Abu Dhabi, UAE(纽约大学Embodied AI和机器人实验室及纽约大学阿布扎赫尔人工智能与机器人中心) Harvard AI and Robotics Lab, Harvard University, Boston, USA(哈佛大学人工智能与机器人实验室) Department of Computer Science, University College London, London, UK(伦敦大学学院计算机科学系) Department of Computer Science, University of Liverpool, UK(利物浦大学计算机科学系) NYUAD Center for Artificial Intelligence and Robotics, New York University Abu Dhabi, Abu Dhabi, UAE(纽约大学阿布扎赫尔人工智能与机器人中心)

专题命中 多模态Agent :multimodal(title,abstract);multimodal foundation model(title);分类 cs.AI

Comments website link: https://humanoid-coa.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03612 2025-10-07 cs.AI cs.CR 83%

Cross-Modal Content Optimization for Steering Web Agent Preferences

Tanqiu Jiang, Min Bai, Nikolaos Pappas, Yanjun Qi, Sandesh Swamy

机构 * Stony Brook University(石溪大学) AWS AI Labs(亚马逊人工智能实验室)

专题命中 多模态Agent :cross-modal(title,abstract);multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04791 2025-10-07 cs.SE 82%

GUISpector: An MLLM Agent Framework for Automated Verification of Natural Language Requirements in GUI Prototypes

Kristian Kolthoff, Felix Kretzer, Simone Paolo Ponzetto, Alexander Maedche, Christian Bartelt

专题命中 多模态Agent :MLLM(title,abstract);multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04560 2025-10-07 cs.AI 79%

ContextNav: Towards Agentic Multimodal In-Context Learning

Honghao Fu, Yuan Ouyang, Kai-Wei Chang, Yiwei Wang, Zi Huang, Yujun Cai

机构 * The University of Queensland(昆士兰大学) Nanjing University(南京大学) University of California, Los Angeles(加州大学洛杉矶分校) University of California, Merced(加州大学默塞德分校)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07634 2025-10-07 cs.RO cs.AI cs.CV 62%

Neural Brain: A Neuroscience-inspired Framework for Embodied Agents

Jian Liu, Xiongtao Shi, Thai Duy Nguyen, Haitian Zhang, Tianxiang Zhang, Wei Sun, Yanjie Li, Athanasios V. Vasilakos, Giovanni Iacca, Arshad Ali Khan, Arvind Kumar, Jae Won Cho, Ajmal Mian, Lihua Xie, Erik Cambria, Lin Wang

机构 * School of Electrical and Electronic Engineering, Nanyang Technological University(南洋理工大学电子与电气工程学院) School of Artificial Intelligence and Robotics, Hunan University(湖南大学人工智能与机器人学院) School of Intelligence Science and Engineering, Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)智能科学与工程学院) Department of Information and Communication Technology, University of Agder(阿格德大学信息与通信技术系) Department of Information Engineering and Computer Science, University of Trento(特伦特大学信息工程与计算机科学系) Elm Company(Elm公司) Division of Computational Science and Technology, KTH Royal Institute of Technology(皇家理工学院计算科学与技术系) School of Artificial Intelligence and Data Science, Sejong University(世宗大学人工智能与数据科学学院) Department of Computer Science of the University of Western Australia(西澳大学计算机科学系) College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.AI

Comments 51 pages, 17 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04039 2025-10-07 cs.CV cs.AI 62%

\textsc{GUI-Spotlight}: Adaptive Iterative Focus Refinement for Enhanced GUI Visual Grounding

Bin Lei, Nuo Xu, Ali Payani, Mingyi Hong, Chunhua Liao, Yu Cao, Caiwen Ding

机构 * University of Minnesota(明尼苏达大学) Cisco Research(思科研究) Lawrence Livermore National Labs(劳伦斯利弗莫尔国家实验室)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04023 2025-10-07 cs.AI cs.CL 62%

LLM-Based Data Science Agents: A Survey of Capabilities, Challenges, and Future Directions

Mizanur Rahman, Amran Bhuiyan, Mohammed Saidul Islam, Md Tahmid Rahman Laskar, Ridwan Mahbub, Ahmed Masry, Shafiq Joty, Enamul Hoque

机构 * York University(约克大学) Vector Institute for AI(人工智能矢量研究所) Dialpad Inc.(Dialpad公司) Nanyang Technological University(南洋理工大学) Salesforce AI Research(Salesforce人工智能研究)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL、cs.AI

Comments Survey paper; 45 data science agents; under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03829 2025-10-07 cs.NI cs.AI 57%

A4FN: an Agentic AI Architecture for Autonomous Flying Networks

André Coelho, Pedro Ribeiro, Helder Fontes, Rui Campos

机构 * FCT – Fundação para a Ciência e a Tecnologia, I.P.(葡萄牙科学与技术基金会)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments This paper has been accepted for presentation in the Auto ML for Zero-Touch Network Management Workshop (WS04-01) at the IEEE International Symposium on Personal, Indoor and Mobile Radio Communications (PIMRC) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00413 2025-10-07 cs.CV 57%

PAL-UI: Planning with Active Look-back for Vision-Based GUI Agents

Zikang Liu, Junyi Li, Wayne Xin Zhao, Dawei Gao, Yaliang Li, Ji-rong Wen

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学全球化人工智能学院) Department of Data Science, City University of Hong Kong(香港城市大学数据科学系) Alibaba Group(阿里巴巴集团)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07778 2025-10-07 cs.CV 57%

A Neurosymbolic Agent System for Compositional Visual Reasoning

Yichang Xu, Gaowen Liu, Ramana Rao Kompella, Sihao Hu, Fatih Ilhan, Selim Furkan Tekin, Zachary Yahn, Ling Liu

机构 * Georgia Institute of Technology(佐治亚理工学院) Cisco Systems(思科系统)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.20140 2025-10-07 cs.AI 57%

MAD-Sherlock: Multi-Agent Debate for Visual Misinformation Detection

Kumud Lakara, Georgia Channing, Christian Rupprecht, Juil Sock, Philip Torr, John Collomosse, Christian Schroeder de Witt

机构 * University of Oxford, Oxford, UK(牛津大学) BBC AI Research, London, UK(BBC人工智能研究) University of Surrey, Guildford, UK(萨里大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏