arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-12 至 2025-11-12 共收录 8 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 8 篇

2511.01310 2025-11-12 cs.MA 78%

From Pixels to Cooperation Multi Agent Reinforcement Learning based on Multimodal World Models

Sureyya Akin, Kavita Srivastava, Prateek B. Kapoor, Pradeep G. Sethi, Sunita Q. Patel, Rahu Srivastava

专题命中 多模态Agent :multimodal(title,abstract)

Comments We have identified critical issues in the code implementation that severely deviate from Algorithm 1, invalidating all experimental results and conclusions. Despite exhaustive efforts to correct these issues, we find they fundamentally undermine the paper's core claims. To uphold academic integrity and prevent misinformation, we are withdrawing this manuscript

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18515 2025-11-12 cs.MA 78%

Socialized Learning and Emergent Behaviors in Multi-Agent Systems based on Multimodal Large Language Models

Sureyya Akin, Shruti T. Tiwari, Ram Bhattacharya, Sagar A. Raman, Kiran Mohanty, Sita Krishnan

专题命中 多模态Agent :multimodal(title,abstract)

Comments We have identified critical issues in the code implementation that severely deviate from Algorithm 1, invalidating all experimental results and conclusions. Despite exhaustive efforts to correct these issues, we find they fundamentally undermine the paper's core claims. To uphold academic integrity and prevent misinformation, we are withdrawing this manuscript

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08098 2025-11-12 cs.RO cs.AI cs.CL cs.HC 73%

PerspAct: Enhancing LLM Situated Collaboration Skills through Perspective Taking and Active Vision

Sabrina Patania, Luca Annese, Anita Pellegrini, Silvia Serino, Anna Lambiase, Luca Pallonetto, Silvia Rossi, Simone Colombani, Tom Foulsham, Azzurra Ruggeri, Dimitri Ognibene

机构 * University of Milan-Bicocca(米兰-比科卡大学) University of Naples Federico II(那不勒斯费德里科二世大学) Oversonic Robotics(Oversonic机器人公司) University of Essex(埃塞克斯大学) TUM School of Social Sciences and Technology(慕尼黑技术大学社会科学与技术学院)

专题命中 多模态Agent :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CL、cs.AI

Comments Accepted at IAS19

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08521 2025-11-12 cs.CV 57%

UniVA: Universal Video Agent towards Open-Source Next-Generation Video Generalist

Zhengyang Liang, Daoan Zhang, Huichi Zhou, Rui Huang, Bobo Li, Yuechen Zhang, Shengqiong Wu, Xiaohan Wang, Jiebo Luo, Lizi Liao, Hao Fei

机构 * Singapore Management University(新加坡管理大学) University of Rochester(罗切斯特大学) University College London(伦敦大学学院) National University of Singapore(新加坡国立大学) The Chinese University of Hong Kong(香港中文大学) Stanford University(斯坦福大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

Comments Technical Report. 24 figures, 37 pages. Website: https://univa.online/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24563 2025-11-12 cs.CV 57%

OSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use Agents

Hongrui Jia, Jitong Liao, Xi Zhang, Haiyang Xu, Tianbao Xie, Chaoya Jiang, Ming Yan, Si Liu, Wei Ye, Fei Huang

机构 * Peking University(北京大学) Tongyi Lab, Alibaba Group(通义实验室,阿里巴巴集团) Beijing Zhongguancun Academy(北京中关村学院)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11075 2025-11-12 cs.AI 57%

Emergence of Goal-Directed Behaviors via Active Inference with Self-Prior

Dongmin Kim, Hoshinori Kanazawa, Naoto Yoshida, Yasuo Kuniyoshi

机构 * Graduate School of Information Science and Technology(信息科学与技术研究生院) The University of Tokyo(东京大学) Graduate School of Informatics(信息学研究生院) Kyoto University(京都大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments 23 pages, 8 figures, Code is available at https://github.com/kim135797531/self-prior

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24030 2025-11-12 cs.MA 50%

Human Machine Social Hybrid Intelligence:A Collaborative Decision Making Framework for Large Model Agent Groups and Human Experts

Ahmet Akkaya Melih, Yamuna Singh, Kunal L. Agarwal, Priya Mukherjee, Kiran Pattnaik, Hanuman Bhatia

专题命中 多模态Agent :multi-modal(abstract)

Comments We have identified critical issues in the code implementation that severely deviate from Algorithm 1, invalidating all experimental results and conclusions. Despite exhaustive efforts to correct these issues, we find they fundamentally undermine the paper's core claims. To uphold academic integrity and prevent misinformation, we are withdrawing this manuscript

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01386 2025-11-12 cs.LG cs.AR 50%

CATransformers: Carbon Aware Transformers Through Joint Model-Hardware Optimization

Irene Wang, Newsha Ardalani, Mostafa Elhoushi, Daniel Jiang, Samuel Hsia, Ekin Sumbul, Divya Mahajan, Carole-Jean Wu, Bilge Acun

机构 * Georgia Institute of Technology(佐治亚理工学院) FAIR at Meta(Meta的FAIR部门) Reality Labs at Meta(Meta的Reality Labs) Meta

专题命中 多模态Agent :multi-modal(abstract)

Journal ref 39th Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏