arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-12 至 2025-11-12 共收录 52 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 11 篇

2507.14686 2025-11-12 cs.CV 70%

From Semantics, Scene to Instance-awareness: Distilling Foundation Model for Grounded Open-vocabulary Situation Recognition

Chen Cai, Tianyi Liu, Jianjun Gao, Wenyang Liu, Kejun Wu, Ruoyu Wang, Yi Wang, Soo Chin Liew

机构 * National University of Singapore(新加坡国立大学) Nanyang Technological University(南洋理工大学) Huazhong University of Science and Technology(华中科技大学) The Hong Kong Polytechnic University(香港理工大学)

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22340 2025-11-12 cs.AI cs.CL cs.CV cs.LG 67%

DynaSolidGeo: A Dynamic Benchmark for Genuine Spatial Mathematical Reasoning of VLMs in Solid Geometry

Changti Wu, Shijie Lian, Zihao Liu, Lei Zhang, Laurence Tianruo Yang, Kai Chen

机构 * East China Normal University(华东师范大学) Zhongguancun Academy(中关村学院) Huazhong University of Science and Technology(华中科技大学) Peking University(北京大学) Zhengzhou University(郑州大学) Zhongguancun Institute of Artificial Intelligence(中关村人工智能研究院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments The code and dataset are available at \href{https://zgca-ai4edu.github.io/DynaSolidGeo/}{DynaSolidGeo}

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08140 2025-11-12 cs.CV 57%

PEOD: A Pixel-Aligned Event-RGB Benchmark for Object Detection under Challenging Conditions

Luoping Cui, Hanqing Liu, Mingjie Liu, Endian Lin, Donghong Jiang, Yuhao Wang, Chuang Zhu

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07897 2025-11-12 cs.AI cs.LG 57%

Data Descriptions from Large Language Models with Influence Estimation

Chaeri Kim, Jaeyeon Bae, Taehwan Kim

机构 * Ulsan National Institute of Science and Technology(UNIST)(乌山国立科学技术研究院)

专题命中 多模态评测 :cross-modal(abstract);分类 cs.AI

Journal ref Published in EMNLP 2025, check our project on this https URL : https://github.com/kimchaeri/Data-Descriptions-from-Large-Language-Models-with-Influence-Estimation

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态Agent 8 篇

2511.01310 2025-11-12 cs.MA 78%

From Pixels to Cooperation Multi Agent Reinforcement Learning based on Multimodal World Models

Sureyya Akin, Kavita Srivastava, Prateek B. Kapoor, Pradeep G. Sethi, Sunita Q. Patel, Rahu Srivastava

专题命中 多模态Agent :multimodal(title,abstract)

Comments We have identified critical issues in the code implementation that severely deviate from Algorithm 1, invalidating all experimental results and conclusions. Despite exhaustive efforts to correct these issues, we find they fundamentally undermine the paper's core claims. To uphold academic integrity and prevent misinformation, we are withdrawing this manuscript

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18515 2025-11-12 cs.MA 78%

Socialized Learning and Emergent Behaviors in Multi-Agent Systems based on Multimodal Large Language Models

Sureyya Akin, Shruti T. Tiwari, Ram Bhattacharya, Sagar A. Raman, Kiran Mohanty, Sita Krishnan

专题命中 多模态Agent :multimodal(title,abstract)

Comments We have identified critical issues in the code implementation that severely deviate from Algorithm 1, invalidating all experimental results and conclusions. Despite exhaustive efforts to correct these issues, we find they fundamentally undermine the paper's core claims. To uphold academic integrity and prevent misinformation, we are withdrawing this manuscript

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08098 2025-11-12 cs.RO cs.AI cs.CL cs.HC 73%

PerspAct: Enhancing LLM Situated Collaboration Skills through Perspective Taking and Active Vision

Sabrina Patania, Luca Annese, Anita Pellegrini, Silvia Serino, Anna Lambiase, Luca Pallonetto, Silvia Rossi, Simone Colombani, Tom Foulsham, Azzurra Ruggeri, Dimitri Ognibene

机构 * University of Milan-Bicocca(米兰-比科卡大学) University of Naples Federico II(那不勒斯费德里科二世大学) Oversonic Robotics(Oversonic机器人公司) University of Essex(埃塞克斯大学) TUM School of Social Sciences and Technology(慕尼黑技术大学社会科学与技术学院)

专题命中 多模态Agent :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CL、cs.AI

Comments Accepted at IAS19

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08521 2025-11-12 cs.CV 57%

UniVA: Universal Video Agent towards Open-Source Next-Generation Video Generalist

Zhengyang Liang, Daoan Zhang, Huichi Zhou, Rui Huang, Bobo Li, Yuechen Zhang, Shengqiong Wu, Xiaohan Wang, Jiebo Luo, Lizi Liao, Hao Fei

机构 * Singapore Management University(新加坡管理大学) University of Rochester(罗切斯特大学) University College London(伦敦大学学院) National University of Singapore(新加坡国立大学) The Chinese University of Hong Kong(香港中文大学) Stanford University(斯坦福大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

Comments Technical Report. 24 figures, 37 pages. Website: https://univa.online/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24563 2025-11-12 cs.CV 57%

OSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use Agents

Hongrui Jia, Jitong Liao, Xi Zhang, Haiyang Xu, Tianbao Xie, Chaoya Jiang, Ming Yan, Si Liu, Wei Ye, Fei Huang

机构 * Peking University(北京大学) Tongyi Lab, Alibaba Group(通义实验室,阿里巴巴集团) Beijing Zhongguancun Academy(北京中关村学院)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11075 2025-11-12 cs.AI 57%

Emergence of Goal-Directed Behaviors via Active Inference with Self-Prior

Dongmin Kim, Hoshinori Kanazawa, Naoto Yoshida, Yasuo Kuniyoshi

机构 * Graduate School of Information Science and Technology(信息科学与技术研究生院) The University of Tokyo(东京大学) Graduate School of Informatics(信息学研究生院) Kyoto University(京都大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments 23 pages, 8 figures, Code is available at https://github.com/kim135797531/self-prior

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24030 2025-11-12 cs.MA 50%

Human Machine Social Hybrid Intelligence:A Collaborative Decision Making Framework for Large Model Agent Groups and Human Experts

Ahmet Akkaya Melih, Yamuna Singh, Kunal L. Agarwal, Priya Mukherjee, Kiran Pattnaik, Hanuman Bhatia

专题命中 多模态Agent :multi-modal(abstract)

Comments We have identified critical issues in the code implementation that severely deviate from Algorithm 1, invalidating all experimental results and conclusions. Despite exhaustive efforts to correct these issues, we find they fundamentally undermine the paper's core claims. To uphold academic integrity and prevent misinformation, we are withdrawing this manuscript

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01386 2025-11-12 cs.LG cs.AR 50%

CATransformers: Carbon Aware Transformers Through Joint Model-Hardware Optimization

Irene Wang, Newsha Ardalani, Mostafa Elhoushi, Daniel Jiang, Samuel Hsia, Ekin Sumbul, Divya Mahajan, Carole-Jean Wu, Bilge Acun

机构 * Georgia Institute of Technology(佐治亚理工学院) FAIR at Meta(Meta的FAIR部门) Reality Labs at Meta(Meta的Reality Labs) Meta

专题命中 多模态Agent :multi-modal(abstract)

Journal ref 39th Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态训练与对齐 6 篇

2511.07941 2025-11-12 cs.CV cs.AI 84%

Libra-MIL: Multimodal Prototypes Stereoscopic Infused with Task-specific Language Priors for Few-shot Whole Slide Image Classification

Zhenfeng Zhuang, Fangyu Zhou, Liansheng Wang

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21486 2025-11-12 cs.CV 83%

Reasoning-Enhanced Domain-Adaptive Pretraining of Multimodal Large Language Models for Short Video Content Governance

Zixuan Wang, Yu Sun, Hongwei Wang, Baoyu Jing, Xiang Shen, Xin Dong, Zhuolin Hao, Hongyu Xiong, Yang Song

机构 * TikTok Inc.(字节跳动公司)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments Camera Ready for EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08152 2025-11-12 cs.CV cs.LG 79%

Boomda: Balanced Multi-objective Optimization for Multimodal Domain Adaptation

Jun Sun, Xinxin Zhang, Simin Hong, Jian Zhu, Xiang Gao

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00598 2025-11-12 cs.CV 57%

DGL-RSIS: Decoupling Global Spatial Context and Local Class Semantics for Training-Free Remote Sensing Image Segmentation

Boyi Li, Ce Zhang, Richard M. Timmerman, Wenxuan Bao

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03565 2025-11-12 cs.CV 57%

Bridged Semantic Alignment for Zero-shot 3D Medical Image Diagnosis

Haoran Lai, Zihang Jiang, Qingsong Yao, Rongsheng Wang, Zhiyang He, Xiaodong Tao, Weifu Lv, Wei Wei, S. Kevin Zhou

机构 * University of Science and Technology of China(中国科学技术大学) Suzhou Institute for Advanced Research(苏州先进研究所) Stanford University(斯坦福大学) iFlytek Co. Ltd.(iFlytek公司) The First Affiliated Hospital of USTC, Division of Life Sciences and Medicine, USTC(中国科学技术大学第一附属医院)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11639 2025-11-12 cs.IR 50%

OneRec-Think: In-Text Reasoning for Generative Recommendation

Zhanyu Liu, Shiyao Wang, Xingmei Wang, Rongzhou Zhang, Jiaxin Deng, Honghui Bao, Jinghao Zhang, Wuchao Li, Pengfei Zheng, Xiangyu Wu, Yifei Hu, Qigen Hu, Xinchen Luo, Lejian Ren, Zixing Zhang, Qianqian Wang, Kuo Cai, Yunfan Wu, Hongtao Cheng, Zexuan Cheng, Lu Ren, Huanjie Wang, Yi Su, Ruiming Tang, Kun Gai, Guorui Zhou

专题命中 多模态训练与对齐 :cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 其他多模态 4 篇

2511.08246 2025-11-12 cs.AI 79%

Where and What Matters: Sensitivity-Aware Task Vectors for Many-Shot Multimodal In-Context Learning

Ziyu Ma, Chenhui Gou, Yiming Hu, Yong Wang, Xiangxiang Chu, Bohan Zhuang, Jianfei Cai

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07912 2025-11-12 cs.AI 57%

Neurophysiological Characteristics of Adaptive Reasoning for Creative Problem-Solving Strategy

Jun-Young Kim, Young-Seok Kweon, Gi-Hwan Shin, Seong-Whan Lee

机构 * Dept. of Artificial Intelligence(人工智能系) Korea University(韩国大学) Dept. of Brain and Cognitive Engineering(脑科学与认知工程系)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

Comments 4 pages, 4 figures, 1 table,

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18108 2025-11-12 cs.CV 57%

Unveiling Visual Perception in Language Models: An Attention Head Analysis Approach

Jing Bi, Junjia Guo, Yunlong Tang, Lianggong Bruce Wen, Zhang Liu, Chenliang Xu

机构 * University of Rochester(罗切斯特大学) Corning Inc(康宁公司)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

Journal ref CVPR 2025 (IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08438 2025-11-12 astro-ph.CO astro-ph.IM cs.LG 50%

Galactification: painting galaxies onto dark matter only simulations using a transformer-based model

Shivam Pandey, Christopher C. Lovell, Chirag Modi, Benjamin D. Wandelt

机构 * Department of Physics and Astronomy, Johns Hopkins University(约翰霍普金斯大学物理与天文学系) Kavli Institute for Cosmology, University of Cambridge(剑桥大学卡弗利宇宙研究所) Center for Cosmology and Particle Physics, New York University(纽约大学宇宙与粒子物理中心)

专题命中 其他多模态 :multi-modal(abstract)

Comments 8 pages, 4 figures. , accepted at Machine Learning and the Physical Sciences Workshop at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏