arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-29 至 2025-10-29 共收录 60 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 12 篇

2507.12841 2025-10-29 cs.CV 83%

AnyCap Project: A Unified Framework, Dataset, and Benchmark for Controllable Omni-modal Captioning

Yiming Ren, Zhiqiang Lin, Yu Li, Gao Meng, Weiyun Wang, Junjie Wang, Zicheng Lin, Jifeng Dai, Yujiu Yang, Wenhai Wang, Ruihang Chu

机构 * Tsinghua University(清华大学) Shanghai AI Laboratory(上海人工智能实验室) Fudan University(复旦大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 多模态评测 :omni-modal(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22672 2025-10-29 cs.CV cs.CL cs.RO 81%

Look and Tell: A Dataset for Multimodal Grounding Across Egocentric and Exocentric Views

Anna Deichler, Jonas Beskow

机构 * KTH Royal Institute of Technology(皇家理工学院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments 10 pages, 6 figures, 2 tables. Accepted to the NeurIPS 2025 Workshop on SPACE in Vision, Language, and Embodied AI (SpaVLE). Dataset: https://huggingface.co/datasets/annadeichler/KTH-ARIA-referential

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23727 2025-10-29 cs.LG cs.CL 79%

MUStReason: A Benchmark for Diagnosing Pragmatic Reasoning in Video-LMs for Multimodal Sarcasm Detection

Anisha Saha, Varsha Suresh, Timothy Hospedales, Vera Demberg

机构 * Max Planck Institute for Informatics, Saarland Informatics Campus(马克斯·普朗克研究所信息学研究所,萨尔兰州信息学校园) Saarland University(萨尔兰州大学) The University of Edinburgh(爱丁堡大学) Samsung AI Center, Cambridge(三星人工智能中心,剑桥)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01271 2025-10-29 cs.LG cs.AI 79%

PULSE: Practical Evaluation Scenarios for Large Multimodal Model Unlearning

Tatsuki Kawakami, Kazuki Egashira, Atsuyuki Miyai, Go Irie, Kiyoharu Aizawa

机构 * The University of Tokyo(东京大学) Tokyo University of Science(东京科学大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Comments Accepted at NeurIPS 2025 Workshop: Evaluating the Evolving LLM Lifecycle

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11842 2025-10-29 cs.CV cs.CL 62%

Video-SafetyBench: A Benchmark for Safety Evaluation of Video LVLMs

Xuannan Liu, Zekun Li, Zheqi He, Peipei Li, Shuhan Xia, Xing Cui, Huaibo Huang, Xi Yang, Ran He

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Beijing Academy of Artificial Intelligence(北京人工智能研究院) University of California, Santa Barbara(加州大学圣芭芭拉分校) Center for Research on Intelligent Perception and Computing, NLPR, CASIA(中国科学院CASIA智能感知与计算中心)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Accepted by NeurIPS 2025 Dataset and Benchmark Track, Project page: https://liuxuannan.github.io/Video-SafetyBench.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04897 2025-10-29 cs.CV 61%

From Objects to Anywhere: A Holistic Benchmark for Multi-level Visual Grounding in 3D Scenes

Tianxu Wang, Zhuofan Zhang, Ziyu Zhu, Yue Fan, Jing Xiong, Pengxiang Li, Xiaojian Ma, Qing Li

机构 * State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室,BIGAI) Tsinghua University(清华大学) Peking University(北京大学) Beijing Institute of Technology(北京理工大学)

专题命中 多模态评测 :multimodal(abstract,comments);分类 cs.CV

Comments Update v3 of the NeurIPS 2025 Datasets and Benchmarks paper (v2), including additional evaluations of state-of-the-art multimodal large language models. Project page: https://anywhere-3d.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17102 2025-10-29 cs.CV 57%

GRASP: Geospatial pixel Reasoning viA Structured Policy learning

Chengjie Jiang, Yunqi Zhou, Jiafeng Yan, Jing Li, Jiayang Li, Yue Zhou, Hongjie He, Jonathan Li

机构 * Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) School of Information, Central University of Finance and Economics(中央财经大学信息学院) Key Laboratory of Geographic Information Science (Ministry of Education), East China Normal University(教育部地理信息科学重点实验室,华东师范大学) School of Electronic and Computer Engineering, Peking University(北京大学电子工程学院) Department of Geography and Environmental Management, University of Waterloo(滑铁卢大学地理与环境管理系)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Comments 15 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20510 2025-10-29 cs.CV 57%

CPathAgent: An Agent-based Foundation Model for Interpretable High-Resolution Pathology Image Analysis Mimicking Pathologists' Diagnostic Logic

Yuxuan Sun, Yixuan Si, Chenglu Zhu, Kai Zhang, Zhongyi Shui, Bowen Ding, Tao Lin, Lin Yang

机构 * College of Computer Science and Technology, Zhejiang University, China(浙江大学计算机科学与技术学院) Research Center for Industries of the Future and School of Engineering, Westlake University, China(未来产业研究中心和西湖大学工程学院) Department of Computer Science and Engineering, The Ohio State University, USA(俄亥俄州立大学计算机科学与工程系)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Comments 52 pages, 34 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.06560 2025-10-29 cs.CL 57%

NeedleInATable: Exploring Long-Context Capability of Large Language Models towards Long-Structured Tables

Lanrui Wang, Mingyu Zheng, Hongyin Tang, Zheng Lin, Yanan Cao, Jingang Wang, Xunliang Cai, Weiping Wang

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22414 2025-10-29 cs.HC 50%

Complementary Human-AI Clinical Reasoning in Ophthalmology

Mertcan Sevgi, Fares Antaki, Abdullah Zafar Khan, Ariel Yuhan Ong, David Adrian Merle, Kuang Hu, Shafi Balal, Sophie-Christin Kornelia Ernst, Josef Huemer, Gabriel T. Kaufmann, Hagar Khalid, Faye Levina, Celeste Limoli, Ana Paula Ribeiro Reis, Samir Touma, Anil Palepu, Khaled Saab, Ryutaro Tanno, Valentin Liévin, Tao Tu, Yong Cheng, Mike Schaekermann, S. Sara Mahdavi, Elahe Vedadi, David Stutz, Vivek Natarajan, Alan Karthikesalingam, Pearse A. Keane, Wei-Hung Weng

专题命中 多模态评测 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09539 2025-10-29 cs.IR 50%

GlobalMood: A cross-cultural benchmark for music emotion recognition

Harin Lee, Elif Çelen, Peter Harrison, Manuel Anglada-Tort, Pol van Rijn, Minsu Park, Marc Schönwiesner, Nori Jacoby

专题命中 多模态评测 :multimodal(abstract)

Comments To be presented at International Society of Music Information Retrieval (ISMIR)

Journal ref International Society for Music Information Retrieval Conference (ISMIR), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态Agent 4 篇

2502.21142 2025-10-29 cs.AI cs.LG q-bio.NC 79%

Multimodal Dreaming: A Global Workspace Approach to World Model-Based Reinforcement Learning

Léopold Maytié, Roland Bertin Johannet, Rufin VanRullen

机构 * Univ Toulouse, CNRS, CerCo, and ANITI, Artificial and Natural Intelligence Toulouse Institute(图卢兹大学、法国国家科学研究中心、CerCo以及ANITI人工智能与自然智能图卢兹研究所)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23691 2025-10-29 cs.AI 79%

Game-TARS: Pretrained Foundation Models for Scalable Generalist Multimodal Game Agents

Zihao Wang, Xujing Li, Yining Ye, Junjie Fang, Haoming Wang, Longxiang Liu, Shihao Liang, Junting Lu, Zhiyong Wu, Jiazhan Feng, Wanjun Zhong, Zili Li, Yu Wang, Yu Miao, Bo Zhou, Yuanfan Li, Hao Wang, Zhongkai Zhao, Faming Wu, Zhengxuan Jiang, Weihao Tan, Heyuan Yao, Shi Yan, Xiangyang Li, Yitao Liang, Yujia Qin, Guang Shi

机构 * Bytedance Seed(字节跳动种子)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20077 2025-10-29 cs.RO cs.CV cs.HC 79%

Queryable 3D Scene Representation: A Multi-Modal Framework for Semantic Reasoning and Robotic Task Planning

Xun Li, Rodrigo Santa Cruz, Mingze Xi, Hu Zhang, Madhawa Perera, Ziwei Wang, Ahalya Ravendran, Brandon J. Matthews, Feng Xu, Matt Adcock, Dadong Wang, Jiajun Liu

机构 * CSIRO(澳大利亚联邦科学与工业研究组织)

专题命中 多模态Agent :multi-modal(title);multimodal(abstract);分类 cs.CV

Journal ref MM '25: Proceedings of the 33rd ACM International Conference on Multimedia (2025) Pages 12492 - 12500

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24161 2025-10-29 cs.AI cs.MM cs.RO 73%

BLM$_1$: A Boundless Large Model for Cross-Space, Cross-Task, and Cross-Embodiment Learning

Wentao Tan, Bowen Wang, Heng Zhi, Chenyu Liu, Zhe Li, Jian Liu, Zengrong Lin, Yukun Dai, Yipeng Chen, Wenjie Yang, Enci Xie, Hao Xue, Baixu Ji, Chen Xu, Zhibin Wang, Tianshi Wang, Lei Zhu, Heng Tao Shen

机构 * School of Computer Science and Technology, Tongji University(同济大学计算机科学与技术学院)

专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态训练与对齐 10 篇

2504.09060 2025-10-29 cs.LG cs.AI q-bio.GN 85%

Multimodal 3D Genome Pre-training

Minghao Yang, Pengteng Li, Yan Liang, Qianyi Cai, Zhihang Zheng, Shichen Zhang, Pengfei Zhang, Zhi-An Huang, Hui Xiong

机构 * Thrust of Artificial Intelligence, The Hong Kong University of Science and Technology (Guangzhou), China(人工智能研究所,香港科技大学(广州)) School of Artificial Intelligence, South China Normal University, China(人工智能学院,华南师范大学) Thrust of Bioscience and Biomedical Engineering, The Hong Kong University of Science and Technology (Guangzhou), China(生物科学与生物医学工程研究所,香港科技大学(广州)) Department of Computer Science, City University of Hong Kong (Dongguan), China(计算机科学系,香港城市大学(东莞)) Department of Computer Science and Engineering, The Hong Kong University of Science and Technology Hong Kong SAR, China(计算机科学与工程系,香港科技大学香港特别行政区)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);multimodal foundation model(abstract);分类 cs.AI

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23640 2025-10-29 cs.LG cs.AI 83%

Structure-Aware Fusion with Progressive Injection for Multimodal Molecular Representation Learning

Zihao Jing, Yan Sun, Yan Yi Li, Sugitha Janarthanan, Alana Deng, Pingzhao Hu

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18672 2025-10-29 cs.CV 83%

CalFuse: Multi-Modal Continual Learning via Feature Calibration and Parameter Fusion

Juncen Guo, Siao Liu, Xiaoguang Zhu, Lianlong Sun, Liangyu Teng, Jingyi Wu, Di Li, Linxiao Gong, Weiwei Jiang, Wei Zhou, Liang Song

机构 * College of Intelligent Robotics and Advanced Manufacturing, Fudan University(智能机器人与先进制造学院,复旦大学) School of Future Science and Engineering, Soochow University(未来科学与工程学院,苏州大学) DataLab: Data Science and Informatics, University of California, Davis(数据实验室:数据科学与信息学,加州大学戴维斯分校) University of Rochester(罗切斯特大学) Ningbo University(宁波大学) Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Beijing University of Posts and Telecommunications(北京邮电大学) Academy for Computer Science and Informatics, Cardiff University(计算机科学与信息学学院,卡迪夫大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05177 2025-10-29 cs.CV 80%

Long-VITA: Scaling Large Multi-modal Models to 1 Million Tokens with Leading Short-Context Accuracy

Yunhang Shen, Chaoyou Fu, Shaoqi Dong, Xiong Wang, Yi-Fan Zhang, Peixian Chen, Mengdan Zhang, Haoyu Cao, Ke Li, Shaohui Lin, Xiawu Zheng, Yan Zhang, Yiyi Zhou, Ran He, Caifeng Shan, Rongrong Ji, Xing Sun

机构 * Tencent Youtu Lab(腾讯云图实验室) Nanjing University(南京大学) East China Normal University(华东师范大学) Xiamen University(厦门大学) CASIA(中国科学院自动化研究所)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV;MLLM(comments)

Comments https://github.com/VITA-MLLM/Long-VITA

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06456 2025-10-29 cs.CV 79%

DynCIM: Dynamic Curriculum for Imbalanced Multimodal Learning

Chengxuan Qian, Kai Han, Jiaxin Liu, Zhenlong Yuan, Zhengzhong Zhu, Jingchao Wang, Chongwen Lyu, Jun Chen, Zhe Liu

机构 * Jiangsu University(江苏大学) UIUC(伊利诺伊大学香槟分校) Alibaba(阿里巴巴) UCAS(中国科学院大学) Sichuan University(四川大学) Peking University(北京大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24551 2025-10-29 cs.AI 57%

Generative AI for Healthcare: Fundamentals, Challenges, and Perspectives

Gang Chen, Changshuo Liu, Gene Anne Ooi, Marcus Tan, Zhongle Xie, Jianwei Yin, James Wei Luen Yip, Wenqiao Zhang, Jiaqi Zhu, Beng Chin Ooi

机构 * College of Computer Science and Technology, Zhejiang University, Hangzhou 310027, China(浙江大学计算机科学与技术学院) College of Software Technology, Zhejiang University, Ningbo 315100, China(浙江大学软件技术学院) School of Computing, National University of Singapore, Singapore 117417(新加坡国立大学计算机学院) Singapore General Hospital, Singapore 169608(新加坡中央医院) National University Hospital, Singapore 119074(新加坡国立医院)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14643 2025-10-29 cs.CV 57%

Multispectral State-Space Feature Fusion: Bridging Shared and Cross-Parametric Interactions for Object Detection

Jifeng Shen, Haibo Zhan, Shaohua Dong, Xin Zuo, Wankou Yang, Haibin Ling

机构 * School of Electrical and Information Engineering, Jiangsu University, Zhenjiang, 212013, China(江苏大学电气与信息工程学院) Department of Computer Science and Engineering, University of North Texas, Denton, TX 76207, USA(德克萨斯大学北卡罗来纳分校计算机科学与工程系) School of Computer Science and Engineering, Jiangsu University of Science and Technology, Zhenjiang, 212003, China(江苏科技大学计算机科学与工程学院) School of Automation, Southeast University, Nanjing, 210096, China(东南大学自动化学院) Bodhi Intelligence Lab, Department of Artificial Intelligence, Westlake University, Hangzhou, Zhejiang 310030, China(西湖大学人工智能研究院)

专题命中 多模态训练与对齐 :cross-modal(abstract);分类 cs.CV

Comments submitted on 30/4/2025, Accepted by Information Fusion

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24214 2025-10-29 cs.CV 57%

SCOPE: Saliency-Coverage Oriented Token Pruning for Efficient Multimodel LLMs

Jinhong Deng, Wen Li, Joey Tianyi Zhou, Yang He

机构 * University of Electronic Science and Technology of China(电子科技大学) Shenzhen Institute for Advanced Study(深圳先进研究 institute) CFAR, Agency for Science, Technology and Research (A*STAR)(科技研究局(A*STAR)认知与人工智能研究中心) IHPC, Agency for Science, Technology and Research (A*STAR)(科技研究局(A*STAR)人工智能中心)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00034 2025-10-29 cs.RO cs.CV 57%

GaussianFusion: Gaussian-Based Multi-Sensor Fusion for End-to-End Autonomous Driving

Shuai Liu, Quanmin Liang, Zefeng Li, Boyang Li, Kai Huang

机构 * School of Computer Science and Engineering, Sun Yat-sen University(计算机科学与工程学院,中山大学)

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV

Comments Accepted at NeurIPS2025 (Spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22304 2025-10-29 q-bio.BM 50%

ODesign: A World Model for Biomolecular Interaction Design

Odin Zhang, Xujun Zhang, Haitao Lin, Cheng Tan, Qinghan Wang, Yuanle Mo, Qiantai Feng, Gang Du, Yuntao Yu, Zichang Jin, Ziyi You, Peicong Lin, Yijie Zhang, Yuyang Tao, Shicheng Chen, Jack Xiaoyu Chen, Chenqing Hua, Weibo Zhao, Runze Ma, Yunpeng Xia, Kejun Ying, Jun Li, Yundian Zeng, Lijun Lang, Peichen Pan, Hanqun Cao, Zihao Song, Bo Qiang, Jiaqi Wang, Pengfei Ji, Lei Bai, Jian Zhang, Chang-yu Hsieh, Pheng Ann Heng, Siqi Sun, Tingjun Hou, Shuangjia Zheng

专题命中 多模态训练与对齐 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 其他多模态 5 篇

2502.07862 2025-10-29 cs.LG cs.AI cs.CV 81%

ADMN: A Layer-Wise Adaptive Multimodal Network for Dynamic Input Noise and Compute Resources

Jason Wu, Yuyang Yuan, Kang Yang, Lance Kaplan, Mani Srivastava

机构 * Electrical and Computer Engineering University of California, Los Angeles(电气与计算机工程大学加州大学洛杉矶分校) DEVCOM Army Research Laboratory(国防部陆军研究实验室) University of California, Los Angeles(加州大学洛杉矶分校) Amazon(亚马逊)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted to Neurips 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24023 2025-10-29 cs.CL 57%

Success and Cost Elicit Convention Formation for Efficient Communication

Saujas Vaduguru, Yilun Hua, Yoav Artzi, Daniel Fried

机构 * Carnegie Mellon University(卡内基梅隆大学) Department of Computer Science and Cornell Tech, Cornell University(计算机科学系和康奈尔科技,康奈尔大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23648 2025-10-29 cs.SI cs.AI 57%

RoGBot: Relationship-Oblivious Graph-based Neural Network with Contextual Knowledge for Bot Detection

Ashutosh Anshul, Mohammad Zia Ur Rehman, Sri Akash Kadali, Nagendra Kumar

机构 * Indian Institute of Technology Indore(印度理工学院印多尔分校) University of Maryland, College Park, USA(美国马里兰大学学院公园分校)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

Comments Submitted to IEEE

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24352 2025-10-29 stat.ME stat.CO 50%

Self-Normalized Quantile Empirical Saddlepoint Approximation

Hou Jian, Meng Tan, Tian Maozai

专题命中 其他多模态 :multimodal(abstract)

Comments 24 pages, 3 figures, 12 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24125 2025-10-29 cs.LG 50%

Causal Convolutional Neural Networks as Finite Impulse Response Filters

Kiran Bacsa, Wei Liu, Xudong Jian, Huangbin Liang, Eleni Chatzi

机构 * Future Resilient Systems, Singapore ETH Centre(新加坡未来韧性系统,新加坡ETH中心) Department of Industrial Systems Engineering and Management(工业系统工程与管理系) ETHZ Department of Civil, Environmental and Geomatic Engineering(土木、环境与测绘工程系)

专题命中 其他多模态 :multimodal(abstract)

Comments 14 pages, 19 figures, Under review

详情

展开后加载摘要…

URL PDF HTML 收藏