arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-11 至 2025-11-11 共收录 105 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 20 篇

2511.06020 2025-11-11 cs.DB 78%

RF-Behavior: A Multimodal Radio-Frequency Dataset for Human Behavior and Emotion Analysis

Si Zuo, Yuqing Song, Sahar Golipoor, Ying Liu, Xujun Ma, Stephan Sigg

专题命中 多模态评测 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05596 2025-11-11 cs.LG physics.comp-ph physics.flu-dyn 78%

AutoHood3D: A Multi-Modal Benchmark for Automotive Hood Design and Fluid-Structure Interaction

Vansh Sharma, Harish Jai Ganesh, Maryam Akram, Wanjiao Liu, Venkat Raman

机构 * University of Michigan Ann Arbor(密歇根大学安娜堡分校) Ford Research and Innovation Center(福特研究与创新中心)

专题命中 多模态评测 :multi-modal(title,abstract)

Journal ref 39th Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03248 2025-11-11 cs.CR 76%

Auditing M-LLMs for Privacy Risks: A Synthetic Benchmark and Evaluation Framework

Junhao Li, Jiahao Chen, Zhou Feng, Chunyi Zhou

专题命中 多模态评测 :multimodal(abstract,comments);multi-modal(abstract);cross-modal(abstract)

Comments 14 pages, 3 figures; Accepted by MMM 2026; Complete version in progress. Dataset available at https://huggingface.co/datasets/xaddh/multimodal-privacy

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10534 2025-11-11 cs.CV cs.LG cs.MM 73%

MCE: Towards a General Framework for Handling Missing Modalities under Imbalanced Missing Rates

Binyu Zhao, Wei Zhang, Zhaonian Zou

机构 * The School of Computer Science and Technology, Harbin Institute of Technology(计算机科学与技术学院,哈尔滨工业大学)

专题命中 多模态评测 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.MM

Comments This is the accepted version of an article that has been published in \textbf{Pattern Recognition}. The final version is available via the DOI, or for 50 days' free access via this Share Link: https://authors.elsevier.com/a/1m40D77nKsBm- (valid until December 28, 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04910 2025-11-11 cs.CL 70%

SDS KoPub VDR: A Benchmark Dataset for Visual Document Retrieval in Korean Public Documents

Jaehoon Lee, Sohyun Kim, Wanggeun Park, Geon Lee, Seungkyung Kim, Minyoung Lee

专题命中 多模态评测 :multimodal(abstract);cross-modal(abstract);分类 cs.CL

Comments 27 pages, 15 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10514 2025-11-11 cs.CV cs.AI cs.CL cs.LG 67%

ColorBench: Can VLMs See and Understand the Colorful World? A Comprehensive Benchmark for Color Perception, Reasoning, and Robustness

Yijun Liang, Ming Li, Chenrui Fan, Ziyue Li, Dang Nguyen, Kwesi Cobbina, Shweta Bhardwaj, Jiuhai Chen, Fuxiao Liu, Tianyi Zhou

机构 * University of Maryland, College Park(马里兰大学学院公园分校)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted by NeurIPS2025. 36 pages, including references and appendix. Code is available at https://github.com/tianyi-lab/ColorBench

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05542 2025-11-11 q-bio.NC cs.AI cs.CV cs.LG 62%

ConnectomeBench: Can LLMs Proofread the Connectome?

Jeff Brown, Andrew Kirjner, Annika Vivekananthan, Ed Boyden

机构 * MIT(麻省理工学院) HHMI(霍华德·休斯医学研究所) McGovern Institute(麦戈文研究所) MIT Departments of Brain and Cognitive Sciences(麻省理工学院脑科学与认知科学系)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

Comments To appear in NeurIPS 2025 Datasets and Benchmarks Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06840 2025-11-11 cs.CV cs.RO 57%

PanoNav: Mapless Zero-Shot Object Navigation with Panoramic Scene Parsing and Dynamic Memory

Qunchao Jin, Yilin Wu, Changhao Chen

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Comments Accepted as a poster in AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06752 2025-11-11 cs.CV 57%

Med-SORA: Symptom to Organ Reasoning in Abdomen CT Images

You-Kyoung Na, Yeong-Jun Cho

机构 * Chonnam National University(全南国立大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06522 2025-11-11 cs.AI cs.LG 57%

FractalBench: Diagnosing Visual-Mathematical Reasoning Through Recursive Program Synthesis

Jan Ondras, Marek Šuppa

机构 * MIT(麻省理工学院) Comenius University in Bratislava(布拉迪斯拉发孔院大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

Comments Accepted to The 5th Workshop on Mathematical Reasoning and AI at the 39th Conference on Neural Information Processing Systems (NeurIPS 2025); 25 pages, 14 figures, 8 tables; Code available at https://github.com/NaiveNeuron/FractalBench

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04307 2025-11-11 cs.AI 57%

GUI-360$^\circ$: A Comprehensive Dataset and Benchmark for Computer-Using Agents

Jian Mu, Chaoyun Zhang, Chiming Ni, Lu Wang, Bo Qiao, Kartik Mathur, Qianhui Wu, Yuhang Xie, Xiaojun Ma, Mengyu Zhou, Si Qin, Liqun Li, Yu Kang, Minghua Ma, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang

机构 * Nanjing University(南京大学) Microsoft(微软) ZJU-UIUC(浙大-UIUC) Peking University(北京大学)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.00527 2025-11-11 eess.IV cs.CV 57%

MAROON: A Dataset for the Joint Characterization of Near-Field High-Resolution Radio-Frequency and Optical Depth Imaging Techniques

Vanessa Wirth, Johanna Bräunig, Nikolai Hofmann, Martin Vossiek, Tim Weyrich, Marc Stamminger

机构 * Friedrich-Alexander-Universität Erlangen-Nürnberg(埃朗根-纽伦堡弗里德里希-亚历山大大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06360 2025-11-11 cs.CV 57%

AesTest: Measuring Aesthetic Intelligence from Perception to Production

Guolong Wang, Heng Huang, Zhiqiang Zhang, Wentian Li, Feilong Ma, Xin Jin

机构 * University of International Business and Economics(国际商务经济大学) University of Science and Technology of China(中国科学技术大学) Huawei Technologies Co., Ltd(华为技术有限公司) Beijing Electronic Science and Technology Institute(北京电子科技研究所) Beijing Institute for General Artificial Intelligence(北京通用人工智能研究院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Comments 10 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态Agent 10 篇

2511.07315 2025-11-11 cs.CR 78%

JPRO: Automated Multimodal Jailbreaking via Multi-Agent Collaboration Framework

Yuxuan Zhou, Yang Bai, Kuofeng Gao, Tao Dai, Shu-Tao Xia

专题命中 多模态Agent :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07219 2025-11-11 q-bio.GN 78%

Integrating Epigenetic and Phenotypic Features for Biological Age Estimation in Cancer Patients via Multimodal Learning

Shuyue Jiang, Wenjing Ma, Shaojun Yu, Chang Su, Runze Yan, Jiaying Lu

专题命中 多模态Agent :multimodal(title,abstract)

Journal ref In Proceedings of The 19th IEEE International Conference on Bioinformatics and Biomedicine (BIBM 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13757 2025-11-11 cs.MA cs.AI cs.CL cs.HC 73%

MobA: Multifaceted Memory-Enhanced Adaptive Planning for Efficient Mobile Task Automation

Zichen Zhu, Hao Tang, Yansi Li, Dingye Liu, Hongshen Xu, Kunyao Lan, Danyang Zhang, Yixuan Jiang, Hao Zhou, Chenrun Wang, Situo Zhang, Liangtai Sun, Yixiao Wang, Yuheng Sun, Lu Chen, Kai Yu

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.CL、cs.AI

Comments NAACL 2025 Demo Track [code] https://github.com/OpenDFM/MobA [dataset] https://huggingface.co/datasets/OpenDFM/MobA-MobBench

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06417 2025-11-11 cs.AI 70%

AUTO-Explorer: Automated Data Collection for GUI Agent

Xiangwu Guo, Difei Gao, Mike Zheng Shou

机构 * Show Lab, National University of Singapore(新加坡国立大学展示实验室)

专题命中 多模态Agent :multimodal(abstract);MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05965 2025-11-11 cs.CV cs.AI 62%

Adaptive Agent Selection and Interaction Network for Image-to-point cloud Registration

Zhixin Cheng, Xiaotian Yin, Jiacheng Deng, Bohao Liao, Yujia Chen, Xu Zhou, Baoqun Yin, Tianzhu Zhang

专题命中 多模态Agent :cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted by AAAI2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06260 2025-11-11 cs.GT cs.AI cs.SY eess.SY 57%

LLM-Guided Reinforcement Learning with Representative Agents for Traffic Modeling

Hanlin Sun, Jiayang Li

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05639 2025-11-11 cs.CL 57%

ECom-Bench: Can LLM Agent Resolve Real-World E-commerce Customer Support Issues?

Haoxin Wang, Xianhan Peng, Xucheng Huang, Yizhe Huang, Ming Gong, Chenghan Yang, Yang Liu, Ling Jiang

机构 * Xiaoduo AI Lab(小多人工智能实验室) Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL

Comments Accepted as a main conference paper at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05912 2025-11-11 eess.SP 50%

RadioSim Agent: Combining Large Language Models and Deterministic EM Simulators for Interactive Radio Map Analysis

Sajjad Hussain, Conor Brennan

专题命中 多模态Agent :multimodal(abstract)

Comments Submitted to EuCAP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05723 2025-11-11 cs.RO 50%

TumorMap: A Laser-based Surgical Platform for 3D Tumor Mapping and Fully-Automated Tumor Resection

Guangshen Ma, Ravi Prakash, Beatrice Schleupner, Jeffrey Everitt, Arpit Mishra, Junqin Chen, Brian Mann, Boyuan Chen, Leila Bridgeman, Pei Zhong, Mark Draelos, William C. Eward, Patrick J. Codd

机构 * Thomas Lord Department of Mechanical Engineering and Materials Science, Duke University(杜克大学机械工程与材料科学系) Department of Robotics, University of Michigan, Ann Arbor(密歇根大学机器人学系) Department of Orthopaedic Surgery, School of Medicine, Duke University(杜克大学骨科手术系) Department of Pathology, School of Medicine, Duke University(杜克大学病理学系) Department of Ophthalmology and Visual Sciences, University of Michigan Medical School, Ann Arbor(密歇根大学医学学院眼科与视觉科学系) Department of Neurosurgery, School of Medicine, Duke University(杜克大学神经外科系)

专题命中 多模态Agent :multimodal(abstract)

Comments 41 pages, 25 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.15922 2025-11-11 cs.LG cs.RO 50%

The Dark Side of Rich Rewards: Understanding and Mitigating Noise in VLM Rewards

Sukai Huang, Shu-Wei Liu, Nir Lipovetzky, Trevor Cohn

机构 * Google DeepMind(谷歌DeepMind)

专题命中 多模态Agent :multimodal(abstract)

Comments accepted by PRL Workshop Series @ ICAPS 2025. 11 main body pages, 21 appendix pages

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态训练与对齐 16 篇

2503.18135 2025-11-11 cs.CV 88%

MLLM-For3D: Adapting Multimodal Large Language Model for 3D Reasoning Segmentation

Jiaxin Huang, Runnan Chen, Ziwen Li, Zhengqing Gao, Xiao He, Yandong Guo, Mingming Gong, Tongliang Liu

机构 * MBZUAI The University of Sydney(悉尼大学) The University of Melbourne(墨尔本大学) AI2Robotic

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(title,abstract);分类 cs.CV

Comments 10 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07274 2025-11-11 cs.LG 82%

Multi-modal Dynamic Proxy Learning for Personalized Multiple Clustering

Jinfeng Xu, Zheyu Chen, Shuo Yang, Jinze Li, Ziyue Peng, Zewei Liu, Hewei Wang, Jiayi Zhang, Edith C. H. Ngai

专题命中 多模态训练与对齐 :multi-modal(title,abstract);cross-modal(abstract)

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06805 2025-11-11 cs.AI cs.LG 79%

MathSE: Improving Multimodal Mathematical Reasoning via Self-Evolving Iterative Reflection and Reward-Guided Fine-Tuning

Jinhao Chen, Zhen Yang, Jianxin Shi, Tianyu Wo, Jie Tang

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.AI

Comments 19 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06593 2025-11-11 cs.CV 79%

Spatial-Frequency Enhanced Mamba for Multi-Modal Image Fusion

Hui Sun, Long Lv, Pingping Zhang, Tongdan Tang, Feng Tian, Weibing Sun, Huchuan Lu

机构 * School of Future Technology, School of Artificial Intelligence, Dalian University of Technology(大连理工大学未来技术学院、人工智能学院) Affiliated Zhongshan Hospital of Dalian University(大连大学附属中山医院) Central Hospital of Dalian University of Technology(大连理工大学中心医院) School of Information and Communication Engineering, Dalian University of Technology(大连理工大学信息与通信工程学院)

专题命中 多模态训练与对齐 :multi-modal(title,abstract);分类 cs.CV

Comments This work is accepted by IEEE Transactions on Image Processing. More modifications may be performed

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06723 2025-11-11 cs.LG 78%

Multi-Modal Continual Learning via Cross-Modality Adapters and Representation Alignment with Knowledge Preservation

Evelyn Chee, Wynne Hsu, Mong Li Lee

机构 * School of Computing, National University of Singapore(computing学院,新加坡国立大学)

专题命中 多模态训练与对齐 :multi-modal(title,abstract)

Comments Accepted to ECAI 2025

Journal ref 28th European Conference on Artificial Intelligence (ECAI), 2025, pp.1083-1090

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05716 2025-11-11 cs.LG 78%

Distributionally Robust Multimodal Machine Learning

Peilin Yang, Yu Ma

机构 * University of Wisconsin, Madison(威斯康星大学麦迪逊分校)

专题命中 多模态训练与对齐 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05553 2025-11-11 cs.CV cs.AI 73%

EVLP:Learning Unified Embodied Vision-Language Planner with Reinforced Supervised Fine-Tuning

Xinyan Cai, Shiguang Wu, Dafeng Chi, Yuzheng Zhuang, Xingyue Quan, Jianye Hao, Qiang Guan

机构 * Institute of Automation, Chinese Academy of Sciences (CASIA)(中国科学院自动化研究所)

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏