arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 9205 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 9205 篇

2501.19339 2025-10-23 cs.CV cs.CL 73%

PixelWorld: How Far Are We from Perceiving Everything as Pixels?

Zhiheng Lyu, Xueguang Ma, Wenhu Chen

机构 * University of Waterloo(滑铁卢大学) Vector Institute, Toronto(多伦多向量研究所)

专题命中 多模态评测 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17040 2025-10-17 cs.CV cs.AI 73%

From Easy to Hard: The MIR Benchmark for Progressive Interleaved Multi-Image Reasoning

Hang Du, Jiayang Zhang, Guoshun Nan, Wendi Deng, Zhenyan Chen, Chenyang Zhang, Wang Xiao, Shan Huang, Yuqi Pan, Tao Qi, Sicong Leng

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Nanyang Technological University(南洋理工大学)

专题命中 多模态评测 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08011 2025-10-10 cs.CV cs.CL 73%

Play to Generalize: Learning to Reason Through Game Play

Yunfei Xie, Yinsong Ma, Shiyi Lan, Alan Yuille, Junfei Xiao, Chen Wei

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.CL

Comments Project Page: https://yunfeixie233.github.io/ViGaL/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20024 2025-09-23 cs.CV cs.AI cs.RO 73%

ReasonPlan: Unified Scene Prediction and Decision Reasoning for Closed-loop Autonomous Driving

Xueyi Liu, Zuodong Zhong, Yuxin Guo, Yun-Fu Liu, Zhiguo Su, Qichao Zhang, Junli Wang, Yinfeng Gao, Yupeng Zheng, Qiao Lin, Huiyong Chen, Dongbin Zhao

机构 * SKL-MAIS, Institute of Automation, Chinese Academy of Sciences, Beijing, China(中国科学院自动化研究所SKL-MAIS部门,北京) School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China(中国科学院大学人工智能学院,北京) EACON, Fujian, China(福建EACON机构,中国) School of Automation and Electrical Engineering, University of Science and Technology Beijing, Beijing, China(北京科技大学自动化与电气工程学院)

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments 18 pages; 9 figures; https://github.com/Liuxueyi/ReasonPlan

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04077 2025-09-12 cs.CL cs.SD eess.AS 73%

A Novel Data Augmentation Approach for Automatic Speaking Assessment on Opinion Expressions

Chung-Chun Wang, Jhen-Ke Lin, Hao-Chien Lu, Hong-Yun Lin, Berlin Chen

机构 * National Taiwan Normal University(台湾国立台湾师范大学)

专题命中 多模态评测 :multimodal(abstract);cross-modal(abstract);分类 cs.CL、eess.AS

Comments submitted to the ISCA SLaTE-2025 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.16822 2025-09-11 cs.CV cs.AI 73%

Integrating Clinical Knowledge Graphs and Gradient-Based Neural Systems for Enhanced Melanoma Diagnosis via the 7-Point Checklist

Yuheng Wang, Tianze Yu, Jiayue Cai, Sunil Kalia, Harvey Lui, Z. Jane Wang, Tim K. Lee

机构 * The University of British Columbia(不列颠哥伦比亚大学) Vancouver Coastal Health Research Institute(温哥华海岸健康研究机构) Department of Dermatology and Skin Science(皮肤科与皮肤病学系) Photomedicine Institute(光医学研究所) Centre for Clinical Epidemiology and Evaluation(临床流行病学与评估中心) School of Biomedical Engineering(生物医学工程学院) Department of Electrical and Computer Engineering(电气与计算机工程系) Department of Population Health Sciences(人群健康科学系) BC Cancer(不列颠哥伦比亚省癌症中心) BC Children’s Hospital Research Institute(不列颠哥伦比亚儿童医院研究机构) Guangdong Key Laboratory of Biomedical Measurements and Ultrasound Imaging(广东生物医学测量与超声成像重点实验室) Shenzhen University Medical School(深圳大学医学院) Shenzhen University(深圳大学)

专题命中 多模态评测 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments The paper was officially accepted for publication in IEEE Transactions on Neural Networks and Learning Systems in August 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21113 2025-09-03 cs.CV cs.AI cs.LG 73%

R-4B: Incentivizing General-Purpose Auto-Thinking Capability in MLLMs via Bi-Mode Annealing and Reinforce Learning

Qi Yang, Bolin Ni, Shiming Xiang, Han Hu, Houwen Peng, Jie Jiang

机构 * Tencent Hunyuan Team(腾讯文言团队) Institute of Automation, CAS(中国科学院自动化研究所)

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments 20 pages, 14 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11738 2025-08-19 cs.CY cs.AI cs.CV 73%

Artificial Intelligence in Rural Healthcare Delivery: Bridging Gaps and Enhancing Equity through Innovation

Kiruthika Balakrishnan, Durgadevi Velusamy, Hana E. Hinkle, Zhi Li, Karthikeyan Ramasamy, Hikmat Khan, Srini Ramaswamy, Pir Masoom Shah

机构 * University of Illinois College of Medicine Rockford(伊利诺伊大学罗克福德医学院) National Center for Rural Health Professions(农村健康专业中心) Sri Sivasubramaniya Nadar College of Engineering(塞维萨布拉曼尼亚那德尔工程学院) M. Kumarasamy College of Engineering(M. Kumarasamy 工程学院) The Ohio State University Wexner Medical Center(俄亥俄州立大学韦克斯纳医疗中心) iWorks Corporation(iWorks 公司) Central South University(中南大学) Bacha Khan University(巴哈汗大学)

专题命中 多模态评测 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.18057 2025-08-18 cs.CV cs.CL 73%

CLEAR: Character Unlearning in Textual and Visual Modalities

Alexey Dontsov, Dmitrii Korzh, Alexey Zhavoronkin, Boris Mikheev, Denis Bobkov, Aibek Alanov, Oleg Y. Rogov, Ivan Oseledets, Elena Tutubalina

机构 * AIRI HSE University(俄罗斯高等经济大学) Skoltech(斯克里普契克技术大学) Sber AI(俄罗斯储蓄银行人工智能实验室) MTUCI(莫斯科国立大学) MIPT(米哈伊尔戈尔斯基理工学院) ISP RAS Research Center for Trusted AI(俄罗斯科学院信息与系统问题研究所可信人工智能研究中心)

专题命中 多模态评测 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL

Journal ref https://aclanthology.org/2025.findings-acl.1058/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09593 2025-08-14 cs.CV cs.AI 73%

Hierarchical Brain Structure Modeling for Predicting Genotype of Glioma

Haotian Tang, Jianwei Chen, Xinrui Tang, Yunjia Wu, Zhengyang Miao, Chao Li

机构 * College of Medicine and Biological Information Engineering, Northeastern University, Liaoning, China(医学与生物信息工程学院,东北大学,辽宁,中国) School of Science and Engineering, University of Dundee, Dundee, UK(科学与工程学院,邓迪大学,邓迪,英国) School of Medicine, University of Dundee, Dundee, UK(医学院,邓迪大学,邓迪,英国) Department of Applied Maths and Theoretical Physics, University of Cambridge, Cambridge, UK(应用数学与理论物理系,剑桥大学,剑桥,英国)

专题命中 多模态评测 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.02255 2025-08-08 eess.AS cs.LG cs.MM cs.SD 73%

MidiCaps: A large-scale MIDI dataset with text captions

Jan Melechovsky, Abhinaba Roy, Dorien Herremans

专题命中 多模态评测 :multi-modal(abstract);cross-modal(abstract);分类 cs.MM、eess.AS

Comments Accepted in ISMIR2024

Journal ref Proceedings of ISMIR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03958 2025-08-07 cs.IR cs.AI cs.CL cs.LG 73%

A Comparative Study of Specialized LLMs as Dense Retrievers

Hengran Zhang, Keping Bi, Jiafeng Guo

机构 * Key Laboratory of Network Data Science and Technology(网络数据科学与技术重点实验室) Institute of Computing Technology(计算技术研究所) Chinese Academy of Sciences(中国科学院) State Key Laboratory of AI Safety(人工智能安全国家重点实验室) University of Chinese Academy of Science(中国科学院大学)

专题命中 多模态评测 :multimodal(abstract);cross-modal(abstract);分类 cs.CL、cs.AI

Comments Accepted by CCIR25 and published by Springer LNCS or LNAI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04757 2025-07-25 cs.CV cs.CL 73%

ELITE: Enhanced Language-Image Toxicity Evaluation for Safety

Wonjun Lee, Doehyeon Lee, Eugene Choi, Sangyoon Yu, Ashkan Yousefpour, Haon Park, Bumsub Ham, Suhyun Kim

机构 * Yonsei University(延世大学) Kyung Hee University(庆熙大学) Korea Institute of Science(韩国科学研究院) Seoul National University(首尔国立大学) Sookmyung Women's University(_sookmyung女子大学)

专题命中 多模态评测 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.CL

Comments ICML 2025. Project page at https://velpegor.github.io/ELITE/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12739 2025-07-18 cs.CV cs.AI 73%

Transformer-based Spatial Grounding: A Comprehensive Survey

Ijazul Haq, Muhammad Saqib, Yingjie Zhang

机构 * School of Intelligent Manufacturing, South China University of Technology(华南理工大学智能制造学院) Department of Software Engineering, University of Engineering & Technology(工程大学软件工程系)

专题命中 多模态评测 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01029 2025-07-03 cs.LG cs.AI cs.CL 73%

PathCoT: Chain-of-Thought Prompting for Zero-shot Pathology Visual Reasoning

Junjie Zhou, Yingli Zuo, Shichang Feng, Peng Wan, Qi Zhu, Daoqiang Zhang, Wei Shao

机构 * The College of Artificial Intelligence, Nanjing University of Aeronautics(人工智能学院,南京航空航天大学) Astronautics The Key Laboratory of Brain-Machine Intelligence Technology, Ministry of Education(航空航天脑机智能技术重点实验室,教育部)

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00891 2025-07-02 cs.CL cs.AI 73%

MemeCMD: An Automatically Generated Chinese Multi-turn Dialogue Dataset with Contextually Retrieved Memes

Yuheng Wang, Xianhe Tang, Pufeng Huang

机构 * Wuhan University(武汉大学)

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17931 2025-06-24 cs.CV cs.AI cs.LG 73%

IDAL: Improved Domain Adaptive Learning for Natural Images Dataset

Ravi Kant Gupta, Shounak Das, Amit Sethi

机构 * Indian Institute of Technology Bombay(印度理工学院班加罗尔)

专题命中 多模态评测 :multimodal(abstract);multi-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted in ICPR'24 (International Conference on Pattern Recognition)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11394 2025-06-16 cs.CV cs.AI 73%

Dynamic Double Space Tower

Weikai Sun, Shijie Song, Han Wang

机构 * Huawei Cloud Computing(华为云计算)

专题命中 多模态评测 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24182 2025-06-02 cs.CV cs.AI 73%

Seeing is Not Reasoning: MVPBench for Graph-based Evaluation of Multi-path Visual Physical CoT

Zhuobai Dong, Junchao Yi, Ziyuan Zheng, Haochen Han, Xiangxi Zheng, Alex Jinpeng Wang, Fangming Liu, Linjie Li

机构 * Central South University(中南大学) University of Electronic Science and Technology of China(电子科技大学) Peng Cheng Laboratory(鹏城实验室) Nanjing University(南京大学) Microsoft(微软公司)

专题命中 多模态评测 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23727 2025-05-30 cs.CV cs.MM 73%

PixelThink: Towards Efficient Chain-of-Pixel Reasoning

Song Wang, Gongfan Fang, Lingdong Kong, Xiangtai Li, Jianyun Xu, Sheng Yang, Qiang Li, Jianke Zhu, Xinchao Wang

机构 * Zhejiang University(浙江大学) National University of Singapore(新加坡国立大学) Nanyang Technological University(南洋理工大学) AD Lab, CaiNiao Inc., Alibaba Group(阿里集团 Cainiao 事业部实验室)

专题命中 多模态评测 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.MM

Comments Project Page: https://PixelThink.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18291 2025-05-28 cs.CV cs.CL cs.RO 73%

InstructPart: Task-Oriented Part Segmentation with Instruction Reasoning

Zifu Wan, Yaqi Xie, Ce Zhang, Zhiqiu Lin, Zihan Wang, Simon Stepputtis, Deva Ramanan, Katia Sycara

机构 * Robotics Institute, Carnegie Mellon University(机器人研究所,卡内基梅隆大学)

专题命中 多模态评测 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV、cs.CL

Comments Accepted by ACL 2025 Main. Project page: https://zifuwan.github.io/InstructPart/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18881 2025-05-27 cs.CV cs.AI cs.RO 73%

SD-OVON: A Semantics-aware Dataset and Benchmark Generation Pipeline for Open-Vocabulary Object Navigation in Dynamic Scenes

Dicong Qiu, Jiadi You, Zeying Gong, Ronghe Qiu, Hui Xiong, Junwei Liang

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 多模态评测 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV、cs.AI

Comments Preprint. 21 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15579 2025-04-29 cs.CV cs.CL 73%

An Explainable Biomedical Foundation Model via Large-Scale Concept-Enhanced Vision-Language Pre-training

Yuxiang Nie, Sunan He, Yequan Bie, Yihui Wang, Zhixuan Chen, Shu Yang, Zhiyuan Cai, Hongmei Wang, Xi Wang, Luyang Luo, Mingxiang Wu, Xian Wu, Ronald Cheong Kin Chan, Yuk Ming Lau, Yefeng Zheng, Pranav Rajpurkar, Hao Chen

机构 * Shenzhen People’s Hospital(深圳人民医院) The Chinese University of Hong Kong(香港中文大学)

专题命中 多模态评测 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.12966 2025-04-18 cs.CV cs.AI 73%

Look Before You Decide: Prompting Active Deduction of MLLMs for Assumptive Reasoning

Yian Li, Wentao Tian, Yang Jiao, Jingjing Chen, Tianwen Qian, Bin Zhu, Na Zhao, Yu-Gang Jiang

机构 * Fudan University(复旦大学) East China Normal University(华东师范大学) Singapore Management University(新加坡管理大学) Singapore University of Technology and Design(新加坡科技设计大学)

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10471 2025-04-15 cs.CV cs.CL 73%

MIEB: Massive Image Embedding Benchmark

Chenghao Xiao, Isaac Chung, Imene Kerboua, Jamie Stirling, Xin Zhang, Márton Kardos, Roman Solomatin, Noura Al Moubayed, Kenneth Enevoldsen, Niklas Muennighoff

机构 * Durham University(杜伦大学) Zendesk(赞德斯克公司) Esker(艾斯克尔公司) INSA Lyon(里昂国立应用科学学院) The Hong Kong Polytechnic University(香港理工大学) Aarhus University(奥胡斯大学) ITMO University(ITMO大学) Contextual AI(康泰克斯特人工智能公司) Stanford University(斯坦福大学)

专题命中 多模态评测 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.11116 2025-04-15 cs.CV cs.AI 73%

PhD: A ChatGPT-Prompted Visual hallucination Evaluation Dataset

Jiazhen Liu, Yuhan Fu, Ruobing Xie, Runquan Xie, Xingwu Sun, Fengzong Lian, Zhanhui Kang, Xirong Li

机构 * Renmin University of China(中国人民大学) Machine Learning Platform Department, Tencent(腾讯机器学习平台部) HKUST(香港科技大学)

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments Accepted by CVPR 2025, Highlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11089 2025-03-17 cs.RO cs.AI cs.CV 73%

EmbodiedVSR: Dynamic Scene Graph-Guided Chain-of-Thought Reasoning for Visual Spatial Tasks

Yi Zhang, Qiang Zhang, Xiaozhu Ju, Zhaoyang Liu, Jilei Mao, Jingkai Sun, Jintao Wu, Shixiong Gao, Shihan Cai, Zhiyuan Qin, Linkai Liang, Jiaxu Wang, Yiqun Duan, Jiahang Cao, Renjing Xu, Jian Tang

机构 * Beijing Innovation Center of Humanoid Robotics(北京人形机器人创新中心) Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Hong Kong University of Science and Technology(香港科技大学) University of Technology Sydney(悉尼科技大学)

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments technical report

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07587 2025-03-11 cs.CV cs.AI cs.RO 73%

Robusto-1 Dataset: Comparing Humans and VLMs on real out-of-distribution Autonomous Driving VQA from Peru

Dunant Cusipuma, David Ortega, Victor Flores-Benites, Arturo Deza

机构 * Artificio

专题命中 多模态评测 :multimodal(abstract);multi-modal(abstract);分类 cs.CV、cs.AI

Comments A pre-print. 26 pages. Link to Code + Data: https://huggingface.co/datasets/Artificio/robusto-1

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.14973 2025-03-07 cs.CL cs.AI cs.LG 73%

GenCeption: Evaluate Vision LLMs with Unlabeled Unimodal Data

Lele Cao, Valentin Buchner, Zineb Senane, Fangkai Yang

机构 * Microsoft Gaming (ABK)(微软游戏(ABK)) KTH Royal Institute of Technology(皇家理工学院(KTH))

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);分类 cs.CL、cs.AI

Comments Published by Computer Speech & Language (https://doi.org/10.1016/j.csl.2025.101785). Source code and Leaderboard: https://github.com/llcresearch/GenCeption

Journal ref Computer Speech & Language 93 (2025) 101785

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.18967 2025-03-03 cs.CV cs.CL cs.LG 73%

Ferret-UI 2: Mastering Universal User Interface Understanding Across Platforms

Zhangheng Li, Keen You, Haotian Zhang, Di Feng, Harsh Agrawal, Xiujun Li, Mohana Prasad Sathya Moorthy, Jeff Nichols, Yinfei Yang, Zhe Gan

机构 * Apple(苹果公司)

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.CL

Comments Accepted to ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏