arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-10 至 2025-10-10 共收录 58 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 16 篇

2504.01444 2025-10-10 cs.CR cs.AI 79%

PiCo: Jailbreaking Multimodal Large Language Models via Pictorial Code Contextualization

Aofan Liu, Lulu Tang, Ting Pan, Yuguo Yin, Bin Wang, Ao Yang

机构 * School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院) School of Electronic and Computer Engineering, Peking University(北京大学电子与计算机工程学院) Beijing Academy of Artificial Intelligence(北京人工智能研究院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Comments Accepted to IEEE International Conference on Multimedia and Expo (ICME) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07325 2025-10-10 cs.LG cs.NE 78%

A Modality-Aware Cooperative Co-Evolutionary Framework for Multimodal Graph Neural Architecture Search

Sixuan Wang, Jiao Yin, Jinli Cao, Mingjian Tang, Yong-Feng Ge

机构 * Department of Computer Science and Information Technology, La Trobe University(计算机科学与信息技术系,拉特罗布大学) Institute for Sustainable Industries and Liveable Cities, Victoria University(可持续产业与宜居城市研究所,维多利亚大学)

专题命中 多模态评测 :multimodal(title,abstract)

Comments 11 pages, 6 figures. This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08011 2025-10-10 cs.CV cs.CL 73%

Play to Generalize: Learning to Reason Through Game Play

Yunfei Xie, Yinsong Ma, Shiyi Lan, Alan Yuille, Junfei Xiao, Chen Wei

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.CL

Comments Project Page: https://yunfeixie233.github.io/ViGaL/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16746 2025-10-10 cs.CV cs.CL cs.LG 66%

Zebra-CoT: A Dataset for Interleaved Vision Language Reasoning

Ang Li, Charles Wang, Deqing Fu, Kaiyu Yue, Zikui Cai, Wang Bill Zhu, Ollie Liu, Peng Guo, Willie Neiswanger, Furong Huang, Tom Goldstein, Micah Goldblum

机构 * Columbia University(哥伦比亚大学) University of Maryland(马里兰大学) University of Southern California(南加州大学) New York University(纽约大学)

专题命中 多模态评测 :multimodal(abstract,comments);分类 cs.CV、cs.CL

Comments dataset link: https://huggingface.co/datasets/multimodal-reasoning-lab/Zebra-CoT

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08202 2025-10-10 cs.HC cs.AI cs.CL cs.ET 62%

Sentiment Matters: An Analysis of 200 Human-SAV Interactions

Lirui Guo, Michael G. Burke, Wynita M. Griggs

机构 * Department of Civil and Environmental Engineering, Monash University(莫纳什大学土木与环境工程系) Department of Electrical and Computer Systems Engineering, Monash University(莫纳什大学电子与计算机系统工程系)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL、cs.AI

Comments Accepted for presentation at IEEE ITSC 2025 and for publication in its Proceedings. \c{opyright} 2025 IEEE. Personal use permitted; other uses require permission from IEEE, including reprinting, republishing, or reuse of any copyrighted component of this work

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13546 2025-10-10 cs.CV 57%

GazeProphet: Software-Only Gaze Prediction for VR Foveated Rendering

Farhaan Ebadulla, Chiraag Mudlapur, Gaurav BV

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV

Comments 8 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08022 2025-10-10 cs.RO cs.AI 57%

FastUMI-100K: Advancing Data-driven Robotic Manipulation with a Large-scale UMI-style Dataset

Kehui Liu, Zhongjie Jia, Yang Li, Zhaxizhuoma, Pengan Chen, Song Liu, Xin Liu, Pingrui Zhang, Haoming Song, Xinyi Ye, Nieqing Cao, Zhigang Wang, Jia Zeng, Dong Wang, Yan Ding, Bin Zhao, Xuelong Li

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Northwestern Polytechnical University(西北工业大学) Shanghai Jiao Tong University(上海交通大学) TongJi University(同济大学) Xi’an Jiaotong-Liverpool University(西安交通大学-利物浦大学) Suzhou OneStar Robotics Corp Ltd(苏州OneStar机器人有限公司) Institute of Artificial Intelligence, China Telecom Corp Ltd(中国电信人工智能研究院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04097 2025-10-10 cs.AI 57%

WebRenderBench: Enhancing Web Interface Generation through Layout-Style Consistency and Reinforcement Learning

Peichao Lai, Jinhui Zhuang, Kexuan Zhang, Ningchang Xiong, Shengjie Wang, Yanwei Xu, Chong Chen, Yilei Wang, Bin Cui

机构 * Peking University(北京大学) Xiamen Huaxia University(厦门华夏大学) Fuzhou University(福州市大学) City University of Hong Kong(香港城市大学) Huawei Cloud BU(华为云业务单元)

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01571 2025-10-10 cs.CV 57%

PainFormer: a Vision Foundation Model for Automatic Pain Assessment

Stefanos Gkikas, Raul Fernandez Rojas, Manolis Tsiknakis

机构 * Hellenic Mediterranean University, Department of Electrical and Computer Engineering(希腊地中海大学电子与计算机工程系) Institute of Computer Science, Foundation for Research & Technology-Hellas(希腊研究所计算机科学研究所) University of Canberra, Faculty of Science and Technology(堪培拉大学科学与技术学院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Journal ref IEEE Transactions on Affective Computing; 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08106 2025-10-10 cs.RO 50%

Beyond hospital reach: Autonomous lightweight ultrasound robot for liver sonography

Zihan Li, Yixiao Xu, Lei Zhang, Taiyu Han, Xinshan Yang, Yingni Wang, Mingxuan Liu, Shenghai Xin, Linxun Liu, Hongen Liao, Guochen Ning

专题命中 多模态评测 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态Agent 3 篇

2508.09736 2025-10-10 cs.CV 83%

Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term Memory

Lin Long, Yichen He, Wentao Ye, Yiyuan Pan, Yuan Lin, Hang Li, Junbo Zhao, Wei Li

机构 * Zhejiang University(浙江大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态Agent :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07709 2025-10-10 cs.AI cs.CL cs.CY cs.MA 81%

Multimodal Safety Evaluation in Generative Agent Social Simulations

Alhim Vera, Karen Sanchez, Carlos Hinojosa, Haidar Bin Hamid, Donghoon Kim, Bernard Ghanem

机构 * University of Cincinnati(辛辛那提大学) King Abdullah University of Science and Technology(国王阿卜杜勒·阿齐兹大学科学与技术)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06223 2025-10-10 cs.HC cs.AI 79%

A Multimodal GUI Architecture for Interfacing with LLM-Based Conversational Assistants

Hans G. W. van Dam

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI

Comments 24 pages, 19 figures, code available at https://github.com/hansvdam/langbar

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态训练与对齐 9 篇

2510.07326 2025-10-10 cs.MM cs.SD 83%

Audio-Visual Separation with Hierarchical Fusion and Representation Alignment

Han Hu, Dongheng Lin, Qiming Huang, Yuqi Hou, Hyung Jin Chang, Jianbo Jiao

专题命中 多模态训练与对齐 :audio-visual(title,abstract);multimodal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08022 2025-10-10 cs.LG cs.AI cs.CL cs.CV 82%

Modality-Balancing Preference Optimization of Large Multimodal Models by Adversarial Negative Mining

Chenxi Liu, Tianyi Xiong, Yanshuo Chen, Ruibo Chen, Yihan Wu, Junfeng Guo, Tianyi Zhou, Heng Huang

机构 * University of Maryland, College Park(马里兰大学学院 park)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07513 2025-10-10 cs.LG cs.AI cs.CV cs.DB 81%

MLLM4TS: Leveraging Vision and Multimodal Language Models for General Time-Series Analysis

Qinghua Liu, Sam Heshmati, Zheda Mai, Zubin Abraham, John Paparrizos, Liu Ren

机构 * The Ohio State University(俄亥俄州立大学) Bosch Research North America(博世北美研究)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08457 2025-10-10 cs.CL 79%

ARES: Multimodal Adaptive Reasoning via Difficulty-Aware Token-Level Entropy Shaping

Shuang Chen, Yue Guo, Yimeng Ye, Shijue Huang, Wenbo Hu, Haoxi Li, Manyuan Zhang, Jiayu Chen, Song Guo, Nanyun Peng

机构 * University of California, Los Angeles(加州大学洛杉矶分校) The Hong Kong University of Science and Technology(香港科技大学) Columbia University(哥伦比亚大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05674 2025-10-10 cs.CV 70%

Context Matters: Learning Global Semantics via Object-Centric Representation

Jike Zhong, Yuxiang Lai, Xiaofeng Yang, Konstantinos Psounis

机构 * University of Southern California(美国南加州大学) Emory University(埃默里大学)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03321 2025-10-10 cs.CV 70%

Empowering Lightweight MLLMs with Reasoning via Long CoT SFT

Linyu Ou, YuYang Yin

机构 * Beijing Institute of Technology(北京理工大学) Basic Algorithm Center, PCG, Tencent(腾讯基本算法中心)

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08470 2025-10-10 cs.AI cs.CL cs.LG 62%

Looking to Learn: Token-wise Dynamic Gating for Low-Resource Vision-Language Modelling

Bianca-Mihaela Ganescu, Suchir Salhan, Andrew Caines, Paula Buttery

机构 * ALTA Institute(ALTA研究院) Department of Computer Science & Technology, University of Cambridge(计算机科学与技术系,剑桥大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CL、cs.AI

Comments Accepted to the EMNLP 2025 BabyLM Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07910 2025-10-10 cs.LG cs.AI cs.CV 62%

MMM: Quantum-Chemical Molecular Representation Learning for Combinatorial Drug Recommendation

Chongmyung Kwon, Yujin Kim, Seoeun Park, Yunji Lee, Charmgil Hong

机构 * Handong Global University(Handong Global大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Medical Image Computing and Computer-Assisted Intervention (MICCAI) Predictive Intelligence in Medicine Workshop (MICCAI PRIME) 2025; 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23663 2025-10-10 cs.CV 57%

HIVTP: A Training-Free Method to Improve VLMs Efficiency via Hierarchical Visual Token Pruning Using Middle-Layer-Based Importance Score

Jingqi Xu, Jingxi Lu, Chenghao Li, Sreetama Sarkar, Peter A. Beerel

机构 * University of Southern California(美国南加州大学)

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 其他多模态 6 篇

2510.08565 2025-10-10 cs.CV 83%

NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints

Changyao Tian, Hao Li, Gen Luo, Xizhou Zhu, Weijie Su, Hanming Deng, Jinguo Zhu, Jie Shao, Ziran Zhu, Yunpeng Liu, Lewei Lu, Wenhai Wang, Hongsheng Li, Jifeng Dai

机构 * Shanghai AI Laboratory(上海人工智能实验室) The Chinese University of Hong Kong(香港中文大学) Tsinghua University(清华大学) Sensetime Research(商汤科技研究院) Nanjing University(南京大学)

专题命中 其他多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025. 22 pages, link: https://github.com/OpenGVLab/NaViL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07684 2025-10-10 astro-ph.GA 78%

Multi-modal Foundation Model for Cosmological Simulation Data

Bin Xia, Nesar Ramachandra, Azton I. Wells, Salman Habib, John Wise

专题命中 其他多模态 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07562 2025-10-10 cs.LG 78%

EBGAN-MDN: An Energy-Based Adversarial Framework for Multi-Modal Behavior Cloning

Yixiao Li, Julia Barth, Thomas Kiefer, Ahmad Fraij

专题命中 其他多模态 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07509 2025-10-10 cs.LG cs.IT math.IT 78%

Efficient Generalization via Multimodal Co-Training under Data Scarcity and Distribution Shift

Tianyu Bell Pan, Damon L. Woodard

机构 * Department of Electrical and Computer Engineering(电气与计算机工程系) Florida Institute for National Security (FINS)(佛罗里达国家安全研究所) Applied Artificial Intelligence Group(应用人工智能小组) University of Florida(佛罗里达大学)

专题命中 其他多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08186 2025-10-10 cond-mat.mtrl-sci 71%

Multimodal Topological Textures Arising from Coupled Structural Orders in SrTiO$_3$

Fernando Gómez-Ortiz, Louis Bastogne, Philippe Ghosez

专题命中 其他多模态 :multimodal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05938 2025-10-10 cond-mat.mtrl-sci 50%

Autonomous interpretation of atomistic scattering data

Andy S. Anker, John L. A. Gardner, Louise A. M. Rosset, Andrew L. Goodwin, Volker L. Deringer

专题命中 其他多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏