arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-07-29 至 2025-07-29 共收录 105 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 16 篇

2503.03492 2025-07-29 cs.CV 57%

Find First, Track Next: Decoupling Identification and Propagation in Referring Video Object Segmentation

Suhwan Cho, Seunghoon Lee, Minhyeok Lee, Jungho Lee, Sangyoun Lee

机构 * GenGenAI Yonsei University(延世大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments ICCVW 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21016 2025-07-29 cs.LG q-bio.NC 50%

Predicting Cognition from fMRI:A Comparative Study of Graph, Transformer, and Kernel Models Across Task and Rest Conditions

Jagruti Patel, Mikkel Schöttner, Thomas A. W. Bolton, Patric Hagmann

专题命中 视频多模态 :multimodal(abstract)

Comments Preliminary version; a revised version will be uploaded later

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20866 2025-07-29 physics.comp-ph 50%

Neuromorphic Photonic Processing and Memory with Spiking Resonant Tunnelling Diode Neurons and Neural Networks

Dafydd Owen-Newns, Joshua Robertson, Giovanni Donati, Jose Figueiredo, Edward Wasige, Kathy Ludge, Bruno Romeira, Antonio Hurtado

专题命中 视频多模态 :multi-modal(abstract)

Comments 19 pages, 11 figures, submitted to Advanced Intelligent Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19566 2025-07-29 eess.IV 50%

SLENet: A Novel Multiscale CNN-Based Network for Detecting the Rats Estrous Cycle

Qinyang Wang, Hoileong Lee, Xiaodi Pu, Yuanming Lai, Yiming Ma

专题命中 视频多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 跨模态检索 6 篇

2504.14348 2025-07-29 cs.CV 88%

Manipulating Multimodal Agents via Cross-Modal Prompt Injection

Le Wang, Zonghao Ying, Tianyuan Zhang, Siyuan Liang, Shengshan Hu, Mingchuan Zhang, Aishan Liu, Xianglong Liu

机构 * Beihang University(北洋大学) National University of Singapore(新加坡国立大学) Huazhong University of Science and Technology(华中科技大学) Henan University of Science and Technology(河南科技大学)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.CV

Comments 16 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20326 2025-07-29 cs.LG cs.AI 83%

MIPS: a Multimodal Infinite Polymer Sequence Pre-training Framework for Polymer Property Prediction

Jiaxi Wang, Yaosen Min, Xun Zhu, Miao Li, Ji Wu

机构 * Department of Electronic Engineering, Tsinghua University Beijing China Zhongguancun Institute of Artificial Intelligence Beijing China Department of Electronic Engineering \& College of AI, Tsinghua University Beijing National Research Center for Information Science Department of Electronic Engineering, Tsinghua University Zhongguancun Institute of Artificial Intelligence

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

Comments 14 pages, 8 figures, accepted by ACM Multimedia 2025 (oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20259 2025-07-29 cs.CV 79%

L-MCAT: Unpaired Multimodal Transformer with Contrastive Attention for Label-Efficient Satellite Image Classification

Mitul Goswami, Mrinal Goswami

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20189 2025-07-29 eess.SP cs.AI cs.LG q-bio.NC 79%

NeuroCLIP: A Multimodal Contrastive Learning Method for rTMS-treated Methamphetamine Addiction Analysis

Chengkai Wang, Di Wu, Yunsheng Liao, Wenyao Zheng, Ziyi Zeng, Xurong Gao, Hemmings Wu, Zhoule Zhu, Jie Yang, Lihua Zhong, Weiwei Cheng, Yun-Hsuan Chen, Mohamad Sawan

机构 * CenBRAIN Neurotech Center of Excellence, School of Engineering, Westlake University(西溪大学工程学院先进神经技术中心) School of Data Science, Xiamen University Malaysia(马来西亚厦门大学数据科学学院) Department of Neurosurgery, Second Affiliated Hospital, School of Medicine, Zhejiang University(浙江大学医学院附属第二医院神经外科) Department of Education and Correction, Zhejiang Gongchen Compulsory Isolated Detoxification Center(浙江省公检强制隔离戒毒所教育矫正部) Zhejiang Liangzhu Compulsory Isolated Detoxification Center(浙江省良渚强制隔离戒毒所)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20110 2025-07-29 cs.CV cs.AI cs.LG 62%

NeuroVoxel-LM: Language-Aligned 3D Perception via Dynamic Voxelization and Meta-Embedding

Shiyu Liu, Lianlei Shan

机构 * School of Electrical and Electronic Engineering(电气与电子工程学院) Nanyang Technological University(南洋理工大学) School of Computer Science and Technology(计算机科学与技术学院) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

Comments **14 pages, 3 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21015 2025-07-29 cs.CV 57%

Learning Transferable Facial Emotion Representations from Large-Scale Semantically Rich Captions

Licai Sun, Xingxun Jiang, Haoyu Chen, Yante Li, Zheng Lian, Biu Liu, Yuan Zong, Wenming Zheng, Jukka M. Leppänen, Guoying Zhao

机构 * University of Oulu(奥卢大学) Southeast University(东南大学) University of Turku(图尔库大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态生成 13 篇

2507.20368 2025-07-29 cs.CV cs.MM 84%

MagicAnime: A Hierarchically Annotated, Multimodal and Multitasking Dataset with Benchmarks for Cartoon Animation Generation

Shuolin Xu, Bingyuan Wang, Zeyu Cai, Fangteng Fu, Yue Ma, Tongyi Lee, Hongchuan Yu, Zeyu Wang

机构 * National Centre for Computer Animation, Bournemouth University(伯恩茅斯大学计算机动画国家中心) Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Hong Kong University of Science and Technology(香港科技大学) Department of Computer Science and Information Engineering, National Cheng Kung University(国立成功大学计算机科学与信息工程系)

专题命中 多模态生成 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.MM

Comments 8 pages,6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.12789 2025-07-29 cs.CV 83%

Efficient Physics Simulation for 3D Scenes via MLLM-Guided Gaussian Splatting

Haoyu Zhao, Hao Wang, Xingyue Zhao, Hao Fei, Hongqiu Wang, Chengjiang Long, Hua Zou

机构 * School of Computer Science, Wuhan University(武汉大学计算机学院) Wuhan National Laboratory for Optoelectronics, Huazhong University of Science and Technology(华中科技大学光电研究院) Meta Reality Lab(Meta现实实验室) Xi’an Jiao Tong University(西安交通大学) National University of Singapore(新加坡国立大学) The Department of Systems Hub, Hong Kong University of Science and Technology (Guangzhou)(香港科技大学系统枢纽部门(广州))

专题命中 多模态生成 :MLLM(title,abstract);multi-modal(abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19939 2025-07-29 cs.CV 79%

LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs

Jiaze Wang, Rui Chen, Haowang Cui

机构 * Tianjin University(天津大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02984 2025-07-29 cs.CL 79%

From Answers to Rationales: Self-Aligning Multimodal Reasoning with Answer-Oriented Chain-of-Thought

Wentao Tan, Qiong Cao, Yibing Zhan, Chao Xue, Changxing Ding

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.21291 2025-07-29 cs.CV 79%

MIGE: Mutually Enhanced Multimodal Instruction-Based Image Generation and Editing

Xueyun Tian, Wei Li, Bingbing Xu, Yige Yuan, Yuanzhuo Wang, Huawei Shen

机构 * State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences(人工智能安全国家重点实验室,计算技术研究所,中国科学院) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments This paper have been accepted by ACM MM25

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03225 2025-07-29 cs.CV 79%

MaterialPicker: Multi-Modal DiT-Based Material Generation

Xiaohe Ma, Valentin Deschaintre, Miloš Hašan, Fujun Luan, Kun Zhou, Hongzhi Wu, Yiwei Hu

机构 * State Key Lab of CAD\&CG, Zhejiang University(浙江大学CAD与CG国家重点实验室) Adobe Research(Adobe研究) State Key Lab of CAD\&CG, Zhejiang University and ZJU-FaceUnity Joint Lab of Intelligent Graphics(浙江大学CAD与CG国家重点实验室) ZJU-FaceUnity Joint Lab of Intelligent Graphics(浙大-面 unity 智能图形联合实验室)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.17046 2025-07-29 cs.CV 70%

Text-to-Image Generation Via Energy-Based CLIP

Roy Ganz, Michael Elad

机构 * Electrical Engineering Department Technion(技术学院电子工程系) Computer Science Department Technion(技术学院计算机科学系)

专题命中 多模态生成 :multimodal(abstract);image-text(abstract);分类 cs.CV

Comments Accepted to TMLR

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19492 2025-07-29 cs.HC cs.AI cs.CV 62%

ChartGen: Scaling Chart Understanding Via Code-Guided Synthetic Chart Generation

Jovana Kondic, Pengyuan Li, Dhiraj Joshi, Zexue He, Shafiq Abedin, Jennifer Sun, Ben Wiesel, Eli Schwartz, Ahmed Nassar, Bo Wu, Assaf Arbelle, Aude Oliva, Dan Gutfreund, Leonid Karlinsky, Rogerio Feris

机构 * MIT(麻省理工学院) MIT-IBM Watson AI Labs(麻省理工-IBM沃森人工智能实验室) IBM Research(IBM研究院)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.00045 2025-07-29 cs.MM cs.AI cs.LG 62%

Detecting Multimedia Generated by Large AI Models: A Survey

Li Lin, Neeraj Gupta, Yue Zhang, Hainan Ren, Chun-Hao Liu, Feng Ding, Xin Wang, Xin Li, Luisa Verdoliva, Shu Hu

机构 * Department of Computer and Information Technology, Purdue University(普渡大学计算机与信息科技系) School of Software, Nanchang University(南昌大学软件学院) Amazon Prime Video(亚马逊Prime视频) Department of Epidemiology and Biostatistics, School of Public Health(公共卫生学院流行病学与生物统计学系) Department of Computer Science, College of Nanotechnology, Science, and Engineering(纳米技术、科学与工程学院计算机科学系) University at Albany, SUNY(阿尔巴尼大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20976 2025-07-29 cs.CV 57%

Adapting Vehicle Detectors for Aerial Imagery to Unseen Domains with Weak Supervision

Xiao Fang, Minhyek Jeon, Zheyang Qin, Stanislav Panev, Celso de Melo, Shuowen Hu, Shayok Chakraborty, Fernando De la Torre

机构 * Carnegie Mellon University(卡内基梅隆大学) DEVCOM Army Research Laboratory(陆军研究实验室) Florida State University(佛罗里达州立大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19882 2025-07-29 cs.AI 57%

Causality-aligned Prompt Learning via Diffusion-based Counterfactual Generation

Xinshu Li, Ruoyu Wang, Erdun Gao, Mingming Gong, Lina Yao

机构 * The University of New South Wales(新南威尔士大学) The University of Adelaide(阿德莱德大学) The University of Melbourne(墨尔本大学) CSIRO’s Data 61(CSIRO数据61)

专题命中 多模态生成 :image-text(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21771 2025-07-29 cs.CV 57%

A Unified Image-Dense Annotation Generation Model for Underwater Scenes

Hongkai Lin, Dingkang Liang, Zhenghao Qi, Xiang Bai

机构 * Huazhong University of Science and Technology(华中科技大学)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

Comments Accepted by CVPR 2025. The code is available at https://github.com/HongkLin/TIDE

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.11824 2025-07-29 cs.CV 57%

KITTEN: A Knowledge-Intensive Evaluation of Image Generation on Visual Entities

Hsin-Ping Huang, Xinyi Wang, Yonatan Bitton, Hagai Taitelbaum, Gaurav Singh Tomar, Ming-Wei Chang, Xuhui Jia, Kelvin C. K. Chan, Hexiang Hu, Yu-Chuan Su, Ming-Hsuan Yang

专题命中 多模态生成 :MLLM(abstract);分类 cs.CV

Comments Project page: https://kitten-project.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 多模态评测 21 篇

2411.17776 2025-07-29 cs.CV cs.MM 84%

Beyond Walking: A Large-Scale Image-Text Benchmark for Text-based Person Anomaly Search

Shuyu Yang, Yaxiong Wang, Li Zhu, Zhedong Zheng

机构 * Xi’an Jiaotong University(西安交通大学) Hefei University of Technology(合肥工业大学) University of Macau(澳门大学)

专题命中 多模态评测 :image-text(title,abstract);cross-modal(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.03328 2025-07-29 cs.CV cs.AI cs.NE 84%

Visual Enumeration Remains Challenging for Multimodal Generative AI

Alberto Testolin, Kuinan Hou, Marco Zorzi

机构 * Department of General Psychology and Department of Mathematics University of Padova(帕多瓦大学心理学系和数学系) Department of General Psychology University of Padova(帕多瓦大学心理学系) Department of General Psychology and Padova Neuroscience Center University of Padova(帕多瓦大学心理学系和帕多瓦神经科学中心) IRCSS San Camillo Hospital, Venice-Lido(威尼斯利多医院IRCSS桑卡莫医院)

专题命中 多模态评测 :multimodal(title,abstract);multimodal foundation model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20388 2025-07-29 cs.CV 83%

ModalFormer: Multimodal Transformer for Low-Light Image Enhancement

Alexandru Brateanu, Raul Balmez, Ciprian Orhei, Codruta Ancuti, Cosmin Ancuti

机构 * Department of Computer Science University of Manchester(计算机科学系曼彻斯特大学) Department of Computer and Information Technology Politehnica University of Timisoara(计算机与信息科技系蒂米什瓦拉工业大学)

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19525 2025-07-29 cs.LG cs.AI 83%

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs

Chenchen Zhao, Zhengyuan Shi, Xiangyu Wen, Chengjie Liu, Yi Liu, Yunhao Zhou, Yuxiang Zhao, Hefei Feng, Yinan Zhu, Gwok-Waa Wan, Xin Cheng, Weiyu Chen, Yongqi Fu, Chujie Chen, Chenhao Xue, Guangyu Sun, Ying Wang, Yibo Lin, Jun Yang, Ning Xu, Xi Wang, Qiang Xu

机构 * Department of Computer Science and Engineering, The Chinese University of Hong Kong(中国香港中文大学计算机科学与工程系) School of Electronic Science and Engineering, Nanjing University(南京大学电子科学与工程学院) School of Integrated Circuits, Peking University(北京大学集成电路学院) School of Intergrated Circuits, Southeast University(东南大学集成电路学院) School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院) Department of Computer Science and Technology, University of Chinese Academy of Sciences(中国科学院大学计算机科学与技术系) National Center of Technology Innovation for EDA(EDA技术创新国家中心)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

Comments 10 pages, 1 figure, 5 tables. To appear in ICCAD 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20872 2025-07-29 cs.CV cs.AI cs.LG 81%

Not Only Grey Matter: OmniBrain for Robust Multimodal Classification of Alzheimer's Disease

Ahmed Sharshar, Yasser Ashraf, Tameem Bakr, Salma Hassan, Hosam Elgendy, Mohammad Yaqub, Mohsen Guizani

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Published in Third Workshop on Computer Vision for Automated Medical Diagnosis CVAMD 2025 in ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20737 2025-07-29 cs.CV cs.AI cs.HC 81%

Multi-Masked Querying Network for Robust Emotion Recognition from Incomplete Multi-Modal Physiological Signals

Geng-Xin Xu, Xiang Zuo, Ye Li

机构 * Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Shenzhen 518055, China(深圳先进技术研究院,中国科学院,深圳518055,中国) Southern University of Science and Technology, Shenzhen 518055, China(南方科技大学,深圳518055,中国)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments MICCAI2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19969 2025-07-29 cs.CL cs.CV 81%

Text2Vis: A Challenging and Diverse Benchmark for Generating Multimodal Visualizations from Text

Mizanur Rahman, Md Tahmid Rahman Laskar, Shafiq Joty, Enamul Hoque

机构 * York University(约克大学) Dialpad Salesforce AI Research(Salesforce人工智能研究)

专题命中 多模态评测 :multimodal(title);cross-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏