arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-30 至 2025-09-30 共收录 142 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 14 篇

2509.23044 2025-09-30 cs.CV cs.AI 81%

MMeViT: Multi-Modal ensemble ViT for Post-Stroke Rehabilitation Action Recognition

Ye-eun Kim, Suhyeon Lim, Andrew J. Choi

机构 * National Rehabilitation Center, Ministry of Health and Welfare, Korea(韩国卫生福利部国家康复中心) Gachon University(高丽大学)

专题命中 视频多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.AI

Comments 9 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18436 2025-09-30 cs.AI cs.CL cs.DB 81%

Memory-QA: Answering Recall Questions Based on Multimodal Memories

Hongda Jiang, Xinyuan Zhang, Siddhant Garg, Rishab Arora, Shiun-Zu Kuo, Jiayang Xu, Ankur Bansal, Christopher Brossman, Yue Liu, Aaron Colak, Ahmed Aly, Anuj Kumar, Xin Luna Dong

机构 * Meta Reality Labs(Meta现实实验室)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23085 2025-09-30 cs.IR cs.AI 79%

Enhancing Live Broadcast Engagement: A Multi-modal Approach to Short Video Recommendations Using MMGCN and User Preferences

Saeid Aghasoleymani Najafabadi, Elaheh Nabavi Nia

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13838 2025-09-30 eess.SP cs.CV cs.IT eess.IV math.IT 74%

Generative Video Semantic Communication via Multimodal Semantic Fusion with Large Model

Hang Yin, Li Qiao, Yu Ma, Shuo Sun, Kan Li, Zhen Gao, Dusit Niyato

机构 * School of Information and Electronics, Beijing Institute of Technology(信息与电子学院,北京理工大学) School of Computer Science and Engineering, Nanyang Technological University(计算机科学与工程学院,南洋理工大学)

专题命中 视频多模态 :multimodal(title);分类 cs.CV

Comments IEEE Transactions on Vehicular Technology

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24871 2025-09-30 cs.CV 70%

StreamForest: Efficient Online Video Understanding with Persistent Event Memory

Xiangyu Zeng, Kefan Qiu, Qingyu Zhang, Xinhao Li, Jing Wang, Jiaxin Li, Ziang Yan, Kun Tian, Meng Tian, Xinhai Zhao, Yi Wang, Limin Wang

机构 * Nanjing University(南京大学) Shanghai AI Laboratory(上海人工智能实验室) Zhejiang University(浙江大学) Noah’s Ark Lab, Huawei(华为诺亚实验室) Yinwang Intelligent Tech.(云网智能科技)

专题命中 视频多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments Accepted as a Spotlight at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25190 2025-09-30 cs.CV 57%

Visual Jigsaw Post-Training Improves MLLMs

Penghao Wu, Yushan Zhang, Haiwen Diao, Bo Li, Lewei Lu, Ziwei Liu

机构 * S-Lab, Nanyang Technological University(南洋理工大学S实验室) Linköping University(_linköping大学) SenseTime Research(商汤科技研究院)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24109 2025-09-30 cs.CV 57%

SVAC: Scaling Is All You Need For Referring Video Object Segmentation

Li Zhang, Haoxiang Gao, Zhihao Zhang, Luoxiao Huang, Tao Zhang

机构 * Columbia University New York, USA(哥伦比亚大学) Carnegie Mellon University Pittsburgh, USA(卡内基梅隆大学) New York University New York, USA(纽约大学) Wuhan University Wuhan, China(武汉大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments This paper is accepted to BMVC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10390 2025-09-30 cs.CV 57%

DART: Differentiable Dynamic Adaptive Region Tokenizer for Vision Foundation Models

Shicheng Yin, Kaixuan Yin, Yang Liu, Weixing Chen, Liang Lin

机构 * Sun Yat-sen University(中山大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Code is available at https://github.com/HCPLab-SYSU/DART

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06856 2025-09-30 cs.CV 57%

Vision-EKIPL: External Knowledge-Infused Policy Learning for Visual Reasoning

Chaoyang Wang, Zeyu Zhang, Meng Meng, Xu Zhou, Haiyun Jiang

机构 * Sangfor Technologies The Australian National University(澳大利亚国立大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02371 2025-09-30 cond-mat.mtrl-sci cond-mat.mes-hall cond-mat.str-el 50%

Time- and Polarization-Resolved Extreme Ultraviolet Momentum Microscopy

Sotirios Fragkos, Quentin Courtade, Olena Tkach, Jérôme Gaudin, Dominique Descamps, Guillaume Barrette, Stéphane Petit, Gerd Schönhense, Yann Mairesse, Samuel Beaulieu

专题命中 视频多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23112 2025-09-30 cs.RO 50%

FTACT: Force Torque aware Action Chunking Transformer for Pick-and-Reorient Bottle Task

Ryo Watanabe, Maxime Alvarez, Pablo Ferreiro, Pavel Savkin, Genki Sano

机构 * TELEXISTENCE Inc, Foundation Model Division(TELEXISTENCE公司,基础模型部门) The University of Tokyo(东京大学)

专题命中 视频多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.16829 2025-09-30 q-bio.NC cond-mat.stat-mech 50%

Neural spikes as rare events

Siddharth Kackar

专题命中 视频多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 跨模态检索 9 篇

2509.23895 2025-09-30 cs.CV cs.AI 88%

Preserving Cross-Modal Stability for Visual Unlearning in Multimodal Scenarios

Jinghan Xu Yuyang Zhang Qixuan Cai Jiancheng Chen Keqiu Li

机构 * Tianjin University(天津大学)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.CV、cs.AI

Comments 9 pages,4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21956 2025-09-30 cs.CV cs.AI cs.CL cs.LG 85%

Cross-modal RAG: Sub-dimensional Text-to-Image Retrieval-Augmented Generation

Mengdan Zhu, Senhao Cheng, Guangji Bai, Yifei Zhang, Liang Zhao

机构 * Emory University(埃默里大学) University of Michigan, Ann Arbor(密歇根大学安娜堡分校)

专题命中 跨模态检索 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25177 2025-09-30 cs.CV 83%

Mitigating Hallucination in Multimodal LLMs with Layer Contrastive Decoding

Bingkui Tong, Jiaer Xia, Kaiyang Zhou

机构 * Mohamed bin Zayed University of Artificial Intelligence(莫扎德·本·扎耶德人工智能大学) Hong Kong Baptist University(香港 Baptist大学)

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19994 2025-09-30 cs.CV 83%

Improving Generalizability and Undetectability for Targeted Adversarial Attacks on Multimodal Pre-trained Models

Zhifang Zhang, Jiahan Zhang, Shengjie Zhou, Qi Wei, Shuo He, Feng Liu, Lei Feng

机构 * Southeast University(东南大学) Johns Hopkins University(约翰霍普金斯大学) Chongqing University(重庆大学) Nanyang Technological University(南洋理工大学) University of Melbourne(墨尔本大学)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18450 2025-09-30 cs.CL 83%

BRIT: Bidirectional Retrieval over Unified Image-Text Graph

Ainulla Khan, Yamada Moyuru, Srinidhi Akella

机构 * Fujitsu Research India(富士通印度研究)

专题命中 跨模态检索 :image-text(title);multi-modal(abstract);cross-modal(abstract);分类 cs.CL

Comments Accepted in EMNLP-2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24823 2025-09-30 cs.CR cs.AI cs.CV cs.LG 62%

Of-SemWat: High-payload text embedding for semantic watermarking of AI-generated images with arbitrary size

Benedetta Tondi, Andrea Costanzo, Mauro Barni

机构 * University of Siena(锡耶纳大学)

专题命中 跨模态检索 :image-text(abstract);分类 cs.CV、cs.AI

Comments 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24370 2025-09-30 cs.CV 57%

DINOReg: Strong Point Cloud Registration with Vision Foundation Model

Congjia Chen, Yufu Qu

机构 * School of Instrumentation and Optoelectronic Engineering(仪器与光电工程学院)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23242 2025-09-30 cs.CV 57%

TATTOO: Training-free AesTheTic-aware Outfit recOmmendation

Yuntian Wu, Xiaonan Hu, Ziqi Zhou, Hao Lu

机构 * Huazhong University of Science and Technology(华中科技大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

Comments 4 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23543 2025-09-30 q-bio.GN cs.NE q-bio.MN 50%

Contrastive Learning Enhances Language Model Based Cell Embeddings for Low-Sample Single Cell Transcriptomics

Luxuan Zhang, Douglas Jiang, Qinglong Wang, Haoqi Sun, Feng Tian

专题命中 跨模态检索 :multimodal(abstract)

Comments 14 pages, 4 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态生成 15 篇

2509.22930 2025-09-30 cs.CV 79%

FishAI 2.0: Marine Fish Image Classification with Multi-modal Few-shot Learning

Chenghan Yang, Peng Zhou, Dong-Sheng Zhang, Yueyun Wang, Hong-Bin Shen, Xiaoyong Pan

专题命中 多模态生成 :multi-modal(title);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23760 2025-09-30 cs.CV 77%

UniAlignment: Semantic Alignment for Unified Image Generation, Understanding, Manipulation and Perception

Xinyang Song, Libin Wang, Weining Wang, Shaozhen Liu, Dandan Zheng, Jingdong Chen, Qi Li, Zhenan Sun

机构 * Ant Group(蚂蚁集团)

专题命中 多模态生成 :multimodal(abstract);multi-modal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23635 2025-09-30 cs.CV 74%

MotionVerse: A Unified Multimodal Framework for Motion Comprehension, Generation and Editing

Ruibing Hou, Mingshuang Luo, Hongyu Pan, Hong Chang, Shiguang Shan

机构 * Key Laboratory of Intelligent Information Processing, Institute of Computing Technology (ICT), Chinese Academy of Sciences (CAS)(智能信息处理重点实验室,计算技术研究所(ICT),中国科学院(CAS)) University of the Chinese Academy of Sciences(中国科学院大学)

专题命中 多模态生成 :multimodal(title);分类 cs.CV

Comments 17 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25047 2025-09-30 cs.AI 70%

Scaling Synthetic Task Generation for Agents via Exploration

Ram Ramrakhya, Andrew Szot, Omar Attia, Yuhao Yang, Anh Nguyen, Bogdan Mazoure, Zhe Gan, Harsh Agrawal, Alexander Toshev

机构 * Apple(苹果公司)

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24427 2025-09-30 cs.CV 70%

UI2V-Bench: An Understanding-based Image-to-video Generation Benchmark

Ailing Zhang, Lina Lei, Dehong Kong, Zhixin Wang, Jiaqi Xu, Fenglong Song, Chun-Le Guo, Chang Liu, Fan Li, Jie Chen

机构 * Peking University(北京大学) Huawei Noah’s Ark Lab(华为诺亚实验室) Nankai University(南开大学) Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区) Tsinghua University(清华大学)

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23625 2025-09-30 cs.CV cs.AI cs.CL cs.LG 67%

RIV: Recursive Introspection Mask Diffusion Vision Language Model

YuQian Li, Limeng Qiao, Lin Ma

机构 * Meituan(美团)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22940 2025-09-30 cs.CL cs.CV 62%

LLMs Behind the Scenes: Enabling Narrative Scene Illustration

Melissa Roemmele, John Joon Young Chung, Taewook Kim, Yuqian Sun, Alex Calderwood, Max Kreminski

机构 * Midjourney Northwestern University(西北大学) University of California, Santa Cruz(加州大学圣克鲁兹分校)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.CL

Comments Accepted at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24903 2025-09-30 cs.RO cs.CV eess.IV 57%

DRCP: Diffusion on Reinforced Cooperative Perception for Perceiving Beyond Limits

Lantao Li, Kang Yang, Rui Song, Chen Sun

机构 * Sony (China) Limited(索尼(中国)有限公司) Renmin University of China(中国人民大学) Fraunhofer IVI(弗劳恩霍夫IVI研究所)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24875 2025-09-30 cs.CV cs.LG 57%

Environment-Aware Satellite Image Generation with Diffusion Models

Nikos Kostagiolas, Pantelis Georgiades, Yannis Panagakis, Mihalis A. Nicolaou

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏