arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-28 至 2025-10-28 共收录 129 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 15 篇

2510.22602 2025-10-28 cs.CL cs.AI cs.CY 62%

Personal Care Utility (PCU): Building the Health Infrastructure for Everyday Insight and Guidance

Mahyar Abbasian, Ramesh Jain

机构 * University of California, Irvine(加州大学尔湾分校)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CL、cs.AI

Comments 22 pages, 2 figures, 1 table, Journal paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21813 2025-10-28 cs.CV cs.AI cs.LG 62%

SITS-DECO: A Generative Decoder Is All You Need For Multitask Satellite Image Time Series Modelling

Samuel J. Barrett, Docko Sow

机构 * LGND AI Tolbi

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments 27 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23569 2025-10-28 cs.CV 57%

EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT

Baoqi Pei, Yifei Huang, Jilan Xu, Yuping He, Guo Chen, Fei Wu, Yu Qiao, Jiangmiao Pang

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Zhejiang University(浙江大学) The University of Tokyo(东京大学) Fudan University(复旦大学) Nanjing University(南京大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23473 2025-10-28 cs.CV 57%

Video-Thinker: Sparking "Thinking with Videos" via Reinforcement Learning

Shijian Wang, Jiarui Jin, Xingjian Wang, Linxin Song, Runhao Fu, Hecheng Wang, Zongyuan Ge, Yuan Lu, Xuelian Cheng

机构 * Southeast University(东南大学) Monash University(墨尔本大学) Xiaohongshu Inc.(小红书公司) University of Southern California(南加州大学) Fudan University(复旦大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23397 2025-10-28 cs.CV 57%

VideoTG-R1: Boosting Video Temporal Grounding via Curriculum Reinforcement Learning on Reflected Boundary Annotations

Lu Dong, Haiyu Zhang, Han Lin, Ziang Yan, Xiangyu Zeng, Hongjie Zhang, Yifei Huang, Yi Wang, Zhen-Hua Ling, Limin Wang, Yali Wang

机构 * University of Science and Technology of China(中国科学技术大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Beihang University(北京航空航天大学) Shanghai Jiao Tong University(上海交通大学) Zhejiang University(浙江大学) State Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室) Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究所)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23253 2025-10-28 cs.CV 57%

A Video Is Not Worth a Thousand Words

Sam Pollard, Michael Wray

机构 * University of Bristol(布里斯托大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17924 2025-10-28 cs.CV 57%

Gaze into the Heart: A Multi-View Video Dataset for rPPG and Health Biomarkers Estimation

Konstantin Egorov, Stepan Botman, Pavel Blinov, Galina Zubkova, Anton Ivaschenko, Alexander Kolsanov, Andrey Savchenko

机构 * Sber AI Lab(Sber AI实验室) Samara State Medical University(萨马拉州医学大学) ISP RAS Research Center for Trusted Artificial Intelligence(俄罗斯科学院信息与通信技术研究所可信人工智能研究中心)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted to ACMMM 2025, Datasets track

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22339 2025-10-28 cs.RO 50%

Estimating Continuum Robot Shape under External Loading using Spatiotemporal Neural Networks

Enyi Wang, Zhen Deng, Chuanchuan Pan, Bingwei He, Jianwei Zhang

机构 * Hamlyn Centre for Robotic Surgery, Institute of Global Health Innovation, Imperial College London(帝国理工学院伦敦校区全球健康创新研究所机器人手术中心) Department of Mechanical Engineering and Automation, Fuzhou University(福州大学机械工程与自动化学院) TAMS Group, Informatics, University of Hamburg(汉堡大学信息学院TAMS集团)

专题命中 视频多模态 :multi-modal(abstract)

Comments 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 跨模态检索 11 篇

2510.02745 2025-10-28 cs.CV 88%

Retrv-R1: A Reasoning-Driven MLLM Framework for Universal and Efficient Multimodal Retrieval

Lanyun Zhu, Deyi Ji, Tianrun Chen, Haiyang Wu, Shiqi Wang

机构 * City University of Hong Kong(香港城市大学) Tencent(腾讯) Zhejiang University(浙江大学)

专题命中 跨模态检索 :multimodal(title,abstract);MLLM(title,abstract);分类 cs.CV

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22694 2025-10-28 cs.CV cs.CL cs.IR 81%

Windsock is Dancing: Adaptive Multimodal Retrieval-Augmented Generation

Shu Zhao, Tianyi Shen, Nilesh Ahuja, Omesh Tickoo, Vijaykrishnan Narayanan

机构 * The Pennsylvania State University(宾夕法尼亚州立大学) Intel(英特尔)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted at NeurIPS 2025 UniReps Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23224 2025-10-28 cs.CV cs.IR 79%

Accurate and Scalable Multimodal Pathology Retrieval via Attentive Vision-Language Alignment

Hongyi Wang, Zhengjie Zhu, Jiabo Ma, Fang Wang, Yue Shi, Bo Luo, Jili Wang, Qiuyu Cai, Xiuming Zhang, Yen-Wei Chen, Lanfen Lin, Hao Chen

机构 * Department of Computer Science and Engineering, The Hong Kong University of Science and Technology(香港科学与技术大学计算机科学与工程系) College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) Department of Radiology, Union Hospital, Tongji Medical College, Huazhong University of Science and Technology(华中科技大学同济医学院附属同济医院放射科) Department of Pathology, Sir Run Run Shaw Hospital, School of Medicine, Zhejiang University(浙江大学医学院附属邵氏医院病理科) Department of Pathology, The Central Hospital of Wuhan, Tongji Medical College, Huazhong University of Science and Technology(华中科技大学同济医学院附属武汉中心医院病理科) Department of Pathology, The First Affiliated Hospital, School of Medicine, Zhejiang University(浙江大学医学院附属第一医院病理科) Department of Chemical and Biological Engineering, The Hong Kong University of Science and Technology(香港科学与技术大学化学与生物工程系) Division of Life Science, The Hong Kong University of Science and Technology(香港科学与技术大学生命科学系)

专题命中 跨模态检索 :multimodal(title);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22880 2025-10-28 cs.LG cs.AI 79%

Learning Reconfigurable Representations for Multimodal Federated Learning with Missing Data

Duong M. Nguyen, Trong Nghia Hoang, Thanh Trung Huynh, Quoc Viet Hung Nguyen, Phi Le Nguyen

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Washington State University(华盛顿州立大学) VinUniversity(文大学) Griffin University(格里芬大学) Hanoi University of Science and Technology(河内科学技术大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

Comments Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05715 2025-10-28 cs.IR cs.MM 79%

From ID-based to ID-free: Rethinking ID Effectiveness in Multimodal Collaborative Filtering Recommendation

Guohao Li, Li Jing, Jia Wu, Xuefei Li, Kai Zhu, Yue He

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.MM

Comments We identified that our current approach achieves its reported performance only under specific data conditions, and its robustness is weaker than we initially expected

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22008 2025-10-28 cs.LG q-bio.MN 78%

A Multimodal Human Protein Embeddings Database: DeepDrug Protein Embeddings Bank (DPEB)

Md Saiful Islam Sajol, Magesh Rajasekaran, Hayden Gemeinhardt, Adam Bess, Chris Alvin, Supratik Mukhopadhyay

机构 * Computer Science, Louisiana State University(计算机科学,路易斯安那州立大学) Computer Science, Furman University(计算机科学,福兰明大学) Environmental Sciences and Center for Computation and Technology, Louisiana State University(环境科学与计算技术中心,路易斯安那州立大学)

专题命中 跨模态检索 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22998 2025-10-28 cs.AI 57%

ProfileXAI: User-Adaptive Explainable AI

Gilber A. Corrales, Carlos Andrés Ferro Sánchez, Reinel Tabares-Soto, Jesús Alfonso López Sotelo, Gonzalo A. Ruz, Johan Sebastian Piña Durán

机构 * Facultad de Ingeniería y Ciencias Básicas, Universidad Autónoma de Occidente(工程与基础科学学院,自治大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

Comments pages, 1 figure, 3 tables. Preprint. Evaluated on UCI Heart Disease (1989) and UCI Differentiated Thyroid Cancer Recurrence (2023). Uses IEEEtran

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22937 2025-10-28 cs.CV cs.LG 57%

Bi-Encoder Contrastive Learning for Fingerprint and Iris Biometrics

Matthew So, Judah Goldfeder, Mark Lis, Hod Lipson

机构 * Department of Computer Science(计算机科学系) Columbia University(哥伦比亚大学) College of Medicine(医学院) SUNY Downstate Health Sciences University(SUNY 下州健康科学大学) Department of Mechanical Engineering(机械工程系)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21717 2025-10-28 cs.HC cs.AI cs.SE 57%

AI-Enhanced Operator Assistance for UNICOS Applications

Bernard Tam, Jean-Charles Tournier, Fernando Varela Rodriguez

机构 * The University of Sydney(悉尼大学)

专题命中 跨模态检索 :multi-modal(abstract);分类 cs.AI

Comments Prepared as part of the CERN openlab programme 2025. Also available on Zenodo, a repository operated by CERN and co-funded by the European Union

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22613 2025-10-28 cs.SE 50%

DynaCausal: Dynamic Causality-Aware Root Cause Analysis for Distributed Microservices

Songhan Zhang, Aoyang Fang, Yifan Yang, Ruiyi Cheng, Xiaoying Tang, Pinjia He

专题命中 跨模态检索 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07759 2025-10-28 cs.IR 50%

A Survey of Long-Document Retrieval in the PLM and LLM Era

Minghan Li, Miyang Luo, Tianrui Lv, Yishuai Zhang, Siqi Zhao, Ercong Nie, Guodong Zhou

专题命中 跨模态检索 :multimodal(abstract)

Comments 32 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态生成 11 篇

2505.22453 2025-10-28 cs.CL cs.AI cs.CV cs.LG 85%

First SFT, Second RL, Third UPT: Continual Improving Multi-Modal LLM Reasoning via Unsupervised Post-Training

Lai Wei, Yuting Li, Chen Wang, Yue Wang, Linghe Kong, Weiran Huang, Lichao Sun

机构 * School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学学院) Zhongguancun Academy(中关村学院) Shanghai Innovation Institute(上海创新研究院) Lehigh University(莱特大学)

专题命中 多模态生成 :multi-modal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23449 2025-10-28 cs.MM cs.CV cs.IR 84%

CMIE: Combining MLLM Insights with External Evidence for Explainable Out-of-Context Misinformation Detection

Fanxiao Li, Jiaying Wu, Canyuan He, Wei Zhou

机构 * School of Information Science and Engineering, Yunnan University(云南大学信息科学与工程学院) National University of Singapore(新加坡国立大学) Engineering Research Center of Cyberspace, Yunnan University(云南大学网络空间研究院)

专题命中 多模态生成 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22684 2025-10-28 cs.CV cs.CL 81%

RoboSVG: A Unified Framework for Interactive SVG Generation with Multi-modal Guidance

Jiuniu Wang, Gongjie Zhang, Quanhao Qian, Junlong Gao, Deli Zhao, Ran Xu

机构 * DAMO Academy, Alibaba Group(达摩院,阿里巴巴集团)

专题命中 多模态生成 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.CL

Comments 15 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22521 2025-10-28 cs.CV cs.AI cs.IR cs.LG 81%

Open Multimodal Retrieval-Augmented Factual Image Generation

Yang Tian, Fan Liu, Jingyuan Zhang, Wei Bi, Yupeng Hu, Liqiang Nie

机构 * Shandong University(山东大学) National University of Singapore(新加坡国立大学) Kuaishou Technology(快手科技) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23382 2025-10-28 cs.CV 79%

An Efficient Remote Sensing Super Resolution Method Exploring Diffusion Priors and Multi-Modal Constraints for Crop Type Mapping

Songxi Yang, Tang Sui, Qunying Huang

机构 * Department of Geography(地理系) University of Wisconsin-Madison(威斯康星大学麦迪逊分校)

专题命中 多模态生成 :multi-modal(title);multimodal(abstract);分类 cs.CV

Comments 41 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21835 2025-10-28 cs.LG cs.AI cs.CL cs.CV 79%

A Multimodal, Multitask System for Generating E Commerce Text Listings from Images

Nayan Kumar Singh

专题命中 多模态生成 :multimodal(title,comments);分类 cs.CV、cs.CL、cs.AI

Comments 24 pages, 10 figures, 11 tables. Code can be found at: https://github.com/SinghNayanKumar/multimodal-product-lister/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14350 2025-10-28 cs.CV cs.AI cs.CL 75%

VEGGIE: Instructional Editing and Reasoning Video Concepts with Grounded Generation

Shoubin Yu, Difan Liu, Ziqiao Ma, Yicong Hong, Yang Zhou, Hao Tan, Joyce Chai, Mohit Bansal

机构 * Adobe Research(Adobe研究机构) University of Michigan(密歇根大学) UNC Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

Comments ICCV 2025; First three authors contributed equally. Project page: https://veggie-gen.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22431 2025-10-28 cs.MA cs.CV 74%

Hollywood Town: Long-Video Generation via Cross-Modal Multi-Agent Orchestration

Zheng Wei, Mingchen Li, Zeqian Zhang, Ruibin Yuan, Pan Hui, Huamin Qu, James Evans, Maneesh Agrawala, Anyi Rao

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) University of Chicago(芝加哥大学) Stanford University(斯坦福大学)

专题命中 多模态生成 :cross-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23020 2025-10-28 cs.CV cs.CL 62%

M$^{3}$T2IBench: A Large-Scale Multi-Category, Multi-Instance, Multi-Relation Text-to-Image Benchmark

Huixuan Zhang, Xiaojun Wan

机构 * Wangxuan Institute of Computer Technology, Peking University(计算机技术研究所,北京大学)

专题命中 多模态生成 :image-text(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14627 2025-10-28 cs.RO cs.CV 57%

GOPLA: Generalizable Object Placement Learning via Synthetic Augmentation of Human Arrangement

Yao Zhong, Hanzhi Chen, Simon Schaefer, Anran Zhang, Stefan Leutenegger

机构 * Technical University of Munich(慕尼黑技术大学) ETH Zurich(苏黎世联邦理工学院)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10776 2025-10-28 cs.CR cs.AI 57%

ME: Trigger Element Combination Backdoor Attack on Copyright Infringement

Feiyu Yang, Siyuan Liang, Aishan Liu, Dacheng Tao

专题命中 多模态生成 :image-text(abstract);分类 cs.AI

Comments Finding unfinished issue in this work , still refining

详情

展开后加载摘要…

URL PDF HTML 收藏