arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-25 至 2025-11-25 共收录 134 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 18 篇

2511.17945 2025-11-25 cs.CV 79%

Test-Time Temporal Sampling for Efficient MLLM Video Understanding

测试时时间采样用于高效多模态大语言模型视频理解

Kaibin Wang, Mingbao Lin

机构 * SenseTime, China(深睿时代,中国) Rakuten, Singapore(拉结恩,新加坡)

专题命中 视频多模态 :MLLM(title);multimodal(abstract);分类 cs.CV

AI总结 T3S通过测试时时间采样提高多模态大语言模型处理长视频的效率和准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10510 2025-11-25 cs.NI cs.AI cs.HC cs.MM 73%

Chat with AI: The Surprising Turn of Real-time Video Communication from Human to AI

与AI聊天:从人类到AI的实时视频通信惊人转变

Jiangkai Wu, Zhiyuan Ren, Liming Liu, Xinggong Zhang

机构 * Peking University(北京大学)

专题命中 视频多模态 :multimodal(abstract);MLLM(abstract);分类 cs.AI、cs.MM

AI总结 本文提出了一种面向AI的实时视频通信方法,通过上下文感知视频流技术降低延迟并提升AI理解视频的准确性。

Comments 9 pages, 10 figures, Proceedings of the 24th ACM Workshop on Hot Topics in Networks (HotNets 2025), College Park, Maryland, USA

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13589 2025-11-25 cs.CV 70%

AdaVideoRAG: Omni-Contextual Adaptive Retrieval-Augmented Efficient Long Video Understanding

AdaVideoRAG:多情境自适应检索增强高效长视频理解

Zhucun Xue, Jiangning Zhang, Xurong Xie, Yuxuan Cai, Yong Liu, Xiangtai Li, Dacheng Tao

机构 * Zhejiang University(浙江大学) Youtu Lab, Tencent(腾讯优图实验室) Huazhong University of Science and Technolog(华中科技大学) Nanyang Technological University(南洋理工大学)

专题命中 视频多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

AI总结 AdaVideoRAG通过自适应检索增强框架提升长视频理解效率与准确性,支持多层级知识检索与深度语义分析。

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18711 2025-11-25 cs.CV cs.AI 62%

Modality-Collaborative Low-Rank Decomposers for Few-Shot Video Domain Adaptation

模态协同低秩分解器用于少样本视频域适应

Yuyang Wanyan, Xiaoshan Yang, Weiming Dong, Changsheng Xu

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) PengCheng Laboratory(鹏城实验室)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出模态协同低秩分解器,通过分解不同领域偏移级别的模态特征,提升少样本视频域适应的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27280 2025-11-25 cs.CV cs.AI cs.LG 62%

FOCUS: Efficient Keyframe Selection for Long Video Understanding

FOCUS: 长视频理解中的高效关键帧选择

Zirui Zhu, Hailun Xu, Yang Luo, Yong Liu, Kanchan Sarkar, Zhenheng Yang, Yang You

机构 * National University of Singapore(新加坡国立大学) TikTok

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 FOCUS通过两阶段探索-利用策略,在严格标记预算下高效选择关键帧,提升长视频理解的准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19236 2025-11-25 cs.RO cs.AI 57%

SENTINEL: A Fully End-to-End Language-Action Model for Humanoid Whole Body Control

SENTINEL:一种用于人形机器人全身控制的端到端语言-动作模型

Yuxuan Wang, Haobin Jiang, Shiqing Yao, Ziluo Ding, Zongqing Lu

机构 * Peking University(北京大学) BeingBeyond

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

AI总结 SENTINEL是一种端到端语言-动作模型,通过直接映射语言指令和本体感觉输入到低层动作,实现人形机器人全身控制,并支持多模态扩展。

Comments 23 pages, 8 figures, 11 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18920 2025-11-25 cs.CV 57%

EventSTU: Event-Guided Efficient Spatio-Temporal Understanding for Video Large Language Models

EventSTU: 基于事件的高效空间-时间理解用于视频大语言模型

Wenhao Xu, Xin Dong, Yue Li, Haoyuan Shi, Zhiwei Xiong

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

AI总结 EventSTU通过事件引导的方法,高效处理视频大语言模型的空间-时间理解,实现显著的计算效率提升和性能优化。

Comments 8 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18814 2025-11-25 cs.CV 57%

DetAny4D: Detect Anything 4D Temporally in a Streaming RGB Video

DetAny4D: 在流式RGB视频中实现任意4D检测

Jiawei Hou, Shenghao Zhang, Can Wang, Zheng Gu, Yonggen Ling, Taiping Zeng, Xiangyang Xue, Jingbo Zhang

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

AI总结 DetAny4D通过端到端框架实现流式RGB视频中任意4D物体检测,提升时间稳定性和检测精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18700 2025-11-25 cs.MM 57%

When Top-ranked Recommendations Fail: Modeling Multi-Granular Negative Feedback for Explainable and Robust Video Recommendation

当顶级推荐失效时:建模多粒度负面反馈以实现可解释且稳健的视频推荐

Siran Chen, Boyu Chen, Chenyun Yu, Yi Ouyang, Cheng Lei, Chengxiang Zhuo, Zang Li, Yali Wang

专题命中 视频多模态 :multimodal(abstract);分类 cs.MM

AI总结 本文提出ENF框架和S-GRPO算法,通过多粒度负面反馈建模提升视频推荐的可解释性和鲁棒性,实验表明在负面反馈预测和用户满意度提升方面效果显著。

Comments Accepted in AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16669 2025-11-25 cs.CV 57%

Video-as-Answer: Predict and Generate Next Video Event with Joint-GRPO

视频作答:联合GRPO预测并生成下一个视频事件

Junhao Cheng, Liang Hou, Xin Tao, Jing Liao

机构 * City University of Hong Kong(香港城市大学) Kling Team, Kuaishou Technology(快手技术 Kling 团队)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

AI总结 VANS通过联合GRPO方法,利用强化学习对齐视觉-语言模型与视频扩散模型,实现视频下一个事件预测任务的高精度视频生成与描述生成。

Comments Project page: https://video-as-answer.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18382 2025-11-25 cs.CV 57%

ViMix-14M: A Curated Multi-Source Video-Text Dataset with Long-Form, High-Quality Captions and Crawl-Free Access

ViMix-14M:一个经过精心整理的多源视频-文本数据集,具有长形式、高质量的描述和无爬虫访问

Timing Yang, Sucheng Ren, Alan Yuille, Feng Wang

机构 * Johns Hopkins University(约翰霍普金斯大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

AI总结 ViMix-14M是一个经过精心整理的多源视频-文本数据集,提供无爬虫访问、高质量描述和长形式文本,用于提升视频生成和多模态任务的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17943 2025-11-25 cs.CV 57%

SciEducator: Scientific Video Understanding and Educating via Deming-Cycle Multi-Agent System

SciEducator: 基于Deming循环多智能体系统的科学视频理解与教育

Zhiyu Xu, Weilong Yan, Yufei Shi, Xin Meng, Tao He, Huiping Zhuang, Ming Li, Hehe Fan

机构 * Jinan University(济南大学) National University of Singapore(新加坡国立大学) Nanyang Technological University(南洋理工大学) Peking University(北京大学) University of Electronic Science and Technology of China(电子科技大学) South China University of Technology(华南理工大学) Guangming Laboratory(光明实验室) Zhejiang University(浙江大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

AI总结 SciEducator通过Deming循环多智能体系统实现科学视频的自演化理解与教育,生成多模态教学内容并超越现有大语言模型和视频智能体。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01248 2025-11-25 cs.HC 50%

FocusView: Understanding and Customizing Informational Video Watching Experiences for Viewers with ADHD

FocusView:为ADHD观众理解并定制信息视频观看体验

Hanxiu 'Hazel' Zhu, Ruijia Chen, Yuhang Zhao

专题命中 视频多模态 :multimodal(abstract)

AI总结 FocusView通过定制信息视频界面,帮助ADHD观众减少干扰,提升视频观看体验。

Comments 15 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17926 2025-11-25 cs.SD cs.HC 50%

Three-Class Emotion Classification for Audiovisual Scenes Based on Ensemble Learning Scheme

基于集成学习方案的音频视觉场景三类情感分类

Xiangrui Xiong, Zhou Zhou, Guocai Nong, Junlin Deng, Ning Wu

专题命中 视频多模态 :multimodal(abstract)

AI总结 本文提出基于音频的集成学习框架,用于对电影场景进行三类情感分类,实验结果显示在现实数据集上达到86%的准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17898 2025-11-25 cs.RO 50%

L1 Sample Flow for Efficient Visuomotor Learning

L1样本流用于高效视觉-运动学习

Weixi Song, Zhetao Chen, Tao Xu, Xianchao Zeng, Xinyu Zhou, Lixin Yang, Donglin Wang, Cewu Lu, Yong-Lu Li

机构 * Zhejiang University(浙江大学) Shanghai Innovation Institute(上海创新研究院) Westlake University(西湖大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 视频多模态 :multi-modal(abstract)

AI总结 L1流通过结合去噪模型的多模分布捕捉能力与L1回归的高效性,实现高效的视觉-运动学习。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 跨模态检索 9 篇

2511.19257 2025-11-25 cs.CR cs.AI cs.LG 88%

Medusa: Cross-Modal Transferable Adversarial Attacks on Multimodal Medical Retrieval-Augmented Generation

Medusa: 跨模态可转移的对抗攻击用于多模态医疗检索增强生成

Yingjia Shang, Yi Liu, Huimin Wang, Furong Li, Wenfang Sun, Wu Chengyu, Yefeng Zheng

机构 * Westlake University(西湖大学) Heilongjiang University(黑龙江大学) City University of Hong Kong(香港城市大学) Tencent(腾讯)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(title,abstract);分类 cs.AI

AI总结 Medusa提出了一种针对多模态医疗检索增强生成系统的跨模态可转移对抗攻击方法,通过优化扰动和双循环策略实现高攻击成功率并抵御主流防御措施。

Comments Accepted at KDD 2026 First Cycle (full version). Authors marked with * contributed equally. Yi Liu is the lead author

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16654 2025-11-25 cs.CL 83%

Comparison of Text-Based and Image-Based Retrieval in Multimodal Retrieval Augmented Generation Large Language Model Systems

多模态检索增强生成大语言模型系统中基于文本和基于图像的检索比较

Elias Lumer, Alex Cardenas, Matt Melich, Myles Mason, Sara Dieter, Vamse Kumar Subbiah, Pradeep Honaganahalli Basavaraju, Roberto Hernandez

机构 * PricewaterhouseCoopers U.S.(普华永道美国公司)

专题命中 跨模态检索 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CL

AI总结 本文比较了多模态RAG系统中基于文本和基于图像的检索方法,发现直接多模态嵌入检索在性能和准确性上优于基于LLM总结的方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19380 2025-11-25 cs.CV 79%

UISearch: Graph-Based Embeddings for Multimodal Enterprise UI Screenshots Retrieval

UISearch: 基于图的多模态企业UI截图检索

Maroun Ayli, Youssef Bakouny, Tushar Sharma, Nader Jalloul, Hani Seifeddine, Rima Kilany

机构 * Center For Computer Science(计算机科学中心) Saint Joseph University of Beirut(贝鲁特圣约瑟夫大学) Faculty of Computer Science(计算机科学学院) Dalhousie University(达尔豪斯大学) Murex

专题命中 跨模态检索 :multimodal(title);multi-modal(abstract);分类 cs.CV

AI总结 UISearch通过基于图的结构嵌入与语义检索结合,实现多模态企业UI截图检索,提升检索准确率与效率。

Comments 12 pages, 2 figures, 3 algorithms, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18983 2025-11-25 cs.CV 79%

UMCL: Unimodal-generated Multimodal Contrastive Learning for Cross-compression-rate Deepfake Detection

UMCL: 单模生成多模对比学习用于跨压缩率深度伪造检测

Ching-Yi Lai, Chih-Yu Jian, Pei-Cheng Chuang, Chia-Ming Lee, Chih-Chung Hsu, Chiou-Ting Hsu, Chia-Wen Lin

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

AI总结 UMCL通过单模生成多模对比学习,提升跨压缩率深度伪造检测的鲁棒性和准确性。

Comments 24-page manuscript accepted to IJCV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18298 2025-11-25 cs.AI 70%

Cross-Disciplinary Knowledge Retrieval and Synthesis: A Compound AI Architecture for Scientific Discovery

跨学科知识检索与综合:一种用于科学发现的复合AI架构

Svitlana Volkova, Peter Bautista, Avinash Hiriyanna, Gabriel Ganberg, Isabel Erickson, Zachary Klinefelter, Nick Abele, Hsien-Te Kao, Grant Engberson

机构 * Aptima, Inc.(Aptima公司)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

AI总结 BioSage通过整合LLMs与RAG,利用专门代理实现跨学科知识检索与综合,提升科学发现效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.07221 2025-11-25 cs.CV cs.AI 62%

Exploring the Use of Contrastive Language-Image Pre-Training for Human Posture Classification: Insights from Yoga Pose Analysis

探索对比语言-图像预训练在人体姿态分类中的应用:从瑜伽姿势分析获得的见解

Andrzej D. Dobrzycki, Ana M. Bernardos, Luca Bergesio, Andrzej Pomirski, Daniel Sáez-Trigueros

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本研究利用CLIP模型在瑜伽姿势分类中取得高准确率,展示其在人体姿态识别中的潜力,并验证其在自动化系统中的应用可行性。

Journal ref Mathematics 2024, 12(1), 76

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24466 2025-11-25 cs.CV 57%

SA-Person: Text-Based Person Retrieval with Scene-aware Re-ranking

SA-Person: 基于文本的人检索与场景感知重排序

Yingjia Xu, Jinlin Wu, Daming Gao, Zhen Chen, Yang Yang, Min Cao, Mang Ye, Zhen Lei

机构 * School of Computer Science and Technology, Soochow University(苏州大学计算机科学与技术学院) Centre for Artificial Intelligence and Robotics, Hong Kong Institute of Science and Innovation, Chinese Academy of Sciences(中国科学院香港创新研究院人工智能与机器人中心) Multimodal Artificial Intelligence Systems (MAIS), Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所多模态人工智能系统(MAIS)) Wuhan University(武汉大学) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

AI总结 SA-Person通过整合个体外观和全局场景上下文,提升基于文本的人检索准确性,提出ScenePerson-13W数据集和两阶段检索框架。

Comments 13 pages, 8 figures. Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18598 2025-11-25 q-bio.OT 50%

Assessing Gaze and Pointing: Human Cue Interpretation by Indian Free-Ranging Dogs in a Food Retrieval Task

评估目光与指认:印度自由放养狗在食物获取任务中的人类提示解读

Srijaya Nandi, Dipanjan Roy, Aesha Lahiri, Anamitra Roy, Anindita Bhadra

专题命中 跨模态检索 :multimodal(abstract)

AI总结 研究发现印度自由放养狗在结合指认和目光提示时能准确找到隐藏食物,但单一或冲突提示下表现无显著差异,且狗的气质影响其参与意愿和接近延迟,但不影响选择准确性。

Comments 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10584 2025-11-25 cs.IR 50%

DAS: Dual-Aligned Semantic IDs Empowered Industrial Recommender System

DAS: 基于双对齐语义ID的工业推荐系统

Wencai Ye, Mingjie Sun, Shaoyun Shi, Peng Wang, Wenjin Wu, Peng Jiang

专题命中 跨模态检索 :multi-modal(abstract)

AI总结 DAS通过双对齐语义ID方法,提升推荐系统中多模态内容整合与协同信号对齐效率,有效解决信息损失与灵活性问题。

Comments Accepted by CIKM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态生成 10 篇

2505.22633 2025-11-25 cs.CL cs.AI cs.CV cs.LG cs.MM 85%

Spatial Knowledge Graph-Guided Multimodal Synthesis

基于空间知识图的多模态合成

Yida Xue, Zhen Bi, Jinnan Yang, Jungang Lou, Kehai Chen, Min Zhang, Huajun Chen, Ningyu Zhang

机构 * Zhejiang University(浙江大学) Nanjing University of Science and Technology(南京理工大学) Huzhou University(湖州大学) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 多模态生成 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 SKG2DATA通过空间知识图引导多模态合成,提升多模态大语言模型的空间感知与推理能力。

Comments IEEE/ACM Transactions on Audio, Speech and Language Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19343 2025-11-25 cs.CV 83%

Syn-GRPO: Self-Evolving Data Synthesis for MLLM Perception Reasoning

Syn-GRPO:面向多模态大语言模型感知推理的自进化数据合成

Qihan Huang, Haofei Zhang, Rong Wei, Yi Wang, Rui Tang, Mingli Song, Jie Song

机构 * Zhejiang University(浙江大学) Manycore Tech Inc.(Manycore科技公司)

专题命中 多模态生成 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV

AI总结 Syn-GRPO通过自进化数据合成提升多模态大语言模型的感知推理能力,采用数据服务器与GRPO工作流协同生成高质量多样化训练数据,实验验证其在视觉感知任务中的优越性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18714 2025-11-25 cs.AI cs.CY 83%

MAGMA-Edu: Multi-Agent Generative Multimodal Framework for Text-Diagram Educational Question Generation

MAGMA-Edu:多智能体生成多模态框架用于文本-图表教育问题生成

Zhenyu Wu, Jian Li, Hua Huang

机构 * School of Artificial Intelligence, Beijing Normal University(人工智能学院,北京师范大学)

专题命中 多模态生成 :multimodal(title,abstract);image-text(abstract);分类 cs.AI

AI总结 MAGMA-Edu通过多智能体框架实现教育问题生成,提升文本与图像的一致性,达到多模态教育内容生成的新水平。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18262 2025-11-25 cs.CV 79%

MammothModa2: A Unified AR-Diffusion Framework for Multimodal Understanding and Generation

MammothModa2: 一种用于多模态理解和生成的统一AR-扩散框架

Tao Shen, Xin Wan, Taicai Chen, Rui Zhang, Junwen Pan, Dawei Lu, Fanding Lei, Zhilin Lu, Yunfei Yang, Chen Cheng, Qi She, Chang Liu, Zhenbang Sun

机构 * ByteDance(字节跳动)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

AI总结 MammothModa2通过统一的AR-扩散框架实现多模态理解和生成,结合自回归语义规划与扩散生成,取得文本到图像和指令编辑的优异表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18131 2025-11-25 cs.CV 77%

Video4Edit: Viewing Image Editing as a Degenerate Temporal Process

Video4Edit: 将图像编辑视为一种退化的时间过程

Xiaofan Li, Yanpeng Sun, Chenming Wu, Fan Duan, YuAn Wang, Weihao Bo, Yumeng Zhang, Dingkang Liang

机构 * Baidu Inc.(百度公司)

专题命中 多模态生成 :multimodal(abstract);cross-modal(abstract);multimodal foundation model(abstract);分类 cs.CV

AI总结 Video4Edit通过将图像编辑视为退化的时间过程,利用视频预训练的单帧演化先验,实现高效的数据微调,从而在性能上与主流模型相当,但仅需1%的监督数据。

Comments 10 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17986 2025-11-25 cs.CV cs.AI 62%

Plan-X: Instruct Video Generation via Semantic Planning

Plan-X: 通过语义规划指导视频生成

Lun Huang, You Xie, Hongyi Xu, Tianpei Gu, Chenxu Zhang, Guoxian Song, Zenan Li, Xiaochen Zhao, Linjie Luo, Guillermo Sapiro

机构 * Duke University(杜克大学) Princeton University(普林斯顿大学) ByteDance Intelligent Creation(字节跳动智能创作) Apple(苹果公司)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 Plan-X通过语义规划框架减少视频生成中的视觉幻觉,实现与多模态上下文一致的精细指令对齐生成。

Comments The project page is at https://byteaigc.github.io/Plan-X

详情

展开后加载摘要…

URL PDF HTML 收藏