arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-09 至 2025-09-09 共收录 11 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 11 篇

2411.19787 2025-09-09 cs.LG cs.AI 83%

CAREL: Instruction-guided reinforcement learning with cross-modal auxiliary objectives

Armin Saghafian, Amirmohammad Izadi, Negin Hashemi Dijujin, Mahdieh Soleymani Baghshah

机构 * Sharif University of Technology(谢里夫理工大学)

专题命中 视频多模态 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.AI

Comments Accepted to TMLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06831 2025-09-09 cs.CV 79%

Leveraging Generic Foundation Models for Multimodal Surgical Data Analysis

Simon Pezold, Jérôme A. Kurylec, Jan S. Liechti, Beat P. Müller, Joël L. Lavanchy

机构 * Department of Biomedical Engineering, University of Basel, Allschwil, Switzerland(巴塞尔大学生物医学工程系) Clarunis – University Digestive Health Care Center Basel, Basel, Switzerland(Clarunis – 巴塞尔大学消化健康医疗中心)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 13 pages, 3 figures; accepted at ML-CDS @ MICCAI 2025, Daejeon, Republic of Korea

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06422 2025-09-09 cs.CV 79%

Phantom-Insight: Adaptive Multi-cue Fusion for Video Camouflaged Object Detection with Multimodal LLM

Hua Zhang, Changjiang Luo, Ruoyu Chen

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院)

专题命中 视频多模态 :multimodal(title);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06389 2025-09-09 cs.SD cs.AI 79%

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation

Xiaoran Yang, Jianxuan Yang, Xinyue Guo, Haoyu Wang, Ningning Pan, Gongping Huang

机构 * School of Electronic Information, Wuhan University, Wuhan, China(武汉大学电子信息学院) MiLM Plus, Xiaomi Inc., Wuhan, China(小米公司MiLM Plus团队)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05898 2025-09-09 cs.HC 78%

Attention, Action, and Memory: How Multi-modal Interfaces and Cognitive Load Alter Information Retention

Omar Elgohary, Zhu-Tien

专题命中 视频多模态 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19153 2025-09-09 cs.RO cs.AI cs.CV cs.SY eess.IV eess.SY 73%

QuadKAN: KAN-Enhanced Quadruped Motion Control via End-to-End Reinforcement Learning

Yinuo Wang, Gavin Tao

机构 * Allen Wang Gavin Tao

专题命中 视频多模态 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments 14pages, 9 figures, Journal paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09632 2025-09-09 cs.CV cs.AI 62%

Preacher: Paper-to-Video Agentic System

Jingwei Liu, Ling Yang, Hao Luo, Fan Wang, Hongyan Li, Mengdi Wang

机构 * School of Intelligence Science and Technology, Peking University(北京理工大学智能科学与技术学院) DAMO Academy, Alibaba group(阿里巴巴集团大模型研究院) Hupan Lab(虎扑实验室) National Key Laboratory of General Artificial Intelligence, Peking University(北京人工智能 general artificial intelligence 国家重点实验室) Department of Electrical and Computer Engineering, Princeton University(普林斯顿大学电气与计算机工程系)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments ICCV 2025. Code: https://github.com/Gen-Verse/Paper2Video

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05298 2025-09-09 cs.HC cs.AI cs.MM 62%

Livia: An Emotion-Aware AR Companion Powered by Modular AI Agents and Progressive Memory Compression

Rui Xi, Xianghan Wang

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI、cs.MM

Comments Accepted to the Proceedings of the 2025 International Conference on Artificial Intelligence and Virtual Reality (AIVR 2025). \c{opyright} 2025 Springer. This is the author-accepted manuscript. Rui Xi and Xianghan Wang contributed equally to this work. The final version will be available via SpringerLink

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06023 2025-09-09 cs.CV 57%

DVLO4D: Deep Visual-Lidar Odometry with Sparse Spatial-temporal Fusion

Mengmeng Liu, Michael Ying Yang, Jiuming Liu, Yunpeng Zhang, Jiangtao Li, Sander Oude Elberink, George Vosselman, Hao Cheng

机构 * University of Twente(特文特大学) University of Bath(巴斯大学) Shanghai Jiao Tong University(上海交通大学) PhiGent Robotics(PhiGent机器人)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted by ICRA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01563 2025-09-09 cs.CV 57%

Kwai Keye-VL 1.5 Technical Report

Biao Yang, Bin Wen, Boyang Ding, Changyi Liu, Chenglong Chu, Chengru Song, Chongling Rao, Chuan Yi, Da Li, Dunju Zang, Fan Yang, Guorui Zhou, Guowang Zhang, Han Shen, Hao Peng, Haojie Ding, Hao Wang, Haonan Fan, Hengrui Ju, Jiaming Huang, Jiangxia Cao, Jiankang Chen, Jingyun Hua, Kaibing Chen, Kaiyu Jiang, Kaiyu Tang, Kun Gai, Muhao Wei, Qiang Wang, Ruitao Wang, Sen Na, Shengnan Zhang, Siyang Mao, Sui Huang, Tianke Zhang, Tingting Gao, Wei Chen, Wei Yuan, Xiangyu Wu, Xiao Hu, Xingyu Lu, Yi-Fan Zhang, Yiping Yang, Yulong Chen, Zeyi Lu, Zhenhua Wu, Zhixin Ling, Zhuoran Yang, Ziming Li, Di Xu, Haixuan Gao, Hang Li, Jing Wang, Lejian Ren, Qigen Hu, Qianqian Wang, Shiyao Wang, Xinchen Luo, Yan Li, Yuhang Hu, Zixing Zhang

机构 * Keye Team, Kuaishou Group(快手集团Keye团队)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Github page: https://github.com/Kwai-Keye/Keye

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09423 2025-09-09 cs.RO cs.CV 57%

Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter

Kechun Xu, Xunlong Xia, Kaixuan Wang, Yifei Yang, Yunxuan Mao, Bing Deng, Jieping Ye, Rong Xiong, Yue Wang

机构 * Zhejiang University and Alibaba Cloud(浙江大学和阿里云) Alibaba Cloud(阿里云) Zhejiang University(浙江大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted by T-ASE and CoRL25 GenPriors Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏