arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4735 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4735 篇

2508.05145 2025-08-08 cs.AI 57%

Graph-based Event Log Repair

Sebastiano Dissegna, Chiara Di Francescomarino, Massimiliano Ronzani

机构 * Department of Computer Science Engineering University of Trento(计算机科学工程系 特伦托大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03651 2025-08-06 cs.HC cs.AI 57%

Probing the Gaps in ChatGPT Live Video Chat for Real-World Assistance for People who are Blind or Visually Impaired

Ruei-Che Chang, Rosiana Natalie, Wenqian Xu, Jovan Zheng Feng Yap, Anhong Guo

机构 * University of Michigan(密歇根大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

Comments ACM ASSETS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01546 2025-08-05 cs.CV 57%

E-VRAG: Enhancing Long Video Understanding with Resource-Efficient Retrieval Augmented Generation

Zeyu Xu, Junkang Zhang, Qiang Wang, Yi Liu

机构 * Zeyu Xu(作者) Junkang Zhang(作者) Qiang Wang(作者) Yi Liu(作者)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21971 2025-07-30 cs.CV 57%

EIFNet: Leveraging Event-Image Fusion for Robust Semantic Segmentation

Zhijiang Li, Haoran He

机构 * Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20629 2025-07-29 cs.CV 57%

DAMS:Dual-Branch Adaptive Multiscale Spatiotemporal Framework for Video Anomaly Detection

Dezhi An, Wenqiang Liu, Kefan Wang, Zening Chen, Jun Lu, Shengcai Zhang

机构 * School of Cyberspace Security,Gansu University of Political Science and Law(网络安全学院、政治学科学校)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments 13 pages,7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20939 2025-07-29 cs.CV 57%

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Yuying Ge, Yixiao Ge, Chen Li, Teng Wang, Junfu Pu, Yizhuo Li, Lu Qiu, Jin Ma, Lisheng Duan, Xinyu Zuo, Jinwen Luo, Weibo Gu, Zexuan Li, Xiaojing Zhang, Yangyu Tao, Han Hu, Di Wang, Ying Shan

机构 * ARC Lab, Tencent PCG(腾讯PCG ARC实验室) Search Application Department, Tencent CSIG(腾讯CSIG搜索应用部门) Tencent Hunyuan(腾讯文生视频) Big Data Platform Department, Tencent PCG(腾讯PCG大数据平台部门)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Project Page: https://tencentarc.github.io/posts/arc-video-announcement/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20763 2025-07-29 cs.CV 57%

KASportsFormer: Kinematic Anatomy Enhanced Transformer for 3D Human Pose Estimation on Short Sports Scene Video

Zhuoer Yin, Calvin Yeung, Tomohiro Suzuki, Ryota Tanaka, Keisuke Fujii

机构 * Nagoya University(名古屋大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments 10 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19948 2025-07-29 cs.CV 57%

UniCT Depth: Event-Image Fusion Based Monocular Depth Estimation with Convolution-Compensated ViT Dual SA Block

Luoxi Jing, Dianxi Shi, Zhe Liu, Songchang Jin, Chunping Qiu, Ziteng Qiao, Yuxian Li, Jianqiang Xia

机构 * School of Computer Science, Peking University(北京大学计算机科学系) Intelligent Game and Decision Lab (IGDL)(智能游戏与决策实验室) College of Computer, National University of Defense Technology(国防科技大学计算机学院) School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学系)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted by IJCAI 2025 (International Joint Conference on Artificial Intelligence)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19863 2025-07-29 cs.MM 57%

Anchoring Trends: Mitigating Social Media Popularity Prediction Drift via Feature Clustering and Expansion

Chia-Ming Lee, Bo-Cheng Qiu, Cheng-Jun Kang, Yi-Hsuan Wu, Jun-Lin Chen, Yu-Fan Lin, Yi-Shiuan Chou, Chih-Chung Hsu

专题命中 视频多模态 :multi-modal(abstract);分类 cs.MM

Comments Accepted by ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01301 2025-07-29 cs.RO cs.AI 57%

Bi-LAT: Bilateral Control-Based Imitation Learning via Natural Language and Action Chunking with Transformers

Takumi Kobayashi, Masato Kobayashi, Thanpimon Buamanee, Yuki Uranishi

机构 * Graduate School of Information Science and Technology, The University of Osaka(信息科学与技术研究生学校,大阪大学) D3 Center, The University of Osaka(大阪大学D3中心) Graduate School of Maritime Sciences, Kobe University(海洋科学研究生学校, Kobe大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19599 2025-07-29 cs.CV 57%

Object-centric Video Question Answering with Visual Grounding and Referring

Haochen Wang, Qirui Chen, Cilin Yan, Jiayin Cai, Xiaolong Jiang, Yao Hu, Weidi Xie, Stratis Gavves

机构 * University of Amsterdam(阿姆斯特丹大学) SAI, Shanghai Jiao Tong University(上海交通大学SAI研究所) Xiaohongshu Inc(小红书公司)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03492 2025-07-29 cs.CV 57%

Find First, Track Next: Decoupling Identification and Propagation in Referring Video Object Segmentation

Suhwan Cho, Seunghoon Lee, Minhyeok Lee, Jungho Lee, Sangyoun Lee

机构 * GenGenAI Yonsei University(延世大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments ICCVW 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18634 2025-07-25 cs.CV 57%

Captain Cinema: Towards Short Movie Generation

Junfei Xiao, Ceyuan Yang, Lvmin Zhang, Shengqu Cai, Yang Zhao, Yuwei Guo, Gordon Wetzstein, Maneesh Agrawala, Alan Yuille, Lu Jiang

机构 * Johns Hopkins University(约翰霍普金斯大学) Stanford University(斯坦福大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Under review. Project page: https://thecinema.ai

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18342 2025-07-25 cs.CV 57%

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs

Yuping He, Yifei Huang, Guo Chen, Baoqi Pei, Jilan Xu, Tong Lu, Jiangmiao Pang

机构 * Nanjing University(南京大学) Shanghai AI Laboratory(上海人工智能实验室) The University of Tokyo(东京大学) Zhejiang University(浙江大学) Fudan University(复旦大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00566 2025-07-25 cs.CV 57%

Zero-Shot Skeleton-Based Action Recognition With Prototype-Guided Feature Alignment

Kai Zhou, Shuhai Zhang, Zeng You, Jinwu Hu, Mingkui Tan, Fei Liu

机构 * School of Software Engineering, South China University of Technology(南方科技大学软件工程学院) South China University of Technology(南方科技大学) Pazhou Lab(琶洲实验室) School of Future Technology, South China University of Technology(未来技术学院) Peng Cheng Laboratory(鹏城实验室) Key Laboratory of Big Data and Intelligent Robot (South China University of Technology), Ministry of Education(大数据与智能机器人重点实验室)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments This paper is accepted by IEEE TIP 2025 (The journal version is available at https://doi.org/10.1109/TIP.2025.3586487). Code is publicly available at https://github.com/kaai520/PGFA

Journal ref IEEE Transactions on Image Processing 34 (2025) 4602-4617

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16363 2025-07-23 cs.LG cs.MM 57%

Bipartite Patient-Modality Graph Learning with Event-Conditional Modelling of Censoring for Cancer Survival Prediction

Hailin Yue, Hulin Kuang, Jin Liu, Junjian Li, Lanlan Wang, Mengshen He, Jianxin Wang

机构 * Hunan Provincial Key Lab on Bioinformatics, School of Computer Science and Engineering, Central South University(湖南省级生物信息学重点实验室,计算机科学与工程学院,中南大学) Xinjiang Engineering Research Center of Big Data and Intelligent Software, School of Software, Xinjiang University(新疆大数据与智能软件工程研究中心,软件学院,新疆大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22139 2025-07-23 cs.CV 57%

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs

Shaojie Zhang, Jiahui Yang, Jianqin Yin, Zhenbo Luo, Jian Luan

机构 * MiLM Plus, Xiaomi Inc.(小米公司)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14301 2025-07-22 cs.IR cs.CV cs.DB 57%

LOVO: Efficient Complex Object Query in Large-Scale Video Datasets

Yuxin Liu, Yuezhang Peng, Hefeng Zhou, Hongze Liu, Xinyu Lu, Jiong Lou, Chentao Wu, Wei Zhao, Jie Li

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments @inproceedings{liu2025lovo,title={LOVO: Efficient Complex Object Query in Large-Scale Video Datasets},author={Liu, Yuxin and Peng, Yuezhang and Zhou, Hefeng and Liu, Hongze and Lu, Xinyu and Lou, Jiong and Wu, Chentao and Zhao, Wei and Li, Jie},booktitle={2025 IEEE 41st International Conference on Data Engineering (ICDE)},pages={1938--1951},year={2025},organization={IEEE Computer Society}}

Journal ref 2025 IEEE 41st International Conference on Data Engineering (ICDE)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12628 2025-07-18 cs.CV 57%

Funnel-HOI: Top-Down Perception for Zero-Shot HOI Detection

Sandipan Sarma, Agney Talwarr, Arijit Sur

机构 * Department of Computer Science and Engineering, Indian Institute of Technology, Guwahati, Assam(计算机科学与工程系,印度理工学院,古瓦哈蒂,阿萨姆)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments 10 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23765 2025-07-18 cs.CV 57%

STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?

Yun Li, Yiming Zhang, Tao Lin, Xiangrui Liu, Wenxiao Cai, Zheng Liu, Bo Zhao

机构 * School of AI, Shanghai Jiao Tong University(上海交通大学人工智能学院) China University of Geosciences(中国地质大学) Nanyang Technological University(南洋理工大学) BAAI(百度人工智能研究院) Stanford University(斯坦福大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11579 2025-07-17 cs.CV 57%

Vamba: Understanding Hour-Long Videos with Hybrid Mamba-Transformers

Weiming Ren, Wentao Ma, Huan Yang, Cong Wei, Ge Zhang, Wenhu Chen

机构 * University of Waterloo(滑铁卢大学) University of Toronto(多伦多大学) Kuaishou Technology(快手科技) Vector Institute(向量研究所) M-A-P

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments ICCV 2025 Camera Ready Version. Project Page: https://tiger-ai-lab.github.io/Vamba/

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10745 2025-07-17 cs.CV 57%

Bridging the Skeleton-Text Modality Gap: Diffusion-Powered Modality Alignment for Zero-shot Skeleton-based Action Recognition

Jeonghyeok Do, Munchurl Kim

机构 * Korea Advanced Institute of Science and Technology(韩国科学技术院)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments ICCV 2025 (camera-ready version). Please visit our project page at https://kaist-viclab.github.io/TDSM_site/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20254 2025-07-16 cs.CV 57%

Recognizing Surgical Phases Anywhere: Few-Shot Test-time Adaptation and Task-graph Guided Refinement

Kun Yuan, Tingxuan Chen, Shi Li, Joel L. Lavanchy, Christian Heiliger, Ege Özsoy, Yiming Huang, Long Bai, Nassir Navab, Vinkle Srivastav, Hongliang Ren, Nicolas Padoy

机构 * University of Strasbourg(斯特拉斯堡大学) CNRS(法国国家科学研究中心) INSERM(法国国家健康与医学研究院) ICube(ICube研究中心) UMR7357(法国大学-研究中心7357) IHU Strasbourg(斯特拉斯堡IHU医院) University Digestive Health Care Center – Clarunis(大学消化健康中心 – Clarunis) Ludwig Maximilian University of Munich(慕尼黑路易斯·马克西米利安大学) Chinese University of Hong Kong(香港中文大学) Munich Center for Machine Learning(慕尼黑机器学习中心)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted by MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09403 2025-07-15 cs.IR cs.MM 57%

Balancing Semantic Relevance and Engagement in Related Video Recommendations

Amit Jaspal, Feng Zhang, Wei Chang, Sumit Kumar, Yubo Wang, Roni Mittleman, Qifan Wang, Weize Mao

专题命中 视频多模态 :multimodal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.05357 2025-07-15 eess.IV cs.CV 57%

Unmixing Optical Signals from Undersampled Volumetric Measurements by Filtering the Pixel Latent Variables

Catherine Bouchard, Andréanne Deschênes, Vincent Boulanger, Jean-Michel Bellavance, Julia Chabbert, Alexy Pelletier-Rioux, Flavie Lavoie-Cardinal, Christian Gagné

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments 42 pages, 9 figures (main paper) + 22 pages, 15 figures (supplementary material)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17224 2025-07-14 cs.CV 57%

Visual and Textual Prompts in VLLMs for Enhancing Emotion Recognition

Zhifeng Wang, Qixuan Zhang, Peter Zhang, Wenjia Niu, Kaihao Zhang, Ramesh Sankaranarayana, Sabrina Caldwell, Tom Gedeon

机构 * School of Computing, Australian National University (ANU)(澳大利亚国立大学计算机学院) Quriosity Pty Ltd Human-Centric Advancements Chair in AI, Curtin University and Australian National University(Curtin大学人工智能人本发展主席职位,澳大利亚国立大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted by IEEE TCSVT

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07415 2025-07-11 cs.CV 57%

EPIC: Efficient Prompt Interaction for Text-Image Classification

Xinyao Yu, Hao Sun, Zeyu Ling, Ziwei Niu, Zhenjia Bai, Rui Qin, Yen-Wei Chen, Lanfen Lin

机构 * College of Computer Science and Technology, Zhejiang University, Hangzhou, China(浙江大学计算机科学与技术学院) College of Information Science, Ritsumeikan University, Shiga, Japan(立命馆大学信息科学学院)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments arXiv admin note: substantial text overlap with arXiv:2401.14856

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07262 2025-07-11 cs.CV 57%

DisenQ: Disentangling Q-Former for Activity-Biometrics

Shehreen Azad, Yogesh S Rawat

机构 * Center for Research in Computer Vision(计算机视觉研究中心) University of Central Florida(中央佛罗里达大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted in ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06531 2025-07-10 cs.CV 57%

ILNet: Trajectory Prediction with Inverse Learning Attention for Enhancing Intention Capture

Mingjin Zeng, Nan Ouyang, Wenkang Wan, Lei Ao, Qing Cai, Kai Sheng

机构 * Key Laboratory of Collaborative Intelligence Systems, Ministry of Education, Xidian University(协同智能系统重点实验室,教育部,西安电子科技大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05336 2025-07-08 cs.CV 57%

VideoMolmo: Spatio-Temporal Grounding Meets Pointing

Ghazi Shazan Ahmad, Ahmed Heakl, Hanan Gani, Abdelrahman Shaker, Zhiqiang Shen, Fahad Shahbaz Khan, Salman Khan

机构 * Mohamed Bin Zayed University of Artificial Intelligence(莫扎德·本·扎耶德人工智能大学) Linköping University(林奈大学) Australian National University(澳大利亚国立大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments 20 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏