arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4735 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4735 篇

2501.13805 2025-07-08 cs.CV 57%

mmEgoHand: Egocentric Hand Pose Estimation and Gesture Recognition with Head-mounted Millimeter-wave Radar and IMU

Yizhe Lv, Tingting Zhang, Zhijian Wang, Yunpeng Song, Han Ding, Jinsong Han, Fei Wang

机构 * School of Software Engineering, Xi’an Jiaotong University(软件工程学院,西安交通大学) School of Cyber Science and Engineering, Xi’an Jiaotong University(网络科学与工程学院,西安交通大学) School of Computer Science and Technology, Xi’an Jiaotong University(计算机科学与技术学院,西安交通大学) College of Computer Science and Technology, Zhejiang University(计算机科学与技术学院,浙江大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments 11 pages, Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01492 2025-07-03 cs.CV 57%

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization

Jiyang Tang, Hengyi Li, Yifan Du, Wayne Xin Zhao

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院) College of Artificial Intelligence, Nankai University(南开大学人工智能学院) School of Computer Science and Technology, Beijing Institute of Technology(北京理工大学计算机科学与技术学院)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.14607 2025-07-01 cs.CV 57%

ReferDINO: Referring Video Object Segmentation with Visual Grounding Foundations

Tianming Liang, Kun-Yu Lin, Chaolei Tan, Jianguo Zhang, Wei-Shi Zheng, Jian-Fang Hu

机构 * Sun Yat-sen University(中山大学) Southern University of Science and Technology(南方科技大学)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted to ICCV 2025. Project page: \url{https://isee-laboratory.github.io/ReferDINO}

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17690 2025-07-01 cs.CV 57%

CountLLM: Towards Generalizable Repetitive Action Counting via Large Language Model

Ziyu Yao, Xuxin Cheng, Zhiqi Huang, Lei Li

机构 * Peking University(北京大学) University of Washington(华盛顿大学) University of Copenhagen(哥本哈根大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted by CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10432 2025-06-30 cs.LG cs.CL 57%

BeamLLM: Vision-Empowered mmWave Beam Prediction with Large Language Models

Can Zheng, Jiguang He, Guofa Cai, Zitong Yu, Chung G. Kang

机构 * School of Electrical Engineering, Korea University(韩国大学电气工程学院) School of Computing and Information Technology, Great Bay University(大湾大学计算与信息科技学院) School of Information Engineering, Guangdong University of Technology(广东技术大学信息工程学院)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CL

Comments 6 pages, 7 figures, conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21317 2025-06-27 cs.CV 57%

LLaVA-Pose: Enhancing Human Pose and Action Understanding via Keypoint-Integrated Instruction Tuning

Dewen Zhang, Tahir Hussain, Wangpeng An, Hayaru Shouno

机构 * Department of Informatics, Graduate School of Informatics and Engineering, The University of Electro-Communications(信息学院、信息工程研究生院、东京电波通信大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments arXiv admin note: substantial text overlap with arXiv:2409.09306

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15808 2025-06-27 cs.CV 57%

ClearSight: Human Vision-Inspired Solutions for Event-Based Motion Deblurring

Xiaopeng Lin, Yulong Huang, Hongwei Ren, Zunchang Liu, Yue Zhou, Haotian Fu, Bojun Cheng

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20601 2025-06-26 cs.CV 57%

Video Perception Models for 3D Scene Synthesis

Rui Huang, Guangyao Zhai, Zuria Bauer, Marc Pollefeys, Federico Tombari, Leonidas Guibas, Gao Huang, Francis Engelmann

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17840 2025-06-24 cs.LG cs.AI 57%

Causal Spherical Hypergraph Networks for Modelling Social Uncertainty

Anoushka Harit, Zhongtian Sun

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17746 2025-06-24 cs.CV 57%

PhysID: Physics-based Interactive Dynamics from a Single-view Image

Sourabh Vasant Gothe, Ayon Chattopadhyay, Gunturi Venkata Sai Phani Kiran, Pratik, Vibhav Agarwal, Jayesh Rajkumar Vachhani, Sourav Ghosh, Parameswaranath VM, Barath Raj KR

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Published in 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Project page: https://physid.github.io/

Journal ref 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Hyderabad, India, 2025, pp. 1-5

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16701 2025-06-23 cs.CV 57%

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition

Xiaodan Hu, Chuhang Zou, Suchen Wang, Jaechul Kim, Narendra Ahuja

机构 * University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Amazon.com LLC(亚马逊公司)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16201 2025-06-23 cs.RO cs.CV 57%

FlowRAM: Grounding Flow Matching Policy with Region-Aware Mamba Framework for Robotic Manipulation

Sen Wang, Le Wang, Sanping Zhou, Jingyi Tian, Jiayi Li, Haowen Sun, Wei Tang

机构 * National Key Laboratory of Human-Machine Hybrid Augmented Intelligence(国家人类-机器混合增强智能重点实验室) National Engineering Research Center for Visual Information and Applications(国家视觉信息与应用工程研究中心) Institute of Artificial Intelligence and Robotics(人工智能与机器人研究所) Xi’an Jiaotong University(西安交通大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08343 2025-06-19 cs.CL 57%

Wait, We Don't Need to "Wait"! Removing Thinking Tokens Improves Reasoning Efficiency

Chenlong Wang, Yuanning Feng, Dongping Chen, Zhaoyang Chu, Ranjay Krishna, Tianyi Zhou

机构 * University College London(伦敦大学学院) University of Washington(华盛顿大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13956 2025-06-19 cs.CV 57%

Improving LLM Video Understanding with 16 Frames Per Second

Yixuan Li, Changli Tang, Jimin Zhuang, Yudong Yang, Guangzhi Sun, Wei Li, Zejun Ma, Chao Zhang

机构 * Tsinghua University(清华大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12811 2025-06-17 cs.LG cs.AI 57%

Flow-Based Policy for Online Reinforcement Learning

Lei Lv, Yunfei Li, Yu Luo, Fuchun Sun, Tao Kong, Jiafeng Xu, Xiao Ma

机构 * Tsinghua University(清华大学) Shanghai Research Institute for Intelligent Autonomous Systems,Tongji University(上海智能自主系统研究所,同济大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.03948 2025-06-17 cs.CV 57%

Reading Between the Lanes: Text VideoQA on the Road

George Tom, Minesh Mathew, Sergi Garcia, Dimosthenis Karatzas, C. V. Jawahar

机构 * Center for Visual Information Technology (CVIT)(视觉信息科技中心)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11040 2025-06-16 cs.LG cs.CL cs.ET 57%

Large Language models for Time Series Analysis: Techniques, Applications, and Challenges

Feifei Shi, Xueyan Yin, Kang Wang, Wanyu Tu, Qifu Sun, Huansheng Ning

机构 * School of Computer and Communication Engineering, University of Science and Technology Beijing(计算机与通信工程学院,北京科技大学)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09565 2025-06-16 cs.CV 57%

SemanticSplat: Feed-Forward 3D Scene Understanding with Language-Aware Gaussian Fields

Qijing Li, Jingxiang Sun, Liang An, Zhaoqi Su, Hongwen Zhang, Yebin Liu

机构 * Beijing Normal University(北京师范大学) Tsinghua University(清华大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15513 2025-06-11 cs.CV 57%

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler

Xingjian Zhang, Xi Weng, Yihao Yue, Zhaoxin Fan, Wenjun Wu, Lei Huang

机构 * SKLCCSE, Institute of Artificial Intelligence, Beihang University, Beijing, China(信息与通信工程学院,人工智能研究院,北京航空航天大学,北京,中国) Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing, Beihang University(未来区块链与隐私计算先进创新中心,北京航空航天大学) Hangzhou International Innovation Institute, Beihang University, Hangzhou, China(杭州国际创新研究院,北京航空航天大学,杭州,中国)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments code and training recipes are available at https://github.com/ZhangXJ199/TinyLLaVA-Video

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.03817 2025-06-11 cs.CV 57%

Markerless Multi-view 3D Human Pose Estimation: a survey

Ana Filipa Rodrigues Nogueira, Hélder P. Oliveira, Luís F. Teixeira

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments 26 pages, 10 tables, 6 figures, accepted at Image and Vision Computing (IMAVIS)

Journal ref In: Image and Vision Computing 155 (2025), p. 105437. issn: 0262-8856

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24718 2025-06-10 cs.CV 57%

Reinforcing Video Reasoning with Focused Thinking

Jisheng Dang, Jingze Wu, Teng Wang, Xuanhui Lin, Nannan Zhu, Hongbo Chen, Wei-Shi Zheng, Meng Wang, Tat-Seng Chua

机构 * Sun Yat-sen University(中山大学) Lanzhou University(兰州大学) University of Hong Kong(香港大学) Hefei University of Technology(合肥工业大学) National University of Singapore(新加坡国立大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16784 2025-06-10 cs.CV 57%

Four Eyes Are Better Than Two: Harnessing the Collaborative Potential of Large Models via Differentiated Thinking and Complementary Ensembles

Jun Xie, Xiongjun Guan, Yingjian Zhu, Zhaoran Zhao, Xinming Wang, Hongzhu Yi, Feng Chen, Zhepeng Wang

机构 * Lenovo Research(联想研究院) Tsinghua University(清华大学) School of Artificial Intelligence, University of Chinese Academy of Sciences (UCAS)(中国科学院大学人工智能学院) Institute of Automation, Chinese Academy of Sciences(CAS)(中国科学院自动化研究所) Zhongguancun Academy(中关村学院) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06918 2025-06-10 cs.CV cs.RO 57%

Reading in the Dark with Foveated Event Vision

Carl Brander, Giovanni Cioffi, Nico Messikommer, Davide Scaramuzza

机构 * Robotics and Perception Group, University of Zurich, Switzerland(苏黎世大学机器人与感知组)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments CVPR 2025 Workshop on Event-based Vision

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06359 2025-06-10 cs.LG cs.AI 57%

From Transformers to Large Language Models: A systematic review of AI applications in the energy sector towards Agentic Digital Twins

Gabriel Antonesi, Tudor Cioara, Ionut Anghel, Vasilis Michalakopoulos, Elissaios Sarmas, Liana Toderean

专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16652 2025-06-10 cs.CV cs.LG 57%

Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal Decoding

Feilong Tang, Chengzhi Liu, Zhongxing Xu, Ming Hu, Zelin Peng, Zhiwei Yang, Jionglong Su, Minquan Lin, Yifan Peng, Xuelian Cheng, Imran Razzak, Zongyuan Ge

机构 * Monash University(蒙纳士大学) MBZUAI XJTLU(西安交通大学) Shanghai Jiaotong University(上海交通大学) Fudan University(复旦大学) University of Minnesota(明尼苏达大学) Cornell University(康奈尔大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Clarification note for the CVPR 2025 paper (FarSight). Prepared by a subset of the original authors; remaining co-authors are acknowledged in the text

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06165 2025-06-09 cs.HC cs.AI 57%

(AI peers) are people learning from the same standpoint: Perception of AI characters in a Collaborative Science Investigation

Eunhye Grace Ko, Soo Hyoung Joo

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

Comments 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06128 2025-06-09 cs.CV 57%

CCLSTM: Coupled Convolutional Long-Short Term Memory Network for Occupancy Flow Forecasting

Peter Lengyel

机构 * aiMotive Budapest, Hungary(aiMotive布达佩斯,匈牙利)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05302 2025-06-06 cs.CV 57%

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos

Weifeng Lin, Xinyu Wei, Ruichuan An, Tianhe Ren, Tingwei Chen, Renrui Zhang, Ziyu Guo, Wentao Zhang, Lei Zhang, Hongsheng Li

机构 * CUHK(香港中文大学) HKU(香港大学) PolyU Peking University(北京大学)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments 19 pages, 13 figures, Website: https://Perceive-Anything.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05163 2025-06-06 cs.CV 57%

FRED: The Florence RGB-Event Drone Dataset

Gabriele Magrini, Niccolò Marini, Federico Becattini, Lorenzo Berlincioni, Niccolò Biondi, Pietro Pala, Alberto Del Bimbo

机构 * University of Florence(佛罗伦萨大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04983 2025-06-06 cs.CV 57%

TextVidBench: A Benchmark for Long Video Scene Text Understanding

Yangyang Zhong, Ji Qi, Yuan Yao, Pengxin Luo, Yunfeng Yan, Donglian Qi, Zhiyuan Liu, Tat-Seng Chua

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏