arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

共收录 1390 信号源:cs.CV, eess.IV, cs.MM

1. 视频理解 1390 篇

1610.01376 2016-11-11 cs.CV 57%

Recognizing and Presenting the Storytelling Video Structure with Deep Multimodal Networks

Lorenzo Baraldi, Costantino Grana, Rita Cucchiara

专题命中 视频理解 :long video(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1607.05177 2016-07-19 cs.CV 57%

Query-Focused Extractive Video Summarization

Aidean Sharghi, Boqing Gong, Mubarak Shah

专题命中 视频理解 :long video(abstract);分类 cs.CV

Comments Accepted to ECCV 2016

详情

展开后加载摘要…

URL PDF HTML 收藏
1603.09439 2016-04-04 cs.CV 57%

The Open World of Micro-Videos

Phuc Xuan Nguyen, Gregory Rogez, Charless Fowlkes, Deva Ramanan

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1512.07314 2015-12-24 cs.CV 57%

Mid-level Representation for Visual Recognition

Moin Nabi

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1505.03825 2015-05-15 cs.CV 57%

Unsupervised Object Discovery and Tracking in Video Collections

Suha Kwak, Minsu Cho, Ivan Laptev, Jean Ponce, Cordelia Schmid

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1501.03069 2015-02-10 cs.CV 57%

Learning from Multiple Sources for Video Summarisation

Xiatian Zhu, Chen Change Loy, Shaogang Gong

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1501.07738 2015-02-02 cs.CV 57%

Co-Regularized Deep Representations for Video Summarization

Olivier Morère, Hanlin Goh, Antoine Veillard, Vijay Chandrasekhar, Jie Lin

专题命中 视频理解 :long video(abstract);分类 cs.CV

Comments Video summarization, deep convolutional neural networks, co-regularized restricted Boltzmann machines

详情

展开后加载摘要…

URL PDF HTML 收藏
1411.0085 2014-11-04 cs.CV 57%

Complex Events Recognition under Uncertainty in a Sensor Network

Atul Kanaujia, Tae Eun Choe, Hongli Deng

专题命中 视频理解 :long video(abstract);分类 cs.CV

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.21267 2026-07-24 cs.AI 新提交 50%

BasketEvent: Understanding Who Did What and When in Basketball Videos

BasketEvent:理解篮球视频中谁在何时做了什么

Yu Zhang, Jiayuan Rao, Haoning Wu, Weidi Xie

机构 * School of Artificial Intelligence, Shanghai Jiao Tong University(上海交通大学人工智能学院)

专题命中 视频理解 :video understanding(abstract)

AI总结 该研究旨在解决篮球视频理解问题,提出以球员为中心的BasketEvent数据集,引入PlayNet推理框架,通过建模多种互动并聚合时间证据进行事件预测,实验证明其在体育视频理解上优于基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15859 2026-06-16 cs.AR 新提交 50%

EPIC: A System Framework for Efficient Egocentric Perception on Embodied AR Glasses

EPIC:面向具身AR眼镜的高效自我中心感知系统框架

Tianhua Xia, Haiyu Wang, Jiajing Zheng, Su Chen, Sai Qian Zhang

专题命中 视频理解 :video understanding(abstract)

AI总结 提出EPIC系统,通过算法-硬件协同优化,利用注视、姿态和惯性信号推断用户意图,仅保留高分辨率感知输入中最具信息量的部分,平均减少27.5倍内存占用和24.3倍能耗,同时保持自我中心视频理解任务的智能辅助精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15330 2026-06-16 cs.IR 新提交 50%

OneBar: An End-to-End Content-Grounded Generative Query Recommendation Framework for E-Commerce Video Feeds

OneBar:一种面向电商视频流的端到端内容驱动生成式查询推荐框架

Yao Tang, Ying Yang, Ben Chen, Yufei Ma, Zihan Liang, Chenyi Lei, Wenwu Ou, Jian Liu

专题命中 视频理解 :video understanding(abstract)

AI总结 提出OneBar框架,通过协同多模态意图对齐、统一端到端架构和渐进偏好学习,解决短视频平台查询推荐中的延迟和偏好漂移问题,显著提升查询曝光和点击。

Comments Any questions feel free to contact: benchen4395@gmail.com

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.07049 2026-06-05 cs.NE cs.LG 50%

GridPE: A Grid Cell-Inspired Unified Position Embedding for Arbitrary-Dimensional Spaces

GridPE: 一种基于网格细胞的统一位置嵌入方法用于任意维度空间

Boyang Li, Yulin Wu, Nuoxian Huang, Wenjia Zhang

机构 * New York University(纽约大学) Peking University(北京大学) Imperial College London(伦敦帝国学院) Tongji University(同济大学)

专题命中 视频理解 :video understanding(abstract)

AI总结 本文提出GridPE,一种受哺乳动物空间认知中六边形周期编码启发的新型位置嵌入框架,旨在解决高维时空任务中位置嵌入的理论保障问题,通过结合计算神经科学原理和调和分析,为任意维度空间提供统一的位置嵌入解决方案。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.04596 2026-06-04 cs.CL 50%

A Systematic Evaluation of Positional Bias in Multi-Video Summarization with MLLMs

多视频摘要中位置偏差的系统评估:基于多模态大语言模型

Huangchen Xu, Yuan Wu, Yi Chang

机构 * School of Artificial Intelligence, Jilin University(吉林大学人工智能学院) Engineering Research Center of Knowledge-Driven Human-Machine Intelligence, Jilin University(知识驱动人机智能工程研究中心) International Center of Future Science, Jilin University(未来科学国际中心)

专题命中 视频理解 :video understanding(abstract)

AI总结 本研究系统评估了多模态大语言模型在多视频摘要任务中的位置偏差,通过构建基准和三种互补指标揭示了领域与模型依赖的偏差特性,并分析了提示级缓解方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04804 2026-05-14 cs.CL 50%

OmniSIFT: Modality-Asymmetric Token Compression for Efficient Omni-modal Large Language Models

OmniSIFT: 多模态非对称令牌压缩用于高效的多模态大语言模型

Yue Ding, Yiyan Ji, Jungang Li, Xuyang Liu, Xinlong Chen, Junfei Wu, Bozhou Li, Bohan Zeng, Yang Shi, Yushuo Guan, Yuanxing Zhang, Jiaheng Liu, Qiang Liu, Pengfei Wan, Liang Wang

机构 * New Laboratory of Pattern Recognition (NLPR), Institute of Automation, Chinese Academy of Sciences (CASIA)(模式识别新实验室(NLPR)、自动化研究所、中国科学院(CASIA)) Nanjing University(南京大学) The Hong Kong University of Science(香港科学大学) Sichuan University(四川大学) Peking University(北京大学)

专题命中 视频理解 :video understanding(abstract)

AI总结 OmniSIFT通过非对称令牌压缩框架提升多模态大语言模型效率,采用时空视频修剪和视觉引导音频选择模块,实现低参数高鲁棒性。

Comments [ICML 2026] Code Link: https://github.com/dingyue772/OmniSIFT

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05650 2026-04-10 cs.CL 50%

See the Forest for the Trees: Loosely Speculative Decoding via Visual-Semantic Guidance for Efficient Inference of Video LLMs

见林木而非树木:通过视觉-语义引导实现松散推测解码以提高视频大语言模型的推理效率

Yicheng Ji, Jun Zhang, Jinpeng Chen, Cong Wang, Lidan Shou, Gang Chen, Huan Li

机构 * ZJU(浙江大学) BUPT(北京邮电大学)

专题命中 视频理解 :video understanding(abstract)

AI总结 本文提出LVSpec,一种无需训练的松散推测解码框架,通过视觉相关锚点识别和位置移位容忍机制,提升视频大语言模型的推理速度和精度。

Comments ACL'2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16397 2026-03-18 cs.CL cs.AI 50%

Fanar 2.0: Arabic Generative AI Stack

Fanar 2.0:阿拉伯生成AI堆栈

FANAR TEAM, Ummar Abbas, Mohammad Shahmeer Ahmad, Minhaj Ahmad, Abdulaziz Al-Homaid, Anas Al-Nuaimi, Enes Altinisik, Ehsaneddin Asgari, Sanjay Chawla, Shammur Chowdhury, Fahim Dalvi, Kareem Darwish, Nadir Durrani, Mohamed Elfeky, Ahmed Elmagarmid, Mohamed Eltabakh, Asim Ersoy, Masoomali Fatehkia, Mohammed Qusay Hashim, Majd Hawasly, Mohamed Hefeeda, Mus'ab Husaini, Keivin Isufaj, Soon-Gyo Jung, Houssam Lachemat, Ji Kim Lucas, Abubakr Mohamed, Tasnim Mohiuddin, Basel Mousi, Hamdy Mubarak, Ahmad Musleh, Mourad Ouzzani, Amin Sadeghi, Husrev Taha Sencar, Mohammed Shinoy, Omar Sinan, Yifan Zhang

机构 * Qatar Computing Research Institute (QCRI)(卡塔尔计算研究 institute) Hamad Bin Khalifa University(哈马德·本·卡西姆大学)

专题命中 视频理解 :video understanding(abstract)

AI总结 Fanar 2.0在资源受限条件下实现卓越性能,通过数据质量优先、持续预训练和模型融合,提升阿拉伯语知识、语言、方言及英语能力,同时引入多项新功能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19184 2026-02-24 cs.RO 50%

Human-to-Robot Interaction: Learning from Video Demonstration for Robot Imitation

人机交互:从视频演示中学习机器人模仿

Thanh Nguyen Canh, Thanh-Tuan Tran, Haolan Zhang, Ziyan Gao, Nak Young Chong, Xiem HoangVan

专题命中 视频理解 :video understanding(abstract)

AI总结 本研究提出了一种基于视频演示的机器人模仿学习方法,通过模块化框架结合时间位移模块和深度强化学习,实现机器人从无结构视频中学习基本操作技能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10986 2026-02-12 cs.LG 50%

TVCACHE: A Stateful Tool-Value Cache for Post-Training LLM Agents

TVCACHE: 一种具有状态的工具-价值缓存用于训练后LLM代理

Abhishek Vijaya Kumar, Bhaskar Kataria, Byungsoo Oh, Emaad Manzoor, Rachee Singh

机构 * cornell(康奈尔大学)

专题命中 视频理解 :video understanding(abstract)

AI总结 TVCACHE通过维护工具调用序列树和最长前缀匹配,实现LLM代理训练后的高效缓存,提高缓存命中率并减少工具调用时间。

Comments Abhishek Vijaya Kumar and Bhaskar Kataria have equal contribution

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10015 2026-02-12 cs.RO cs.AI 50%

RoboSubtaskNet: Temporal Sub-task Segmentation for Human-to-Robot Skill Transfer in Real-World Environments

RoboSubtaskNet: 人类-机器人技能转移中的时间子任务分割用于现实环境

Dharmendra Sharma, Archit Sharma, John Rebeiro, Vaibhav Kesharwani, Peeyush Thakur, Narendra Kumar Dhar, Laxmidhar Behera

专题命中 视频理解 :video understanding(abstract)

AI总结 RoboSubtaskNet通过结合增强注意力的I3D特征和改进的MS-TCN,实现了人类-机器人技能转移中的时间子任务分割,提升了现实环境中的机器人操作性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16231 2026-01-26 cs.SD cs.AI cs.CL cs.LG eess.AS 50%

SoundBreak: A Systematic Study of Audio-Only Adversarial Attacks on Trimodal Models

SoundBreak: 对三模态模型上仅音频对抗攻击的系统研究

Aafiya Hussain, Gaurav Srivastava, Alvi Ishmam, Zaber Hakim, Chris Thomas

机构 * Department of Computer Science, Virginia Tech, USA(计算机科学系,弗吉尼亚理工大学)

专题命中 视频理解 :video-language(abstract)

AI总结 SoundBreak研究了对三模态模型的仅音频对抗攻击,发现音频扰动可导致严重多模态失败,攻击成功率高达96%,并揭示了多模态系统中被忽视的单模态攻击面。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08871 2026-01-15 cs.SD cs.AI eess.AS 50%

Semantic visually-guided acoustic highlighting with large vision-language models

语义视觉引导的音频突出与大型视觉-语言模型

Junhua Huang, Chao Huang, Chenliang Xu

机构 * University of Rochester(罗切斯特大学)

专题命中 视频理解 :video understanding(abstract)

AI总结 本文提出利用大型视觉-语言模型提取视觉-语义特征,以提升音频混音质量,发现摄像机焦点、语气和场景背景对感知混音质量提升最显著。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20876 2026-01-13 cs.RO 50%

Proprioception Enhances Vision Language Model in Generating Captions and Subtask Segmentations for Robot Task

本体感知增强视觉语言模型在为机器人任务生成描述和子任务分割中的应用

Kanata Suzuki, Shota Shimizu, Tetsuya Ogata

机构 * Faculty of Science and Engineering, Waseda University(工学部,早稻田大学) Artificial Intelligence Laboratory, Fujitsu Limited(Fujitsu 人工智能实验室) National Institute of Advanced Industrial Science and Technology(国家先进工业科学与技术研究院)

专题命中 视频理解 :video understanding(abstract)

AI总结 本研究通过引入本体感知数据,提升视觉语言模型在机器人任务描述和子任务分割中的性能,以增强机器人模仿学习效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05472 2025-12-08 cs.NE 50%

Unleashing Temporal Capacity of Spiking Neural Networks through Spatiotemporal Separation

通过时空分离释放脉冲神经网络的时间能力

Yiting Dong, Zhaofei Yu, Jianhao Ding, Zijie Xu, Tiejun Huang

专题命中 视频理解 :video understanding(abstract)

AI总结 本文提出STSep网络,通过解耦空间与时间分支,提升脉冲神经网络在视频理解中的时空建模能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11664 2025-11-03 cs.AI 50%

VRoPE: Rotary Position Embedding for Video Large Language Models

Zikang Liu, Longteng Guo, Yepeng Tang, Tongtian Yue, Junxian Cai, Kai Ma, Qingbin Liu, Xi Chen, Jing Liu

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) School of Computer Science and Technology, Beijing Jiaotong University(北京交通大学计算机科学与技术学院) Basic Algorithm Center, Tencent(腾讯基础算法中心)

专题命中 视频理解 :video understanding(abstract)

Comments EMNLP 2025 Main Camera Ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21026 2025-10-27 cs.RO 50%

HRT1: One-Shot Human-to-Robot Trajectory Transfer for Mobile Manipulation

Sai Haneesh Allu, Jishnu Jaykumar P, Ninad Khargonkar, Tyler Summers, Jian Yao, Yu Xiang

机构 * University of Texas at Dallas(德克萨斯大学达拉斯分校) XPeng(小鹏)

专题命中 视频理解 :video understanding(abstract)

Comments 14 pages, 11 figures and 3 tables. Project page is available at \url{https://irvlutd.github.io/HRT1/}

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23870 2025-08-12 cs.LG cs.AI 50%

MaCP: Minimal yet Mighty Adaptation via Hierarchical Cosine Projection

Yixian Shen, Qi Bi, Jia-Hong Huang, Hongyi Zhu, Andy D. Pimentel, Anuj Pathania

专题命中 视频理解 :video understanding(abstract)

Comments This work was intended as a replacement of arXiv:2410.09103 and any subsequent updates will appear there

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.09103 2025-08-12 cs.LG cs.AI 50%

MaCP: Minimal yet Mighty Adaptation via Hierarchical Cosine Projection

Yixian Shen, Qi Bi, Jia-Hong Huang, Hongyi Zhu, Andy D. Pimentel, Anuj Pathania

机构 * University of Amsterdam(阿姆斯特丹大学)

专题命中 视频理解 :video understanding(abstract)

Comments 17 pages; Previously this version appeared as arXiv:2505.23870 which was submitted as a new work by accident

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18805 2025-04-29 cs.CL cs.AI cs.LG 50%

Stealing Creator's Workflow: A Creator-Inspired Agentic Framework with Iterative Feedback Loop for Improved Scientific Short-form Generation

Jong Inn Park, Maanas Taneja, Qianwen Wang, Dongyeop Kang

机构 * University of Minnesota(明尼苏达大学)

专题命中 视频理解 :video generation(abstract)

Comments Project page: https://minnesotanlp.github.io/scitalk-project-page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10742 2025-04-25 cs.LG cs.CL 50%

Keyframe-oriented Vision Token Pruning: Enhancing Efficiency of Large Vision Language Models on Long-Form Video Processing

Yudong Liu, Jingwei Sun, Yueqian Lin, Jingyang Zhang, Ming Yin, Qinsi Wang, Jianyi Zhang, Hai Li, Yiran Chen

机构 * Duke University(杜克大学) Duke Kunshan University(杜克昆山大学)

专题命中 视频理解 :video understanding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14692 2025-04-22 cs.CL 50%

OmniV-Med: Scaling Medical Vision-Language Model for Universal Visual Understanding

Songtao Jiang, Yuan Wang, Sibo Song, Yan Zhang, Zijie Meng, Bohan Lei, Jian Wu, Jimeng Sun, Zuozhu Liu

机构 * Zhejiang University(浙江大学) Alibaba Group(阿里巴巴集团) UIUC(伊利诺伊大学香槟分校)

专题命中 视频理解 :video understanding(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏