arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

2025-09-05 至 2025-09-05 共收录 4 信号源:cs.CV, eess.IV, cs.MM

1. 视频生成 1 篇

2410.13287 2025-09-05 cs.LG 50%

PAK-UCB Contextual Bandit: An Online Learning Approach to Prompt-Aware Selection of Generative Models and LLMs

Xiaoyan Hu, Ho-fung Leung, Farzan Farnia

机构 * Department of Computer Science and Engineering, The Chinese University of Hong Kong(香港中文大学计算机科学与工程系) Independent Researcher(独立研究者)

专题命中 视频生成 :video generation(abstract)

Comments accepted to ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视频扩散模型 1 篇

2509.03680 2025-09-05 cs.GR cs.AI cs.CV 79%

LuxDiT: Lighting Estimation with Video Diffusion Transformer

Ruofan Liang, Kai He, Zan Gojcic, Igor Gilitschenski, Sanja Fidler, Nandita Vijaykumar, Zian Wang

机构 * NVIDIA University of Toronto(多伦多大学) Vector Institute(向量研究所)

专题命中 视频扩散模型 :video diffusion(title,abstract);分类 cs.CV

Comments Project page: https://research.nvidia.com/labs/toronto-ai/LuxDiT/

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 动作与事件理解 1 篇

2509.03883 2025-09-05 cs.CV cs.MM 81%

Human Motion Video Generation: A Survey

Haiwei Xue, Xiangyang Luo, Zhanghao Hu, Xin Zhang, Xunzhi Xiang, Yuqin Dai, Jianzhuang Liu, Zhensong Zhang, Minglei Li, Jian Yang, Fei Ma, Zhiyong Wu, Changpeng Yang, Zonghong Dai, Fei Richard Yu

机构 * Tsinghua University(清华大学) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究所) Huawei Noah’s Ark Lab(华为诺亚实验室) School of Mathematics and Statistics, Xi’an Jiaotong University(西安交通大学数学与统计学学院) Artificial Intelligence Innovation and Incubation (Al’) Institute of Fudan University(复旦大学人工智能创新与孵化院) University of Chinese Academy of Sciences(中国科学院大学) PCA Lab, Nanjing University of Science and Technology(南京理工大学PCA实验室) Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(广东人工智能与数字经济实验室(深圳)) Shenzhen University and Carleton University(深圳大学和卡尔顿大学)

专题命中 动作与事件理解 :video generation(title,abstract);分类 cs.CV、cs.MM

Comments Accepted by TPAMI. Github Repo: https://github.com/Winn1y/Awesome-Human-Motion-Video-Generation IEEE Access: https://ieeexplore.ieee.org/document/11106267

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 长视频与时序推理 1 篇

2509.04450 2025-09-05 cs.CV cs.LG 83%

Virtual Fitting Room: Generating Arbitrarily Long Videos of Virtual Try-On from a Single Image -- Technical Preview

Jun-Kun Chen, Aayush Bansal, Minh Phuoc Vo, Yu-Xiong Wang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) SpreeAI

专题命中 长视频与时序推理 :long video(title,abstract);video generation(abstract);分类 cs.CV

Comments Project Page: https://immortalco.github.io/VirtualFittingRoom/

详情

展开后加载摘要…

URL PDF HTML 收藏