TV2TV: A Unified Framework for Interleaved Language and Video Generation
TV2TV:一种用于交错语言和视频生成的统一框架
Xiaochuang Han, Youssef Emad, Melissa Hall, John Nguyen, Karthik Padthe, Liam Robbins, Amir Bar, Delong Chen, Michal Drozdzal, Maha Elbayad, Yushi Hu, Shang-Wen Li, Sreya Dutta Roy, Jakob Verbeek, XuDong Wang, Marjan Ghazvininejad, Luke Zettlemoyer, Emily Dinan
Less Is More, but Where? Dynamic Token Compression via LLM-Guided Keyframe Prior
少即是多,但在哪里?通过LLM引导的关键帧先验实现动态令牌压缩
Yulin Li, Haokun Gui, Ziyang Fan, Junjie Wang, Bin Kang, Bin Chen, Zhuotao Tian
机构
*
Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))
;
Shenzhen Loop Area Institute(深圳河套学院)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
The Hong Kong University of Science and Technology(香港科技大学)
Sampling-Based Model Predictive Control for Dexterous Manipulation on a Biomimetic Tendon-Driven Hand
基于采样的模型预测控制在仿生腱驱动手的灵巧操作中的应用
Adrian Hess, Alexander M. Kübler, Benedek Forrai, Mehmet Dogar, Robert K. Katzschmann
机构
*
Soft Robotics Lab, Department of Mechanical and Process Engineering, ETH Zurich(苏黎世联邦理工学院机械与过程工程系软机器人实验室)
;
School of Computer Science, University of Leeds(利兹大学计算机科学学院)
专题命中
VLM训练与架构
:VLM(abstract);visual language model(abstract)
DREAM: Drafting with Refined Target Features and Entropy-Adaptive Cross-Attention Fusion for Multimodal Speculative Decoding
Yunhai Hu, Tianhua Xia, Zining Liu, Rahul Raman, Xingyu Liu, Bo Bao, Eric Sather, Vithursan Thangarasa, Sai Qian Zhang
机构
*
Courant Institute of Mathematical Sciences, New York University(纽约大学数学科学学院)
;
Tandon School of Engineering, New York University(纽约大学工程学院)
;
Cerebras Systems Inc.(Cerebras Systems公司)
;
University of Pennsylvania(宾夕法尼亚大学)
机构
*
Beijing University of Posts and Telecommunications(北京邮电大学)
;
Westlake University(西湖大学)
;
Zhejiang University(浙江大学)
;
OpenHelix Team(OpenHelix团队)
;
State Key Laboratory of Networking and Switching Technology(网络与交换技术国家重点实验室)
;
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
机构
*
University of Houston(德克萨斯大学)
;
University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
University of Connecticut(康涅狄格大学)
;
NEC Laboratories America(日本 NEC 美国实验室)
;
Florida International University(佛罗里达国际大学)
专题命中
VLM训练与架构
:vision language model(abstract);分类 cs.CV、cs.AI、cs.LG