VARPose: Flexible 2D Pose Densification via Visual Autoregressive Modeling for Enhanced 3D Lifting
VARPose:基于视觉自回归建模的灵活2D姿态密集化以提升3D姿态提升性能
Kaiyuan Pu, Tiantian Yang, Dan Zeng
机构
*
School of Artificial Intelligence, Sun Yat-sen University(中山大学人工智能学院)
;
The Technology Innovation Center for Collaborative Applications of Natural Resources Data in GBA, MNR(自然资源部粤港澳大湾区自然资源数据协同应用技术创新中心)
Geometry-guided Emotion Modulation for Controllable and Photorealistic Emotional Talking Face Generation
几何引导的情感调制用于可控且照片级真实感的情感说话人脸生成
Chenggong Hu, Shaoyin Ma, Yi Wang, Li Sun, Mingli Song, Jie Song
机构
*
School of Software Technology, Zhejiang University(浙江大学软件学院)
;
College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院)
;
Ningbo Global Innovation Center, Zhejiang University(浙江大学宁波全球创新中心)
Retrieval-Driven Training-Free AI-Generated Video Attribution
检索驱动的无训练AI生成视频溯源
Renxi Cheng, Chaolei Han, Jie Gui, Hongsong Wang
机构
*
School of Cyber Science and Engineering, Southeast University(东南大学网络空间科学与工程学院)
;
Purple Mountain Laboratories(紫金山实验室)
;
School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院)
TAP-RAG: Task-Aware Policy Control for Long-Document Multimodal Question Answering
TAP-RAG:用于长文档多模态问答的任务感知策略控制
Zhong Ji, Keqi Jin, Yan Zhang, Jiasheng Li
机构
*
School of Electrical and Information Engineering, Tianjin University(天津大学电气与信息工程学院)
;
The International Joint Institute of Tianjin University(天津大学国际联合学院)
;
Department of Automation, Tsinghua University(清华大学自动化系)
机构
*
School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院)
;
Shenzhen Loop Area Institute(深圳河套学院)
;
Guangdong Key Laboratory of Big Data Analysis and Processing(广东省大数据分析与处理重点实验室)
;
Baidu Inc.(百度公司)
;
LightSpeed Studios, Tencent America(美国光速工作室,腾讯)
;
School of Computer Science and Artificial Intelligence, Lanzhou University of Technology(兰州理工大学计算机科学与人工智能学院)
Video Generation Models are General-Purpose Vision Learners
视频生成模型是通用视觉学习者
Letian Wang, Chuhan Zhang, Rishabh Kabra, Jasper Uijlings, Steven Waslander, Andrew Zisserman, Joao Carreira, Kaiming He, Misha Andriluka, Eduard Gabriel Bazavan, Andrei Zanfir, Cristian Sminchisescu
机构
*
Faculty of Applied Sciences, Macao Polytechnic University(澳门理工学院应用科学学院)
;
The Chinese University of Hong Kong(香港中文大学)
;
AgentecFusion Limited(AgentecFusion有限公司)
;
School of Information and Communications Engineering, Xi’an Jiaotong University(西安交通大学信息与通信工程学院)
;
School of Computer Science, Northwestern Polytechnical University(西北工业大学计算机学院)
机构
*
Southern University of Science and Technology(南方科技大学)
;
Spatialtemporal AI(时空人工智能)
;
Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))
;
University of Sheffield(谢菲尔德大学)