Teaching Vision-Language-Action Models What to See and Where to Look
教视觉-语言-动作模型看什么和看哪里
Yuguang Yang, Canyu Chen, Zhewen Tan, Yizhi Wang, Zichao Feng, Chunyang Liu, Kehua Sheng, Juan Zhang, Linlin Yang, Baochang Zhang, Yan Wang, Bo Zhang, Xianbin Cao
机构
*
School of Electronic Information Engineering, Beihang University(北京航空航天大学电子信息工程学院)
;
National College for Excellent Engineers, Beihang University(北京航空航天大学卓越工程师学院)
;
Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院)
;
DiDi(滴滴出行)
;
State Key Laboratory of Media Convergence and Communication, Communication University of China(中国传媒大学媒体融合与传播国家重点实验室)
;
School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院)
;
School of Cyber Science and Technology, Beihang University(北京航空航天大学网络安全科学与技术学院)
;
School of Artificial Intelligence, Beihang University(北京航空航天大学人工智能学院)
CommentsWe are pleased to announce that this paper has been accepted by the 19th European Conference on Computer Vision (ECCV 2026). We appreciate the valuable feedback from the reviewers and look forward to sharing our findings with the community
DeWorldSG: Depth-Aware 3D Semantic Scene Graph Generation via World-Model Priors
DeWorldSG: 基于世界模型先验的深度感知3D语义场景图生成
Seok-Young Kim, Abdelrahman Elskhawy, Taewook Ha, Dooyoung Kim, Eunjae Shin, Benjamin Busam, Woontack Woo
机构
*
Korea Advanced Institute of Science and Technology(韩国科学技术院)
;
Technical University of Munich(慕尼黑工业大学)
;
Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心)
;
La Trobe University(拉筹伯大学)
Changsheng Lu, Yuxin Chen, Haokun Gui, Rong Wang, Jie Yang, Harry Yang, Anton van den Hengel, Jiaya Jia
机构
*
Hong Kong University of Science and Technology(香港科技大学)
;
Australian National University(澳大利亚国立大学)
;
Tencent Inc.(腾讯公司)
;
Adelaide University(阿德莱德大学)
Towards Memory-Efficient Autoregressive Video Generation via Instance-Specific Parametric Absorption
面向内存高效的自回归视频生成:基于实例特定参数吸收
Xiaomeng Fu, Jia Li, Yiming Hu, Yong Wang, Hayden Kwok-Hay So, Jiao Dai, Xiangxiang Chu, Jizhong Han
机构
*
Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
;
The University of Hong Kong(香港大学)
;
AMAP, Alibaba Group(阿里巴巴集团高德地图)
SPECSIA: Stylization Dataset for Novel-View Enhancement in Drawing-based 3D Animation
SPECSIA:用于基于绘图的3D动画中新颖视角增强的风格化数据集
Kyuwon Kim, Sunjae Yoon, Chang D. Yoo
机构
*
School of Electrical Engineering, KAIST, Daejeon, Republic of Korea(韩国成均馆大学电子工程学院)
;
Department of AI, Chung-Ang University, Seoul, Republic of Korea(韩国 Chung-Ang 大学人工智能系)