Align-KD: Distilling Cross-Modal Alignment Knowledge for Mobile Vision-Language Model Enhancement
Align-KD:为移动视觉语言模型增强提取跨模态对齐知识
Qianhan Feng, Wenshuo Li, Tong Lin, Xinghao Chen
机构
*
State Key Laboratory of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University, China(通用人工智能国家重点实验室,智能科学与技术学院,北京大学,中国)
;
Huawei Noah’s Ark Lab, China(华为诺亚方舟实验室,中国)
Training-Free Composed Video Retrieval via Visual Representation-Guided Video-LLM Reasoning
基于视觉表示引导的视频-大语言模型推理的无训练组合视频检索
Yang Liu, Qianqian Xu, Peisong Wen, Siran Dai, Qingming Huang
机构
*
School of Computer Science and Technology, University of Chinese Academy of Sciences(中国科学院大学计算机科学与技术学院)
;
State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences(中国科学院人工智能安全国家重点实验室)
;
Beijing Academy of Artificial Intelligence(北京人工智能研究院)
;
Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
;
School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院)
机构
*
The University of Hong Kong(香港大学)
;
Shenyang Institute of Automation, Chinese Academy of Sciences(中国科学院沈阳自动化研究所)
;
The Chinese University of Hong Kong(香港中文大学)
;
University of California, Santa Cruz(加州大学圣克鲁兹分校)
Generative Diffusion Priors for 3D Mapping of the Dark Universe
用于暗宇宙三维映射的生成扩散先验
Brandon Zhao, Diana Scognamiglio, Olivier Doré, Katherine L. Bouman
机构
*
Department of Computing and Mathematical Sciences, California Institute of Technology(加州理工学院计算与数学科学系)
;
Jet Propulsion Laboratory, California Institute of Technology(加州理工学院喷气推进实验室)
;
Department of Physics, Duke University(杜克大学物理系)
;
Cahill Center for Astronomy and Astrophysics, California Institute of Technology(加州理工学院卡希尔天文与天体物理中心)
Physical Object Understanding with a Physically Controllable World Model
基于物理可控世界模型的物理对象理解
Rahul Venkatesh, Klemen Kotar, Lilian Naing Chen, Wanhee Lee, Gia Ancone, Seungwoo Kim, Luca Thomas Wheeler, Jared Watrous, Honglin Chen, Daniel Bear, Stefan Stojanov, Daniel LK Yamins
Closed-Loop Neural Activation Control in Vision-Language-Action Models
视觉-语言-动作模型中的闭环神经激活控制
Abhijith Babu, Ramneet Kaur, Nathaniel D. Bastian, Olivera Kotevska, Susmit Jha, Yanzhao Wu, Sumit Kumar Jha, Anirban Roy
机构
*
Florida International University(佛罗里达国际大学)
;
SRI International(美国桑尼沃德国际研究机构)
;
United States Military Academy(美国军事学院)
;
Oak Ridge National Laboratory(橡树岭国家实验室)
;
University of Florida(佛罗里达大学)
CoCoVideo: The High-Quality Commercial-Model-Based Contrastive Benchmark for AI-Generated Video Detection
CoCoVideo: 基于商业模型的高质量对比基准用于AI生成视频检测
Huidong Feng, Wentao Chen, Jie Chen, Xinqi Cai, Ruolong Ma, Yinglin Zheng, Yuxin Lin, Ming Zeng
机构
*
School of Informatics, Xiamen University(厦门大学信息学院)
;
China Academy of Information and Communications Technology(中国信息通信技术研究院)
;
AI Transcend Pte. Ltd.(AI Transcend有限公司)