Token-Wise Latent Streaming from Slow Reasoners to Fast Planners for Dynamic Vision Language Navigation
用于动态视觉语言导航的从慢速推理器到快速规划器的逐令牌潜在流
Tianshuai Hu, Yangyi Zhong, Zeying Gong, Lingdong Kong, Xiaodong Mei, Guoyang Zhao, Xiaolu Liu, Song Wang, Rong Li, Junwei Liang
机构
*
The Hong Kong University of Science and Technology(香港科技大学)
;
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
National University of Singapore(新加坡国立大学)
;
Zhejiang University(浙江大学)
机构
*
Intelligent Space Robotics Laboratory, Center for Digital Engineering, Skolkovo Institute of Science and Technology(智能空间机器人实验室,数字工程中心,斯克尔科沃科学与技术研究所)
DM-KG: A Novel Method for Boosting Spatial Cognition of Vision-Language Models in Street View Imagery
DM-KG:一种提升街景图像中视觉语言模型空间认知的新方法
Xinyue Xu, Zheng Zhang, Kunyang Ma, Ge Zhu, Lianshuai Cao, Lei Wang, Zixuan Li, Yi Cheng
机构
*
Institute of Surveying and Mapping, Information Engineering University(信息工程大学测绘学院)
;
Institute of Geographic Sciences and Natural Resources Research, Chinese Academy of Sciences(中国科学院地理科学与资源研究所)
SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving
SpaceDrive: 在基于视觉语言模型的自动驾驶中引入空间感知
Peizheng Li, Zhenghao Zhang, David Holtz, Hang Yu, Yutong Yang, Yuzhi Lai, Rui Song, Andreas Geiger, Andreas Zell
机构
*
Mercedes-Benz AG(梅赛德斯-奔驰集团)
;
University of Tübingen(图宾根大学)
;
Tübingen AI Center(图宾根人工智能中心)
;
TU Munich(慕尼黑工业大学)
;
Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院)
;
University of Stuttgart(斯图加特大学)
;
UCLA(加州大学洛杉矶分校)
What Matters in RL-Based Methods for Object-Goal Navigation? An Empirical Study and A Unified Framework
基于强化学习的目标导航方法中什么重要?实证研究与统一框架
Hongze Wang, Boyang Sun, Jiaxu Xing, Fan Yang, Marco Hutter, Dhruv Shah, Davide Scaramuzza, Marc Pollefeys
机构
*
Computer Vision and Geometry Group, ETH Zurich, Switzerland(苏黎世联邦理工学院计算机视觉与几何组)
;
Robotics and Perception Group, University of Zurich, Switzerland(苏黎世大学机器人与感知组)
;
Robotic Systems Lab, ETH Zurich, Switzerland(苏黎世联邦理工学院机器人系统实验室)
;
Princeton Robotic Intelligence and SysteMs group, Princeton University, USA(普林斯顿大学机器人智能与系统组)
;
Microsoft, Switzerland(微软公司)
CoMind: Understanding Collaborative Human Activity from Multiple Minds and Views
CoMind:从多视角理解人类协作活动
Alexey Gavryushin, Dingxi Zhang, Zhao Huang, Alexandros Delitzas, Jiaqi Chen, Ben Ellis, Cedric Zöllner, Manthan Patel, Manuel Kaufmann, Marc Pollefeys, Xi Wang
机构
*
ETH Zurich(苏黎世联邦理工学院)
;
MPI for Informatics(马克斯·普朗克信息研究所)
;
Microsoft Switzerland(微软瑞士公司)
;
TU Munich(慕尼黑工业大学)
机构
*
MoE Key Lab of Artificial Intelligence, Institute of AI, Shanghai Jiao Tong University(教育部人工智能重点实验室,上海交通大学人工智能研究院)
;
Central Research Institute, Huawei(华为中央研究院)
Latent Noise Mask for Reducing Visual Redundancy in Multimodal Large Language Models
潜在噪声掩码:减少多模态大语言模型中的视觉冗余
Kai Jiang, Ruishu Zhu, Siqi Huang, Hongyuan Zhang, Xuelong Li
机构
*
School of Artificial Intelligence, OPtics and ElectroNics (iOPEN), Northwestern Polytechnical University(人工智能学院、光学和电子学(iOPEN)、西北工业大学)
;
Institute of Artificial Intelligence, China Telecom (TeleAI)(人工智能研究院、中国电信(TeleAI))
;
Fudan University(复旦大学)
;
The University of Hong Kong(香港大学)
机构
*
The University of Tokyo, Japan(东京大学)
;
National Institute of Informatics, Japan(日本信息处理学会)
;
The University of Osaka, Japan(大阪大学)
;
Advanced Telecommunications Research Institute International, Japan(国际先进电信研究机构)
OmniSpace: Efficient Geometry Awareness for Autonomous Vehicles MLLMs
OmniSpace: 自动驾驶多模态大语言模型的高效几何感知
Hao Vo, Phu Loc Nguyen, Khoa Vo, Sieu Tran, Duc Minh Nguyen, Ngo Xuan Cuong, Nghi D. Q. Bui, Anh Nguyen, Duy Minh Ho Nguyen, Ngan Le
机构
*
University of Arkansas(阿肯色大学)
;
Google Research, Google(谷歌研究院)
;
University of Liverpool(利物浦大学)
;
Max Planck Research School for Intelligent Systems(马克斯·普朗克智能系统研究所)
机构
*
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
Beijing Academy of Artificial Intelligence (BAAI)(北京人工智能研究院)
;
Peking University(北京大学)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
SleepWalk: A Three-Tier Benchmark for Stress-Testing Instruction-Guided Vision-Language Navigation
SleepWalk:一种三层压力测试基准,用于指导的视觉-语言导航
Niyati Rawal, Sushant Ravva, Shah Alam Abir, Saksham Jain, Aman Chadha, Vinija Jain, Suranjana Trivedy, Amitava Das
机构
*
Indian AI Research Organization (IAIRO)(印度人工智能研究组织)
;
ĸragya Lab, BITS Pilani Goa(BITS Pilani Goa 的 ĸragya 实验室)
;
University of Dhaka(达卡大学)
;
Delhi Technological University(德里技术大学)
;
Apple(苹果公司)
;
Meta