机构
*
SKLCCSE Lab Beihang University Beijing China(SKLCCSE实验室 北航)
;
Department of Data Science City University of Hong Kong Hong Kong China(数据科学系 香港城市大学)
;
Beijing Advanced Innovation Center Beihang University Beijing China(北京先进创新中心 北航)
;
Beihang University(北航)
;
City University of Hong Kong(香港城市大学)
机构
*
Tsinghua University(清华大学)
;
State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室,字节跳动公司人工智能研究院)
;
Harbin Institute of Technology(哈尔滨工业大学)
;
Peking University(北京大学)
专题命中
GUI与屏幕智能体
:grounding(abstract);multimodal large language model(abstract);MLLM(abstract)
机构
*
College of Intelligent Robotics and Advanced Manufacturing, Fudan University(复旦大学智能机器人与先进制造学院)
;
Shanghai Key Lab of Intelligent Information Processing, College of Computer Science and Artificial Intelligence, Fudan University(复旦大学计算机科学与人工智能学院智能信息处理上海市重点实验室)
;
TeleAI, China Telecom(中国电信天翼人工智能公司)
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations
XR-1:通过学习统一的视觉-运动表示实现多功能的视觉-语言-动作模型
Shichao Fan, Kun Wu, Zhengping Che, Xinhua Wang, Di Wu, Fei Liao, Ning Liu, Yixue Zhang, Zhen Zhao, Zhiyuan Xu, Meng Li, Qingjie Liu, Shanghang Zhang, Min Wan, Jian Tang
机构
*
Beijing Innovation Center of Humanoid Robotics, Beijing, China(北京人形机器人创新中心,北京,中国)
;
School of Mechanical Engineering and Automation, Beihang University, Beijing, China(北京航空航天大学机械工程及自动化学院,北京,中国)
;
State Key Laboratory of Virtual Reality Technology and Systems, SCSE, Beihang University, Beijing, China(虚拟现实技术与系统国家重点实验室,SCSE,北京航空航天大学,北京,中国)
;
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University, Beijing, China(多媒体信息处理国家重点实验室,计算机科学学院,北京大学,北京,中国)
机构
*
School of Cyber Science and Engineering, Huazhong University of Science and Technology(华中科技大学网络空间安全学院)
;
College of Computer Science, Chongqing University(重庆大学计算机科学学院)
;
School of Software and engineering, Huazhong University of Science and Technology(华中科技大学软件工程学院)
;
School of Information and Communication Technology, Griffith University(格里菲斯大学信息与通信技术学院)
专题命中
幻觉与鲁棒性
:vision-language model(title);vision language model(abstract);分类 cs.CV
SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving
SpaceDrive: 在基于视觉语言模型的自动驾驶中引入空间感知
Peizheng Li, Zhenghao Zhang, David Holtz, Hang Yu, Yutong Yang, Yuzhi Lai, Rui Song, Andreas Geiger, Andreas Zell
机构
*
Mercedes-Benz AG(梅赛德斯-奔驰集团)
;
University of Tübingen(图宾根大学)
;
Tübingen AI Center(图宾根人工智能中心)
;
TU Munich(慕尼黑工业大学)
;
Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院)
;
University of Stuttgart(斯图加特大学)
;
UCLA(加州大学洛杉矶分校)
专题命中
VLM训练与架构
:VLM(title,summary_cn);vision language model(abstract);分类 cs.CV
Zero-Effort Image-to-Music Generation: An Interpretable RAG-based VLM Approach
零努力图像到音乐生成:一种可解释的基于RAG的视觉语言模型方法
Zijian Zhao, Dian Jin, Zijing Zhou
机构
*
The Hong Kong University of Science and Technology(香港科学与技术大学)
;
The Hong Kong Polytechnic University(香港理工大学)
;
The University of Hong Kong(香港大学)
;
The Hong Kong University of Science(香港科学大学)
;
The Hong Kong Polytechnic University Hong Kong China(香港理工大学香港中国)
;
The University of Hong Kong Hong Kong China(香港大学香港中国)
专题命中
VLM训练与架构
:VLM(title,abstract);vision language model(abstract);分类 cs.AI
On the Efficiency of LoRA Fine-Tuning for Vision-Language-Action Models in Industrial Robotic Manipulation
工业机器人操作中视觉语言动作模型的LoRA微调效率研究
Finn Ferchau, Daniel Pommer, Cristian Axenie
机构
*
Technische Hochschule Nürnberg Georg Simon Ohm(纽伦堡乔治·西蒙·欧姆应用技术大学)
;
Siemens AG(西门子股份公司)
;
Fraunhofer Institute for Integrated Circuits (IIS)(弗劳恩霍夫集成电路研究所)