arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉与机器人

VLA / 视觉-语言-动作模型

视觉-语言-动作模型、机器人基础模型和语言条件机器人控制。

2025-09-29 至 2025-09-29 共收录 6 信号源:cs.RO, cs.CV, cs.AI, cs.LG

1. VLA模型 6 篇

2509.21986 2025-09-29 cs.RO cs.AI 90%

Developing Vision-Language-Action Model from Egocentric Videos

Tomoya Yoshida, Shuhei Kurita, Taichi Nishimura, Shinsuke Mori

机构 * Kyoto University(京都大学) National Institute of Informatics(国家信息研究所) Institute of Science Tokyo(东京科学研究所) NII LLMC(日本信息机构语言模型中心) Sony Interactive Entertainment(索尼互动娱乐)

专题命中 VLA模型 :vision-language-action(title,abstract);action model(title,abstract);VLA(abstract);分类 cs.RO、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22093 2025-09-29 cs.RO cs.AI 86%

Action-aware Dynamic Pruning for Efficient Vision-Language-Action Manipulation

Xiaohuan Pei, Yuxing Chen, Siyu Xu, Yunke Wang, Yuheng Shi, Chang Xu

机构 * School of Computer Science, The University of Sydney(悉尼大学计算机科学学院)

专题命中 VLA模型 :vision-language-action(title,abstract);VLA(abstract);action model(abstract);分类 cs.RO、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05667 2025-09-29 cs.CV cs.AI 84%

DriveAction: A Benchmark for Exploring Human-like Driving Decisions in VLA Models

Yuhan Hao, Zhengning Li, Lei Sun, Weilong Wang, Naixin Yi, Sheng Song, Caihong Qin, Mofan Zhou, Yifei Zhan, Xianpeng Lang

机构 * Li Auto Inc.

专题命中 VLA模型 :VLA(title,abstract);vision-language-action(abstract);分类 cs.CV、cs.AI

Comments Benchmark: https://huggingface.co/datasets/LiAuto-DriveAction/drive-action

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22441 2025-09-29 cs.RO 84%

UnderwaterVLA: Dual-brain Vision-Language-Action architecture for Autonomous Underwater Navigation

Zhangyuan Wang, Yunpeng Zhu, Yuqi Yan, Xiaoyuan Tian, Xinhao Shao, Meixuan Li, Weikun Li, Guangsheng Su, Weicheng Cui, Dixia Fan

机构 * School of Engineering, Westlake University(西lake大学工程学院) College of Information Science & Electronic Engineering, Zhejiang University(浙江大学信息科学与电子工程学院) College of Environmental and Resource Sciences, Zhejiang University(浙江大学环境与资源科学学院) Australian National University(澳大利亚国立大学)

专题命中 VLA模型 :vision-language-action(title,abstract);VLA(abstract,comments);分类 cs.RO

Comments This paper introduces the first VLA framework for AUVs, featuring a dual-brain architecture and zero-data MPC for real-world underwater navigation

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15098 2025-09-29 cs.HC cs.SY eess.SY 71%

Optimal Behavior Planning for Implicit Communication using a Probabilistic Vehicle-Pedestrian Interaction Model

Markus Amann, Malte Probst, Raphael Wenzel, Thomas H. Weisswange, Miguel Ángel Sotelo

专题命中 VLA模型 :action model(title)

Comments 8 pages, 5 figures, conference article

Journal ref Proceedings of 36th IEEE Intelligent Vehicles Symposium 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21607 2025-09-29 cs.LG 57%

Causal Abstraction Inference under Lossy Representations

Kevin Xia, Elias Bareinboim

机构 * CausalAI Lab, Columbia University(因果推理实验室,哥伦比亚大学)

专题命中 VLA模型 :action model(abstract);分类 cs.LG

Comments 35 pages, 8 figures, published at ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏