arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

VLA / 视觉-语言-动作模型

视觉-语言-动作模型、机器人基础模型和语言条件机器人控制。

2025-08-12 至 2025-08-12 共收录 8 信号源:cs.RO, cs.CV, cs.AI, cs.LG

1. VLA模型 6 篇

2508.06547 2025-08-12 cs.RO 88%

A tutorial note on collecting simulated data for vision-language-action models

Heran Wu, Zirun Zhou, Jingfeng Zhang

机构 * School of Computer Science, The University of Auckland(计算机科学学院,奥克兰大学)

专题命中 VLA模型 :vision-language-action(title,abstract);action model(title);VLA(abstract);分类 cs.RO

Comments This is a tutorial note for educational purposes

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07728 2025-08-12 math.OC math.AP 78%

Optimization of a Nonlinear Acoustics -- Structure Interaction Model

Barbara Kaltenbacher, Amjad Tuffaha

专题命中 VLA模型 :action model(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07318 2025-08-12 cs.CV 57%

RORPCap: Retrieval-based Objects and Relations Prompt for Image Captioning

Jinjing Gu, Tianbao Qin, Yuanyuan Pu, Zhengpeng Zhao

机构 * School of Information Science and Engineering(信息科学与工程学院)

专题命中 VLA模型 :action model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08333 2025-08-12 cs.CV 57%

DAViD: Modeling Dynamic Affordance of 3D Objects Using Pre-trained Video Diffusion Models

Hyeonwoo Kim, Sangwon Baik, Hanbyul Joo

机构 * Seoul National University(首尔国立大学)

专题命中 VLA模型 :action model(abstract);分类 cs.CV

Comments Project Page: https://snuvclab.github.io/david/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08056 2025-08-12 astro-ph.IM hep-ex 50%

AugerPrime: Status and first results

David Schmidt

专题命中 VLA模型 :action model(abstract)

Comments Presented at the 39th International Cosmic Ray Conference (ICRC 2025). 8 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07105 2025-08-12 astro-ph.HE hep-ph 50%

EPOS.LHC-R : a global approach to solve the muon puzzle

Tanguy Pierog, Klaus Werner

专题命中 VLA模型 :action model(abstract)

Comments 8 pages, 3 figures, Proceeding of the 39th ICRC conference in Geneva (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 动作表示与策略 1 篇

2502.05855 2025-08-12 cs.RO cs.CV 74%

DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control

Junjie Wen, Yichen Zhu, Jinming Li, Zhibin Tang, Chaomin Shen, Feifei Feng

专题命中 动作表示与策略 :VLA(abstract,comments);vision-language-action(abstract);分类 cs.RO、cs.CV

Comments The webpage is at https://dex-vla.github.io/. DexVLA is accepted by CoRL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 数据集与评测 1 篇

2508.06553 2025-08-12 cs.CV 70%

Static and Plugged: Make Embodied Evaluation Simple

Jiahao Xiao, Jianbo Zhang, BoWen Yan, Shengyu Guo, Tongrui Ye, Kaiwei Zhang, Zicheng Zhang, Xiaohong Liu, Zhengxue Cheng, Lei Fan, Chuyi Li, Guangtao Zhai

专题命中 数据集与评测 :vision-language-action(abstract);action model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏