arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉与机器人

VLA / 视觉-语言-动作模型

视觉-语言-动作模型、机器人基础模型和语言条件机器人控制。

2025-09-30 至 2025-09-30 共收录 18 信号源:cs.RO, cs.CV, cs.AI, cs.LG

1. VLA模型 15 篇

2509.24768 2025-09-30 cs.RO 92%

IA-VLA: Input Augmentation for Vision-Language-Action models in settings with semantically complex tasks

Eric Hannus, Miika Malin, Tran Nguyen Le, Ville Kyrki

机构 * Intelligent Robotics Group at the Department of Electrical Engineering and Automation, School of Electrical Engineering, Aalto University(Aalto大学电气工程学院电气工程与自动化系智能机器人组) Biomimetics and Intelligent Systems Group at the Faculty of Information Technology and Electrical Engineering, University of Oulu(奥卢大学信息科技与电气工程学院仿生学与智能系统组) Section of Mechanical Technology at the Department of Engineering Technology and Didactics, Technical University of Denmark(丹麦技术大学工程技术与教学系机械技术部门)

专题命中 VLA模型 :vision-language-action(title,abstract);VLA(title,abstract);action model(title,abstract);分类 cs.RO

Comments Under review for ICRA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13888 2025-09-30 cs.RO 89%

InSpire: Vision-Language-Action Models with Intrinsic Spatial Reasoning

Ji Zhang, Shihan Wu, Xu Luo, Hao Wu, Lianli Gao, Heng Tao Shen, Jingkuan Song

机构 * Southwest Jiaotong University(西南交通大学) University of Electronic Science and Technology of China(电子科技大学) Tongji University(同济大学)

专题命中 VLA模型 :vision-language-action(title,abstract);action model(title,abstract);VLA(abstract);分类 cs.RO

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23655 2025-09-30 cs.RO cs.AI cs.CV cs.LG 89%

Focusing on What Matters: Object-Agent-centric Tokenization for Vision Language Action models

Rokas Bendikas, Daniel Dijkman, Markus Peschl, Sanjay Haresh, Pietro Mazzaglia

机构 * Centre for Artificial Intelligence, UCL(人工智能中心,伦敦大学学院) Qualcomm AI Research(高通人工智能研究)

专题命中 VLA模型 :vision language action(title);action model(title);vision-language-action(abstract);VLA(abstract)

Comments Presented at 9th Conference on Robot Learning (CoRL 2025), Seoul, Korea

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23121 2025-09-30 cs.AI 86%

Transferring Vision-Language-Action Models to Industry Applications: Architectures, Performance, and Challenges

Shuai Li, Chen Yizhe, Li Dong, Liu Sichao, Lan Dapeng, Liu Yu, Zhibo Pang

机构 * Shenyang Institute of Automation(沈阳自动化研究所) Chinese Academy of Sciences(中国科学院) Shandong Normal University(山东师范大学) Department of Production Engineering(生产工程系) Royal Institute of Technology (KTH)(皇家理工学院(KTH)) University of Chinese Academy of Sciences(中国科学院大学) Department of Intelligent Systems(智能系统系)

专题命中 VLA模型 :vision-language-action(title);action model(title);VLA(abstract);分类 cs.AI

Comments Accepted to IAI 2025 (International Conference on Industrial Artificial Intelligence), Shenyang, China, Aug 21 - 24, 2025. Preprint (before IEEE copyright transfer)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23224 2025-09-30 cs.RO cs.AI cs.CV cs.SY eess.SY 85%

Leave No Observation Behind: Real-time Correction for VLA Action Chunks

Kohei Sendai, Maxime Alvarez, Tatsuya Matsushima, Yutaka Matsuo, Yusuke Iwasawa

机构 * The University of Tokyo(东京大学)

专题命中 VLA模型 :VLA(title,abstract);vision-language-action(abstract);分类 cs.RO、cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22199 2025-09-30 cs.RO cs.AI 84%

MimicDreamer: Aligning Human and Robot Demonstrations for Scalable VLA Training

Haoyun Li, Ivan Zhang, Runqi Ouyang, Xiaofeng Wang, Zheng Zhu, Zhiqin Yang, Zhentao Zhang, Boyuan Wang, Chaojun Ni, Wenkang Qin, Xinze Chen, Yun Ye, Guan Huang, Zhenbo Song, Xingang Wang

机构 * GigaAI CASIA NJUST(南京工业大学) Tsinghua University(清华大学)

专题命中 VLA模型 :VLA(title,abstract);vision language action(abstract);分类 cs.RO、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19752 2025-09-30 cs.RO 83%

Beyond Human Demonstrations: Diffusion-Based Reinforcement Learning to Generate Data for VLA Training

Rushuai Yang, Hangxing Wei, Ran Zhang, Zhiyuan Feng, Xiaoyu Chen, Tong Li, Chuheng Zhang, Li Zhao, Jiang Bian, Xiu Su, Yi Chen

机构 * Hong Kong University of Science and Technology(香港科技大学) Microsoft Research Asia(微软亚洲研究院) Wuhan University(武汉大学) University of Chinese Academy of Sciences(中国科学院大学) Tsinghua University(清华大学) Big Data Institute, Central South University(中南大学大数据研究院)

专题命中 VLA模型 :VLA(title,abstract);vision-language-action(abstract);分类 cs.RO

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.19400 2025-09-30 cs.LG cs.AI cs.RO 82%

Vintix: Action Model via In-Context Reinforcement Learning

Andrey Polubarov, Nikita Lyubaykin, Alexander Derevyagin, Ilya Zisman, Denis Tarasov, Alexander Nikulin, Vladislav Kurenkov

机构 * Innopolis University(因诺波利斯大学) Research Center for Trusted Artificial Intelligence, ISP RAS(可信人工智能研究所以及ISP俄罗斯科学院)

专题命中 VLA模型 :action model(title,abstract);分类 cs.RO、cs.AI、cs.LG

Comments ICML 2025, Poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24524 2025-09-30 cs.RO cs.AI cs.SY eess.SY 73%

PhysiAgent: An Embodied Agent Framework in Physical World

Zhihao Wang, Jianxiong Li, Jinliang Zheng, Wencong Zhang, Dongxiu Liu, Yinan Zheng, Haoyi Niu, Junzhi Yu, Xianyuan Zhan

机构 * AIR, Tsinghua University(清华大学) Peking University(北京大学) University of California, Berkeley(加州大学伯克利分校)

专题命中 VLA模型 :vision-language-action(abstract);VLA(abstract);分类 cs.RO、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24559 2025-09-30 cs.LG 70%

Emergent World Representations in OpenVLA

Marco Molinari, Leonardo Nevali, Saharsha Navani, Omar G. Younis

机构 * London School of Economics(伦敦经济学院) ETH Zurich(苏黎世联邦理工学院) Princeton University(普林斯顿大学) Department of Computer Science(计算机科学系) Mila - Quebec AI Institute(魁北克人工智能研究所)

专题命中 VLA模型 :vision language action(abstract);action model(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24447 2025-09-30 astro-ph.HE 50%

A Simulation Study on the Cosmic Ray Energy Spectra of Elemental Mass Groups using the Tibet Air Shower and Muon Detector Arrays through the Bayesian Unfolding Method

G. Imaizumi, M. Anzorena, K. Fujita, Y. Katayose, S. Kato, T. Kawashima, K. Kawata, A. Mizuno, M. Ohnishi, R. Garcia, T. Sako, F. Sugimoto, M. Takita, Y. Yokoe

专题命中 VLA模型 :action model(abstract)

Journal ref Progress of Theoretical and Experimental Physics, ptaf133, 25 September 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23547 2025-09-30 astro-ph.HE 50%

Particle Acceleration along Magnetic Fields as the Origin of Ear-like Structures in Supernova Remnants

Huan Yu, Jun Fang

专题命中 VLA模型 :action model(abstract)

Comments 6 pages, 4 figures, Accepted for publication in MNRAS

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23295 2025-09-30 astro-ph.HE 50%

On the properties of turbulence in the remnant of Tycho supernova

Oleh Petruk, Taras Kuzyo

专题命中 VLA模型 :VLA(abstract)

Comments Accepted by Astrophysical Journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22950 2025-09-30 q-bio.QM 50%

Twin Peaks: Dual-Head Architecture for Structure-Free Prediction of Protein-Protein Binding Affinity and Mutation Effects

Supantha Dey, Ratul Chowdhury

专题命中 VLA模型 :action model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22789 2025-09-30 astro-ph.GA 50%

Masses, Star-Formation Efficiencies, and Dynamical Evolution of 18,000 HII Regions

Debosmita Pathak, Adam K. Leroy, Ashley. T. Barnes, Todd A. Thompson, Laura A. Lopez, Karin M. Sandstrom, Jiayi Sun, Simon C. O. Glover, Ralf S. Klessen, Eric W. Koch, Kirsten L. Larson, Janice Lee, Sharon Meidt, Patricia Sanchez-Blazquez, Eva Schinnerer, Zein Bazzi, Francesco Belfiore, Médéric Boquien, Ryan Chown, Dario Colombo, Enrico Congiu, Oleg V. Egorov, Cosima Eibensteiner, Sushma Kurapati, Miguel Querejeta, Daniel A. Dale, Timo Kravtsov, Mansi Padave, D. J. Pisano, Erik Rosolowsky, Sumit K. Sarbadhicary, Thomas G. Williams, Remy Indebetouw, Hsi-An Pan, Leonardo Úbeda, Amirnezam Amiri, Frank Bigiel, Guillermo A. Blanc, Kathryn Grasha

专题命中 VLA模型 :VLA(abstract)

Comments Accepted for publication in ApJL; main text: 13 pages, 4 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 动作表示与策略 2 篇

2506.08440 2025-09-30 cs.RO cs.AI 87%

TGRPO :Fine-tuning Vision-Language-Action Model via Trajectory-wise Group Relative Policy Optimization

Zengjue Chen, Runliang Niu, He Kong, Qi Wang, Qianli Xing, Zipei Fan

机构 * School of Artificial Intelligence, Jilin University(人工智能学院,吉林大学)

专题命中 动作表示与策略 :vision-language-action(title);action model(title);VLA(abstract);分类 cs.RO、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24020 2025-09-30 cs.CV 57%

Hazy Pedestrian Trajectory Prediction via Physical Priors and Graph-Mamba

Jian Chen, Zhuoran Zheng, Han Hu, Guijuan Zhang, Dianjie Lu, Liang Li, Chen Lyu

机构 * Shandong Normal University(山东师范大学) Sun Yat-sen University(中山大学) Shandong Jiaotong University(山东交通大学)

专题命中 动作表示与策略 :action model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 数据集与评测 1 篇

2509.25032 2025-09-30 cs.RO cs.AI cs.CV 75%

AIRoA MoMa Dataset: A Large-Scale Hierarchical Dataset for Mobile Manipulation

Ryosuke Takanami, Petr Khrapchenkov, Shu Morikuni, Jumpei Arima, Yuta Takaba, Shunsuke Maeda, Takuya Okubo, Genki Sano, Satoshi Sekioka, Aoi Kadoya, Motonari Kambara, Naoya Nishiura, Haruto Suzuki, Takanori Yoshimoto, Koya Sakamoto, Shinnosuke Ono, Hu Yang, Daichi Yashima, Aoi Horo, Tomohiro Motoda, Kensuke Chiyoma, Hiroshi Ito, Koki Fukuda, Akihito Goto, Kazumi Morinaga, Yuya Ikeda, Riko Kawada, Masaki Yoshikawa, Norio Kosuge, Yuki Noguchi, Kei Ota, Tatsuya Matsushima, Yusuke Iwasawa, Yutaka Matsuo, Tetsuya Ogata

机构 * The University of Tokyo(东京大学) AI Robot Association (AIRoA)(人工智能机器人协会) Toyota Motor Corporation(丰田汽车公司) Telexistence, Inc.(Telexistence公司) National Institute of Advanced Industrial Science and Technology (AIST)(国家先进工业科学与技术研究院) Waseda University(早稻田大学)

专题命中 数据集与评测 :vision-language-action(abstract);action model(abstract);分类 cs.RO、cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏