arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

VLA / 视觉-语言-动作模型

视觉-语言-动作模型、机器人基础模型和语言条件机器人控制。

2026-01-30 至 2026-01-30 共收录 9 信号源:cs.RO, cs.CV, cs.AI, cs.LG
2601.22153 2026-01-30 cs.RO cs.CV 89%

DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation

DynamicVLA: 一种用于动态物体操控的视觉-语言-动作模型

Haozhe Xie, Beichen Wen, Jiarui Zheng, Zhaoxi Chen, Fangzhou Hong, Haiwen Diao, Ziwei Liu

机构 * S-Lab, Nanyang Technological University(南洋理工大学S实验室)

专题命中 VLA模型 :vision-language-action(title,abstract);action model(title);VLA(abstract,comments);分类 cs.RO、cs.CV

AI总结 DynamicVLA通过整合时间推理和闭环适应,提出了一种用于动态物体操控的视觉-语言-动作模型,提升了响应速度和泛化能力。

Comments Project Page: https://www.infinitescript.com/project/dynamic-vla/ GitHub: https://github.com/hzxie/DynamicVLA

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25889 2026-01-30 cs.LG 88%

$π_\texttt{RL}$: Online RL Fine-tuning for Flow-based Vision-Language-Action Models

$π_\ exttt{RL}$: 流基于视觉-语言-动作模型的在线强化学习微调

Kang Chen, Zhihao Liu, Tonghe Zhang, Zhen Guo, Si Xu, Hao Lin, Hongzhi Zang, Xiang Li, Quanlu Zhang, Zhaofei Yu, Guoliang Fan, Tiejun Huang, Yu Wang, Chao Yu

机构 * Tsinghua University(清华大学) Peking University(北京大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Carnegie Mellon University(卡内基梅隆大学) Infinigence AI Zhongguancun Academy(中关村学院)

专题命中 VLA模型 :vision-language-action(title,abstract);action model(title);VLA(abstract);分类 cs.LG

AI总结 本文提出$π_\ exttt{RL}$方法,通过流噪声和流SDE技术,解决大规模流基于VLA模型中强化学习微调的挑战,提升模型在分布内和分布外任务中的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20966 2026-01-30 cs.RO cs.AI 84%

Parallels Between VLA Model Post-Training and Human Motor Learning: Progress, Challenges, and Trends

VLA模型后训练与人类运动学习的类比:进展、挑战与趋势

Tian-Yu Xiang, Ao-Qun Jin, Xiao-Hu Zhou, Mei-Jiang Gui, Xiao-Liang Xie, Shi-Qi Liu, Shuang-Yi Wang, Sheng-Bin Duan, Fu-Chao Xie, Wen-Kai Wang, Si-Cheng Wang, Ling-Yun Li, Tian Tu, Zeng-Guang Hou

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) The Grainger College of Engineering, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校格拉inger工程学院) CAS Center for Excellence in Brain Science and Intelligence Technology(中国科学院脑科学与智能技术卓越创新中心) Joint Laboratory of Intelligence Science and Technology, Institute of Systems Engineering, Macau University of Science and Technology(澳门科技大学系统工程学院智能科学与技术联合实验室)

专题命中 VLA模型 :VLA(title,abstract);vision-language-action(abstract);分类 cs.RO、cs.AI

AI总结 本文从人类运动学习角度综述了VLA模型后训练的进展、挑战与趋势,提出四类后训练方法并探讨了其在机器人操作中的应用与未来方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12142 2026-01-30 eess.AS cs.MM cs.RO 83%

Listen, Look, Drive: Coupling Audio Instructions for User-aware VLA-based Autonomous Driving

Listen, Look, Drive: 通过用户意识的VLA基于自主驾驶的音频指令耦合

Ziang Guo, Feng Yang, Xuefeng Zhang, Jiaqi Guo, Kun Zhao, Yixiao Zhou, Peng Lu, Sifa Zheng, Zufeng Zhang

机构 * SuZhou Automotive Research Institute of Tsinghua University(清华大学苏州汽车研究院) Department of Electrical and Electronic Engineering, The University of Hong Kong(香港大学电子与电气工程系) Hyundai Motor Advanced Tech. R&D Center School of Vehicle and Mobility, Tsinghua University(清华大学车辆与移动系统学院)

专题命中 VLA模型 :VLA(title,abstract);vision language action(abstract);分类 cs.RO

AI总结 EchoVLA通过结合音频指令与视觉信息,提升自动驾驶对用户意图和情绪的感知能力,显著降低误差和碰撞率。

Comments Accepted by IV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.00564 2026-01-30 cs.LG cs.AI 81%

Scaling Offline Model-Based RL via Jointly-Optimized World-Action Model Pretraining

通过联合优化的世界-动作模型扩展离线模型基于的强化学习

Jie Cheng, Ruixi Qiao, Yingwei Ma, Binhua Li, Gang Xiong, Qinghai Miao, Yongbin Li, Yisheng Lv

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Alibaba Group(阿里巴巴集团)

专题命中 VLA模型 :action model(title,abstract);分类 cs.AI、cs.LG

AI总结 JOWA通过联合优化的世界-动作模型扩展离线RL,实现高效泛化和高性能

Comments Accepted by ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21506 2026-01-30 cs.RO cs.SY eess.SY 70%

IROS: A Dual-Process Architecture for Real-Time VLM-Based Indoor Navigation

IROS: 一种用于实时基于视觉语言模型的室内导航的双过程架构

Joonhee Lee, Hyunseung Shin, Jeonggil Ko

机构 * Yonsei University(延世大学)

专题命中 VLA模型 :vision-language-action(abstract);VLA(abstract);分类 cs.RO

AI总结 IROS是一种结合视觉语言模型上下文推理与轻量级感知模块的双过程架构,实现低延迟的室内导航。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19859 2026-01-30 cs.RO 70%

Unifying Perception and Action: A Hybrid-Modality Pipeline with Implicit Visual Chain-of-Thought for Robotic Action Generation

统一感知与行动:一种融合模态的管道,具有隐式视觉推理链的机器人行动生成

Xiangkai Ma, Lekai Xing, Han Zhang, Wenzhong Li, Sanglu Lu

机构 * State Key Laboratory for Novel Software Technology, Nanjing University(新型软件技术国家重点实验室,南京大学)

专题命中 VLA模型 :vision-language-action(abstract);VLA(abstract);分类 cs.RO

AI总结 VITA通过融合视觉和动作的共享潜在空间,实现机器人行动生成的统一感知与行动,提升了多个任务的成功率和性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21480 2026-01-30 q-bio.PE 50%

Long-term evolution of regulatory DNA sequences. Part 2: Theory and future challenges

调控DNA序列的长期演化。第二部分:理论与未来挑战

Elia Mascolo, Réka Borbély, Noa Ottilie Borst, Nicholas H Barton, Justin Crocker, Gašper Tkačik

专题命中 VLA模型 :action model(abstract)

AI总结 本文探讨了基因调控序列长期进化的理论基础与未来挑战,分析了进化概念在调控序列演化中的应用及统一理论的潜力。

Comments Invited review (Part II of a two-part series), submitted to Current Opinion in Genetics & Development. Part I is available at arXiv:2601.19681

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20982 2026-01-30 astro-ph.GA 50%

Properties of Polarized Radio Sources in the Wide Chandra Deep Field South from 2 to 4GHz

2至4GHz范围内宽切片Chandra深场南的极化射电源性质

Samantha Adams, Mark Lacy, Preshanth Jagannathan, Jose Afonso, William Nielsen Brandt, B. M. Gaensler, Evanthia Hatziminaoglou, Anna Kapinska, Josh Marvil, Hugo Messias, Steve Myers, Ray Norris, Kristina Nyland, Wiphu Rujopakarn, Nick Seymour, Mattia Vaccari, Rick White

专题命中 VLA模型 :VLA(abstract)

AI总结 本研究分析了2-4GHz范围内宽切片Chandra深场南的极化射电源性质,探讨了极化分数、法拉第旋转及光谱指数的关系,并验证了VLASky Survey的偏振数据。

Comments 18 pages, six figures. Catalogs and Stokes I image available at https://doi.org/10.5281/zenodo.18202043

Journal ref Universe, 2026, 12(2), 38

详情

展开后加载摘要…

URL PDF HTML 收藏