arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

VLA / 视觉-语言-动作模型

视觉-语言-动作模型、机器人基础模型和语言条件机器人控制。

2026-02-02 至 2026-02-02 共收录 6 信号源:cs.RO, cs.CV, cs.AI, cs.LG

1. VLA模型 6 篇

2601.20321 2026-02-02 cs.RO 91%

TaF-VLA: Tactile-Force Alignment in Vision-Language-Action Models for Force-aware Manipulation

TaF-VLA:面向力感知操控的视觉-语言-动作模型中的触觉-力对齐

Yuzhe Huang, Pei Lin, Wanlin Li, Daohan Li, Jiajun Li, Jiaming Jiang, Chenxi Xiao, Ziyuan Jiao

机构 * Beihang University(北航大学) ShanghaiTech University(上海科技大学) Beijing Institute for General Artificial Intelligence(北京一般人工智能研究院) The University of Hong Kong(香港大学)

专题命中 VLA模型 :vision-language-action(title,abstract);VLA(title,abstract);action model(title);分类 cs.RO

AI总结 TaF-VLA通过触觉-力对齐提升视觉-语言-动作模型在力感知操控中的性能。

Comments 17pages,9fig

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04769 2026-02-02 cs.CV 90%

Vision-Language-Action (VLA) Models: Concepts, Progress, Applications and Challenges

视觉-语言-动作(VLA)模型:概念、进展、应用与挑战

Ranjan Sapkota, Yang Cao, Konstantinos I. Roumeliotis, Manoj Karkee

机构 * Cornell University(康奈尔大学) The Hong Kong University of Science and Technology(香港科学与技术大学) University of the Peloponnese(希腊皮洛斯大学)

专题命中 VLA模型 :vision-language-action(title,abstract);VLA(title,abstract);vision language action(abstract);action model(abstract)

AI总结 本文综述了视觉-语言-动作(VLA)模型的概念、进展、应用与挑战,探讨了其在自动驾驶、医疗机器人等领域的应用及未来发展方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19236 2026-02-02 cs.RO cs.CV 87%

MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation

MemoryVLA: 视觉-语言-动作模型中用于机器人操作的感知-认知记忆

Hao Shi, Bin Xie, Yingfei Liu, Lin Sun, Fengrong Liu, Tiancai Wang, Erjin Zhou, Haoqiang Fan, Xiangyu Zhang, Gao Huang

机构 * Department of Automation, BNRist, Tsinghua University(自动化系、BNRist、清华大学) Dexmal MEGVII Technology(MEGVII技术) Tianjin University(天津大学) Harbin Institute of Technology(哈尔滨工业大学) StepFun

专题命中 VLA模型 :vision-language-action(title);action model(title);VLA(abstract);分类 cs.RO、cs.CV

AI总结 MemoryVLA通过结合感知与认知记忆机制,提升机器人操作中长时间跨度任务的性能,实现对时间依赖任务的高效处理。

Comments ICLR 2026 | The project is available at https://shihao1895.github.io/MemoryVLA

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22714 2026-02-02 cs.LG cs.AI cs.CV 75%

Vision-Language Models Unlock Task-Centric Latent Actions

视觉-语言模型解锁任务导向的潜在动作

Alexander Nikulin, Ilya Zisman, Albina Klepach, Denis Tarasov, Alexander Derevyagin, Andrei Polubarov, Lyubaykin Nikita, Vladislav Kurenkov

机构 * Innopolis University(因诺波利斯大学) Research Center for Trusted Artificial Intelligence, ISP RAS(可信人工智能研究所以及ISP俄罗斯科学院)

专题命中 VLA模型 :vision-language-action(abstract);action model(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 本文提出利用视觉-语言模型的常识推理能力,通过可提示的表示分离可控变化与噪声,提升潜在动作质量,从而提高下游任务性能。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14133 2026-02-02 cs.RO cs.CV 73%

TwinBrainVLA: Unleashing the Potential of Generalist VLMs for Embodied Tasks via Asymmetric Mixture-of-Transformers

TwinBrainVLA: 通过非对称Transformer混合释放通用视觉语言模型在具身任务中的潜力

Bin Yu, Shijie Lian, Xiaopeng Lin, Yuliang Wei, Zhaolong Shen, Changti Wu, Yuzhuo Miao, Xinming Wang, Bailing Wang, Cong Huang, Kai Chen

机构 * HIT(哈尔滨工业大学) ZGCA(中钢集团人工智能研究院) ZGCI(中钢集团智能计算研究院) HUST(华中科技大学) HKUST(GZ)(香港科技大学(广州)) BUAA(北京航空航天大学) ECNU(华东师范大学) CASIA(中国科学院自动化研究所) DeepCybo

专题命中 VLA模型 :vision-language-action(abstract);VLA(abstract);分类 cs.RO、cs.CV

AI总结 TwinBrainVLA通过非对称Transformer混合机制,利用冻结的通用视觉语言模型和可训练的专家路径,实现机器人任务中的具身智能提升。

Comments GitHub: https://github.com/ZGC-EmbodyAI/TwinBrainVLA

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18466 2026-02-02 gr-qc 50%

Study of dynamical systems and large-scale structure

动力系统与大尺度结构的研究

Dumiso Mithi, Saikat Charkraborty, Shambel Sahlu, Amare Abebe

专题命中 VLA模型 :action model(abstract)

AI总结 本文通过动力系统方法研究大尺度结构,探讨了暗能量模型与暗物质相互作用的理论可行性。

Comments The 68th Annual Conference of the South African Institute of Physics (SAIP) proceedings 2024

详情

展开后加载摘要…

URL PDF HTML 收藏