arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

VLA / 视觉-语言-动作模型

视觉-语言-动作模型、机器人基础模型和语言条件机器人控制。

2025-12-09 至 2025-12-09 共收录 13 信号源:cs.RO, cs.CV, cs.AI, cs.LG
2512.04952 2025-12-09 cs.CV cs.RO 88%

FASTer: Toward Efficient Autoregressive Vision Language Action Modeling via Neural Action Tokenization

FASTer:通过神经动作分词实现高效的自回归视觉语言动作建模

Yicheng Liu, Shiduo Zhang, Zibin Dong, Baijun Ye, Tianyuan Yuan, Xiaopeng Yu, Linqi Yin, Chenhao Lu, Junhao Shi, Luca Jiang-Tao Yu, Liangtao Zheng, Tao Jiang, Jingjing Gong, Xipeng Qiu, Hang Zhao

机构 * Tsinghua University(清华大学) Fudan University(复旦大学) Shanghai Innovation Institute(上海创新研究院) Galaxea AI Tianjin University(天津大学) Hong Kong University(香港大学) UCSD(加州大学圣地亚哥分校)

专题命中 VLA模型 :vision language action(title);action model(title);vision-language-action(abstract);VLA(abstract)

AI总结 FASTer通过神经动作分词实现高效自回归视觉语言动作建模,提升机器人学习的推理效率和任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07582 2025-12-09 cs.RO 88%

See Once, Then Act: Vision-Language-Action Model with Task Learning from One-Shot Video Demonstrations

一次观看,随后行动:基于单次视频示范的任务学习视觉-语言-行动模型

Guangyan Chen, Meiling Wang, Qi Shao, Zichen Zhou, Weixin Mao, Te Cui, Minzhao Zhu, Yinan Deng, Luojie Yang, Zhanqi Zhang, Yi Yang, Hua Chen, Yufeng Yue

机构 * Beijing Institute of Technology(北京理工大学) LimX Dynamics

专题命中 VLA模型 :vision-language-action(title,abstract);action model(title);VLA(abstract);分类 cs.RO

AI总结 ViVLA通过单次视频示范高效学习机器人操控任务,实现跨任务和跨身体的显著性能提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03724 2025-12-09 cs.CV cs.RO 84%

PosA-VLA: Enhancing Action Generation via Pose-Conditioned Anchor Attention

PosA-VLA: 通过姿态条件化锚点注意力增强动作生成

Ziwen Li, Xin Wang, Hanlue Zhang, Runnan Chen, Runqi Lin, Xiao He, Han Huang, Yandong Guo, Fakhri Karray, Tongliang Liu, Mingming Gong

机构 * MBZUAI AI2 Robotics The University of Sydney(悉尼大学) The University of Melbourne(墨尔本大学)

专题命中 VLA模型 :VLA(title,abstract);vision-language-action(abstract);分类 cs.RO、cs.CV

AI总结 PosA-VLA通过姿态条件化锚点注意力机制提升动作生成的精度和效率,适用于多样化的机器人任务和复杂环境。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07595 2025-12-09 cs.FL cs.SC 78%

Specializing anti-unification for interaction models composition via gate connections

通过门连接专门化交互模型组合的反统一

Joel Nguetoum, Boutheina Bannour, Pascale Le Gall, Erwan Mahe

专题命中 VLA模型 :action model(title,abstract)

AI总结 本文提出通过门连接专门化交互模型组合的反统一方法,旨在通过保持常数和一般化结构来实现全局模型的统一。

Comments 26 pages (21 in the article, 5 pages in appendices), 7 figures, 6 tables, submitted to Formal Methods 2026 (FM 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06963 2025-12-09 cs.RO cs.AI cs.CV 75%

VideoVLA: Video Generators Can Be Generalizable Robot Manipulators

VideoVLA: 视频生成器可以成为通用的机器人操作器

Yichao Shen, Fangyun Wei, Zhiying Du, Yaobo Liang, Yan Lu, Jiaolong Yang, Nanning Zheng, Baining Guo

机构 * IAIR, Xi’an Jiaotong University(人工智能研究院、西安交通大学) Microsoft Research Asia(微软亚洲研究院) Fudan University(复旦大学)

专题命中 VLA模型 :vision-language-action(abstract);VLA(abstract);分类 cs.RO、cs.CV、cs.AI

AI总结 VideoVLA通过将视频生成模型转化为机器人VLA操作器,实现了动作与视觉后果的双预测,提升机器人操作的泛化能力。

Comments Project page: https://videovla-nips2025.github.io

Journal ref The Thirty-ninth Annual Conference on Neural Information Processing Systems(NeurIPS2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00975 2025-12-09 cs.CV cs.LG cs.RO 75%

MM-ACT: Learn from Multimodal Parallel Generation to Act

MM-ACT: 从多模态并行生成中学习以行动

Haotian Liang, Xinyi Chen, Bin Wang, Mingkang Chen, Yitian Liu, Yuhao Zhang, Zanxin Chen, Tianshuo Yang, Yilun Chen, Jiangmiao Pang, Dong Liu, Xiaokang Yang, Yao Mu, Wenqi Shao, Ping Luo

机构 * Shanghai AI Laboratory(上海人工智能实验室) Shanghai Jiao Tong University(上海交通大学) The University of Hong Kong(香港大学) University of Science and Technology of China(中国科学技术大学) Fudan University(复旦大学) Zhejiang University(浙江大学)

专题命中 VLA模型 :vision-language-action(abstract);VLA(abstract);分类 cs.RO、cs.CV、cs.LG

AI总结 MM-ACT通过多模态并行生成提升机器人任务执行能力,实现96.3%的成功率。

Comments 17 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07472 2025-12-09 cs.RO cs.LG 73%

Affordance Field Intervention: Enabling VLAs to Escape Memory Traps in Robotic Manipulation

可及场干预:使VLAs在机器人操作中摆脱记忆陷阱

Siyu Xu, Zijian Wang, Yunke Wang, Chenghao Xia, Tao Huang, Chang Xu

机构 * School of Computer Science, The University of Sydney(悉尼大学计算机科学学院) John Hopcropt Center for Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学中心)

专题命中 VLA模型 :vision-language-action(abstract);VLA(abstract);分类 cs.RO、cs.LG

AI总结 本研究提出可及场干预(AFI)方法,通过引入3D空间可及场提升VLA在机器人操作中对分布变化的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06294 2025-12-09 q-bio.MN cs.LG math.PR q-bio.QM stat.ML 57%

Interpretable Neural Approximation of Stochastic Reaction Dynamics with Guaranteed Reliability

可解释的神经近似随机反应动力学:具有保证可靠性的方法

Quentin Badolle, Arthur Theuer, Zhou Fang, Ankit Gupta, Mustafa Khammash

机构 * Department of Biosystems Science and Engineering, ETH Zurich(1 生物系统科学与工程系,苏黎世联邦理工学院)

专题命中 VLA模型 :action model(abstract);分类 cs.LG

AI总结 DeepSKA通过结合可解释性、可靠性保证和高效计算,为随机反应动力学提供了一种新的神经近似方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06290 2025-12-09 cs.CV 57%

StrokeNet: Unveiling How to Learn Fine-Grained Interactions in Online Handwritten Stroke Classification

StrokeNet: 解析如何在在线手写体识别中学习细粒度交互

Yiheng Huang, Shuang She, Zewei Wei, Jianmin Lin, Ming Yang, Wenyin Liu

机构 * College of Computer Science and Technology(计算机科学与技术学院) Guangdong University of Technology(广东技术大学) CVTE Research(CVTE研究院)

专题命中 VLA模型 :action model(abstract);分类 cs.CV

AI总结 StrokeNet通过参考点对表示和动态选择参考点,有效捕捉手写体中细粒度交互,提升识别准确率。

Comments 17 pages, 5 figures

Journal ref ICDAR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20064 2025-12-09 astro-ph.HE 50%

A Radio-quiet AGN as a candidate counterpart to neutrino event IceCube-200615A

一个无线电安静的AGN作为IceCube-200615A中微子事件的候选对应物

F. McBride, N. Schettino, J. D. O'Brien, W. Harwood, L. Perot, G. Temple, H. Ayalo Solares, A. Corsi, A. Coleiro, D. Cowen, D. B. Fox, Y. Li, K. Murase, A. Pellegrino, T. D. Russell, S. Wissel

专题命中 VLA模型 :VLA(abstract)

AI总结 通过X射线观测和多波段分析,研究发现1RXS J093117.6+033146可能是IceCube-200615A中微子事件的候选源,支持无线电安静AGN与中微子事件的关联。

Comments Accepted for publication by MNRAS

Journal ref Mon.Not.Roy.Astron.Soc. 541 (2025) 1613

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.05466 2025-12-09 cs.LO cs.GT 50%

Seven kinds of equivalent models for generalized coalition logics

七类等价模型用于广义联盟逻辑

Zixuan Chen, Fengkui Ju

专题命中 VLA模型 :action model(abstract)

AI总结 本文提出七类等价模型用于广义联盟逻辑,通过分析不同假设条件下的模型,揭示了八种联盟逻辑的有效公式集由六种其他模型决定。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06641 2025-12-09 cs.IR cs.CL 50%

An Index-based Approach for Efficient and Effective Web Content Extraction

基于索引的方法用于高效有效的网络内容提取

Yihan Chen, Benfeng Xu, Xiaorui Wang, Zhendong Mao

机构 * University of Science and Technology of China(中国科学技术大学) Metastone Technology(Metastone技术)

专题命中 VLA模型 :action model(abstract)

AI总结 本文提出基于索引的方法,用于高效提取网络相关内容,提升RAG系统查询准确性与处理速度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06066 2025-12-09 astro-ph.HE astro-ph.SR 50%

Revealing the accelerating wind in the inner region of the colliding-wind binary WR 112

揭示碰撞风双星WR 112内区的加速风

John D. Monnier, Yinuo Han, Michael F. Corcoran, Sanne Bloot, Joseph R. Callingham, William Danchi, Philip G. Edwards, Lincoln Greenhill, Kenji Hamaguchi, Matthew J. Hankins, Ryan Lau, Jon M. Miller, Anthony F. J. Moffat, Garreth Ruane, Christopher M. P. Russell, Anthony Soulain, Samaporn Tinyanont, Peter Tuthill, Jason J. Wang, Peredur M. Williams

专题命中 VLA模型 :VLA(abstract)

AI总结 WR 112通过揭示碰撞风双星中加速风的运动学特性,为测试多样化轨道架构下的碰撞风物理提供了基准系统。

Comments 28 pages; 12 figures; 6 tables; Accepted Astronomical Journal (AJ)

Journal ref Astronomical Journal (AJ) 170, 218 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏