arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

VLA / 视觉-语言-动作模型

视觉-语言-动作模型、机器人基础模型和语言条件机器人控制。

2026-01-27 至 2026-01-27 共收录 6 信号源:cs.RO, cs.CV, cs.AI, cs.LG

1. VLA模型 6 篇

2509.11480 2026-01-27 cs.AI cs.CV cs.ET cs.LG cs.RO 89%

Cross-Platform Scaling of Vision-Language-Action Models from Edge to Cloud GPUs

从边缘到云GPU的视觉-语言-动作模型跨平台扩展

Amir Taherin, Juyi Lin, Arash Akbari, Arman Akbari, Pu Zhao, Weiwei Chen, David Kaeli, Yanzhi Wang

机构 * Department of Electrical and Computer Engineering(电气与计算机工程系)

专题命中 VLA模型 :vision-language-action(title,abstract);action model(title);VLA(abstract);分类 cs.RO、cs.CV、cs.AI

AI总结 本文研究了视觉-语言-动作模型在边缘到云GPU平台上的跨平台扩展,揭示了架构选择和电力限制对性能的影响,并挑战了数据中心硬件的优越性假设。

Comments To appear in the Asilomar Conference on Signals, Systems, and Computers 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17885 2026-01-27 cs.CV cs.AI cs.RO 87%

PEAfowl: Perception-Enhanced Multi-View Vision-Language-Action for Bimanual Manipulation

PEAfowl:感知增强的多视角视觉-语言-动作用于双臂操作

Qingyu Fan, Zhaoxiang Li, Yi Lu, Wang Chen, Qiu Shen, Xiao-xiao Long, Yinghao Cai, Tao Lu, Shuo Wang, Xun Cao

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Nanjing University(南京大学)

专题命中 VLA模型 :vision-language-action(title,abstract);VLA(abstract);action model(abstract);分类 cs.RO、cs.CV、cs.AI

AI总结 PEAfowl通过增强感知的多视角视觉-语言-动作策略,提升双臂操作在复杂环境中的稳定性和成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11908 2026-01-27 cs.RO 77%

Safe Learning for Contact-Rich Robot Tasks: A Survey from Classical Learning-Based Methods to Safe Foundation Models

接触丰富机器人任务的安全学习:从经典学习方法到安全基础模型的综述

Heng Zhang, Rui Dai, Gokhan Solak, Pokuang Zhou, Yu She, Arash Ajoudani

机构 * Human-Robot Interfaces and Interaction Lab, Istituto Italiano di Tecnologia, Genova, Italy(人类-机器人接口与交互实验室,意大利技术研究院,热那亚,意大利) Ph.D. program of national interest in Robotics and Intelligent Machines (DRIM) and Università di Genova, Genoa, Italy(机器人与智能机器国家利益博士项目(DRIM)和热那亚大学,热那亚,意大利) Edwardson School of Industrial Engineering, Purdue University, West Lafayette, IN, USA(工业工程埃德华森学校,普渡大学,西拉法伊斯,美国)

专题命中 VLA模型 :vision-language-action(abstract);VLA(abstract);action model(abstract);分类 cs.RO

AI总结 本文综述了接触丰富机器人任务的安全学习方法,探讨了从经典学习方法到安全基础模型的发展,分析了安全探索与执行的关键技术及未来方向。

Comments version 2

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18323 2026-01-27 cs.RO 70%

TC-IDM: Grounding Video Generation for Executable Zero-shot Robot Motion

TC-IDM:为可执行零样本机器人运动实现视频生成

Weishi Mi, Yong Bao, Xiaowei Chi, Xiaozhu Ju, Zhiyuan Qin, Kuangzhi Ge, Kai Tang, Peidong Jia, Shanghang Zhang, Jian Tang

机构 * Beijing Innovation Center of Humanoid Robotics(北京人形机器人创新中心) State Key Laboratory of Multimedia Information Processing(多媒体信息处理国家重点实验室) School of Computer Science, Peking University(北京大学计算机科学学院)

专题命中 VLA模型 :vision-language-action(abstract);VLA(abstract);分类 cs.RO

AI总结 TC-IDM通过工具中心逆动力学模型实现视频生成,提升机器人零样本任务的执行能力与泛化性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.01482 2026-01-27 astro-ph.GA 50%

NGC3521 as the Milky Way analogue: spectral energy distributon from UV to Radio and photometric variability

NGC3521作为银河系的类比:从紫外到射电的光谱能量分布及光度变化

O. S. Pastoven, O. V. Kompaniiets, I. B. Vavilova, I. O. Izviekova

专题命中 VLA模型 :VLA(abstract)

AI总结 研究NGC 3521的光谱能量分布及光度变化,确认其核活动为LINER,发现弱光度变化并确定其恒星质量、尘埃质量和恒星形成率。

Comments 17 pages, 9 figures

Journal ref Space Sci. & Technol. 2024; 30(6):67-83

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06615 2026-01-27 math.AP 50%

A Revisiting of the Pressure Elimination for a Fluid-Structure PDE Interaction and Its Implications

对流体-结构PDE相互作用中压力消除的重新审视及其影响

George Avalos, Yuhao Mu

专题命中 VLA模型 :action model(abstract)

AI总结 本文提出了一种新的压力消除技术,用于流体-结构相互作用模型,通过显式半群生成元证明了连续PDE的well-posedness,并展示了有限元方法在一般几何中的应用。

Comments 28 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏