arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

VLA / 视觉-语言-动作模型

视觉-语言-动作模型、机器人基础模型和语言条件机器人控制。

2025-12-15 至 2025-12-15 共收录 8 信号源:cs.RO, cs.CV, cs.AI, cs.LG
2512.11769 2025-12-15 cs.RO 88%

BLURR: A Boosted Low-Resource Inference for Vision-Language-Action Models

BLURR: 一种提升低资源推理的视觉-语言-动作模型

Xiaoyu Ma, Zhengqing Yuan, Zheyuan Zhang, Kaiwen Shi, Lichao Sun, Yanfang Ye

机构 * University of Notre Dame(诺丁汉大学) Lehigh University(莱斯大学)

专题命中 VLA模型 :vision-language-action(title,abstract);action model(title);VLA(abstract);分类 cs.RO

AI总结 BLURR通过轻量级推理包装器提升低资源下的视觉-语言-动作模型推理效率,保持任务成功率的同时降低计算开销。

Comments 10 pages, 3 figures. Code and integration scripts will be released at this http URL: https://github.com/JijiKing-Sam/BLURR-A-Boosted-Low-Resource-Inference-for-Vision-Language-Action-Model

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11315 2025-12-15 cs.LG 88%

Benchmarking the Generality of Vision-Language-Action Models

对视觉-语言-动作模型通用性的基准测试

Pranav Guruprasad, Sudipta Chowdhury, Harsh Sikka, Mridul Sharma, Helen Lu, Sean Rivera, Aryan Khurana, Hangliang Ren, Yangyue Wang

机构 * Manifold Research Metarch AI Georgia Tech(佐治亚理工学院) Tufts University(塔夫茨大学) Northeastern University(东北大学) Birla Institute of Technology and Science, Pilani(比拉理工学院,帕利尼) Institute for Research and Innovation in Intelligent Systems (IRIIS)(智能系统研究与创新研究所)

专题命中 VLA模型 :action model(title,abstract);vision-language-action(title);vision language action(abstract);分类 cs.LG

AI总结 本文提出MultiNet v1.0基准,评估视觉-语言-动作模型在六个基础能力领域的跨领域泛化能力,发现现有模型在未见领域和模态转移时表现显著退化。

Comments 23 pages, 7 figures, and 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11584 2025-12-15 cs.LG cs.AI cs.RO 85%

Atomic Action Slicing: Planner-Aligned Options for Generalist VLA Agents

原子动作切片:为通用VLA代理的规划对齐选项

Stefan Tabakov, Asen Popov, Dimitar Dimitrov, S. Ensiye Kiyamousavi, Vladimir Hristov, Boris Kraychev

机构 * Sofia University 'St. Kliment Ohridski(索菲亚大学 '圣克莱门特·欧里基斯基') Technical University of Sofia(索菲亚技术大学) EEMCS, University of Twente(EEMCS,特文特大学) GATE Institute, Sofia University 'St. Kliment Ohridski(GATE研究所,索菲亚大学 '圣克莱门特·欧里基斯基')

专题命中 VLA模型 :VLA(title,abstract);vision-language-action(abstract);分类 cs.RO、cs.AI、cs.LG

AI总结 本文提出原子动作切片方法,通过分解长周期演示为短的原子动作,提升VLA代理的任务成功率,并公开发布相关数据集。

Comments The 41st ACM/SIGAPP Symposium On Applied Computing

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11218 2025-12-15 cs.RO cs.CV 84%

Seeing to Act, Prompting to Specify: A Bayesian Factorization of Vision Language Action Policy

视觉与语言动作政策的贝叶斯分解:看见以行动,提示以指定

Kechun Xu, Zhenjie Zhu, Anzhe Chen, Shuqi Zhao, Qing Huang, Yifei Yang, Haojian Lu, Rong Xiong, Masayoshi Tomizuka, Yue Wang

机构 * Zhejiang University(浙江大学) UC Berkeley(伯克利大学)

专题命中 VLA模型 :vision language action(title);vision-language-action(abstract);VLA(abstract);分类 cs.RO、cs.CV

AI总结 本文提出BayesVLA,通过贝叶斯分解策略,解决VLA模型中因模态不平衡导致的灾难性遗忘问题,提升泛化能力和指令遵循性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11620 2025-12-15 cs.RO cs.SY eess.SY 79%

Architecting Large Action Models for Human-in-the-Loop Intelligent Robots

为具有人类在环的智能机器人构建大型动作模型

Kanisorn Sangchai, Methasit Boonpun, Withawin Kraipetchara, Paulo Garcia

专题命中 VLA模型 :action model(title,abstract);分类 cs.RO

AI总结 本文提出通过整合符号方法与现成模型构建可验证的神经符号智能机器人动作模型,以提升可靠性和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11612 2025-12-15 cs.CV eess.IV 70%

Embodied Image Compression

具身图像压缩

Chunyi Li, Rui Qing, Jianbo Zhang, Yuan Tian, Xiangyang Zhu, Zicheng Zhang, Xiaohong Liu, Weisi Lin, Guangtao Zhai

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai AI Lab(上海人工智能实验室) Nanyang Technological University(南洋理工大学)

专题命中 VLA模型 :vision-language-action(abstract);action model(abstract);分类 cs.CV

AI总结 本文提出具身图像压缩问题,建立EmbodiedComp基准测试,证明现有模型在超低比特率下无法完成简单任务,推动具身AI在现实中的应用。

Comments 15 pages, 12 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07410 2025-12-15 cs.CV 57%

InterAgent: Physics-based Multi-agent Command Execution via Diffusion on Interaction Graphs

InterAgent:基于物理的多智能体命令执行通过交互图上的扩散

Bin Li, Ruichi Zhang, Han Liang, Jingyan Zhang, Juze Zhang, Xin Chen, Lan Xu, Jingyi Yu, Jingya Wang

机构 * ShanghaiTech University(上海科技大学) University of Pennsylvania(宾夕法尼亚大学) ByteDance(字节跳动) Stanford University(斯坦福大学) InstAdapt

专题命中 VLA模型 :action model(abstract);分类 cs.CV

AI总结 InterAgent通过交互图上的扩散模型实现了基于物理的多智能体协调控制,能够从文本提示中生成连贯且物理合理的多代理行为。

Comments Project page: https://binlee26.github.io/InterAgent-Page

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18333 2025-12-15 astro-ph.IM hep-ex 50%

Estimation of Temporal Muon Signals in Water-Cherenkov Detectors of the Surface Detector of the Pierre Auger Observatory

水切连检测器中地表探测器中缪子信号的估计

Margita Kubátová

专题命中 VLA模型 :action model(abstract)

AI总结 本文利用循环神经网络估计水切连检测器中缪子信号,以研究超高能宇宙射线的质量组成。

Comments Presented at the 2nd European AI for Fundamental Physics Conference (EuCAIFCon2025). Submission to SciPost Physics Proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏