arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

VLA / 视觉-语言-动作模型

视觉-语言-动作模型、机器人基础模型和语言条件机器人控制。

2025-12-04 至 2025-12-04 共收录 8 信号源:cs.RO, cs.CV, cs.AI, cs.LG

1. VLA模型 5 篇

2512.03913 2025-12-04 cs.RO cs.AI 90%

Hierarchical Vision Language Action Model Using Success and Failure Demonstrations

基于成功与失败示范的分层视觉语言行动模型

Jeongeun Park, Jihwan Yoon, Byungwoo Jeon, Juhan Park, Jinwoo Shin, Namhoon Cho, Kyungjae Lee, Sangdoo Yun, Sungjoon Choi

机构 * Department of Artificial Intelligence, Korea University(韩国大学人工智能系) Kim Jaechul Graduate School of AI, KAIST(金在拙人工智能研究生院,韩国科学技术院) Department of Aerospace Engineering, Seoul National University(首尔国立大学航空航天工程系) Departmnet of Statistics, Korea University(韩国大学统计系) NAVER AI Lab(NAVER人工智能实验室)

专题命中 VLA模型 :action model(title,abstract);vision language action(title);VLA(abstract,comments);vision-language-action(abstract)

AI总结 VINE通过分层强化学习框架,利用失败数据提升视觉语言行动模型的鲁棒性和执行能力。

Comments https://vine-vla.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04018 2025-12-04 cs.RO 88%

FPC-VLA: A Vision-Language-Action Framework with a Supervisor for Failure Prediction and Correction

FPC-VLA:一个带有监督器的视觉-语言-动作框架,用于故障预测和纠正

Yifan Yang, Zhixiang Duan, Tianshi Xie, Fuyu Cao, Pinxi Shen, Peili Song, Piaopiao Jin, Guokang Sun, Shaoqing Xu, Yangwei You, Jingtai Liu

机构 * The Institute of Robotics and Automatic Information System(机器人与自动信息系统研究所) Tianjin Key Laboratory of Intelligent Robotics(智能机器人天津重点实验室) TBI Center, Nankai University, Tianjin 300350, China(南开大学天津中心) Faculty of Robot Science and Engineering, Northeastern University, Shenyang 110819, China(机器人科学与工程学院) The State Key Laboratory of Internet of Things for Smart City(智能城市物联网国家重点实验室) Centre for Artificial Intelligence(人工智能中心) Department of Electromechanical Engineering, University of Macau, Macau SAR, China(机电工程系,澳门大学)

专题命中 VLA模型 :vision-language-action(title,abstract);VLA(title,abstract);分类 cs.RO

AI总结 FPC-VLA 提出了一种双模型框架,结合视觉-语言-动作模块与监督器,用于预测和纠正机器人操作中的故障,提升了自主系统的可靠性和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03828 2025-12-04 cs.RO 74%

IM HERE: Interaction Model for Human Effort Based Robot Engagement

IM HERE: 以人为中心的机器人互动模型

Dominykas Strazdas, Magnus Jung, Jan Marquenie, Ingo Siegert, Ayoub Al-Hamadi

机构 * Neuro-Information Technology Otto von Guericke University(神经信息技术奥托·冯·格里克大学) Mobile Dialog Systems Otto von Guericke University(移动对话系统奥托·冯·格里克大学)

专题命中 VLA模型 :action model(title);分类 cs.RO

AI总结 IM HERE提出了一种基于人类努力的机器人互动模型,旨在通过建模社交行为来实现自主系统与社会规范的协调。

Comments 8 pages, 5 figures

Journal ref 2025 IEEE Conference on Cognitive and Computational Aspects of Situation Management (CogSIMA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03308 2025-12-04 astro-ph.EP astro-ph.IM 50%

Roman coronagraph simulations of exozodi observations in the presence of wavefront errors

罗马冠状仪在存在波前误差情况下的系外尘埃观测模拟

Jorge Llop-Sayson, Vanessa P. Bailey, Justin Hom, John Krist, Bertrand Mennesson, Samantha N. Hasler, Alexandra Z. Greenbaum, A J Eldorado Riggs, Geoffrey Bryden

专题命中 VLA模型 :action model(abstract)

AI总结 罗马冠状仪在存在波前误差的情况下,通过模拟研究抖动对系外尘埃检测的影响,揭示了抖动与尘埃结构的退化问题及观测灵敏度变化。

Comments Accepted for publication in JATIS

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21004 2025-12-04 cs.SD 50%

CoHear: Conversation Enhancement via Multi-Earphone Collaboration

CoHear:通过多耳机协作提升对话

Lixing He, Yunqi Guo, Zhenyu Yan, Guoliang Xing

机构 * The Chinese University of Hong Kong(香港中文大学)

专题命中 VLA模型 :action model(abstract)

AI总结 CoHear通过多耳机协作提升对话质量,采用对话驱动网络协议和稳健的目标对话提取模型,实现高准确率和实时语音增强。

Comments Submitted to IMWUT

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 动作表示与策略 2 篇

2410.16411 2025-12-04 cs.RO cs.LG 62%

The Duality of Generative AI and Reinforcement Learning in Robotics: A Review

生成式人工智能与强化学习在机器人中的双重性:综述

Angelo Moroncelli, Vishal Soni, Marco Forgione, Dario Piga, Blerina Spahiu, Loris Roveda

机构 * SUPSI(瑞士苏黎世联邦理工学院)

专题命中 动作表示与策略 :VLA(abstract);分类 cs.RO、cs.LG

AI总结 本文综述了生成式人工智能与强化学习在机器人中的双重性,探讨了两者的整合方法、挑战及未来研究方向。

Comments Submitted for publication to Information Fusion

Journal ref Information Fusion Volume 129, May 2026, 104003

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03538 2025-12-04 cs.RO 57%

AdaPower: Specializing World Foundation Models for Predictive Manipulation

AdaPower: 为预测操纵专门化世界基础模型

Yuhang Huang, Shilong Zou, Jiazhao Zhang, Xinwang Liu, Ruizhen Hu, Kai Xu

机构 * National University of Defense Technology(国防科技大学) Peking University(北京大学) Shenzhen University(深圳大学)

专题命中 动作表示与策略 :VLA(abstract);分类 cs.RO

AI总结 AdaPower通过时空测试时间训练和记忆持久性机制,将通用世界基础模型转化为专门模型,提升机器人任务成功率41%

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 数据集与评测 1 篇

2504.08396 2025-12-04 stat.AP 50%

Fairness is in the details: Face Dataset Auditing

细节中的公平性:人脸数据集审计

Valentin Lafargue, Emmanuelle Claeys, Jean-Michel Loubes

专题命中 数据集与评测 :action model(abstract)

AI总结 本文提出了一种基于敏感特征的自动方法,用于审计人脸数据集以确保公平性。

Journal ref Machine Learning and Knowledge Discovery in Databases. Applied Data Science Track and Demo Track. ECML PKDD 2025. Lecture Notes in Computer Science(), vol 16022. Springer, Cham

详情

展开后加载摘要…

URL PDF HTML 收藏