arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

VLA / 视觉-语言-动作模型

视觉-语言-动作模型、机器人基础模型和语言条件机器人控制。

2025-12-04 至 2025-12-04 共收录 5 信号源:cs.RO, cs.CV, cs.AI, cs.LG

1. VLA模型 5 篇

2512.03913 2025-12-04 cs.RO cs.AI 90%

Hierarchical Vision Language Action Model Using Success and Failure Demonstrations

基于成功与失败示范的分层视觉语言行动模型

Jeongeun Park, Jihwan Yoon, Byungwoo Jeon, Juhan Park, Jinwoo Shin, Namhoon Cho, Kyungjae Lee, Sangdoo Yun, Sungjoon Choi

机构 * Department of Artificial Intelligence, Korea University(韩国大学人工智能系) Kim Jaechul Graduate School of AI, KAIST(金在拙人工智能研究生院,韩国科学技术院) Department of Aerospace Engineering, Seoul National University(首尔国立大学航空航天工程系) Departmnet of Statistics, Korea University(韩国大学统计系) NAVER AI Lab(NAVER人工智能实验室)

专题命中 VLA模型 :action model(title,abstract);vision language action(title);VLA(abstract,comments);vision-language-action(abstract)

AI总结 VINE通过分层强化学习框架,利用失败数据提升视觉语言行动模型的鲁棒性和执行能力。

Comments https://vine-vla.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04018 2025-12-04 cs.RO 88%

FPC-VLA: A Vision-Language-Action Framework with a Supervisor for Failure Prediction and Correction

FPC-VLA:一个带有监督器的视觉-语言-动作框架,用于故障预测和纠正

Yifan Yang, Zhixiang Duan, Tianshi Xie, Fuyu Cao, Pinxi Shen, Peili Song, Piaopiao Jin, Guokang Sun, Shaoqing Xu, Yangwei You, Jingtai Liu

机构 * The Institute of Robotics and Automatic Information System(机器人与自动信息系统研究所) Tianjin Key Laboratory of Intelligent Robotics(智能机器人天津重点实验室) TBI Center, Nankai University, Tianjin 300350, China(南开大学天津中心) Faculty of Robot Science and Engineering, Northeastern University, Shenyang 110819, China(机器人科学与工程学院) The State Key Laboratory of Internet of Things for Smart City(智能城市物联网国家重点实验室) Centre for Artificial Intelligence(人工智能中心) Department of Electromechanical Engineering, University of Macau, Macau SAR, China(机电工程系,澳门大学)

专题命中 VLA模型 :vision-language-action(title,abstract);VLA(title,abstract);分类 cs.RO

AI总结 FPC-VLA 提出了一种双模型框架,结合视觉-语言-动作模块与监督器,用于预测和纠正机器人操作中的故障,提升了自主系统的可靠性和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03828 2025-12-04 cs.RO 74%

IM HERE: Interaction Model for Human Effort Based Robot Engagement

IM HERE: 以人为中心的机器人互动模型

Dominykas Strazdas, Magnus Jung, Jan Marquenie, Ingo Siegert, Ayoub Al-Hamadi

机构 * Neuro-Information Technology Otto von Guericke University(神经信息技术奥托·冯·格里克大学) Mobile Dialog Systems Otto von Guericke University(移动对话系统奥托·冯·格里克大学)

专题命中 VLA模型 :action model(title);分类 cs.RO

AI总结 IM HERE提出了一种基于人类努力的机器人互动模型,旨在通过建模社交行为来实现自主系统与社会规范的协调。

Comments 8 pages, 5 figures

Journal ref 2025 IEEE Conference on Cognitive and Computational Aspects of Situation Management (CogSIMA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03308 2025-12-04 astro-ph.EP astro-ph.IM 50%

Roman coronagraph simulations of exozodi observations in the presence of wavefront errors

罗马冠状仪在存在波前误差情况下的系外尘埃观测模拟

Jorge Llop-Sayson, Vanessa P. Bailey, Justin Hom, John Krist, Bertrand Mennesson, Samantha N. Hasler, Alexandra Z. Greenbaum, A J Eldorado Riggs, Geoffrey Bryden

专题命中 VLA模型 :action model(abstract)

AI总结 罗马冠状仪在存在波前误差的情况下,通过模拟研究抖动对系外尘埃检测的影响,揭示了抖动与尘埃结构的退化问题及观测灵敏度变化。

Comments Accepted for publication in JATIS

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21004 2025-12-04 cs.SD 50%

CoHear: Conversation Enhancement via Multi-Earphone Collaboration

CoHear:通过多耳机协作提升对话

Lixing He, Yunqi Guo, Zhenyu Yan, Guoliang Xing

机构 * The Chinese University of Hong Kong(香港中文大学)

专题命中 VLA模型 :action model(abstract)

AI总结 CoHear通过多耳机协作提升对话质量,采用对话驱动网络协议和稳健的目标对话提取模型,实现高准确率和实时语音增强。

Comments Submitted to IMWUT

详情

展开后加载摘要…

URL PDF HTML 收藏