arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

机器人 / 具身智能

机器人、具身智能、机器人学习、操作、导航和具身世界模型。

共收录 4124 信号源:cs.RO, cs.AI, cs.CV, cs.LG

1. 模仿学习与强化学习 4124 篇

2011.07778 2020-11-17 cs.RO 72%

Towards Autonomous Eye Surgery by Combining Deep Imitation Learning with Optimal Control

Ji Woong Kim, Peiyao Zhang, Peter Gehlbach, Iulian Iordachita, Marin Kobilarov

专题命中 模仿学习与强化学习 :manipulation(abstract);navigation(abstract);分类 cs.RO;robot learning(comments)

Comments Accepted to Conference on Robot Learning (CoRL) 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.04142 2019-10-10 cs.RO cs.AI cs.CV cs.LG cs.NE 72%

Imagined Value Gradients: Model-Based Policy Optimization with Transferable Latent Dynamics Models

Arunkumar Byravan, Jost Tobias Springenberg, Abbas Abdolmaleki, Roland Hafner, Michael Neunert, Thomas Lampe, Noah Siegel, Nicolas Heess, Martin Riedmiller

专题命中 模仿学习与强化学习 :manipulation(abstract);分类 cs.RO、cs.AI、cs.CV;robot learning(comments)

Comments To appear at the 3rd annual Conference on Robot Learning, Osaka, Japan (CoRL 2019). 24 pages including appendix (main paper - 8 pages)

详情

展开后加载摘要…

URL PDF HTML 收藏
1809.09369 2018-10-11 cs.LG stat.ML 72%

S-RL Toolbox: Environments, Datasets and Evaluation Metrics for State Representation Learning

Antonin Raffin, Ashley Hill, René Traoré, Timothée Lesort, Natalia Díaz-Rodríguez, David Filliat

专题命中 模仿学习与强化学习 :robotics(abstract,comments);robotic(abstract);分类 cs.LG

Comments Github repo: https://github.com/araffin/robotics-rl-srl Documentation: https://s-rl-toolbox.readthedocs.io/en/latest/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14959 2026-06-23 cs.RO cs.AI cs.LG cs.SY eess.SY 71%

CBF-RL: Safety Filtering Reinforcement Learning in Training with Control Barrier Functions

CBF-RL: 基于控制屏障函数的安全过滤强化学习

Lizhi Yang, Blake Werner, Massimiliano de Sa, Aaron D. Ames

机构 * Caltech MCE(加州理工学院机械工程系)

专题命中 模仿学习与强化学习 :navigation(abstract,comments);分类 cs.RO、cs.AI、cs.LG;robotics(comments)

AI总结 本文提出CBF-RL框架,通过在训练过程中强制控制屏障函数以生成安全行为,使强化学习策略内在化安全约束,实现无需在线安全过滤的鲁棒安全部署。

Comments Accepted to the 2026 IEEE International Conference on Robotics and Automation (ICRA 2026). Copyright transferred to IEEE. Sample code for the navigation example with CBF-RL reward core construction can be found at https://github.com/lzyang2000/cbf-rl-navigation-demo

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04054 2026-05-26 eess.SY cs.SY 71%

Necessary and Sufficient Conditions for the Optimization-Based Concurrent Execution of Learned Robotic Tasks

基于优化的学习型机器人任务并发执行的充分必要条件

Sheikh A. Tahmid, Gennaro Notomista

专题命中 模仿学习与强化学习 :robotic(title)

AI总结 本文提出定理,给出在状态空间子集内使用最小范数控制器并发执行一组学习任务(通过强化学习学习的值函数编码)的充分必要条件,并扩展框架以处理折扣因子训练的值函数。

Comments To be presented at ACC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03897 2026-04-17 cs.RO cs.AI cs.CL cs.HC cs.LG 71%

IROSA: Interactive Robot Skill Adaptation using Natural Language

IROSA: 基于自然语言的交互机器人技能适应

Markus Knauer, Samuel Bustamante, Thomas Eiband, Alin Albu-Schäffer, Freek Stulp, João Silvério

机构 * German Aerospace Center (DLR), Institute of Robotics and Mechatronics (RMC)(德国航空航天中心(DLR)机器人与机电研究所) School of Computation, Information and Technology (CIT), Technical University of Munich (TUM)(计算、信息与技术学院(CIT),慕尼黑技术大学)

专题命中 模仿学习与强化学习 :robotics(abstract,comments);分类 cs.RO、cs.AI、cs.LG

AI总结 本文提出IROSA框架,利用预训练大语言模型实现开放词汇技能适应,通过工具架构在语言模型与机器人硬件间保持抽象层,无需微调即可通过自然语言指令调整机器人技能。

Comments Accepted IEEE Robotics and Automation Letters (RA-L) journal, 8 pages, 5 figures, 3 tables, 1 listing. Code available: https://github.com/DLR-RM/IROSA

Journal ref IEEE Robotics and Automation Letters (RA-L), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.11687 2025-08-29 eess.SY cs.SY 71%

Coevolution of Opinion Dynamics and Recommendation System: Modeling, Analysis and Reinforcement Learning Based Manipulation

Yuhong Chen, Xiaobing Dai, Martin Buss, Fangzhou Liu

专题命中 模仿学习与强化学习 :manipulation(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.07185 2024-10-11 cs.RO cs.AI cs.LG 71%

Reward Learning from Suboptimal Demonstrations with Applications in Surgical Electrocautery

Zohre Karimi, Shing-Hei Ho, Bao Thach, Alan Kuntz, Daniel S. Brown

专题命中 模仿学习与强化学习 :robotic(abstract);robotics(comments,journal_ref);分类 cs.RO、cs.AI、cs.LG

Comments In proceedings of the International Symposium on Medical Robotics (ISMR) 2024. Equal contribution from two first authors

Journal ref 2024 International Symposium on Medical Robotics (ISMR), pp. 1-7, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.14139 2023-08-29 eess.SY cs.SY 71%

Reinforcement Learning-based Optimal Control and Software Rejuvenation for Safe and Efficient UAV Navigation

Angela Chen, Konstantinos Mitsopoulos, Raffaele Romagnoli

专题命中 模仿学习与强化学习 :navigation(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.02603 2022-12-07 cs.RO cs.AI cs.LG cs.SY eess.SY 71%

Learning to Optimize in Model Predictive Control

Jacob Sacks, Byron Boots

专题命中 模仿学习与强化学习 :robotics(abstract,comments);分类 cs.RO、cs.AI、cs.LG

Comments Proceedings of the IEEE Conference on Robotics and Automation (ICRA), 2022. Paper is 6 pages with 2 figures and 2 tables

Journal ref In 2022 International Conference on Robotics and Automation (ICRA), pp. 10549-10556. IEEE, 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.14711 2022-01-20 cs.LG cs.AI cs.RO 71%

Explanation-Aware Experience Replay in Rule-Dense Environments

Francesco Sovrano, Alex Raymond, Amanda Prorok

专题命中 模仿学习与强化学习 :navigation(abstract);robotics(comments,journal_ref);分类 cs.RO、cs.AI、cs.LG

Comments To appear in IEEE Robotics and Automation Letters (IEEE RA-L). Please cite the published version

Journal ref IEEE Robotics and Automation Letters ( Volume: 7, Issue: 2, April 2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.07971 2021-05-11 cs.AI cs.LG cs.RO 71%

Super-Human Performance in Gran Turismo Sport Using Deep Reinforcement Learning

Florian Fuchs, Yunlong Song, Elia Kaufmann, Davide Scaramuzza, Peter Duerr

专题命中 模仿学习与强化学习 :robotics(abstract,comments);分类 cs.RO、cs.AI、cs.LG

Comments Accepted for Publication at the IEEE Robotics and Automation Letters (RA-L) 2021, and International Conference on Robots and Automation (ICRA) 2021

Journal ref IEEE Robotics and Automation Letters (RAL) 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.05149 2020-07-15 cond-mat.quant-gas cond-mat.dis-nn 71%

Creation and manipulation of quantized vortices in Bose-Einstein condensates using reinforcement learning

Hiroki Saito

专题命中 模仿学习与强化学习 :manipulation(title)

Comments 7 pages, 5 figures, 2 movies

详情

展开后加载摘要…

URL PDF HTML 收藏
1901.08748 2019-02-21 quant-ph cond-mat.quant-gas 71%

Manipulation of Spin Dynamics by Deep Reinforcement Learning Agent

Jun-Jie Chen, Ming Xue

专题命中 模仿学习与强化学习 :manipulation(title)

Comments 8pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1707.06354 2018-02-07 cs.AI cs.HC cs.LG cs.RO 71%

Pragmatic-Pedagogic Value Alignment

Jaime F. Fisac, Monica A. Gates, Jessica B. Hamrick, Chang Liu, Dylan Hadfield-Menell, Malayandi Palaniappan, Dhruv Malik, S. Shankar Sastry, Thomas L. Griffiths, Anca D. Dragan

专题命中 模仿学习与强化学习 :robotics(abstract,comments);分类 cs.RO、cs.AI、cs.LG

Comments Published at the International Symposium on Robotics Research (ISRR 2017)

Journal ref International Symposium on Robotics Research, 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.21899 2026-08-25 cs.RO 新提交 70%

CIDER: Continual Interactive Distillation for Embodied Reinforcement Learning

CIDER:面向具身强化学习的持续交互式蒸馏

Houlin Li, Minghui Xu, Guo Xu, Xuan Du, Xiaohan Yan, Chun Wang, Yuxiang Yan, Shukai Yang, Yongcheng Liu, Wei Shan, Maoqing Yao

机构 * AgiBot Shanghai Jiao Tong University(上海交通大学)

专题命中 模仿学习与强化学习 :manipulation(abstract);robotic(abstract);分类 cs.RO

AI总结 本文提出CIDER框架,通过冻结历史策略为教师策略并引入梯度路由,解决具身强化学习的灾难性遗忘问题,在6个现实世界操作任务上实现了新技能获取与旧技能保留的平衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12743 2026-08-14 cs.AI 新提交 70%

Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence

空间记忆智能体:用于空间智能的基于经验的过程记忆

Haokai Zhang, Yuhang Ding, Yunshu Zhou, Xinze Du, Shengtao Zhang, Zhiyue Zhao, Yuling Xi, Hao Chen

机构 * Zhejiang University(浙江大学) Shanghai Jiao Tong University(上海交通大学) Shanghai Innovation Institute(上海创新研究院)

专题命中 模仿学习与强化学习 :embodied agent(abstract);robotic(abstract);分类 cs.AI

AI总结 该研究提出SMA框架,让冻结VLM智能体无需外部空间工具,通过无参数更新自进化提升空间推理,在多基准和模型上表现最优,提供了空间自进化的实用路径。

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07065 2026-08-10 cs.RO cs.AI cs.CV cs.HC cs.LG 新提交 70%

AutoIntervene: Calibrated Intervention for Action-Chunking Imitation Learning Policies

AutoIntervene:用于动作分块模仿学习策略的校准干预方法

Jinhe Tang, Weiming Zhi

机构 * Australian Center For Robotics, The University of Sydney(悉尼大学澳大利亚机器人中心) College of Connected Computing, Vanderbilt University(范德堡大学连接计算学院)

专题命中 模仿学习与强化学习 :manipulation(abstract);分类 cs.RO、cs.AI、cs.CV

AI总结 AutoIntervene是一种在线框架,通过校准的双向切换阈值在动作分块模仿学习策略与操作员间切换控制,提升了真实世界双臂操作任务的适应后成功率并降低了操作员控制时间。

Comments 9 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.02034 2026-08-04 cs.LG 新提交 70%

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning

用于离策略强化学习的上期望多步Q学习

Abdelghani Ghanem, Mounir Ghogho

专题命中 模仿学习与强化学习 :manipulation(abstract);navigation(abstract);分类 cs.LG

AI总结 针对离策略强化学习中多步回报引发的悲观偏差问题,提出ENQ算法,证明其理论性质,在27个任务上与LQL性能相当且吞吐量更高,集成多评论家时获益更多。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03438 2026-08-03 cs.AI 版本更新 70%

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents

具有自适应验证器的多模态强化学习

Reuben Tan, Baolin Peng, Zhengyuan Yang, Hao Cheng, Oier Mees, Theodore Zhao, Andrea Tupini, Isar Meijer, Qianhui Wu, Yuncong Yang, Lars Liden, Yu Gu, Sheng Zhang, Xiaodong Liu, Lijuan Wang, Marc Pollefeys, Yong Jae Lee, Jianfeng Gao

机构 * Microsoft Research(微软研究院) UMass Amherst(马萨诸塞大学阿默斯特分校) ETH Zurich(苏黎世联邦理工学院) UW–Madison(威斯康星大学麦迪逊分校)

专题命中 模仿学习与强化学习 :robotics(abstract);embodied AI(abstract);分类 cs.AI

AI总结 本文提出Argos验证器,通过结合SFT数据筛选与RL训练,提升多模态推理模型在空间推理、视觉幻觉及机器人任务中的性能,同时减少奖励黑客问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.26417 2026-07-30 cs.LG 新提交 70%

SCOUT: Per-Context Reset Curricula for Sparse-Reward Reinforcement Learning

SCOUT:面向稀疏奖励强化学习的每上下文重置课程

Siddharth Aphale, Ayushman Singh

专题命中 模仿学习与强化学习 :manipulation(abstract);navigation(abstract);分类 cs.LG

AI总结 SCOUT是为每个上下文提供独立重置课程的在线控制器,在六类任务中解决了全局进度无法适配学习差异的问题,无需组标签即可提升稀疏奖励强化学习效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06954 2026-07-24 cs.RO cs.SY eess.SY 版本更新 70%

Is Your Safe Controller Actually Safe? A Critical Review of CBF Tautologies and Hidden Assumptions

你的安全控制器真的安全吗?CBF永真式和隐藏假设的批判性审查

Taekyung Kim

机构 * Department of Robotics, University of Michigan(机器人学系,密歇根大学)

专题命中 模仿学习与强化学习 :navigation(abstract);robotic(abstract);分类 cs.RO

AI总结 本文批判性地审查了CBF在机器人安全中的应用,揭示了安全控制器理论与实际实现之间的差距,并提供了构建实际安全论证的实用指南。

Comments Technical Report. Interactive web demo: https://cbf.taekyung.me

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.18362 2026-07-22 cs.RO 新提交 70%

FARO: Feasibility-Aware Robot Motion Optimization

FARO:可行性感知机器人运动优化

Michal Ciebielski, Shafeef Omar, Aaron Johnson, Majid Khadiv

机构 * Munich Institute of Robotics and Machine Intelligence (MIRMI), Technical University of Munich (TUM)(慕尼黑工业大学慕尼黑机器人与机器智能研究所) Institute for Advanced Study, Technical University of Munich(慕尼黑工业大学高级研究所) Carnegie Mellon University(卡内基梅隆大学) SIEMENS AG(西门子公司) Huawei-TUM joint laboratory(华为-慕尼黑工业大学联合实验室)

专题命中 模仿学习与强化学习 :robotics(abstract);manipulation(abstract);分类 cs.RO

AI总结 针对机器人在未知场景快速规划新行为的挑战,提出嵌套运动动力学框架,结合可行性引导树搜索和大语言模型采样策略,能改善搜索过程,生成的轨迹可用强化学习控制器跟踪,质量高可用于实际运动操作场景。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07530 2026-07-21 cs.RO 版本更新 70%

ICLR: In-Context Imitation Learning with Visual Reasoning

ICLR: 基于视觉推理的上下文模仿学习

Toan Nguyen, Weiduo Yuan, Songlin Wei, Hui Li, Daniel Seita, Yue Wang

机构 * University of Southern California(南加州大学) Autodesk Research(Autodesk研究)

专题命中 模仿学习与强化学习 :manipulation(abstract);robotic(abstract);分类 cs.RO

AI总结 ICLR通过结合视觉推理与上下文模仿学习,提升机器人在复杂任务中的适应性和泛化能力。

Comments Accepted to IROS 2026. Project website: https://toannguyen1904.github.io/ICLR

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02390 2026-07-03 cs.LG 新提交 70%

DecompRL: Solving Harder Problems by Learning Modular Code Generation

DecompRL: 通过学习模块化代码生成解决更难的问题

Juliette Decugis, Fabian Gloeckle, Francis Bach, Taco Cohen, Gabriel Synnaeve

机构 * FAIR at Meta Inria, \'Ecole Normale Sup\'erieure

专题命中 模仿学习与强化学习 :world model(abstract,abstract_cn);分类 cs.LG

AI总结 提出DecompRL算法,通过强化学习将问题分解为可复用的子函数,重组模块生成候选解,将GPU瓶颈转移至CPU评估,在LiveCodeBench和CodeContests上超越标准RL基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01721 2026-07-03 cs.RO 新提交 70%

CoRe: Combined Rewards with Vision-Language Model Feedback for Preference-Aligned Reinforcement Learning

CoRe: 结合视觉语言模型反馈的复合奖励用于偏好对齐强化学习

Hexian Ni, Tao Lu, Yinghao Cai

专题命中 模仿学习与强化学习 :manipulation(abstract);robotic(abstract);分类 cs.RO

AI总结 提出CoRe框架,将奖励分解为基于任务知识的正式奖励和从观察中学习的残差奖励,利用视觉语言模型迭代优化奖励,实现无需人工的偏好对齐,在机器人操作任务中优于现有方法。

Comments ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.00808 2026-07-02 cs.LG 新提交 70%

Local Motion Matters: A Deconstruct-Recompose Paradigm for Reinforcement Learning Pre-training from Videos

局部运动至关重要:一种用于从视频中进行强化学习预训练的解构-重组范式

Jinwen Wang, Youfang Lin, Xiaobo Hu, Shuo Wang, Kai Lv

机构 * Beijing Jiaotong University(北京交通大学) Beijing Key Laboratory of Traffic Data Mining and Embodied Intelligence(北京交通数据挖掘与具身智能重点实验室)

专题命中 模仿学习与强化学习 :manipulation(abstract);robotic(abstract);分类 cs.LG

AI总结 提出解构-重组范式(DRP),通过解构全局运动为原子动作学习局部运动表示,再重组以加速下游策略学习,在机器人控制任务中显著提升样本效率。

Comments 20 pages, 16 figures

Journal ref Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026, pages 9859-9868

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.00796 2026-07-02 cs.LG 新提交 70%

Task-Relevant Representation Decoupling for Visual Reinforcement Learning Generalization

面向视觉强化学习泛化的任务相关表示解耦

Jinwen Wang, Youfang Lin, Xiaobo Hu, Qian Xu, Shuo Wang, Zhuo Chen, Kai Lv

机构 * Beijing Key Laboratory of Traffic Data Mining and Embodied Intelligence, School of Computer Science & Technology, Beijing Jiaotong University(北京交通大学计算机科学与技术学院交通数据挖掘与具身智能北京市重点实验室) CSSC Intelligent Innovation Research Institute(中国船舶集团智能创新研究院) Zhejiang University(浙江大学)

专题命中 模仿学习与强化学习 :manipulation(abstract);robotic(abstract);分类 cs.LG

AI总结 提出自监督任务相关表示解耦算法T2RD,通过表示一致性、交叉重建和交叉动态预测分离任务相关与无关特征,在控制任务中实现SOTA泛化性能。

Comments 23 pages, 13 figures

Journal ref ACM Transactions on Multimedia Computing, Communications and Applications (TOMM), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.30192 2026-06-30 cs.AI 70%

Domain Adaptation with Adaptive Imagination for Visual Reinforcement Learning under Limited Target Data

面向有限目标数据的视觉强化学习自适应想象域适应

Hyunwoo Park, Sang-Hyun Lee

机构 * STRADVISION Department of Mobility Engineering, Ajou University(移动工程系,阿乔大学)

专题命中 模仿学习与强化学习 :robotics(abstract,abstract_cn);分类 cs.AI

AI总结 提出AIDA框架,通过自适应想象生成可靠轨迹增强有限目标数据,并利用自一致性损失学习语义状态表示,在少数据下提升视觉强化学习迁移性能。

Comments 28 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26341 2026-06-26 cs.RO 新提交 70%

Scaling Nonlinear Optimization: Many Problems One GPU

扩展非线性优化:多问题单GPU

John Viljoen, Johanna Haffner, Masayoshi Tomizuka, Negar Mehr

机构 * Department of Mechanical Engineering, University of California, Berkeley(加州大学伯克利分校机械工程系) Department of Biosystems Science and Engineering, ETH Zürich(苏黎世联邦理工学院生物系统科学与工程系)

专题命中 模仿学习与强化学习 :robotics(abstract);navigation(abstract);分类 cs.RO

AI总结 提出首个基于JAX的GPU批处理NLP求解器jaxipm,通过异构迭代融合和迭代级批处理实现多问题并发求解,在四旋翼控制任务中吞吐量提升达32.85倍。

Comments 8 pages, 6 figures, ieeeconf style, submission for RA-L

详情

展开后加载摘要…

URL PDF HTML 收藏