Towards Autonomous Eye Surgery by Combining Deep Imitation Learning with Optimal Control
专题命中 模仿学习与强化学习 :manipulation(abstract);navigation(abstract);分类 cs.RO;robot learning(comments)
Comments Accepted to Conference on Robot Learning (CoRL) 2020
视觉与机器人
机器人、具身智能、机器人学习、操作、导航和具身世界模型。
专题命中 模仿学习与强化学习 :manipulation(abstract);navigation(abstract);分类 cs.RO;robot learning(comments)
Comments Accepted to Conference on Robot Learning (CoRL) 2020
专题命中 模仿学习与强化学习 :manipulation(abstract);分类 cs.RO、cs.AI、cs.CV;robot learning(comments)
Comments To appear at the 3rd annual Conference on Robot Learning, Osaka, Japan (CoRL 2019). 24 pages including appendix (main paper - 8 pages)
专题命中 模仿学习与强化学习 :robotics(abstract,comments);robotic(abstract);分类 cs.LG
Comments Github repo: https://github.com/araffin/robotics-rl-srl Documentation: https://s-rl-toolbox.readthedocs.io/en/latest/
CBF-RL: 基于控制屏障函数的安全过滤强化学习
机构 * Caltech MCE(加州理工学院机械工程系)
专题命中 模仿学习与强化学习 :navigation(abstract,comments);分类 cs.RO、cs.AI、cs.LG;robotics(comments)
AI总结 本文提出CBF-RL框架,通过在训练过程中强制控制屏障函数以生成安全行为,使强化学习策略内在化安全约束,实现无需在线安全过滤的鲁棒安全部署。
Comments Accepted to the 2026 IEEE International Conference on Robotics and Automation (ICRA 2026). Copyright transferred to IEEE. Sample code for the navigation example with CBF-RL reward core construction can be found at https://github.com/lzyang2000/cbf-rl-navigation-demo
基于优化的学习型机器人任务并发执行的充分必要条件
专题命中 模仿学习与强化学习 :robotic(title)
AI总结 本文提出定理,给出在状态空间子集内使用最小范数控制器并发执行一组学习任务(通过强化学习学习的值函数编码)的充分必要条件,并扩展框架以处理折扣因子训练的值函数。
Comments To be presented at ACC 2026
IROSA: 基于自然语言的交互机器人技能适应
机构 * German Aerospace Center (DLR), Institute of Robotics and Mechatronics (RMC)(德国航空航天中心(DLR)机器人与机电研究所) ; School of Computation, Information and Technology (CIT), Technical University of Munich (TUM)(计算、信息与技术学院(CIT),慕尼黑技术大学)
专题命中 模仿学习与强化学习 :robotics(abstract,comments);分类 cs.RO、cs.AI、cs.LG
AI总结 本文提出IROSA框架,利用预训练大语言模型实现开放词汇技能适应,通过工具架构在语言模型与机器人硬件间保持抽象层,无需微调即可通过自然语言指令调整机器人技能。
Comments Accepted IEEE Robotics and Automation Letters (RA-L) journal, 8 pages, 5 figures, 3 tables, 1 listing. Code available: https://github.com/DLR-RM/IROSA
Journal ref IEEE Robotics and Automation Letters (RA-L), 2026
专题命中 模仿学习与强化学习 :manipulation(title)
专题命中 模仿学习与强化学习 :robotic(abstract);robotics(comments,journal_ref);分类 cs.RO、cs.AI、cs.LG
Comments In proceedings of the International Symposium on Medical Robotics (ISMR) 2024. Equal contribution from two first authors
Journal ref 2024 International Symposium on Medical Robotics (ISMR), pp. 1-7, 2024
专题命中 模仿学习与强化学习 :navigation(title)
专题命中 模仿学习与强化学习 :robotics(abstract,comments);分类 cs.RO、cs.AI、cs.LG
Comments Proceedings of the IEEE Conference on Robotics and Automation (ICRA), 2022. Paper is 6 pages with 2 figures and 2 tables
Journal ref In 2022 International Conference on Robotics and Automation (ICRA), pp. 10549-10556. IEEE, 2022
专题命中 模仿学习与强化学习 :navigation(abstract);robotics(comments,journal_ref);分类 cs.RO、cs.AI、cs.LG
Comments To appear in IEEE Robotics and Automation Letters (IEEE RA-L). Please cite the published version
Journal ref IEEE Robotics and Automation Letters ( Volume: 7, Issue: 2, April 2022)
专题命中 模仿学习与强化学习 :robotics(abstract,comments);分类 cs.RO、cs.AI、cs.LG
Comments Accepted for Publication at the IEEE Robotics and Automation Letters (RA-L) 2021, and International Conference on Robots and Automation (ICRA) 2021
Journal ref IEEE Robotics and Automation Letters (RAL) 2021
专题命中 模仿学习与强化学习 :manipulation(title)
Comments 7 pages, 5 figures, 2 movies
专题命中 模仿学习与强化学习 :manipulation(title)
Comments 8pages, 8 figures
专题命中 模仿学习与强化学习 :robotics(abstract,comments);分类 cs.RO、cs.AI、cs.LG
Comments Published at the International Symposium on Robotics Research (ISRR 2017)
Journal ref International Symposium on Robotics Research, 2017
CIDER:面向具身强化学习的持续交互式蒸馏
机构 * AgiBot ; Shanghai Jiao Tong University(上海交通大学)
专题命中 模仿学习与强化学习 :manipulation(abstract);robotic(abstract);分类 cs.RO
AI总结 本文提出CIDER框架,通过冻结历史策略为教师策略并引入梯度路由,解决具身强化学习的灾难性遗忘问题,在6个现实世界操作任务上实现了新技能获取与旧技能保留的平衡。
空间记忆智能体:用于空间智能的基于经验的过程记忆
机构 * Zhejiang University(浙江大学) ; Shanghai Jiao Tong University(上海交通大学) ; Shanghai Innovation Institute(上海创新研究院)
专题命中 模仿学习与强化学习 :embodied agent(abstract);robotic(abstract);分类 cs.AI
AI总结 该研究提出SMA框架,让冻结VLM智能体无需外部空间工具,通过无参数更新自进化提升空间推理,在多基准和模型上表现最优,提供了空间自进化的实用路径。
Comments Under Review
AutoIntervene:用于动作分块模仿学习策略的校准干预方法
机构 * Australian Center For Robotics, The University of Sydney(悉尼大学澳大利亚机器人中心) ; College of Connected Computing, Vanderbilt University(范德堡大学连接计算学院)
专题命中 模仿学习与强化学习 :manipulation(abstract);分类 cs.RO、cs.AI、cs.CV
AI总结 AutoIntervene是一种在线框架,通过校准的双向切换阈值在动作分块模仿学习策略与操作员间切换控制,提升了真实世界双臂操作任务的适应后成功率并降低了操作员控制时间。
Comments 9 pages, 7 figures
用于离策略强化学习的上期望多步Q学习
专题命中 模仿学习与强化学习 :manipulation(abstract);navigation(abstract);分类 cs.LG
AI总结 针对离策略强化学习中多步回报引发的悲观偏差问题,提出ENQ算法,证明其理论性质,在27个任务上与LQL性能相当且吞吐量更高,集成多评论家时获益更多。
具有自适应验证器的多模态强化学习
机构 * Microsoft Research(微软研究院) ; UMass Amherst(马萨诸塞大学阿默斯特分校) ; ETH Zurich(苏黎世联邦理工学院) ; UW–Madison(威斯康星大学麦迪逊分校)
专题命中 模仿学习与强化学习 :robotics(abstract);embodied AI(abstract);分类 cs.AI
AI总结 本文提出Argos验证器,通过结合SFT数据筛选与RL训练,提升多模态推理模型在空间推理、视觉幻觉及机器人任务中的性能,同时减少奖励黑客问题。
SCOUT:面向稀疏奖励强化学习的每上下文重置课程
专题命中 模仿学习与强化学习 :manipulation(abstract);navigation(abstract);分类 cs.LG
AI总结 SCOUT是为每个上下文提供独立重置课程的在线控制器,在六类任务中解决了全局进度无法适配学习差异的问题,无需组标签即可提升稀疏奖励强化学习效果。
你的安全控制器真的安全吗?CBF永真式和隐藏假设的批判性审查
机构 * Department of Robotics, University of Michigan(机器人学系,密歇根大学)
专题命中 模仿学习与强化学习 :navigation(abstract);robotic(abstract);分类 cs.RO
AI总结 本文批判性地审查了CBF在机器人安全中的应用,揭示了安全控制器理论与实际实现之间的差距,并提供了构建实际安全论证的实用指南。
Comments Technical Report. Interactive web demo: https://cbf.taekyung.me
FARO:可行性感知机器人运动优化
机构 * Munich Institute of Robotics and Machine Intelligence (MIRMI), Technical University of Munich (TUM)(慕尼黑工业大学慕尼黑机器人与机器智能研究所) ; Institute for Advanced Study, Technical University of Munich(慕尼黑工业大学高级研究所) ; Carnegie Mellon University(卡内基梅隆大学) ; SIEMENS AG(西门子公司) ; Huawei-TUM joint laboratory(华为-慕尼黑工业大学联合实验室)
专题命中 模仿学习与强化学习 :robotics(abstract);manipulation(abstract);分类 cs.RO
AI总结 针对机器人在未知场景快速规划新行为的挑战,提出嵌套运动动力学框架,结合可行性引导树搜索和大语言模型采样策略,能改善搜索过程,生成的轨迹可用强化学习控制器跟踪,质量高可用于实际运动操作场景。
ICLR: 基于视觉推理的上下文模仿学习
机构 * University of Southern California(南加州大学) ; Autodesk Research(Autodesk研究)
专题命中 模仿学习与强化学习 :manipulation(abstract);robotic(abstract);分类 cs.RO
AI总结 ICLR通过结合视觉推理与上下文模仿学习,提升机器人在复杂任务中的适应性和泛化能力。
Comments Accepted to IROS 2026. Project website: https://toannguyen1904.github.io/ICLR
DecompRL: 通过学习模块化代码生成解决更难的问题
机构 * FAIR at Meta ; Inria, \'Ecole Normale Sup\'erieure
专题命中 模仿学习与强化学习 :world model(abstract,abstract_cn);分类 cs.LG
AI总结 提出DecompRL算法,通过强化学习将问题分解为可复用的子函数,重组模块生成候选解,将GPU瓶颈转移至CPU评估,在LiveCodeBench和CodeContests上超越标准RL基线。
CoRe: 结合视觉语言模型反馈的复合奖励用于偏好对齐强化学习
专题命中 模仿学习与强化学习 :manipulation(abstract);robotic(abstract);分类 cs.RO
AI总结 提出CoRe框架,将奖励分解为基于任务知识的正式奖励和从观察中学习的残差奖励,利用视觉语言模型迭代优化奖励,实现无需人工的偏好对齐,在机器人操作任务中优于现有方法。
Comments ICML 2026
局部运动至关重要:一种用于从视频中进行强化学习预训练的解构-重组范式
机构 * Beijing Jiaotong University(北京交通大学) ; Beijing Key Laboratory of Traffic Data Mining and Embodied Intelligence(北京交通数据挖掘与具身智能重点实验室)
专题命中 模仿学习与强化学习 :manipulation(abstract);robotic(abstract);分类 cs.LG
AI总结 提出解构-重组范式(DRP),通过解构全局运动为原子动作学习局部运动表示,再重组以加速下游策略学习,在机器人控制任务中显著提升样本效率。
Comments 20 pages, 16 figures
Journal ref Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026, pages 9859-9868
面向视觉强化学习泛化的任务相关表示解耦
机构 * Beijing Key Laboratory of Traffic Data Mining and Embodied Intelligence, School of Computer Science & Technology, Beijing Jiaotong University(北京交通大学计算机科学与技术学院交通数据挖掘与具身智能北京市重点实验室) ; CSSC Intelligent Innovation Research Institute(中国船舶集团智能创新研究院) ; Zhejiang University(浙江大学)
专题命中 模仿学习与强化学习 :manipulation(abstract);robotic(abstract);分类 cs.LG
AI总结 提出自监督任务相关表示解耦算法T2RD,通过表示一致性、交叉重建和交叉动态预测分离任务相关与无关特征,在控制任务中实现SOTA泛化性能。
Comments 23 pages, 13 figures
Journal ref ACM Transactions on Multimedia Computing, Communications and Applications (TOMM), 2026
面向有限目标数据的视觉强化学习自适应想象域适应
机构 * STRADVISION ; Department of Mobility Engineering, Ajou University(移动工程系,阿乔大学)
专题命中 模仿学习与强化学习 :robotics(abstract,abstract_cn);分类 cs.AI
AI总结 提出AIDA框架,通过自适应想象生成可靠轨迹增强有限目标数据,并利用自一致性损失学习语义状态表示,在少数据下提升视觉强化学习迁移性能。
Comments 28 pages, 10 figures
扩展非线性优化:多问题单GPU
机构 * Department of Mechanical Engineering, University of California, Berkeley(加州大学伯克利分校机械工程系) ; Department of Biosystems Science and Engineering, ETH Zürich(苏黎世联邦理工学院生物系统科学与工程系)
专题命中 模仿学习与强化学习 :robotics(abstract);navigation(abstract);分类 cs.RO
AI总结 提出首个基于JAX的GPU批处理NLP求解器jaxipm,通过异构迭代融合和迭代级批处理实现多问题并发求解,在四旋翼控制任务中吞吐量提升达32.85倍。
Comments 8 pages, 6 figures, ieeeconf style, submission for RA-L