arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

机器人 / 具身智能

机器人、具身智能、机器人学习、操作、导航和具身世界模型。

至 收录 60418 信号源:cs.RO, cs.AI, cs.CV, cs.LG
2511.07416 2025-11-11 cs.RO cs.AI cs.CV 92%

Robot Learning from a Physical World Model

Jiageng Mao, Sicheng He, Hao-Ning Wu, Yang You, Shuyang Sun, Zhicheng Wang, Yanan Bao, Huizhong Chen, Leonidas Guibas, Vitor Guizilini, Howard Zhou, Yue Wang

机构 * Google DeepMind(谷歌DeepMind) USC(美国斯克利普斯大学) Stanford(斯坦福大学) Toyota Research Institute(丰田研究机构)

专题命中 机器人学习 :robot learning(title,abstract);world model(title,abstract);robotics(abstract);manipulation(abstract)

Comments Project page: https://pointscoder.github.io/PhysWorld_Web/

详情
英文摘要

We introduce PhysWorld, a framework that enables robot learning from video generation through physical world modeling. Recent video generation models can synthesize photorealistic visual demonstrations from language commands and images, offering a powerful yet underexplored source of training signals for robotics. However, directly retargeting pixel motions from generated videos to robots neglects physics, often resulting in inaccurate manipulations. PhysWorld addresses this limitation by coupling video generation with physical world reconstruction. Given a single image and a task command, our method generates task-conditioned videos and reconstructs the underlying physical world from the videos, and the generated video motions are grounded into physically accurate actions through object-centric residual reinforcement learning with the physical world model. This synergy transforms implicit visual guidance into physically executable robotic trajectories, eliminating the need for real robot data collection and enabling zero-shot generalizable robotic manipulation. Experiments on diverse real-world tasks demonstrate that PhysWorld substantially improves manipulation accuracy compared to previous approaches. Visit \href{https://pointscoder.github.io/PhysWorld_Web/}{the project webpage} for details.

URL PDF HTML 收藏
2605.00080 2026-05-04 cs.RO cs.CV 91%

World Model for Robot Learning: A Comprehensive Survey

机器人学习的世界模型:全面综述

Bohan Hou, Gen Li, Jindou Jia, Tuo An, Xinying Guo, Sicong Leng, Haoran Geng, Yanjie Ze, Tatsuya Harada, Philip Torr, Oier Mees, Marc Pollefeys, Zhuang Liu, Jiajun Wu, Pieter Abbeel, Jitendra Malik, Yilun Du, Jianfei Yang

机构 * Nanyang Technological University(南洋理工大学) University of California, Berkeley(加州大学伯克利分校) Stanford University(斯坦福大学) The University of Tokyo(东京大学) University of Oxford(牛津大学) Microsoft(微软公司) ETH Zurich(苏黎世联邦理工学院) Princeton University(普林斯顿大学) Harvard University(哈佛大学)

专题命中 机器人学习 :robot learning(title,abstract);world model(title,abstract);embodied agent(abstract);navigation(abstract)

AI总结 综述从机器人学习角度系统回顾世界模型的快速发展,探讨其与机器人策略的耦合、作为强化学习模拟器的作用以及机器人视频世界模型的发展,总结关键范式与挑战。

Comments 43 pages, 6 figures

详情
AI中文摘要

世界模型作为预测环境演变的表示,已成为机器人学习的核心组件。它们支持策略学习、规划、模拟、评估、数据生成,并随着基础模型和大规模视频生成的兴起而迅速发展。然而,文献在架构、功能角色和具身应用领域仍碎片化。为此,本文从机器人学习角度进行全面综述,探讨世界模型如何与机器人策略耦合,如何作为强化学习和评估的学得模拟器,以及机器人视频世界模型如何从基于想象的生成发展到可控、结构化和基础规模的公式化。进一步将这些理念联系到导航和自动驾驶,并总结代表性数据集、基准和评估协议。总体而言,本文系统回顾了机器人学习中世界模型快速发展的文献,澄清了关键范式和应用,并突出了预测建模在具身代理中的主要挑战和未来方向。为方便持续访问新兴工作、基准和资源,本文将维护并定期更新配套的GitHub仓库。

英文摘要

World models, which are predictive representations of how environments evolve under actions, have become a central component of robot learning. They support policy learning, planning, simulation, evaluation, data generation, and have advanced rapidly with the rise of foundation models and large-scale video generation. However, the literature remains fragmented across architectures, functional roles, and embodied application domains. To address this gap, we present a comprehensive review of world models from a robot-learning perspective. We examine how world models are coupled with robot policies, how they serve as learned simulators for reinforcement learning and evaluation, and how robotic video world models have progressed from imagination-based generation to controllable, structured, and foundation-scale formulations. We further connect these ideas to navigation and autonomous driving, and summarize representative datasets, benchmarks, and evaluation protocols. Overall, this survey systematically reviews the rapidly growing literature on world models for robot learning, clarifies key paradigms and applications, and highlights major challenges and future directions for predictive modeling in embodied agents. To facilitate continued access to newly emerging works, benchmarks, and resources, we will maintain and regularly update the accompanying GitHub repository alongside this survey.

URL PDF HTML 收藏
2604.19683 2026-04-23 cs.RO 91%

Mask World Model: Predicting What Matters for Robust Robot Policy Learning

遮罩世界模型:为鲁棒机器人策略学习预测重要的内容

Yunfan Lou, Xiaowei Chi, Xiaojie Zhang, Zezhong Qian, Chengxuan Li, Rongyu Zhang, Yaoxu Lyu, Guoyu Song, Chuyao Fu, Haoxuan Xu, Pengwei Wang, Shanghang Zhang

机构 * Peking University, Beijing, China State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University Beijing Academy of Artificial Intelligence, Beijing, China National University of Singapore, Singapore The Hong Kong University of Science Nanjing University, Nanjing, China

专题命中 机器人学习 :world model(title,summary_cn);robot policy(title,abstract);分类 cs.RO

AI总结 本文提出Mask World Model,通过预测语义遮罩而非像素,提升机器人策略学习的鲁棒性与泛化能力,实验证明其在仿真和现实任务中均优于现有方法。

Comments 16 pages,5 figures

详情
AI中文摘要

从大规模视频生成预训练中衍生出的世界模型已成为通用机器人策略学习有前景的范式。然而,标准方法往往聚焦于高保真RGB视频预测,这会导致过度拟合无关因素,如动态背景和光照变化。这些干扰会降低模型的泛化能力,最终导致不可靠且脆弱的控制策略。为了解决这个问题,我们引入了Mask World Model (MWM),利用视频扩散架构预测语义遮罩的演变而非像素。这种转变施加了一个几何信息瓶颈,迫使模型捕捉必要的物理动态和接触关系,同时过滤掉视觉噪声。我们无缝集成此遮罩动态骨干与基于扩散的策略头,以实现鲁棒的端到端控制。广泛的评估显示,MWM在LIBERO和RLBench仿真基准上表现优异,显著优于基于RGB的世界模型。此外,现实世界实验和鲁棒性评估(通过随机令牌剪枝)表明,MWM在泛化能力和对纹理信息丢失的鲁棒性方面表现更优。

英文摘要

World models derived from large-scale video generative pre-training have emerged as a promising paradigm for generalist robot policy learning. However, standard approaches often focus on high-fidelity RGB video prediction, this can result in overfitting to irrelevant factors, such as dynamic backgrounds and illumination changes. These distractions reduce the model's ability to generalize, ultimately leading to unreliable and fragile control policies. To address this, we introduce the Mask World Model (MWM), which leverages video diffusion architectures to predict the evolution of semantic masks instead of pixels. This shift imposes a geometric information bottleneck, forcing the model to capture essential physical dynamics and contact relations while filtering out visual noise. We seamlessly integrate this mask dynamics backbone with a diffusion-based policy head to enable robust end-to-end control. Extensive evaluations demonstrate the superiority of MWM on the LIBERO and RLBench simulation benchmarks, significantly outperforming the state-of-the-art RGB-based world models. Furthermore, real-world experiments and robustness evaluation (via random token pruning) reveal that MWM exhibits superior generalization capabilities and robust resilience to texture information loss.

URL PDF HTML 收藏
2505.15659 2025-05-22 cs.RO cs.LG 91%

FLARE: Robot Learning with Implicit World Modeling

Ruijie Zheng, Jing Wang, Scott Reed, Johan Bjorck, Yu Fang, Fengyuan Hu, Joel Jang, Kaushil Kundalia, Zongyu Lin, Loic Magne, Avnish Narayan, You Liang Tan, Guanzhi Wang, Qi Wang, Jiannan Xiang, Yinzhen Xu, Seonghyeon Ye, Jan Kautz, Furong Huang, Yuke Zhu, Linxi Fan

机构 * NVIDIA University of Maryland, College Park(马里兰大学) Nanyang Technological University(南洋理工大学) University of Texas, Austin(德克萨斯大学)

专题命中 机器人学习 :world model(title,abstract);robot learning(title);manipulation(abstract);robot policy(abstract)

Comments Project Webpage / Blogpost: https://research.nvidia.com/labs/gear/flare

详情
英文摘要

We introduce $\textbf{F}$uture $\textbf{LA}$tent $\textbf{RE}$presentation Alignment ($\textbf{FLARE}$), a novel framework that integrates predictive latent world modeling into robot policy learning. By aligning features from a diffusion transformer with latent embeddings of future observations, $\textbf{FLARE}$ enables a diffusion transformer policy to anticipate latent representations of future observations, allowing it to reason about long-term consequences while generating actions. Remarkably lightweight, $\textbf{FLARE}$ requires only minimal architectural modifications -- adding a few tokens to standard vision-language-action (VLA) models -- yet delivers substantial performance gains. Across two challenging multitask simulation imitation learning benchmarks spanning single-arm and humanoid tabletop manipulation, $\textbf{FLARE}$ achieves state-of-the-art performance, outperforming prior policy learning baselines by up to 26%. Moreover, $\textbf{FLARE}$ unlocks the ability to co-train with human egocentric video demonstrations without action labels, significantly boosting policy generalization to a novel object with unseen geometry with as few as a single robot demonstration. Our results establish $\textbf{FLARE}$ as a general and scalable approach for combining implicit world modeling with high-frequency robotic control.

URL PDF HTML 收藏
2402.02385 2024-02-07 cs.RO cs.AI 91%

A Survey on Robotics with Foundation Models: toward Embodied AI

Zhiyuan Xu, Kun Wu, Junjie Wen, Jinming Li, Ning Liu, Zhengping Che, Jian Tang

专题命中 机器人学习 :robotics(title,abstract);embodied AI(title,abstract);robot learning(abstract);manipulation(abstract)

详情
英文摘要

While the exploration for embodied AI has spanned multiple decades, it remains a persistent challenge to endow agents with human-level intelligence, including perception, learning, reasoning, decision-making, control, and generalization capabilities, so that they can perform general-purpose tasks in open, unstructured, and dynamic environments. Recent advances in computer vision, natural language processing, and multi-modality learning have shown that the foundation models have superhuman capabilities for specific tasks. They not only provide a solid cornerstone for integrating basic modules into embodied AI systems but also shed light on how to scale up robot learning from a methodological perspective. This survey aims to provide a comprehensive and up-to-date overview of foundation models in robotics, focusing on autonomous manipulation and encompassing high-level planning and low-level control. Moreover, we showcase their commonly used datasets, simulators, and benchmarks. Importantly, we emphasize the critical challenges intrinsic to this field and delineate potential avenues for future research, contributing to advancing the frontier of academic and industrial discourse.

URL PDF HTML 收藏
2210.12278 2022-10-25 cs.RO cs.AI 91%

Sample Efficient Robot Learning with Structured World Models

Tuluhan Akbulut, Max Merlin, Shane Parr, Benedict Quartey, Skye Thompson

专题命中 机器人学习 :world model(title,abstract);robot learning(title);robotics(abstract);manipulation(abstract)

详情
英文摘要

Reinforcement learning has been demonstrated as a flexible and effective approach for learning a range of continuous control tasks, such as those used by robots to manipulate objects in their environment. But in robotics particularly, real-world rollouts are costly, and sample efficiency can be a major limiting factor when learning a new skill. In game environments, the use of world models has been shown to improve sample efficiency while still achieving good performance, especially when images or other rich observations are provided. In this project, we explore the use of a world model in a deformable robotic manipulation task, evaluating its effect on sample efficiency when learning to fold a cloth in simulation. We compare the use of RGB image observation with a feature space leveraging built-in structure (keypoints representing the cloth configuration), a common approach in robot skill learning, and compare the impact on task performance and learning efficiency with and without the world model. Our experiments showed that the usage of keypoints increased the performance of the best model on the task by 50%, and in general, the use of a learned or constructed reduced feature space improved task performance and sample efficiency. The use of a state transition predictor(MDN-RNN) in our world models did not have a notable effect on task performance.

URL PDF HTML 收藏
1612.07139 2026-06-04 cs.RO cs.AI cs.LG cs.SY eess.SY 90%

A Survey of Deep Network Solutions for Learning Control in Robotics: From Reinforcement to Imitation

深度网络在机器人学习控制中的应用综述:从强化到模仿

Lei Tai, Jingwei Zhang, Ming Liu, Joschka Boedecker, Wolfram Burgard

机构 * University of Freiburg(弗赖堡大学)

专题命中 机器人学习 :manipulation(summary_cn,abstract);robotics(title,abstract);navigation(abstract);分类 cs.RO、cs.AI、cs.LG

AI总结 本文综述了深度学习在机器人学习控制中的应用,探讨了深度强化学习和模仿学习两大主流方法,分析了其在导航、 manipulation 任务中的应用及现实差距挑战。

Comments 19 pages, 1 figures

详情
AI中文摘要

深度学习技术已广泛应用于各种研究领域,取得了最先进的成果。本文综述了针对机器人应用的学习控制策略的深度学习解决方案。我们讨论了深度学习在学习控制中的两大主要范式:深度强化学习和模仿学习。对于深度强化学习(DRL),我们从传统强化学习算法开始,展示了如何将其扩展到深度领域,并介绍了在机器人导航和 manipulation 任务中使用 DRL 的代表性工作。我们继续讨论了解决现实差距挑战的方法,即如何将仿真中训练的 DRL 策略转移到现实世界场景,并总结了用于 DRL 研究的机器人仿真平台。对于模仿学习,我们探讨了其三个主要类别:行为克隆、逆强化学习和生成对抗模仿学习,介绍了它们的公式及其在机器人应用中的对应情况。最后,我们讨论了开放挑战和研究前沿。

英文摘要

Deep learning techniques have been widely applied, achieving state-of-the-art results in various fields of study. This survey focuses on deep learning solutions that target learning control policies for robotics applications. We carry out our discussions on the two main paradigms for learning control with deep networks: deep reinforcement learning and imitation learning. For deep reinforcement learning (DRL), we begin from traditional reinforcement learning algorithms, showing how they are extended to the deep context and effective mechanisms that could be added on top of the DRL algorithms. We then introduce representative works that utilize DRL to solve navigation and manipulation tasks in robotics. We continue our discussion on methods addressing the challenge of the reality gap for transferring DRL policies trained in simulation to real-world scenarios, and summarize robotics simulation platforms for conducting DRL research. For imitation leaning, we go through its three main categories, behavior cloning, inverse reinforcement learning and generative adversarial imitation learning, by introducing their formulations and their corresponding robotics applications. Finally, we discuss the open challenges and research frontiers.

URL PDF HTML 收藏
2206.14176 2022-06-29 cs.RO cs.AI cs.LG 90%

DayDreamer: World Models for Physical Robot Learning

Philipp Wu, Alejandro Escontrela, Danijar Hafner, Ken Goldberg, Pieter Abbeel

专题命中 机器人学习 :robot learning(title,abstract);world model(title,abstract);robotic(abstract);分类 cs.RO、cs.AI、cs.LG

Comments Website: https://danijar.com/daydreamer

详情
英文摘要

To solve tasks in complex environments, robots need to learn from experience. Deep reinforcement learning is a common approach to robot learning but requires a large amount of trial and error to learn, limiting its deployment in the physical world. As a consequence, many advances in robot learning rely on simulators. On the other hand, learning inside of simulators fails to capture the complexity of the real world, is prone to simulator inaccuracies, and the resulting behaviors do not adapt to changes in the world. The Dreamer algorithm has recently shown great promise for learning from small amounts of interaction by planning within a learned world model, outperforming pure reinforcement learning in video games. Learning a world model to predict the outcomes of potential actions enables planning in imagination, reducing the amount of trial and error needed in the real environment. However, it is unknown whether Dreamer can facilitate faster learning on physical robots. In this paper, we apply Dreamer to 4 robots to learn online and directly in the real world, without simulators. Dreamer trains a quadruped robot to roll off its back, stand up, and walk from scratch and without resets in only 1 hour. We then push the robot and find that Dreamer adapts within 10 minutes to withstand perturbations or quickly roll over and stand back up. On two different robotic arms, Dreamer learns to pick and place multiple objects directly from camera images and sparse rewards, approaching human performance. On a wheeled robot, Dreamer learns to navigate to a goal position purely from camera images, automatically resolving ambiguity about the robot orientation. Using the same hyperparameters across all experiments, we find that Dreamer is capable of online learning in the real world, establishing a strong baseline. We release our infrastructure for future applications of world models to robot learning.

URL PDF HTML 收藏
2606.09499 2026-06-09 cs.RO cs.AI cs.CR 新提交 90%

Targeting World Models to Compromise Robot Learning Pipelines

针对世界模型以破坏机器人学习流程

Ethan Rathbun, Ahmed Agha, Saaduddin Mahmud, Christopher Amato, Alina Oprea, Eugene Bagdasarian

机构 * Northeastern University(东北大学) University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)

专题命中 机器人学习 :robot learning(title,abstract);world model(title,abstract);robotic(abstract);分类 cs.RO、cs.AI

AI总结 本文提出针对世界模型的新型数据投毒攻击方法,通过注入恶意提示或转换动态,在看似安全的数据中生成危险训练轨迹,导致下游策略不安全。

Comments 8 Pages, CoRL Preprint

详情
AI中文摘要

世界模型近来在流行度和能力上迅速增长,成为生成机器人训练数据或模拟真实环境的更高效工具,许多工作提议将其集成到机器人学习流程中。尽管非常实用,但本文证明世界模型引入了机器人学习供应链中一种独特隐蔽且有效的数据投毒入口,可能导致部署不安全或受损的机器人策略,尽管训练数据看似安全。与传统数据投毒技术直接向已售或上传数据集中植入危险轨迹不同,我们的新型攻击方法将恶意提示或受损转换动态注入到视觉安全的遥操作数据集中,这些数据仅当通过世界模型作为输入时才会被激活。这可能导致生成合成的危险机器人训练轨迹,进而产生不安全或受损的机器人策略。我们展示了针对最先进的行动条件和文本条件世界模型的攻击有效性,展示了在下游DRL策略上的完整端到端后门攻击,以及针对VLA设置的概念验证。总体而言,这些发现需要研究更安全的世界模型,并重新评估其在机器人学习供应链中的地位。

英文摘要

World models have recently seen a rapid growth in both their popularity and capability as more data efficient tools for generating robot training data or simulating real world environments, with many works proposing their integration into the robot learning pipeline. While highly practical, in this work we demonstrate that world models introduce a uniquely stealthy and effective data poisoning entry point into the robot learning supply chain that can result in the deployment of unsafe or otherwise compromised robotic policies despite training on seemingly safe ground truth training data. In contrast to traditional data poisoning techniques which directly implant dangerous trajectories into sold or uploaded datasets, our novel attack methods inject malicious prompts or compromising transition dynamics into visibly safe teleoperated datasets which are only activated once fed through a world model as input. This can result in the generation of synthetic, dangerous robot training trajectories and subsequently unsafe or compromised robot policies. We demonstrate the effectiveness of our attacks against both state of the art action conditioned and text conditioned world models, showing a full end-to-end backdoor on a downstream DRL policy and a proof-of-concept for the VLA setting. Overall these findings necessitate research into more secure world models and reevaluating their position within the robot learning supply chain.

URL PDF HTML 收藏
2305.04718 2023-09-21 cs.RO cs.AI cs.CV 89%

The Treachery of Images: Bayesian Scene Keypoints for Deep Policy Learning in Robotic Manipulation

Jan Ole von Hartz, Eugenio Chisari, Tim Welschehold, Wolfram Burgard, Joschka Boedecker, Abhinav Valada

专题命中 机器人学习 :manipulation(title,abstract);robotic(title,abstract);分类 cs.RO、cs.AI、cs.CV;robotics(journal_ref)

Journal ref IEEE Robotics and Automation Letters, vol. 8, no. 11, pp. 6931-6938, Nov. 2023

详情
英文摘要

In policy learning for robotic manipulation, sample efficiency is of paramount importance. Thus, learning and extracting more compact representations from camera observations is a promising avenue. However, current methods often assume full observability of the scene and struggle with scale invariance. In many tasks and settings, this assumption does not hold as objects in the scene are often occluded or lie outside the field of view of the camera, rendering the camera observation ambiguous with regard to their location. To tackle this problem, we present BASK, a Bayesian approach to tracking scale-invariant keypoints over time. Our approach successfully resolves inherent ambiguities in images, enabling keypoint tracking on symmetrical objects and occluded and out-of-view objects. We employ our method to learn challenging multi-object robot manipulation tasks from wrist camera observations and demonstrate superior utility for policy learning compared to other representation learning techniques. Furthermore, we show outstanding robustness towards disturbances such as clutter, occlusions, and noisy depth measurements, as well as generalization to unseen objects both in simulation and real-world robotic experiments.

URL PDF HTML 收藏
2503.05231 2026-04-27 cs.RO cs.AI 89%

Kaiwu: A Multimodal Manipulation Dataset and Framework for Robot Learning and Human-Robot Interaction

Kaiwu:一种用于机器人学习和人机交互的多模态操控数据集和框架

Shuo Jiang, Haonan Li, Ruochen Ren, Yanmin Zhou, Zhipeng Wang, Bin He

专题命中 机器人学习 :robot learning(title,abstract);manipulation(title,abstract);robotics(comments,journal_ref);分类 cs.RO、cs.AI

AI总结 本文提出Kaiwu多模态数据集,解决复杂装配场景中缺失的真实同步多模态数据问题,通过20名受试者和30个交互对象,记录11,664个整合动作实例,支持机器人学习、精细操作、人类意图研究和人机协作。

Comments 8 pages, 5 figures, Submitted to IEEE Robotics and Automation Letters (RAL)

Journal ref IEEE Robotics and Automation Letters, vol. 10, no. 11, pp. 11482-11489, Nov. 2025

详情
AI中文摘要

本文提出Kaiwu多模态数据集,解决复杂装配场景中缺失的真实同步多模态数据问题,通过20名受试者和30个交互对象,记录11,664个整合动作实例,支持机器人学习、精细操作、人类意图研究和人机协作。

英文摘要

Cutting-edge robot learning techniques including foundation models and imitation learning from humans all pose huge demands on large-scale and high-quality datasets which constitute one of the bottleneck in the general intelligent robot fields. This paper presents the Kaiwu multimodal dataset to address the missing real-world synchronized multimodal data problems in the sophisticated assembling scenario,especially with dynamics information and its fine-grained labelling. The dataset first provides an integration of human,environment and robot data collection framework with 20 subjects and 30 interaction objects resulting in totally 11,664 instances of integrated actions. For each of the demonstration,hand motions,operation pressures,sounds of the assembling process,multi-view videos, high-precision motion capture information,eye gaze with first-person videos,electromyography signals are all recorded. Fine-grained multi-level annotation based on absolute timestamp,and semantic segmentation labelling are performed. Kaiwu dataset aims to facilitate robot learning,dexterous manipulation,human intention investigation and human-robot collaboration research.

URL PDF HTML 收藏
2412.04835 2024-12-09 cs.RO cs.AI cs.CV cs.LG 89%

Maximizing Alignment with Minimal Feedback: Efficiently Learning Rewards for Visuomotor Robot Policy Alignment

Ran Tian, Yilin Wu, Chenfeng Xu, Masayoshi Tomizuka, Jitendra Malik, Andrea Bajcsy

专题命中 机器人学习 :robot policy(title,abstract);robotics(abstract);manipulation(abstract);robotic(abstract)

Comments Submitted to IJRR, this paper is an extended journal version of the conference paper arXiv:2310.07932 with new results and discussion. arXiv admin note: substantial text overlap with arXiv:2310.07932

详情
英文摘要

Visuomotor robot policies, increasingly pre-trained on large-scale datasets, promise significant advancements across robotics domains. However, aligning these policies with end-user preferences remains a challenge, particularly when the preferences are hard to specify. While reinforcement learning from human feedback (RLHF) has become the predominant mechanism for alignment in non-embodied domains like large language models, it has not seen the same success in aligning visuomotor policies due to the prohibitive amount of human feedback required to learn visual reward functions. To address this limitation, we propose Representation-Aligned Preference-based Learning (RAPL), an observation-only method for learning visual rewards from significantly less human preference feedback. Unlike traditional RLHF, RAPL focuses human feedback on fine-tuning pre-trained vision encoders to align with the end-user's visual representation and then constructs a dense visual reward via feature matching in this aligned representation space. We first validate RAPL through simulation experiments in the X-Magical benchmark and Franka Panda robotic manipulation, demonstrating that it can learn rewards aligned with human preferences, more efficiently uses preference data, and generalizes across robot embodiments. Finally, our hardware experiments align pre-trained Diffusion Policies for three object manipulation tasks. We find that RAPL can fine-tune these policies with 5x less real human preference data, taking the first step towards minimizing human feedback while maximizing visuomotor robot policy alignment.

URL PDF HTML 收藏
1903.09589 2019-03-25 cs.RO 88%

A Fog Robotics Approach to Deep Robot Learning: Application to Object Recognition and Grasp Planning in Surface Decluttering

Ajay Kumar Tanwani, Nitesh Mor, John Kubiatowicz, Joseph E. Gonzalez, Ken Goldberg

专题命中 机器人学习 :robotics(title,abstract);robot learning(title,abstract);分类 cs.RO

Comments IEEE International Conference on Robotics and Automation, ICRA, 2019

详情
英文摘要

The growing demand of industrial, automotive and service robots presents a challenge to the centralized Cloud Robotics model in terms of privacy, security, latency, bandwidth, and reliability. In this paper, we present a `Fog Robotics' approach to deep robot learning that distributes compute, storage and networking resources between the Cloud and the Edge in a federated manner. Deep models are trained on non-private (public) synthetic images in the Cloud; the models are adapted to the private real images of the environment at the Edge within a trusted network and subsequently, deployed as a service for low-latency and secure inference/prediction for other robots in the network. We apply this approach to surface decluttering, where a mobile robot picks and sorts objects from a cluttered floor by learning a deep object recognition and a grasp planning model. Experiments suggest that Fog Robotics can improve performance by sim-to-real domain adaptation in comparison to exclusively using Cloud or Edge resources, while reducing the inference cycle time by 4\times to successfully declutter 86% of objects over 213 attempts.

URL PDF HTML 收藏
2601.07060 2026-04-07 cs.RO 88%

PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation

PALM:通过具身理由推理进行长时间 horizon 肢体操作的进展感知策略学习

Yuanzhe Liu, Jingyuan Zhu, Yuchen Mo, Gen Li, Xu Cao, Jin Jin, Yifan Shen, Zhengyuan Li, Tianjiao Yu, Wenzhen Yuan, Fangqiang Ding, Ismini Lourentzou

机构 * University of Pennsylvania(宾夕法尼亚大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Nanyang Technological University(南洋理工大学) University of Oxford(牛津大学) Massachusetts Institute of Technology(麻省理工学院)

专题命中 机器人学习 :manipulation(title,abstract);robotic(title,abstract);分类 cs.RO

AI总结 PALM通过具身理由推理和子任务进展线索,改进长horizon机器人操作的策略学习,实现91.8%的成功率和更长的任务长度。

Comments CVPR 2026

详情
AI中文摘要

最近在视觉-语言-动作(VLA)模型方面的进展在机器人操作中显示出潜力,但仍然在长horizon、多步骤任务上遇到困难。现有方法缺乏内部推理机制,无法识别任务相关交互线索或跟踪子任务中的进展,导致关键执行错误,如重复动作、遗漏步骤和提前终止。为了解决这些挑战,我们引入了PALM,一个VLA框架,围绕交互中心的具身理由推理和子任务进展线索结构化策略学习。PALM提炼互补的具身表示,捕捉物体相关性、接触几何、空间位置和运动动力学,并作为任务相关的锚点用于视觉-运动控制。为进一步稳定长horizon执行,PALM预测连续的子任务进展,实现无缝的子任务转换。在广泛的模拟和现实世界实验中,PALM一致优于基线,实现了在LIBERO-LONG上的91.8%成功率,在CALVIN ABC->D上的平均长度提高了12.5%,并在三个长horizon泛化设置中实现了两倍于现实世界基线的改进。

英文摘要

Recent advancements in vision-language-action (VLA) models have shown promise in robotic manipulation, yet they continue to struggle with long-horizon, multi-step tasks. Existing methods lack internal reasoning mechanisms that can identify task-relevant interaction cues or track progress within a subtask, leading to critical execution errors such as repeated actions, missed steps, and premature termination. To address these challenges, we introduce PALM, a VLA framework that structures policy learning around interaction-centric affordance reasoning and subtask progress cues. PALM distills complementary affordance representations that capture object relevance, contact geometry, spatial placements, and motion dynamics, and serve as task-relevant anchors for visuomotor control. To further stabilize long-horizon execution, PALM predicts continuous within-subtask progress, enabling seamless subtask transitions. Across extensive simulation and real-world experiments, PALM consistently outperforms baselines, achieving a 91.8% success rate on LIBERO-LONG, a 12.5% improvement in average length on CALVIN ABC->D, and a 2x improvement over real-world baselines across three long-horizon generalization settings.

URL PDF HTML 收藏
2603.01474 2026-03-09 cs.RO 88%

ROSER: Few-Shot Robotic Sequence Retrieval for Scalable Robot Learning

ROSER:用于可扩展机器人学习的少样本机器人序列检索

Zillur Rahman, Eddison Pham, Alejandro Daniel Noel, Cristian Meo

机构 * University of Nevada, Las Vegas(内华达大学拉斯维加斯分校) University of Toronto(多伦多大学) LatentWorlds AI

专题命中 机器人学习 :robot learning(title,abstract);robotic(title,abstract);分类 cs.RO

AI总结 ROSER通过少样本检索框架有效解决机器人序列数据稀缺问题,提升机器人学习的数据可用性。

Comments 2026 ICLR DATA-FM Workshop

详情
AI中文摘要

机器人学习中的关键瓶颈是任务标记、分段训练数据的稀缺性,尽管存在大量长连续交互日志记录的大规模机器人数据集。现有数据集包含大量多样化的行为,但结构上与需要干净分段、任务特定轨迹的现代学习框架不兼容。我们通过正式化机器人序列检索来解决这一数据利用危机:即使用仅几个参考示例从未标记的日志中提取可重用的、以任务为中心的片段。我们介绍了ROSER,一个轻量级的少样本检索框架,它在时间窗口上学习任务无关的度量空间,使得在仅需3-5个演示的情况下就能实现准确的检索,而无需任何任务特定的训练。为了验证我们的方法,我们建立了全面的评估协议,并在三个大规模数据集(如LIBERO、DROID和nuScenes)上将ROSER与经典对齐方法、学习嵌入和语言模型基线进行比较。我们的实验表明,ROSER在准确性和效率上均优于所有先前方法,实现每匹配亚毫秒的推理时间,同时保持优越的分布对齐。通过将数据整理重新框架化为少样本检索,ROSER为解锁未充分利用的机器人数据集提供了实用路径,从根本上提高了机器人学习的数据可用性。

英文摘要

A critical bottleneck in robot learning is the scarcity of task-labeled, segmented training data, despite the abundance of large-scale robotic datasets recorded as long, continuous interaction logs. Existing datasets contain vast amounts of diverse behaviors, yet remain structurally incompatible with modern learning frameworks that require cleanly segmented, task-specific trajectories. We address this data utilization crisis by formalizing robotic sequence retrieval: the task of extracting reusable, task-centric segments from unlabeled logs using only a few reference examples. We introduce ROSER, a lightweight few-shot retrieval framework that learns task-agnostic metric spaces over temporal windows, enabling accurate retrieval with as few as 3-5 demonstrations, without any task-specific training required. To validate our approach, we establish comprehensive evaluation protocols and benchmark ROSER against classical alignment methods, learned embeddings, and language model baselines across three large-scale datasets (e.g., LIBERO, DROID, and nuScenes). Our experiments demonstrate that ROSER consistently outperforms all prior methods in both accuracy and efficiency, achieving sub-millisecond per-match inference while maintaining superior distributional alignment. By reframing data curation as few-shot retrieval, ROSER provides a practical pathway to unlock underutilized robotic datasets, fundamentally improving data availability for robot learning.

URL PDF HTML 收藏
2507.10543 2025-12-04 cs.RO 88%

MP1: MeanFlow Tames Policy Learning in 1-step for Robotic Manipulation

MP1: 均值流使政策学习在1步中适用于机器人操作

Juyi Sheng, Ziyi Wang, Peiming Li, Mengyuan Liu

专题命中 机器人学习 :manipulation(title,abstract);robotic(title);robot learning(abstract);分类 cs.RO

AI总结 MP1通过MeanFlow范式和轻量级分散损失,在1-NFE中实现更精确且可控的动作轨迹生成,优于现有方法。

Comments This paper has been accepted by AAAI 2026

详情
AI中文摘要

在机器人操作中,机器人学习已成为主流方法。然而,该领域内的生成模型面临着扩散模型缓慢迭代采样的根本权衡,以及基于流的方法的架构约束,后者通常依赖于显式的一致性损失。为了解决这些限制,我们引入了MP1,它将3D点云输入与MeanFlow范式结合,以在单个网络函数评估(1-NFE)中生成动作轨迹。通过直接学习通过“MeanFlow身份”获得的区间平均速度,我们的策略避免了任何额外的一致性约束。这种形式化消除了推理过程中的数值微分方程求解器误差,从而产生更精确的轨迹。MP1进一步结合CFG以提高轨迹可控性,同时保持1-NFE推理而不重新引入结构约束。由于细微的场景上下文变化对机器人学习至关重要,尤其是在少样本学习中,我们引入了一种轻量级的分散损失,它在训练期间排斥状态嵌入,从而提升泛化能力而不影响推理速度。我们在Adroit和Meta-World基准以及现实世界场景中验证了我们的方法。实验结果表明,MP1在平均任务成功率方面表现优异,比DP3高出10.2%,比FlowPolicy高出7.3%。其平均推理时间仅为DP3的6.8毫秒-19倍快,几乎比FlowPolicy快2倍。我们的项目页面可在https://mp1-2254.github.io/上获得,代码可在https://github.com/LogSSim/MP1上获取。

英文摘要

In robot manipulation, robot learning has become a prevailing approach. However, generative models within this field face a fundamental trade-off between the slow, iterative sampling of diffusion models and the architectural constraints of faster Flow-based methods, which often rely on explicit consistency losses. To address these limitations, we introduce MP1, which pairs 3D point-cloud inputs with the MeanFlow paradigm to generate action trajectories in one network function evaluation (1-NFE). By directly learning the interval-averaged velocity via the "MeanFlow Identity", our policy avoids any additional consistency constraints. This formulation eliminates numerical ODE-solver errors during inference, yielding more precise trajectories. MP1 further incorporates CFG for improved trajectory controllability while retaining 1-NFE inference without reintroducing structural constraints. Because subtle scene-context variations are critical for robot learning, especially in few-shot learning, we introduce a lightweight Dispersive Loss that repels state embeddings during training, boosting generalization without slowing inference. We validate our method on the Adroit and Meta-World benchmarks, as well as in real-world scenarios. Experimental results show MP1 achieves superior average task success rates, outperforming DP3 by 10.2% and FlowPolicy by 7.3%. Its average inference time is only 6.8 ms-19x faster than DP3 and nearly 2x faster than FlowPolicy. Our project page is available at https://mp1-2254.github.io/, and the code can be accessed at https://github.com/LogSSim/MP1.

URL PDF HTML 收藏
2404.18201 2025-11-11 cs.RO 88%

What Foundation Models can Bring for Robot Learning in Manipulation : A Survey

Dingzhe Li, Yixiang Jin, Yuhao Sun, Yong A, Hongze Yu, Jun Shi, Xiaoshuai Hao, Peng Hao, Huaping Liu, Xiang Li, Xinde Li, Fuchun Sun, Jianwei Zhang, Bin Fang

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Tsinghua University(清华大学) Southeast University(东南大学) Universität Hamburg(汉堡大学)

专题命中 机器人学习 :robot learning(title,abstract);manipulation(title,abstract);分类 cs.RO

详情
英文摘要

The realization of universal robots is an ultimate goal of researchers. However, a key hurdle in achieving this goal lies in the robots' ability to manipulate objects in their unstructured surrounding environments according to different tasks. The learning-based approach is considered an effective way to address generalization. The impressive performance of foundation models in the fields of computer vision and natural language suggests the potential of embedding foundation models into manipulation tasks as a viable path toward achieving general manipulation capability. However, we believe achieving general manipulation capability requires an overarching framework akin to auto driving. This framework should encompass multiple functional modules, with different foundation models assuming distinct roles in facilitating general manipulation capability. This survey focuses on the contributions of foundation models to robot learning for manipulation. We propose a comprehensive framework and detail how foundation models can address challenges in each module of the framework. What's more, we examine current approaches, outline challenges, suggest future research directions, and identify potential risks associated with integrating foundation models into this domain.

URL PDF HTML 收藏
2312.05323 2025-08-18 cs.RO 88%

BaRiFlex: A Robotic Gripper with Versatility and Collision Robustness for Robot Learning

Gu-Cheol Jeong, Arpit Bahety, Gabriel Pedraza, Ashish D. Deshpande, Roberto Martín-Martín

机构 * Department of Mechanical Engineering, The University of Texas at Austin(机械工程系,德克萨斯大学奥斯汀分校) Department of Computer Science, The University of Texas at Austin(计算机科学系,德克萨斯大学奥斯汀分校)

专题命中 机器人学习 :robot learning(title,abstract);robotic(title,abstract);分类 cs.RO

Comments 8 pages, 6 figures, project website: https://robin-lab.cs.utexas.edu/bariflex/

详情
英文摘要

We present a new approach to robot hand design specifically suited for successfully implementing robot learning methods to accomplish tasks in daily human environments. We introduce BaRiFlex, an innovative gripper design that alleviates the issues caused by unexpected contact and collisions during robot learning, offering robustness, grasping versatility, task versatility, and simplicity to the learning processes. This achievement is enabled by the incorporation of low-inertia actuators, providing high Back-drivability, and the strategic combination of Rigid and Flexible materials which enhances versatility and the gripper's resilience against unpredicted collisions. Furthermore, the integration of flexible Fin-Ray linkages and rigid linkages allows the gripper to execute compliant grasping and precise pinching. We conducted rigorous performance tests to characterize the novel gripper's compliance, durability, grasping and task versatility, and precision. We also integrated the BaRiFlex with a 7 Degree of Freedom (DoF) Franka Emika's Panda robotic arm to evaluate its capacity to support a trial-and-error (reinforcement learning) training procedure. The results of our experimental study are then compared to those obtained using the original rigid Franka Hand and a reference Fin-Ray soft gripper, demonstrating the superior capabilities and advantages of our developed gripper system.

URL PDF HTML 收藏
2607.29302 2026-08-03 cs.RO cs.CV 新提交 88%

BWM: A Low-Cost High-Fidelity World Simulator for Robot Learning

BWM:面向机器人学习的低成本高保真世界模拟器

BWM Team

专题命中 机器人学习 :robot learning(title,abstract);world model(abstract,abstract_cn);manipulation(abstract);分类 cs.RO、cs.CV

AI总结 本文提出面向机器人操纵的开源世界模拟器BWM,通过动作条件化自回归预测实现高保真,兼具数据引擎与策略评估器功能,在WorldArena挑战赛中表现最优并开源相关资源。

详情
AI中文摘要

可靠的机器人学习需要一种世界模拟器,能在物理硬件执行前预测动作后果,包括高风险和易失败的结果。现有物理模拟器需要大量资产构建与校准,仍存在现实差距;视频生成器往往无法精确控制对机器人细粒度动作的响应。本文提出Boundless World Model(BWM),一种面向机器人操纵的开源、低成本、高保真世界模拟器。BWM是一种动作条件化的世界模型,结合初始环境引导、动态视觉历史以及时间对齐的机器人动作条件,对未来观测进行有状态自回归预测。我们通过轨迹回放、重叠片段采样和初始观测增强构建动作对齐的训练片段。BWM兼具数据引擎和策略评估器的功能:作为数据引擎,它用动作对齐的rollout扩充模仿学习数据;作为策略评估器,它用于闭环评估、风险预判和策略排序。在WorldArena基准测试和物理机器人上的实验表明,BWM在数据引擎和策略评估器设置下,模拟器保真度和功能效用均有所提升,在WorldArena挑战赛的Track 1及两个Track 2应用中均排名第一。我们发布了BWM开源生态系统,包括模型检查点、训练与推理代码,以及数据生成和策略评估接口。

英文摘要

Reliable robot learning requires a world simulator that can predict action consequences before execution on physical hardware, including risky and failure-prone outcomes. Existing physics simulators require substantial asset construction and calibration and still face a sim-to-real gap, while video generators often lack precise control over their responses to fine-grained robot actions. In this paper, we present the Boundless World Model (BWM), an open-source, low-cost, high-fidelity world simulator for robot manipulation. BWM is an action-conditioned world model that combines initial-environment guidance, dynamic visual history, and temporally aligned robot-action conditioning for stateful autoregressive prediction of future observations. We construct action-aligned training clips through trajectory replay, overlapping clip sampling, and initial-observation enhancement. BWM serves as a data engine that augments imitation-learning data with action-aligned rollouts, and as a policy evaluator for closed-loop assessment, risk anticipation, and policy ranking. Experiments on the WorldArena benchmark and physical robots demonstrate improved simulator fidelity and functional utility across the data-engine and policy-evaluator settings. BWM ranks first overall in the WorldArena Challenge across Track 1 and its two Track 2 applications. We release the BWM open-source ecosystem, including model checkpoints, training and inference code, and interfaces for data generation and policy evaluation.

URL PDF HTML 收藏
2606.10371 2026-06-10 cs.RO cs.AI 新提交 88%

Test-time Adversarial Takeover: A Real-time Hijacking Interface against Robotic Diffusion Policies

测试时对抗接管:针对机器人扩散策略的实时劫持接口

Zi Yin, Peilin Chai, Siyuan Huang, Zhanhao Hu

机构 * Tsinghua University(清华大学) Independent Researcher(独立研究员) Johns Hopkins University(约翰霍普金斯大学) UC Berkeley(加州大学伯克利分校)

专题命中 机器人学习 :robotic(title);embodied AI(abstract);manipulation(abstract);navigation(abstract)

AI总结 提出测试时对抗接管(TAKO)方法,通过可微扩散推理学习可重复使用的通用补丁,在测试时切换补丁以劫持机器人策略,实现远程操控,在多种任务和模型上达到100%接管成功率。

详情
AI中文摘要

基于扩散的动作生成已成为具身AI的基础组件,但其对视觉条件的依赖使得部署的视觉运动策略容易受到对抗性操纵。大多数先前的攻击侧重于破坏:它们扰动观测流以降低任务成功率或引发异常行为。我们研究了一种更强的威胁,即测试时对抗接管(TAKO),其中攻击者获得对冻结机器人策略的实时转向接口,并将其转变为远程操控仪器。TAKO通过可微扩散推理学习一个小的可重用通用补丁词汇表;在测试时,攻击者在摄像头流中切换这些补丁以组合攻击者选择的轨迹。这种方法之所以有效,是因为扰动作用于视觉条件路径,其中诱导的偏差可以通过迭代生成推理持续存在。我们进一步表明,自然的目标基线——目标策略匹配——会失败,因为受害者策略无法可靠地在分布外目标偏移上监督自身。在四个任务(2D操作、模拟空中递送、模拟地面导航和物理世界地面导航)、两个视觉编码器(ResNet-18和EfficientNet-B0 + Transformer)以及三个生成推理族(DDPM、DDIM和流匹配)中,人类操作员在每个评估设置中均实现了100%的接管成功率,满足攻击者定义的目标。项目页面可在此https URL获取。

英文摘要

Diffusion-based action generation has become a foundational component of embodied AI, but its reliance on visual conditioning leaves deployed visuomotor policies vulnerable to adversarial manipulation. Most prior attacks focus on disruption: they perturb the observation stream to reduce task success or induce erratic behavior. We study a stronger threat, Test-time Adversarial Takeover (TAKO), in which an attacker obtains a real-time steering interface over a frozen robot policy and turns it into a remotely piloted instrument. TAKO learns a small vocabulary of reusable universal patches through differentiable diffusion inference; at test time, the attacker switches among these patches in the camera stream to compose attacker-chosen trajectories. This works because the perturbation acts on the visual conditioning pathway, where the induced bias can persist through iterative generative inference. We further show that the natural targeted baseline, target-policy matching, fails because the victim policy cannot reliably supervise itself on out-of-distribution target shifts. Across four tasks (2D manipulation, simulated aerial delivery, simulated ground navigation, and physical-world ground navigation), two visual encoders (ResNet-18 and EfficientNet-B0 + Transformer), and three generative inference families (DDPM, DDIM, and flow matching), human operators achieve 100\% takeover success on attacker-defined objectives in every evaluated setting. The project page is available at https://tako-attack.github.io.

URL PDF HTML 收藏
2604.27621 2026-05-01 cs.RO cs.CV 88%

Robot Learning from Human Videos: A Survey

从人类视频学习机器人:综述

Junyi Ma, Erhang Zhang, Haoran Yang, Ditao Li, Chenyang Xu, Guangming Wang, Hesheng Wang

机构 * Shanghai Jiao Tong University(上海交通大学) University of Cambridge(剑桥大学)

专题命中 机器人学习 :robot learning(title);robotics(abstract);embodied AI(abstract);manipulation(abstract)

AI总结 本文综述了通过人类视频学习机器人技能的方法,探讨了政策学习基础、人类视频接口、技能转移层次分类及数据基础,分析了挑战与未来研究方向。

Comments Paper list: https://github.com/IRMVLab/awesome-robot-learning-from-human-videos

详情
AI中文摘要

从人类视频学习机器人领域的一个关键瓶颈是机器人数据的扩展问题。为解决此问题,近年来该领域迅速受到关注,得益于人类活动视频的丰富性和计算机视觉的进步。本文综述了人类视频基于学习技术在机器人中的应用,重点探讨了人类-机器人技能转移和数据基础。首先回顾了机器人中的政策学习基础,然后描述了将人类视频整合到系统中的基本接口。随后介绍了一种层次化的分类方法,涵盖任务、观察和动作导向的路径,以及不同数据配置和学习范式之间的耦合分析。此外,本文还探讨了数据基础,包括广泛使用的视频数据集和视频生成方案,并提供了大规模的统计趋势。最后,强调了该领域固有的挑战和限制,并指出了未来研究的潜在方向。本文的综述列表可在https://github.com/IRMVLab/awesome-robot-learning-from-human-videos上找到。

英文摘要

A critical bottleneck hindering further advancement in embodied AI and robotics is the challenge of scaling robot data. To address this, the field of learning robot manipulation skills from human video data has attracted rapidly growing attention in recent years, driven by the abundance of human activity videos and advances in computer vision. This line of research promises to enable robots to acquire skills passively from the vast and readily available resource of human demonstrations, substantially favoring scalable learning for generalist robotic systems. Therefore, we present this survey to provide a comprehensive and up-to-date review of human-video-based learning techniques in robotics, focusing on both human-robot skill transfer and data foundations. We first review the policy learning foundations in robotics, and then describe the fundamental interfaces to incorporate human videos. Subsequently, we introduce a hierarchical taxonomy of transferring human videos to robot skills, covering task-, observation-, and action-oriented pathways, along with a cross-family analysis of their couplings with different data configurations and learning paradigms. In addition, we investigate the data foundations including widely-used human video datasets and video generation schemes, and provide large-scale statistical trends in dataset development and utilization. Ultimately, we emphasize the challenges and limitations intrinsic to this field, and delineate potential avenues for future research. The paper list of our survey is available at https://github.com/IRMVLab/awesome-robot-learning-from-human-videos.

URL PDF HTML 收藏
2507.09117 2025-07-15 cs.RO cs.AI 88%

Towards Human-level Dexterity via Robot Learning

Gagan Khandate

专题命中 机器人学习 :robot learning(title,abstract);robotics(abstract);manipulation(abstract);robotic(abstract)

Comments PhD thesis

详情
英文摘要

Dexterous intelligence -- the ability to perform complex interactions with multi-fingered hands -- is a pinnacle of human physical intelligence and emergent higher-order cognitive skills. However, contrary to Moravec's paradox, dexterous intelligence in humans appears simple only superficially. Many million years were spent co-evolving the human brain and hands including rich tactile sensing. Achieving human-level dexterity with robotic hands has long been a fundamental goal in robotics and represents a critical milestone toward general embodied intelligence. In this pursuit, computational sensorimotor learning has made significant progress, enabling feats such as arbitrary in-hand object reorientation. However, we observe that achieving higher levels of dexterity requires overcoming very fundamental limitations of computational sensorimotor learning. I develop robot learning methods for highly dexterous multi-fingered manipulation by directly addressing these limitations at their root cause. Chiefly, through key studies, this disseration progressively builds an effective framework for reinforcement learning of dexterous multi-fingered manipulation skills. These methods adopt structured exploration, effectively overcoming the limitations of random exploration in reinforcement learning. The insights gained culminate in a highly effective reinforcement learning that incorporates sampling-based planning for direct exploration. Additionally, this thesis explores a new paradigm of using visuo-tactile human demonstrations for dexterity, introducing corresponding imitation learning techniques.

URL PDF HTML 收藏
2409.12061 2024-09-19 cs.RO cs.AI 88%

Generalized Robot Learning Framework

Jiahuan Yan, Zhouyang Hong, Yu Zhao, Yu Tian, Yunxin Liu, Travis Davies, Luhui Hu

专题命中 机器人学习 :robot learning(title,abstract);robotics(abstract);manipulation(abstract);robotic(abstract)

Comments 6 pages, 2 figures. cs.RO

详情
英文摘要

Imitation based robot learning has recently gained significant attention in the robotics field due to its theoretical potential for transferability and generalizability. However, it remains notoriously costly, both in terms of hardware and data collection, and deploying it in real-world environments demands meticulous setup of robots and precise experimental conditions. In this paper, we present a low-cost robot learning framework that is both easily reproducible and transferable to various robots and environments. We demonstrate that deployable imitation learning can be successfully applied even to industrial-grade robots, not just expensive collaborative robotic arms. Furthermore, our results show that multi-task robot learning is achievable with simple network architectures and fewer demonstrations than previously thought necessary. As the current evaluating method is almost subjective when it comes to real-world manipulation tasks, we propose Voting Positive Rate (VPR) - a novel evaluation strategy that provides a more objective assessment of performance. We conduct an extensive comparison of success rates across various self-designed tasks to validate our approach. To foster collaboration and support the robot learning community, we have open-sourced all relevant datasets and model checkpoints, available at huggingface.co/ZhiChengAI.

URL PDF HTML 收藏
2505.20795 2026-06-01 cs.RO 88%

Learning Generalizable Robot Policy with Human Demonstration Video as a Prompt

以人类演示视频为提示学习可泛化的机器人策略

Xiang Zhu, Yichen Liu, Hezhong Li, Jianyu Chen

机构 * Tsinghua University, China(清华大学,中国) Shanghai Qi Zhi Institute, China(上海启智研究院,中国)

专题命中 机器人学习 :robot policy(title,abstract);robot learning(abstract);manipulation(abstract);robotic(abstract)

AI总结 提出两阶段框架,利用人类演示视频学习可泛化机器人策略,无需遥操作数据或微调即可执行新任务。

Comments Accepted to the IEEE International Conference on Robotics and Automation (ICRA), 2026

详情
AI中文摘要

最近的机器人学习方法通常依赖于通过遥操作收集的大规模机器人数据集的模仿学习。面对新任务时,这些方法通常需要收集一组新的遥操作数据并微调策略。此外,遥操作数据收集流程也繁琐且昂贵。相反,人类能够通过观察他人操作高效学习新任务。在本文中,我们介绍了一种新颖的两阶段框架,利用人类演示学习可泛化的机器人策略。该策略可以直接以人类演示视频为提示,执行新任务,无需任何新的遥操作数据和模型微调。在第一阶段,我们训练视频生成模型,通过交叉预测捕获人类和机器人演示视频数据的联合表示。在第二阶段,我们使用新颖的原型对比损失将学习到的表示与人类和机器人之间的共享动作空间融合。在真实世界灵巧操作任务上的实证评估显示了所提出方法的有效性和泛化能力。

英文摘要

Recent robot learning methods commonly rely on imitation learning from massive robotic dataset collected with teleoperation. When facing a new task, such methods generally require collecting a set of new teleoperation data and finetuning the policy. Furthermore, the teleoperation data collection pipeline is also tedious and expensive. Instead, human is able to efficiently learn new tasks by just watching others do. In this paper, we introduce a novel two-stage framework that utilizes human demonstrations to learn a generalizable robot policy. Such policy can directly take human demonstration video as a prompt and perform new tasks without any new teleoperation data and model finetuning at all. In the first stage, we train video generation model that captures a joint representation for both the human and robot demonstration video data using cross-prediction. In the second stage, we fuse the learned representation with a shared action space between human and robot using a novel prototypical contrastive loss. Empirical evaluations on real-world dexterous manipulation tasks show the effectiveness and generalization capabilities of our proposed method.

URL PDF HTML 收藏
2604.25788 2026-04-29 cs.RO 88%

KinDER: A Physical Reasoning Benchmark for Robot Learning and Planning

KinDER:机器人学习与规划的运动学与动力学推理基准

Yixuan Huang, Bowen Li, Vaibhav Saxena, Yichao Liang, Utkarsh Aashu Mishra, Liang Ji, Lihan Zha, Jimmy Wu, Nishanth Kumar, Sebastian Scherer, Danfei Xu, Tom Silver

机构 * Princeton University(普林斯顿大学) Carnegie Mellon University(卡内基梅隆大学) Georgia Tech(佐治亚理工学院) University of Cambridge(剑桥大学) NVIDIA(英伟达) MIT(麻省理工学院)

专题命中 机器人学习 :robot learning(title,abstract);robotics(abstract,comments);manipulation(abstract);robotic(abstract)

AI总结 KinDER基准针对机器人学习与规划中的物理推理挑战,包含25个生成环境、Python库和评估套件,旨在通过系统比较不同方法提升机器人物理推理能力。

Comments Project website: https://prpl-group.com/kinder-site/. 21 pages, 8 figures. Accepted to Robotics Science and Systems (RSS), 2026

详情
AI中文摘要

与物理世界交互的机器人系统必须考虑自身躯体、环境和任务所施加的运动学和动力学约束。我们介绍了KinDER,一个针对机器人学习和规划中物理推理挑战的基准。KinDER包含25个程序生成的环境、一个兼容Gymnasium的Python库以及标准化的评估套件,包含13种基线方法。环境设计旨在隔离五个核心物理推理挑战:基本空间关系、非抓取多物体操作、工具使用、组合几何约束和动态约束,这些挑战与感知、语言理解和应用特定复杂性分离。实证评估显示现有方法难以解决许多环境,表明当前物理推理方法存在显著差距。我们还包含在移动机械臂上的真实到仿真到真实实验,以评估仿真与现实物理交互的一致性。KinDER完全开源,旨在通过跨不同范式系统比较,推动机器人物理推理的发展。网站和代码:https://prpl-group.com/kinder-site/

英文摘要

Robotic systems that interact with the physical world must reason about kinematic and dynamic constraints imposed by their own embodiment, their environment, and the task at hand. We introduce KinDER, a benchmark for Kinematic and Dynamic Embodied Reasoning that targets physical reasoning challenges arising in robot learning and planning. KinDER comprises 25 procedurally generated environments, a Gymnasium-compatible Python library with parameterized skills and demonstrations, and a standardized evaluation suite with 13 implemented baselines spanning task and motion planning, imitation learning, reinforcement learning, and foundation-model-based approaches. The environments are designed to isolate five core physical reasoning challenges: basic spatial relations, nonprehensile multi-object manipulation, tool use, combinatorial geometric constraints, and dynamic constraints, disentangled from perception, language understanding, and application-specific complexity. Empirical evaluation shows that existing methods struggle to solve many of the environments, indicating substantial gaps in current approaches to physical reasoning. We additionally include real-to-sim-to-real experiments on a mobile manipulator to assess the correspondence between simulation and real-world physical interaction. KinDER is fully open-sourced and intended to enable systematic comparison across diverse paradigms for advancing physical reasoning in robotics. Website and code: https://prpl-group.com/kinder-site/

URL PDF HTML 收藏
2411.16959 2025-05-21 cs.RO cs.AI cs.CV cs.LG 88%

RoCoDA: Counterfactual Data Augmentation for Data-Efficient Robot Learning from Demonstrations

Ezra Ameperosa, Jeremy A. Collins, Mrinal Jain, Animesh Garg

机构 * Georgia Institute of Technology(佐治亚理工学院) Apptronik, Inc.(Apptronik公司)

专题命中 机器人学习 :robot learning(title);robotics(abstract,comments);manipulation(abstract);robotic(abstract)

Comments Accepted to 2025 IEEE International Conference on Robotics and Automation (ICRA)

详情
英文摘要

Imitation learning in robotics faces significant challenges in generalization due to the complexity of robotic environments and the high cost of data collection. We introduce RoCoDA, a novel method that unifies the concepts of invariance, equivariance, and causality within a single framework to enhance data augmentation for imitation learning. RoCoDA leverages causal invariance by modifying task-irrelevant subsets of the environment state without affecting the policy's output. Simultaneously, we exploit SE(3) equivariance by applying rigid body transformations to object poses and adjusting corresponding actions to generate synthetic demonstrations. We validate RoCoDA through extensive experiments on five robotic manipulation tasks, demonstrating improvements in policy performance, generalization, and sample efficiency compared to state-of-the-art data augmentation methods. Our policies exhibit robust generalization to unseen object poses, textures, and the presence of distractors. Furthermore, we observe emergent behavior such as re-grasping, indicating policies trained with RoCoDA possess a deeper understanding of task dynamics. By leveraging invariance, equivariance, and causality, RoCoDA provides a principled approach to data augmentation in imitation learning, bridging the gap between geometric symmetries and causal reasoning. Project Page: https://rocoda.github.io

URL PDF HTML 收藏
2310.08864 2025-05-15 cs.RO 88%

Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Embodiment Collaboration, Abby O'Neill, Abdul Rehman, Abhinav Gupta, Abhiram Maddukuri, Abhishek Gupta, Abhishek Padalkar, Abraham Lee, Acorn Pooley, Agrim Gupta, Ajay Mandlekar, Ajinkya Jain, Albert Tung, Alex Bewley, Alex Herzog, Alex Irpan, Alexander Khazatsky, Anant Rai, Anchit Gupta, Andrew Wang, Andrey Kolobov, Anikait Singh, Animesh Garg, Aniruddha Kembhavi, Annie Xie, Anthony Brohan, Antonin Raffin, Archit Sharma, Arefeh Yavary, Arhan Jain, Ashwin Balakrishna, Ayzaan Wahid, Ben Burgess-Limerick, Beomjoon Kim, Bernhard Schölkopf, Blake Wulfe, Brian Ichter, Cewu Lu, Charles Xu, Charlotte Le, Chelsea Finn, Chen Wang, Chenfeng Xu, Cheng Chi, Chenguang Huang, Christine Chan, Christopher Agia, Chuer Pan, Chuyuan Fu, Coline Devin, Danfei Xu, Daniel Morton, Danny Driess, Daphne Chen, Deepak Pathak, Dhruv Shah, Dieter Büchler, Dinesh Jayaraman, Dmitry Kalashnikov, Dorsa Sadigh, Edward Johns, Ethan Foster, Fangchen Liu, Federico Ceola, Fei Xia, Feiyu Zhao, Felipe Vieira Frujeri, Freek Stulp, Gaoyue Zhou, Gaurav S. Sukhatme, Gautam Salhotra, Ge Yan, Gilbert Feng, Giulio Schiavi, Glen Berseth, Gregory Kahn, Guangwen Yang, Guanzhi Wang, Hao Su, Hao-Shu Fang, Haochen Shi, Henghui Bao, Heni Ben Amor, Henrik I Christensen, Hiroki Furuta, Homanga Bharadhwaj, Homer Walke, Hongjie Fang, Huy Ha, Igor Mordatch, Ilija Radosavovic, Isabel Leal, Jacky Liang, Jad Abou-Chakra, Jaehyung Kim, Jaimyn Drake, Jan Peters, Jan Schneider, Jasmine Hsu, Jay Vakil, Jeannette Bohg, Jeffrey Bingham, Jeffrey Wu, Jensen Gao, Jiaheng Hu, Jiajun Wu, Jialin Wu, Jiankai Sun, Jianlan Luo, Jiayuan Gu, Jie Tan, Jihoon Oh, Jimmy Wu, Jingpei Lu, Jingyun Yang, Jitendra Malik, João Silvério, Joey Hejna, Jonathan Booher, Jonathan Tompson, Jonathan Yang, Jordi Salvador, Joseph J. Lim, Junhyek Han, Kaiyuan Wang, Kanishka Rao, Karl Pertsch, Karol Hausman, Keegan Go, Keerthana Gopalakrishnan, Ken Goldberg, Kendra Byrne, Kenneth Oslund, Kento Kawaharazuka, Kevin Black, Kevin Lin, Kevin Zhang, Kiana Ehsani, Kiran Lekkala, Kirsty Ellis, Krishan Rana, Krishnan Srinivasan, Kuan Fang, Kunal Pratap Singh, Kuo-Hao Zeng, Kyle Hatch, Kyle Hsu, Laurent Itti, Lawrence Yunliang Chen, Lerrel Pinto, Li Fei-Fei, Liam Tan, Linxi "Jim" Fan, Lionel Ott, Lisa Lee, Luca Weihs, Magnum Chen, Marion Lepert, Marius Memmel, Masayoshi Tomizuka, Masha Itkina, Mateo Guaman Castro, Max Spero, Maximilian Du, Michael Ahn, Michael C. Yip, Mingtong Zhang, Mingyu Ding, Minho Heo, Mohan Kumar Srirama, Mohit Sharma, Moo Jin Kim, Muhammad Zubair Irshad, Naoaki Kanazawa, Nicklas Hansen, Nicolas Heess, Nikhil J Joshi, Niko Suenderhauf, Ning Liu, Norman Di Palo, Nur Muhammad Mahi Shafiullah, Oier Mees, Oliver Kroemer, Osbert Bastani, Pannag R Sanketi, Patrick "Tree" Miller, Patrick Yin, Paul Wohlhart, Peng Xu, Peter David Fagan, Peter Mitrano, Pierre Sermanet, Pieter Abbeel, Priya Sundaresan, Qiuyu Chen, Quan Vuong, Rafael Rafailov, Ran Tian, Ria Doshi, Roberto Martín-Martín, Rohan Baijal, Rosario Scalise, Rose Hendrix, Roy Lin, Runjia Qian, Ruohan Zhang, Russell Mendonca, Rutav Shah, Ryan Hoque, Ryan Julian, Samuel Bustamante, Sean Kirmani, Sergey Levine, Shan Lin, Sherry Moore, Shikhar Bahl, Shivin Dass, Shubham Sonawani, Shubham Tulsiani, Shuran Song, Sichun Xu, Siddhant Haldar, Siddharth Karamcheti, Simeon Adebola, Simon Guist, Soroush Nasiriany, Stefan Schaal, Stefan Welker, Stephen Tian, Subramanian Ramamoorthy, Sudeep Dasari, Suneel Belkhale, Sungjae Park, Suraj Nair, Suvir Mirchandani, Takayuki Osa, Tanmay Gupta, Tatsuya Harada, Tatsuya Matsushima, Ted Xiao, Thomas Kollar, Tianhe Yu, Tianli Ding, Todor Davchev, Tony Z. Zhao, Travis Armstrong, Trevor Darrell, Trinity Chung, Vidhi Jain, Vikash Kumar, Vincent Vanhoucke, Vitor Guizilini, Wei Zhan, Wenxuan Zhou, Wolfram Burgard, Xi Chen, Xiangyu Chen, Xiaolong Wang, Xinghao Zhu, Xinyang Geng, Xiyuan Liu, Xu Liangwei, Xuanlin Li, Yansong Pang, Yao Lu, Yecheng Jason Ma, Yejin Kim, Yevgen Chebotar, Yifan Zhou, Yifeng Zhu, Yilin Wu, Ying Xu, Yixuan Wang, Yonatan Bisk, Yongqiang Dou, Yoonyoung Cho, Youngwoon Lee, Yuchen Cui, Yue Cao, Yueh-Hua Wu, Yujin Tang, Yuke Zhu, Yunchu Zhang, Yunfan Jiang, Yunshuang Li, Yunzhu Li, Yusuke Iwasawa, Yutaka Matsuo, Zehan Ma, Zhuo Xu, Zichen Jeff Cui, Zichen Zhang, Zipeng Fu, Zipeng Lin

专题命中 机器人学习 :robotic(title,abstract);robotics(abstract,comments);manipulation(abstract);robot policy(abstract)

Comments Project website: https://robotics-transformer-x.github.io

详情
英文摘要

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for many applications. Can such a consolidation happen in robotics? Conventionally, robotic learning methods train a separate model for every application, every robot, and even every environment. Can we instead train generalist X-robot policy that can be adapted efficiently to new robots, tasks, and environments? In this paper, we provide datasets in standardized data formats and models to make it possible to explore this possibility in the context of robotic manipulation, alongside experimental results that provide an example of effective X-robot policies. We assemble a dataset from 22 different robots collected through a collaboration between 21 institutions, demonstrating 527 skills (160266 tasks). We show that a high-capacity model trained on this data, which we call RT-X, exhibits positive transfer and improves the capabilities of multiple robots by leveraging experience from other platforms. More details can be found on the project website https://robotics-transformer-x.github.io.

URL PDF HTML 收藏
2405.03164 2024-05-07 cs.RO cs.AI cs.CV 88%

The Role of Predictive Uncertainty and Diversity in Embodied AI and Robot Learning

Ransalu Senanayake

专题命中 机器人学习 :robot learning(title);embodied AI(title);robotics(abstract);分类 cs.RO、cs.AI、cs.CV

详情
英文摘要

Uncertainty has long been a critical area of study in robotics, particularly when robots are equipped with analytical models. As we move towards the widespread use of deep neural networks in robots, which have demonstrated remarkable performance in research settings, understanding the nuances of uncertainty becomes crucial for their real-world deployment. This guide offers an overview of the importance of uncertainty and provides methods to quantify and evaluate it from an applications perspective.

URL PDF HTML 收藏
2209.15428 2026-07-22 cs.RO 87%

PyPose: A Library for Robot Learning with Physics-based Optimization

PyPose:一个结合物理优化的机器人学习库

Chen Wang, Dasong Gao, Kuan Xu, Junyi Geng, Yaoyu Hu, Yuheng Qiu, Bowen Li, Fan Yang, Brady Moon, Abhinav Pandey, Aryan, Jiahe Xu, Tianhao Wu, Haonan He, Daning Huang, Zhongqiang Ren, Shibo Zhao, Taimeng Fu, Pranay Reddy, Xiao Lin, Wenshan Wang, Jingnan Shi, Rajat Talak, Kun Cao, Yi Du, Han Wang, Huai Yu, Shanzhao Wang, Siyu Chen, Ananth Kashyap, Rohan Bandaru, Karthik Dantu, Jiajun Wu, Lihua Xie, Luca Carlone, Marco Hutter, Sebastian Scherer

专题命中 机器人学习 :robot learning(title,abstract);robotics(abstract);navigation(abstract);robotic(abstract)

AI总结 PyPose结合深度感知模型与物理优化,提供高效的机器人学习库,实现计算速度提升超过10倍,适用于SLAM、规划、控制等任务。

Comments Project Website: https://pypose.org Documentation: https://pypose.org/docs/ Tutorial: https://pypose.org/tutorials/ Source code: https://github.com/pypose/pypose

Journal ref IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023

详情
AI中文摘要

深度学习在机器人感知领域取得了显著成功,但其数据驱动的特性在面对不断变化的环境时表现不佳。相比之下,基于物理的优化具有更好的泛化能力,但在复杂任务中由于缺乏高层语义信息和依赖手动参数调优而表现较差。为利用这两个互补的世界,我们提出了PyPose:一个面向机器人、基于PyTorch的库,将深度感知模型与基于物理的优化相结合。PyPose的架构整洁且组织良好,具有命令式风格的接口,并且高效且易于使用,使其能够轻松集成到实际的机器人应用中。此外,它支持任何阶次的李群和李代数的并行计算以及二阶优化器,如信任区域方法。实验表明,PyPose在计算速度上比最先进的库快了超过10倍。为了促进未来研究,我们为机器人学习的几个领域提供了具体的示例,包括SLAM、规划、控制和惯性导航。

英文摘要

Deep learning has had remarkable success in robotic perception, but its data-centric nature suffers when it comes to generalizing to ever-changing environments. By contrast, physics-based optimization generalizes better, but it does not perform as well in complicated tasks due to the lack of high-level semantic information and reliance on manual parametric tuning. To take advantage of these two complementary worlds, we present PyPose: a robotics-oriented, PyTorch-based library that combines deep perceptual models with physics-based optimization. PyPose's architecture is tidy and well-organized, it has an imperative style interface and is efficient and user-friendly, making it easy to integrate into real-world robotic applications. Besides, it supports parallel computing of any order gradients of Lie groups and Lie algebras and $2^{\text{nd}}$-order optimizers, such as trust region methods. Experiments show that PyPose achieves more than $10\times$ speedup in computation compared to the state-of-the-art libraries. To boost future research, we provide concrete examples for several fields of robot learning, including SLAM, planning, control, and inertial navigation.

URL PDF HTML 收藏
2512.02020 2025-12-16 cs.RO cs.AI cs.CV cs.LG 87%

EfficientFlow: Efficient Equivariant Flow Policy Learning for Embodied AI

EfficientFlow: 为具身AI高效的等价流策略学习

Jianlei Chang, Ruofeng Mei, Wei Ke, Xiangyu Xu

机构 * Xi'an Jiaotong University(西安交通大学)

专题命中 机器人学习 :embodied AI(title,abstract);manipulation(abstract);robotic(abstract);分类 cs.RO、cs.AI、cs.CV

AI总结 EfficientFlow通过引入等价性与加速正则化策略,提升具身AI的效率与性能。

Comments Accepted by AAAI 2026. Project Page: https://efficientflow.github.io/

详情
AI中文摘要

生成建模最近在视觉-运动策略学习中展示了显著的潜力,能够实现跨多样具身AI任务的灵活和表达性控制。然而,现有的生成策略往往在数据效率方面存在问题,需要大规模的示范数据,并且在采样效率方面也存在问题,在推理过程中导致动作生成缓慢。我们引入EfficientFlow,一种统一的框架,用于高效的具身AI流基于策略学习。为了提高数据效率,我们将在流匹配中引入等价性。我们理论证明,当使用各向同性高斯先验和等价速度预测网络时,所得到的动作分布保持等价性,从而提高泛化能力和大幅减少数据需求。为了加速采样,我们提出了一种新的加速正则化策略。由于直接计算加速度对于边际流轨迹来说是不可行的,我们推导出一种新的替代损失,使得仅使用条件轨迹即可实现稳定和可扩展的训练。在广泛的机器人操作基准测试中,所提出的算法在数据有限的情况下实现了竞争或优越的性能,同时提供了显著更快的推理速度。这些结果突显了EfficientFlow作为高性能具身AI强大而高效范式的潜力。

英文摘要

Generative modeling has recently shown remarkable promise for visuomotor policy learning, enabling flexible and expressive control across diverse embodied AI tasks. However, existing generative policies often struggle with data inefficiency, requiring large-scale demonstrations, and sampling inefficiency, incurring slow action generation during inference. We introduce EfficientFlow, a unified framework for efficient embodied AI with flow-based policy learning. To enhance data efficiency, we bring equivariance into flow matching. We theoretically prove that when using an isotropic Gaussian prior and an equivariant velocity prediction network, the resulting action distribution remains equivariant, leading to improved generalization and substantially reduced data demands. To accelerate sampling, we propose a novel acceleration regularization strategy. As direct computation of acceleration is intractable for marginal flow trajectories, we derive a novel surrogate loss that enables stable and scalable training using only conditional trajectories. Across a wide range of robotic manipulation benchmarks, the proposed algorithm achieves competitive or superior performance under limited data while offering dramatically faster inference. These results highlight EfficientFlow as a powerful and efficient paradigm for high-performance embodied AI.

URL PDF HTML 收藏