arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.11175cs.ROcs.AI

高阶动作监督造就强策略类

Higher-Order Action Supervision Makes A Strong Policy Class

Peng Cheng, Yunxian Hou, Zhi Zhou, Qian Zhang, Chang Huang, Xianyuan Zhan

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对数据驱动决策方法在现实应用中的控制不稳定问题,提出同时监督零阶与一阶动作的损失方案,作为即插即用模块集成于离线RL框架,在OGBench、D4RL等数据集上显著提升策略性能、鲁棒性及少数据下的分布外泛化能力。

中文摘要 AI 辅助

现代数据驱动的决策方法,如模仿学习(IL)和强化学习(RL),在解决诸多复杂任务时已取得巨大成功。然而,这些方法应用于机器人、自动驾驶等现实场景时,常存在严重的控制不稳定性和鲁棒性问题,给实际部署带来显著挑战。我们认为,这种不稳定性主要源于它们仅监督和优化零阶动作(即动作标签),未考虑高阶动作动态性和时间一致性。本文表明,同时监督零阶和一阶动作可大幅提升策略的性能与控制鲁棒性。为实现这一点,我们提出一种新颖简洁的损失方案,该方案有正式理论保证支持,无需任何结构修改即可为任何现成策略模型(如确定性策略、随机策略或流策略)赋予高阶动作监督能力。此外,我们提出的方法可作为轻量即插即用模块,无缝集成到广泛的现有离线RL框架中。在OGBench和D4RL上的大量评估表明,我们的方法在多种连续控制环境中均带来显著的性能和鲁棒性提升。值得注意的是,我们的方法还能在极具挑战性的少数据场景中增强策略的分布外(OOD)泛化能力,使其成为解决诸多现实控制问题的理想工具。

英文摘要

Modern data-driven decision-making methods, such as imitation learning (IL) and reinforcement learning (RL), have achieved great success in solving many complex tasks. However, these methods often suffer from serious control instability and robustness issues when applied in real-world applications such as robotics and autonomous driving, posing notable challenges for their practical deployment. We argue that this instability issue stems largely from their limitations in solely supervising and optimizing zeroth-order actions (i.e., the action labels), failing to account for higher-order action dynamics and temporal consistency. In this paper, we show that simultaneously supervising both zeroth- and first-order actions can dramatically enhance policies' performance and control robustness. To achieve this, we introduce a novel and elegant loss scheme supported by formal theoretical guarantees that can equip any off-the-shelf policy model (e.g., deterministic, stochastic, or flow policies) with the capability for higher-order action supervision, without requiring any structural modifications. Moreover, our proposed method can serve as a lightweight plug-and-play module that seamlessly integrates with a broad spectrum of existing offline RL frameworks. Extensive evaluations on OGBench and D4RL demonstrate that our approach yields substantial performance and robustness improvements across a wide range of continuous control environments. Notably, our method can also enhance policies' out-of-distribution (OOD) generalization capability in the challenging low-data regime, making it an ideal tool in tackling many real-world control problems.

发表机构

  • Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院(AIR))
  • Zenfex AI Lab(Zenfex人工智能实验室)
  • University of Electronic Science and Technology of China(电子科技大学)
  • Horizon Robotics(地平线机器人)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑