发表机构
State Key Laboratory of Robotics and Systems, Harbin Institute of Technology(机器人技术与系统国家重点实验室,哈尔滨工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究机器人操作流策略中动作选择问题,提出HCPG-Flow方法,通过分层、以对象为中心的接触进展引导增强SAC-Flow,在模拟和物理任务中提高成功率,减少完成时间。
AI 中文摘要
流策略可表示机器人操作的多模态动作分布,但机器人在每个控制步骤必须执行一个动作。当采样多个提议时,基于评论家的排序使数据收集依赖于对重放中可能表示较弱的候选动作的值估计。我们引入了HCPG-Flow,这是一种分析性的展开时间选择器,它在保留其演员和评论家目标的同时,用分层的、以对象为中心的接触进展引导增强了SAC-Flow。HCPG在接触后从末端执行器方法切换到任务进展,通过与任务相关距离的一阶减少对每个提议进行评分,在候选集中标准化分数,并执行温度控制的动作嵌入。在十个模拟任务中,HCPG在两个基准上均提高了SAC-Flow的平均成功率,包括在Maniskill上提高了9.5个百分点。四个物理任务进一步显示出高成功率,成功完成时间减少了17.4%。
英文摘要
Flow policies can represent multimodal action distributions for robot manipulation, yet a robot must execute one action at each control step. When several proposals are sampled, critic-based ranking makes data collection depend on value estimates over candidate actions that may be weakly represented in replay. We introduce HCPG-Flow, an analytic rollout-time selector that augments SAC-Flow with hierarchical, object-centric contact-progress guidance while preserving its actor and critic objectives. HCPG switches from end-effector approach to task progress after contact, scores each proposal by the first-order reduction of a task-relevant distance, standardizes scores within the candidate set, and executes a temperature-controlled action embedding. Across ten simulated tasks, HCPG improves mean success over SAC-Flow on both benchmarks, including a 9.5 percentage-point gain on Maniskill. Four physical tasks further show high success with a 17.4% reduction in successful completion steps.Project page: https://hitxraz.github.io/HCPG-Flow/