发表机构
Beijing Academy of Agriculture and Forestry Sciences; Huazhong Agricultural University(北京市农林科学院; 华中农业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对机器人草莓采摘中复杂接触动力学问题,开发含启发式阶段协调的强化学习框架,通过共享策略、观察空间及分层架构实现任务,采用相关技术支持模拟到现实转移,实验验证了该方法在不同平台的有效性。
AI 中文摘要
严重遮挡和可变形植物结构带来复杂接触动力学,给机器人草莓采摘带来挑战。开发了具有启发式阶段协调的策略驱动强化学习框架,将障碍物分离、果实采摘和放置制定为顺序决策任务。共享交互感知策略生成笛卡尔运动,轻量级启发式逻辑协调任务进展和夹爪事件。使用共享结构化观察空间表示目标、障碍物、末端执行器和任务上下文信息。分层架构将高级策略与低级笛卡尔阻抗控制相结合以实现柔顺交互。采用可行性优先观察对齐和域随机化以支持零样本模拟到现实的转移。该策略在模拟中成功率达89.7%,在实际实验中达82.0%。随着遮挡水平从1增加到5,平均执行时间从12.99秒增加到21.73秒,反映出更大的交互复杂性。这些结果证明了交互感知收获行为有效转移到结构不同的机器人平台。
英文摘要
Severe occlusions and deformable plant structures introduce complex contact dynamics that challenge robotic strawberry harvesting. A policy-driven reinforcement learning (RL) framework with heuristic phase coordination was developed, in which obstacle separation, fruit detachment, and placement were formulated as a sequential decision-making task. A shared interaction-aware policy generated Cartesian motions across all task phases, while lightweight heuristic logic coordinated task progression and gripper events. A shared structured observation space was used to represent target, obstacle, end-effector, and task-context information. A hierarchical architecture combined the high-level policy with low-level Cartesian impedance control for compliant interaction. To support zero-shot sim-to-real transfer, feasibility-first observation alignment and domain randomization were adopted. The policy achieved success rates of 89.7% in simulation and 82.0% in real-world experiments. As the occlusion level increased from 1 to 5, the average execution time increased from 12.99 s to 21.73 s, reflecting greater interaction complexity. These results demonstrated effective transfer of interaction-aware harvesting behaviors to a structurally different robotic platform.
CommentsAccepted to IROS 2026