AI 中文总结
该研究提出基于可供性识别与动作效果预测的操纵规划系统,通过多模态目标匹配模块评估候选计划,能跟踪遮挡物体位置。还用图像转换模块辅助,分别评估模块性能并展示集成系统在模拟与硬件上的规划能力。
AI 中文摘要
我们提出了一个基于可供性识别和动作效果预测的操纵规划系统。该系统以视觉形式对可能的未来进行推理,并使用多模态目标匹配模块,通过将预测结果与运行时设置的基于文本的目标进行匹配来评估候选计划。即使目标文本中命名的物体被遮挡,也能通过预测跟踪其位置,从而即使物体被遮挡或其初始描述符在未来状态中无法识别它们时,也能生成行动计划。我们还通过一个图像转换模块扩展了该系统,将具有不同形状和视觉外观物体的真实世界状态图像转换为一致的视觉外观,以促进物理机器人设置中的操纵规划。我们分别评估了系统模块的性能,并在模拟和硬件上的一组具有挑战性的任务中展示了集成系统的操纵规划能力。
英文摘要
We present a manipulation planning system based on affordance recognition and action effect prediction. The system reasons through possible futures in visual form, and evaluates candidate plans by agreement of predicted outcomes with text-based goals set at run-time, using a multi-modal goal-matching module. Positions of objects named in the goal text are tracked through predictions even when occluded, making it possible to generate action plans even when objects become occluded, or when their initial descriptors cease to identify them in future states. We further expand the system with an image conversion module for translating real-world state images with objects of varied shapes and visual appearances into a consistent visual appearance, to facilitate manipulation planning in a physical robot setup. We evaluate performance of the system's modules in isolation and demonstrate the integrated system's manipulation planning capabilities on a set of challenging tasks in both simulation and on hardware.
Comments14 pages, 10 figures