发表机构
University of Pennsylvania(宾夕法尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现实世界多物体杂乱环境处理难题,提出LENS方法,通过大语言模型引导自动生成特定场景和任务的自适应更新抽象,经实验验证其能改进多种模型在高度杂乱操纵场景中的表现。
AI 中文摘要
尽管通用机器人操纵最近取得了进展,但现实世界中的多物体杂乱对当前流行方法来说仍具挑战性。问题因物体增多、碰撞、接触物理不可预测、干扰因素和任务模糊性而愈发复杂。将其应用于实际部署需要有效的场景抽象,但目前生成此类抽象需要大量特定任务的人工工程,且难以扩展。我们提出一种即插即用的解决方案,在现有规划和控制堆栈之上自动生成特定场景、特定任务且自适应更新的抽象。大语言模型引导的环境简化(LENS)通过合并(如堆叠物体)或修剪(如远处物体)场景实体,响应任务进展,在闭环中生成去杂乱的抽象场景表示。这些动态、与任务相关的抽象通用且易用。实验表明,LENS在各种高度杂乱的操纵场景中改进了经典规划、基于模型的控制和视觉语言动作模型。
英文摘要
Despite recent advances in general-purpose robotic manipulation, real-world multi-object clutter remains challenging to handle for today's prevalent approaches. The problem scales in complexity due to more objects and collisions, more unpredictable contact physics, distractors, and task ambiguity. Bridging this gap to real-world deployment requires effective scene abstractions; yet today, producing such abstractions requires extensive task-specific manual engineering, which does not scale. These abstractions are costly to generate and difficult to adjust or fine-tune. We instead propose a plug-and-play fix to automatically generate scene-specific, task-specific, adaptively updating abstractions on top of existing planning and control stacks. LLM-guided Environment Simplification (LENS) produces a de-cluttered abstracted scene representation by merging (e.g., stacked objects) or pruning (e.g., distant objects) scene entities in a closed loop in response to task progress. These dynamic, task-relevant abstractions are versatile and easy to use. In our experiments, we show that LENS improves classical planning, model-based control, and a vision-language-action model, across a diverse set of highly cluttered manipulation scenes. Project website: https://lens-2026.github.io/.