arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Steer2Grasp:推理时具身感知引导用于多样化物理可行抓取扩散

Steer2Grasp: Inference-Time Embodiment-Aware Steering for Diverse Physically Feasible Grasp Diffusion

Vignesh Vembar, Ayush Kaura, A Padmaprabhan, Siddharth Sinha, Kailash Nagarajan, Keshab Patra, Md Faizal Karim, K Madhava Krishna

arXiv 2609.33546首次发表:更新:

发表机构

Robotics Research Center, IIIT Hyderabad; Johns Hopkins University(国际信息技术研究所海得拉巴机器人研究中心; 约翰斯·霍普金斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Steer2Grasp提出一种无需训练的推理时具身感知引导框架,通过Feynman-Kac粒子重加权与重采样,将冻结的扩散模型引导至可行抓取模式,显著提升单双臂在受限环境中的可行抓取生成。

AI 中文摘要

当前的抓取扩散模型为生成提供了丰富的先验,但其以物体为中心的方法可能违反由具身和环境施加的运动学和碰撞约束。现有的具身感知方法主要通过梯度引导或优化对生成的抓取进行局部修正,这使得难以从根本不可行的模式中恢复。我们提出了Steer2Grasp,一个无需训练、具身无关的推理时抓取引导框架,它利用部署特定的奖励来调整冻结的笛卡尔抓取扩散模型。通过基于Feynman-Kac(FK)启发的粒子重加权和重采样,该方法将群体质量从不可行模式重新分配到高奖励抓取模式,从而在不修改预训练扩散模型或不需要可微约束的情况下实现群体级别的模式转换。该框架通过可达性和碰撞感知奖励实现对单臂和双臂抓取的统一处理,随后进行无梯度的夹爪级局部细化。在多样化的物体、机器人具身和受限环境中,我们的方法大幅提高了可行抓取的生成,同时保持与底层抓取先验的接近性。

英文摘要

Current grasp diffusion models provide rich priors for generation, yet their object-centric approach can violate the kinematic and collision constraints imposed by the embodiment and the environment. Existing embodiment-aware methods primarily perform local corrections around generated grasps through gradient guidance or optimization, making it difficult to recover from fundamentally infeasible modes. We present Steer2Grasp, a training-free, embodiment-agnostic framework for inference-time grasp steering that adapts a frozen Cartesian grasp diffusion model using deployment-specific rewards. Through Feynman-Kac (FK) inspired particle reweighting and resampling, the method reallocates population mass from infeasible to high-reward grasp modes, enabling population-level mode transitions without modifying the pretrained diffusion model or requiring differentiable constraints. The framework enables a unified treatment for single and dual arm grasping through reachability and collision aware rewards, followed by gradient free gripper level local refinement. Across diverse objects, robot embodiments, and constrained environments, our method substantially improves feasible grasp generation while maintaining proximity to the underlying grasp prior.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑