arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AgenticTactileVLA:无需VLA重训练的接触引导执行时监督实现通用灵巧操作

AgenticTactileVLA: Contact-Guided Execution-Time Supervision for Generalizable Dexterous Manipulation without VLA Retraining

Elizaveta Semenyakina, Ivan Snegirev, Mikhail Kiselev, Miguel Altamirano Cabrera, Artem Lykov, Hajira Amjad, Dzmitry Tsetserukou

arXiv 2610.04391首次发表:更新:

发表机构

Intelligent Space Robotics Lab, Skolkovo Institute of Science and Technology; R&D Center, MWS(斯科尔科沃科学技术学院智能空间机器人实验室; MWS研发中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出AgenticTactileVLA执行时监督器,利用手指位置和力矩反馈在物理交互中调整固定VLA策略,无需重训练,在留出物体上将操作完成率从61.3%提升至84.0%。

AI 中文摘要

视觉-语言-动作策略可能预测出可迁移的操作策略,但在遇到的物体上却无法可靠地实现该策略:与相同抓取兼容的物体在几何形状和柔顺性上存在差异,并且视觉反馈在闭合遮挡下会退化。本文提出AgenticTactileVLA作为一种执行时监督器,将部分物体特定适应性从预测转移到物理交互中。固定的VLA提供接近方式和手部目标;监督器决定是保持透明、调整手指弯曲、保留或释放修正后的配置、将控制权交还给VLA进行重试,还是选择柔顺的手部控制模式。它利用手指位置和电机力矩反馈作为本体感觉接触证据,既不需要触觉传感器,也不需要VLA重训练。在配备BrainCo Revo2手部的Unitree G1机器人上,对从VLA训练中留出的五个物体进行随机匹配块评估,基础VLA的完成率为61.3%,无条件接近失速控制的完成率为72.0%,而在共享预算下监督器的完成率为84.0%;增益在每个物体上均为正,并在中等姿态扰动下持续存在。消融实验表明,该增益不能仅由延长闭合时间解释,选择性触发将修正次数减少了65.7%,且未检测到完成率变化。保留审计显示,在88.9%的留出案例中,接受预测了保留,而柔顺物体则暴露出保守的误拒绝。一项薄壁杯实验展示了上下文路由到柔顺控制,与始终柔顺的参考方法相匹配。这些结果表明,接触引导的执行时适应可以通过调整物理实现,在不进行物体特定重训练的情况下,提高固定VLA对留出物体的物体级泛化能力。

英文摘要

Vision-language-action policies may predict a transferable manipulation strategy yet fail to realize it reliably on the encountered object: objects compatible with the same grasp differ in geometry and compliance, and visual feedback degrades under closure occlusion. AgenticTactileVLA is presented as an execution-time supervisor that shifts part of object-specific adaptation from prediction to physical interaction. A fixed VLA provides the approach and hand targets; the supervisor decides whether to remain transparent, refine finger flexion, retain or release the corrected configuration, return control to the VLA for retry, or select a compliant hand-control regime. It uses finger-position and motor-effort feedback as proprioceptive contact evidence and requires neither tactile sensors nor VLA retraining. On a Unitree G1 with a BrainCo Revo2 hand, a randomized matched-block evaluation on five objects held out from VLA training yields 61.3% completion for the base VLA, 72.0% for unconditional close-to-stall control, and 84.0% for the supervisor under a shared budget; the gain is positive on every object and persists under moderate pose perturbations. Ablations show the gain is not explained by extended closure alone, and that selective triggering reduces correction episodes by 65.7% with no detected change in completion. A retention audit shows acceptance predicts retention in 88.9% of held-out cases, while compliant objects expose conservative false rejection. A thin-walled-cup study demonstrates contextual routing to compliant control, matching an always-compliant reference. These results suggest that contact-guided execution-time adaptation can improve the object-level generalization of a fixed VLA to held-out objects by adapting physical realization without object-specific retraining.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑