AI 中文总结
ViTaR将触觉反馈转为有界残差修正的执行调制器,在冻结VLA上实现自适应,在UniVTAC基准上成功率提升30.6个百分点,适配真实机器人场景。
AI 中文摘要
随着视觉-语言-动作(VLA)模型向实际部署扩展,接触丰富的操作任务暴露出一个关键盲区:这些策略编码了广泛的视觉语义先验,但对局部接触事件毫无感知,在接触建立、丢失或不稳定时会产生相同的动作。现有解决方案要么修改VLA内部结构,存在灾难性遗忘的风险;要么需要在接近失败的接触条件下进行在线强化学习,这两种方案都赋予触觉对动作生成的无限制影响力,与VLA的可泛化先验相冲突。我们提出ViTaR,将触觉反馈从动作生成的感知输入重新定义为执行调制器,在冻结的VLA之上选择并缩放有界残差修正,从结构上保留预训练能力。ViTaR将自适应过程分解为两个阶段:效果引导建模通过基于结果的偏好证据确定局部是否需要修正以及哪种修正合理;残差动作调制将该证据转化为残差选择,利用实时视觉-触觉观测提供连续缩放的增益。在涵盖7项接触丰富任务的UniVTAC基准上,ViTaR的平均成功率达61.3%,相比其冻结的VLA基准提升了30.6个百分点,同时优于专用的触觉基线。物理机器人实验证实,有界触觉调制可适配真实传感器噪声和动力学。
英文摘要
As Vision-Language-Action (VLA) models scale toward real-world deployment, contact-rich manipulation exposes a critical blind spot: these policies encode broad visual-semantic priors yet remain unaware of local contact events, producing identical actions whether contact is established, lost, or destabilized. Existing remedies either modify VLA internals, risking catastrophic forgetting, or demand online reinforcement under near-failure contact conditions. Both grant tactile unbounded influence over action generation, conflicting with the priors that make VLAs generalizable. We introduce ViTaR, which reframes tactile feedback from an action-generating perceptual input to an execution modulator that selects and scales bounded residual corrections atop a frozen VLA, preserving pretrained capabilities by construction. ViTaR decomposes adaptation into two stages: Effect-Guided Modeling determines whether and which correction is locally justified via outcome-grounded preference evidence, and Residual Action Modulation converts this evidence into a residual choice with continuously scaled gain from real-time visuotactile observations. On the UniVTAC benchmark spanning seven contact-rich tasks, ViTaR achieves 61.3% average success, a 30.6 percentage-point improvement over its frozen VLA base that also surpasses purpose-built tactile baselines. Physical-robot experiments confirm that bounded tactile modulation transfers to real sensor noise and dynamics.