arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.23224cs.ROcs.AIcs.CV

按需思考:用于视觉-语言-动作操纵中选择性慢路径干预的提示权限控制

Think Only When Needed: Prompt-Authority Control for Selective Slow-Path Intervention in Vision-Language-Action Manipulation

Zhiruo Zhou, Zelin Li, Xiwen Chen, Jiazhuo Li, Chenwei Wang, Huiming Chen, Xiaojun Zhu

首次发表
浏览论文内容

中文总结 AI 辅助

针对视觉-语言-动作操纵中检索文本导致的提示形式崩溃问题,提出TOWN-VLA提示权限接口,在LIBERO-Plus评估和PiPER机械臂实验中提升了任务成功率。

中文摘要 AI 辅助

检索能够在不重新训练的情况下有效增强冻结的视觉-语言-动作(VLA)策略,但检索到的文本一旦进入执行提示就会成为控制干预。在匹配审计中,原始附加文本将平均成功率从92.47%降至3.00%,而有意义且长度匹配的无意义附加文本在全部500个状态上均失败。该结果揭示了“提示形式崩溃”:改变指令形式而非添加有用语义会主导执行。我们引入TOWN-VLA(Think Only When Needed),一种将候选生成与更改策略输入的权限分离的提示权限接口。固定兼容性规则授权使用规范的简洁指令;否则,接口会完全恢复原始基础(Base)提示。在900条审计路线中,每条路线都遵循此约定:525条路线通过匹配哈希值恢复基础提示,所有375条授权提示均保留任务特征。在匹配的4×7 LIBERO-Plus评估中,每种方法有10030个回合,成功率从69.5%升至73.1%(增加362个回合;95%置信区间为1.89-5.45个百分点),在6个扰动轴和全部4个套件上均有提升。在配备冻结pizerofive检查点的物理PiPER机械臂上,每种方法各150次试验,成功率从52.7%升至78.7%(p=3.16×10⁻⁶)。提示权限可对冻结控制器强制执行;无需神示的准入校准是下一部署目标。

英文摘要

Retrieval can efficiently and effectively augment a frozen vision--language--action (VLA) policy without retraining, yet retrieved text becomes a control intervention once it enters the executed prompt. In a matched audit, raw appended text reduces mean success from 92.47\% to 3.00\%, while meaningful and length-matched meaningless appends both fail on all 500 states. This result identifies \emph{prompt-form collapse}: changing the instruction form, rather than adding useful semantics, can dominate execution. We introduce TOWN-VLA (Think Only When Needed), a prompt-authority interface that separates candidate generation from permission to alter the policy input. A fixed compatibility rule authorizes a canonical compact instruction; otherwise, the interface restores the original Base prompt exactly. Across 900 audited routes, every route follows this contract: 525 routes recover Base with matching hashes, and all 375 authorized prompts preserve the task signature. On a matched $4\times7$ LIBERO-Plus evaluation with 10{,}030 episodes per method, success rises from 69.5\% to 73.1\% ($+362$ episodes; 95\% CI 1.89--5.45 points), improving on six perturbation axes and all four suites. On a physical PiPER arm with a frozen \pizerofive{} checkpoint, success rises from 52.7\% to 78.7\% over 150 trials per method ($p=3.16\times10^{-6}$). Prompt authority is enforceable for a frozen controller; oracle-free admission calibration is the next deployment target.

↑